1 Commits
Author SHA1 Message Date
sneak 3d78aacac2 Carry file size, thumbnail size, and deletion flag through decryptFile (closes #37)
check / check (push) Successful in 25s
Foundation unit for the cache/API design. Three fields arrived on the wire
but decryptFile dropped them:

- `file.size` from `info.fileSize` and `thumbnail.size` from `info.thumbSize`,
  left `undefined` when the server omits `info`.
- `isDeleted` carried from the diff row onto `EnteFile`.

No caller change: `listFiles` still filters deleted rows before decrypt, so no
deleted row reaches decryptFile here. Surfacing a deleted file through
decryption belongs to the enumeration unit, issue 38, where the return shape
can be designed around where such a row actually flows.

Model: opus-4-8
2026-09-22 09:43:28 +00:00
41 changed files with 1032 additions and 11067 deletions
+32 -215
View File
@@ -38,47 +38,34 @@ yarn quak get 67890 --out ./photo.jpg
yarn quak backup ./my-backup yarn quak backup ./my-backup
``` ```
For library use, the primary surface is the cache-backed `Library`: For library use:
```ts ```ts
import { Client, Library } from "quak"; import { Client } from "quak";
// Log in once; the client satisfies the library's client interface.
const client = await Client.login({ const client = await Client.login({
email: "you@example.com", email: "you@example.com",
password: "your-password", password: "your-password",
}); });
// Open a cache-backed library. On an empty cache this awaits one server for (const c of await client.listCollections()) {
// refresh; on an existing cache it returns immediately and refreshes in the console.log(c.id, c.name);
// background every `refreshIntervalSeconds` (default 3). const files = await client.listFiles(c.id, c.key);
const lib = await Library.open({ client }); for (const f of files) {
console.log(` ${f.metadata.title} [${f.metadata.fileType}]`);
// Default reads answer synchronously from the local cache — no network.
for (const album of lib.albums.list()) {
console.log(album.collectionID, album.name);
for (const photo of album.photos.list()) {
console.log(` ${photo.title} [${photo.fileType}]`);
} }
} }
// Fresh reads await a server round-trip and answer with current state. // Download a file
const { albums } = await lib.fresh(); const files = await client.listFiles(collectionID, collectionKey);
console.log(`${albums.list().length} albums as of now`); await client.downloadFile(files[0], "./photo.jpg");
// Fetch (and cache) one photo's full-resolution bytes. // Serialize session for later (consumer handles persistence)
const photo = lib.photos.byID({ fileID: 12345 }); const snapshot = client.toJSON();
if (photo) { // ... later:
const { path } = await photo.original(); const restored = Client.fromJSON(snapshot);
console.log(`original at ${path}`);
}
lib.close();
``` ```
The lower-level `Client` (login, session serialization, and the raw
enumeration/download calls) is exported too and documented under Design below.
## Entrypoints ## Entrypoints
This repository adheres to the This repository adheres to the
@@ -433,7 +420,6 @@ you would treat the password itself.
### CLI surface ### CLI surface
``` ```
quak [--cache-dir <path>] <command> global: local metadata/content cache location
quak login interactive or QUAK_EMAIL/QUAK_PASSWORD quak login interactive or QUAK_EMAIL/QUAK_PASSWORD
quak whoami print logged-in account as JSON quak whoami print logged-in account as JSON
quak logout delete saved session quak logout delete saved session
@@ -442,28 +428,13 @@ quak files --collection <id> [--json] list files in a collection
quak get <fileID> [--out path] [--collection] download and decrypt a file quak get <fileID> [--out path] [--collection] download and decrypt a file
quak get-thumb <fileID> [--out] [--collection] download and decrypt a thumbnail quak get-thumb <fileID> [--out] [--collection] download and decrypt a thumbnail
quak backup <dir> [--json] full incremental backup quak backup <dir> [--json] full incremental backup
quak backup-metadata <dir> [--exif] dump all decrypted metadata as JSON
quak helper list-missing-thumbnails [--json] find files with missing thumbnails quak helper list-missing-thumbnails [--json] find files with missing thumbnails
quak helper fix-missing-thumbnails [--file ids] generate + upload missing thumbnails quak helper fix-missing-thumbnails [--file ids] generate + upload missing thumbnails
``` ```
Every command runs on the same cache-backed library. The read commands — `get` and `get-thumb` search all collections for the file ID when `--collection`
`collections`, `files`, `get`, and `get-thumb` — force a fresh server round-trip is not specified. All listing and backup commands support `--json` for
before they answer, so they report current account state rather than whatever machine-readable output.
the cache last held. `--cache-dir` overrides where the cache lives; without it
each account gets its own directory under the per-user cache path.
`get` and `get-thumb` resolve the file by ID directly, so `--collection` is
accepted for backward compatibility but ignored. `backup-metadata --exif` (alias
`--all`) additionally downloads each file to extract full EXIF/IPTC/XMP
metadata. The listing and backup commands support `--json` for machine-readable
output.
`helper fix-missing-thumbnails` regenerates thumbnails for baseline JPEG images
only, because the bundled decoder (`jpeg-js`) decodes only JPEG. A non-JPEG
image (PNG, HEIC) or a video is reported as `skipped` (unsupported format), kept
distinct from a `failed` repair, and does not affect the exit code; a genuine
failure still exits non-zero.
### Backup layout ### Backup layout
@@ -489,7 +460,7 @@ code is non-zero if any files failed.
- [x] Retry policy: no retry on 4xx, exponential backoff on 5xx and network - [x] Retry policy: no retry on 4xx, exponential backoff on 5xx and network
errors errors
- [x] Update the API reference section below to match the current implementation - [ ] Update the API reference section below to match the current implementation
- [x] `make docker` green - [x] `make docker` green
- [ ] Tag `v1.0.0` - [ ] Tag `v1.0.0`
@@ -503,179 +474,25 @@ Future (desktop client, separate repo):
## API reference ## API reference
The library's primary surface is the cache-backed `Library`; the lower-level The API reference section below is from an earlier draft and does not fully
`Client` sits underneath it and is covered by the Design sections above. The reflect the current implementation. The authoritative API documentation is in
test suite is the canonical, executable documentation — `test/library/` and the test files, particularly `test/client/usage.test.ts` which is a literate
`test/client/usage.test.ts` walk every operation, and `yarn test` verifies them. tutorial walking through every operation. Run `yarn test` to verify the examples
are correct.
### Opening a library The key types and their actual signatures can be found in:
`Library.open(options)` loads the on-disk cache, starts the background refresh
loop, and resolves to a `Library`. On an empty cache it awaits the first refresh
so it never opens onto empty data; on an existing cache it returns immediately
and refreshes in the background, so an unreachable server does not block
opening.
`LibraryOptions`:
| Option | Default | Meaning |
| ------------------------ | --------------------------- | --------------------------------------------------------------------- |
| `client` | required | the account client (a `Client`, or any `LibraryClient`) |
| `cacheDirectory` | `<XDG cache>/quak/<userID>` | where `metadata.json` and the content cache live |
| `downloadDirectory` | none | backup destination; an original already stored there counts as cached |
| `refreshIntervalSeconds` | `3` | background refresh cadence |
| `precacheThumbnails` | `true` | prefetch every thumbnail, newest first |
| `precacheOriginals` | `true` | prefetch the favorites album and the latest-window originals |
| `precacheOriginalsDays` | `7` | length in days of that latest window |
| `cacheOriginalsMaxBytes` | 100 GiB | hard ceiling on the originals cache |
| `freeBelowBytes` | 50 GiB | free space to protect on the volume; the effective limit adapts down |
| `isOriginalPinned` | none | extra predicate for originals that must never be evicted |
| `pools` | fresh `RequestPools` | the bounded request pools (sets concurrency) |
| `onProgress` | none | refresh/ML/precache progress callback (`RefreshEvent`) |
| `contentSource` | the client's own | override the byte source (mainly for tests) |
Concurrency is set through `pools`: construct
`new RequestPools({ metadataConcurrency, contentConcurrency, thumbnailConcurrency })`
and pass it. The three pools default to 10 / 5 / 25 (see Request pools below).
`lib.status()` returns a `LibraryStatus` (collection/file counts, last
refresh/ML times and errors, originals usage and effective limit, precache
progress, and `closed`). `lib.close()` stops the background timer; it is
idempotent, and an in-flight refresh is left to finish.
### Default reads vs. fresh reads
Default reads — `lib.albums`, `lib.photos`, `lib.timeline` — answer
synchronously from the last refreshed copy held in RAM and never touch the
network. The background timer refreshes that copy every
`refreshIntervalSeconds`, so a default read is immediate but may be up to one
interval stale.
`await lib.fresh()` forces a refresh, waits for it to complete and persist, and
returns the same `{ albums, photos, timeline }` namespaces — now guaranteed to
reflect a completed server round-trip. Concurrent `fresh()` calls coalesce onto
one refresh, and a refresh that fails rejects the caller (default reads stay
silent and keep serving the last good copy). The CLI's read commands use fresh
reads (issue https://git.eeqj.de/sneak/quak/issues/75).
### Read surface
- `lib.albums.list()``Album[]`, newest-updated first.
`lib.albums.byID({ collectionID })` and `byName({ albumName })`
`Album | undefined`.
- `lib.photos.byID({ fileID })``Photo | undefined`.
`lib.photos.records({ fileIDs })``PhotoRecord[]` in the requested order,
each id once, unknown ids dropped.
- `lib.timeline.groups({ groupBy, filter? })``TimelineGroup[]`, grouped by
`"day" | "week" | "month"` (keys `YYYY-MM-DD`, ISO `YYYY-Www`, `YYYY-MM`),
newest group first. A `PhotoFilter` combines `albumID`, `text`
(title/caption/album-name substring), `fileTypes`, `hasLocation`, and
`includeArchived`; hidden photos are always excluded.
An `Album` exposes its record fields and `album.photos.list()``Photo[]`
(newest first). A `Photo` exposes its record fields, `photo.record()`
`PhotoRecord`, and two content methods:
- `await photo.original(opts?)``{ path, bytes }` — the full-resolution file.
- `await photo.thumbnail(opts?)``{ path, bytes }`.
Both serve from the on-disk content cache when the bytes are present and
otherwise fetch through the pools; `opts.onProgress` reports per-file progress.
They throw when the library was opened without a content source.
Lower-level accessors that return decrypted model objects (which hold key
material) are also available: `listCollections()`, `getCollection(id)`,
`listFiles(collectionID)`, `getFile(collectionID, fileID)`, and
`getFileByID(fileID)`.
### Records and change notifications
The GUI-facing records hold no key material and no binary, so they survive
`structuredClone`/JSON across the Electron IPC boundary:
- `PhotoRecord`: `fileID`, `albumIDs`, `title`, `takenAt` (milliseconds),
`fileType`, optional `caption` / `width` / `height` / `latitude` /
`longitude`, `isArchived`, `isHidden`, and `thumbnailPath` / `originalPath`
once the bytes are cached.
- `AlbumRecord`: `collectionID`, `name`, `type`, `isShared`, `updationTime`, and
`fileIDs` (newest first).
- `LibrarySnapshot`: `{ albums, photos, takenAt }`.
`lib.snapshot()` returns a `LibrarySnapshot` (albums newest-updated first,
photos newest first). `lib.subscribe({ onChange })` delivers a `LibraryChange`
(`albumsChanged`, `photosChanged`, `fileIDsRemoved`, `albumIDsRemoved`,
`refreshedAt`) whenever a refresh alters the projection, and returns
`{ unsubscribe }`; a refresh that changes nothing delivers nothing.
### Thumbnails, ML search, and backup
- `lib.thumbnails.ensure({ fileIDs, priority, signal?, onProgress? })`
prefetches thumbnails through the thumbnail pool, deduped by fileID, returning
one `EnsureResult` (`{ fileID, path?, error? }`) per file. `priority` is
`"visible" | "ahead" | "background"`; only `"visible"` preempts background
work.
- `lib.mldata` searches the CLIP index built from Ente's per-file ML data:
`forFile({ fileID })``Promise<MLData | undefined>` (the whole stored
payload — face boxes, landmarks, embedding — read from disk on demand);
`similar({ fileID, limit? })` and `searchByEmbedding({ embedding, limit? })`
`SimilarResult[]` (`{ fileID, score }`, cosine similarity, most similar first,
default limit 20). quak bundles no text encoder, so `searchByEmbedding` takes
a query vector the caller produced elsewhere.
- `await lib.backup(opts?)``BackupResult`. It refreshes, fetches every
in-scope original (and, with `includeThumbnails`, thumbnails) through the
content cache, and rebuilds the on-disk backup tree with a durable failure
ledger. `BackupOptions`: `downloadDirectory` (falls back to the one `open()`
was given), `includeOriginals` (default `true`), `includeThumbnails` (default
`false`), `onlyAlbumNames`, and `onProgress`. See Backup layout above for the
tree it writes.
### Request pools
`RequestPools` holds three independent bounded pools — metadata (10), content
(5), thumbnails (25) — because Ente meters these traffic classes differently.
Each pool orders on-demand work ahead of background/precache work and dedups
in-flight fetches by key, and an idle pool never lends its slots to a busy one.
### On-disk cache layout
Under `cacheDirectory`:
```
<cacheDirectory>/
metadata.json decrypted account state + refresh cursor
originals/<fileID>.<ext> cached full-resolution files
thumbnails/<fileID>.jpg cached thumbnails
mldata/
<fileID>.json one decrypted ML payload per file
clip.f32, clip.json the packed CLIP index and its id list
fetched.json per-file fetch bookkeeping
```
A stored file appears only via an atomic temp-then-rename, so its presence means
it is complete. The design also calls for a content-hash comparison against
`FileMetadata.hash` on each fetched original; that check is deferred (issue
https://git.eeqj.de/sneak/quak/issues/68) because the exact hash construction
cannot yet be confirmed against the repo's fixtures.
### Key types by source file
- `src/library/index.ts`: `Library`, `LibraryOptions`, `LibraryStatus`,
`LibraryClient`, `RefreshEvent`
- `src/library/read.ts`: `Album`, `Photo`, `AlbumsAPI`, `PhotosAPI`,
`TimelineAPI`, `PhotoFilter`, `TimelineGroup`, `GroupBy`
- `src/library/content.ts`: `ContentResult`, `ContentOptions`, `ThumbnailsAPI`,
`EnsureOptions`, `EnsureResult`, `ContentSource`
- `src/library/records.ts`: `PhotoRecord`, `AlbumRecord`, `LibrarySnapshot`,
`LibraryChange`
- `src/library/mlsearch.ts`: `MLDataAPI`, `SimilarResult`
- `src/library/pools.ts`: `RequestPools`, `RequestPoolsOptions`, `BoundedPool`
- `src/backup.ts`: `BackupOptions`, `BackupResult`, `BackupError`
- `src/client.ts`: `Client`, `LoginOptions`, `ClientSnapshot` - `src/client.ts`: `Client`, `LoginOptions`, `ClientSnapshot`
- `src/api/client.ts`: `ApiClient`, `ApiClientOptions`, `StreamOptions` - `src/api/client.ts`: `ApiClient`, `ApiClientOptions`, `ApiError`,
`StreamOptions`
- `src/errors.ts`: `ApiError`, `TruncatedStreamError` - `src/errors.ts`: `ApiError`, `TruncatedStreamError`
- `src/retry.ts`: `withRetry`, `isRetryable`, `isSafeToReplay`, `RetryOptions` - `src/retry.ts`: `withRetry`, `isRetryable`, `isSafeToReplay`, `RetryOptions`
- `src/model/types.ts`: `Collection`, `EnteFile`, `FileMetadata`, `FileType`, - `src/auth/types.ts`: `KeyAttributes`, `SRPAttributes`,
`CollectionType`, `RawCollection`, `RawEnteFile` `AuthorizationResponse`, `LoginChallenge`
- `src/model/types.ts`: `Collection`, `EnteFile`, `FileMetadata`, `FileBlob`,
`RawCollection`, `RawEnteFile`, `RawMagicMetadata`
- `src/download/index.ts`: `DownloadResult`
- `src/backup.ts`: `BackupResult`, `BackupError`
- `src/thumbnails.ts`: `MissingThumbnailInfo`, `ThumbnailFixResult` - `src/thumbnails.ts`: `MissingThumbnailInfo`, `ThumbnailFixResult`
## Source attribution ## Source attribution
+2 -19
View File
@@ -14,28 +14,10 @@ pre-1.0
# Next Step # Next Step
Tag v1.0.0. Update the README API reference section to match the current implementation.
# Completed Steps # Completed Steps
- 2026-09-22: Rewrote the README API reference (and the Getting Started / usage
snippets) to match the shipped cache/API library on `next` (issue 53, issue
13). Documented `Library.open` and its options, the default-read vs `fresh()`
distinction (and that the CLI's read commands are fresh), the record types and
`snapshot()`/`subscribe()`, the `albums`/`photos`/`timeline` read surface,
`Photo` content methods, `thumbnails.ensure`, the `mldata` search surface,
`backup()`, the three request pools (10/5/25), and the on-disk cache layout;
noted the deferred content-hash integrity check (issue 68). Docs-only; no code
changed.
- 2026-09-22: Added resumable, deletion-aware enumeration to `Client` (issue 38,
closes issue 7). `collectionsSince`/`filesSince` take a starting cursor,
decrypt live records, surface tombstoned ids in a separate `deleted` list (a
tombstone has nothing to decrypt, so it is a bare id, not a hollow record),
and return the max `updationTime` seen as the cursor to resume from.
`filesSince` refuses to loop when the diff reports `hasMore` without advancing
the cursor (issue 7). `listCollections`/`listFiles` are now thin wrappers that
enumerate from `sinceTime: 0` and drop deletions, so existing callers are
unaffected.
- 2026-09-22: Carried file size, thumbnail size, and the deletion flag through - 2026-09-22: Carried file size, thumbnail size, and the deletion flag through
`decryptFile` (issue 37, foundation for the cache/API design). Live files now `decryptFile` (issue 37, foundation for the cache/API design). Live files now
populate `file.size`/`thumbnail.size` from the server's `info` (left populate `file.size`/`thumbnail.size` from the server's `info` (left
@@ -131,6 +113,7 @@ Tag v1.0.0.
# Future Steps # Future Steps
- Tag v1.0.0.
- Future desktop client, separate repo: - Future desktop client, separate repo:
- Electron app skeleton consuming this library. - Electron app skeleton consuming this library.
- Local SQLite cache keyed on (collectionID, fileID, updationTime). - Local SQLite cache keyed on (collectionID, fileID, updationTime).
+124 -190
View File
@@ -2,26 +2,13 @@
import { input, password as passwordPrompt } from "@inquirer/prompts"; import { input, password as passwordPrompt } from "@inquirer/prompts";
import { stdout, stderr } from "node:process"; import { stdout, stderr } from "node:process";
import { import { existsSync, mkdirSync, readFileSync, writeFileSync } from "node:fs";
copyFileSync,
existsSync,
mkdirSync,
readFileSync,
writeFileSync,
} from "node:fs";
import { join } from "node:path"; import { join } from "node:path";
import { Command } from "commander"; import { Command } from "commander";
import envPaths from "env-paths"; import envPaths from "env-paths";
import { Client, type ClientSnapshot } from "../src/client.js"; import { Client, type ClientSnapshot } from "../src/client.js";
import { init } from "../src/crypto/index.js"; import { init } from "../src/crypto/index.js";
import { Library, type LibraryClient } from "../src/library/index.js"; import { runBackup } from "../src/backup.js";
import {
fileListRow,
fileListLine,
originalName,
thumbnailName,
} from "../src/cli-output.js";
import { freshCollections, freshFiles, freshFile } from "../src/cli-read.js";
import { runMetadataBackup } from "../src/metadata-backup.js"; import { runMetadataBackup } from "../src/metadata-backup.js";
import { import {
listMissingThumbnails, listMissingThumbnails,
@@ -68,61 +55,7 @@ const program = new Command();
program program
.name("quak") .name("quak")
.description("CLI for the Ente end-to-end encrypted photo service") .description("CLI for the Ente end-to-end encrypted photo service")
.version("0.0.0") .version("0.0.0");
.option(
"--cache-dir <path>",
"Directory for the local metadata/content cache " +
"(default: the per-user cache directory)",
);
// The `--cache-dir` global, or undefined to let the library pick its per-user
// default keyed by the account id.
const cacheDirOption = (): string | undefined =>
program.opts<{ cacheDir?: string }>().cacheDir;
// A library client that omits `fetchMLData`, so the point commands below do not
// kick the library's background ML backfill: they read metadata, or fetch one
// file's content, and exit. `backup` and `backup-metadata` handle ML on their
// own terms. The content source is kept so `get`/`get-thumb`/`--exif` can fetch
// originals through the on-disk cache.
const readLibraryClient = (client: Client): LibraryClient => ({
whoami: () => client.whoami(),
collectionsSince: (args) => client.collectionsSince(args),
filesSince: (args) => client.filesSince(args),
contentSource: () => client.contentSource(),
});
// Open a library for a single point command: the aggressive background precache
// (issue #48) is off — a one-shot `collections` or `get` must not start
// downloading the whole account — and the refresh interval is long so no second
// refresh fires mid-command.
const openReadLibrary = (client: Client): Promise<Library> =>
Library.open({
client: readLibraryClient(client),
cacheDirectory: cacheDirOption(),
refreshIntervalSeconds: 3600,
precacheThumbnails: false,
precacheOriginals: false,
});
// Close the library and exit once stdout/stderr have drained. `process.exit`
// alone can truncate buffered piped output, and the library keeps the event
// loop alive with a background refresh, so a plain return could hang; this does
// neither.
const finish = (lib: Library | undefined, code: number): void => {
lib?.close();
const pending = [stdout, stderr].filter((s) => s.writableLength > 0);
if (pending.length === 0) {
process.exit(code);
return;
}
let remaining = pending.length;
for (const s of pending) {
s.once("drain", () => {
if (--remaining === 0) process.exit(code);
});
}
};
program program
.command("login") .command("login")
@@ -183,11 +116,7 @@ program
.action(async (opts: { json?: boolean }) => { .action(async (opts: { json?: boolean }) => {
await init(); await init();
const client = requireSession(); const client = requireSession();
const lib = await openReadLibrary(client); const collections = await client.listCollections();
// Force a server round-trip and list in enumeration order (issue #36
// amendment, issue #52): the pre-library CLI printed current state in
// this order, not the albums projection's newest-first order.
const collections = await freshCollections(lib);
if (opts.json) { if (opts.json) {
stdout.write( stdout.write(
@@ -211,7 +140,6 @@ program
); );
} }
} }
finish(lib, 0);
}); });
program program
@@ -231,29 +159,36 @@ program
process.exit(1); process.exit(1);
} }
const lib = await openReadLibrary(client); const collections = await client.listCollections();
// Force a server round-trip and list in enumeration order (issue #36 const col = collections.find((c) => c.id === collectionID);
// amendment, issue #52). Each file prints from its own decrypted if (!col) {
// metadata (raw title, microsecond creationTime) via cli-output, and in
// the pre-library CLI's enumeration order, not the projection's
// newest-first order.
const files = await freshFiles(lib, collectionID);
if (!files) {
stderr.write(`Collection ${collectionID} not found\n`); stderr.write(`Collection ${collectionID} not found\n`);
finish(lib, 1); process.exit(1);
return;
} }
const files = await client.listFiles(col.id, col.key);
if (opts.json) { if (opts.json) {
stdout.write( stdout.write(
JSON.stringify(files.map(fileListRow), null, 2) + "\n", JSON.stringify(
files.map((f) => ({
id: f.id,
title: f.metadata.title,
fileType: f.metadata.fileType,
creationTime: f.metadata.creationTime,
collectionID: f.collectionID,
})),
null,
2,
) + "\n",
); );
} else { } else {
for (const file of files) { for (const f of files) {
stdout.write(fileListLine(file) + "\n"); stdout.write(
`${f.id}\t${f.metadata.fileType}\t${f.metadata.title}\n`,
);
} }
} }
finish(lib, 0);
}); });
program program
@@ -261,70 +196,98 @@ program
.description("Download and decrypt a single file") .description("Download and decrypt a single file")
.argument("<fileID>", "File ID (from `quak files`)") .argument("<fileID>", "File ID (from `quak files`)")
.option("--out <path>", "Output file path") .option("--out <path>", "Output file path")
.option("--collection <id>", "Accepted for compatibility; ignored") .option(
.action(async (fileIDStr: string, opts: { out?: string }) => { "--collection <id>",
await init(); "Collection ID (required to look up the file key)",
const client = requireSession(); )
const fileID = Number(fileIDStr); .action(
if (!Number.isFinite(fileID)) { async (
stderr.write("Invalid file ID\n"); fileIDStr: string,
process.exit(1); opts: { out?: string; collection?: string },
} ) => {
await init();
const client = requireSession();
const fileID = Number(fileIDStr);
if (!Number.isFinite(fileID)) {
stderr.write("Invalid file ID\n");
process.exit(1);
}
const lib = await openReadLibrary(client); const collections = await client.listCollections();
// Force a server round-trip so the file resolves against current state let targetCol;
// (issue #36 amendment, issue #52). if (opts.collection) {
const resolved = await freshFile(lib, fileID); targetCol = collections.find(
if (!resolved) { (c) => c.id === Number(opts.collection),
);
}
// Search all collections (or the specified one) for the file
const searchCols = targetCol ? [targetCol] : collections;
for (const col of searchCols) {
const files = await client.listFiles(col.id, col.key);
const file = files.find((f) => f.id === fileID);
if (file) {
const result = await client.downloadFile(file, opts.out);
stderr.write(
`${result.bytesWritten} bytes -> ${result.path}\n`,
);
return;
}
}
stderr.write(`File ${fileID} not found\n`); stderr.write(`File ${fileID} not found\n`);
finish(lib, 1); process.exit(1);
return; },
} );
const { photo, file } = resolved;
const result = await photo.original();
// Default name is the file's own title, as the pre-library CLI used
// (not the editedName-preferring projection title) (issue #52).
const outPath = opts.out ?? originalName(file);
copyFileSync(result.path, outPath);
stderr.write(`${result.bytes} bytes -> ${outPath}\n`);
finish(lib, 0);
});
program program
.command("get-thumb") .command("get-thumb")
.description("Download and decrypt a thumbnail") .description("Download and decrypt a thumbnail")
.argument("<fileID>", "File ID (from `quak files`)") .argument("<fileID>", "File ID (from `quak files`)")
.option("--out <path>", "Output file path") .option("--out <path>", "Output file path")
.option("--collection <id>", "Accepted for compatibility; ignored") .option(
.action(async (fileIDStr: string, opts: { out?: string }) => { "--collection <id>",
await init(); "Collection ID (required to look up the file key)",
const client = requireSession(); )
const fileID = Number(fileIDStr); .action(
if (!Number.isFinite(fileID)) { async (
stderr.write("Invalid file ID\n"); fileIDStr: string,
process.exit(1); opts: { out?: string; collection?: string },
} ) => {
await init();
const client = requireSession();
const fileID = Number(fileIDStr);
if (!Number.isFinite(fileID)) {
stderr.write("Invalid file ID\n");
process.exit(1);
}
const lib = await openReadLibrary(client); const collections = await client.listCollections();
// Force a server round-trip so the file resolves against current state let targetCol;
// (issue #36 amendment, issue #52). if (opts.collection) {
const resolved = await freshFile(lib, fileID); targetCol = collections.find(
if (!resolved) { (c) => c.id === Number(opts.collection),
);
}
const searchCols = targetCol ? [targetCol] : collections;
for (const col of searchCols) {
const files = await client.listFiles(col.id, col.key);
const file = files.find((f) => f.id === fileID);
if (file) {
const result = await client.downloadThumbnail(
file,
opts.out,
);
stderr.write(
`${result.bytesWritten} bytes -> ${result.path}\n`,
);
return;
}
}
stderr.write(`File ${fileID} not found\n`); stderr.write(`File ${fileID} not found\n`);
finish(lib, 1); process.exit(1);
return; },
} );
const { photo, file } = resolved;
const result = await photo.thumbnail();
// Default name is thumb_<file's own title>, as the pre-library CLI
// used (not the projection title) (issue #52).
const outPath = opts.out ?? thumbnailName(file);
copyFileSync(result.path, outPath);
stderr.write(`${result.bytes} bytes -> ${outPath}\n`);
finish(lib, 0);
});
program program
.command("backup-metadata") .command("backup-metadata")
@@ -340,12 +303,10 @@ program
.action(async (dir: string, opts: { exif?: boolean; all?: boolean }) => { .action(async (dir: string, opts: { exif?: boolean; all?: boolean }) => {
await init(); await init();
const client = requireSession(); const client = requireSession();
const lib = await openReadLibrary(client); await runMetadataBackup(client, dir, {
await runMetadataBackup(lib, client, dir, {
exif: opts.exif || opts.all, exif: opts.exif || opts.all,
onProgress: (msg) => stderr.write(msg + "\n"), onProgress: (msg) => stderr.write(msg + "\n"),
}); });
finish(lib, 0);
}); });
program program
@@ -360,16 +321,8 @@ program
const client = requireSession(); const client = requireSession();
stderr.write("Starting backup...\n"); stderr.write("Starting backup...\n");
const lib = await Library.open({ const result = await runBackup(client, dir, (msg) => {
client, if (!opts.json) stderr.write(msg + "\n");
downloadDirectory: dir,
cacheDirectory: cacheDirOption(),
});
const result = await lib.backup({
downloadDirectory: dir,
onProgress: (msg) => {
if (!opts.json) stderr.write(msg + "\n");
},
}); });
if (opts.json) { if (opts.json) {
@@ -390,7 +343,7 @@ program
} }
} }
finish(lib, result.failed > 0 ? 1 : 0); process.exit(result.failed > 0 ? 1 : 0);
}); });
const helper = program const helper = program
@@ -404,8 +357,7 @@ helper
.action(async (opts: { json?: boolean }) => { .action(async (opts: { json?: boolean }) => {
await init(); await init();
const client = requireSession(); const client = requireSession();
const lib = await openReadLibrary(client); const missing = await listMissingThumbnails(client, (msg) => {
const missing = await listMissingThumbnails(lib, client, (msg) => {
if (!opts.json) stderr.write(msg + "\n"); if (!opts.json) stderr.write(msg + "\n");
}); });
@@ -425,7 +377,6 @@ helper
} }
} }
} }
finish(lib, 0);
}); });
helper helper
@@ -441,61 +392,44 @@ helper
.action(async (opts: { file?: string[]; json?: boolean }) => { .action(async (opts: { file?: string[]; json?: boolean }) => {
await init(); await init();
const client = requireSession(); const client = requireSession();
const lib = await openReadLibrary(client);
let fileIDs: number[]; let fileIDs: number[];
if (opts.file && opts.file.length > 0) { if (opts.file && opts.file.length > 0) {
fileIDs = opts.file.map(Number).filter(Number.isFinite); fileIDs = opts.file.map(Number).filter(Number.isFinite);
} else { } else {
stderr.write("Scanning for missing thumbnails...\n"); stderr.write("Scanning for missing thumbnails...\n");
const missing = await listMissingThumbnails(lib, client, (msg) => { const missing = await listMissingThumbnails(client, (msg) => {
if (!opts.json) stderr.write(msg + "\n"); if (!opts.json) stderr.write(msg + "\n");
}); });
fileIDs = missing.map((m) => m.fileID); fileIDs = missing.map((m) => m.fileID);
if (fileIDs.length === 0) { if (fileIDs.length === 0) {
stderr.write("No missing thumbnails found.\n"); stderr.write("No missing thumbnails found.\n");
finish(lib, 0);
return; return;
} }
stderr.write(`Found ${fileIDs.length} file(s) to fix.\n`); stderr.write(`Found ${fileIDs.length} file(s) to fix.\n`);
} }
const results = await fixMissingThumbnails( const results = await fixMissingThumbnails(client, fileIDs, (msg) => {
lib, if (!opts.json) stderr.write(msg + "\n");
client, });
fileIDs,
(msg) => {
if (!opts.json) stderr.write(msg + "\n");
},
);
if (opts.json) { if (opts.json) {
stdout.write(JSON.stringify(results, null, 2) + "\n"); stdout.write(JSON.stringify(results, null, 2) + "\n");
} else { } else {
const fixed = results.filter((r) => r.status === "fixed").length; const ok = results.filter((r) => r.success).length;
const skipped = results.filter( const fail = results.filter((r) => !r.success).length;
(r) => r.status === "skipped",
).length;
const failed = results.filter((r) => r.status === "failed").length;
stderr.write(`\n--- Done ---\n`); stderr.write(`\n--- Done ---\n`);
stderr.write(` Fixed: ${fixed}\n`); stderr.write(` Fixed: ${ok}\n`);
stderr.write(` Skipped: ${skipped}\n`); stderr.write(` Failed: ${fail}\n`);
stderr.write(` Failed: ${failed}\n`); if (fail > 0) {
if (skipped > 0) {
stderr.write("\nSkipped (unsupported format):\n");
for (const r of results.filter((r) => r.status === "skipped")) {
stderr.write(` ${r.fileID}\t${r.title}\t${r.reason}\n`);
}
}
if (failed > 0) {
stderr.write("\nFailed files:\n"); stderr.write("\nFailed files:\n");
for (const r of results.filter((r) => r.status === "failed")) { for (const r of results.filter((r) => !r.success)) {
stderr.write(` ${r.fileID}\t${r.title}\t${r.reason}\n`); stderr.write(` ${r.fileID}\t${r.title}\t${r.error}\n`);
} }
} }
} }
finish(lib, results.some((r) => r.status === "failed") ? 1 : 0); process.exit(results.some((r) => !r.success) ? 1 : 0);
}); });
await init(); await init();
+98 -369
View File
@@ -1,65 +1,13 @@
// The backup command, rebuilt on the library API (issue #51).
//
// `lib.backup()` refreshes the library, then, for every file in scope, gets its
// original bytes onto disk under `downloadDirectory` and rebuilds the derived
// views (per-file sidecars, per-collection symlink trees, per-collection JSON)
// from the model. The on-disk layout is the historical one, unchanged:
//
// <downloadDirectory>/
// originals/<fileID>.<ext> the decrypted bytes
// originals/<fileID>.json per-file metadata sidecar
// collections/<name>/<title> symlink into ../../originals
// collections/<name>.json per-collection metadata
// failures.json durable ledger of unresolved failures
//
// Crash-safety rests on two properties. Bytes are present-means-complete: an
// original appears under `originals/` only via the content layer's atomic
// temp-then-rename, so a file that exists is whole and is never re-fetched — an
// interrupted run resumes by listing the directory. The derived views hold no
// unique state, so they are rebuilt every run; that repairs stale sidecars and
// missing or broken symlinks left by an earlier crash.
//
// Resilience (issue #8): no per-file condition aborts the run. A failed
// download or a failed symlink is caught, recorded in `failures.json` with a
// classification, a running attempt count, and the last-tried time, and the run
// continues. `result.failed` — and thus the CLI's exit code — stays non-zero
// while any failure remains unresolved and clears once every one succeeds. Each
// run reconciles the ledger against the files it attempted, so an entry for a
// file that has since left the library (deleted) or this run's scope is dropped
// rather than counted forever, which would poison a scheduled backup's exit code.
import { import {
copyFileSync, existsSync,
lstatSync,
mkdirSync, mkdirSync,
readFileSync,
readlinkSync,
renameSync,
rmSync,
statSync, statSync,
symlinkSync, symlinkSync,
writeFileSync, writeFileSync,
} from "node:fs"; } from "node:fs";
import { basename, dirname, extname, join, relative } from "node:path"; import { join, relative, extname } from "node:path";
import type { Client } from "./client.js";
import type { Collection, EnteFile } from "./model/types.js"; import type { EnteFile } from "./model/types.js";
export type ProgressCallback = (message: string) => void;
export interface BackupOptions {
// Where the backup tree lives. Required: with none, `backup()` throws
// before any network traffic. A library opened with a `downloadDirectory`
// supplies the default.
downloadDirectory?: string;
// Fetch and store full-resolution originals. Default true.
includeOriginals?: boolean;
// Also fetch and store thumbnails under `thumbnails/<fileID>.jpg`. Default
// false.
includeThumbnails?: boolean;
// Restrict the backup to albums with these names; others are left untouched.
onlyAlbumNames?: string[];
onProgress?: ProgressCallback;
}
export interface BackupError { export interface BackupError {
fileID: number; fileID: number;
@@ -69,358 +17,139 @@ export interface BackupError {
} }
export interface BackupResult { export interface BackupResult {
// Distinct files in scope this run.
totalFiles: number; totalFiles: number;
// Originals fetched (or copied from the cache) this run.
downloaded: number; downloaded: number;
// Originals already present and left untouched.
skipped: number; skipped: number;
// Files with an unresolved failure after this run (the ledger size); the
// CLI exits non-zero while this is above zero. A file can be both
// downloaded and failed if its bytes landed but its symlink did not.
failed: number; failed: number;
// This run's per-file errors, in encounter order.
errors: BackupError[]; errors: BackupError[];
} }
// The slice of the library that backup drives. `Library` implements it; a test export type ProgressCallback = (message: string) => void;
// can drive backup with a stand-in.
export interface BackupLibrary {
refresh(): Promise<void>;
listCollections(): Collection[];
listFiles(collectionID: number): EnteFile[];
// Get an original's bytes onto disk through the content cache/pools,
// returning where they landed (the cache, or a prior backup).
original(fileID: number): Promise<{ path: string }>;
thumbnail(fileID: number): Promise<{ path: string }>;
}
type FailureClass = "transient" | "permanent" | "unknown";
interface FailureEntry {
fileID: number;
title: string;
classification: FailureClass;
attempts: number;
lastTriedAt: number;
error: string;
}
const LEDGER_VERSION = 1;
const sanitizePath = (name: string): string => const sanitizePath = (name: string): string =>
name.replace(/[/\\:*?"<>|]/g, "_").replace(/^\.+/, "_"); name.replace(/[/\\:*?"<>|]/g, "_").replace(/^\.+/, "_");
// The originals/ filename for a file: `<id><ext>`, the extension taken from the const originalFileName = (file: EnteFile): string => {
// title (or `.bin`). Matches the content cache's own naming so a present check
// lines up with what a fetch would write.
const originalName = (file: EnteFile): string => {
const ext = extname(file.metadata.title || "") || ".bin"; const ext = extname(file.metadata.title || "") || ".bin";
return `${file.id}${ext}`; return `${file.id}${ext}`;
}; };
// A regular file with content is treated as complete. A zero-byte file is not:
// it is the shape an aborted write leaves and must be re-fetched.
const isPresent = (path: string): boolean => {
try {
const s = statSync(path);
return s.isFile() && s.size > 0;
} catch {
return false;
}
};
// Best-effort classification for the ledger. Retryable server/network problems
// are transient; refusals and local filesystem/decrypt errors are permanent;
// anything else is unknown. Both the error code and message are inspected.
const classify = (err: unknown): FailureClass => {
const e = err as NodeJS.ErrnoException;
const text =
`${e?.code ?? ""} ${err instanceof Error ? err.message : String(err)}`.toLowerCase();
if (
/timeout|timed out|econnreset|econnrefused|econnaborted|network|socket|eai_again|throttl|temporarily|429|500|502|503|504/.test(
text,
)
) {
return "transient";
}
if (
/enoent|eacces|eperm|eexist|eisdir|enotempty|erofs|enospc|not found|forbidden|unauthor|decrypt|truncat|401|403|404/.test(
text,
)
) {
return "permanent";
}
return "unknown";
};
const errorMessage = (err: unknown): string =>
err instanceof Error ? err.message : String(err);
// Copy bytes into `dest` via a temp file in the same directory plus rename, so
// `dest` appears only once it is whole ("present means complete").
const copyAtomic = (src: string, dest: string): void => {
if (src === dest) return;
const tmp = join(
dirname(dest),
`.quak-backup-${basename(dest)}-${process.pid}-${Math.random()
.toString(36)
.slice(2)}.tmp`,
);
try {
copyFileSync(src, tmp);
renameSync(tmp, dest);
} finally {
rmSync(tmp, { force: true });
}
};
// Ensure `linkPath` is a symlink to `target`, rebuilding a missing, wrong, or
// non-symlink entry. Throws on failure (a directory in the way, no permission)
// so the caller records it and moves on rather than aborting the run.
const rebuildSymlink = (linkPath: string, target: string): void => {
try {
const st = lstatSync(linkPath);
if (st.isSymbolicLink() && readlinkSync(linkPath) === target) return;
} catch {
// Nothing there (or unreadable): fall through to create it.
}
// Remove a wrong symlink or stray file. `force` ignores a missing path but
// still refuses a directory (no `recursive`), which surfaces as a failure.
rmSync(linkPath, { force: true });
symlinkSync(target, linkPath);
};
const loadLedger = (path: string): Map<number, FailureEntry> => {
const ledger = new Map<number, FailureEntry>();
try {
const parsed = JSON.parse(readFileSync(path, "utf-8")) as {
files?: Record<string, FailureEntry>;
};
for (const entry of Object.values(parsed.files ?? {})) {
if (entry && typeof entry.fileID === "number") {
ledger.set(entry.fileID, entry);
}
}
} catch {
// No ledger yet, or an unreadable one: start clean.
}
return ledger;
};
const saveLedger = (path: string, ledger: Map<number, FailureEntry>): void => {
if (ledger.size === 0) {
rmSync(path, { force: true });
return;
}
const files: Record<string, FailureEntry> = {};
for (const [fileID, entry] of ledger) files[String(fileID)] = entry;
writeFileSync(
path,
JSON.stringify({ version: LEDGER_VERSION, files }, null, 2),
);
};
const writeSidecar = (path: string, file: EnteFile): void => {
const meta: Record<string, unknown> = {
id: file.id,
collectionID: file.collectionID,
ownerID: file.ownerID,
metadata: file.metadata,
};
if (file.magicMetadata) meta.magicMetadata = file.magicMetadata;
if (file.pubMagicMetadata) meta.pubMagicMetadata = file.pubMagicMetadata;
writeFileSync(path, JSON.stringify(meta, null, 2));
};
export const runBackup = async ( export const runBackup = async (
lib: BackupLibrary, client: Client,
opts: BackupOptions, outDir: string,
onProgress?: ProgressCallback,
): Promise<BackupResult> => { ): Promise<BackupResult> => {
const downloadDirectory = opts.downloadDirectory; const log = onProgress ?? (() => {});
if (!downloadDirectory) {
throw new Error(
"backup requires a downloadDirectory (pass one to backup() or " +
"open the library with one)",
);
}
const includeOriginals = opts.includeOriginals ?? true;
const includeThumbnails = opts.includeThumbnails ?? false;
const log = opts.onProgress ?? (() => {});
const only = opts.onlyAlbumNames ? new Set(opts.onlyAlbumNames) : undefined;
log("Refreshing library..."); mkdirSync(outDir, { recursive: true });
await lib.refresh(); const originalsDir = join(outDir, "originals");
const originalsDir = join(downloadDirectory, "originals");
const collectionsDir = join(downloadDirectory, "collections");
const thumbnailsDir = join(downloadDirectory, "thumbnails");
mkdirSync(originalsDir, { recursive: true }); mkdirSync(originalsDir, { recursive: true });
const collectionsDir = join(outDir, "collections");
mkdirSync(collectionsDir, { recursive: true }); mkdirSync(collectionsDir, { recursive: true });
if (includeThumbnails) mkdirSync(thumbnailsDir, { recursive: true });
const ledgerPath = join(downloadDirectory, "failures.json"); log("Fetching collections...");
const ledger = loadLedger(ledgerPath); const collections = await client.listCollections();
const now = Date.now(); const downloadedIDs = new Set<number>();
// Collections in scope, and the distinct files across them (a file shared let totalFiles = 0;
// by two albums is one original).
const collections = lib
.listCollections()
.filter((c) => (only ? only.has(c.name) : true));
const collectionName = new Map<number, string>();
for (const c of collections) collectionName.set(c.id, c.name);
const distinct = new Map<number, EnteFile>();
const filesByCollection = new Map<number, EnteFile[]>();
for (const c of collections) {
const files = lib.listFiles(c.id);
filesByCollection.set(c.id, files);
for (const f of files) if (!distinct.has(f.id)) distinct.set(f.id, f);
}
const errors: BackupError[] = [];
const failedThisRun = new Set<number>();
let downloaded = 0; let downloaded = 0;
let skipped = 0; let skipped = 0;
let failed = 0;
const errors: BackupError[] = [];
const recordFailure = ( for (const col of collections) {
file: EnteFile, const colDirName = sanitizePath(col.name || `collection-${col.id}`);
collection: string,
err: unknown,
): void => {
// Count at most one attempt per file per run: a file whose original
// and thumbnail both fail this run must not double its attempt count
// or appear twice in errors.
if (failedThisRun.has(file.id)) return;
const error = errorMessage(err);
errors.push({
fileID: file.id,
title: file.metadata.title,
collection,
error,
});
const prior = ledger.get(file.id);
ledger.set(file.id, {
fileID: file.id,
title: file.metadata.title,
classification: classify(err),
attempts: (prior?.attempts ?? 0) + 1,
lastTriedAt: now,
error,
});
failedThisRun.add(file.id);
};
// Phase 1: get the bytes. Fetch each pending original (and optional
// thumbnail) through the content cache/pools and place it under the backup
// tree; a present file is left as is.
if (includeOriginals) {
for (const [fileID, file] of distinct) {
const dest = join(originalsDir, originalName(file));
if (isPresent(dest)) {
skipped++;
continue;
}
try {
log(`Fetching original ${file.metadata.title} (${fileID})...`);
const { path } = await lib.original(fileID);
copyAtomic(path, dest);
downloaded++;
} catch (err) {
log(
`FAILED original ${file.metadata.title}: ${errorMessage(err)}`,
);
recordFailure(
file,
collectionName.get(file.collectionID) ?? "",
err,
);
}
}
}
if (includeThumbnails) {
for (const [fileID, file] of distinct) {
const dest = join(thumbnailsDir, `${fileID}.jpg`);
if (isPresent(dest)) continue;
try {
const { path } = await lib.thumbnail(fileID);
copyAtomic(path, dest);
} catch (err) {
recordFailure(
file,
collectionName.get(file.collectionID) ?? "",
err,
);
}
}
}
// Phase 2: rebuild the derived views from the model. Sidecars first, for
// every present original (this repairs stale ones).
if (includeOriginals) {
for (const [fileID, file] of distinct) {
const orig = join(originalsDir, originalName(file));
if (isPresent(orig)) {
writeSidecar(join(originalsDir, `${fileID}.json`), file);
}
}
}
// Then the per-collection symlink trees and JSON.
for (const c of collections) {
const colDirName = sanitizePath(c.name || `collection-${c.id}`);
const colDir = join(collectionsDir, colDirName); const colDir = join(collectionsDir, colDirName);
mkdirSync(colDir, { recursive: true }); mkdirSync(colDir, { recursive: true });
const files = filesByCollection.get(c.id) ?? []; log(`[${col.name}] Fetching file list...`);
const metaFiles: { id: number; metadata: EnteFile["metadata"] }[] = []; const files = await client.listFiles(col.id, col.key);
log(`[${col.name}] ${files.length} file(s)`);
const collectionMeta: {
id: number;
name: string;
type: string;
files: { id: number; metadata: EnteFile["metadata"] }[];
} = {
id: col.id,
name: col.name,
type: col.type,
files: [],
};
for (const file of files) { for (const file of files) {
metaFiles.push({ id: file.id, metadata: file.metadata }); totalFiles++;
if (!includeOriginals) continue; const origName = originalFileName(file);
const orig = join(originalsDir, originalName(file)); const origPath = join(originalsDir, origName);
if (!isPresent(orig)) continue;
const linkName = sanitizePath( const linkName = sanitizePath(
file.metadata.title || `file-${file.id}`, file.metadata.title || `file-${file.id}`,
); );
const linkPath = join(colDir, linkName); const linkPath = join(colDir, linkName);
try {
rebuildSymlink(linkPath, relative(colDir, orig)); if (!downloadedIDs.has(file.id)) {
} catch (err) { if (existsSync(origPath) && statSync(origPath).size > 0) {
log( skipped++;
`FAILED symlink ${c.name}/${linkName}: ${errorMessage(err)}`, downloadedIDs.add(file.id);
); } else {
recordFailure(file, c.name, err); try {
log(`[${col.name}] Downloading ${linkName}...`);
await client.downloadFile(file, origPath);
downloaded++;
downloadedIDs.add(file.id);
} catch (err) {
log(
`[${col.name}] FAILED ${linkName}: ${err instanceof Error ? err.message : err}`,
);
failed++;
errors.push({
fileID: file.id,
title: file.metadata.title,
collection: col.name,
error:
err instanceof Error
? err.message
: String(err),
});
continue;
}
}
} }
// Write per-file metadata JSON alongside the original
const metaJsonPath = join(originalsDir, `${file.id}.json`);
if (!existsSync(metaJsonPath)) {
const fileMeta: Record<string, unknown> = {
id: file.id,
collectionID: file.collectionID,
ownerID: file.ownerID,
metadata: file.metadata,
};
if (file.magicMetadata) {
fileMeta.magicMetadata = file.magicMetadata;
}
if (file.pubMagicMetadata) {
fileMeta.pubMagicMetadata = file.pubMagicMetadata;
}
writeFileSync(metaJsonPath, JSON.stringify(fileMeta, null, 2));
}
if (!existsSync(linkPath) && existsSync(origPath)) {
const target = relative(colDir, origPath);
symlinkSync(target, linkPath);
}
collectionMeta.files.push({
id: file.id,
metadata: file.metadata,
});
} }
writeFileSync( writeFileSync(
join(collectionsDir, `${colDirName}.json`), join(collectionsDir, `${colDirName}.json`),
JSON.stringify( JSON.stringify(collectionMeta, null, 2),
{ id: c.id, name: c.name, type: c.type, files: metaFiles },
null,
2,
),
); );
} }
// Reconcile the ledger against what this run actually attempted: an entry return { totalFiles, downloaded, skipped, failed, errors };
// survives only for a file that failed this run. A file that succeeded had
// its failure resolved; a file gone from the library (deleted) or outside
// this run's scope is not something this run can resolve, so keeping its
// stale entry would keep the exit code non-zero forever — a single
// since-deleted photo would fail every future scheduled backup.
for (const fileID of [...ledger.keys()]) {
if (!failedThisRun.has(fileID)) ledger.delete(fileID);
}
saveLedger(ledgerPath, ledger);
return {
totalFiles: distinct.size,
downloaded,
skipped,
failed: ledger.size,
errors,
};
}; };
-40
View File
@@ -1,40 +0,0 @@
// How the CLI presents a file's identity in `files`, `get`, and `get-thumb`.
//
// These read the file's own decrypted metadata — the raw title and the
// creationTime in microseconds — rather than the `PhotoRecord` projection the
// rest of the library exposes. The projection prefers `editedName`/`editedTime`
// and reports time in milliseconds, which is right for a photo browser but
// would change the CLI's externally-visible output. The pre-library CLI printed
// `metadata.title` and `metadata.creationTime` and named downloads after
// `metadata.title`, and issue #52 requires that output stay byte-identical, so
// the commands shape their output from the raw `EnteFile` through here.
import type { EnteFile, FileType, Microseconds } from "./model/types.js";
// One row of `quak files --json`.
export interface FileListRow {
id: number;
title: string;
fileType: FileType;
creationTime: Microseconds;
collectionID: number;
}
export const fileListRow = (file: EnteFile): FileListRow => ({
id: file.id,
title: file.metadata.title,
fileType: file.metadata.fileType,
creationTime: file.metadata.creationTime,
collectionID: file.collectionID,
});
// One line of `quak files` in its human, tab-separated form.
export const fileListLine = (file: EnteFile): string =>
`${file.id}\t${file.metadata.fileType}\t${file.metadata.title}`;
// Default output path for `quak get` when `--out` is not given.
export const originalName = (file: EnteFile): string => file.metadata.title;
// Default output path for `quak get-thumb` when `--out` is not given.
export const thumbnailName = (file: EnteFile): string =>
`thumb_${file.metadata.title}`;
-62
View File
@@ -1,62 +0,0 @@
// How the CLI's read commands obtain current data.
//
// `collections`, `files --collection`, `get`, and `get-thumb` must answer for
// the account's state at the moment the command runs, not for whatever the
// local cache last happened to hold (owner amendment, issue #36). Each helper
// therefore forces a server round-trip through `Library.fresh()` and only then
// reads — so a collection, file, or metadata change made elsewhere is visible.
//
// `collections` and `files` also list in the library's own enumeration order —
// `listCollections()`/`listFiles()`, the order the pre-library CLI printed —
// rather than the `albums`/`photos` projection's newest-first order, which
// re-sorts the rows. The field values still come from each record's raw
// metadata via `cli-output.ts`.
import type { Collection, EnteFile } from "./model/types.js";
import type { Photo, PhotosAPI } from "./library/index.js";
// The slice of `Library` these helpers read. `Library` satisfies it
// structurally; a test can drive them with a stand-in that records the
// `fresh()` call and serves records in a known enumeration order.
export interface FreshReadLibrary {
fresh(): Promise<unknown>;
listCollections(): Collection[];
getCollection(id: number): Collection | undefined;
listFiles(collectionID: number): EnteFile[];
getFileByID(fileID: number): EnteFile | undefined;
photos: Pick<PhotosAPI, "byID">;
}
// Every live collection, current as of a forced refresh, in enumeration order.
export const freshCollections = async (
lib: FreshReadLibrary,
): Promise<Collection[]> => {
await lib.fresh();
return lib.listCollections();
};
// The files of one collection, current as of a forced refresh, in enumeration
// order. `undefined` (not an empty list) when the collection does not exist, so
// the caller can tell "no such collection" from "an empty collection".
export const freshFiles = async (
lib: FreshReadLibrary,
collectionID: number,
): Promise<EnteFile[] | undefined> => {
await lib.fresh();
if (!lib.getCollection(collectionID)) return undefined;
return lib.listFiles(collectionID);
};
// One file, current as of a forced refresh, resolved to both its content
// handle (`Photo`, for fetching bytes) and its raw record (`EnteFile`, for the
// default output name and field values). `undefined` when the file is unknown.
export const freshFile = async (
lib: FreshReadLibrary,
fileID: number,
): Promise<{ photo: Photo; file: EnteFile } | undefined> => {
await lib.fresh();
const photo = lib.photos.byID({ fileID });
const file = lib.getFileByID(fileID);
if (!photo || !file) return undefined;
return { photo, file };
};
+38 -138
View File
@@ -7,16 +7,11 @@ import {
} from "./auth/login.js"; } from "./auth/login.js";
import { unwrapAuth } from "./auth/unwrap.js"; import { unwrapAuth } from "./auth/unwrap.js";
import { init, fromBase64, toBase64 } from "./crypto/index.js"; import { init, fromBase64, toBase64 } from "./crypto/index.js";
import { fetchMLDataBatch, type MLData } from "./mldata-fetch.js";
import { decryptCollection, decryptFile } from "./model/index.js"; import { decryptCollection, decryptFile } from "./model/index.js";
import { import {
downloadFile as dlFile, downloadFile as dlFile,
downloadThumbnail as dlThumb, downloadThumbnail as dlThumb,
} from "./download/index.js"; } from "./download/index.js";
import {
makeDownloadContentSource,
type ContentSource,
} from "./library/content.js";
import type { import type {
Collection, Collection,
EnteFile, EnteFile,
@@ -42,22 +37,6 @@ export interface ClientSnapshot {
publicKey: string; publicKey: string;
} }
// The result of a resumable enumeration. Live decrypted records and deleted
// ids are kept apart on purpose: a tombstone carries no key or metadata to
// decrypt, so it is a bare id rather than a hollowed-out record. `cursor` is
// the max `updationTime` seen, to pass back into the next call.
export interface CollectionsPage {
collections: Collection[];
deleted: number[];
cursor: number;
}
export interface FilesPage {
files: EnteFile[];
deleted: number[];
cursor: number;
}
export class Client { export class Client {
private readonly api: ApiClient; private readonly api: ApiClient;
private readonly email: string; private readonly email: string;
@@ -145,14 +124,6 @@ export class Client {
return this.api; return this.api;
} }
// The content-cache byte source over this client's API: each fetch is the
// download layer's request + streaming decrypt + atomic write. `Library`
// calls this to enable the on-disk content cache.
contentSource(): ContentSource {
this.assertLoggedIn();
return makeDownloadContentSource(this.api);
}
private assertLoggedIn(): void { private assertLoggedIn(): void {
if (this.loggedOut) throw new Error("Client has been logged out"); if (this.loggedOut) throw new Error("Client has been logged out");
} }
@@ -179,123 +150,52 @@ export class Client {
this.api.clearAuthToken(); this.api.clearAuthToken();
} }
// Enumerate collections changed since `sinceTime`. Live collections are
// decrypted; tombstoned ones (isDeleted) are surfaced as bare ids. The
// returned cursor is the max `updationTime` seen — including tombstones, so
// the next sync resumes past them — falling back to `sinceTime` when the
// response is empty. `/collections/v2` returns the whole changed set in one
// response, so there is no pagination here.
async collectionsSince(args: {
sinceTime: number;
}): Promise<CollectionsPage> {
this.assertLoggedIn();
const { collections: raws } = await this.api.getJSON<{
collections: RawCollection[];
}>("/collections/v2", { sinceTime: args.sinceTime });
const collections: Collection[] = [];
const deleted: number[] = [];
let cursor = args.sinceTime;
for (const raw of raws) {
if (raw.isDeleted) {
deleted.push(raw.id);
} else {
collections.push(
decryptCollection(
raw,
{
masterKey: this.masterKey,
publicKey: this.publicKey,
secretKey: this.secretKey,
},
this.userID,
),
);
}
if (raw.updationTime > cursor) cursor = raw.updationTime;
}
return { collections, deleted, cursor };
}
// Enumerate a collection's files changed since `sinceTime`, paginating the
// diff from that cursor. Live rows are decrypted; tombstoned ones are
// surfaced as bare ids. Returns the final cursor to resume from.
async filesSince(args: {
collectionID: number;
collectionKey: Uint8Array;
sinceTime: number;
}): Promise<FilesPage> {
this.assertLoggedIn();
const { collectionID, collectionKey } = args;
const files: EnteFile[] = [];
const deleted: number[] = [];
let cursor = args.sinceTime;
for (;;) {
const { diff, hasMore } = await this.api.getJSON<{
diff: RawEnteFile[];
hasMore: boolean;
}>("/collections/v2/diff", { collectionID, sinceTime: cursor });
let pageMax = cursor;
for (const raw of diff) {
if (raw.isDeleted) {
deleted.push(raw.id);
} else {
files.push(decryptFile(raw, collectionKey));
}
if (raw.updationTime > pageMax) pageMax = raw.updationTime;
}
if (!hasMore) {
cursor = pageMax;
break;
}
// The server says there is more, but this page did not advance the
// cursor: following hasMore would refetch the same page forever
// (#7). Stop with a clear error instead of looping.
if (pageMax <= cursor) {
throw new Error(
`/collections/v2/diff for collection ${collectionID} ` +
`returned hasMore with a cursor that did not advance ` +
`(stuck at ${cursor}); refusing to loop`,
);
}
cursor = pageMax;
}
return { files, deleted, cursor };
}
// Whole-account listing: every live collection, deletions hidden. A thin
// wrapper over `collectionsSince` from the beginning of time.
async listCollections(): Promise<Collection[]> { async listCollections(): Promise<Collection[]> {
const { collections } = await this.collectionsSince({ sinceTime: 0 }); this.assertLoggedIn();
return collections; const { collections } = await this.api.getJSON<{
collections: RawCollection[];
}>("/collections/v2", { sinceTime: 0 });
// The sync API keeps returning deleted collections as tombstones
// (isDeleted: true); their diff endpoint 404s, so drop them.
return collections
.filter((raw) => !raw.isDeleted)
.map((raw) =>
decryptCollection(
raw,
{
masterKey: this.masterKey,
publicKey: this.publicKey,
secretKey: this.secretKey,
},
this.userID,
),
);
} }
// Every live file in a collection, deletions hidden. A thin wrapper over
// `filesSince` from the beginning of time.
async listFiles( async listFiles(
collectionID: number, collectionID: number,
collectionKey: Uint8Array, collectionKey: Uint8Array,
): Promise<EnteFile[]> { ): Promise<EnteFile[]> {
const { files } = await this.filesSince({
collectionID,
collectionKey,
sinceTime: 0,
});
return files;
}
// Fetch machine-learning data (face detections + CLIP embeddings) for up
// to a batch of files, each decrypted with its own key. One request; the
// library batches at `MLDATA_BATCH_SIZE` and schedules each batch through
// its metadata request pool.
async fetchMLData(args: {
fileIDs: number[];
fileKeys: Map<number, Uint8Array>;
}): Promise<Map<number, MLData>> {
this.assertLoggedIn(); this.assertLoggedIn();
return fetchMLDataBatch(this.api, args.fileIDs, args.fileKeys); const allFiles: EnteFile[] = [];
let sinceTime = 0;
for (;;) {
const { diff, hasMore } = await this.api.getJSON<{
diff: RawEnteFile[];
hasMore: boolean;
}>("/collections/v2/diff", { collectionID, sinceTime });
for (const raw of diff) {
if (!raw.isDeleted) {
allFiles.push(decryptFile(raw, collectionKey));
}
if (raw.updationTime > sinceTime) {
sinceTime = raw.updationTime;
}
}
if (!hasMore) break;
}
return allFiles;
} }
async downloadFile( async downloadFile(
+55 -179
View File
@@ -1,6 +1,5 @@
import { randomUUID } from "node:crypto"; import { randomUUID } from "node:crypto";
import { open, rename, rm } from "node:fs/promises"; import { rename, rm, writeFile } from "node:fs/promises";
import type { FileHandle } from "node:fs/promises";
import { dirname, join } from "node:path"; import { dirname, join } from "node:path";
import { import {
fromBase64, fromBase64,
@@ -20,101 +19,42 @@ export interface DownloadResult {
bytesWritten: number; bytesWritten: number;
} }
// Fired as decrypted plaintext accumulates, with the running total of
// plaintext bytes recovered so far. Within one download it is non-decreasing
// and its last value equals the final `bytesWritten`. A retry restarts the
// file from byte zero (see `fetchAndDecrypt`), so a fresh attempt begins its
// own count from zero.
export type ProgressCallback = (bytesDone: number) => void;
const ENC_CHUNK_SIZE = STREAM_CHUNK_SIZE + STREAM_CHUNK_OVERHEAD; const ENC_CHUNK_SIZE = STREAM_CHUNK_SIZE + STREAM_CHUNK_OVERHEAD;
// Decrypt a secretstream body, handing each plaintext chunk to `sink` as it is
// produced rather than accumulating the whole file. Peak memory is one
// ciphertext chunk of network buffer plus one plaintext chunk — bounded by
// `STREAM_CHUNK_SIZE` regardless of the file's size — so a multi-gigabyte video
// no longer needs its size again in RAM. Returns the total plaintext length.
//
// The truncation contract is exactly the buffered version's, only the sink is
// new: a body cut short still decrypts and authenticates up to its last whole
// chunk, so the absence of TAG_FINAL is the sole evidence it was cut short, and
// this throws rather than let a caller keep a short file. The sink has already
// seen those chunks by then; the caller (`decryptToTemp`) stages them in a temp
// file that is renamed into place only on a clean return, so a throw leaves
// nothing on disk.
const streamDecrypt = async ( const streamDecrypt = async (
stream: ReadableStream<Uint8Array>, stream: ReadableStream<Uint8Array>,
header: Uint8Array, header: Uint8Array,
key: Uint8Array, key: Uint8Array,
sink: (plaintext: Uint8Array) => Promise<void>, ): Promise<Uint8Array> => {
onProgress?: ProgressCallback,
): Promise<number> => {
const state = initStreamPull(header, key); const state = initStreamPull(header, key);
const reader = stream.getReader(); const reader = stream.getReader();
// Incoming reads are held as-is and only stitched into a contiguous chunk let buffer = new Uint8Array(0);
// at each `ENC_CHUNK_SIZE` boundary, so every received byte is copied once. const plainChunks: Uint8Array[] = [];
// Concatenating on each read instead — reallocating the whole accumulator
// per read — is O(n^2) in the bytes buffered, and for a 4 MiB chunk that
// memory churn dwarfs the libsodium decryption itself.
const pending: Uint8Array[] = [];
let pendingBytes = 0;
let totalPlain = 0; let totalPlain = 0;
let chunksPulled = 0; let chunksPulled = 0;
let lastTag = -1; let lastTag = -1;
// Remove the first `size` bytes from `pending` as one contiguous buffer.
// A read that straddles the boundary is split with `subarray` (a view, no
// copy); its tail stays queued for the next chunk. `size` never exceeds
// `pendingBytes`, so the queue always holds enough.
const takeContiguous = (size: number): Uint8Array => {
const out = new Uint8Array(size);
let offset = 0;
while (offset < size) {
const piece = pending[0]!;
const need = size - offset;
if (piece.length <= need) {
out.set(piece, offset);
offset += piece.length;
pending.shift();
} else {
out.set(piece.subarray(0, need), offset);
pending[0] = piece.subarray(need);
offset += need;
}
}
pendingBytes -= size;
return out;
};
const consume = async (
plaintext: Uint8Array,
tag: number,
): Promise<void> => {
await sink(plaintext);
totalPlain += plaintext.length;
chunksPulled++;
lastTag = tag;
onProgress?.(totalPlain);
};
for (;;) { for (;;) {
const { done, value } = await reader.read(); const { done, value } = await reader.read();
if (value && value.length > 0) { if (value) {
pending.push(value); const merged = new Uint8Array(buffer.length + value.length);
pendingBytes += value.length; merged.set(buffer);
merged.set(value, buffer.length);
buffer = merged;
} }
while (pendingBytes >= ENC_CHUNK_SIZE) { while (buffer.length >= ENC_CHUNK_SIZE) {
const encChunk = takeContiguous(ENC_CHUNK_SIZE); const encChunk = buffer.slice(0, ENC_CHUNK_SIZE);
// A whole chunk that fails to authenticate while the stream carries buffer = buffer.slice(ENC_CHUNK_SIZE);
// on is corruption, not truncation; that error propagates unchanged.
const { plaintext, tag } = pullStreamChunk(state, encChunk); const { plaintext, tag } = pullStreamChunk(state, encChunk);
await consume(plaintext, tag); plainChunks.push(plaintext);
totalPlain += plaintext.length;
chunksPulled++;
lastTag = tag;
} }
if (done) { if (done) {
if (pendingBytes > 0) { if (buffer.length > 0) {
const buffer = takeContiguous(pendingBytes);
// Whatever is left over once every whole chunk has been // Whatever is left over once every whole chunk has been
// consumed must be the stream's final chunk, and a final // consumed must be the stream's final chunk, and a final
// chunk that actually arrived in full authenticates. If it // chunk that actually arrived in full authenticates. If it
@@ -122,9 +62,7 @@ const streamDecrypt = async (
// ordinary shape of a dropped connection. Poly1305 cannot // ordinary shape of a dropped connection. Poly1305 cannot
// tell a partial chunk from a corrupt one, so this is // tell a partial chunk from a corrupt one, so this is
// reported as the truncation it almost always is, with the // reported as the truncation it almost always is, with the
// authentication failure kept as the error's cause. Only the // authentication failure kept as the error's cause.
// pull is guarded: a sink failure on a chunk that did
// authenticate is a disk error, not a truncation.
let pulled; let pulled;
try { try {
pulled = pullStreamChunk(state, buffer); pulled = pullStreamChunk(state, buffer);
@@ -134,7 +72,10 @@ const streamDecrypt = async (
{ cause: err }, { cause: err },
); );
} }
await consume(pulled.plaintext, pulled.tag); plainChunks.push(pulled.plaintext);
totalPlain += pulled.plaintext.length;
chunksPulled++;
lastTag = pulled.tag;
} }
break; break;
} }
@@ -143,6 +84,8 @@ const streamDecrypt = async (
// Only the last chunk of a secretstream carries TAG_FINAL. Everything a // Only the last chunk of a secretstream carries TAG_FINAL. Everything a
// dropped connection did deliver still decrypts and authenticates, so the // dropped connection did deliver still decrypts and authenticates, so the
// absence of TAG_FINAL is the only evidence that the body was cut short. // absence of TAG_FINAL is the only evidence that the body was cut short.
// Returning a short plaintext here would put a corrupt file on disk that
// later backup runs would treat as complete.
if (chunksPulled === 0) { if (chunksPulled === 0) {
throw new TruncatedStreamError( throw new TruncatedStreamError(
"download: stream truncated: response body contained no secretstream chunks", "download: stream truncated: response body contained no secretstream chunks",
@@ -154,54 +97,31 @@ const streamDecrypt = async (
`download: stream truncated: last chunk tag ${lastTag}, expected TAG_FINAL (${tagFinal})`, `download: stream truncated: last chunk tag ${lastTag}, expected TAG_FINAL (${tagFinal})`,
); );
} }
return totalPlain;
const result = new Uint8Array(totalPlain);
let offset = 0;
for (const chunk of plainChunks) {
result.set(chunk, offset);
offset += chunk.length;
}
return result;
}; };
// Stage a write to `destination` atomically and durably, then rename it into // Write `plaintext` to `destination` atomically: stage it in a temporary
// place. `fill` writes the contents into the open temp file handle — either the // sibling file (same directory, so the rename cannot cross a filesystem
// whole buffer at once (`writeAtomic`) or chunk by chunk as they decrypt // boundary) and rename it into place. Callers therefore never observe a
// (`decryptToTemp`). The temp file is a sibling of the destination (same // partially written destination, and a pre-existing file at that path is
// directory, so the rename cannot cross a filesystem boundary), so callers
// never observe a partially written destination, and a pre-existing file is
// replaced only once the new contents are complete on disk. // replaced only once the new contents are complete on disk.
// const writeAtomic = async (
// Durability against a power cut needs two fsyncs. Without them the write can
// return while the data or the rename is still only in the kernel's page
// cache, and a crash then resurrects an empty renamed file — exactly the
// corruption a later backup run treats as a complete download. So the temp
// file's contents are fsynced before the rename, and the containing directory
// is fsynced after it, so both the bytes and the new directory entry are on
// stable storage before this returns.
//
// On any failure — including a `fill` that throws because the stream was
// truncated — the temp file is removed, so the destination is untouched and no
// scratch file is left to fill the disk on repeated failures.
const stageAtomic = async (
destination: string, destination: string,
fill: (handle: FileHandle) => Promise<void>, plaintext: Uint8Array,
): Promise<void> => { ): Promise<void> => {
const dir = dirname(destination);
// The random suffix keeps concurrent downloads of the same destination // The random suffix keeps concurrent downloads of the same destination
// from stepping on each other's temporary file. // from stepping on each other's temporary file.
const tmpPath = join(dir, `.quak-${randomUUID()}.tmp`); const tmpPath = join(dirname(destination), `.quak-${randomUUID()}.tmp`);
try { try {
const handle = await open(tmpPath, "w"); await writeFile(tmpPath, plaintext);
try {
await fill(handle);
await handle.sync();
} finally {
await handle.close();
}
await rename(tmpPath, destination); await rename(tmpPath, destination);
// Fsync the directory so the rename itself survives a crash: renaming
// over a synced temp file still leaves the new directory entry in the
// page cache until the directory is synced.
const dirHandle = await open(dir, "r");
try {
await dirHandle.sync();
} finally {
await dirHandle.close();
}
} catch (err) { } catch (err) {
// Best-effort cleanup. A failure to remove the temporary file must // Best-effort cleanup. A failure to remove the temporary file must
// never replace the error that actually explains what went wrong. // never replace the error that actually explains what went wrong.
@@ -210,44 +130,7 @@ const stageAtomic = async (
} }
}; };
// Write `plaintext` to `destination` atomically and durably. Exported so the // Fetch a stream and decrypt it, retrying the whole sequence.
// metadata store can reuse the same durable write for small whole-buffer
// payloads; originals go through `decryptToTemp` instead so they never buffer.
export const writeAtomic = async (
destination: string,
plaintext: Uint8Array,
): Promise<void> =>
stageAtomic(destination, (handle) => handle.writeFile(plaintext));
// Decrypt `stream` straight to `destination`, one plaintext chunk at a time,
// under the atomic writer's temp-then-rename discipline. Memory stays bounded
// by the chunk size: each decrypted chunk is written to the temp file and
// dropped. The rename happens only after the stream authenticates as terminated
// on TAG_FINAL; a truncated stream throws and leaves the destination untouched.
// Returns the plaintext length written.
const decryptToTemp = async (
destination: string,
stream: ReadableStream<Uint8Array>,
header: Uint8Array,
key: Uint8Array,
onProgress?: ProgressCallback,
): Promise<number> => {
let bytesWritten = 0;
await stageAtomic(destination, async (handle) => {
bytesWritten = await streamDecrypt(
stream,
header,
key,
async (plaintext) => {
await handle.write(plaintext);
},
onProgress,
);
});
return bytesWritten;
};
// Fetch a stream and decrypt it to `destination`, retrying the whole sequence.
// //
// The request is only the first third of a download. `getXStream` returns as // The request is only the first third of a download. `getXStream` returns as
// soon as headers arrive, and the bytes are pulled here, so a socket reset // soon as headers arrive, and the bytes are pulled here, so a socket reset
@@ -260,60 +143,53 @@ const decryptToTemp = async (
// four attempts would mean sixteen requests for one file. The policy comes // four attempts would mean sixteen requests for one file. The policy comes
// from the client so a caller that configured one gets it here too. // from the client so a caller that configured one gets it here too.
// //
// Because the plaintext is streamed to disk rather than buffered, the atomic // A retry starts the file over from byte zero: the secretstream pull state is
// write is part of the retried unit. A retry starts the file over from byte // not resumable and there is no Range support on these endpoints.
// zero — the secretstream pull state is not resumable and there is no Range
// support — staging into a fresh temp file each time: a failed attempt writes
// and then removes its own temp file, and only the attempt that reaches
// TAG_FINAL renames one into place, so a download that needed three tries still
// performs exactly one rename over the destination.
const fetchAndDecrypt = async ( const fetchAndDecrypt = async (
api: ApiClient, api: ApiClient,
openStream: () => Promise<ReadableStream<Uint8Array>>, openStream: () => Promise<ReadableStream<Uint8Array>>,
header: Uint8Array, header: Uint8Array,
key: Uint8Array, key: Uint8Array,
destination: string, ): Promise<Uint8Array> =>
onProgress?: ProgressCallback,
): Promise<number> =>
withRetry(async () => { withRetry(async () => {
const stream = await openStream(); const stream = await openStream();
return decryptToTemp(destination, stream, header, key, onProgress); return streamDecrypt(stream, header, key);
}, api.getRetryOptions()); }, api.getRetryOptions());
export const downloadFile = async ( export const downloadFile = async (
api: ApiClient, api: ApiClient,
file: EnteFile, file: EnteFile,
outPath?: string, outPath?: string,
onProgress?: ProgressCallback,
): Promise<DownloadResult> => { ): Promise<DownloadResult> => {
const resolvedPath = outPath ?? file.metadata.title; const resolvedPath = outPath ?? file.metadata.title;
const header = fromBase64(file.file.decryptionHeader); const header = fromBase64(file.file.decryptionHeader);
const bytesWritten = await fetchAndDecrypt( const plaintext = await fetchAndDecrypt(
api, api,
() => api.getFileStream(file.id, { retry: false }), () => api.getFileStream(file.id, { retry: false }),
header, header,
file.key, file.key,
resolvedPath,
onProgress,
); );
return { path: resolvedPath, bytesWritten }; // Outside the retry, deliberately: only the attempt that produced a
// complete, authenticated plaintext gets to stage a temporary file, so a
// download that needed three tries still performs exactly one write and
// one rename.
await writeAtomic(resolvedPath, plaintext);
return { path: resolvedPath, bytesWritten: plaintext.length };
}; };
export const downloadThumbnail = async ( export const downloadThumbnail = async (
api: ApiClient, api: ApiClient,
file: EnteFile, file: EnteFile,
outPath?: string, outPath?: string,
onProgress?: ProgressCallback,
): Promise<DownloadResult> => { ): Promise<DownloadResult> => {
const resolvedPath = outPath ?? `thumb_${file.metadata.title}`; const resolvedPath = outPath ?? `thumb_${file.metadata.title}`;
const header = fromBase64(file.thumbnail.decryptionHeader); const header = fromBase64(file.thumbnail.decryptionHeader);
const bytesWritten = await fetchAndDecrypt( const plaintext = await fetchAndDecrypt(
api, api,
() => api.getThumbnailStream(file.id, { retry: false }), () => api.getThumbnailStream(file.id, { retry: false }),
header, header,
file.key, file.key,
resolvedPath,
onProgress,
); );
return { path: resolvedPath, bytesWritten }; await writeAtomic(resolvedPath, plaintext);
return { path: resolvedPath, bytesWritten: plaintext.length };
}; };
+1 -54
View File
@@ -1,12 +1,6 @@
export const VERSION = "0.0.0"; export const VERSION = "0.0.0";
export { export { Client, type LoginOptions, type ClientSnapshot } from "./client.js";
Client,
type LoginOptions,
type ClientSnapshot,
type CollectionsPage,
type FilesPage,
} from "./client.js";
export { export {
ApiClient, ApiClient,
ApiError, ApiError,
@@ -33,53 +27,6 @@ export {
requestEmailOTP, requestEmailOTP,
submitEmailOTP, submitEmailOTP,
} from "./auth/login.js"; } from "./auth/login.js";
export {
Library,
DEFAULT_REFRESH_INTERVAL_SECONDS,
Album,
Photo,
type LibraryClient,
type LibraryOptions,
type LibraryStatus,
type RefreshEvent,
type RefreshProgressCallback,
type AlbumsAPI,
type PhotosAPI,
type TimelineAPI,
type PhotoFilter,
type TimelineGroup,
type GroupBy,
type ContentSource,
type ContentResult,
type ContentEvent,
type ContentOptions,
type PhotoContent,
type ThumbnailsAPI,
type ThumbnailPriority,
type EnsureOptions,
type EnsureResult,
type EnsureEvent,
runBackup,
type BackupOptions,
type BackupResult,
type BackupError,
} from "./library/index.js";
export {
RequestPools,
BoundedPool,
DEFAULT_METADATA_CONCURRENCY,
DEFAULT_CONTENT_CONCURRENCY,
DEFAULT_THUMBNAIL_CONCURRENCY,
type RequestPoolsOptions,
type Priority,
type RunOptions,
} from "./library/pools.js";
export type {
AlbumRecord,
PhotoRecord,
LibrarySnapshot,
LibraryChange,
} from "./library/records.js";
export { decryptCollection, decryptFile } from "./model/index.js"; export { decryptCollection, decryptFile } from "./model/index.js";
export { downloadFile, downloadThumbnail } from "./download/index.js"; export { downloadFile, downloadThumbnail } from "./download/index.js";
export type { export type {
-676
View File
@@ -1,676 +0,0 @@
// The on-disk content and thumbnail cache keyed by fileID (issue #46).
//
// Layout under `cacheDirectory`: `originals/<fileID>.<ext>` and
// `thumbnails/<fileID>.<ext>`, flat directories at 0700 with files at 0600.
// Content appears only by the streaming atomic writer's rename (the download
// layer, #40), so a file that exists is whole — "present means complete". The
// directory listing taken at `open()` is the record of what is cached, and the
// orphan temp files a crashed write may have left are reaped there.
//
// A fetch goes through the shared request pools (#45): the content pool for
// originals, the thumbnail pool for thumbnails. The pool limits concurrency,
// orders on-demand work ahead of background, and dedups by key so a fileID
// requested twice while the first is still in flight downloads once.
//
// Integrity. The reused streaming decrypt is the enforced guarantee: every
// chunk is authenticated and the writer renames the file into place only once
// the stream ends on TAG_FINAL, so a truncated or corrupt fetch throws and
// nothing is stored. On top of that this module refuses to record a stored file
// that came out empty. The design also asks for a content-hash comparison
// against `FileMetadata.hash` (with a `fileSize` fallback); that is deferred —
// see the PR — because the exact hash construction cannot be confirmed against
// the repo's fixtures and `FileBlob.size` is the encrypted object size, not the
// decrypted length this layer has.
import { existsSync, statSync } from "node:fs";
import {
chmod,
mkdir,
readdir,
rm,
stat,
statfs,
utimes,
} from "node:fs/promises";
import { dirname, extname, join } from "node:path";
import type { ApiClient } from "../api/client.js";
import {
downloadFile,
downloadThumbnail,
type ProgressCallback,
} from "../download/index.js";
import type { EnteFile } from "../model/types.js";
import type { Priority, RequestPools } from "./pools.js";
const DIR_MODE = 0o700;
const FILE_MODE = 0o600;
const TEMP_PREFIX = ".quak-";
const TEMP_SUFFIX = ".tmp";
const GIB = 1024 * 1024 * 1024;
// Owner ruling (#36): bound the originals cache at 100 GiB, but back off when
// the volume has under 50 GiB free so the cache never crowds the disk.
export const DEFAULT_ORIGINALS_MAX_BYTES = 100 * GIB;
export const DEFAULT_FREE_BELOW_BYTES = 50 * GIB;
// Ente thumbnails are always JPEG, so the cache stores them with a fixed
// extension rather than deriving one from the (image or video) title.
const THUMBNAIL_EXT = ".jpg";
type Kind = "original" | "thumbnail";
// An original write in progress, with the IDs of the concurrent original writes
// it overlaps (recorded both ways as writes begin, cleared when the write ends).
interface OriginalWrite {
fileID: number;
overlaps: Set<number>;
}
// The priority a caller attaches to a thumbnail prefetch. The pool has two
// tiers, so this three-value surface collapses onto them: only a currently
// visible thumbnail preempts (on-demand); "ahead" prefetch and speculative
// "background" work both yield to it.
export type ThumbnailPriority = "visible" | "ahead" | "background";
const poolPriorityOf = (priority: ThumbnailPriority): Priority =>
priority === "visible" ? "on-demand" : "background";
export interface ContentResult {
path: string;
bytes: number;
}
// Progress for a single `original`/`thumbnail` call. A present file emits one
// `skipped` event and nothing else; a fetched file emits `downloading` as
// plaintext lands and a final `done`.
export type ContentEvent =
| { status: "skipped"; bytes: number }
| { status: "downloading"; bytesDone: number }
| { status: "done"; bytes: number };
export interface ContentOptions {
onProgress?: (event: ContentEvent) => void;
}
// The Photo-facing content surface (the read wrappers call these). The cache
// implements it; a library opened without a content source leaves it absent.
export interface PhotoContent {
original(fileID: number, opts?: ContentOptions): Promise<ContentResult>;
thumbnail(fileID: number, opts?: ContentOptions): Promise<ContentResult>;
}
export interface EnsureResult {
fileID: number;
path?: string;
error?: string;
}
export interface EnsureEvent {
fileID: number;
status: "skipped" | "done" | "failed" | "aborted";
path?: string;
error?: string;
}
export interface EnsureOptions {
fileIDs: number[];
priority: ThumbnailPriority;
signal?: AbortSignal;
onProgress?: (event: EnsureEvent) => void;
}
export interface ThumbnailsAPI {
ensure(args: EnsureOptions): Promise<EnsureResult[]>;
}
// The byte source the cache fetches through. The real implementation streams
// and decrypts to the destination via the download layer; tests inject a
// stand-in so the cache logic runs with no crypto and no network. Pool routing,
// dedup, present-checks and integrity live in the cache, not here.
export interface ContentSource {
original(args: {
file: EnteFile;
destination: string;
onProgress?: ProgressCallback;
}): Promise<{ bytesWritten: number }>;
thumbnail(args: {
file: EnteFile;
destination: string;
onProgress?: ProgressCallback;
}): Promise<{ bytesWritten: number }>;
}
// The production source: each fetch is the download layer's request +
// streaming decrypt + atomic write + retry as one unit.
export const makeDownloadContentSource = (api: ApiClient): ContentSource => ({
original: ({ file, destination, onProgress }) =>
downloadFile(api, file, destination, onProgress),
thumbnail: ({ file, destination, onProgress }) =>
downloadThumbnail(api, file, destination, onProgress),
});
export interface CachedPaths {
originalPath?: string;
thumbnailPath?: string;
}
// The slice of `fs.statfs` the eviction limit needs: `bavail` is the blocks
// available to an unprivileged writer and `bsize` their size, so
// `bavail * bsize` is the free byte count. Injectable so tests drive the
// adaptive limit without a real volume.
export interface StatFsResult {
bsize: number;
bavail: number;
}
export type StatFsFn = (path: string) => Promise<StatFsResult>;
const realStatFs: StatFsFn = async (path) => {
const s = await statfs(path);
return { bsize: s.bsize, bavail: s.bavail };
};
// The current usage and effective limit of the originals cache, in bytes.
// `limitBytes` is the adaptive ceiling last computed (see `originalsLimit`).
export interface OriginalsStatus {
usedBytes: number;
limitBytes?: number;
}
export interface ContentCacheOptions {
pools: RequestPools;
source: ContentSource;
cacheDirectory: string;
// The backup destination (issue-level `downloadDirectory`). An original
// already stored there by a backup counts as present, so the cache serves
// it rather than fetching a second copy.
downloadDirectory?: string;
// Resolve any membership of a file; every membership shares the underlying
// content key, so any one decrypts the same bytes.
getFile: (fileID: number) => EnteFile | undefined;
// Hard ceiling on `cacheDirectory/originals` (default 100 GiB) and the free
// space to protect on the volume (default 50 GiB). The effective limit is
// the lesser of the ceiling and what fits above the protected free space.
cacheOriginalsMaxBytes?: number;
freeBelowBytes?: number;
// Whether an original is pinned (favorites + latest week; the precache unit
// #48 supplies the set). Pinned originals are never evicted; when only
// pinned originals remain the cache runs over-limit until the set shrinks.
isPinned?: (fileID: number) => boolean;
// Free-space probe on the volume holding `cacheDirectory`; defaults to the
// real `fs.statfs`.
statfs?: StatFsFn;
}
// Thrown inside a pooled task to drop a queued fetch that was aborted before it
// started running. Never escapes `ensureThumbnails`.
class AbortDrop extends Error {
constructor() {
super("aborted");
this.name = "AbortDrop";
}
}
const originalName = (file: EnteFile): string => {
const ext = extname(file.metadata.title || "") || ".bin";
return `${file.id}${ext}`;
};
// The fileID a cache filename encodes, or undefined when the name is not one
// the cache writes (`<digits><ext>`).
const fileIDFromName = (name: string): number | undefined => {
const base = name.slice(0, name.length - extname(name).length);
if (!/^\d+$/.test(base)) return undefined;
const id = Number(base);
return Number.isSafeInteger(id) ? id : undefined;
};
// Size of a regular file, or undefined if it is absent (or not a regular file).
const fileSize = (path: string): number | undefined => {
try {
const s = statSync(path);
return s.isFile() ? s.size : undefined;
} catch {
return undefined;
}
};
export class ContentCache implements PhotoContent, ThumbnailsAPI {
private readonly pools: RequestPools;
private readonly source: ContentSource;
private readonly downloadDirectory?: string;
private readonly getFile: (fileID: number) => EnteFile | undefined;
private readonly originalsDir: string;
private readonly thumbnailsDir: string;
// fileID -> absolute path of the cached bytes, seeded from the directory
// listing at open() and extended as fetches store new files.
private readonly originals = new Map<number, string>();
private readonly thumbnails = new Map<number, string>();
private readonly maxOriginalsBytes: number;
private readonly freeBelowBytes: number;
private readonly isPinned: (fileID: number) => boolean;
private readonly statfs: StatFsFn;
// The last measured usage and effective limit, refreshed at open() and after
// every original write; exposed through `originalsStatus`.
private originalsUsedBytes = 0;
private originalsLimitBytes?: number;
// Serializes limit enforcement so concurrent original writes never race on
// the map or delete each other's just-freed room.
private enforcing: Promise<void> = Promise.resolve();
// Original writes in progress. Writes for different files run concurrently
// (the content pool), so an eviction pass must never delete a file whose
// fetch has not yet returned. Each entry records the IDs of the concurrent
// original writes it overlaps — noted both ways as writes begin — and a
// write's eviction pass spares them all. Bounded by the pool's concurrency,
// so eviction is never deferred beyond the active working set.
private readonly inFlightOriginals = new Set<OriginalWrite>();
constructor(opts: ContentCacheOptions) {
this.pools = opts.pools;
this.source = opts.source;
this.downloadDirectory = opts.downloadDirectory;
this.getFile = opts.getFile;
this.originalsDir = join(opts.cacheDirectory, "originals");
this.thumbnailsDir = join(opts.cacheDirectory, "thumbnails");
this.maxOriginalsBytes =
opts.cacheOriginalsMaxBytes ?? DEFAULT_ORIGINALS_MAX_BYTES;
this.freeBelowBytes = opts.freeBelowBytes ?? DEFAULT_FREE_BELOW_BYTES;
this.isPinned = opts.isPinned ?? (() => false);
this.statfs = opts.statfs ?? realStatFs;
}
// Prepare the cache directories, reap orphan temp files, and take the
// record of what is already cached. Called once before the cache serves.
async open(): Promise<void> {
await this.ensureDir(this.originalsDir);
await this.ensureDir(this.thumbnailsDir);
await this.scan(this.originalsDir, this.originals);
await this.scan(this.thumbnailsDir, this.thumbnails);
// Publish the current usage and limit without evicting; a restart
// reuses whatever survived on disk. Eviction only ever fires on a write.
await this.refreshOriginalsLimit();
}
// The current originals usage and effective limit, both in bytes, as of the
// last write or open. `status().originalsLimitBytes` surfaces this.
originalsStatus(): OriginalsStatus {
return {
usedBytes: this.originalsUsedBytes,
limitBytes: this.originalsLimitBytes,
};
}
// The cache paths known for a file, for the record projection to expose as
// `originalPath`/`thumbnailPath`.
pathsFor(fileID: number): CachedPaths {
const out: CachedPaths = {};
const original = this.originals.get(fileID);
if (original !== undefined) out.originalPath = original;
const thumbnail = this.thumbnails.get(fileID);
if (thumbnail !== undefined) out.thumbnailPath = thumbnail;
return out;
}
async original(
fileID: number,
opts?: ContentOptions,
): Promise<ContentResult> {
return this.get(fileID, "original", "on-demand", opts?.onProgress);
}
async thumbnail(
fileID: number,
opts?: ContentOptions,
): Promise<ContentResult> {
return this.get(fileID, "thumbnail", "on-demand", opts?.onProgress);
}
async ensure(args: EnsureOptions): Promise<EnsureResult[]> {
return this.ensureThumbnails(args);
}
async ensureThumbnails(args: EnsureOptions): Promise<EnsureResult[]> {
return this.ensureMany(
"thumbnail",
poolPriorityOf(args.priority),
args.fileIDs,
args.signal,
args.onProgress,
);
}
// Fill originals through the content pool for the precache (#48), always at
// background priority so an on-demand `original()` preempts the fill. A
// present original is a map lookup and no fetch; a per-file failure is
// returned, not thrown, so one bad file never halts a background sweep.
async ensureOriginals(args: {
fileIDs: number[];
signal?: AbortSignal;
onProgress?: (event: EnsureEvent) => void;
}): Promise<EnsureResult[]> {
return this.ensureMany(
"original",
"background",
args.fileIDs,
args.signal,
args.onProgress,
);
}
private async ensureMany(
kind: Kind,
priority: Priority,
fileIDs: number[],
signal: AbortSignal | undefined,
onProgress: ((event: EnsureEvent) => void) | undefined,
): Promise<EnsureResult[]> {
// Dedup the request list so a repeated fileID is fetched once and
// reported once, in first-requested order.
const seen = new Set<number>();
const unique: number[] = [];
for (const id of fileIDs) {
if (!seen.has(id)) {
seen.add(id);
unique.push(id);
}
}
return Promise.all(
unique.map((fileID) =>
this.ensureOne(fileID, kind, priority, signal, onProgress),
),
);
}
private async ensureOne(
fileID: number,
kind: Kind,
priority: Priority,
signal: AbortSignal | undefined,
onProgress: ((event: EnsureEvent) => void) | undefined,
): Promise<EnsureResult> {
try {
const result = await this.acquire(fileID, kind, priority, signal);
const status = result.cached ? "skipped" : "done";
onProgress?.({ fileID, status, path: result.path });
return { fileID, path: result.path };
} catch (err) {
if (err instanceof AbortDrop) {
onProgress?.({ fileID, status: "aborted" });
return { fileID, error: "aborted" };
}
const error = err instanceof Error ? err.message : String(err);
onProgress?.({ fileID, status: "failed", error });
return { fileID, error };
}
}
private async get(
fileID: number,
kind: Kind,
priority: Priority,
onProgress: ((event: ContentEvent) => void) | undefined,
): Promise<ContentResult> {
const onByte: ProgressCallback | undefined = onProgress
? (bytesDone) => onProgress({ status: "downloading", bytesDone })
: undefined;
const result = await this.acquire(fileID, kind, priority, undefined, {
onByte,
});
onProgress?.(
result.cached
? { status: "skipped", bytes: result.bytes }
: { status: "done", bytes: result.bytes },
);
return { path: result.path, bytes: result.bytes };
}
// The core: return the cached path if present, else fetch through the pool,
// store, and return it. `cached` distinguishes a present hit (no network,
// no download event) from a fresh fetch.
private async acquire(
fileID: number,
kind: Kind,
priority: Priority,
signal: AbortSignal | undefined,
opts?: { onByte?: ProgressCallback },
): Promise<{ path: string; bytes: number; cached: boolean }> {
const file = this.getFile(fileID);
if (!file) throw new Error(`content cache: unknown file ${fileID}`);
const known = kind === "original" ? this.originals : this.thumbnails;
const cached = known.get(fileID);
if (cached !== undefined) {
const size = fileSize(cached);
if (size !== undefined && size > 0) {
// Returning an original's path is a use: bump its mtime so LRU
// order reflects it and survives a restart with no ledger.
if (
kind === "original" &&
dirname(cached) === this.originalsDir
)
await this.touch(cached);
return { path: cached, bytes: size, cached: true };
}
// A recorded file that has since gone re-fetches below.
known.delete(fileID);
}
// An original a backup already stored counts as present.
if (kind === "original" && this.downloadDirectory !== undefined) {
const backupPath = join(
this.downloadDirectory,
"originals",
originalName(file),
);
const size = fileSize(backupPath);
if (size !== undefined && size > 0) {
this.originals.set(fileID, backupPath);
return { path: backupPath, bytes: size, cached: true };
}
}
const dir =
kind === "original" ? this.originalsDir : this.thumbnailsDir;
const dest =
kind === "original"
? join(dir, originalName(file))
: join(dir, `${fileID}${THUMBNAIL_EXT}`);
const pool =
kind === "original" ? this.pools.content : this.pools.thumbnails;
return pool.run(
async () => {
// Dropping queued work on abort: a task still waiting for a slot
// when the signal fired sees it here and never touches the
// network. A task already past this point is in flight and runs
// to completion.
if (signal?.aborted) throw new AbortDrop();
// Register this original among those in flight, linking it with
// every sibling already writing so neither evicts the other's
// file. Non-null iff this is an original.
const write =
kind === "original"
? this.beginOriginalWrite(fileID)
: null;
try {
await this.download(file, dest, kind, opts?.onByte);
await chmod(dest, FILE_MODE);
const size = (await stat(dest)).size;
if (size === 0) {
throw new Error(
`content cache: ${kind} ${fileID} stored empty`,
);
}
known.set(fileID, dest);
// A fresh original may have crossed the limit; make room by
// evicting least-recently-used originals. An over-budget
// fetch keeps the file it returns, and no overlapping
// sibling is evicted. Thumbnails are never bounded.
if (write) await this.enforceOriginalsLimit(write);
return { path: dest, bytes: size, cached: false };
} finally {
if (write) this.inFlightOriginals.delete(write);
}
},
{ priority, key: fileID },
);
}
private async download(
file: EnteFile,
destination: string,
kind: Kind,
onProgress: ProgressCallback | undefined,
): Promise<number> {
const args = { file, destination, onProgress };
const result =
kind === "original"
? await this.source.original(args)
: await this.source.thumbnail(args);
return result.bytesWritten;
}
// Best-effort bump of a file's mtime to now; a failed touch must never fail
// the read it accompanies.
private async touch(path: string): Promise<void> {
const now = new Date();
await utimes(path, now, now).catch(() => undefined);
}
// Every stored original that lives under `originalsDir` (a backup-directory
// hit recorded in the map is excluded), with its size and mtime. Entries
// whose file has vanished are dropped from the map. Backups and thumbnails
// are never counted.
private async measureOriginals(): Promise<{
entries: {
fileID: number;
path: string;
size: number;
mtimeMs: number;
}[];
used: number;
}> {
const entries: {
fileID: number;
path: string;
size: number;
mtimeMs: number;
}[] = [];
let used = 0;
for (const [fileID, path] of this.originals) {
if (dirname(path) !== this.originalsDir) continue;
try {
const s = await stat(path);
entries.push({
fileID,
path,
size: s.size,
mtimeMs: s.mtimeMs,
});
used += s.size;
} catch {
this.originals.delete(fileID);
}
}
return { entries, used };
}
// The effective ceiling on originals: the configured max, but no more than
// what fits once the protected free space is set aside. `used + free` is the
// volume space the cache could occupy; subtracting `freeBelowBytes` leaves
// the reserve untouched. Clamped at zero.
private async originalsLimit(used: number): Promise<number> {
const { bsize, bavail } = await this.statfs(this.originalsDir);
const free = bsize * bavail;
const adaptive = used + free - this.freeBelowBytes;
return Math.max(0, Math.min(this.maxOriginalsBytes, adaptive));
}
// Recompute and publish usage and limit without evicting (used at open()).
private refreshOriginalsLimit(): Promise<void> {
return this.serializeEnforce(async () => {
const { used } = await this.measureOriginals();
this.originalsUsedBytes = used;
this.originalsLimitBytes = await this.originalsLimit(used);
});
}
// Record a starting original write among those in flight, linking it with
// every sibling already writing so neither can evict the other's file.
private beginOriginalWrite(fileID: number): OriginalWrite {
const write: OriginalWrite = { fileID, overlaps: new Set() };
for (const other of this.inFlightOriginals) {
write.overlaps.add(other.fileID);
other.overlaps.add(fileID);
}
this.inFlightOriginals.add(write);
return write;
}
// Evict least-recently-used originals until usage fits the limit. Skipped:
// pinned originals, the file `write` just stored, and every original whose
// write overlaps it (`write.overlaps`). The last two spare any fetch whose
// lifetime overlaps this one, so concurrent over-budget fetches all keep the
// paths they return; when only such originals remain the cache stays
// over-limit until they settle, and a later, non-overlapping write finds
// them eligible again.
private enforceOriginalsLimit(write: OriginalWrite): Promise<void> {
return this.serializeEnforce(async () => {
const { entries, used } = await this.measureOriginals();
const limit = await this.originalsLimit(used);
let remaining = used;
if (remaining > limit) {
const evictable = entries
.filter(
(e) =>
e.fileID !== write.fileID &&
!write.overlaps.has(e.fileID) &&
!this.isPinned(e.fileID),
)
.sort((a, b) => a.mtimeMs - b.mtimeMs);
for (const e of evictable) {
if (remaining <= limit) break;
await rm(e.path, { force: true });
this.originals.delete(e.fileID);
remaining -= e.size;
}
}
this.originalsUsedBytes = remaining;
this.originalsLimitBytes = limit;
});
}
// Run limit work one at a time; failures are swallowed so a transient
// statfs or unlink error never rejects the read or write that triggered it.
private serializeEnforce(work: () => Promise<void>): Promise<void> {
const next = this.enforcing.then(work).catch(() => undefined);
this.enforcing = next;
return next;
}
private async ensureDir(dir: string): Promise<void> {
// chmod after mkdir so the mode is tightened even when the directory
// already existed with a looser one; mkdir alone would not.
await mkdir(dir, { recursive: true, mode: DIR_MODE });
await chmod(dir, DIR_MODE);
}
private async scan(dir: string, into: Map<number, string>): Promise<void> {
let entries: string[];
try {
entries = await readdir(dir);
} catch {
return;
}
for (const name of entries) {
if (name.startsWith(TEMP_PREFIX) && name.endsWith(TEMP_SUFFIX)) {
await rm(join(dir, name), { force: true }).catch(
() => undefined,
);
continue;
}
const id = fileIDFromName(name);
const path = join(dir, name);
if (id !== undefined && existsSync(path)) into.set(id, path);
}
}
}
-844
View File
@@ -1,844 +0,0 @@
// The library surface over the local cache.
//
// `Library.open()` loads the on-disk metadata store (issue #41), then starts
// the refresh loop. When the cache loaded empty it awaits the first refresh,
// so the library never opens onto an empty store it could have filled; when an
// existing copy loaded, that first refresh runs in the background and `open()`
// returns as soon as the cached data is ready to serve — a slow or unreachable
// server no longer stalls opening. A background timer then refreshes every
// `refreshIntervalSeconds`. Every default read is answered from RAM — no
// default read touches the network. There is deliberately no `sync()`, no
// `refresh()`, no `serverReachable` flag, and no "before each read" mode
// (design #36).
//
// `fresh()` is the one exception (issue #75, an owner amendment to #36): it
// forces a refresh, awaits it, and only then hands back the read namespaces, so
// a caller that needs server-current data can ask for it. Concurrent `fresh()`
// calls coalesce onto one in-flight refresh, and a refresh that fails rejects
// the caller (the default reads stay silent and serve the last good copy). The
// default methods and the background loop are unchanged.
//
// A refresh stages all of its network work first and only mutates the store
// once every fetch has succeeded. A refresh that fails partway therefore never
// becomes visible to reads: the last good snapshot stays in place, and the
// failure surfaces through `onProgress` and `status()` instead. A commit that
// mutates RAM but then fails to persist keeps `status().lastError` set and the
// store marked unsaved until a later save actually lands, so a stuck disk is
// never masked by a subsequent empty refresh.
import { join } from "node:path";
import envPaths from "env-paths";
import { MetadataStore } from "./store.js";
import { MLDataStore } from "./mldata.js";
import { RequestPools } from "./pools.js";
import {
deriveRecords,
snapshotFrom,
diffRecords,
type DerivedRecords,
type LibrarySnapshot,
type LibraryChange,
} from "./records.js";
import {
makeAlbumsAPI,
makePhotosAPI,
makeTimelineAPI,
type AlbumsAPI,
type PhotosAPI,
type TimelineAPI,
type FreshReads,
} from "./read.js";
import {
ContentCache,
type ContentSource,
type ThumbnailsAPI,
type EnsureOptions,
type EnsureResult,
} from "./content.js";
import { makeMLDataAPI, type MLDataAPI } from "./mlsearch.js";
import { Precache } from "./precache.js";
export {
Album,
Photo,
type AlbumsAPI,
type PhotosAPI,
type TimelineAPI,
type FreshReads,
type PhotoFilter,
type TimelineGroup,
type GroupBy,
} from "./read.js";
export {
type ContentSource,
type ContentResult,
type ContentEvent,
type ContentOptions,
type PhotoContent,
type ThumbnailsAPI,
type ThumbnailPriority,
type EnsureOptions,
type EnsureResult,
type EnsureEvent,
} from "./content.js";
export { type MLDataAPI, type SimilarResult } from "./mlsearch.js";
import type { CollectionsPage, FilesPage } from "../client.js";
import { MLDATA_BATCH_SIZE, type MLData } from "../mldata-fetch.js";
import type { Collection, EnteFile } from "../model/types.js";
import { runBackup, type BackupOptions, type BackupResult } from "../backup.js";
export {
runBackup,
type BackupOptions,
type BackupResult,
type BackupError,
} from "../backup.js";
export const DEFAULT_REFRESH_INTERVAL_SECONDS = 3;
// Project a metadata store into by-id records, filling each record's cache
// paths from the content cache when one is given. Shared by the live read
// projection and the precache's initial seeding at open().
const deriveRecordsFromStore = (
store: MetadataStore,
cache?: ContentCache,
): DerivedRecords => {
const collections = store.listCollections();
const files: EnteFile[] = [];
for (const c of collections) files.push(...store.listFiles(c.id));
return deriveRecords(
collections,
files,
cache ? (fileID) => cache.pathsFor(fileID) : undefined,
);
};
// The slice of `Client` the library depends on. Narrowing to an interface lets
// tests drive a mock with no crypto or network; the real `Client` satisfies it
// structurally.
export interface LibraryClient {
whoami(): { email: string; userID: number };
collectionsSince(args: { sinceTime: number }): Promise<CollectionsPage>;
filesSince(args: {
collectionID: number;
collectionKey: Uint8Array;
sinceTime: number;
}): Promise<FilesPage>;
// Fetch ML data (face detections + CLIP embeddings) for up to a batch of
// files. Optional: a client without it simply disables ML fetching, leaving
// the metadata refresh untouched.
fetchMLData?(args: {
fileIDs: number[];
fileKeys: Map<number, Uint8Array>;
}): Promise<Map<number, MLData>>;
// The byte source for the on-disk content cache. Optional so a mock client
// that only serves metadata still satisfies the interface; when absent (and
// no explicit `contentSource` is passed to `open`) the content cache is
// disabled and `Photo.original`/`thumbnail` and `thumbnails.ensure` throw.
contentSource?(): ContentSource;
}
// A progress event for one unit of background work. A metadata "refresh" or an
// ML "fetchMLData" pass each fire "started" before their network work and then
// exactly one of "done" or "failed"; "failed" carries the error message and
// an ML "done" reports how many payloads it stored. The precache fills
// ("precacheThumbnails"/"precacheOriginals", #48) fire "started"/"done" around
// each sweep that has work, "done" reporting the count newly cached.
export interface RefreshEvent {
operation:
| "refresh"
| "fetchMLData"
| "precacheThumbnails"
| "precacheOriginals";
status: "started" | "done" | "failed";
error?: string;
fetched?: number;
}
export type RefreshProgressCallback = (event: RefreshEvent) => void;
export interface LibraryOptions {
client: LibraryClient;
// Where `metadata.json` lives. Defaults to the env-paths cache directory
// plus the user id, so each account has its own cache.
cacheDirectory?: string;
// Persistent backup destination. The refresh loop does not use it; the
// content cache treats an original already stored there as present.
downloadDirectory?: string;
refreshIntervalSeconds?: number;
onProgress?: RefreshProgressCallback;
// The bounded request pools (issue #45), shared by the ML-data fetch (the
// metadata pool) and the content cache. Defaults to a fresh set at the
// design's caps.
pools?: RequestPools;
// Overrides the client's own `contentSource()`; mainly for tests that drive
// the cache with a stand-in source.
contentSource?: ContentSource;
// Bound on `cacheDirectory/originals` (default 100 GiB) and the free space
// to protect on its volume (default 50 GiB). The effective limit adapts
// down as the disk fills; `status().originalsLimitBytes` reports it.
cacheOriginalsMaxBytes?: number;
freeBelowBytes?: number;
// An extra pinned predicate OR-ed with the precache's own pinned set
// (favorites + latest week, #48). Pinned originals are never evicted.
isOriginalPinned?: (fileID: number) => boolean;
// The aggressive local precache (#48), all starting inside `open()` with no
// caller input. Thumbnails: every file, newest first, until all are on
// disk. Originals: the favorites album then the latest `precacheOriginalsDays`
// window (the days ending at the newest file). Both default on; the days
// default to 7.
precacheThumbnails?: boolean;
precacheOriginals?: boolean;
precacheOriginalsDays?: number;
}
export interface LibraryStatus {
userID: number;
collections: number;
files: number;
// Wall-clock ms of the last refresh that succeeded, or undefined if none
// has yet.
lastRefreshAt?: number;
// The message from the most recent refresh, set only while that refresh
// failed; cleared by the next success.
lastError?: string;
// Wall-clock ms of the last ML fetch pass that succeeded, or undefined if
// none has yet (or ML fetching is disabled).
lastMLFetchAt?: number;
// The most recent ML fetch pass's error, set only while it failed.
lastMLError?: string;
// ML payloads stored on disk and CLIP embeddings in the index; undefined
// when ML fetching is disabled.
mlStored?: number;
mlIndexed?: number;
// Bytes stored in the originals cache and the effective size limit as of the
// last write or open; undefined when no content cache is open.
originalsUsedBytes?: number;
originalsLimitBytes?: number;
// Precache progress (#48); undefined when no content cache is open. Totals
// are the files targeted (0 when a fill is disabled); "cached" is how many
// of them are on disk.
thumbnailsCached?: number;
thumbnailsTotal?: number;
originalsCached?: number;
originalsPinned?: number;
closed: boolean;
}
export class Library {
readonly cacheDirectory: string;
readonly downloadDirectory?: string;
// The in-process read surface (issue #44). Each namespace answers
// synchronously from the live record projection; no read touches the
// network.
readonly albums: AlbumsAPI;
readonly photos: PhotosAPI;
readonly timeline: TimelineAPI;
// The thumbnail-prefetch surface (issue #46): drives the thumbnail pool
// with priority, dedup, and abort.
readonly thumbnails: ThumbnailsAPI;
// The content-similarity search surface over the CLIP index (issue #50).
// Present whether or not ML fetching is enabled; with no ML store it
// returns empty results.
readonly mldata: MLDataAPI;
private readonly client: LibraryClient;
private readonly store: MetadataStore;
// The on-disk content cache, or undefined when no content source is
// available (a metadata-only client with no explicit source).
private readonly cache?: ContentCache;
private readonly userID: number;
private readonly intervalMs: number;
private readonly onProgress?: RefreshProgressCallback;
private readonly pools: RequestPools;
// The ML-data cache, present only when the client can fetch ML data.
private readonly mlStore?: MLDataStore;
// The local precache (#48), present only when the content cache is.
private readonly precache?: Precache;
private timer?: ReturnType<typeof setTimeout>;
// The in-flight refresh cycle, or undefined when none runs. One slot serves
// both paths: the background loop skips when it is set, and a fresh read
// (issue #75) coalesces onto it or starts one. The promise carries the
// cycle's real outcome (it rejects on failure); the background loop ignores
// that, a fresh read propagates it.
private cycle?: Promise<void>;
// Guards the ML fetch pass so a slow backfill never runs twice at once; a
// refresh whose pass is still running kicks nothing new.
private mlFetching = false;
private closed = false;
private lastRefreshAt?: number;
private lastError?: string;
private lastMLFetchAt?: number;
private lastMLError?: string;
// The plain-record projection as of the last refresh, and the GUI change
// subscribers. A refresh that alters the projection notifies each with the
// delta; `lastRecords` is kept current every refresh so a subscriber that
// joins later diffs against the state its own `snapshot()` already returned.
private readonly subscribers = new Set<(change: LibraryChange) => void>();
private lastRecords: DerivedRecords;
// RAM holds changes disk has not yet accepted (an earlier save failed).
// Cleared only when a save actually succeeds; keeps the store trying to
// persist and the failure visible in `status()` until then.
private unsaved = false;
private constructor(args: {
client: LibraryClient;
store: MetadataStore;
userID: number;
cacheDirectory: string;
downloadDirectory?: string;
intervalMs: number;
onProgress?: RefreshProgressCallback;
pools: RequestPools;
mldata?: MLDataStore;
cache?: ContentCache;
precache?: Precache;
}) {
this.client = args.client;
this.store = args.store;
this.userID = args.userID;
this.cacheDirectory = args.cacheDirectory;
this.downloadDirectory = args.downloadDirectory;
this.intervalMs = args.intervalMs;
this.onProgress = args.onProgress;
this.pools = args.pools;
this.mlStore = args.mldata;
this.cache = args.cache;
this.precache = args.precache;
this.lastRecords = this.deriveNow();
// The read namespaces derive fresh from the store on each call, so they
// always reflect the latest refresh.
const derive = (): DerivedRecords => this.deriveNow();
this.albums = makeAlbumsAPI(derive, this.cache);
this.photos = makePhotosAPI(derive, this.cache);
this.timeline = makeTimelineAPI(derive);
this.thumbnails = {
ensure: (opts: EnsureOptions): Promise<EnsureResult[]> => {
if (!this.cache) {
return Promise.reject(
new Error(
"thumbnails.ensure requires a library opened with a content cache",
),
);
}
return this.cache.ensureThumbnails(opts);
},
};
// Reads the ML store live so results grow as ML data is fetched.
this.mldata = makeMLDataAPI(() => this.mlStore);
}
// Load the cache and start the refresh loop. With an empty cache the first
// refresh is awaited, so `open()` resolves onto populated data whenever the
// server is reachable; that awaited refresh may still fail, and the library
// then opens empty with the failure recorded in `status()`. With an
// existing cache the first refresh runs in the background and `open()`
// returns as soon as the cached data is ready — an unreachable server does
// not block opening.
static async open(opts: LibraryOptions): Promise<Library> {
const { userID } = opts.client.whoami();
const cacheDirectory =
opts.cacheDirectory ??
join(envPaths("quak", { suffix: "" }).cache, String(userID));
const store = await MetadataStore.load(
join(cacheDirectory, "metadata.json"),
);
const intervalMs =
(opts.refreshIntervalSeconds ?? DEFAULT_REFRESH_INTERVAL_SECONDS) *
1000;
// One request-pool set serves both the ML-data fetch and the content
// cache, so both honour the same concurrency caps.
const pools = opts.pools ?? new RequestPools();
// The ML cache only earns its keep when the client can fetch ML data;
// a client without that capability opens no `mldata/` directory.
const mldata = opts.client.fetchMLData
? await MLDataStore.open(join(cacheDirectory, "mldata"))
: undefined;
// Build the content cache from an explicit source or the client's own,
// and take its record of what is already cached (and reap orphan temp
// files) before the first projection, so cached paths are present from
// the start and the first refresh raises no spurious path-change diff.
const source = opts.contentSource ?? opts.client.contentSource?.();
let cache: ContentCache | undefined;
let precache: Precache | undefined;
if (source) {
// The precache owns the pinned set (favorites + latest week), which
// is the cache's eviction predicate. It is built and seeded from
// the loaded store first so the cache can wire `isPinned` to it, and
// then bound to the cache it fills. A caller-supplied predicate is
// OR-ed in so both survive.
precache = new Precache({
thumbnails: opts.precacheThumbnails,
originals: opts.precacheOriginals,
originalsDays: opts.precacheOriginalsDays,
onEvent: opts.onProgress,
});
precache.update(deriveRecordsFromStore(store));
const extraPinned = opts.isOriginalPinned;
cache = new ContentCache({
pools,
source,
cacheDirectory,
downloadDirectory: opts.downloadDirectory,
getFile: (fileID) => store.getFileByID(fileID),
cacheOriginalsMaxBytes: opts.cacheOriginalsMaxBytes,
freeBelowBytes: opts.freeBelowBytes,
isPinned: (fileID) =>
precache!.isPinned(fileID) ||
(extraPinned?.(fileID) ?? false),
});
await cache.open();
precache.bind(cache);
}
const lib = new Library({
client: opts.client,
store,
userID,
cacheDirectory,
downloadDirectory: opts.downloadDirectory,
intervalMs,
onProgress: opts.onProgress,
pools,
mldata,
cache,
precache,
});
// Start filling from whatever the loaded store already holds; each
// refresh below re-kicks with the new files (and retries any that
// failed). An empty store starts empty here and fills after its first
// refresh.
precache?.start();
if (store.loadedFromDisk) {
// An existing copy already answers reads; refresh in the background
// and start the interval once that first cycle settles.
void lib.runRefresh().then(() => lib.scheduleNext());
} else {
// Nothing was cached: wait for the first refresh to fill the store
// (or fail) rather than resolve onto an empty library.
await lib.runRefresh();
lib.scheduleNext();
}
return lib;
}
listCollections(): Collection[] {
return this.store.listCollections();
}
getCollection(id: number): Collection | undefined {
return this.store.getCollection(id);
}
listFiles(collectionID: number): EnteFile[] {
return this.store.listFiles(collectionID);
}
getFile(collectionID: number, fileID: number): EnteFile | undefined {
return this.store.getFile(collectionID, fileID);
}
// Any membership of a file, addressed by file id alone. A file's own
// metadata (title, creationTime) is identical across the collections it
// belongs to, so this serves the point commands that hold only a fileID.
getFileByID(fileID: number): EnteFile | undefined {
return this.store.getFileByID(fileID);
}
// A synchronous, RAM-only projection of the whole library into plain
// records (no keys), the surface the GUI reads across IPC. Photos are
// deduplicated to one record per file and ordered newest first.
snapshot(): LibrarySnapshot {
return snapshotFrom(this.deriveNow(), Date.now());
}
// Deliver a `LibraryChange` whenever a refresh alters the projection. A
// refresh that changes nothing delivers nothing. The returned handle's
// `unsubscribe` stops delivery.
subscribe(args: { onChange: (change: LibraryChange) => void }): {
unsubscribe: () => void;
} {
const { onChange } = args;
this.subscribers.add(onChange);
return {
unsubscribe: () => {
this.subscribers.delete(onChange);
},
};
}
status(): LibraryStatus {
let files = 0;
const collections = this.store.listCollections();
for (const c of collections) {
files += this.store.listFiles(c.id).length;
}
const ml = this.mlStore?.stats();
const originals = this.cache?.originalsStatus();
const pre = this.precache?.status();
return {
userID: this.store.userID,
collections: collections.length,
files,
lastRefreshAt: this.lastRefreshAt,
lastError: this.lastError,
lastMLFetchAt: this.lastMLFetchAt,
lastMLError: this.lastMLError,
mlStored: ml?.stored,
mlIndexed: ml?.indexed,
originalsUsedBytes: originals?.usedBytes,
originalsLimitBytes: originals?.limitBytes,
thumbnailsCached: pre?.thumbnailsCached,
thumbnailsTotal: pre?.thumbnailsTotal,
originalsCached: pre?.originalsCached,
originalsPinned: pre?.originalsPinned,
closed: this.closed,
};
}
// Fresh reads (issue #75, owner amendment to design #36). Force a refresh,
// wait for it to complete and persist, then hand back the same
// `albums`/`photos`/`timeline` namespaces — now guaranteed to reflect a
// completed server round-trip. Concurrent calls coalesce onto one refresh;
// a refresh that fails rejects here, where the default namespaces would
// instead stay silent and serve the last good copy.
async fresh(): Promise<FreshReads> {
await this.refreshNow();
return {
albums: this.albums,
photos: this.photos,
timeline: this.timeline,
};
}
// Back up every in-scope file to `downloadDirectory` in the historical
// on-disk layout, with a durable failure ledger (issue #51). Refreshes
// first, fetches pending originals (and optional thumbnails) through the
// content cache and pools, then rebuilds the derived symlink/JSON views
// from the model. Throws before any network work when no download directory
// is available or no content cache backs the originals it must fetch.
backup(opts?: BackupOptions): Promise<BackupResult> {
const downloadDirectory =
opts?.downloadDirectory ?? this.downloadDirectory;
const includeOriginals = opts?.includeOriginals ?? true;
const includeThumbnails = opts?.includeThumbnails ?? false;
if (!downloadDirectory) {
return Promise.reject(
new Error(
"backup requires a downloadDirectory (pass one to " +
"backup() or open the library with one)",
),
);
}
if ((includeOriginals || includeThumbnails) && !this.cache) {
return Promise.reject(
new Error(
"backup requires a library opened with a content cache",
),
);
}
const cache = this.cache;
return runBackup(
{
refresh: () => this.runRefresh(),
listCollections: () => this.store.listCollections(),
listFiles: (id) => this.store.listFiles(id),
original: (fileID) => cache!.original(fileID),
thumbnail: (fileID) => cache!.thumbnail(fileID),
},
{ ...opts, downloadDirectory },
);
}
// Stop the background timer. Idempotent. An in-flight refresh is left to
// finish; it will not schedule another cycle once closed.
close(): void {
this.closed = true;
this.precache?.close();
if (this.timer !== undefined) {
clearTimeout(this.timer);
this.timer = undefined;
}
}
private scheduleNext(): void {
if (this.closed) return;
this.timer = setTimeout(() => {
void this.runRefresh().then(() => this.scheduleNext());
}, this.intervalMs);
// Do not keep the process alive for the sake of the timer.
this.timer.unref?.();
}
// The background loop's refresh: run a cycle unless one is already in flight
// (or the library is closed), and never let a failure escape — the
// background path reports errors through `status()`/`onProgress`, it does
// not throw. Resolves once the cycle it started (or skipped past) settles.
private runRefresh(): Promise<void> {
if (this.closed || this.cycle) return Promise.resolve();
return this.startCycle().catch(() => {});
}
// A fresh read's refresh (issue #75): force a cycle and await it, rejecting
// if it fails. Concurrent fresh reads coalesce onto the one in-flight cycle
// — the background loop's included — so they never fan out into redundant
// server round-trips.
private refreshNow(): Promise<void> {
if (this.closed) {
return Promise.reject(new Error("the library is closed"));
}
return this.cycle ?? this.startCycle();
}
// Start one refresh cycle and record it as the in-flight cycle so every
// caller coalesces onto it. The returned promise carries the cycle's real
// outcome; each caller attaches the handling its own path needs, and the
// slot is cleared once the cycle settles.
private startCycle(): Promise<void> {
const cycle = this.refreshCycle();
this.cycle = cycle;
void cycle.then(
() => {
if (this.cycle === cycle) this.cycle = undefined;
},
() => {
if (this.cycle === cycle) this.cycle = undefined;
},
);
return cycle;
}
// One refresh cycle: the network fetch and commit, wrapped in the progress
// events and status bookkeeping. Throws when the refresh fails so a fresh
// read can reject; `runRefresh` swallows that throw for the background loop.
private async refreshCycle(): Promise<void> {
this.emit({ operation: "refresh", status: "started" });
try {
await this.refreshOnce();
this.lastRefreshAt = Date.now();
this.lastError = undefined;
this.emit({ operation: "refresh", status: "done" });
// Backfill ML data for the files this refresh knows about. It runs
// outside the refresh's success/failure so a fetch or disk problem
// there never marks the metadata refresh failed, and it is not
// awaited so it never stalls the refresh interval.
void this.runMLFetch();
} catch (err) {
const error = err instanceof Error ? err.message : String(err);
this.lastError = error;
this.emit({ operation: "refresh", status: "failed", error });
throw err;
}
}
// Fetch every change since the stored cursor, then commit. All network
// reads happen before any store mutation, so a fetch that throws leaves the
// store untouched and the previous snapshot intact.
private async refreshOnce(): Promise<void> {
const page = await this.client.collectionsSince({
sinceTime: this.store.collectionsSinceTime,
});
// Stage per-collection file diffs. A collection's files are
// re-enumerated only when its updationTime has advanced past the cached
// copy; an unchanged album's file list cannot have changed. New
// collections enumerate from the beginning of time.
const filePages: { collectionID: number; page: FilesPage }[] = [];
for (const collection of page.collections) {
const known = this.store.getCollection(collection.id);
if (known && collection.updationTime <= known.updationTime)
continue;
const filePage = await this.client.filesSince({
collectionID: collection.id,
collectionKey: collection.key,
sinceTime: known ? known.updationTime : 0,
});
filePages.push({ collectionID: collection.id, page: filePage });
}
// Network work done; commit to the store and persist only if something
// actually changed.
let changed = false;
if (this.store.userID !== this.userID) {
this.store.userID = this.userID;
changed = true;
}
for (const id of page.deleted) {
if (this.store.getCollection(id)) {
this.store.deleteCollection(id);
changed = true;
}
}
for (const collection of page.collections) {
this.store.putCollection(collection);
changed = true;
}
for (const { collectionID, page: filePage } of filePages) {
for (const id of filePage.deleted) {
if (this.store.getFile(collectionID, id)) {
this.store.deleteFile(collectionID, id);
changed = true;
}
}
for (const f of filePage.files) {
this.store.putFile(f);
changed = true;
}
}
if (page.cursor !== this.store.collectionsSinceTime) {
this.store.collectionsSinceTime = page.cursor;
changed = true;
}
if (changed) this.unsaved = true;
// Reproject and notify subscribers of the delta. This tracks RAM (what
// reads see), so it fires whether or not the save below succeeds; a
// save failure surfaces separately through `status().lastError`.
// `lastRecords` advances every changed refresh so the next diff is
// against current state.
if (changed) {
const next = this.deriveNow();
if (this.subscribers.size > 0) {
const change = diffRecords(this.lastRecords, next, Date.now());
if (change) this.notify(change);
}
this.lastRecords = next;
// Recompute the fill orders and pinned set against the new library.
this.precache?.update(next);
}
// Re-kick the fills every cycle: a finished sweep starts afresh to pick
// up new files and retry any that failed, and a running one is left be.
this.precache?.start();
// Persist whenever RAM holds changes disk has not accepted — including
// changes an earlier cycle staged whose save failed. `unsaved` clears
// only once a save lands, so a save failure both stays visible through
// `status().lastError` (the throw below records it) and keeps being
// retried, instead of a later empty refresh silently clearing it while
// the on-disk cache is still behind RAM.
if (this.unsaved) {
await this.store.save();
this.unsaved = false;
}
}
// One ML fetch pass: fetch, decrypt and store the ML data for every file
// the store knows about that is not cached (or whose `updationTime` has
// advanced), through the metadata pool, and update the CLIP index. Guarded
// so passes never overlap; a failure is reported, not thrown.
private async runMLFetch(): Promise<void> {
const mldata = this.mlStore;
// Bind so the call keeps the client as its receiver when invoked
// through the pool below.
const fetchMLData = this.client.fetchMLData?.bind(this.client);
if (!mldata || !fetchMLData || this.closed || this.mlFetching) return;
const files = this.uniqueFiles();
const needed = mldata.neededFor(files);
if (needed.length === 0) return;
this.mlFetching = true;
this.emit({ operation: "fetchMLData", status: "started" });
try {
const fileKeys = new Map<number, Uint8Array>();
const updation = new Map<number, number>();
for (const f of files) {
fileKeys.set(f.id, f.key);
updation.set(f.id, f.updationTime);
}
let stored = 0;
for (let i = 0; i < needed.length; i += MLDATA_BATCH_SIZE) {
if (this.closed) break;
const batch = needed.slice(i, i + MLDATA_BATCH_SIZE);
const payloads = await this.pools.metadata.run(
() => fetchMLData({ fileIDs: batch, fileKeys }),
{ priority: "background" },
);
stored += (await mldata.storeFetched(payloads, updation))
.stored;
}
this.lastMLFetchAt = Date.now();
this.lastMLError = undefined;
this.emit({
operation: "fetchMLData",
status: "done",
fetched: stored,
});
} catch (err) {
const error = err instanceof Error ? err.message : String(err);
this.lastMLError = error;
this.emit({ operation: "fetchMLData", status: "failed", error });
} finally {
this.mlFetching = false;
}
}
// The distinct files the store holds, one entry per fileID (a file in
// several collections shares its ML data), each carrying the key and the
// newest `updationTime` seen across its memberships.
private uniqueFiles(): {
id: number;
key: Uint8Array;
updationTime: number;
}[] {
const byID = new Map<
number,
{ id: number; key: Uint8Array; updationTime: number }
>();
for (const collection of this.store.listCollections()) {
for (const f of this.store.listFiles(collection.id)) {
const seen = byID.get(f.id);
if (seen === undefined || f.updationTime > seen.updationTime)
byID.set(f.id, {
id: f.id,
key: f.key,
updationTime: f.updationTime,
});
}
}
return [...byID.values()];
}
// Gather every file membership and project the store into by-id records,
// filling each record's cache paths from the content cache when present.
private deriveNow(): DerivedRecords {
return deriveRecordsFromStore(this.store, this.cache);
}
private notify(change: LibraryChange): void {
for (const onChange of this.subscribers) {
// A misbehaving subscriber must not break the loop or its peers.
try {
onChange(change);
} catch {
// ignore
}
}
}
private emit(event: RefreshEvent): void {
if (!this.onProgress) return;
// A misbehaving callback must not break the refresh loop.
try {
this.onProgress(event);
} catch {
// ignore
}
}
}
-378
View File
@@ -1,378 +0,0 @@
// The on-disk cache of Ente's per-file machine-learning data and the CLIP
// index derived from it (issue #49).
//
// Under `<cacheDirectory>/mldata/` this keeps:
//
// - `<fileID>.json` — one decrypted, gunzipped payload per file, written by
// rename. Its presence means it is complete: a torn write never leaves a
// half-file, so the set of these files is the source of truth for what is
// cached. The full payload (face boxes, landmarks, embeddings) is read back
// from here on demand and never held in RAM.
//
// - `clip.f32` + `clip.json` — the derived index the content search runs on.
// `clip.json` lists the indexed fileIDs in order plus the embedding length;
// `clip.f32` is those CLIP embeddings packed as one `Float32Array`, so the
// index loads in a single read with no per-vector parse. The index is
// rebuilt from the payloads whenever it is missing or structurally
// disagrees with the files present, and appended to as new payloads arrive.
//
// - `fetched.json` — a small map of fileID to the `updationTime` it was
// fetched at. This is best-effort bookkeeping for refetch decisions (a file
// whose `updationTime` later advances is refetched); the payloads, not this
// file, remain the record of what is cached, so losing it only forgoes
// update-driven refetch until the next fetch rewrites it.
//
// In RAM this holds only the id list and the packed `Float32Array`.
import { mkdir, readFile, readdir } from "node:fs/promises";
import { join } from "node:path";
import { writeAtomic } from "../download/index.js";
import type { MLData } from "../mldata-fetch.js";
const CLIP_VECTORS = "clip.f32";
const CLIP_INDEX = "clip.json";
const FETCHED = "fetched.json";
// A payload file is named for its fileID alone; the derived files above are
// not, so this pattern picks out payloads and nothing else.
const PAYLOAD_RE = /^(\d+)\.json$/;
const BYTES_PER_FLOAT = 4;
// The on-disk form of `clip.json`.
interface ClipIndexFile {
fileIDs: number[];
embeddingLength: number;
}
// A file the model knows about, for deciding what to fetch.
export interface MLDataFile {
id: number;
updationTime: number;
}
// The RAM index the search reads: `fileIDs[i]` owns the `embeddingLength`
// floats of `embeddings` starting at `i * embeddingLength`.
export interface MLIndex {
fileIDs: number[];
embeddingLength: number;
embeddings: Float32Array;
}
// Pull the CLIP embedding out of a payload, or undefined when it is absent or
// misshapen. Kept strict so a bad payload is skipped rather than corrupting the
// packed index.
const clipEmbedding = (payload: MLData): number[] | undefined => {
const clip = payload.clip;
if (typeof clip !== "object" || clip === null) return undefined;
const embedding = (clip as { embedding?: unknown }).embedding;
if (!Array.isArray(embedding)) return undefined;
if (embedding.some((v) => typeof v !== "number" || !Number.isFinite(v)))
return undefined;
return embedding as number[];
};
export class MLDataStore {
readonly dir: string;
// fileIDs whose payload JSON is present on disk (present means complete).
private readonly present = new Set<number>();
// fileID -> updationTime it was fetched at.
private readonly fetched = new Map<number, number>();
// The packed index and where each id sits in it.
private ids: number[] = [];
private embeddingLength = 0;
private embeddings = new Float32Array(0);
private readonly pos = new Map<number, number>();
private constructor(dir: string) {
this.dir = dir;
}
// Open (creating the directory) and load the id list and packed index into
// RAM, rebuilding the index from the payloads when it is missing or does
// not match the files present.
static async open(dir: string): Promise<MLDataStore> {
const store = new MLDataStore(dir);
await mkdir(dir, { recursive: true });
await store.loadPresent();
await store.loadFetched();
if (!(await store.tryLoadIndex())) await store.rebuildIndex();
return store;
}
// The fileIDs among `files` that must be fetched: every file with no
// payload yet (first run, then new files), plus any whose `updationTime`
// has advanced past the one its cached payload was fetched at. Returned
// sorted and unique.
neededFor(files: MLDataFile[]): number[] {
const latest = new Map<number, number>();
for (const f of files) {
const seen = latest.get(f.id);
if (seen === undefined || f.updationTime > seen)
latest.set(f.id, f.updationTime);
}
const needed: number[] = [];
for (const [id, updationTime] of latest) {
if (!this.present.has(id)) {
needed.push(id);
continue;
}
const at = this.fetched.get(id);
if (at !== undefined && updationTime > at) needed.push(id);
}
return needed.sort((a, b) => a - b);
}
// Store a batch of fetched payloads: write one file per id, fold their CLIP
// embeddings into the packed index (in place for a refetch, appended for a
// new file), and persist the derived files. Returns how many payloads were
// stored and how many ids the index now holds.
async storeFetched(
payloads: Map<number, MLData>,
updation: Map<number, number>,
): Promise<{ stored: number; indexed: number }> {
if (payloads.size === 0) return { stored: 0, indexed: this.ids.length };
for (const [id, payload] of payloads) {
await this.writePayload(id, payload);
this.present.add(id);
const at = updation.get(id);
if (at !== undefined) this.fetched.set(id, at);
}
const updates: { at: number; vector: number[] }[] = [];
const appends: { id: number; vector: number[] }[] = [];
for (const [id, payload] of payloads) {
const vector = clipEmbedding(payload);
if (!vector) continue;
if (this.embeddingLength === 0 && this.ids.length === 0)
this.embeddingLength = vector.length;
// The index is fixed-width; a vector of another length (never seen
// from Ente's CLIP model) is stored but left out of the index.
if (vector.length !== this.embeddingLength) continue;
const at = this.pos.get(id);
if (at !== undefined) updates.push({ at, vector });
else appends.push({ id, vector });
}
for (const { at, vector } of updates)
this.embeddings.set(vector, at * this.embeddingLength);
if (appends.length > 0) {
const length = this.embeddingLength;
const grown = new Float32Array(
this.embeddings.length + appends.length * length,
);
grown.set(this.embeddings);
let offset = this.embeddings.length;
for (const { id, vector } of appends) {
grown.set(vector, offset);
this.pos.set(id, this.ids.length);
this.ids.push(id);
offset += length;
}
this.embeddings = grown;
}
await this.persistIndex();
await this.persistFetched();
return { stored: payloads.size, indexed: this.ids.length };
}
// The packed index the search runs on. The id list is copied so callers
// cannot disturb the store's own order; the embeddings are the live buffer.
getIndex(): MLIndex {
return {
fileIDs: [...this.ids],
embeddingLength: this.embeddingLength,
embeddings: this.embeddings,
};
}
// The full payload for a file, read from disk, or undefined when it is not
// cached or does not parse.
async readPayload(fileID: number): Promise<MLData | undefined> {
if (!this.present.has(fileID)) return undefined;
let raw: string;
try {
raw = await readFile(this.payloadPath(fileID), "utf-8");
} catch {
return undefined;
}
try {
return JSON.parse(raw) as MLData;
} catch {
return undefined;
}
}
stats(): { stored: number; indexed: number } {
return { stored: this.present.size, indexed: this.ids.length };
}
private payloadPath(id: number): string {
return join(this.dir, `${id}.json`);
}
private async writePayload(id: number, payload: MLData): Promise<void> {
await writeAtomic(
this.payloadPath(id),
new TextEncoder().encode(JSON.stringify(payload)),
);
}
private async loadPresent(): Promise<void> {
let names: string[];
try {
names = await readdir(this.dir);
} catch {
return;
}
for (const name of names) {
const match = PAYLOAD_RE.exec(name);
if (match) this.present.add(Number(match[1]));
}
}
private async loadFetched(): Promise<void> {
let raw: string;
try {
raw = await readFile(join(this.dir, FETCHED), "utf-8");
} catch {
return;
}
try {
const parsed = JSON.parse(raw) as Record<string, unknown>;
for (const [key, value] of Object.entries(parsed)) {
const id = Number(key);
if (
Number.isInteger(id) &&
typeof value === "number" &&
this.present.has(id)
)
this.fetched.set(id, value);
}
} catch {
// Corrupt bookkeeping degrades refetch decisions, never fails open.
}
}
// Load the packed index if it is present and agrees with the payloads in
// both directions: every id it names must still be present, its vector file
// must be exactly the size the id count and embedding length imply, and no
// embedding-bearing payload on disk may be missing from it. Returns whether
// it loaded.
private async tryLoadIndex(): Promise<boolean> {
let metaRaw: string;
try {
metaRaw = await readFile(join(this.dir, CLIP_INDEX), "utf-8");
} catch {
return false;
}
let meta: ClipIndexFile;
try {
meta = JSON.parse(metaRaw) as ClipIndexFile;
} catch {
return false;
}
if (
!Array.isArray(meta.fileIDs) ||
typeof meta.embeddingLength !== "number"
)
return false;
if (meta.fileIDs.some((id) => !this.present.has(id))) return false;
// The reverse must hold too. A payload carrying an embedding but absent
// from the index means the index is stale — realistically the process
// died after storeFetched renamed the payloads into place but before it
// rewrote clip.json/clip.f32. Loading such an index as "consistent"
// would drop those embeddings for good (neededFor sees the payloads
// present and never refetches), so treat it as a disagreement and
// rebuild. Only present ids the index omits are read; a payload
// legitimately without an embedding stays out and forces no rebuild.
const indexed = new Set(meta.fileIDs);
for (const id of this.present) {
if (indexed.has(id)) continue;
const payload = await this.readPayload(id);
if (payload && clipEmbedding(payload)) return false;
}
let bytes: Buffer;
try {
bytes = await readFile(join(this.dir, CLIP_VECTORS));
} catch {
return false;
}
const expected =
meta.fileIDs.length * meta.embeddingLength * BYTES_PER_FLOAT;
if (bytes.byteLength !== expected) return false;
// One read, no parse: copy into an aligned buffer and view it as
// floats. The copy is needed because a Buffer from the pool can start
// at an offset a Float32Array cannot be laid over.
const aligned = new Uint8Array(bytes.byteLength);
aligned.set(bytes);
this.embeddings = new Float32Array(aligned.buffer);
this.embeddingLength = meta.embeddingLength;
this.ids = [...meta.fileIDs];
this.pos.clear();
this.ids.forEach((id, i) => this.pos.set(id, i));
return true;
}
// Rebuild the packed index by reading every payload present, then persist
// it. Payloads without a CLIP embedding (or of an unexpected length) are
// simply not indexed.
private async rebuildIndex(): Promise<void> {
this.ids = [];
this.pos.clear();
this.embeddingLength = 0;
const vectors: number[][] = [];
for (const id of [...this.present].sort((a, b) => a - b)) {
const payload = await this.readPayload(id);
if (!payload) continue;
const vector = clipEmbedding(payload);
if (!vector) continue;
if (this.embeddingLength === 0)
this.embeddingLength = vector.length;
if (vector.length !== this.embeddingLength) continue;
this.pos.set(id, this.ids.length);
this.ids.push(id);
vectors.push(vector);
}
const length = this.embeddingLength;
const packed = new Float32Array(this.ids.length * length);
vectors.forEach((vector, i) => packed.set(vector, i * length));
this.embeddings = packed;
await this.persistIndex();
}
private async persistIndex(): Promise<void> {
const meta: ClipIndexFile = {
fileIDs: this.ids,
embeddingLength: this.embeddingLength,
};
await writeAtomic(
join(this.dir, CLIP_INDEX),
new TextEncoder().encode(JSON.stringify(meta)),
);
await writeAtomic(
join(this.dir, CLIP_VECTORS),
new Uint8Array(
this.embeddings.buffer,
this.embeddings.byteOffset,
this.embeddings.byteLength,
),
);
}
private async persistFetched(): Promise<void> {
const record: Record<string, number> = {};
for (const [id, at] of this.fetched) record[id] = at;
await writeAtomic(
join(this.dir, FETCHED),
new TextEncoder().encode(JSON.stringify(record)),
);
}
}
-129
View File
@@ -1,129 +0,0 @@
// The content-similarity search surface over the CLIP index (issue #50).
//
// This is `lib.mldata`. It answers three questions against the ML-data cache
// (#49) without touching the network:
//
// - `forFile` returns the whole stored payload (face boxes, landmarks,
// embedding) for a file, read from disk on demand — the only method here
// that touches the disk, and the only one that is async.
// - `similar` and `searchByEmbedding` rank fileIDs by cosine similarity over
// the packed `Float32Array` index alone. That index (~50k×512) already
// lives in RAM, so each query is a plain loop over it and nothing else.
//
// quak bundles no text encoder (owner-deferred), so `searchByEmbedding` takes
// the query vector the caller has produced elsewhere; `similar` uses the
// query file's own indexed embedding.
import type { MLData } from "../mldata-fetch.js";
import type { MLDataStore, MLIndex } from "./mldata.js";
// How many nearest files a query returns when the caller names no limit.
const DEFAULT_LIMIT = 20;
// One ranked result: a fileID and its cosine similarity to the query, in
// [-1, 1]. Callers wanting only the ids read `.fileID`.
export interface SimilarResult {
fileID: number;
score: number;
}
export interface MLDataAPI {
// The whole stored ML payload for a file, or undefined when it is not
// cached. Reads the payload from disk, so it is async.
forFile(args: { fileID: number }): Promise<MLData | undefined>;
// The files nearest the given file by cosine over their CLIP embeddings,
// most similar first, excluding the file itself. Empty when the file has
// no indexed embedding.
similar(args: { fileID: number; limit?: number }): SimilarResult[];
// The files nearest a caller-supplied query embedding by cosine, most
// similar first. Empty when the query is the wrong length for the index,
// has zero magnitude, or the index is empty.
searchByEmbedding(args: {
embedding: ArrayLike<number>;
limit?: number;
}): SimilarResult[];
}
// Rank the packed index by cosine similarity to `query`, most similar first,
// and return the top `limit`. `skip` (a query file's own id) is left out. Both
// each row's magnitude and the query's are computed here rather than cached:
// the index mutates as ML data is fetched, and one plain pass over ~50k×512
// floats is fast enough that a norm cache would only add a staleness bug. A
// zero-magnitude vector has no direction, so it is dropped rather than divided
// by zero.
const topByCosine = (
index: MLIndex,
query: ArrayLike<number>,
limit: number,
skip?: number,
): SimilarResult[] => {
const { fileIDs, embeddingLength, embeddings } = index;
if (embeddingLength === 0 || query.length !== embeddingLength) return [];
// Every indexed read below is in range: the inner loops run to
// `embeddingLength`, the query is exactly that long (checked above), and
// the packed buffer holds `fileIDs.length * embeddingLength` floats.
// `noUncheckedIndexedAccess` still widens each read to `number | undefined`,
// so they are asserted non-null rather than paying a per-element guard in
// this hot ~50k×512 loop.
let queryNorm = 0;
for (let k = 0; k < embeddingLength; k++) {
const q = query[k]!;
queryNorm += q * q;
}
queryNorm = Math.sqrt(queryNorm);
if (queryNorm === 0) return [];
const results: SimilarResult[] = [];
for (let i = 0; i < fileIDs.length; i++) {
const id = fileIDs[i]!;
if (id === skip) continue;
const base = i * embeddingLength;
let dot = 0;
let norm = 0;
for (let k = 0; k < embeddingLength; k++) {
const v = embeddings[base + k]!;
dot += query[k]! * v;
norm += v * v;
}
if (norm === 0) continue;
results.push({
fileID: id,
score: dot / (queryNorm * Math.sqrt(norm)),
});
}
// Descending score, ties broken by ascending fileID for a stable order.
results.sort((a, b) => b.score - a.score || a.fileID - b.fileID);
return results.slice(0, Math.max(0, Math.trunc(limit)));
};
// Build the search surface over a store the library supplies lazily (the store
// is absent when the client cannot fetch ML data). Reading it per call keeps
// the surface current as the index grows.
export const makeMLDataAPI = (
store: () => MLDataStore | undefined,
): MLDataAPI => ({
forFile: ({ fileID }): Promise<MLData | undefined> => {
const s = store();
return s ? s.readPayload(fileID) : Promise.resolve(undefined);
},
similar: ({ fileID, limit }): SimilarResult[] => {
const s = store();
if (!s) return [];
const index = s.getIndex();
const pos = index.fileIDs.indexOf(fileID);
if (pos < 0) return [];
const base = pos * index.embeddingLength;
const query = index.embeddings.subarray(
base,
base + index.embeddingLength,
);
return topByCosine(index, query, limit ?? DEFAULT_LIMIT, fileID);
},
searchByEmbedding: ({ embedding, limit }): SimilarResult[] => {
const s = store();
if (!s) return [];
return topByCosine(s.getIndex(), embedding, limit ?? DEFAULT_LIMIT);
},
});
-181
View File
@@ -1,181 +0,0 @@
// Three bounded request pools for metadata, content, and thumbnails (issue
// #45).
//
// Ente meters differently by traffic class, so quak keeps three independent
// pools instead of one global limit: metadata is cheap and chatty, original
// content is heavy, thumbnails are small but numerous. Each pool is a
// `BoundedPool` — a plain concurrency limiter — and the three run at the
// design's caps (10 / 5 / 25) unless the caller overrides them.
//
// Two behaviours beyond a bare limiter, both per pool:
//
// - Priority: work waiting for a slot is ordered on-demand before
// background/precache, so a slot that frees up serves the request a user is
// waiting on ahead of speculative prefetch. Within one priority the order
// is first-come-first-served.
//
// - In-flight dedup: a task submitted under a `key` that a still-pending task
// already carries is not run a second time; both callers await the one
// result. The key is released the moment that task settles — success or
// failure — so a later request for the same key runs afresh. Callers key by
// the id whose fetch must not be duplicated (a fileID, say).
//
// The pool holds a task's slot for that task's entire lifetime. A task that
// retries internally is doing so inside its slot: the slot is not freed between
// attempts, which is what keeps a retrying request counted against the cap. The
// pool knows nothing of the retry policy; it only holds the slot until the
// task's promise settles.
//
// This module is self-contained infrastructure. Routing a given `Client` call
// to the right pool belongs to the unit that wires the pools into the cache;
// here there is only the machinery.
export const DEFAULT_METADATA_CONCURRENCY = 10;
export const DEFAULT_CONTENT_CONCURRENCY = 5;
export const DEFAULT_THUMBNAIL_CONCURRENCY = 25;
// On-demand work is served before background/precache work waiting in the same
// pool.
export type Priority = "on-demand" | "background";
export interface RunOptions {
// Defaults to "background": an unmarked request yields to on-demand work.
priority?: Priority;
// When set, a task already pending under this key is shared instead of run
// again. Omit for work that must always execute.
key?: string | number;
}
// One queued submission awaiting a slot. `start` runs the task and holds the
// slot until it settles.
interface Waiter {
priority: Priority;
// Submission order, used to break ties within a priority (FIFO).
seq: number;
start: () => void;
}
export class BoundedPool {
readonly concurrency: number;
private active = 0;
private nextSeq = 0;
private readonly waiting: Waiter[] = [];
// Keyed by a caller-supplied dedup key; holds the shared promise for as
// long as that task is pending, cleared when it settles.
private readonly pending = new Map<string | number, Promise<unknown>>();
constructor(concurrency: number) {
if (!Number.isInteger(concurrency) || concurrency < 1) {
throw new RangeError(
`concurrency must be a positive integer, got ${concurrency}`,
);
}
this.concurrency = concurrency;
}
// Submit `task` to the pool. It runs once a slot is free, subject to
// priority; the returned promise settles with the task's result. With a
// `key`, a still-pending submission under the same key is returned instead
// of running `task` again.
run<T>(task: () => Promise<T>, opts: RunOptions = {}): Promise<T> {
const { key } = opts;
if (key !== undefined) {
const shared = this.pending.get(key);
if (shared !== undefined) return shared as Promise<T>;
}
const promise = this.enqueue(task, opts.priority ?? "background");
if (key !== undefined) {
this.pending.set(key, promise);
const release = (): void => {
// Only clear our own entry: a fresh submission under the same
// key after this one settled must not be evicted here.
if (this.pending.get(key) === promise) this.pending.delete(key);
};
promise.then(release, release);
}
return promise;
}
private enqueue<T>(task: () => Promise<T>, priority: Priority): Promise<T> {
return new Promise<T>((resolve, reject) => {
const start = (): void => {
this.active++;
// Hold the slot until the task fully settles — every internal
// retry included — then admit the next waiter.
void (async () => {
try {
resolve(await task());
} catch (err) {
reject(err);
} finally {
this.active--;
this.pump();
}
})();
};
this.waiting.push({ priority, seq: this.nextSeq++, start });
this.pump();
});
}
// Admit waiters until the pool is full or the queue is empty.
private pump(): void {
while (this.active < this.concurrency) {
const next = this.takeNext();
if (next === undefined) return;
next.start();
}
}
// Remove and return the highest-priority waiter: on-demand before
// background, earliest submission first within a priority.
private takeNext(): Waiter | undefined {
let bestIndex = -1;
let best: Waiter | undefined;
for (let i = 0; i < this.waiting.length; i++) {
const w = this.waiting[i];
if (w === undefined) continue;
if (best === undefined || this.precedes(w, best)) {
best = w;
bestIndex = i;
}
}
if (best === undefined) return undefined;
this.waiting.splice(bestIndex, 1);
return best;
}
private precedes(a: Waiter, b: Waiter): boolean {
if (a.priority !== b.priority) return a.priority === "on-demand";
return a.seq < b.seq;
}
}
export interface RequestPoolsOptions {
metadataConcurrency?: number;
contentConcurrency?: number;
thumbnailConcurrency?: number;
}
// The three pools the design calls for, each independent: an idle pool never
// lends its slots to a busy one.
export class RequestPools {
readonly metadata: BoundedPool;
readonly content: BoundedPool;
readonly thumbnails: BoundedPool;
constructor(opts: RequestPoolsOptions = {}) {
this.metadata = new BoundedPool(
opts.metadataConcurrency ?? DEFAULT_METADATA_CONCURRENCY,
);
this.content = new BoundedPool(
opts.contentConcurrency ?? DEFAULT_CONTENT_CONCURRENCY,
);
this.thumbnails = new BoundedPool(
opts.thumbnailConcurrency ?? DEFAULT_THUMBNAIL_CONCURRENCY,
);
}
}
-265
View File
@@ -1,265 +0,0 @@
// The aggressive local precache (issue #48), started from `Library.open` with
// no caller input.
//
// Two background fills run concurrently through the shared request pools (#45):
//
// - Thumbnails: every file in the account, newest first, through the
// thumbnail pool until all are on disk. Never evicted. The pool is the same
// one `thumbnails.ensure` uses, so a visible or ahead request always jumps
// ahead of this background fill and a fileID both want is fetched once.
//
// - Originals (the pinned set): through the content pool, the favorites album
// first, then every file whose `takenAt` falls in the latest
// `originalsDays` window — the days ending at the newest file in the
// account. The pinned set is the eviction predicate (#47): a pinned
// original is never evicted, and a file that leaves the set (a favorite
// removed, or the window moving past it on a later refresh) becomes an
// ordinary, evictable original with its bytes left in place.
//
// Both fills yield to on-demand work: every fetch goes to its pool at
// background priority, which the pool serves only after on-demand requests. A
// file already cached costs one map lookup (`pathsFor`) and no fetch. Each
// sweep is driven in bounded chunks so the pool's waiting queue never grows to
// the whole account, keeping on-demand preemption and per-admit cost cheap on a
// large library. Failures are not fatal: an uncached file is retried on the
// next sweep, which `Library` re-kicks after every refresh.
import type { EnsureResult } from "./content.js";
import type { DerivedRecords } from "./records.js";
const ONE_DAY_MS = 24 * 60 * 60 * 1000;
export const DEFAULT_PRECACHE_ORIGINALS_DAYS = 7;
// How many files a sweep submits to a pool before awaiting them. Bounds the
// pool's waiting queue so on-demand work is never stuck behind the whole
// account; the values track each pool's concurrency (#45).
const THUMBNAIL_CHUNK = 25;
const ORIGINAL_CHUNK = 5;
// The precache metrics `Library.status()` surfaces.
export interface PrecacheStatus {
thumbnailsCached: number;
thumbnailsTotal: number;
originalsCached: number;
originalsPinned: number;
}
// A background-fill progress event. Each active sweep fires "started" before
// its fetches and "done" (with the count newly on disk) after; a sweep with
// nothing left to fetch is silent.
export interface PrecacheEvent {
operation: "precacheThumbnails" | "precacheOriginals";
status: "started" | "done";
fetched?: number;
}
// The slice of the content cache the precache drives. The real `ContentCache`
// satisfies it; tests inject a fake.
export interface PrecacheCache {
pathsFor(fileID: number): { originalPath?: string; thumbnailPath?: string };
ensureThumbnails(args: {
fileIDs: number[];
priority: "background";
signal?: AbortSignal;
}): Promise<EnsureResult[]>;
ensureOriginals(args: {
fileIDs: number[];
signal?: AbortSignal;
}): Promise<EnsureResult[]>;
}
export interface PrecacheOptions {
// Default true; false disables the fill and, for originals, the pinning.
thumbnails?: boolean;
originals?: boolean;
// The latest-week window length in days; default 7.
originalsDays?: number;
onEvent?: (event: PrecacheEvent) => void;
}
export class Precache {
private readonly doThumbnails: boolean;
private readonly doOriginals: boolean;
private readonly originalsDays: number;
private readonly onEvent?: (event: PrecacheEvent) => void;
private cache?: PrecacheCache;
// Every file in the account, newest first (the thumbnail fill order).
private thumbOrder: number[] = [];
// The pinned originals in fetch order: favorites first, then the window.
private originalsOrder: number[] = [];
private pinned = new Set<number>();
// A sweep runs at most once per fill at a time; a re-kick while one runs is
// a no-op, and the next refresh re-kicks after it finishes.
private thumbRunning = false;
private originalsRunning = false;
private readonly aborter = new AbortController();
private closed = false;
constructor(opts: PrecacheOptions = {}) {
this.doThumbnails = opts.thumbnails ?? true;
this.doOriginals = opts.originals ?? true;
this.originalsDays =
opts.originalsDays ?? DEFAULT_PRECACHE_ORIGINALS_DAYS;
this.onEvent = opts.onEvent;
}
// Attach the cache the fills fetch through. `isPinned` works before this is
// called, so the cache can be constructed with `isPinned` wired in and then
// bound here.
bind(cache: PrecacheCache): void {
this.cache = cache;
}
// Whether an original is pinned (favorites + the latest-week window), the
// eviction predicate (#47). False for every file when originals precaching
// is disabled.
isPinned(fileID: number): boolean {
return this.pinned.has(fileID);
}
// Recompute the fill orders and the pinned set from the current projection.
// Called at open and after every refresh that changes the library.
update(records: DerivedRecords): void {
const photos = [...records.photos.values()].sort(
(a, b) => b.takenAt - a.takenAt || b.fileID - a.fileID,
);
this.thumbOrder = this.doThumbnails ? photos.map((p) => p.fileID) : [];
const pinned = new Set<number>();
const order: number[] = [];
if (this.doOriginals) {
// Favorites first, in the album's own newest-first order.
for (const album of records.albums.values()) {
if (album.type !== "favorites") continue;
for (const id of album.fileIDs) {
if (!pinned.has(id)) {
pinned.add(id);
order.push(id);
}
}
}
// Then the latest-week window, ending at the newest file. Photos
// are newest first, so stop once one falls before the window start.
const newest = photos[0];
if (newest !== undefined) {
const windowStart =
newest.takenAt - this.originalsDays * ONE_DAY_MS;
for (const p of photos) {
if (p.takenAt < windowStart) break;
if (!pinned.has(p.fileID)) {
pinned.add(p.fileID);
order.push(p.fileID);
}
}
}
}
this.pinned = pinned;
this.originalsOrder = order;
}
status(): PrecacheStatus {
const cache = this.cache;
let thumbnailsCached = 0;
let originalsCached = 0;
if (cache) {
for (const id of this.thumbOrder)
if (cache.pathsFor(id).thumbnailPath !== undefined)
thumbnailsCached++;
for (const id of this.originalsOrder)
if (cache.pathsFor(id).originalPath !== undefined)
originalsCached++;
}
return {
thumbnailsCached,
thumbnailsTotal: this.thumbOrder.length,
originalsCached,
originalsPinned: this.originalsOrder.length,
};
}
// Kick both fills. Idempotent: a fill already sweeping is left alone. Safe
// to call after every refresh; a finished fill starts a fresh sweep that
// picks up new files and retries any that failed before.
start(): void {
if (this.closed || !this.cache) return;
if (this.doThumbnails) this.kickThumbnails();
if (this.doOriginals) this.kickOriginals();
}
// Stop the fills. In-flight fetches are left to settle; queued ones drop.
close(): void {
this.closed = true;
this.aborter.abort();
}
private kickThumbnails(): void {
if (this.thumbRunning) return;
this.thumbRunning = true;
void this.sweep(
"precacheThumbnails",
() => this.thumbOrder,
(id) => this.cache!.pathsFor(id).thumbnailPath !== undefined,
THUMBNAIL_CHUNK,
(chunk) =>
this.cache!.ensureThumbnails({
fileIDs: chunk,
priority: "background",
signal: this.aborter.signal,
}),
).finally(() => {
this.thumbRunning = false;
});
}
private kickOriginals(): void {
if (this.originalsRunning) return;
this.originalsRunning = true;
void this.sweep(
"precacheOriginals",
() => this.originalsOrder,
(id) => this.cache!.pathsFor(id).originalPath !== undefined,
ORIGINAL_CHUNK,
(chunk) =>
this.cache!.ensureOriginals({
fileIDs: chunk,
signal: this.aborter.signal,
}),
).finally(() => {
this.originalsRunning = false;
});
}
// One fill sweep: skip files already on disk (one lookup each), fetch the
// rest in bounded chunks, and report progress only when there was work.
private async sweep(
operation: PrecacheEvent["operation"],
order: () => number[],
present: (fileID: number) => boolean,
chunkSize: number,
fetch: (chunk: number[]) => Promise<EnsureResult[]>,
): Promise<void> {
const todo = order().filter((id) => !present(id));
if (todo.length === 0) return;
this.emit({ operation, status: "started" });
let fetched = 0;
for (let i = 0; i < todo.length && !this.closed; i += chunkSize) {
const results = await fetch(todo.slice(i, i + chunkSize));
for (const r of results) if (r.path !== undefined) fetched++;
}
this.emit({ operation, status: "done", fetched });
}
private emit(event: PrecacheEvent): void {
if (!this.onEvent) return;
// A misbehaving callback must not break the fill loop.
try {
this.onEvent(event);
} catch {
// ignore
}
}
}
-388
View File
@@ -1,388 +0,0 @@
// The in-process read surface over the local cache (issue #44).
//
// A CLI or an in-process script reads albums, photos, and a grouped timeline
// through `lib.albums`, `lib.photos`, and `lib.timeline`. Every call is
// answered synchronously from the same plain-record projection the GUI reads
// (`deriveRecords`, issue #43); nothing here touches the network. Every method
// takes a single named-argument object.
//
// The `Album` and `Photo` classes are thin, in-process-only wrappers over
// those records: a caller that holds an object reference gets typed field
// access and, for an album, its photos. They are not sent across IPC — the
// plain records are the serializable surface, and `record()` returns one.
//
// A `Photo` also fetches its own bytes: `original()` and `thumbnail()` go
// through the on-disk content cache (issue #46), the one place in this module
// that is not synchronous and RAM-only. A library opened without a content
// source leaves that cache absent, and those two methods then throw.
import type { CollectionType, FileType } from "../model/types.js";
import type { ContentOptions, ContentResult, PhotoContent } from "./content.js";
import type { AlbumRecord, PhotoRecord, DerivedRecords } from "./records.js";
// Newest first, with fileID as a stable tiebreak so equal-timed files order
// deterministically — the same order the record projection uses.
const byNewest = (a: PhotoRecord, b: PhotoRecord): number =>
b.takenAt - a.takenAt || b.fileID - a.fileID;
// Albums newest updated first, collection id breaking ties. This is the order
// `albums.list` returns and the order `byName` resolves a name collision in.
const byNewestAlbum = (a: AlbumRecord, b: AlbumRecord): number =>
b.updationTime - a.updationTime || b.collectionID - a.collectionID;
// A single photo. Field access mirrors `PhotoRecord`; `record()` returns the
// underlying plain record for callers that need the IPC-safe value.
export class Photo {
constructor(
private readonly rec: PhotoRecord,
private readonly content?: PhotoContent,
) {}
get fileID(): number {
return this.rec.fileID;
}
get albumIDs(): number[] {
return this.rec.albumIDs;
}
get title(): string {
return this.rec.title;
}
get takenAt(): number {
return this.rec.takenAt;
}
get fileType(): FileType {
return this.rec.fileType;
}
get caption(): string | undefined {
return this.rec.caption;
}
get width(): number | undefined {
return this.rec.width;
}
get height(): number | undefined {
return this.rec.height;
}
get latitude(): number | undefined {
return this.rec.latitude;
}
get longitude(): number | undefined {
return this.rec.longitude;
}
get isArchived(): boolean {
return this.rec.isArchived;
}
get isHidden(): boolean {
return this.rec.isHidden;
}
record(): PhotoRecord {
return this.rec;
}
// Fetch and cache the full-resolution original, returning its on-disk path
// and byte length. Served from the cache (or the backup download directory)
// when already present, otherwise fetched through the content pool.
async original(opts?: ContentOptions): Promise<ContentResult> {
return this.contentOrThrow().original(this.rec.fileID, opts);
}
// As `original`, for the thumbnail, through the thumbnail pool.
async thumbnail(opts?: ContentOptions): Promise<ContentResult> {
return this.contentOrThrow().thumbnail(this.rec.fileID, opts);
}
private contentOrThrow(): PhotoContent {
if (!this.content) {
throw new Error(
"Photo content requires a library opened with a content cache",
);
}
return this.content;
}
}
// A single album. `photos.list()` returns the album's photos as wrappers,
// newest first (the record already stores `fileIDs` in that order).
export class Album {
constructor(
private readonly rec: AlbumRecord,
private readonly records: DerivedRecords,
private readonly content?: PhotoContent,
) {}
get collectionID(): number {
return this.rec.collectionID;
}
get name(): string {
return this.rec.name;
}
get type(): CollectionType {
return this.rec.type;
}
get isShared(): boolean {
return this.rec.isShared;
}
get updationTime(): number {
return this.rec.updationTime;
}
get fileIDs(): number[] {
return this.rec.fileIDs;
}
get photos(): { list: () => Photo[] } {
return { list: (): Photo[] => this.listPhotos() };
}
record(): AlbumRecord {
return this.rec;
}
private listPhotos(): Photo[] {
const out: Photo[] = [];
for (const id of this.rec.fileIDs) {
const p = this.records.photos.get(id);
if (p) out.push(new Photo(p, this.content));
}
return out;
}
}
export interface AlbumsAPI {
list(): Album[];
byName(args: { albumName: string }): Album | undefined;
byID(args: { collectionID: number }): Album | undefined;
}
export interface PhotosAPI {
byID(args: { fileID: number }): Photo | undefined;
// Plain records for the requested ids, in the order requested, each id at
// most once, unknown ids dropped.
records(args: { fileIDs: number[] }): PhotoRecord[];
}
export type GroupBy = "day" | "week" | "month";
// A filter over the timeline. All fields are optional and combine with AND.
// Hidden photos are never included, regardless of this filter.
export interface PhotoFilter {
// Keep only photos that belong to this album.
albumID?: number;
// Case-insensitive substring of the title, caption, or any album name the
// photo belongs to.
text?: string;
// Keep only photos of one of these types.
fileTypes?: FileType[];
// `true` keeps only geotagged photos; `false` keeps only those without a
// location; omitted places no constraint.
hasLocation?: boolean;
// Archived photos are excluded unless this is `true`. Defaults to `false`.
includeArchived?: boolean;
}
export interface TimelineGroup {
// The period's identity: `YYYY-MM-DD` for day, `YYYY-Www` (ISO 8601 week,
// e.g. `2025-W32`) for week, and `YYYY-MM` for month.
key: string;
// Local-time milliseconds at the start of the period.
startsAt: number;
// The period's files, newest first, each file once.
fileIDs: number[];
}
export interface TimelineAPI {
groups(args: { groupBy: GroupBy; filter?: PhotoFilter }): TimelineGroup[];
}
// The surface `Library.fresh()` resolves to (issue #75). It is the same three
// read namespaces as the default `albums`/`photos`/`timeline`, handed back only
// after a forced refresh has brought the local copy current.
export interface FreshReads {
albums: AlbumsAPI;
photos: PhotosAPI;
timeline: TimelineAPI;
}
export const makeAlbumsAPI = (
derive: () => DerivedRecords,
content?: PhotoContent,
): AlbumsAPI => ({
list: (): Album[] => {
const records = derive();
return [...records.albums.values()]
.sort(byNewestAlbum)
.map((rec) => new Album(rec, records, content));
},
byID: ({ collectionID }): Album | undefined => {
const records = derive();
const rec = records.albums.get(collectionID);
return rec ? new Album(rec, records, content) : undefined;
},
byName: ({ albumName }): Album | undefined => {
const records = derive();
// Names are not unique in Ente; resolve a collision deterministically
// to the newest-updated album, matching `list` order.
const match = [...records.albums.values()]
.sort(byNewestAlbum)
.find((rec) => rec.name === albumName);
return match ? new Album(match, records, content) : undefined;
},
});
export const makePhotosAPI = (
derive: () => DerivedRecords,
content?: PhotoContent,
): PhotosAPI => ({
byID: ({ fileID }): Photo | undefined => {
const rec = derive().photos.get(fileID);
return rec ? new Photo(rec, content) : undefined;
},
records: ({ fileIDs }): PhotoRecord[] => {
const { photos } = derive();
const seen = new Set<number>();
const out: PhotoRecord[] = [];
for (const id of fileIDs) {
if (seen.has(id)) continue;
const rec = photos.get(id);
if (rec) {
out.push(rec);
seen.add(id);
}
}
return out;
},
});
export const makeTimelineAPI = (derive: () => DerivedRecords): TimelineAPI => ({
groups: ({ groupBy, filter }): TimelineGroup[] => {
const records = derive();
return groupPhotos(filterPhotos(records, filter), groupBy);
},
});
// Apply a `PhotoFilter` to the projection. Hidden photos are always dropped;
// archived photos are dropped unless `includeArchived` asks for them.
const filterPhotos = (
records: DerivedRecords,
filter?: PhotoFilter,
): PhotoRecord[] => {
const f = filter ?? {};
const includeArchived = f.includeArchived ?? false;
const needle = f.text?.toLowerCase();
const out: PhotoRecord[] = [];
for (const rec of records.photos.values()) {
if (rec.isHidden) continue;
if (rec.isArchived && !includeArchived) continue;
if (f.albumID !== undefined && !rec.albumIDs.includes(f.albumID))
continue;
if (f.fileTypes !== undefined && !f.fileTypes.includes(rec.fileType))
continue;
if (f.hasLocation !== undefined) {
const has =
rec.latitude !== undefined && rec.longitude !== undefined;
if (has !== f.hasLocation) continue;
}
if (needle !== undefined && !matchesText(rec, needle, records))
continue;
out.push(rec);
}
return out;
};
const matchesText = (
rec: PhotoRecord,
needle: string,
records: DerivedRecords,
): boolean => {
if (rec.title.toLowerCase().includes(needle)) return true;
if (rec.caption !== undefined && rec.caption.toLowerCase().includes(needle))
return true;
for (const id of rec.albumIDs) {
const album = records.albums.get(id);
if (album && album.name.toLowerCase().includes(needle)) return true;
}
return false;
};
// Bucket photos into periods, groups newest first, members newest first.
const groupPhotos = (
photos: PhotoRecord[],
groupBy: GroupBy,
): TimelineGroup[] => {
const buckets = new Map<
string,
{ startsAt: number; recs: PhotoRecord[] }
>();
for (const rec of photos) {
const { key, startsAt } = periodOf(rec.takenAt, groupBy);
const bucket = buckets.get(key);
if (bucket) bucket.recs.push(rec);
else buckets.set(key, { startsAt, recs: [rec] });
}
const groups: TimelineGroup[] = [];
for (const [key, bucket] of buckets) {
bucket.recs.sort(byNewest);
groups.push({
key,
startsAt: bucket.startsAt,
fileIDs: bucket.recs.map((r) => r.fileID),
});
}
groups.sort((a, b) => b.startsAt - a.startsAt);
return groups;
};
const pad = (n: number): string => String(n).padStart(2, "0");
const dateKey = (d: Date): string =>
`${d.getFullYear()}-${pad(d.getMonth() + 1)}-${pad(d.getDate())}`;
const WEEK_MS = 7 * 24 * 60 * 60 * 1000;
// The ISO 8601 week key `YYYY-Www` for the week starting at the given Monday.
// The week-year is the year of that week's Thursday, so it can differ from the
// calendar year at the January/December boundary (e.g. 2024-12-30 is 2025-W01).
const isoWeekKey = (monday: Date): string => {
const thursday = new Date(
monday.getFullYear(),
monday.getMonth(),
monday.getDate() + 3,
);
const isoYear = thursday.getFullYear();
// Thursday of ISO week 1 is the Thursday of the week containing January 4.
const jan4 = new Date(isoYear, 0, 4);
const week1Thursday = new Date(
isoYear,
0,
4 + 3 - ((jan4.getDay() + 6) % 7),
);
const week =
1 +
Math.round((thursday.getTime() - week1Thursday.getTime()) / WEEK_MS);
return `${isoYear}-W${pad(week)}`;
};
// The period a millisecond instant falls in, in local time. Weeks start on
// Monday. `Date` normalizes out-of-range day arguments, so the week's Monday
// is correct across month and year boundaries.
const periodOf = (
takenAt: number,
groupBy: GroupBy,
): { key: string; startsAt: number } => {
const d = new Date(takenAt);
const year = d.getFullYear();
const month = d.getMonth();
const day = d.getDate();
if (groupBy === "month") {
const start = new Date(year, month, 1);
return { key: `${year}-${pad(month + 1)}`, startsAt: start.getTime() };
}
if (groupBy === "week") {
// getDay(): 0=Sunday..6=Saturday; shift so Monday is the week start.
const fromMonday = (d.getDay() + 6) % 7;
const start = new Date(year, month, day - fromMonday);
return { key: isoWeekKey(start), startsAt: start.getTime() };
}
const start = new Date(year, month, day);
return { key: dateKey(start), startsAt: start.getTime() };
};
-270
View File
@@ -1,270 +0,0 @@
// Plain records projected from the decrypted store, and the diff between two
// projections. These are the library's GUI-facing surface: they hold no key
// material and no binary, so they survive `structuredClone`/JSON across the
// Electron IPC boundary where methods and file keys cannot go (design #36,
// owner ruling 5). The decrypted `Collection`/`EnteFile` objects stay in RAM in
// the main process; the window only ever sees these records.
//
// Ente holds edited/basic times in microseconds; records expose `takenAt` in
// milliseconds. The magic-metadata field names below are the ones the Ente
// clients write, confirmed against the repo's own fixtures: `w`/`h` in
// test/cli/metadata-backup.test.ts, `visibility` in test/library/store.test.ts.
import type {
Collection,
CollectionType,
EnteFile,
FileType,
} from "../model/types.js";
// Ente private-magic-metadata visibility values.
const VISIBILITY_ARCHIVED = 1;
const VISIBILITY_HIDDEN = 2;
// A single photo, deduplicated across the collections it belongs to. No key,
// no binary: safe to send to a window.
export interface PhotoRecord {
fileID: number;
// Every collection this file is a member of, ascending.
albumIDs: number[];
// `pubMagicMetadata.editedName` when the user renamed the file, else the
// basic-metadata title.
title: string;
// Milliseconds. `pubMagicMetadata.editedTime` when the user edited the
// date, else basic-metadata `creationTime`.
takenAt: number;
fileType: FileType;
caption?: string;
width?: number;
height?: number;
latitude?: number;
longitude?: number;
isArchived: boolean;
isHidden: boolean;
// Local cache paths, set once a later phase caches the bytes; unset here.
thumbnailPath?: string;
originalPath?: string;
}
export interface AlbumRecord {
collectionID: number;
name: string;
// `favorites` identifies the account's favorites album.
type: CollectionType;
isShared: boolean;
updationTime: number;
// The album's files, newest first.
fileIDs: number[];
}
export interface LibrarySnapshot {
albums: AlbumRecord[];
photos: PhotoRecord[];
// Wall-clock milliseconds when the snapshot was taken.
takenAt: number;
}
export interface LibraryChange {
// Full records for albums/photos added or changed by the refresh.
albumsChanged: AlbumRecord[];
photosChanged: PhotoRecord[];
fileIDsRemoved: number[];
albumIDsRemoved: number[];
// Wall-clock milliseconds of the refresh that produced this change.
refreshedAt: number;
}
// The by-id projection of the store at one moment; the source for both
// `snapshotFrom` (sorted arrays for the GUI) and `diffRecords` (change sets).
export interface DerivedRecords {
albums: Map<number, AlbumRecord>;
photos: Map<number, PhotoRecord>;
}
const asString = (v: unknown): string | undefined =>
typeof v === "string" && v.length > 0 ? v : undefined;
const asNumber = (v: unknown): number | undefined =>
typeof v === "number" && Number.isFinite(v) ? v : undefined;
const microsToMillis = (micros: number): number => Math.floor(micros / 1000);
// Newest first, with fileID as a stable tiebreak so equal-timed files order
// deterministically.
const byNewestPhoto = (a: PhotoRecord, b: PhotoRecord): number =>
b.takenAt - a.takenAt || b.fileID - a.fileID;
// Build one PhotoRecord from every membership of a file. The memberships share
// the same underlying file, so metadata is read from a single representative
// (the most recently synced, lowest collection id to break ties); `albumIDs`
// gathers them all.
const toPhotoRecord = (
fileID: number,
memberships: EnteFile[],
): PhotoRecord => {
const albumIDs = memberships
.map((m) => m.collectionID)
.sort((a, b) => a - b);
const rep = memberships.reduce((best, m) =>
m.updationTime > best.updationTime ||
(m.updationTime === best.updationTime &&
m.collectionID < best.collectionID)
? m
: best,
);
const pub = rep.pubMagicMetadata ?? {};
const priv = rep.magicMetadata ?? {};
const takenAtMicros = asNumber(pub.editedTime) ?? rep.metadata.creationTime;
const visibility = asNumber(priv.visibility);
const record: PhotoRecord = {
fileID,
albumIDs,
title: asString(pub.editedName) ?? rep.metadata.title,
takenAt: microsToMillis(takenAtMicros),
fileType: rep.metadata.fileType,
isArchived: visibility === VISIBILITY_ARCHIVED,
isHidden: visibility === VISIBILITY_HIDDEN,
};
const caption = asString(pub.caption);
if (caption !== undefined) record.caption = caption;
const width = asNumber(pub.w);
if (width !== undefined) record.width = width;
const height = asNumber(pub.h);
if (height !== undefined) record.height = height;
if (rep.metadata.latitude !== undefined)
record.latitude = rep.metadata.latitude;
if (rep.metadata.longitude !== undefined)
record.longitude = rep.metadata.longitude;
return record;
};
const toAlbumRecord = (
collection: Collection,
files: EnteFile[],
takenAtByFile: Map<number, number>,
): AlbumRecord => {
const fileIDs = files
.filter((f) => f.collectionID === collection.id)
.map((f) => f.id)
.sort(
(a, b) =>
(takenAtByFile.get(b) ?? 0) - (takenAtByFile.get(a) ?? 0) ||
b - a,
);
return {
collectionID: collection.id,
name: collection.name,
type: collection.type,
isShared: collection.isShared,
updationTime: collection.updationTime,
fileIDs,
};
};
// The cache paths known for a file, so the projection can expose them on the
// record without the read layer reaching into the content cache itself.
export type CachedPathLookup = (fileID: number) => {
originalPath?: string;
thumbnailPath?: string;
};
// Project the decrypted collections and file memberships into by-id records.
// `files` is every membership (a file appears once per collection it is in).
// `cachedPaths`, when given, fills each record's cache paths.
export const deriveRecords = (
collections: Collection[],
files: EnteFile[],
cachedPaths?: CachedPathLookup,
): DerivedRecords => {
const byFileID = new Map<number, EnteFile[]>();
for (const f of files) {
const arr = byFileID.get(f.id);
if (arr) arr.push(f);
else byFileID.set(f.id, [f]);
}
const photos = new Map<number, PhotoRecord>();
const takenAtByFile = new Map<number, number>();
for (const [fileID, memberships] of byFileID) {
const record = toPhotoRecord(fileID, memberships);
if (cachedPaths) {
const paths = cachedPaths(fileID);
if (paths.originalPath !== undefined)
record.originalPath = paths.originalPath;
if (paths.thumbnailPath !== undefined)
record.thumbnailPath = paths.thumbnailPath;
}
photos.set(fileID, record);
takenAtByFile.set(fileID, record.takenAt);
}
const albums = new Map<number, AlbumRecord>();
for (const c of collections) {
albums.set(c.id, toAlbumRecord(c, files, takenAtByFile));
}
return { albums, photos };
};
// Sorted, GUI-ready arrays: albums newest updated first, photos newest first.
export const snapshotFrom = (
records: DerivedRecords,
takenAt: number,
): LibrarySnapshot => ({
albums: [...records.albums.values()].sort(
(a, b) =>
b.updationTime - a.updationTime || b.collectionID - a.collectionID,
),
photos: [...records.photos.values()].sort(byNewestPhoto),
takenAt,
});
// Records compare by value; they are plain and built with a fixed key order, so
// a serialized form is a sound equality key.
const same = (a: unknown, b: unknown): boolean =>
JSON.stringify(a) === JSON.stringify(b);
const diffMap = <T>(
prev: Map<number, T>,
next: Map<number, T>,
): { changed: T[]; removed: number[] } => {
const changed: T[] = [];
for (const [id, record] of next) {
const before = prev.get(id);
if (before === undefined || !same(before, record)) changed.push(record);
}
const removed: number[] = [];
for (const id of prev.keys()) if (!next.has(id)) removed.push(id);
removed.sort((a, b) => a - b);
return { changed, removed };
};
// The change between two projections, or undefined when nothing changed.
export const diffRecords = (
prev: DerivedRecords,
next: DerivedRecords,
refreshedAt: number,
): LibraryChange | undefined => {
const albums = diffMap(prev.albums, next.albums);
const photos = diffMap(prev.photos, next.photos);
if (
albums.changed.length === 0 &&
albums.removed.length === 0 &&
photos.changed.length === 0 &&
photos.removed.length === 0
) {
return undefined;
}
return {
albumsChanged: albums.changed,
photosChanged: photos.changed,
fileIDsRemoved: photos.removed,
albumIDsRemoved: albums.removed,
refreshedAt,
};
};
-205
View File
@@ -1,205 +0,0 @@
// On-disk JSON metadata store for the local cache.
//
// The store keeps one `metadata.json` file holding the account's server
// state: the user id, a schema version, the cursor for the incremental
// collections listing, and the decrypted collection and file records. The
// whole file is read into RAM on load and rewritten as a whole on save; there
// is no partial update and no lock file. A separate refresh unit populates the
// store from the server — this module only stores what it is given.
//
// The file is a cache, so it is never trusted to exist or to be intact: a
// missing or unreadable file loads as an empty store rather than an error, and
// the refresh unit then repopulates it.
import { mkdir, chmod, readFile } from "node:fs/promises";
import { dirname } from "node:path";
import { writeAtomic } from "../download/index.js";
import type { Collection, EnteFile, Microseconds } from "../model/types.js";
// Bumped only when the on-disk shape changes incompatibly. A file written
// under a different version is discarded on load (see `load`): re-fetching
// from the server is always safe and cheaper than migrating a cache.
export const METADATA_SCHEMA_VERSION = 1;
// Directory and file modes match `session.json`: the records hold decrypted
// key material, so on a shared machine only the owner may read them.
const DIR_MODE = 0o700;
const FILE_MODE = 0o600;
// On-disk shapes. They mirror the in-memory model exactly except for the
// binary `key`, which JSON cannot hold and which is stored as base64.
type StoredCollection = Omit<Collection, "key"> & { key: string };
type StoredFile = Omit<EnteFile, "key"> & { key: string };
interface StoredMetadata {
schemaVersion: number;
userID: number;
collectionsSinceTime: Microseconds;
collections: StoredCollection[];
files: StoredFile[];
}
const encodeKey = (key: Uint8Array): string =>
Buffer.from(key).toString("base64");
const decodeKey = (encoded: string): Uint8Array =>
new Uint8Array(Buffer.from(encoded, "base64"));
// A file membership is identified by the pair (collectionID, fileID): the same
// underlying file can belong to several collections, each a distinct record
// with its own key.
const fileKey = (collectionID: number, fileID: number): string =>
`${collectionID}:${fileID}`;
export class MetadataStore {
readonly path: string;
readonly schemaVersion = METADATA_SCHEMA_VERSION;
userID = 0;
collectionsSinceTime: Microseconds = 0;
// True when `load` populated this store from a valid existing file; false
// on a first run or a missing/corrupt/wrong-version file that loaded empty.
// `Library.open` reads it to decide whether the first refresh may run in
// the background (an existing copy already serves reads) or must be awaited.
loadedFromDisk = false;
private readonly collections = new Map<number, Collection>();
private readonly files = new Map<string, EnteFile>();
private constructor(path: string) {
this.path = path;
}
// Load the store at `path`. A missing file, an unreadable one, unparseable
// contents, or a mismatched schema version all yield an empty store bound
// to that path — never a thrown error, because the file is only a cache.
static async load(path: string): Promise<MetadataStore> {
const store = new MetadataStore(path);
let raw: string;
try {
raw = await readFile(path, "utf8");
} catch {
return store;
}
try {
const parsed = JSON.parse(raw) as StoredMetadata;
if (parsed.schemaVersion !== METADATA_SCHEMA_VERSION) {
return store;
}
store.loadedFromDisk = true;
store.userID = parsed.userID ?? 0;
store.collectionsSinceTime = parsed.collectionsSinceTime ?? 0;
for (const stored of parsed.collections ?? []) {
const collection: Collection = {
...stored,
key: decodeKey(stored.key),
};
store.collections.set(collection.id, collection);
}
for (const stored of parsed.files ?? []) {
const file: EnteFile = {
...stored,
key: decodeKey(stored.key),
};
store.files.set(fileKey(file.collectionID, file.id), file);
}
} catch {
// Any corruption discards the partial result: a half-read cache is
// worse than an empty one, since the refresh unit will rebuild it.
return new MetadataStore(path);
}
return store;
}
// Rewrite the whole file. The directory is created 0700 and the file left
// 0600; the write itself is the download layer's durable atomic writer
// (temp file, fsync, rename, dir fsync), so a reader never sees a partial
// file and a crash cannot leave a truncated one. There is no lock file and
// no `sync()` beyond the writer's own fsyncs.
async save(): Promise<void> {
const model: StoredMetadata = {
schemaVersion: METADATA_SCHEMA_VERSION,
userID: this.userID,
collectionsSinceTime: this.collectionsSinceTime,
collections: [...this.collections.values()].map((c) => ({
...c,
key: encodeKey(c.key),
})),
files: [...this.files.values()].map((f) => ({
...f,
key: encodeKey(f.key),
})),
};
const dir = dirname(this.path);
// chmod after mkdir so the mode is 0700 even when the directory
// already existed with a looser mode; mkdir alone would not tighten
// an existing directory.
await mkdir(dir, { recursive: true, mode: DIR_MODE });
await chmod(dir, DIR_MODE);
const payload = new TextEncoder().encode(
JSON.stringify(model, null, 2),
);
await writeAtomic(this.path, payload);
// The atomic writer's temp file inherits the default mode; tighten the
// renamed file to 0600. The 0700 directory already keeps other users
// out during the brief window before this runs.
await chmod(this.path, FILE_MODE);
}
getCollection(id: number): Collection | undefined {
return this.collections.get(id);
}
listCollections(): Collection[] {
return [...this.collections.values()];
}
putCollection(collection: Collection): void {
this.collections.set(collection.id, collection);
}
// Removing a collection also drops its file memberships: a file record is
// only meaningful as part of a collection the cache still knows about.
deleteCollection(id: number): void {
this.collections.delete(id);
for (const [key, file] of this.files) {
if (file.collectionID === id) {
this.files.delete(key);
}
}
}
getFile(collectionID: number, fileID: number): EnteFile | undefined {
return this.files.get(fileKey(collectionID, fileID));
}
// Any membership of a file, or undefined. Every membership re-wraps the
// same underlying content key, so any one is enough to fetch the bytes;
// the content cache resolves a fileID to a file this way.
getFileByID(fileID: number): EnteFile | undefined {
for (const file of this.files.values()) {
if (file.id === fileID) return file;
}
return undefined;
}
listFiles(collectionID: number): EnteFile[] {
return [...this.files.values()].filter(
(f) => f.collectionID === collectionID,
);
}
putFile(file: EnteFile): void {
this.files.set(fileKey(file.collectionID, file.id), file);
}
deleteFile(collectionID: number, fileID: number): void {
this.files.delete(fileKey(collectionID, fileID));
}
}
+73 -35
View File
@@ -1,10 +1,17 @@
import { mkdirSync, readFileSync, writeFileSync } from "node:fs"; import { gunzipSync } from "node:zlib";
import {
mkdirSync,
mkdtempSync,
readFileSync,
rmSync,
writeFileSync,
} from "node:fs";
import { join } from "node:path"; import { join } from "node:path";
import { tmpdir } from "node:os";
import * as jpeg from "jpeg-js"; import * as jpeg from "jpeg-js";
import exifReader from "exif-reader"; import exifReader from "exif-reader";
import type { Client } from "./client.js"; import type { Client } from "./client.js";
import type { Library, Photo } from "./library/index.js"; import { decryptBlob, fromBase64 } from "./crypto/index.js";
import { fetchMLData } from "./mldata-fetch.js";
import type { EnteFile } from "./model/types.js"; import type { EnteFile } from "./model/types.js";
export type ProgressCallback = (message: string) => void; export type ProgressCallback = (message: string) => void;
@@ -17,6 +24,50 @@ export interface MetadataBackupOptions {
const sanitizePath = (name: string): string => const sanitizePath = (name: string): string =>
name.replace(/[/\\:*?"<>|]/g, "_").replace(/^\.+/, "_"); name.replace(/[/\\:*?"<>|]/g, "_").replace(/^\.+/, "_");
interface RawRemoteFileData {
fileID: number;
encryptedData: string;
decryptionHeader: string;
updatedAt?: number;
}
const fetchMLDataForFiles = async (
client: Client,
fileIDs: number[],
fileKeys: Map<number, Uint8Array>,
): Promise<Map<number, Record<string, unknown>>> => {
const api = client.getApiClient();
const result = new Map<number, Record<string, unknown>>();
const batchSize = 200;
for (let i = 0; i < fileIDs.length; i += batchSize) {
const batch = fileIDs.slice(i, i + batchSize);
const { data } = await api.postJSON<{ data: RawRemoteFileData[] }>(
"/files/data/fetch",
{ type: "mldata", fileIDs: batch },
);
for (const entry of data ?? []) {
const key = fileKeys.get(entry.fileID);
if (!key) continue;
try {
const decrypted = decryptBlob(
fromBase64(entry.encryptedData),
fromBase64(entry.decryptionHeader),
key,
);
const jsonStr = gunzipSync(Buffer.from(decrypted)).toString(
"utf-8",
);
result.set(entry.fileID, JSON.parse(jsonStr));
} catch {
// Corrupted ML data for this file; skip it
}
}
}
return result;
};
// Extract the raw EXIF APP1 segment from JPEG bytes. Returns the EXIF // Extract the raw EXIF APP1 segment from JPEG bytes. Returns the EXIF
// data buffer (starting after the APP1 length field, at the "Exif\0\0" // data buffer (starting after the APP1 length field, at the "Exif\0\0"
// header) or undefined if no APP1 marker is found. // header) or undefined if no APP1 marker is found.
@@ -98,29 +149,24 @@ const extractImageMetadata = (
} }
}; };
// Read a file's original bytes through the library's content cache and extract
// its embedded image metadata. The bytes come from `photo.original()` — the
// same on-disk cache the rest of the library fills — rather than a fresh
// per-call download to a throwaway temp file.
const extractExif = async ( const extractExif = async (
photo: Photo, client: Client,
file: EnteFile,
): Promise<Record<string, unknown> | undefined> => { ): Promise<Record<string, unknown> | undefined> => {
const tmpDir = mkdtempSync(join(tmpdir(), "quak-exif-"));
try { try {
const { path } = await photo.original(); const origPath = join(tmpDir, "original");
const fileBytes = new Uint8Array(readFileSync(path)); await client.downloadFile(file, origPath);
const fileBytes = new Uint8Array(readFileSync(origPath));
return extractImageMetadata(fileBytes); return extractImageMetadata(fileBytes);
} catch { } catch {
return undefined; return undefined;
} finally {
rmSync(tmpDir, { recursive: true, force: true });
} }
}; };
// Dump every decrypted metadata layer the account holds into a directory tree
// of plain JSON: account, per-collection, and per-file records including the
// private and public magic metadata and (by default) the ML data. Collections
// and files are enumerated from the library's cache rather than a fresh server
// scan; the ML fetch and EXIF extraction are unchanged.
export const runMetadataBackup = async ( export const runMetadataBackup = async (
lib: Library,
client: Client, client: Client,
outDir: string, outDir: string,
opts?: MetadataBackupOptions, opts?: MetadataBackupOptions,
@@ -138,19 +184,13 @@ export const runMetadataBackup = async (
); );
log("Fetching collections..."); log("Fetching collections...");
const collections = await client.listCollections();
// Enumerate through the library's read surface. Each album carries its const allFiles: { file: EnteFile; colDirName: string }[] = [];
// photos, but the full decrypted `Collection`/`EnteFile` records (with the
// magic-metadata layers this dump exists to preserve) come from the
// library's by-id accessors.
const allFiles: { file: EnteFile; photo: Photo; colDirName: string }[] = [];
const fileKeys = new Map<number, Uint8Array>(); const fileKeys = new Map<number, Uint8Array>();
const seenFileIDs = new Set<number>(); const seenFileIDs = new Set<number>();
for (const album of lib.albums.list()) { for (const col of collections) {
const col = lib.getCollection(album.collectionID);
if (!col) continue;
const dirName = `${col.id}-${sanitizePath(col.name || "unnamed")}`; const dirName = `${col.id}-${sanitizePath(col.name || "unnamed")}`;
const colDir = join(outDir, "collections", dirName); const colDir = join(outDir, "collections", dirName);
mkdirSync(colDir, { recursive: true }); mkdirSync(colDir, { recursive: true });
@@ -175,13 +215,11 @@ export const runMetadataBackup = async (
); );
log(`[${col.name}] Fetching files...`); log(`[${col.name}] Fetching files...`);
const photos = album.photos.list(); const files = await client.listFiles(col.id, col.key);
log(`[${col.name}] ${photos.length} file(s)`); log(`[${col.name}] ${files.length} file(s)`);
for (const photo of photos) { for (const file of files) {
const file = lib.getFile(col.id, photo.fileID); allFiles.push({ file, colDirName: dirName });
if (!file) continue;
allFiles.push({ file, photo, colDirName: dirName });
if (!seenFileIDs.has(file.id)) { if (!seenFileIDs.has(file.id)) {
fileKeys.set(file.id, file.key); fileKeys.set(file.id, file.key);
seenFileIDs.add(file.id); seenFileIDs.add(file.id);
@@ -190,15 +228,15 @@ export const runMetadataBackup = async (
} }
log("Fetching ML data (face detections, CLIP embeddings)..."); log("Fetching ML data (face detections, CLIP embeddings)...");
const mlDataMap = await fetchMLData( const mlDataMap = await fetchMLDataForFiles(
client.getApiClient(), client,
[...fileKeys.keys()], [...fileKeys.keys()],
fileKeys, fileKeys,
); );
log(`Got ML data for ${mlDataMap.size} file(s)`); log(`Got ML data for ${mlDataMap.size} file(s)`);
const writtenFileIDs = new Set<number>(); const writtenFileIDs = new Set<number>();
for (const { file, photo, colDirName } of allFiles) { for (const { file, colDirName } of allFiles) {
const colDir = join(outDir, "collections", colDirName); const colDir = join(outDir, "collections", colDirName);
const fileMeta: Record<string, unknown> = { const fileMeta: Record<string, unknown> = {
@@ -217,7 +255,7 @@ export const runMetadataBackup = async (
if (wantExif && !writtenFileIDs.has(file.id)) { if (wantExif && !writtenFileIDs.has(file.id)) {
log(`[${file.metadata.title}] Extracting EXIF...`); log(`[${file.metadata.title}] Extracting EXIF...`);
const exifData = await extractExif(photo); const exifData = await extractExif(client, file);
if (exifData) fileMeta.imageMetadata = exifData; if (exifData) fileMeta.imageMetadata = exifData;
} }
writtenFileIDs.add(file.id); writtenFileIDs.add(file.id);
-93
View File
@@ -1,93 +0,0 @@
// Fetch and decrypt Ente's per-file machine-learning data ("magic" search
// data: face detections + CLIP embeddings).
//
// The data lives behind `/files/data/fetch` with `type: "mldata"`. Each entry
// comes back encrypted under the file's own key and gzipped; decrypting and
// gunzipping yields the JSON payload
// `{ face: { faces: [...] }, clip: { embedding } }`. Ente caps a request at 200
// ids, so `fetchMLData` batches for callers that want many at once while
// `fetchMLDataBatch` is the single-request unit the library submits to its
// request pool.
import { gunzipSync } from "node:zlib";
import type { ApiClient } from "./api/client.js";
import { decryptBlob, fromBase64 } from "./crypto/index.js";
// The most ids one `/files/data/fetch` request may carry.
export const MLDATA_BATCH_SIZE = 200;
// The decrypted, gunzipped per-file payload. Its concrete shape is Ente's; the
// store keeps the whole object verbatim and each consumer reads the fields it
// needs, so it stays an open record rather than a fixed interface.
export type MLData = Record<string, unknown>;
interface RawRemoteFileData {
fileID: number;
encryptedData: string;
decryptionHeader: string;
updatedAt?: number;
}
// Decrypt one entry with its file key and gunzip the JSON payload. Returns
// undefined when the key is unknown or the entry does not decrypt/parse, so one
// corrupt file never fails a whole batch.
const decodeEntry = (
entry: RawRemoteFileData,
key: Uint8Array | undefined,
): MLData | undefined => {
if (!key) return undefined;
try {
const decrypted = decryptBlob(
fromBase64(entry.encryptedData),
fromBase64(entry.decryptionHeader),
key,
);
const json = gunzipSync(Buffer.from(decrypted)).toString("utf-8");
return JSON.parse(json) as MLData;
} catch {
return undefined;
}
};
// Fetch ML data for up to `MLDATA_BATCH_SIZE` ids in a single request. This is
// the unit the request pools schedule; callers with more ids split them into
// batches and submit each batch to the pool.
export const fetchMLDataBatch = async (
api: ApiClient,
fileIDs: number[],
fileKeys: Map<number, Uint8Array>,
): Promise<Map<number, MLData>> => {
const { data } = await api.postJSON<{ data: RawRemoteFileData[] }>(
"/files/data/fetch",
{ type: "mldata", fileIDs },
);
const result = new Map<number, MLData>();
for (const entry of data ?? []) {
const payload = decodeEntry(entry, fileKeys.get(entry.fileID));
if (payload) result.set(entry.fileID, payload);
}
return result;
};
// Fetch ML data for arbitrarily many ids, batching at `MLDATA_BATCH_SIZE`. Used
// by the one-shot metadata backup; the library fetches through its request pool
// with `fetchMLDataBatch` instead.
export const fetchMLData = async (
api: ApiClient,
fileIDs: number[],
fileKeys: Map<number, Uint8Array>,
): Promise<Map<number, MLData>> => {
const result = new Map<number, MLData>();
for (let i = 0; i < fileIDs.length; i += MLDATA_BATCH_SIZE) {
const batch = fileIDs.slice(i, i + MLDATA_BATCH_SIZE);
for (const [id, payload] of await fetchMLDataBatch(
api,
batch,
fileKeys,
)) {
result.set(id, payload);
}
}
return result;
};
+71 -128
View File
@@ -2,10 +2,13 @@ import { createHash } from "node:crypto";
import { readFileSync } from "node:fs"; import { readFileSync } from "node:fs";
import * as jpeg from "jpeg-js"; import * as jpeg from "jpeg-js";
import type { Client } from "./client.js"; import type { Client } from "./client.js";
import type { Library } from "./library/index.js";
import { ApiError } from "./api/client.js"; import { ApiError } from "./api/client.js";
import { encryptBlob, toBase64 } from "./crypto/index.js"; import { encryptBlob, toBase64 } from "./crypto/index.js";
import { downloadFile } from "./download/index.js";
import type { EnteFile } from "./model/types.js"; import type { EnteFile } from "./model/types.js";
import { mkdtempSync, rmSync } from "node:fs";
import { join } from "node:path";
import { tmpdir } from "node:os";
const THUMB_MAX_DIMENSION = 720; const THUMB_MAX_DIMENSION = 720;
const THUMB_JPEG_QUALITY = 50; const THUMB_JPEG_QUALITY = 50;
@@ -17,51 +20,34 @@ export interface MissingThumbnailInfo {
reason: string; reason: string;
} }
// Three outcomes, not two. "fixed": a thumbnail was generated and uploaded.
// "failed": something went wrong (download, encode, upload) and the file still
// has no thumbnail. "skipped": the file is a format this helper cannot
// regenerate — a video, or an image that is not a baseline JPEG. Skipped is a
// deliberate, expected outcome, not an error (issue #17): the repair path is
// JPEG-only because `jpeg-js` is, and a PNG or HEIC is left for a format-aware
// tool rather than reported as a failure.
export type ThumbnailFixStatus = "fixed" | "skipped" | "failed";
export interface ThumbnailFixResult { export interface ThumbnailFixResult {
fileID: number; fileID: number;
title: string; title: string;
collection: string; collection: string;
status: ThumbnailFixStatus; success: boolean;
// Why the file was skipped or failed; unset when it was fixed. error?: string;
reason?: string;
} }
export type ProgressCallback = (message: string) => void; export type ProgressCallback = (message: string) => void;
// Enumerate every file the library knows about, newest album first, each file
// once, and report those whose server-side thumbnail is missing. "Missing" is
// only two answers: an empty body, or a 404. Any other error reaching this
// point has already exhausted its retries — a failing server, a dropped
// connection, a deadline — and says nothing about whether the thumbnail
// exists, so it is logged and the file is left unreported. That distinction is
// what stops `fix-missing-thumbnails` from regenerating and uploading over
// thumbnails that were fine all along while the CDN was briefly returning 500s.
export const listMissingThumbnails = async ( export const listMissingThumbnails = async (
lib: Library,
client: Client, client: Client,
onProgress?: ProgressCallback, onProgress?: ProgressCallback,
): Promise<MissingThumbnailInfo[]> => { ): Promise<MissingThumbnailInfo[]> => {
const log = onProgress ?? (() => {}); const log = onProgress ?? (() => {});
const api = client.getApiClient();
const missing: MissingThumbnailInfo[] = []; const missing: MissingThumbnailInfo[] = [];
const seen = new Set<number>(); const seen = new Set<number>();
for (const album of lib.albums.list()) { const collections = await client.listCollections();
log(`[${album.name}] Checking thumbnails...`); for (const col of collections) {
for (const photo of album.photos.list()) { log(`[${col.name}] Checking thumbnails...`);
if (seen.has(photo.fileID)) continue; const files = await client.listFiles(col.id, col.key);
seen.add(photo.fileID); for (const file of files) {
if (seen.has(file.id)) continue;
seen.add(file.id);
try { try {
const stream = await api.getThumbnailStream(photo.fileID); const api = client.getApiClient();
const stream = await api.getThumbnailStream(file.id);
const reader = stream.getReader(); const reader = stream.getReader();
let totalBytes = 0; let totalBytes = 0;
for (;;) { for (;;) {
@@ -71,23 +57,35 @@ export const listMissingThumbnails = async (
} }
if (totalBytes === 0) { if (totalBytes === 0) {
missing.push({ missing.push({
fileID: photo.fileID, fileID: file.id,
title: photo.title, title: file.metadata.title,
collection: album.name, collection: col.name,
reason: "empty thumbnail (0 bytes)", reason: "empty thumbnail (0 bytes)",
}); });
} }
} catch (err) { } catch (err) {
// A 404 is the server stating the thumbnail is not there:
// that, and an empty body, are the only two answers that mean
// "missing". Anything else reaching this point is a failure
// that already exhausted its retries — a failing server, a
// dropped connection, a deadline — and says nothing about
// whether the thumbnail exists.
//
// The distinction is what stops `helper
// fix-missing-thumbnails` from downloading originals,
// regenerating thumbnails and uploading them over thumbnails
// that were fine all along, because the CDN was briefly
// returning 500s while this ran.
if (err instanceof ApiError && err.status === 404) { if (err instanceof ApiError && err.status === 404) {
missing.push({ missing.push({
fileID: photo.fileID, fileID: file.id,
title: photo.title, title: file.metadata.title,
collection: album.name, collection: col.name,
reason: "thumbnail not found (HTTP 404)", reason: "thumbnail not found (HTTP 404)",
}); });
} else { } else {
log( log(
`[${album.name}] Could not check ${photo.title}: ${err instanceof Error ? err.message : String(err)} (not reported as missing)`, `[${col.name}] Could not check ${file.metadata.title}: ${err instanceof Error ? err.message : String(err)} (not reported as missing)`,
); );
} }
} }
@@ -96,7 +94,7 @@ export const listMissingThumbnails = async (
return missing; return missing;
}; };
// Bilinear resize of an RGBA pixel buffer. // Bilinear resize of RGBA pixel buffer
const resizeRGBA = ( const resizeRGBA = (
src: Uint8Array, src: Uint8Array,
srcW: number, srcW: number,
@@ -163,33 +161,7 @@ const generateThumbnail = (fileBytes: Uint8Array): Uint8Array => {
return new Uint8Array(encoded.data); return new Uint8Array(encoded.data);
}; };
// A baseline/JFIF JPEG starts with the SOI marker 0xFFD8. `jpeg-js` decodes
// only JPEG, so this signature check is what separates a file the helper can
// regenerate from one it must skip: a PNG, HEIC, or the odd non-image byte
// stream all fail this and are reported as skipped rather than crashing the
// decoder into an opaque failure (issue #17).
const isJpeg = (bytes: Uint8Array): boolean =>
bytes.length >= 2 && bytes[0] === 0xff && bytes[1] === 0xd8;
// The reason a file cannot have a JPEG thumbnail regenerated for it from its
// metadata alone, before any bytes are fetched, or undefined when it might. A
// non-image (video, live photo) is unsupported outright; a still image still
// has to be checked against its actual bytes once downloaded.
const unsupportedByType = (file: EnteFile): string | undefined => {
if (file.metadata.fileType !== "image") {
return `unsupported file type: ${file.metadata.fileType} (only JPEG images can be regenerated)`;
}
return undefined;
};
// Regenerate and upload a thumbnail for each requested file. Originals are read
// through the library's content cache (`photo.original()`); the generated
// thumbnail is JPEG-encoded, encrypted under the file's own key, and registered
// with the server — the encrypt-and-upload path is unchanged. Each file is
// resolved to one outcome (fixed / skipped / failed) and a failure on one file
// never stops the others.
export const fixMissingThumbnails = async ( export const fixMissingThumbnails = async (
lib: Library,
client: Client, client: Client,
fileIDs: number[], fileIDs: number[],
onProgress?: ProgressCallback, onProgress?: ProgressCallback,
@@ -198,24 +170,19 @@ export const fixMissingThumbnails = async (
const results: ThumbnailFixResult[] = []; const results: ThumbnailFixResult[] = [];
const api = client.getApiClient(); const api = client.getApiClient();
// Resolve each requested fileID to its file record and owning album by const collections = await client.listCollections();
// enumerating the library, each file taken from the first album that holds
// it. The raw `EnteFile` carries the per-file key the thumbnail is
// encrypted under, which the projected records deliberately do not.
const wanted = new Set(fileIDs);
const fileMap = new Map< const fileMap = new Map<
number, number,
{ file: EnteFile; collectionName: string } { file: EnteFile; collectionName: string }
>(); >();
for (const album of lib.albums.list()) {
for (const photo of album.photos.list()) { for (const col of collections) {
if (!wanted.has(photo.fileID) || fileMap.has(photo.fileID)) const files = await client.listFiles(col.id, col.key);
continue; for (const file of files) {
const file = lib.getFile(album.collectionID, photo.fileID); if (fileIDs.includes(file.id) && !fileMap.has(file.id)) {
if (file) { fileMap.set(file.id, {
fileMap.set(photo.fileID, {
file, file,
collectionName: album.name, collectionName: col.name,
}); });
} }
} }
@@ -228,61 +195,33 @@ export const fixMissingThumbnails = async (
fileID, fileID,
title: "unknown", title: "unknown",
collection: "unknown", collection: "unknown",
status: "failed", success: false,
reason: "file not found in any collection", error: "file not found in any collection",
}); });
continue; continue;
} }
const { file, collectionName } = entry; const { file, collectionName } = entry;
const title = file.metadata.title; const tmpDir = mkdtempSync(join(tmpdir(), "quak-thumb-"));
const typeReason = unsupportedByType(file);
if (typeReason) {
log(`[${collectionName}] Skipping ${title}: ${typeReason}`);
results.push({
fileID,
title,
collection: collectionName,
status: "skipped",
reason: typeReason,
});
continue;
}
try { try {
const photo = lib.photos.byID({ fileID }); log(
if (!photo) { `[${collectionName}] Downloading ${file.metadata.title} for thumbnail generation...`,
throw new Error("file not present in the library cache"); );
} const origPath = join(tmpDir, "original");
await downloadFile(api, file, origPath);
log( log(
`[${collectionName}] Downloading ${title} for thumbnail generation...`, `[${collectionName}] Generating thumbnail for ${file.metadata.title}...`,
); );
const { path } = await photo.original(); const fileBytes = readFileSync(origPath);
const fileBytes = new Uint8Array(readFileSync(path)); const thumbJpeg = generateThumbnail(new Uint8Array(fileBytes));
if (!isJpeg(fileBytes)) {
const reason =
"unsupported image format (only baseline JPEG can be regenerated)";
log(`[${collectionName}] Skipping ${title}: ${reason}`);
results.push({
fileID,
title,
collection: collectionName,
status: "skipped",
reason,
});
continue;
}
log(`[${collectionName}] Generating thumbnail for ${title}...`);
const thumbJpeg = generateThumbnail(fileBytes);
log( log(
`[${collectionName}] Encrypting and uploading thumbnail (${thumbJpeg.length} bytes)...`, `[${collectionName}] Encrypting and uploading thumbnail (${thumbJpeg.length} bytes)...`,
); );
const { header, ciphertext } = encryptBlob(thumbJpeg, file.key); const { header, ciphertext } = encryptBlob(thumbJpeg, file.key);
const md5 = createHash("md5").update(ciphertext).digest("base64"); const md5 = createHash("md5").update(ciphertext).digest("base64");
const { objectKey, url } = await api.getUploadURL( const { objectKey, url } = await api.getUploadURL(
ciphertext.length, ciphertext.length,
@@ -291,24 +230,28 @@ export const fixMissingThumbnails = async (
await api.putFile(url, ciphertext); await api.putFile(url, ciphertext);
await api.updateThumbnail(file.id, objectKey, toBase64(header)); await api.updateThumbnail(file.id, objectKey, toBase64(header));
log(`[${collectionName}] Thumbnail uploaded for ${title}`);
results.push({
fileID,
title,
collection: collectionName,
status: "fixed",
});
} catch (err) {
log( log(
`[${collectionName}] FAILED ${title}: ${err instanceof Error ? err.message : err}`, `[${collectionName}] Thumbnail uploaded for ${file.metadata.title}`,
); );
results.push({ results.push({
fileID, fileID,
title, title: file.metadata.title,
collection: collectionName, collection: collectionName,
status: "failed", success: true,
reason: err instanceof Error ? err.message : String(err),
}); });
} catch (err) {
log(
`[${collectionName}] FAILED ${file.metadata.title}: ${err instanceof Error ? err.message : err}`,
);
results.push({
fileID,
title: file.metadata.title,
collection: collectionName,
success: false,
error: err instanceof Error ? err.message : String(err),
});
} finally {
rmSync(tmpDir, { recursive: true, force: true });
} }
} }
+386 -399
View File
@@ -1,204 +1,368 @@
/** /**
* Tests for the `quak backup` logic, now built on the library API (issue #51). * Tests for the `quak backup` command's core logic.
* *
* `lib.backup({ downloadDirectory })` refreshes the library, fetches each * `quak backup <dir>` downloads every file from every collection into
* pending file's original through the content cache/pools, and materialises the * a local directory tree:
* unchanged on-disk layout:
* *
* <downloadDirectory>/ * <dir>/
* originals/ * <collection-name>/
* <fileID>.<ext> the decrypted bytes ("present means complete") * <file-title>
* <fileID>.json per-file metadata sidecar (rebuilt each run) * <file-title>
* collections/ * <collection-name>/
* <name>/<title> symlink into ../originals (rebuilt each run) * ...
* <name>.json per-collection metadata (rebuilt each run) * metadata.json (all decrypted collection + file metadata)
* failures.json durable ledger of unresolved failures
* *
* The properties that distinguish backup from a naive download loop, and that * The backup command has two properties that distinguish it from a naive
* these tests lock down: * "download everything" loop:
* *
* 1. Present-means-complete: an original already on disk is not re-fetched, so * 1. **Skip existing files.** If `<dir>/<collection>/<title>` already
* runs are idempotent and interrupted runs resume. * exists on disk and its size matches the decrypted content length
* 2. Per-file resilience: a download failure or a symlink failure is recorded * recorded in metadata.json from a prior run, the file is not
* and the run continues (issue #8); the derived symlink/JSON views are * re-downloaded. This makes interrupted backups resumable and
* rebuilt from the model every run. * incremental runs fast.
* 3. A durable `failures.json` records each unresolved failure's classification,
* attempt count, and last-tried time; the exit code (result.failed) is
* non-zero while any failure remains and clears once every one is resolved.
* *
* The cache and download layers are covered elsewhere (content.test.ts, * 2. **Never crash on a single file failure.** If a file download or
* download tests); here a mock library client and a stand-in content source * decryption fails, the error is logged and the backup continues
* drive the backup logic with no crypto and no network. * with the next file. At the end, the exit code is non-zero if any
* files failed, and the summary lists them. The Ente first-party
* CLI crashes entirely when a single file can't be retrieved,
* which defeats the purpose of a backup tool.
*
* These tests exercise the backup logic (in src/backup.ts) using the
* same mock server from the Client usage tests. The CLI binary itself
* is a thin wrapper around this module.
*/ */
import { import {
existsSync, existsSync,
lstatSync, lstatSync,
mkdirSync,
mkdtempSync, mkdtempSync,
readFileSync, readFileSync,
readlinkSync, readlinkSync,
rmSync, rmSync,
writeFileSync,
} from "node:fs"; } from "node:fs";
import { join } from "node:path"; import { join } from "node:path";
import { tmpdir } from "node:os"; import { tmpdir } from "node:os";
import { describe, it, expect, beforeEach, afterEach } from "vitest"; import sodium from "libsodium-wrappers-sumo";
import { SRP, SrpServer } from "fast-srp-hap";
import { beforeAll, afterAll, describe, expect, it } from "vitest";
import {
init,
toBase64,
deriveKEK,
deriveLoginSubkey,
} from "../../src/crypto/index.js";
import { Client } from "../../src/client.js";
import { runBackup } from "../../src/backup.js";
import type { KeyAttributes } from "../../src/auth/types.js";
import { Library } from "../../src/library/index.js"; // ---------------------------------------------------------------------------
import type { ContentSource } from "../../src/library/content.js"; // Mock server (condensed from usage.test.ts)
import type { CollectionsPage, FilesPage } from "../../src/client.js"; // ---------------------------------------------------------------------------
import type { Collection, EnteFile } from "../../src/model/types.js";
const USER_ID = 42; const TEST_EMAIL = "backup@example.com";
const TEST_PASSWORD = "backuppass";
const TEST_OPS = 2;
const TEST_MEM = 64 * 1024 * 1024;
// Decrypted-byte length each stub original writes, keyed by fileID. interface MockState {
const SIZE_BY_ID: Record<number, number> = { 100: 3000, 101: 2000, 200: 1500 }; verifier: Buffer;
srpAttributes: Record<string, unknown>;
const collection = (id: number, name: string): Collection => ({ keyAttributes: KeyAttributes;
id, encryptedToken: string;
ownerID: USER_ID, collections: Record<string, unknown>[];
key: new Uint8Array([id & 0xff]), files: Record<
name, number,
type: "album", { raw: Record<string, unknown>; plaintext: Uint8Array }
updationTime: 1, >;
isShared: false,
});
const file = (id: number, collectionID: number, title: string): EnteFile => ({
id,
collectionID,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
metadata: {
title,
fileType: "image",
creationTime: 1,
modificationTime: 1,
},
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime: 1,
});
// A metadata-only client: two albums, three files, served once. No ML.
class MockClient {
private served = false;
whoami(): { email: string; userID: number } {
return { email: "backup@example.com", userID: USER_ID };
}
async collectionsSince(): Promise<CollectionsPage> {
if (this.served) return { collections: [], deleted: [], cursor: 1 };
this.served = true;
return {
collections: [collection(1, "Vacation"), collection(2, "Work")],
deleted: [],
cursor: 1,
};
}
async filesSince(args: { collectionID: number }): Promise<FilesPage> {
const files =
args.collectionID === 1
? [file(100, 1, "beach.jpg"), file(101, 1, "sunset.jpg")]
: args.collectionID === 2
? [file(200, 2, "diagram.png")]
: [];
return { files, deleted: [], cursor: 1 };
}
} }
// A content source that writes byte buffers of the expected length and can be let mock: MockState;
// told to fail one fileID's original, to exercise per-file resilience. let testDir: string;
interface StubSource extends ContentSource {
failID?: number;
failThumbID?: number;
originalCalls: number;
}
const stubSource = (): StubSource => { const buildMock = async (): Promise<MockState> => {
const s: StubSource = { const kekSalt = sodium.randombytes_buf(sodium.crypto_pwhash_SALTBYTES);
originalCalls: 0, const kek = await deriveKEK(TEST_PASSWORD, kekSalt, TEST_OPS, TEST_MEM);
original: async ({ file: f, destination }) => { const loginSubKeyBytes = deriveLoginSubkey(kek);
s.originalCalls++;
if (s.failID === f.id) throw new Error("HTTP 500 from server");
const size = SIZE_BY_ID[f.id] ?? 10;
writeFileSync(destination, Buffer.alloc(size));
return { bytesWritten: size };
},
thumbnail: async ({ file: f, destination }) => {
if (s.failThumbID === f.id) throw new Error("HTTP 500 from server");
writeFileSync(destination, Buffer.alloc(5));
return { bytesWritten: 5 };
},
};
return s;
};
let root: string; const srpUserID = "backup-srp";
const srpSalt = sodium.randombytes_buf(16);
const openLibrary = (source: ContentSource): Promise<Library> => const verifier = SRP.computeVerifier(
Library.open({ SRP.params["4096"],
client: new MockClient(), Buffer.from(srpSalt),
cacheDirectory: join(root, "cache"), Buffer.from(srpUserID),
contentSource: source, Buffer.from(loginSubKeyBytes),
refreshIntervalSeconds: 3600,
// These tests count exact fetches; the background precache (#48) would
// add its own, so it is off here (it is covered in precache.test.ts).
precacheThumbnails: false,
precacheOriginals: false,
});
const readLedger = (
outDir: string,
): { files: Record<string, Record<string, unknown>> } =>
JSON.parse(readFileSync(join(outDir, "failures.json"), "utf-8"));
// Write a durable ledger holding one prior failure, to exercise pruning of
// entries the current run cannot resolve.
const seedLedger = (outDir: string, fileID: number, title: string): void => {
mkdirSync(outDir, { recursive: true });
writeFileSync(
join(outDir, "failures.json"),
JSON.stringify({
version: 1,
files: {
[String(fileID)]: {
fileID,
title,
classification: "transient",
attempts: 1,
lastTriedAt: Date.now(),
error: "HTTP 500 from server",
},
},
}),
); );
const masterKey = sodium.randombytes_buf(32);
const keyNonce = sodium.randombytes_buf(sodium.crypto_secretbox_NONCEBYTES);
const encryptedKey = sodium.crypto_secretbox_easy(masterKey, keyNonce, kek);
const kp = sodium.crypto_box_keypair();
const skNonce = sodium.randombytes_buf(sodium.crypto_secretbox_NONCEBYTES);
const encSK = sodium.crypto_secretbox_easy(
kp.privateKey,
skNonce,
masterKey,
);
const tokenBytes = sodium.randombytes_buf(32);
const encToken = sodium.crypto_box_seal(tokenBytes, kp.publicKey);
const keyAttributes: KeyAttributes = {
kekSalt: toBase64(kekSalt),
encryptedKey: toBase64(encryptedKey),
keyDecryptionNonce: toBase64(keyNonce),
publicKey: toBase64(kp.publicKey),
encryptedSecretKey: toBase64(encSK),
secretKeyDecryptionNonce: toBase64(skNonce),
memLimit: TEST_MEM,
opsLimit: TEST_OPS,
};
const makeCollection = (id: number, name: string) => {
const ck = sodium.crypto_secretbox_keygen();
const ckN = sodium.randombytes_buf(sodium.crypto_secretbox_NONCEBYTES);
const encCK = sodium.crypto_secretbox_easy(ck, ckN, masterKey);
const nameBytes = new TextEncoder().encode(name);
const cnN = sodium.randombytes_buf(sodium.crypto_secretbox_NONCEBYTES);
const encCN = sodium.crypto_secretbox_easy(nameBytes, cnN, ck);
return {
raw: {
id,
owner: { id: 42 },
encryptedKey: toBase64(encCK),
keyDecryptionNonce: toBase64(ckN),
encryptedName: toBase64(encCN),
nameDecryptionNonce: toBase64(cnN),
type: "album",
updationTime: 1700000000000000,
},
key: ck,
};
};
const makeFile = (
id: number,
collKey: Uint8Array,
title: string,
plaintext: Uint8Array,
collID: number,
) => {
const fk = sodium.crypto_secretstream_xchacha20poly1305_keygen();
const fkN = sodium.randombytes_buf(sodium.crypto_secretbox_NONCEBYTES);
const encFK = sodium.crypto_secretbox_easy(fk, fkN, collKey);
const meta = JSON.stringify({
title,
fileType: 0,
creationTime: 1700000000000000,
modificationTime: 1700000000000000,
});
const metaPush =
sodium.crypto_secretstream_xchacha20poly1305_init_push(fk);
const encMeta = sodium.crypto_secretstream_xchacha20poly1305_push(
metaPush.state,
new TextEncoder().encode(meta),
null,
sodium.crypto_secretstream_xchacha20poly1305_TAG_FINAL,
);
const filePush =
sodium.crypto_secretstream_xchacha20poly1305_init_push(fk);
const encFile = sodium.crypto_secretstream_xchacha20poly1305_push(
filePush.state,
plaintext,
null,
sodium.crypto_secretstream_xchacha20poly1305_TAG_FINAL,
);
return {
raw: {
id,
collectionID: collID,
ownerID: 42,
encryptedKey: toBase64(encFK),
keyDecryptionNonce: toBase64(fkN),
metadata: {
encryptedData: toBase64(encMeta),
decryptionHeader: toBase64(metaPush.header),
},
file: { decryptionHeader: toBase64(filePush.header) },
thumbnail: {
decryptionHeader: toBase64(sodium.randombytes_buf(24)),
},
updationTime: 1700000000000000,
},
plaintext,
ciphertext: encFile,
};
};
const col1 = makeCollection(1, "Vacation");
const col2 = makeCollection(2, "Work");
const file1 = makeFile(
100,
col1.key,
"beach.jpg",
sodium.randombytes_buf(3000),
1,
);
const file2 = makeFile(
101,
col1.key,
"sunset.jpg",
sodium.randombytes_buf(2000),
1,
);
const file3 = makeFile(
200,
col2.key,
"diagram.png",
sodium.randombytes_buf(1500),
2,
);
return {
verifier,
srpAttributes: {
srpUserID,
srpSalt: toBase64(srpSalt),
memLimit: TEST_MEM,
opsLimit: TEST_OPS,
kekSalt: toBase64(kekSalt),
isEmailMFAEnabled: false,
},
keyAttributes,
encryptedToken: toBase64(encToken),
collections: [col1.raw, col2.raw],
files: {
1: {
raw: [file1.raw, file2.raw],
ciphertexts: { 100: file1.ciphertext, 101: file2.ciphertext },
},
2: { raw: [file3.raw], ciphertexts: { 200: file3.ciphertext } },
100: { plaintext: file1.plaintext },
101: { plaintext: file2.plaintext },
200: { plaintext: file3.plaintext },
} as Record<number, unknown>,
};
}; };
beforeEach(() => { const buildMockFetch = (m: MockState, opts?: { failFileID?: number }) => {
root = mkdtempSync(join(tmpdir(), "quak-backup-test-")); let srpServer: SrpServer;
return (async (
input: RequestInfo | URL,
init?: RequestInit,
): Promise<Response> => {
const url =
typeof input === "string"
? input
: input instanceof URL
? input.href
: input.url;
const parsed = new URL(url);
const path = parsed.pathname;
const json = (body: unknown) =>
new Response(JSON.stringify(body), {
status: 200,
headers: { "content-type": "application/json" },
});
if (path === "/users/srp/attributes")
return json({ attributes: m.srpAttributes });
if (path === "/users/srp/create-session") {
const body = JSON.parse(init?.body as string);
const serverKey = await SRP.genKey();
srpServer = new SrpServer(
SRP.params["4096"],
m.verifier,
serverKey,
);
const B = srpServer.computeB();
srpServer.setA(Buffer.from(body.srpA, "base64"));
return json({ sessionID: "s1", srpB: B.toString("base64") });
}
if (path === "/users/srp/verify-session") {
const body = JSON.parse(init?.body as string);
srpServer.checkM1(Buffer.from(body.srpM1, "base64"));
return json({
srpM2: srpServer.computeM2().toString("base64"),
id: 42,
keyAttributes: m.keyAttributes,
encryptedToken: m.encryptedToken,
});
}
if (path === "/collections/v2")
return json({ collections: m.collections });
if (path === "/collections/v2/diff") {
const collID = Number(parsed.searchParams.get("collectionID"));
const collData = m.files[collID] as { raw: unknown[] } | undefined;
return json({ diff: collData?.raw ?? [], hasMore: false });
}
// File download
if (url.includes("fileID=") || path.startsWith("/files/download/")) {
const fileID = Number(
parsed.searchParams.get("fileID") ?? path.split("/").pop(),
);
if (opts?.failFileID === fileID) {
return new Response("Internal Server Error", { status: 500 });
}
const collData = Object.values(m.files).find(
(v: unknown) =>
v &&
typeof v === "object" &&
"ciphertexts" in (v as Record<string, unknown>) &&
fileID in
(v as Record<string, Record<number, unknown>>)
.ciphertexts,
) as { ciphertexts: Record<number, Uint8Array> } | undefined;
if (collData) {
return new Response(collData.ciphertexts[fileID], {
status: 200,
});
}
return new Response("not found", { status: 404 });
}
return new Response("not found", { status: 404 });
}) as typeof globalThis.fetch;
};
// ---------------------------------------------------------------------------
// Setup / teardown
// ---------------------------------------------------------------------------
beforeAll(async () => {
await init();
await sodium.ready;
mock = await buildMock();
testDir = mkdtempSync(join(tmpdir(), "quak-backup-test-"));
}); });
afterEach(() => { afterAll(() => {
if (root && existsSync(root)) if (testDir && existsSync(testDir))
rmSync(root, { recursive: true, force: true }); rmSync(testDir, { recursive: true, force: true });
}); });
describe("lib.backup", () => { // ---------------------------------------------------------------------------
it("throws before any network when no downloadDirectory is given", async () => { // Tests
const source = stubSource(); // ---------------------------------------------------------------------------
const lib = await openLibrary(source);
await expect(lib.backup()).rejects.toThrow(/downloadDirectory/i);
expect(source.originalCalls).toBe(0);
lib.close();
});
it("writes the expected on-disk layout for every file", async () => { describe("quak backup", () => {
const source = stubSource(); it("downloads all files organized by collection name", async () => {
const lib = await openLibrary(source); const outDir = join(testDir, "full-backup");
const outDir = join(root, "backup"); const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMockFetch(mock) },
});
const result = await lib.backup({ downloadDirectory: outDir }); const result = await runBackup(client, outDir);
expect(result.totalFiles).toBe(3); expect(result.totalFiles).toBe(3);
expect(result.downloaded).toBe(3); expect(result.downloaded).toBe(3);
@@ -206,7 +370,7 @@ describe("lib.backup", () => {
expect(result.failed).toBe(0); expect(result.failed).toBe(0);
expect(result.errors).toEqual([]); expect(result.errors).toEqual([]);
// Originals under originals/<fileID>.<ext>. // Originals are under <outDir>/originals/<fileID>.<ext>
expect(readFileSync(join(outDir, "originals", "100.jpg")).length).toBe( expect(readFileSync(join(outDir, "originals", "100.jpg")).length).toBe(
3000, 3000,
); );
@@ -217,57 +381,55 @@ describe("lib.backup", () => {
1500, 1500,
); );
// Per-file metadata sidecar. // Collection dirs under collections/ contain symlinks to originals
const sidecar = JSON.parse( const beachLink = join(outDir, "collections", "Vacation", "beach.jpg");
readFileSync(join(outDir, "originals", "100.json"), "utf-8"), expect(lstatSync(beachLink).isSymbolicLink()).toBe(true);
); expect(readlinkSync(beachLink)).toContain("originals");
expect(sidecar.id).toBe(100); expect(readFileSync(beachLink).length).toBe(3000);
expect(sidecar.metadata.title).toBe("beach.jpg");
// Collection dirs contain symlinks into ../originals.
const beach = join(outDir, "collections", "Vacation", "beach.jpg");
expect(lstatSync(beach).isSymbolicLink()).toBe(true);
expect(readlinkSync(beach)).toContain("originals");
expect(readFileSync(beach).length).toBe(3000);
// Per-collection metadata JSON.
const vacation = JSON.parse(
readFileSync(join(outDir, "collections", "Vacation.json"), "utf-8"),
);
expect(vacation.name).toBe("Vacation");
expect(vacation.files.length).toBe(2);
expect(vacation.files[0].metadata.title).toBeDefined();
// A clean run leaves no failure ledger behind.
expect(existsSync(join(outDir, "failures.json"))).toBe(false);
lib.close();
}); });
it("is an idempotent no-op when every original is already present", async () => { it("skips files that already exist on disk with matching size", async () => {
const source = stubSource(); const outDir = join(testDir, "incremental");
const lib = await openLibrary(source); const client = await Client.login({
const outDir = join(root, "backup"); email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMockFetch(mock) },
});
const first = await lib.backup({ downloadDirectory: outDir }); // First run: download everything
const first = await runBackup(client, outDir);
expect(first.downloaded).toBe(3); expect(first.downloaded).toBe(3);
const callsAfterFirst = source.originalCalls;
const second = await lib.backup({ downloadDirectory: outDir }); // Second run: everything should be skipped
const second = await runBackup(client, outDir);
expect(second.downloaded).toBe(0); expect(second.downloaded).toBe(0);
expect(second.skipped).toBe(3); expect(second.skipped).toBe(3);
expect(second.failed).toBe(0); expect(second.failed).toBe(0);
// A present original is neither fetched nor copied again.
expect(source.originalCalls).toBe(callsAfterFirst);
lib.close();
}); });
it("continues past a download failure and records it in failures.json", async () => { it("continues after a single file download failure", async () => {
const source = stubSource(); // File 101 (sunset.jpg) will return HTTP 500. The other two
source.failID = 101; // files must still download. The result must report the failure
const lib = await openLibrary(source); // without throwing.
const outDir = join(root, "backup"); //
// A 500 is retryable, so this file now costs several requests before
// it is given up on — that is the point of the retry policy, and
// `runBackup`'s own resilience is unchanged by it: the retry lives
// strictly below this loop, and an exhausted file is still logged,
// counted, and stepped over rather than aborting the run. The
// injected `sleep` is what keeps the suite from actually waiting out
// the backoff.
const outDir = join(testDir, "partial-failure");
const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: {
fetch: buildMockFetch(mock, { failFileID: 101 }),
retry: { sleep: () => Promise.resolve(), random: () => 0 },
},
});
const result = await lib.backup({ downloadDirectory: outDir }); const result = await runBackup(client, outDir);
expect(result.totalFiles).toBe(3); expect(result.totalFiles).toBe(3);
expect(result.downloaded).toBe(2); expect(result.downloaded).toBe(2);
@@ -276,207 +438,32 @@ describe("lib.backup", () => {
expect(result.errors[0]!.fileID).toBe(101); expect(result.errors[0]!.fileID).toBe(101);
expect(result.errors[0]!.title).toBe("sunset.jpg"); expect(result.errors[0]!.title).toBe("sunset.jpg");
// The two good files are on disk; the failed one is not. // The two successful originals are on disk
expect(existsSync(join(outDir, "originals", "100.jpg"))).toBe(true); expect(existsSync(join(outDir, "originals", "100.jpg"))).toBe(true);
expect(existsSync(join(outDir, "originals", "200.png"))).toBe(true); expect(existsSync(join(outDir, "originals", "200.png"))).toBe(true);
// The failed file has no original and no symlink
expect(existsSync(join(outDir, "originals", "101.jpg"))).toBe(false); expect(existsSync(join(outDir, "originals", "101.jpg"))).toBe(false);
expect( expect(
existsSync(join(outDir, "collections", "Vacation", "sunset.jpg")), existsSync(join(outDir, "collections", "Vacation", "sunset.jpg")),
).toBe(false); ).toBe(false);
// Durable ledger with classification, attempts, last-tried.
const ledger = readLedger(outDir);
const entry = ledger.files["101"]!;
expect(entry.attempts).toBe(1);
expect(entry.classification).toBeDefined();
expect(typeof entry.lastTriedAt).toBe("number");
lib.close();
}); });
it("increments the attempt count across runs and clears the ledger once resolved", async () => { it("writes per-collection JSON metadata", async () => {
const source = stubSource(); const outDir = join(testDir, "metadata-check");
source.failID = 101; const client = await Client.login({
const lib = await openLibrary(source); email: TEST_EMAIL,
const outDir = join(root, "backup"); password: TEST_PASSWORD,
apiOptions: { fetch: buildMockFetch(mock) },
const r1 = await lib.backup({ downloadDirectory: outDir });
expect(r1.failed).toBe(1);
expect(readLedger(outDir).files["101"]!.attempts).toBe(1);
// Second run: the two good files are present, only 101 is retried.
const r2 = await lib.backup({ downloadDirectory: outDir });
expect(r2.failed).toBe(1);
expect(r2.skipped).toBe(2);
expect(readLedger(outDir).files["101"]!.attempts).toBe(2);
// Resume with a healthy source: 101 downloads, the rest are skipped.
source.failID = undefined;
const r3 = await lib.backup({ downloadDirectory: outDir });
expect(r3.failed).toBe(0);
expect(r3.skipped).toBe(2);
expect(existsSync(join(outDir, "originals", "101.jpg"))).toBe(true);
// A ledger with no remaining failures is removed.
expect(existsSync(join(outDir, "failures.json"))).toBe(false);
lib.close();
});
it("does not abort when a symlink cannot be created (issue #8)", async () => {
const source = stubSource();
const lib = await openLibrary(source);
const outDir = join(root, "backup");
// Occupy beach.jpg's symlink path with a directory so symlink creation
// fails for that one file.
mkdirSync(join(outDir, "collections", "Vacation", "beach.jpg"), {
recursive: true,
}); });
const result = await lib.backup({ downloadDirectory: outDir }); await runBackup(client, outDir);
// Every original still downloads despite the symlink failure. // Each collection gets a <name>.json next to its image dir
expect(existsSync(join(outDir, "originals", "100.jpg"))).toBe(true); const vacationMeta = join(outDir, "collections", "Vacation.json");
expect(existsSync(join(outDir, "originals", "200.png"))).toBe(true); expect(existsSync(vacationMeta)).toBe(true);
// The other symlinks are still built. const meta = JSON.parse(readFileSync(vacationMeta, "utf-8"));
expect( expect(meta.name).toBe("Vacation");
lstatSync( expect(meta.files.length).toBeGreaterThan(0);
join(outDir, "collections", "Vacation", "sunset.jpg"), expect(meta.files[0].metadata.title).toBeDefined();
).isSymbolicLink(),
).toBe(true);
expect(
lstatSync(
join(outDir, "collections", "Work", "diagram.png"),
).isSymbolicLink(),
).toBe(true);
// The symlink failure is recorded, not thrown.
expect(result.failed).toBeGreaterThanOrEqual(1);
const err = result.errors.find((e) => e.fileID === 100);
expect(err).toBeDefined();
expect(err!.collection).toBe("Vacation");
expect(readLedger(outDir).files["100"]).toBeDefined();
lib.close();
});
it("rebuilds a stale sidecar and a missing symlink on a later run", async () => {
const source = stubSource();
const lib = await openLibrary(source);
const outDir = join(root, "backup");
await lib.backup({ downloadDirectory: outDir });
// Corrupt a sidecar and delete a symlink between runs.
writeFileSync(join(outDir, "originals", "100.json"), "not json");
rmSync(join(outDir, "collections", "Vacation", "beach.jpg"));
const result = await lib.backup({ downloadDirectory: outDir });
expect(result.failed).toBe(0);
// The derived views are repaired from the model.
const sidecar = JSON.parse(
readFileSync(join(outDir, "originals", "100.json"), "utf-8"),
);
expect(sidecar.metadata.title).toBe("beach.jpg");
expect(
lstatSync(
join(outDir, "collections", "Vacation", "beach.jpg"),
).isSymbolicLink(),
).toBe(true);
lib.close();
});
it("backs up only the named albums when onlyAlbumNames is given", async () => {
const source = stubSource();
const lib = await openLibrary(source);
const outDir = join(root, "backup");
const result = await lib.backup({
downloadDirectory: outDir,
onlyAlbumNames: ["Work"],
});
expect(result.totalFiles).toBe(1);
expect(result.downloaded).toBe(1);
expect(existsSync(join(outDir, "originals", "200.png"))).toBe(true);
expect(existsSync(join(outDir, "originals", "100.jpg"))).toBe(false);
expect(existsSync(join(outDir, "collections", "Work.json"))).toBe(true);
expect(existsSync(join(outDir, "collections", "Vacation.json"))).toBe(
false,
);
lib.close();
});
it("prunes a ledger entry for a file no longer in the library and exits zero", async () => {
const source = stubSource();
const lib = await openLibrary(source);
const outDir = join(root, "backup");
// A prior failure for a file that has since left the library (deleted
// from the account). This run has no way to resolve it, so it must not
// keep the exit code non-zero forever.
seedLedger(outDir, 999, "gone.jpg");
const result = await lib.backup({ downloadDirectory: outDir });
// Everything still present is backed up cleanly, and the stale entry is
// dropped rather than counted.
expect(result.downloaded).toBe(3);
expect(result.failed).toBe(0);
expect(existsSync(join(outDir, "failures.json"))).toBe(false);
lib.close();
});
it("prunes an out-of-scope ledger entry on a scoped run and exits zero", async () => {
const source = stubSource();
const lib = await openLibrary(source);
const outDir = join(root, "backup");
// A prior failure for a Vacation file; this run is scoped to Work and
// never attempts it, so it must not poison the scoped run's exit code.
seedLedger(outDir, 100, "beach.jpg");
const result = await lib.backup({
downloadDirectory: outDir,
onlyAlbumNames: ["Work"],
});
expect(result.totalFiles).toBe(1);
expect(result.failed).toBe(0);
expect(existsSync(join(outDir, "originals", "200.png"))).toBe(true);
expect(existsSync(join(outDir, "failures.json"))).toBe(false);
lib.close();
});
it("also stores thumbnails when includeThumbnails is set", async () => {
const source = stubSource();
const lib = await openLibrary(source);
const outDir = join(root, "backup");
await lib.backup({
downloadDirectory: outDir,
includeThumbnails: true,
});
expect(existsSync(join(outDir, "thumbnails", "100.jpg"))).toBe(true);
expect(existsSync(join(outDir, "thumbnails", "200.jpg"))).toBe(true);
lib.close();
});
it("counts one attempt when a file fails both its original and thumbnail in a run", async () => {
const source = stubSource();
source.failID = 101;
source.failThumbID = 101;
const lib = await openLibrary(source);
const outDir = join(root, "backup");
const result = await lib.backup({
downloadDirectory: outDir,
includeThumbnails: true,
});
// Both kinds fail for 101, but the run counts it once.
const errs = result.errors.filter((e) => e.fileID === 101);
expect(errs.length).toBe(1);
expect(readLedger(outDir).files["101"]!.attempts).toBe(1);
lib.close();
}); });
}); });
+64 -64
View File
@@ -2,17 +2,13 @@
* Tests for `quak backup-metadata <dir>`. * Tests for `quak backup-metadata <dir>`.
* *
* This command dumps all decrypted account metadata into a directory * This command dumps all decrypted account metadata into a directory
* tree of plain JSON files, without downloading any file content (unless * tree of plain JSON files, without downloading any file content. It
* `--exif` is given). It is fast and produces a complete plaintext record of * is fast (no multi-megabyte downloads) and produces a complete
* every collection name, file title, creation date, GPS coordinate, camera * plaintext record of every collection name, file title, creation
* model, caption, face label, and any other metadata the Ente clients have * date, GPS coordinate, camera model, caption, face label, and any
* attached. * other metadata the Ente clients have attached.
* *
* As of issue #52 it runs on the library API: `runMetadataBackup(lib, client, * Layout:
* dir)` enumerates collections and files from the library's cache rather than
* scanning the client directly, and `--exif` reads each original through the
* library's content cache (`photo.original()`). The ML fetch is unchanged. The
* output tree is identical:
* *
* <dir>/ * <dir>/
* account.json { email, userID } * account.json { email, userID }
@@ -48,11 +44,7 @@ import {
} from "../../src/crypto/index.js"; } from "../../src/crypto/index.js";
import * as jpegJs from "jpeg-js"; import * as jpegJs from "jpeg-js";
import { Client } from "../../src/client.js"; import { Client } from "../../src/client.js";
import { Library } from "../../src/library/index.js"; import { runMetadataBackup } from "../../src/metadata-backup.js";
import {
runMetadataBackup,
type MetadataBackupOptions,
} from "../../src/metadata-backup.js";
import type { KeyAttributes } from "../../src/auth/types.js"; import type { KeyAttributes } from "../../src/auth/types.js";
const TEST_EMAIL = "metabackup@example.com"; const TEST_EMAIL = "metabackup@example.com";
@@ -439,50 +431,16 @@ afterAll(() => {
rmSync(testDir, { recursive: true, force: true }); rmSync(testDir, { recursive: true, force: true });
}); });
// Log in against the mock and open a library over its cache. The point commands
// open the library with the background precache off and a long refresh interval;
// the same here keeps the test deterministic (no thumbnail/original prefetch it
// did not ask for, no second refresh mid-test).
const openLib = async (client: Client): Promise<Library> =>
Library.open({
// The library client omits `fetchMLData`, matching how the CLI opens
// point commands: `runMetadataBackup` fetches ML data itself through
// the client, so the library's background backfill would only be a
// redundant second pass over the same endpoint.
client: {
whoami: () => client.whoami(),
collectionsSince: (args) => client.collectionsSince(args),
filesSince: (args) => client.filesSince(args),
contentSource: () => client.contentSource(),
},
cacheDirectory: mkdtempSync(join(testDir, "cache-")),
refreshIntervalSeconds: 3600,
precacheThumbnails: false,
precacheOriginals: false,
});
// Run one metadata backup end to end: fresh client, fresh library, then close.
const runBackup = async (
outDir: string,
opts?: MetadataBackupOptions,
): Promise<void> => {
const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMetaFetch(mock) },
});
const lib = await openLib(client);
try {
await runMetadataBackup(lib, client, outDir, opts);
} finally {
lib.close();
}
};
describe("quak backup-metadata", () => { describe("quak backup-metadata", () => {
it("writes account.json with email and userID", async () => { it("writes account.json with email and userID", async () => {
const outDir = join(testDir, "full"); const outDir = join(testDir, "full");
await runBackup(outDir); const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMetaFetch(mock) },
});
await runMetadataBackup(client, outDir);
const account = JSON.parse( const account = JSON.parse(
readFileSync(join(outDir, "account.json"), "utf-8"), readFileSync(join(outDir, "account.json"), "utf-8"),
@@ -493,7 +451,13 @@ describe("quak backup-metadata", () => {
it("creates per-collection directories with _collection.json", async () => { it("creates per-collection directories with _collection.json", async () => {
const outDir = join(testDir, "collections"); const outDir = join(testDir, "collections");
await runBackup(outDir); const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMetaFetch(mock) },
});
await runMetadataBackup(client, outDir);
const collDirs = readdirSync(join(outDir, "collections")); const collDirs = readdirSync(join(outDir, "collections"));
expect(collDirs.length).toBe(2); expect(collDirs.length).toBe(2);
@@ -514,7 +478,13 @@ describe("quak backup-metadata", () => {
it("decrypts collection-level pubMagicMetadata", async () => { it("decrypts collection-level pubMagicMetadata", async () => {
const outDir = join(testDir, "coll-magic"); const outDir = join(testDir, "coll-magic");
await runBackup(outDir); const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMetaFetch(mock) },
});
await runMetadataBackup(client, outDir);
const collDirs = readdirSync(join(outDir, "collections")); const collDirs = readdirSync(join(outDir, "collections"));
const vacDir = collDirs.find((d) => d.includes("Vacation"))!; const vacDir = collDirs.find((d) => d.includes("Vacation"))!;
@@ -531,7 +501,13 @@ describe("quak backup-metadata", () => {
it("writes per-file JSON with all three metadata layers", async () => { it("writes per-file JSON with all three metadata layers", async () => {
const outDir = join(testDir, "file-meta"); const outDir = join(testDir, "file-meta");
await runBackup(outDir); const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMetaFetch(mock) },
});
await runMetadataBackup(client, outDir);
const collDirs = readdirSync(join(outDir, "collections")); const collDirs = readdirSync(join(outDir, "collections"));
const vacDir = collDirs.find((d) => d.includes("Vacation"))!; const vacDir = collDirs.find((d) => d.includes("Vacation"))!;
@@ -550,7 +526,13 @@ describe("quak backup-metadata", () => {
it("handles files with no magic metadata gracefully", async () => { it("handles files with no magic metadata gracefully", async () => {
const outDir = join(testDir, "no-magic"); const outDir = join(testDir, "no-magic");
await runBackup(outDir); const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMetaFetch(mock) },
});
await runMetadataBackup(client, outDir);
const collDirs = readdirSync(join(outDir, "collections")); const collDirs = readdirSync(join(outDir, "collections"));
const workDir = collDirs.find((d) => d.includes("Work"))!; const workDir = collDirs.find((d) => d.includes("Work"))!;
@@ -568,8 +550,14 @@ describe("quak backup-metadata", () => {
it("is incremental: second run does not fail", async () => { it("is incremental: second run does not fail", async () => {
const outDir = join(testDir, "incremental"); const outDir = join(testDir, "incremental");
await runBackup(outDir); const client = await Client.login({
await runBackup(outDir); email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMetaFetch(mock) },
});
await runMetadataBackup(client, outDir);
await runMetadataBackup(client, outDir);
const account = JSON.parse( const account = JSON.parse(
readFileSync(join(outDir, "account.json"), "utf-8"), readFileSync(join(outDir, "account.json"), "utf-8"),
@@ -579,7 +567,13 @@ describe("quak backup-metadata", () => {
it("fetches and decrypts ML data by default", async () => { it("fetches and decrypts ML data by default", async () => {
const outDir = join(testDir, "ml-data"); const outDir = join(testDir, "ml-data");
await runBackup(outDir); const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMetaFetch(mock) },
});
await runMetadataBackup(client, outDir);
const collDirs = readdirSync(join(outDir, "collections")); const collDirs = readdirSync(join(outDir, "collections"));
const vacDir = collDirs.find((d) => d.includes("Vacation"))!; const vacDir = collDirs.find((d) => d.includes("Vacation"))!;
@@ -603,7 +597,13 @@ describe("quak backup-metadata", () => {
it("extracts EXIF from downloaded files when --exif is set", async () => { it("extracts EXIF from downloaded files when --exif is set", async () => {
const outDir = join(testDir, "exif-data"); const outDir = join(testDir, "exif-data");
await runBackup(outDir, { exif: true }); const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildMetaFetch(mock) },
});
await runMetadataBackup(client, outDir, { exif: true });
const collDirs = readdirSync(join(outDir, "collections")); const collDirs = readdirSync(join(outDir, "collections"));
const vacDir = collDirs.find((d) => d.includes("Vacation"))!; const vacDir = collDirs.find((d) => d.includes("Vacation"))!;
-76
View File
@@ -1,76 +0,0 @@
// The CLI presents a file by its own decrypted metadata, not the PhotoRecord
// projection (issue #52). For a renamed file the two disagree: the projection
// prefers `editedName` and reports `editedTime` in milliseconds, while the CLI
// must print the raw `metadata.title` and `metadata.creationTime` (microseconds)
// and name downloads after the raw title, byte-identical to the pre-library CLI.
//
// This locks in that contrast: the shared output helpers emit the raw values,
// and the projection of the same file emits the edited ones — so a regression
// that re-sourced the CLI from the projection would fail here.
import { describe, it, expect } from "vitest";
import {
fileListRow,
fileListLine,
originalName,
thumbnailName,
} from "../../src/cli-output.js";
import { deriveRecords } from "../../src/library/records.js";
import type { EnteFile } from "../../src/model/types.js";
// Microseconds, as Ente stores times.
const RAW_CREATION = 1700000000000000;
const EDITED_TIME = 1710000000000000;
const RAW_TITLE = "IMG_0001.HEIC";
const EDITED_NAME = "Sunset.heic";
// A file the user has renamed and re-dated: basic metadata holds the original
// title and capture time; public magic metadata holds the edits.
const renamedFile: EnteFile = {
id: 100,
collectionID: 10,
ownerID: 42,
key: new Uint8Array(),
metadata: {
title: RAW_TITLE,
fileType: "image",
creationTime: RAW_CREATION,
modificationTime: RAW_CREATION,
},
pubMagicMetadata: { editedName: EDITED_NAME, editedTime: EDITED_TIME },
file: { decryptionHeader: "" },
thumbnail: { decryptionHeader: "" },
updationTime: RAW_CREATION,
};
describe("CLI file output (issue #52)", () => {
it("emits the raw title and microsecond creationTime for --json", () => {
expect(fileListRow(renamedFile)).toEqual({
id: 100,
title: RAW_TITLE,
fileType: "image",
creationTime: RAW_CREATION,
collectionID: 10,
});
});
it("emits the raw title in the human column", () => {
expect(fileListLine(renamedFile)).toBe(`100\timage\t${RAW_TITLE}`);
});
it("names downloads after the raw title", () => {
expect(originalName(renamedFile)).toBe(RAW_TITLE);
expect(thumbnailName(renamedFile)).toBe(`thumb_${RAW_TITLE}`);
});
it("does not use the editedName/editedTime projection", () => {
const record = deriveRecords([], [renamedFile]).photos.get(100);
// The projection prefers the edits and reports milliseconds; the CLI
// helpers above deliberately do not.
expect(record?.title).toBe(EDITED_NAME);
expect(record?.takenAt).toBe(Math.floor(EDITED_TIME / 1000));
expect(fileListRow(renamedFile).title).not.toBe(record?.title);
expect(fileListRow(renamedFile).creationTime).not.toBe(record?.takenAt);
});
});
-195
View File
@@ -1,195 +0,0 @@
/**
* Tests for the CLI read helpers (`src/cli-read.ts`, owner amendment to
* issue #36, issue #52).
*
* The `collections`, `files`, `get`, and `get-thumb` commands must answer for
* current server state, not the local cache, so each helper forces a
* `Library.fresh()` round-trip before it reads. The stand-in library below
* serves nothing until `fresh()` has been awaited, so a helper that read
* without refreshing would come back empty and fail here.
*
* `collections` and `files` also list in the library's enumeration order
* (`listCollections`/`listFiles`) — the order the pre-library CLI printed — not
* the albums/photos projection's newest-first order. The fixtures are seeded in
* an enumeration order that a newest-first sort would rearrange, so a
* regression to the projection order would fail here too. Field values still
* come from the raw metadata via `cli-output.ts`.
*/
import { describe, it, expect } from "vitest";
import {
freshCollections,
freshFiles,
freshFile,
type FreshReadLibrary,
} from "../../src/cli-read.js";
import { fileListRow } from "../../src/cli-output.js";
import type { Photo } from "../../src/library/index.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
const collection = (id: number, updationTime: number): Collection => ({
id,
ownerID: 42,
key: new Uint8Array(),
name: `album-${id}`,
type: "album",
updationTime,
isShared: false,
});
// Microseconds, as Ente stores times.
const file = (
id: number,
collectionID: number,
creationTime: number,
): EnteFile => ({
id,
collectionID,
ownerID: 42,
key: new Uint8Array(),
metadata: {
title: `file-${id}.jpg`,
fileType: "image",
creationTime,
modificationTime: creationTime,
},
file: { decryptionHeader: "" },
thumbnail: { decryptionHeader: "" },
updationTime: creationTime,
});
// A library that reveals its records only after `fresh()` has been awaited, and
// serves them in the enumeration order it was given. `photos.byID` returns a
// stand-in `Photo` carrying just the fileID the helper passes through.
class FakeLibrary implements FreshReadLibrary {
freshCalls = 0;
private refreshed = false;
constructor(
private readonly collections: Collection[],
private readonly files: EnteFile[],
) {}
async fresh(): Promise<unknown> {
this.freshCalls++;
this.refreshed = true;
return {};
}
listCollections(): Collection[] {
return this.refreshed ? this.collections : [];
}
getCollection(id: number): Collection | undefined {
return this.listCollections().find((c) => c.id === id);
}
listFiles(collectionID: number): EnteFile[] {
return this.refreshed
? this.files.filter((f) => f.collectionID === collectionID)
: [];
}
getFileByID(fileID: number): EnteFile | undefined {
if (!this.refreshed) return undefined;
return this.files.find((f) => f.id === fileID);
}
photos = {
byID: ({ fileID }: { fileID: number }): Photo | undefined => {
if (!this.refreshed) return undefined;
if (!this.files.some((f) => f.id === fileID)) return undefined;
return { fileID } as unknown as Photo;
},
};
}
describe("CLI read helpers (issue #36 amendment, issue #52)", () => {
it("freshCollections refreshes first, then lists in enumeration order", async () => {
// Enumeration order 2, 1, 3; a newest-first sort would be 3, 2, 1.
const lib = new FakeLibrary(
[collection(2, 200), collection(1, 300), collection(3, 100)],
[],
);
const rows = await freshCollections(lib);
expect(lib.freshCalls).toBe(1);
expect(rows.map((c) => c.id)).toEqual([2, 1, 3]);
// The projection's newest-first order is a different sequence, so this
// is not accidentally that order.
const newestFirst = [...rows]
.sort((a, b) => b.updationTime - a.updationTime)
.map((c) => c.id);
expect(newestFirst).toEqual([1, 2, 3]);
expect(rows.map((c) => c.id)).not.toEqual(newestFirst);
});
it("freshFiles refreshes first, lists in enumeration order, keeps raw fields", async () => {
// Enumeration order by id 10, 11, 12; creationTimes ascending, so a
// newest-first sort would reverse them.
const files = [
file(10, 1, 1_700_000_000_000_000),
file(11, 1, 1_700_000_000_000_001),
file(12, 1, 1_700_000_000_000_002),
];
const lib = new FakeLibrary([collection(1, 100)], files);
const rows = await freshFiles(lib, 1);
expect(lib.freshCalls).toBe(1);
expect(rows?.map((f) => f.id)).toEqual([10, 11, 12]);
// Field values come from raw metadata: microsecond creationTime and the
// raw title, unchanged.
expect(rows?.map(fileListRow)).toEqual([
{
id: 10,
title: "file-10.jpg",
fileType: "image",
creationTime: 1_700_000_000_000_000,
collectionID: 1,
},
{
id: 11,
title: "file-11.jpg",
fileType: "image",
creationTime: 1_700_000_000_000_001,
collectionID: 1,
},
{
id: 12,
title: "file-12.jpg",
fileType: "image",
creationTime: 1_700_000_000_000_002,
collectionID: 1,
},
]);
});
it("freshFiles returns undefined for an unknown collection", async () => {
const lib = new FakeLibrary([collection(1, 100)], []);
const rows = await freshFiles(lib, 999);
expect(lib.freshCalls).toBe(1);
expect(rows).toBeUndefined();
});
it("freshFile refreshes first, then resolves the photo and its raw record", async () => {
const f = file(10, 1, 1_700_000_000_000_000);
const lib = new FakeLibrary([collection(1, 100)], [f]);
const resolved = await freshFile(lib, 10);
expect(lib.freshCalls).toBe(1);
expect(resolved?.photo.fileID).toBe(10);
expect(resolved?.file.metadata.title).toBe("file-10.jpg");
expect(resolved?.file.metadata.creationTime).toBe(
1_700_000_000_000_000,
);
});
it("freshFile returns undefined for an unknown file", async () => {
const lib = new FakeLibrary([collection(1, 100)], []);
const resolved = await freshFile(lib, 404);
expect(lib.freshCalls).toBe(1);
expect(resolved).toBeUndefined();
});
});
-367
View File
@@ -1,367 +0,0 @@
/**
* Tests for the resumable, deletion-aware enumeration variants on `Client`:
* `collectionsSince` and `filesSince`.
*
* The whole-account methods `listCollections` / `listFiles` always start at
* `sinceTime: 0` and hide deletions. The cache refresh needs the opposite:
* start from a saved cursor, learn what was deleted, and get back a cursor to
* resume from next time. These two methods provide that.
*
* The return shape keeps live records and tombstones apart — `collections` /
* `files` are decrypted live records, `deleted` is a plain list of the ids the
* server tombstoned. A tombstone carries no decryptable key or metadata, so it
* is a bare id rather than a hollowed-out `Collection` / `EnteFile`.
*
* All tests inject a fake `fetch` and drive a real `Client` (built with
* `Client.fromJSON`) so the decryption path runs for real. Live rows are built
* with libsodium exactly as the server would encrypt them; tombstone rows carry
* only the fields the code reads (`id`, `updationTime`, `isDeleted`), because
* they are never decrypted.
*/
import sodium from "libsodium-wrappers-sumo";
import { beforeAll, describe, expect, it } from "vitest";
import { init, toBase64 } from "../../src/crypto/index.js";
import { Client, type ClientSnapshot } from "../../src/client.js";
// ---------------------------------------------------------------------------
// Fixtures
// ---------------------------------------------------------------------------
const USER_ID = 42;
interface Keys {
masterKey: Uint8Array;
publicKey: Uint8Array;
secretKey: Uint8Array;
}
const buildKeys = (): Keys => {
const kp = sodium.crypto_box_keypair();
return {
masterKey: sodium.crypto_secretbox_keygen(),
publicKey: kp.publicKey,
secretKey: kp.privateKey,
};
};
const snapshotFor = (keys: Keys): ClientSnapshot => ({
email: "user@example.com",
userID: USER_ID,
token: "test-token",
masterKey: toBase64(keys.masterKey),
secretKey: toBase64(keys.secretKey),
publicKey: toBase64(keys.publicKey),
});
const secretboxEncrypt = (
plaintext: Uint8Array,
key: Uint8Array,
): { ciphertext: Uint8Array; nonce: Uint8Array } => {
const nonce = sodium.randombytes_buf(sodium.crypto_secretbox_NONCEBYTES);
return {
ciphertext: sodium.crypto_secretbox_easy(plaintext, nonce, key),
nonce,
};
};
/** An owned collection row as the server sends it, keyed under the master key. */
const ownedCollectionRow = (
masterKey: Uint8Array,
opts: { id: number; name: string; updationTime: number },
): Record<string, unknown> => {
const collectionKey = sodium.crypto_secretbox_keygen();
const { ciphertext: encKey, nonce: keyNonce } = secretboxEncrypt(
collectionKey,
masterKey,
);
const { ciphertext: encName, nonce: nameNonce } = secretboxEncrypt(
new TextEncoder().encode(opts.name),
collectionKey,
);
return {
id: opts.id,
owner: { id: USER_ID },
encryptedKey: toBase64(encKey),
keyDecryptionNonce: toBase64(keyNonce),
encryptedName: toBase64(encName),
nameDecryptionNonce: toBase64(nameNonce),
type: "album",
updationTime: opts.updationTime,
};
};
/** A live file row inside a collection, keyed under that collection's key. */
const fileRow = (
collectionKey: Uint8Array,
opts: { id: number; title: string; updationTime: number },
): Record<string, unknown> => {
const fileKey = sodium.crypto_secretbox_keygen();
const { ciphertext: encFileKey, nonce: fileKeyNonce } = secretboxEncrypt(
fileKey,
collectionKey,
);
const metadata = {
title: opts.title,
fileType: 0,
creationTime: opts.updationTime,
modificationTime: opts.updationTime,
};
const push =
sodium.crypto_secretstream_xchacha20poly1305_init_push(fileKey);
const encMeta = sodium.crypto_secretstream_xchacha20poly1305_push(
push.state,
new TextEncoder().encode(JSON.stringify(metadata)),
null,
sodium.crypto_secretstream_xchacha20poly1305_TAG_FINAL,
);
return {
id: opts.id,
collectionID: 1,
ownerID: USER_ID,
encryptedKey: toBase64(encFileKey),
keyDecryptionNonce: toBase64(fileKeyNonce),
metadata: {
encryptedData: toBase64(encMeta),
decryptionHeader: toBase64(push.header),
},
file: { decryptionHeader: toBase64(sodium.randombytes_buf(24)) },
thumbnail: { decryptionHeader: toBase64(sodium.randombytes_buf(24)) },
updationTime: opts.updationTime,
};
};
/** A tombstone row. Never decrypted, so only these fields are ever read. */
const tombstoneRow = (
id: number,
updationTime: number,
): Record<string, unknown> => ({
id,
updationTime,
isDeleted: true,
});
const jsonResponse = (body: unknown): Response =>
new Response(JSON.stringify(body), {
status: 200,
headers: { "content-type": "application/json" },
});
/**
* A fetch that serves canned responses in order and records the `sinceTime`
* query parameter each request carried, so tests can prove the cursor is
* threaded from one page (and one call) to the next.
*/
const recordingFetch = (
...responses: Response[]
): { fetch: typeof globalThis.fetch; sinceTimes: (string | null)[] } => {
const sinceTimes: (string | null)[] = [];
let i = 0;
const fake = async (input: RequestInfo | URL): Promise<Response> => {
const url =
typeof input === "string"
? input
: input instanceof URL
? input.href
: input.url;
sinceTimes.push(new URL(url).searchParams.get("sinceTime"));
if (i >= responses.length) {
throw new Error(`recordingFetch: no response for call #${i}`);
}
return responses[i++]!;
};
return { fetch: fake as typeof globalThis.fetch, sinceTimes };
};
// ---------------------------------------------------------------------------
// Tests
// ---------------------------------------------------------------------------
describe("Client.filesSince", () => {
beforeAll(async () => {
await init();
await sodium.ready;
});
it("pages from the given cursor, decrypts live rows, and collects tombstones", async () => {
const keys = buildKeys();
const collectionKey = sodium.crypto_secretbox_keygen();
// Page 1 mixes a live file and a tombstone; the tombstone has the
// higher updationTime, so it — not the live row — sets the cursor the
// second page must be fetched from.
const { fetch, sinceTimes } = recordingFetch(
jsonResponse({
diff: [
fileRow(collectionKey, {
id: 1001,
title: "first.jpg",
updationTime: 100,
}),
tombstoneRow(1002, 150),
],
hasMore: true,
}),
jsonResponse({
diff: [
fileRow(collectionKey, {
id: 1003,
title: "second.jpg",
updationTime: 200,
}),
],
hasMore: false,
}),
);
const client = Client.fromJSON(snapshotFor(keys), { fetch });
const { files, deleted, cursor } = await client.filesSince({
collectionID: 1,
collectionKey,
sinceTime: 0,
});
expect(files.map((f) => f.id)).toEqual([1001, 1003]);
expect(files.map((f) => f.metadata.title)).toEqual([
"first.jpg",
"second.jpg",
]);
expect(deleted).toEqual([1002]);
expect(cursor).toBe(200);
// First request started at the caller's cursor; the second resumed
// from the max updationTime seen on the first page (the tombstone's).
expect(sinceTimes).toEqual(["0", "150"]);
});
it("fetches only newer rows when the returned cursor is passed back in", async () => {
const keys = buildKeys();
const collectionKey = sodium.crypto_secretbox_keygen();
const { fetch, sinceTimes } = recordingFetch(
jsonResponse({ diff: [], hasMore: false }),
);
const client = Client.fromJSON(snapshotFor(keys), { fetch });
const result = await client.filesSince({
collectionID: 1,
collectionKey,
sinceTime: 200,
});
expect(result.files).toEqual([]);
expect(result.deleted).toEqual([]);
// An empty diff advances nothing: the cursor falls back to the input.
expect(result.cursor).toBe(200);
expect(sinceTimes).toEqual(["200"]);
});
it("stops and throws when the server claims more but does not advance (#7)", async () => {
const keys = buildKeys();
const collectionKey = sodium.crypto_secretbox_keygen();
// hasMore is true, but the page's max updationTime (50) does not exceed
// the cursor the request was made with (50). Following hasMore here
// would refetch this same page forever.
const { fetch, sinceTimes } = recordingFetch(
jsonResponse({ diff: [tombstoneRow(1, 50)], hasMore: true }),
jsonResponse({ diff: [tombstoneRow(1, 50)], hasMore: true }),
);
const client = Client.fromJSON(snapshotFor(keys), { fetch });
await expect(
client.filesSince({
collectionID: 1,
collectionKey,
sinceTime: 50,
}),
).rejects.toThrow(/not advance|non-advancing/i);
// It gave up after the first page rather than looping.
expect(sinceTimes).toEqual(["50"]);
});
});
describe("Client.collectionsSince", () => {
beforeAll(async () => {
await init();
await sodium.ready;
});
it("decrypts live collections, collects tombstones, and returns a cursor", async () => {
const keys = buildKeys();
const { fetch, sinceTimes } = recordingFetch(
jsonResponse({
collections: [
ownedCollectionRow(keys.masterKey, {
id: 1,
name: "Vacation",
updationTime: 100,
}),
tombstoneRow(3, 150),
],
}),
);
const client = Client.fromJSON(snapshotFor(keys), { fetch });
const { collections, deleted, cursor } = await client.collectionsSince({
sinceTime: 0,
});
expect(collections.map((c) => c.id)).toEqual([1]);
expect(collections[0]!.name).toBe("Vacation");
expect(deleted).toEqual([3]);
// The tombstone's updationTime advances the cursor too, so the next
// sync starts after it rather than seeing it again.
expect(cursor).toBe(150);
expect(sinceTimes).toEqual(["0"]);
});
it("falls back to the input cursor on an empty response", async () => {
const keys = buildKeys();
const { fetch, sinceTimes } = recordingFetch(
jsonResponse({ collections: [] }),
);
const client = Client.fromJSON(snapshotFor(keys), { fetch });
const result = await client.collectionsSince({ sinceTime: 150 });
expect(result.collections).toEqual([]);
expect(result.deleted).toEqual([]);
expect(result.cursor).toBe(150);
expect(sinceTimes).toEqual(["150"]);
});
});
describe("Client list wrappers still hide deletions", () => {
beforeAll(async () => {
await init();
await sodium.ready;
});
it("listFiles drops tombstones and returns only live files", async () => {
const keys = buildKeys();
const collectionKey = sodium.crypto_secretbox_keygen();
const { fetch, sinceTimes } = recordingFetch(
jsonResponse({
diff: [
fileRow(collectionKey, {
id: 7,
title: "keep.jpg",
updationTime: 100,
}),
tombstoneRow(8, 150),
],
hasMore: false,
}),
);
const client = Client.fromJSON(snapshotFor(keys), { fetch });
const files = await client.listFiles(1, collectionKey);
expect(files.map((f) => f.id)).toEqual([7]);
// The wrapper starts a full enumeration from zero.
expect(sinceTimes).toEqual(["0"]);
});
});
+5 -348
View File
@@ -72,11 +72,7 @@ import { init, toBase64, STREAM_CHUNK_SIZE } from "../../src/crypto/index.js";
import { ApiClient } from "../../src/api/client.js"; import { ApiClient } from "../../src/api/client.js";
import { ApiError, TruncatedStreamError } from "../../src/errors.js"; import { ApiError, TruncatedStreamError } from "../../src/errors.js";
import type { RetryOptions } from "../../src/retry.js"; import type { RetryOptions } from "../../src/retry.js";
import { import { downloadFile, downloadThumbnail } from "../../src/download/index.js";
downloadFile,
downloadThumbnail,
writeAtomic,
} from "../../src/download/index.js";
import type { EnteFile, FileMetadata } from "../../src/model/types.js"; import type { EnteFile, FileMetadata } from "../../src/model/types.js";
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -105,79 +101,17 @@ const renameHook = vi.hoisted(() => ({
failWith: null as Error | null, failWith: null as Error | null,
})); }));
/**
* `open` is wrapped so the tests can observe the durability fsyncs the atomic
* writer performs — which are otherwise invisible: an fsync leaves no trace in
* the file's contents. Each `FileHandle.sync()` is recorded, and rename and
* sync events are appended to a single ordered `events` log so a test can pin
* the sequence "fsync the temp file, rename, fsync the directory" that makes a
* write survive a power cut. The flag the handle was opened with distinguishes
* the temp file (`w`) from its containing directory (`r`).
*/
const durabilityHook = vi.hoisted(() => ({
events: [] as string[],
}));
/**
* `FileHandle.write` is wrapped so the tests can watch the streaming decrypt
* path put plaintext on disk one chunk at a time. This is the direct evidence
* that memory is bounded by the chunk size and not the file size: a buffered
* downloader would hand the whole file to a single write, whereas the streaming
* one issues one write per secretstream chunk, none larger than
* `STREAM_CHUNK_SIZE`. Each write records the temp path it targeted and its
* length. `writeFile` (which the whole-buffer `writeAtomic` uses) is a distinct
* native call and does not go through this method, so only the streaming path
* is observed here.
*/
const writeHook = vi.hoisted(() => ({
writes: [] as { path: string; length: number }[],
}));
vi.mock("node:fs/promises", async (importOriginal) => { vi.mock("node:fs/promises", async (importOriginal) => {
const actual = await importOriginal<typeof import("node:fs/promises")>(); const actual = await importOriginal<typeof import("node:fs/promises")>();
const { existsSync: sourceExists } = await import("node:fs"); const { existsSync: sourceExists } = await import("node:fs");
return { return {
...actual, ...actual,
open: async (
path: Parameters<typeof actual.open>[0],
flags?: Parameters<typeof actual.open>[1],
...rest: unknown[]
): Promise<Awaited<ReturnType<typeof actual.open>>> => {
const handle = await actual.open(
path,
flags as Parameters<typeof actual.open>[1],
...(rest as []),
);
const realSync = handle.sync.bind(handle);
handle.sync = async (): Promise<void> => {
durabilityHook.events.push(`sync:${String(flags)}:${path}`);
await realSync();
};
const realWrite = handle.write.bind(handle);
handle.write = (async (
data: unknown,
...rest2: unknown[]
): Promise<unknown> => {
if (data instanceof Uint8Array) {
writeHook.writes.push({
path: String(path),
length: data.length,
});
}
return (realWrite as (...a: unknown[]) => Promise<unknown>)(
data,
...rest2,
);
}) as typeof handle.write;
return handle;
},
rename: async (from: string, to: string): Promise<void> => { rename: async (from: string, to: string): Promise<void> => {
renameHook.calls.push({ renameHook.calls.push({
from, from,
to, to,
sourceExisted: sourceExists(from), sourceExisted: sourceExists(from),
}); });
durabilityHook.events.push(`rename:${to}`);
if (renameHook.failWith !== null) { if (renameHook.failWith !== null) {
throw renameHook.failWith; throw renameHook.failWith;
} }
@@ -189,8 +123,6 @@ vi.mock("node:fs/promises", async (importOriginal) => {
beforeEach(() => { beforeEach(() => {
renameHook.calls.length = 0; renameHook.calls.length = 0;
renameHook.failWith = null; renameHook.failWith = null;
durabilityHook.events.length = 0;
writeHook.writes.length = 0;
}); });
let testDir: string; let testDir: string;
@@ -979,13 +911,10 @@ describe.each(entryPoints)("$name retries", ({ name, download }) => {
}); });
it("stages one temp file for the attempt that succeeded, not one per attempt", async () => { it("stages one temp file for the attempt that succeeded, not one per attempt", async () => {
// Each streaming attempt stages into its own temp file, but a retried // The atomic write stays outside the retry loop. A retried download
// download must not leave a trail of half-written scratch files: a // must not leave a trail of half-written scratch files, and the
// failed attempt removes its temp file, and the destination is renamed // destination must be touched exactly once — by the attempt that
// into place exactly once — by the attempt that produced a complete, // produced a complete, authenticated plaintext.
// authenticated plaintext. (Here the two failed attempts reset before a
// whole chunk is pulled, so they write nothing; the point stands either
// way — see the retry-restart test below, where they do write.)
const { key, header, ciphertext } = smallFixture(42); const { key, header, ciphertext } = smallFixture(42);
const { fetch } = scriptedCdnFetch( const { fetch } = scriptedCdnFetch(
{ kind: "reset", bytes: ciphertext.slice(0, 16) }, { kind: "reset", bytes: ciphertext.slice(0, 16) },
@@ -1099,99 +1028,6 @@ describe.each(entryPoints)("$name retries", ({ name, download }) => {
}); });
}); });
// ---------------------------------------------------------------------------
// Streaming decrypt to disk
//
// The plaintext is never held whole in memory: each secretstream chunk is
// written to the temp file as it is decrypted, so peak memory is bounded by the
// chunk size rather than the file size. These tests watch the writes directly
// (see `writeHook`) rather than infer memory behaviour from the final file.
// ---------------------------------------------------------------------------
describe.each(entryPoints)("$name streams to disk", ({ name, download }) => {
const freshDir = (): string =>
mkdtempSync(join(testDir, `${name}-stream-`));
/** Writes recorded against staged temp files (not the `writeFile` path). */
const tempWrites = (): { path: string; length: number }[] =>
writeHook.writes.filter((w) => w.path.endsWith(".tmp"));
it("writes one chunk at a time, none larger than STREAM_CHUNK_SIZE", async () => {
// The multi-chunk fixture decrypts to one full 4 MiB chunk plus a small
// final chunk. A streaming writer therefore issues exactly two writes,
// of STREAM_CHUNK_SIZE and then the final chunk's length — never a
// single write carrying the whole 4 MiB + 1 KiB file. That per-chunk
// shape is what "memory bounded by chunk size" means in practice: the
// plaintext is handed to the filesystem and dropped, chunk by chunk.
const { api, file } = fixtureFor(
multiChunkKey,
multiChunk.header,
multiChunk.body,
);
const outPath = join(freshDir(), "streamed.bin");
const result = await download(api, file, outPath);
const writes = tempWrites();
expect(writes.map((w) => w.length)).toEqual([
STREAM_CHUNK_SIZE,
multiChunk.plaintext.length - STREAM_CHUNK_SIZE,
]);
// No single write ever carried the whole file, and every write fits in
// one chunk's worth of memory.
for (const w of writes) {
expect(w.length).toBeLessThanOrEqual(STREAM_CHUNK_SIZE);
}
expect(result.bytesWritten).toBe(multiChunk.plaintext.length);
expectSameBytes(readFileSync(outPath), multiChunk.plaintext);
});
it("restarts from byte zero on a retry, replacing the temp file cleanly", async () => {
// The secretstream pull state is not resumable, so a retry cannot
// continue a half-written file — it must start over. The first attempt
// here delivers a complete leading chunk and then stops before the
// TAG_FINAL chunk: 4 MiB of plaintext lands in a temp file, then the
// download is rejected as truncated and that temp file is discarded.
// The retry streams the whole body into a *fresh* temp file, so the
// destination ends up with exactly the plaintext once — never the
// leading chunk twice, and never a stale temp file left behind.
const truncatedBody = multiChunk.body.slice(
0,
multiChunk.finalChunkOffset,
);
const { fetch, requests } = scriptedCdnFetch(
{ kind: "body", bytes: truncatedBody },
{ kind: "body", bytes: multiChunk.body },
);
const api = new ApiClient({
fetch,
retry: { ...noWait, attempts: 4 },
});
const file = buildMockEnteFile(
multiChunkKey,
multiChunk.header,
multiChunk.header,
);
const dir = freshDir();
const outPath = join(dir, "retry-restart.bin");
const result = await download(api, file, outPath);
expect(requests()).toBe(2);
// Both attempts streamed to disk, each into its own temp file: the
// truncated first attempt wrote before it failed, proving the retry did
// not resume a partial file but replaced it.
const distinctTemps = new Set(tempWrites().map((w) => w.path));
expect(distinctTemps.size).toBe(2);
// The destination holds the complete plaintext exactly once, and no
// temp file survives.
expect(result.bytesWritten).toBe(multiChunk.plaintext.length);
expectSameBytes(readFileSync(outPath), multiChunk.plaintext);
expect(renameHook.calls).toHaveLength(1);
expect(readdirSync(dir)).toEqual(["retry-restart.bin"]);
});
});
describe("download retries: corruption is not retried", () => { describe("download retries: corruption is not retried", () => {
it("gives up immediately on a chunk that failed to authenticate", async () => { it("gives up immediately on a chunk that failed to authenticate", async () => {
// A whole chunk that failed to authenticate while the stream // A whole chunk that failed to authenticate while the stream
@@ -1225,182 +1061,3 @@ describe("download retries: corruption is not retried", () => {
expect(requests()).toBe(1); expect(requests()).toBe(1);
}); });
}); });
// ---------------------------------------------------------------------------
// Fragmented network reads
//
// A CDN does not hand the body over one secretstream chunk at a time; it
// arrives in whatever pieces the socket produces, many of them far smaller than
// a chunk and most straddling a chunk boundary. `streamDecrypt` reassembles
// those pieces before decrypting, copying each received byte once rather than
// recopying the whole accumulator on every read. This is the path the other
// fixtures never take — their mock fetch delivers each body as a single
// `Response` value, i.e. one read — so it is exercised explicitly here.
// ---------------------------------------------------------------------------
/**
* A fetch that serves `body` through a `ReadableStream` sliced into many
* fixed-size pieces, imitating a socket that trickles bytes in. `pieceSize` is
* chosen not to divide the chunk framing evenly, so pieces straddle the
* `ENC_CHUNK_SIZE` boundary the downloader splits on — the case a single-value
* body can never produce. `emitted` reports how many pieces were yielded, so a
* test can assert the body really was fragmented and not delivered whole.
*/
const mockFetchForFragmentedBody = (
body: Uint8Array,
pieceSize: number,
): { fetch: typeof globalThis.fetch; emitted: () => number } => {
let pieces = 0;
const fake = async (): Promise<Response> =>
new Response(
new ReadableStream<Uint8Array>({
start(controller) {
for (let off = 0; off < body.length; off += pieceSize) {
controller.enqueue(body.subarray(off, off + pieceSize));
pieces++;
}
controller.close();
},
}),
{ status: 200 },
);
return { fetch: fake as typeof globalThis.fetch, emitted: () => pieces };
};
describe("streamDecrypt fragmented reads", () => {
it("decrypts a multi-chunk body delivered in many small pieces", async () => {
// The multi-chunk fixture (one full 4 MiB chunk plus a small final
// chunk) delivered in 1000-byte pieces: several thousand reads, with
// the piece that spans the 4 MiB + 17 byte chunk boundary split across
// two chunks by the reassembler. The plaintext must come out
// byte-identical to the single-read case, and the chunk framing must be
// untouched: exactly two writes, `STREAM_CHUNK_SIZE` then the final
// chunk, the same as when the body arrives whole. If the boundary
// handling were off by a byte under fragmentation, either the pull
// would fail to authenticate or the write sizes would shift.
const { fetch, emitted } = mockFetchForFragmentedBody(
multiChunk.body,
1000,
);
const api = new ApiClient({ fetch });
const file = buildMockEnteFile(
multiChunkKey,
multiChunk.header,
multiChunk.header,
);
const dir = mkdtempSync(join(testDir, "fragmented-"));
const outPath = join(dir, "fragmented.bin");
const result = await downloadFile(api, file, outPath);
// The body really was trickled in, not handed over whole.
expect(emitted()).toBeGreaterThan(1000);
const writes = writeHook.writes.filter((w) => w.path.endsWith(".tmp"));
expect(writes.map((w) => w.length)).toEqual([
STREAM_CHUNK_SIZE,
multiChunk.plaintext.length - STREAM_CHUNK_SIZE,
]);
expect(result.bytesWritten).toBe(multiChunk.plaintext.length);
expectSameBytes(readFileSync(outPath), multiChunk.plaintext);
});
});
// ---------------------------------------------------------------------------
// Durable atomic writes
//
// `writeAtomic` is exported so the metadata store can reuse the same
// power-cut-safe write. Its durability is the point: the bytes and the new
// directory entry must both be on stable storage before it returns, so a crash
// immediately afterwards cannot resurrect an empty renamed file (#22 area 1).
// ---------------------------------------------------------------------------
describe("writeAtomic", () => {
it("fsyncs the temp file before the rename and the directory after", async () => {
const dir = mkdtempSync(join(testDir, "atomic-"));
const dest = join(dir, "durable.bin");
const bytes = patternBytes(2048, 71);
await writeAtomic(dest, bytes);
expect(readFileSync(dest)).toEqual(Buffer.from(bytes));
// The order is the durability contract: fsync the staged temp file so
// its contents are on disk, rename it into place, then fsync the
// directory so that new entry is on disk too. Do the directory fsync
// before the rename, or skip it, and a crash can lose the rename.
expect(durabilityHook.events).toHaveLength(3);
expect(durabilityHook.events[0]).toMatch(/^sync:w:.*\.tmp$/);
expect(durabilityHook.events[1]).toBe(`rename:${dest}`);
expect(durabilityHook.events[2]).toBe(`sync:r:${dir}`);
});
it("leaves no temp file behind when the write cannot be renamed", async () => {
const dir = mkdtempSync(join(testDir, "atomic-fail-"));
const dest = join(dir, "unrenamable.bin");
renameHook.failWith = new Error("simulated rename failure");
await expect(writeAtomic(dest, patternBytes(64, 72))).rejects.toThrow(
"simulated rename failure",
);
// The staged temp file was fsynced, then the rename failed; the cleanup
// path must remove it so a repeatedly failing write cannot fill the disk.
expect(existsSync(dest)).toBe(false);
expect(readdirSync(dir)).toEqual([]);
});
});
// ---------------------------------------------------------------------------
// Per-chunk progress
//
// Callers streaming a large file want bytes-written as it lands, not only the
// final total. The hook fires as decrypted plaintext accumulates; its values
// are non-decreasing and its last value is exactly `bytesWritten`.
// ---------------------------------------------------------------------------
describe.each(entryPoints)("$name progress", ({ name, download }) => {
it("reports monotonic progress ending at bytesWritten", async () => {
// The multi-chunk fixture pulls one full 4 MiB chunk and then a small
// final chunk, so the callback fires more than once and monotonicity is
// actually observable rather than trivially true for a single fire.
const { api, file } = fixtureFor(
multiChunkKey,
multiChunk.header,
multiChunk.body,
);
const outPath = join(
mkdtempSync(join(testDir, `${name}-progress-`)),
"p.bin",
);
const seen: number[] = [];
const result = await download(api, file, outPath, (bytesDone) => {
seen.push(bytesDone);
});
expect(seen.length).toBeGreaterThan(1);
for (let i = 1; i < seen.length; i++) {
expect(seen[i]!).toBeGreaterThan(seen[i - 1]!);
}
expect(seen[seen.length - 1]).toBe(result.bytesWritten);
expect(result.bytesWritten).toBe(multiChunk.plaintext.length);
});
it("downloads normally when no progress callback is given", async () => {
// The callback is optional and its absence must be side-effect-free:
// the download succeeds exactly as it does elsewhere in this file.
const plaintext = patternBytes(300, 73);
const key = sodium.crypto_secretstream_xchacha20poly1305_keygen();
const { header, ciphertext } = encryptFileBody(plaintext, key);
const { api, file } = fixtureFor(key, header, ciphertext);
const outPath = join(
mkdtempSync(join(testDir, `${name}-noprog-`)),
"n.bin",
);
const result = await download(api, file, outPath);
expect(result.bytesWritten).toBe(plaintext.length);
expectSameBytes(readFileSync(outPath), plaintext);
});
});
-358
View File
@@ -1,358 +0,0 @@
/**
* Tests for the originals cache size limit and LRU eviction (issue #47),
* layered on the on-disk content cache (#46).
*
* Only `cacheDirectory/originals` is bounded and evicted. Before each original
* write the effective limit is
* min(cacheOriginalsMaxBytes, bytesUsedByOriginals + bytesFree - freeBelowBytes)
* with `bytesFree` read from `statfs` on the volume holding `cacheDirectory`.
* The limit therefore falls as the disk fills and rises as space returns. When
* a write would cross the limit, least-recently-used originals are removed until
* it fits; pinned files are skipped, and if only pinned files remain the write
* proceeds over-limit. Last-use is the file `mtime`, bumped whenever a read
* returns an original's path, so ordering survives a restart with no ledger.
*
* `statfs` is injected so the adaptive limit is exercised deterministically:
* `bsize` is 1, so `bavail` is the free byte count the formula sees. The source
* writes a controllable number of bytes per file, and tests set each stored
* file's `mtime` explicitly so LRU order does not depend on wall-clock timing.
*/
import { describe, it, expect, beforeEach, afterEach } from "vitest";
import { mkdtempSync, rmSync, existsSync, statSync, utimesSync } from "node:fs";
import { writeFile } from "node:fs/promises";
import { tmpdir } from "node:os";
import { join } from "node:path";
import {
ContentCache,
type ContentSource,
type StatFsFn,
} from "../../src/library/content.js";
import { RequestPools } from "../../src/library/pools.js";
import type { EnteFile } from "../../src/model/types.js";
const file = (id: number): EnteFile => ({
id,
collectionID: 1,
ownerID: 1,
key: new Uint8Array([id & 0xff]),
metadata: {
title: `file-${id}.jpg`,
fileType: "image",
creationTime: 0,
modificationTime: 0,
},
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime: 0,
});
// A source that writes a controllable number of bytes per file (default 10).
class SizedSource implements ContentSource {
sizeFor = new Map<number, number>();
async original(args: {
file: EnteFile;
destination: string;
}): Promise<{ bytesWritten: number }> {
const n = this.sizeFor.get(args.file.id) ?? 10;
await writeFile(args.destination, Buffer.alloc(n, 1));
return { bytesWritten: n };
}
async thumbnail(args: {
file: EnteFile;
destination: string;
}): Promise<{ bytesWritten: number }> {
await writeFile(args.destination, Buffer.alloc(10, 1));
return { bytesWritten: 10 };
}
}
// A source whose original() writes 10 bytes but blocks on a gate before
// returning, so two fetches can be held in flight together. `bothStarted`
// resolves once both originals have entered, and `release()` lets them finish.
class GatedSource implements ContentSource {
private openGate!: () => void;
private readonly gate = new Promise<void>((r) => (this.openGate = r));
private inFlight = 0;
private reachedTwo!: () => void;
readonly bothStarted = new Promise<void>((r) => (this.reachedTwo = r));
async original(args: {
file: EnteFile;
destination: string;
}): Promise<{ bytesWritten: number }> {
this.inFlight += 1;
if (this.inFlight === 2) this.reachedTwo();
await this.gate;
await writeFile(args.destination, Buffer.alloc(10, 1));
return { bytesWritten: 10 };
}
async thumbnail(args: {
file: EnteFile;
destination: string;
}): Promise<{ bytesWritten: number }> {
await writeFile(args.destination, Buffer.alloc(10, 1));
return { bytesWritten: 10 };
}
release(): void {
this.openGate();
}
}
let root: string;
let cacheDir: string;
beforeEach(() => {
root = mkdtempSync(join(tmpdir(), "quak-eviction-"));
cacheDir = join(root, "cache");
});
afterEach(() => {
if (root && existsSync(root))
rmSync(root, { recursive: true, force: true });
});
const originalPath = (id: number): string =>
join(cacheDir, "originals", `${id}.jpg`);
// Pin an original's mtime to a fixed second-resolution instant so LRU order is
// deterministic. Lower `seconds` = older = evicted first.
const setMtime = (id: number, seconds: number): void => {
utimesSync(originalPath(id), seconds, seconds);
};
const buildCache = (args: {
source?: ContentSource;
files?: EnteFile[];
statfs: StatFsFn;
cacheOriginalsMaxBytes?: number;
freeBelowBytes?: number;
isPinned?: (fileID: number) => boolean;
}): { cache: ContentCache; source: SizedSource } => {
const source = (args.source as SizedSource) ?? new SizedSource();
const byID = new Map<number, EnteFile>();
for (const f of args.files ?? [file(1), file(2), file(3), file(4), file(5)])
byID.set(f.id, f);
const cache = new ContentCache({
pools: new RequestPools(),
source,
cacheDirectory: cacheDir,
getFile: (id) => byID.get(id),
statfs: args.statfs,
cacheOriginalsMaxBytes: args.cacheOriginalsMaxBytes,
freeBelowBytes: args.freeBelowBytes,
isPinned: args.isPinned,
});
return { cache, source };
};
// Plenty of free space, so the configured max governs the limit unless a test
// dials it down. `bsize` of 1 makes `bavail` the free byte count.
const abundantFree: StatFsFn = async () => ({
bsize: 1,
bavail: 1_000_000_000,
});
describe("originals eviction", () => {
it("removes least-recently-used originals to fit the limit", async () => {
// Each original is 10 bytes; a 25-byte cap holds two.
const { cache } = buildCache({
statfs: abundantFree,
cacheOriginalsMaxBytes: 25,
freeBelowBytes: 0,
});
await cache.open();
await cache.original(1);
setMtime(1, 1000);
await cache.original(2);
setMtime(2, 2000);
// Writing the third (total 30 > 25) evicts the oldest, file 1.
await cache.original(3);
expect(existsSync(originalPath(1))).toBe(false);
expect(existsSync(originalPath(2))).toBe(true);
expect(existsSync(originalPath(3))).toBe(true);
expect(cache.pathsFor(1).originalPath).toBeUndefined();
expect(cache.originalsStatus().usedBytes).toBe(20);
});
it("skips pinned files, evicting the oldest unpinned instead", async () => {
const { cache } = buildCache({
statfs: abundantFree,
cacheOriginalsMaxBytes: 25,
freeBelowBytes: 0,
isPinned: (id) => id === 1,
});
await cache.open();
await cache.original(1);
setMtime(1, 1000); // oldest, but pinned
await cache.original(2);
setMtime(2, 2000);
await cache.original(3);
// File 1 is oldest but pinned, so file 2 is evicted instead.
expect(existsSync(originalPath(1))).toBe(true);
expect(existsSync(originalPath(2))).toBe(false);
expect(existsSync(originalPath(3))).toBe(true);
});
it("proceeds over-limit when only pinned files remain", async () => {
// A 5-byte cap cannot hold even one 10-byte original.
const { cache } = buildCache({
statfs: abundantFree,
cacheOriginalsMaxBytes: 5,
freeBelowBytes: 0,
isPinned: () => true,
});
await cache.open();
const result = await cache.original(1);
expect(existsSync(result.path)).toBe(true);
const status = cache.originalsStatus();
expect(status.usedBytes).toBe(10);
expect(status.usedBytes).toBeGreaterThan(status.limitBytes ?? 0);
});
it("keeps a just-written original larger than the limit, evicting it only on a later write", async () => {
// A 5-byte cap cannot hold even one 10-byte original, but a fetch must
// never evict the file it just wrote and is about to return.
const { cache } = buildCache({
statfs: abundantFree,
cacheOriginalsMaxBytes: 5,
freeBelowBytes: 0,
});
await cache.open();
const result = await cache.original(1);
// The just-written original survives its own over-budget write and the
// returned path exists on disk.
expect(existsSync(result.path)).toBe(true);
expect(existsSync(originalPath(1))).toBe(true);
expect(cache.originalsStatus().usedBytes).toBe(10);
setMtime(1, 1000);
// A later write finds file 1 eligible and evicts it to make room, while
// the newly written file 2 is itself kept over-limit.
await cache.original(2);
expect(existsSync(originalPath(1))).toBe(false);
expect(existsSync(originalPath(2))).toBe(true);
expect(cache.originalsStatus().usedBytes).toBe(10);
});
it("adapts the limit down as the disk fills and up as space returns", async () => {
// The configured max is generous; free space drives the limit. Eviction
// only kicks in once free space falls below the protected reserve.
let free = 1_000_000;
const statfs: StatFsFn = async () => ({ bsize: 1, bavail: free });
const { cache } = buildCache({
statfs,
cacheOriginalsMaxBytes: 1000,
freeBelowBytes: 100,
});
await cache.open();
// Ample free space: the limit is the configured max and nothing is
// evicted as three 10-byte originals accumulate.
await cache.original(1);
setMtime(1, 1000);
await cache.original(2);
setMtime(2, 2000);
await cache.original(3);
setMtime(3, 3000);
expect(cache.originalsStatus().limitBytes).toBe(1000);
expect(cache.originalsStatus().usedBytes).toBe(30);
// The disk fills: only 95 bytes free, below the 100-byte reserve. The
// next write (used 40) sees limit = min(1000, 40 + 95 - 100) = 35 and
// evicts the oldest original (file 1) to fit.
free = 95;
await cache.original(4);
expect(cache.originalsStatus().limitBytes).toBe(35);
expect(existsSync(originalPath(1))).toBe(false);
expect(cache.originalsStatus().usedBytes).toBe(30);
setMtime(4, 4000);
// Space returns: the limit rises back to the configured max and the
// next write is kept without eviction.
free = 1_000_000;
await cache.original(5);
expect(cache.originalsStatus().limitBytes).toBe(1000);
expect(existsSync(originalPath(4))).toBe(true);
expect(existsSync(originalPath(5))).toBe(true);
});
it("bumps an original's mtime when a read returns its path", async () => {
const { cache } = buildCache({
statfs: abundantFree,
cacheOriginalsMaxBytes: 1000,
freeBelowBytes: 0,
});
await cache.open();
await cache.original(1);
// Age the file well into the past.
setMtime(1, 1000);
expect(statSync(originalPath(1)).mtimeMs).toBeLessThan(2_000_000);
// A second read is a cache hit that must touch the file.
const before = Date.now();
await cache.original(1);
expect(statSync(originalPath(1)).mtimeMs).toBeGreaterThanOrEqual(
before - 2000,
);
});
it("keeps both originals when two over-budget fetches race", async () => {
// A 15-byte cap holds one 10-byte original but not two. Two fetches are
// held in flight together; when both stored files cross the limit,
// neither eviction pass may delete the sibling whose path has not yet
// been returned, so both survive over-limit.
const source = new GatedSource();
const { cache } = buildCache({
source,
statfs: abundantFree,
cacheOriginalsMaxBytes: 15,
freeBelowBytes: 0,
});
await cache.open();
const p1 = cache.original(1);
const p2 = cache.original(2);
await source.bothStarted; // both downloads are in flight before either stores
source.release();
const [r1, r2] = await Promise.all([p1, p2]);
// Both returned paths exist on disk even though together they exceed the
// cap; nothing was evicted out from under a fetch still in progress.
expect(existsSync(r1.path)).toBe(true);
expect(existsSync(r2.path)).toBe(true);
expect(existsSync(originalPath(1))).toBe(true);
expect(existsSync(originalPath(2))).toBe(true);
expect(cache.originalsStatus().usedBytes).toBe(20);
});
it("never counts or evicts thumbnails", async () => {
const { cache } = buildCache({
statfs: abundantFree,
cacheOriginalsMaxBytes: 5,
freeBelowBytes: 0,
});
await cache.open();
await cache.thumbnail(1);
await cache.thumbnail(2);
expect(cache.pathsFor(1).thumbnailPath).toBeDefined();
expect(cache.pathsFor(2).thumbnailPath).toBeDefined();
expect(cache.originalsStatus().usedBytes).toBe(0);
});
});
-160
View File
@@ -1,160 +0,0 @@
/**
* Integration between `Library` and the content cache (issue #46).
*
* The cache itself is covered in `content.test.ts`; this file locks the wiring:
* `Library.open` builds the cache from a content source, `lib.photos` hands out
* `Photo` objects that fetch through it, `lib.thumbnails.ensure` drives it, and
* a cached path shows up on the projected record. A library opened without a
* content source leaves those methods throwing rather than silently doing
* nothing.
*/
import { describe, it, expect, beforeEach, afterEach } from "vitest";
import { mkdtempSync, rmSync, existsSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { Library } from "../../src/library/index.js";
import type { ContentSource } from "../../src/library/content.js";
import type { CollectionsPage, FilesPage } from "../../src/client.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
const USER_ID = 7;
const collection = (id: number): Collection => ({
id,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
name: `album-${id}`,
type: "album",
updationTime: 1,
isShared: false,
});
const file = (id: number, collectionID: number): EnteFile => ({
id,
collectionID,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
metadata: {
title: `file-${id}.jpg`,
fileType: "image",
creationTime: 1,
modificationTime: 1,
},
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime: 1,
});
// A metadata-only client serving one album with one file, once.
class MockClient {
served = false;
whoami(): { email: string; userID: number } {
return { email: "u@example.com", userID: USER_ID };
}
async collectionsSince(): Promise<CollectionsPage> {
if (this.served) return { collections: [], deleted: [], cursor: 1 };
this.served = true;
return { collections: [collection(1)], deleted: [], cursor: 1 };
}
async filesSince(): Promise<FilesPage> {
return { files: [file(1, 1)], deleted: [], cursor: 1 };
}
}
// A content source that writes a marker file and counts thumbnail fetches.
const stubSource = (): ContentSource & { thumbCalls: () => number } => {
let thumbCalls = 0;
return {
thumbCalls: () => thumbCalls,
original: async ({ destination }) => {
writeFileSync(destination, "orig-bytes");
return { bytesWritten: 10 };
},
thumbnail: async ({ destination }) => {
thumbCalls++;
writeFileSync(destination, "thumb");
return { bytesWritten: 5 };
},
};
};
let root: string;
beforeEach(() => {
root = mkdtempSync(join(tmpdir(), "quak-content-lib-"));
});
afterEach(() => {
if (root && existsSync(root))
rmSync(root, { recursive: true, force: true });
});
describe("Library content wiring", () => {
it("fetches a thumbnail through a Photo and records its cache path", async () => {
const source = stubSource();
const lib = await Library.open({
client: new MockClient(),
cacheDirectory: join(root, "cache"),
contentSource: source,
refreshIntervalSeconds: 3600,
// On-demand wiring only; the background precache (#48) is covered
// in precache.test.ts and would race the exact-count assertions.
precacheThumbnails: false,
precacheOriginals: false,
});
const photo = lib.photos.byID({ fileID: 1 });
expect(photo).toBeDefined();
const result = await photo!.thumbnail();
expect(source.thumbCalls()).toBe(1);
expect(result.path).toBe(join(root, "cache", "thumbnails", "1.jpg"));
expect(existsSync(result.path)).toBe(true);
// The cached path is now on the projected record.
expect(lib.photos.byID({ fileID: 1 })!.record().thumbnailPath).toBe(
result.path,
);
lib.close();
});
it("drives thumbnails.ensure through the cache", async () => {
const source = stubSource();
const lib = await Library.open({
client: new MockClient(),
cacheDirectory: join(root, "cache"),
contentSource: source,
refreshIntervalSeconds: 3600,
// On-demand wiring only; the background precache (#48) is covered
// in precache.test.ts and would race the exact-count assertions.
precacheThumbnails: false,
precacheOriginals: false,
});
const results = await lib.thumbnails.ensure({
fileIDs: [1],
priority: "visible",
});
expect(results).toEqual([
{ fileID: 1, path: join(root, "cache", "thumbnails", "1.jpg") },
]);
lib.close();
});
it("throws from content methods when opened without a content source", async () => {
const lib = await Library.open({
client: new MockClient(),
cacheDirectory: join(root, "cache"),
refreshIntervalSeconds: 3600,
});
await expect(
lib.photos.byID({ fileID: 1 })!.thumbnail(),
).rejects.toThrow(/content cache/i);
await expect(
lib.thumbnails.ensure({ fileIDs: [1], priority: "visible" }),
).rejects.toThrow(/content cache/i);
lib.close();
});
});
-423
View File
@@ -1,423 +0,0 @@
/**
* Tests for the on-disk content and thumbnail cache (issue #46).
*
* The cache keys stored bytes by `fileID` under `cacheDirectory`:
* `originals/<fileID>.<ext>` and `thumbnails/<fileID>.<ext>`. Its contract:
*
* 1. **Fetch once, then serve from disk.** The first `original`/`thumbnail`
* fetches through the request pool and stores the bytes; the next finds the
* file present and returns its path with a single `skipped` event and no
* network. A file already sitting in the backup `downloadDirectory` counts
* as present too.
* 2. **Present-means-complete.** Content appears only by the streaming atomic
* writer's rename, so a file that exists is whole. The directory listing
* taken at `open()` is the record of what is cached, and the orphan temp
* files a crashed write may have left are reaped there.
* 3. **`thumbnails.ensure` drives the thumbnail pool with priority, dedup, and
* abort.** A `fileID` asked for twice downloads once; a visible request is
* served ahead of a background one; and an `AbortSignal` drops work still
* queued while letting an in-flight fetch finish.
*
* The `ContentSource` is a stand-in: it writes deterministic bytes to the
* destination and returns the count, so the cache logic is exercised with no
* crypto and no network. Ordering tests gate the stand-in on explicit deferreds
* and assert the persisted result, never a bare call or a timer.
*/
import { describe, it, expect, beforeEach, afterEach } from "vitest";
import {
mkdtempSync,
rmSync,
existsSync,
writeFileSync,
mkdirSync,
statSync,
} from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import {
ContentCache,
type ContentSource,
type EnsureEvent,
} from "../../src/library/content.js";
import { RequestPools } from "../../src/library/pools.js";
import type { EnteFile } from "../../src/model/types.js";
const file = (id: number, title = `file-${id}.jpg`): EnteFile => ({
id,
collectionID: 1,
ownerID: 1,
key: new Uint8Array([id & 0xff]),
metadata: {
title,
fileType: "image",
creationTime: 0,
modificationTime: 0,
},
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime: 0,
});
// A deferred with externally callable resolve, used to gate the stand-in source
// so ordering is controlled by the test rather than by timing.
const deferred = (): { promise: Promise<void>; resolve: () => void } => {
let resolve!: () => void;
const promise = new Promise<void>((r) => {
resolve = r;
});
return { promise, resolve };
};
// A ContentSource that writes `${kind}:${fileID}` bytes to the destination and
// records every call. `gate` optionally blocks a call until released, and
// `completed` records the order in which fetches finished — the observable used
// by the priority and abort tests instead of a timer.
class StubSource implements ContentSource {
originalCalls: number[] = [];
thumbnailCalls: number[] = [];
completed: number[] = [];
emptyFor = new Set<number>();
gates = new Map<number, Promise<void>>();
private async run(
kind: "original" | "thumbnail",
file: EnteFile,
destination: string,
): Promise<{ bytesWritten: number }> {
const gate = this.gates.get(file.id);
if (gate) await gate;
const bytes = this.emptyFor.has(file.id)
? new Uint8Array(0)
: new TextEncoder().encode(`${kind}:${file.id}`);
writeFileSync(destination, bytes);
this.completed.push(file.id);
return { bytesWritten: bytes.length };
}
async original(args: {
file: EnteFile;
destination: string;
}): Promise<{ bytesWritten: number }> {
this.originalCalls.push(args.file.id);
return this.run("original", args.file, args.destination);
}
async thumbnail(args: {
file: EnteFile;
destination: string;
}): Promise<{ bytesWritten: number }> {
this.thumbnailCalls.push(args.file.id);
return this.run("thumbnail", args.file, args.destination);
}
}
let root: string;
let cacheDir: string;
beforeEach(() => {
root = mkdtempSync(join(tmpdir(), "quak-content-"));
cacheDir = join(root, "cache");
});
afterEach(() => {
if (root && existsSync(root))
rmSync(root, { recursive: true, force: true });
});
const buildCache = (
args: {
source?: ContentSource;
files?: EnteFile[];
pools?: RequestPools;
downloadDirectory?: string;
} = {},
): { cache: ContentCache; source: StubSource } => {
const source = (args.source as StubSource) ?? new StubSource();
const byID = new Map<number, EnteFile>();
for (const f of args.files ?? [file(1), file(2), file(3)])
byID.set(f.id, f);
const cache = new ContentCache({
pools: args.pools ?? new RequestPools(),
source,
cacheDirectory: cacheDir,
downloadDirectory: args.downloadDirectory,
getFile: (id) => byID.get(id),
});
return { cache, source };
};
describe("ContentCache.open", () => {
it("creates the cache directories with 0700 permissions", async () => {
const { cache } = buildCache();
await cache.open();
const originals = join(cacheDir, "originals");
const thumbnails = join(cacheDir, "thumbnails");
expect(existsSync(originals)).toBe(true);
expect(existsSync(thumbnails)).toBe(true);
expect(statSync(originals).mode & 0o777).toBe(0o700);
expect(statSync(thumbnails).mode & 0o777).toBe(0o700);
});
it("reaps orphan temp files but keeps complete content", async () => {
const originals = join(cacheDir, "originals");
const thumbnails = join(cacheDir, "thumbnails");
mkdirSync(originals, { recursive: true });
mkdirSync(thumbnails, { recursive: true });
const orphan = join(originals, ".quak-abc123.tmp");
const complete = join(originals, "1.jpg");
const thumb = join(thumbnails, "2.jpg");
writeFileSync(orphan, "half-written");
writeFileSync(complete, "whole");
writeFileSync(thumb, "whole-thumb");
const { cache } = buildCache();
await cache.open();
expect(existsSync(orphan)).toBe(false);
expect(existsSync(complete)).toBe(true);
expect(existsSync(thumb)).toBe(true);
});
it("records already-cached files so their paths appear in pathsFor", async () => {
const originals = join(cacheDir, "originals");
const thumbnails = join(cacheDir, "thumbnails");
mkdirSync(originals, { recursive: true });
mkdirSync(thumbnails, { recursive: true });
writeFileSync(join(originals, "1.jpg"), "orig");
writeFileSync(join(thumbnails, "1.jpg"), "thumb");
const { cache } = buildCache();
await cache.open();
expect(cache.pathsFor(1)).toEqual({
originalPath: join(originals, "1.jpg"),
thumbnailPath: join(thumbnails, "1.jpg"),
});
expect(cache.pathsFor(2)).toEqual({});
});
});
describe("ContentCache.original / thumbnail", () => {
it("fetches once, then serves the cached file with a single skipped event", async () => {
const { cache, source } = buildCache();
await cache.open();
const events: string[] = [];
const first = await cache.original(1, {
onProgress: (e) => events.push(e.status),
});
expect(source.originalCalls).toEqual([1]);
expect(first.path).toBe(join(cacheDir, "originals", "1.jpg"));
expect(first.bytes).toBe("original:1".length);
expect(existsSync(first.path)).toBe(true);
expect(statSync(first.path).mode & 0o777).toBe(0o600);
expect(cache.pathsFor(1).originalPath).toBe(first.path);
const skips: string[] = [];
const second = await cache.original(1, {
onProgress: (e) => skips.push(e.status),
});
// No second download, and exactly one skipped event.
expect(source.originalCalls).toEqual([1]);
expect(second.path).toBe(first.path);
expect(skips).toEqual(["skipped"]);
});
it("serves a file already present in the download directory without fetching", async () => {
const downloadDirectory = join(root, "backup");
mkdirSync(join(downloadDirectory, "originals"), { recursive: true });
const backupPath = join(downloadDirectory, "originals", "1.jpg");
writeFileSync(backupPath, "from-backup");
const { cache, source } = buildCache({ downloadDirectory });
await cache.open();
const events: EnsureEvent["status"][] = [];
const result = await cache.original(1, {
onProgress: (e) => events.push(e.status),
});
expect(source.originalCalls).toEqual([]);
expect(result.path).toBe(backupPath);
expect(result.bytes).toBe("from-backup".length);
expect(events).toEqual(["skipped"]);
});
it("fetches and caches a thumbnail", async () => {
const { cache, source } = buildCache();
await cache.open();
const result = await cache.thumbnail(2);
expect(source.thumbnailCalls).toEqual([2]);
expect(result.path).toBe(join(cacheDir, "thumbnails", "2.jpg"));
expect(existsSync(result.path)).toBe(true);
expect(cache.pathsFor(2).thumbnailPath).toBe(result.path);
});
it("shares one download between concurrent callers for the same file", async () => {
const { cache, source } = buildCache();
await cache.open();
const gate = deferred();
source.gates.set(1, gate.promise);
const a = cache.original(1);
const b = cache.original(1);
gate.resolve();
const [ra, rb] = await Promise.all([a, b]);
expect(source.originalCalls).toEqual([1]);
expect(ra.path).toBe(rb.path);
});
it("does not record a path when the fetched file is empty", async () => {
const { cache, source } = buildCache();
source.emptyFor.add(1);
await cache.open();
await expect(cache.original(1)).rejects.toThrow(/empty/i);
expect(cache.pathsFor(1).originalPath).toBeUndefined();
});
it("rejects an unknown file", async () => {
const { cache } = buildCache({ files: [] });
await cache.open();
await expect(cache.original(999)).rejects.toThrow(/unknown file/i);
});
});
describe("ContentCache.ensureThumbnails", () => {
it("downloads once for a file listed twice and reports every id", async () => {
const { cache, source } = buildCache();
await cache.open();
const results = await cache.ensureThumbnails({
fileIDs: [1, 1, 2],
priority: "visible",
});
expect(source.thumbnailCalls.sort()).toEqual([1, 2]);
expect(results).toEqual([
{ fileID: 1, path: join(cacheDir, "thumbnails", "1.jpg") },
{ fileID: 2, path: join(cacheDir, "thumbnails", "2.jpg") },
]);
});
it("skips present files and reports a skipped event", async () => {
const thumbnails = join(cacheDir, "thumbnails");
mkdirSync(thumbnails, { recursive: true });
writeFileSync(join(thumbnails, "1.jpg"), "present");
const { cache, source } = buildCache();
await cache.open();
const events: EnsureEvent[] = [];
const results = await cache.ensureThumbnails({
fileIDs: [1, 2],
priority: "ahead",
onProgress: (e) => events.push(e),
});
expect(source.thumbnailCalls).toEqual([2]);
expect(results).toEqual([
{ fileID: 1, path: join(thumbnails, "1.jpg") },
{ fileID: 2, path: join(thumbnails, "2.jpg") },
]);
expect(events).toContainEqual({
fileID: 1,
status: "skipped",
path: join(thumbnails, "1.jpg"),
});
});
it("serves a visible request ahead of an already-queued background one", async () => {
// One thumbnail slot, so exactly one fetch runs at a time and the rest
// wait in the pool. A background fetch takes the slot; a background and
// a visible fetch queue behind it. When the slot frees, the pool must
// pick the visible (on-demand) request ahead of the background one that
// was submitted first. The completion order is the observable.
const pools = new RequestPools({ thumbnailConcurrency: 1 });
const { cache, source } = buildCache({ pools });
await cache.open();
const gateA = deferred();
const gateB = deferred();
const gateC = deferred();
source.gates.set(1, gateA.promise);
source.gates.set(2, gateB.promise);
source.gates.set(3, gateC.promise);
const bgFirst = cache.ensureThumbnails({
fileIDs: [1],
priority: "background",
});
// Let fetch 1 take the only slot before the others queue.
await Promise.resolve();
const bgSecond = cache.ensureThumbnails({
fileIDs: [2],
priority: "background",
});
const visible = cache.ensureThumbnails({
fileIDs: [3],
priority: "visible",
});
gateA.resolve();
gateC.resolve();
gateB.resolve();
await Promise.all([bgFirst, bgSecond, visible]);
// 1 ran first (it held the slot). Of the two that were queued, the
// visible id 3 was served before the background id 2.
expect(source.completed).toEqual([1, 3, 2]);
});
it("drops queued work on abort but keeps an in-flight fetch", async () => {
const pools = new RequestPools({ thumbnailConcurrency: 1 });
const { cache, source } = buildCache({ pools });
await cache.open();
const gate = deferred();
source.gates.set(1, gate.promise);
const controller = new AbortController();
const pending = cache.ensureThumbnails({
fileIDs: [1, 2],
priority: "ahead",
signal: controller.signal,
});
// Fetch 1 is in flight (holds the slot); 2 is queued.
await Promise.resolve();
controller.abort();
gate.resolve();
const results = await pending;
// The in-flight fetch finished and is kept; the queued one was dropped
// before it ran.
expect(source.thumbnailCalls).toEqual([1]);
expect(results).toEqual([
{ fileID: 1, path: join(cacheDir, "thumbnails", "1.jpg") },
{ fileID: 2, error: "aborted" },
]);
});
it("captures a per-file failure without failing the batch", async () => {
const { cache } = buildCache({ files: [file(1)] });
await cache.open();
const results = await cache.ensureThumbnails({
fileIDs: [1, 2],
priority: "background",
});
expect(results[0]).toEqual({
fileID: 1,
path: join(cacheDir, "thumbnails", "1.jpg"),
});
expect(results[1]?.fileID).toBe(2);
expect(results[1]?.error).toMatch(/unknown file/i);
});
});
-272
View File
@@ -1,272 +0,0 @@
/**
* Tests for fresh reads (issue #75, an owner amendment to design #36).
*
* The default read namespaces answer from RAM and never touch the network; a
* background loop keeps the local copy current. `fresh()` adds an awaited path:
* it forces a refresh, waits for it to complete and persist, and only then
* hands back the `albums`/`photos`/`timeline` namespaces, guaranteeing the
* local copy reflects a completed server round-trip. The contracts here:
*
* 1. A fresh read observes a server change that a same-instant default read
* would miss (the default path has not refreshed yet).
* 2. Concurrent fresh reads coalesce onto one in-flight refresh — three
* concurrent `fresh()` calls make exactly one collections round-trip, not
* three.
* 3. A refresh that fails rejects the fresh read (currency was unavailable),
* while the default reads stay silent and keep serving the last good copy.
*
* The client is the same metadata-only mock the background-refresh tests use:
* no crypto, no network, scripted pages, and a record of each call. A long
* refresh interval keeps the background timer out of the way so each test's
* refreshes are exactly the ones it triggers.
*/
import { describe, it, expect, beforeEach, afterEach } from "vitest";
import { mkdtempSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { Library } from "../../src/library/index.js";
import type { CollectionsPage, FilesPage } from "../../src/client.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
const USER_ID = 42;
// Long enough that no background tick fires during a test; each test's
// refreshes are only the ones its own `fresh()` calls force.
const SLOW_INTERVAL = 3600;
const collection = (id: number, updationTime: number): Collection => ({
id,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
name: `album-${id}`,
type: "album",
updationTime,
isShared: false,
});
const file = (
id: number,
collectionID: number,
updationTime: number,
): EnteFile => ({
id,
collectionID,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
metadata: {
title: `file-${id}.jpg`,
fileType: "image",
creationTime: updationTime,
modificationTime: updationTime,
},
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime,
});
class MockClient {
userID = USER_ID;
failCollections = false;
collectionsQueue: CollectionsPage[] = [];
filesByCollection = new Map<number, FilesPage[]>();
collectionsSinceTimes: number[] = [];
filesCalls: { collectionID: number; sinceTime: number }[] = [];
whoami(): { email: string; userID: number } {
return { email: "user@example.com", userID: this.userID };
}
async collectionsSince(args: {
sinceTime: number;
}): Promise<CollectionsPage> {
this.collectionsSinceTimes.push(args.sinceTime);
if (this.failCollections) throw new Error("network down");
return (
this.collectionsQueue.shift() ?? {
collections: [],
deleted: [],
cursor: args.sinceTime,
}
);
}
async filesSince(args: {
collectionID: number;
collectionKey: Uint8Array;
sinceTime: number;
}): Promise<FilesPage> {
this.filesCalls.push({
collectionID: args.collectionID,
sinceTime: args.sinceTime,
});
const queue = this.filesByCollection.get(args.collectionID);
return (
queue?.shift() ?? {
files: [],
deleted: [],
cursor: args.sinceTime,
}
);
}
filesFor(collectionID: number, ...pages: FilesPage[]): void {
this.filesByCollection.set(collectionID, pages);
}
}
// A client seeded with one collection and one file, opened with an empty cache
// so the initial refresh is awaited and post-`open()` state is deterministic.
const openSeeded = async (
cacheDirectory: string,
): Promise<{ client: MockClient; lib: Library }> => {
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 90)],
deleted: [],
cursor: 90,
});
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: SLOW_INTERVAL,
});
return { client, lib };
};
describe("Library.fresh", () => {
let dir: string;
let cacheDirectory: string;
beforeEach(() => {
dir = mkdtempSync(join(tmpdir(), "quak-fresh-"));
cacheDirectory = join(dir, "cache");
});
afterEach(() => {
rmSync(dir, { recursive: true, force: true });
});
it("observes a server change a same-instant default read would miss", async () => {
const { client, lib } = await openSeeded(cacheDirectory);
try {
// A new file appears on the server after open, advancing its
// collection so the next refresh re-enumerates it.
client.collectionsQueue.push({
collections: [collection(1, 200)],
deleted: [],
cursor: 200,
});
client.filesFor(1, {
files: [file(1002, 1, 190)],
deleted: [],
cursor: 190,
});
// A default read at this instant has not refreshed: it misses 1002.
expect(lib.photos.byID({ fileID: 1002 })).toBeUndefined();
// A fresh read forces the round-trip and sees it.
const reads = await lib.fresh();
expect(reads.photos.byID({ fileID: 1002 })?.fileID).toBe(1002);
// And the change is now live for the default namespaces too.
expect(lib.photos.byID({ fileID: 1002 })?.fileID).toBe(1002);
} finally {
lib.close();
}
});
it("coalesces concurrent fresh reads onto one in-flight refresh", async () => {
const { client, lib } = await openSeeded(cacheDirectory);
try {
client.collectionsQueue.push({
collections: [collection(1, 200)],
deleted: [],
cursor: 200,
});
client.filesFor(1, {
files: [file(1002, 1, 190)],
deleted: [],
cursor: 190,
});
const collectionsBefore = client.collectionsSinceTimes.length;
const filesBefore = client.filesCalls.length;
// Gate the next collections fetch so all three fresh reads are in
// flight together before any of them completes.
let release: () => void = () => {};
const gate = new Promise<void>((r) => {
release = r;
});
const inner = client.collectionsSince.bind(client);
client.collectionsSince = async (args: { sinceTime: number }) => {
await gate;
return inner(args);
};
const all = Promise.all([lib.fresh(), lib.fresh(), lib.fresh()]);
release();
const [a, b, c] = await all;
// Exactly one collections round-trip and one file round-trip served
// all three fresh reads.
expect(client.collectionsSinceTimes.length).toBe(
collectionsBefore + 1,
);
expect(client.filesCalls.length).toBe(filesBefore + 1);
// All three observed the change.
for (const reads of [a, b, c]) {
expect(reads.photos.byID({ fileID: 1002 })?.fileID).toBe(1002);
}
} finally {
lib.close();
}
});
it("rejects the fresh read when the refresh fails, leaving defaults intact", async () => {
const { client, lib } = await openSeeded(cacheDirectory);
try {
client.failCollections = true;
// Two concurrent fresh reads both reject, and share one failed
// round-trip rather than each making its own.
const collectionsBefore = client.collectionsSinceTimes.length;
const first = lib.fresh();
const second = lib.fresh();
await expect(first).rejects.toThrow(/network down/);
await expect(second).rejects.toThrow(/network down/);
expect(client.collectionsSinceTimes.length).toBe(
collectionsBefore + 1,
);
// The default reads never rejected: they still serve the last good
// copy, and the failure surfaced through status().
expect(lib.photos.byID({ fileID: 1001 })?.fileID).toBe(1001);
expect(lib.status().lastError).toMatch(/network down/);
// Recovery: once the server answers, a fresh read resolves current.
client.failCollections = false;
client.collectionsQueue.push({
collections: [collection(2, 300)],
deleted: [],
cursor: 300,
});
const reads = await lib.fresh();
expect(reads.albums.byID({ collectionID: 2 })?.collectionID).toBe(
2,
);
expect(lib.status().lastError).toBeUndefined();
} finally {
lib.close();
}
});
});
-669
View File
@@ -1,669 +0,0 @@
/**
* Tests for `Library.open()` and its transparent background refresh loop.
*
* The library keeps the account's server state in a `MetadataStore` (issue
* #41) and pulls changes with the resumable, tombstone-aware enumerators on
* `Client` (issue #38: `collectionsSince` / `filesSince`). `open()` loads the
* cache, does one refresh, then refreshes again every `refreshIntervalSeconds`
* on a background timer. The design (#36) forbids an exposed `sync()`, a
* `serverReachable` flag, a `lib.refresh()` method, and a "before each read"
* mode. The contracts exercised here:
*
* 1. Reads are answered from RAM. A read never calls the client.
* 2. `open()` does an initial refresh, then the interval keeps refreshing;
* each refresh resumes from the stored cursor and applies diffs + tombstones.
* 3. The cache is rewritten only when a refresh actually changes something.
* 4. A failed refresh is invisible to reads: the last good data stays, the
* failure surfaces via `onProgress` ("failed") and `status()`, and a later
* success clears the error. `open()` itself resolves even when the first
* refresh fails (offline start from cache).
* 5. `close()` stops the timer and is idempotent.
* 6. `cacheDirectory` defaults to the env-paths cache dir plus the user id.
* 7. `open()` branches on the cache: an empty cache awaits the first refresh
* (it has nothing to serve yet); an existing cache serves its copy at once
* and refreshes in the background, so a slow or dead server never stalls
* opening.
* 8. A save failure that leaves RAM ahead of disk keeps `status().lastError`
* set and keeps retrying the write; a later empty refresh does not clear it.
*
* The client is a mock: no crypto, no network. It serves scripted pages and
* records the `sinceTime` each call carried so cursor threading is provable.
*
* On an empty cache `open()` awaits the initial refresh (including its cache
* write), so state right after `open()` is deterministic; the tests that
* inspect post-`open()` state seed no cache and rely on that. Tests for an
* existing-cache open seed a store first and prove `open()` returns without
* waiting for the network. The interval tests then use real timers with a
* short interval and `vi.waitFor`: a fake clock cannot settle the real
* fsync-and-rename cache write, and empty diffs never write, so the eventual
* state is stable to poll for.
*/
import { describe, it, expect, beforeEach, afterEach, vi } from "vitest";
import { mkdtempSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import envPaths from "env-paths";
import { Library, type RefreshEvent } from "../../src/library/index.js";
import { MetadataStore } from "../../src/library/store.js";
import type { CollectionsPage, FilesPage } from "../../src/client.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
const USER_ID = 42;
// Short enough that a couple of ticks pass within a test, long enough not to
// spin; interval tests poll for the eventual state rather than counting ticks.
const FAST_INTERVAL = 0.02;
const collection = (
id: number,
updationTime: number,
name = `album-${id}`,
): Collection => ({
id,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
name,
type: "album",
updationTime,
isShared: false,
});
const file = (
id: number,
collectionID: number,
updationTime: number,
): EnteFile => ({
id,
collectionID,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
metadata: {
title: `file-${id}.jpg`,
fileType: "image",
creationTime: updationTime,
modificationTime: updationTime,
},
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime,
});
/**
* A mock `Client`. `collectionsSince` shifts one page off `collectionsQueue`
* per call (an empty diff that advances nothing when the queue runs dry);
* `filesSince` shifts from a per-collection queue. `failCollections` makes the
* next and all further collection fetches throw, to simulate an offline server.
*/
class MockClient {
userID = USER_ID;
failCollections = false;
collectionsQueue: CollectionsPage[] = [];
filesByCollection = new Map<number, FilesPage[]>();
collectionsSinceTimes: number[] = [];
filesCalls: { collectionID: number; sinceTime: number }[] = [];
whoami(): { email: string; userID: number } {
return { email: "user@example.com", userID: this.userID };
}
async collectionsSince(args: {
sinceTime: number;
}): Promise<CollectionsPage> {
this.collectionsSinceTimes.push(args.sinceTime);
if (this.failCollections) throw new Error("network down");
return (
this.collectionsQueue.shift() ?? {
collections: [],
deleted: [],
cursor: args.sinceTime,
}
);
}
async filesSince(args: {
collectionID: number;
collectionKey: Uint8Array;
sinceTime: number;
}): Promise<FilesPage> {
this.filesCalls.push({
collectionID: args.collectionID,
sinceTime: args.sinceTime,
});
const queue = this.filesByCollection.get(args.collectionID);
return (
queue?.shift() ?? {
files: [],
deleted: [],
cursor: args.sinceTime,
}
);
}
filesFor(collectionID: number, ...pages: FilesPage[]): void {
this.filesByCollection.set(collectionID, pages);
}
}
describe("Library.open and background refresh", () => {
let dir: string;
let cacheDirectory: string;
beforeEach(() => {
dir = mkdtempSync(join(tmpdir(), "quak-library-"));
cacheDirectory = join(dir, "cache");
});
afterEach(() => {
rmSync(dir, { recursive: true, force: true });
});
it("does an initial refresh and answers reads from the cache", async () => {
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 90), file(1002, 1, 95)],
deleted: [],
cursor: 95,
});
const lib = await Library.open({ client, cacheDirectory });
try {
expect(lib.listCollections().map((c) => c.id)).toEqual([1]);
expect(lib.listFiles(1).map((f) => f.id)).toEqual([1001, 1002]);
expect(lib.getFile(1, 1001)?.metadata.title).toBe("file-1001.jpg");
const status = lib.status();
expect(status.userID).toBe(USER_ID);
expect(status.collections).toBe(1);
expect(status.files).toBe(2);
expect(status.lastRefreshAt).toBeGreaterThan(0);
expect(status.lastError).toBeUndefined();
// The initial refresh persisted the cache to disk.
const reloaded = await MetadataStore.load(
join(cacheDirectory, "metadata.json"),
);
expect(reloaded.getFile(1, 1001)?.id).toBe(1001);
expect(reloaded.collectionsSinceTime).toBe(100);
} finally {
lib.close();
}
});
it("reads never call the client", async () => {
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 90)],
deleted: [],
cursor: 90,
});
const lib = await Library.open({ client, cacheDirectory });
try {
const collectionCalls = client.collectionsSinceTimes.length;
const fileCalls = client.filesCalls.length;
lib.listCollections();
lib.getCollection(1);
lib.listFiles(1);
lib.getFile(1, 1001);
lib.status();
expect(client.collectionsSinceTimes.length).toBe(collectionCalls);
expect(client.filesCalls.length).toBe(fileCalls);
} finally {
lib.close();
}
});
it("resumes each refresh from the stored cursor", async () => {
// Seed a cache with a cursor and a collection, as a prior run left it.
const path = join(cacheDirectory, "metadata.json");
const seed = await MetadataStore.load(path);
seed.userID = USER_ID;
seed.collectionsSinceTime = 500;
seed.putCollection(collection(1, 400));
seed.putFile(file(1001, 1, 400));
await seed.save();
const client = new MockClient();
// The collection's updationTime advances (400 -> 600), so its files are
// re-enumerated from the collection's stored updationTime (400).
client.collectionsQueue.push({
collections: [collection(1, 600)],
deleted: [],
cursor: 600,
});
client.filesFor(1, {
files: [file(1002, 1, 550)],
deleted: [],
cursor: 550,
});
// Opening from an existing cache serves the seeded copy at once and
// refreshes in the background, so the refresh's effects are polled for.
const lib = await Library.open({ client, cacheDirectory });
try {
await vi.waitFor(
() => {
// Collections resumed from the stored cursor, and files were
// re-enumerated from the stored collection updationTime.
expect(client.collectionsSinceTimes[0]).toBe(500);
expect(client.filesCalls).toEqual([
{ collectionID: 1, sinceTime: 400 },
]);
expect(lib.listFiles(1).map((f) => f.id)).toEqual([
1001, 1002,
]);
},
{ timeout: 2000, interval: 5 },
);
} finally {
lib.close();
}
});
it("does not re-enumerate a collection whose updationTime did not advance", async () => {
const path = join(cacheDirectory, "metadata.json");
const seed = await MetadataStore.load(path);
seed.userID = USER_ID;
seed.collectionsSinceTime = 100;
seed.putCollection(collection(1, 400));
await seed.save();
const client = new MockClient();
// The collection comes back in the diff (its metadata changed) but at
// the same updationTime, so its files must not be re-fetched.
client.collectionsQueue.push({
collections: [collection(1, 400, "renamed")],
deleted: [],
cursor: 400,
});
// Existing cache: the rename lands via the background refresh.
const lib = await Library.open({ client, cacheDirectory });
try {
await vi.waitFor(
() => expect(lib.getCollection(1)?.name).toBe("renamed"),
{ timeout: 2000, interval: 5 },
);
// The collection's updationTime did not advance, so its files were
// never re-fetched.
expect(client.filesCalls).toEqual([]);
} finally {
lib.close();
}
});
it("applies diffs and tombstones on the interval", async () => {
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100), collection(2, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 90)],
deleted: [],
cursor: 90,
});
client.filesFor(2, {
files: [file(2001, 2, 90)],
deleted: [],
cursor: 90,
});
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: FAST_INTERVAL,
});
try {
expect(lib.listCollections().map((c) => c.id)).toEqual([1, 2]);
expect(lib.listFiles(2).map((f) => f.id)).toEqual([2001]);
// Next refresh: collection 2 is tombstoned; collection 1 gains a
// file and loses its old one.
client.filesFor(1, {
files: [file(1002, 1, 190)],
deleted: [1001],
cursor: 190,
});
client.collectionsQueue.push({
collections: [collection(1, 200)],
deleted: [2],
cursor: 200,
});
await vi.waitFor(
() => {
expect(lib.listCollections().map((c) => c.id)).toEqual([1]);
expect(lib.listFiles(1).map((f) => f.id)).toEqual([1002]);
// Collection 2's files went with it.
expect(lib.listFiles(2)).toEqual([]);
},
{ timeout: 2000, interval: 5 },
);
} finally {
lib.close();
}
});
it("rewrites the cache only when a refresh changes something", async () => {
const saveSpy = vi.spyOn(MetadataStore.prototype, "save");
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 90)],
deleted: [],
cursor: 90,
});
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: FAST_INTERVAL,
});
try {
// The initial refresh changed everything, so it saved once.
expect(saveSpy).toHaveBeenCalledTimes(1);
// Several empty-diff ticks pass; none of them may rewrite the file.
await new Promise((r) => setTimeout(r, FAST_INTERVAL * 1000 * 4));
expect(saveSpy).toHaveBeenCalledTimes(1);
// A real change triggers exactly one more rewrite; later empty ticks
// still do not, so the count settles at two.
client.collectionsQueue.push({
collections: [collection(2, 200)],
deleted: [],
cursor: 200,
});
await vi.waitFor(() => expect(saveSpy).toHaveBeenCalledTimes(2), {
timeout: 2000,
interval: 5,
});
await new Promise((r) => setTimeout(r, FAST_INTERVAL * 1000 * 4));
expect(saveSpy).toHaveBeenCalledTimes(2);
} finally {
lib.close();
saveSpy.mockRestore();
}
});
it("keeps a failed refresh invisible to reads and recovers later", async () => {
const events: RefreshEvent[] = [];
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 90)],
deleted: [],
cursor: 90,
});
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: FAST_INTERVAL,
onProgress: (e) => events.push(e),
});
try {
expect(lib.listFiles(1).map((f) => f.id)).toEqual([1001]);
// The server goes away; refreshes now fail.
client.failCollections = true;
await vi.waitFor(
() => expect(lib.status().lastError).toMatch(/network down/),
{ timeout: 2000, interval: 5 },
);
// Reads still see the last good data; the failure was reported.
expect(lib.listFiles(1).map((f) => f.id)).toEqual([1001]);
expect(
events.some(
(e) => e.operation === "refresh" && e.status === "failed",
),
).toBe(true);
// Recovery: a later refresh succeeds and clears the error.
client.failCollections = false;
client.collectionsQueue.push({
collections: [collection(2, 300)],
deleted: [],
cursor: 300,
});
await vi.waitFor(
() => {
expect(lib.status().lastError).toBeUndefined();
expect(lib.listCollections().map((c) => c.id)).toEqual([
1, 2,
]);
},
{ timeout: 2000, interval: 5 },
);
} finally {
lib.close();
}
});
it("resolves open() even when the first refresh fails", async () => {
const client = new MockClient();
client.failCollections = true;
const events: RefreshEvent[] = [];
const lib = await Library.open({
client,
cacheDirectory,
onProgress: (e) => events.push(e),
});
try {
// Nothing was cached and the server is unreachable: reads are empty,
// but the library opened and the failure is on record.
expect(lib.listCollections()).toEqual([]);
expect(lib.status().lastError).toMatch(/network down/);
expect(lib.status().lastRefreshAt).toBeUndefined();
expect(
events.some(
(e) => e.operation === "refresh" && e.status === "failed",
),
).toBe(true);
} finally {
lib.close();
}
});
it("opens from an existing cache without waiting for the first refresh", async () => {
// Seed a cache as a prior run left it.
const path = join(cacheDirectory, "metadata.json");
const seed = await MetadataStore.load(path);
seed.userID = USER_ID;
seed.collectionsSinceTime = 500;
seed.putCollection(collection(1, 400));
seed.putFile(file(1001, 1, 400));
await seed.save();
// The server never answers this run's first refresh.
const client = new MockClient();
client.collectionsSince = () => new Promise<CollectionsPage>(() => {});
// open() must resolve from the cache without blocking on the network,
// and reads must serve the seeded copy.
const lib = await Library.open({ client, cacheDirectory });
try {
expect(lib.listCollections().map((c) => c.id)).toEqual([1]);
expect(lib.listFiles(1).map((f) => f.id)).toEqual([1001]);
// The first refresh is still outstanding: nothing has completed or
// failed yet.
expect(lib.status().lastRefreshAt).toBeUndefined();
expect(lib.status().lastError).toBeUndefined();
} finally {
lib.close();
}
});
it("awaits the first refresh on a first run with an empty cache", async () => {
// No cache on disk: open() must not resolve until the first fetch does,
// so it never hands back an empty library it could have filled.
let releaseFirstFetch: (page: CollectionsPage) => void = () => {};
const gate = new Promise<CollectionsPage>((resolve) => {
releaseFirstFetch = resolve;
});
const client = new MockClient();
client.filesFor(1, {
files: [file(1001, 1, 90)],
deleted: [],
cursor: 90,
});
client.collectionsSince = async (args: { sinceTime: number }) => {
client.collectionsSinceTimes.push(args.sinceTime);
return gate;
};
let opened = false;
const openPromise = Library.open({ client, cacheDirectory }).then(
(l) => {
opened = true;
return l;
},
);
// While the first fetch is outstanding, open() has not resolved.
await new Promise((r) => setTimeout(r, 20));
expect(opened).toBe(false);
// Completing the fetch lets open() resolve with the data in place.
releaseFirstFetch({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
const lib = await openPromise;
try {
expect(opened).toBe(true);
expect(lib.listCollections().map((c) => c.id)).toEqual([1]);
expect(lib.listFiles(1).map((f) => f.id)).toEqual([1001]);
expect(lib.status().lastRefreshAt).toBeGreaterThan(0);
} finally {
lib.close();
}
});
it("keeps a save failure visible until a save actually succeeds", async () => {
const saveSpy = vi
.spyOn(MetadataStore.prototype, "save")
.mockRejectedValue(new Error("disk full"));
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 90)],
deleted: [],
cursor: 90,
});
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: FAST_INTERVAL,
});
try {
// The initial refresh mutated RAM but its save failed, so the error
// is on record and no refresh has counted as successful.
expect(lib.status().lastError).toMatch(/disk full/);
expect(lib.status().lastRefreshAt).toBeUndefined();
// Empty-diff ticks pass. Each still retries the unsaved write and
// still fails, so the error never silently clears and the refresh
// clock never advances — RAM must not run ahead of disk unnoticed.
const savesBefore = saveSpy.mock.calls.length;
await new Promise((r) => setTimeout(r, FAST_INTERVAL * 1000 * 4));
expect(saveSpy.mock.calls.length).toBeGreaterThan(savesBefore);
expect(lib.status().lastError).toMatch(/disk full/);
expect(lib.status().lastRefreshAt).toBeUndefined();
// Once the disk recovers, the next tick persists the pending change
// and only then clears the error and advances the clock.
saveSpy.mockRestore();
await vi.waitFor(
() => {
expect(lib.status().lastError).toBeUndefined();
expect(lib.status().lastRefreshAt).toBeGreaterThan(0);
},
{ timeout: 2000, interval: 5 },
);
const reloaded = await MetadataStore.load(
join(cacheDirectory, "metadata.json"),
);
expect(reloaded.getFile(1, 1001)?.id).toBe(1001);
} finally {
lib.close();
saveSpy.mockRestore();
}
});
it("close() stops the timer and is idempotent", async () => {
const client = new MockClient();
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: FAST_INTERVAL,
});
const callsAfterOpen = client.collectionsSinceTimes.length;
lib.close();
lib.close(); // second close must not throw
expect(lib.status().closed).toBe(true);
// No further refreshes fire once closed.
await new Promise((r) => setTimeout(r, FAST_INTERVAL * 1000 * 5));
expect(client.collectionsSinceTimes.length).toBe(callsAfterOpen);
});
it("defaults cacheDirectory to the env-paths cache dir plus user id", async () => {
const xdg = join(dir, "xdg-cache");
const prev = process.env.XDG_CACHE_HOME;
process.env.XDG_CACHE_HOME = xdg;
try {
const client = new MockClient();
const lib = await Library.open({ client });
try {
const expected = join(
envPaths("quak", { suffix: "" }).cache,
String(USER_ID),
);
expect(lib.cacheDirectory).toBe(expected);
expect(lib.cacheDirectory.startsWith(xdg)).toBe(true);
expect(lib.cacheDirectory.endsWith(String(USER_ID))).toBe(true);
} finally {
lib.close();
}
} finally {
if (prev === undefined) delete process.env.XDG_CACHE_HOME;
else process.env.XDG_CACHE_HOME = prev;
}
});
});
-449
View File
@@ -1,449 +0,0 @@
/**
* Tests for the ML-data cache and its derived CLIP index (issue #49).
*
* Two layers are exercised:
*
* 1. `MLDataStore` on its own: storing one payload file per fileID (present
* means complete), building a `clip.f32` + `clip.json` index that reloads
* in a single read, rebuilding that index from the payloads when it is
* missing or disagrees with the files present, appending as new payloads
* arrive, overwriting a refetched file in place, and deciding what to
* (re)fetch as `updationTime` advances.
*
* 2. `Library` wiring: after each refresh the library fetches ML data through
* the metadata pool for every known file not yet cached, is incremental on
* later refreshes, and refetches a file whose `updationTime` advanced.
*
* Embedding values are chosen to be exactly representable as float32 so the
* round-trip through `clip.f32` compares equal.
*/
import { describe, it, expect, beforeEach, afterEach, vi } from "vitest";
import { existsSync, mkdtempSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { MLDataStore } from "../../src/library/mldata.js";
import { Library } from "../../src/library/index.js";
import type { CollectionsPage, FilesPage } from "../../src/client.js";
import type { MLData } from "../../src/mldata-fetch.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
// A payload shaped like Ente's: a CLIP embedding plus face data that only the
// on-disk payload carries (never the RAM index).
const payload = (embedding: number[]): MLData => ({
face: {
faces: [{ faceID: "f", detection: { box: { x: 0.5 } } }],
},
clip: { embedding },
});
describe("MLDataStore", () => {
let dir: string;
beforeEach(() => {
dir = mkdtempSync(join(tmpdir(), "quak-mldata-"));
});
afterEach(() => {
rmSync(dir, { recursive: true, force: true });
});
it("stores one payload file per fileID and builds a one-read index", async () => {
const store = await MLDataStore.open(dir);
const res = await store.storeFetched(
new Map([
[100, payload([0.5, 0.25, 0.75])],
[200, payload([1, -2, 0.5])],
]),
new Map([
[100, 10],
[200, 20],
]),
);
expect(res).toEqual({ stored: 2, indexed: 2 });
// One payload file per fileID, and the derived index files.
expect(existsSync(join(dir, "100.json"))).toBe(true);
expect(existsSync(join(dir, "200.json"))).toBe(true);
expect(existsSync(join(dir, "clip.f32"))).toBe(true);
expect(existsSync(join(dir, "clip.json"))).toBe(true);
// Reopening loads the index from disk in one read.
const reopened = await MLDataStore.open(dir);
const index = reopened.getIndex();
expect(index.fileIDs).toEqual([100, 200]);
expect(index.embeddingLength).toBe(3);
expect([...index.embeddings]).toEqual([0.5, 0.25, 0.75, 1, -2, 0.5]);
// The full payload (face boxes) is read back from disk on demand.
const full = await reopened.readPayload(100);
expect(full?.face).toBeDefined();
expect(await reopened.readPayload(999)).toBeUndefined();
});
it("rebuilds the index from payloads when it is missing", async () => {
const store = await MLDataStore.open(dir);
await store.storeFetched(
new Map([[100, payload([0.5, 0.25, 0.75])]]),
new Map([[100, 10]]),
);
// The derived index is lost but the payloads survive.
rmSync(join(dir, "clip.f32"));
rmSync(join(dir, "clip.json"));
const reopened = await MLDataStore.open(dir);
const index = reopened.getIndex();
expect(index.fileIDs).toEqual([100]);
expect([...index.embeddings]).toEqual([0.5, 0.25, 0.75]);
expect(existsSync(join(dir, "clip.f32"))).toBe(true);
});
it("rebuilds the index when it disagrees with the files present", async () => {
const store = await MLDataStore.open(dir);
await store.storeFetched(
new Map([
[100, payload([0.5, 0.25, 0.75])],
[200, payload([1, -2, 0.5])],
]),
new Map([
[100, 10],
[200, 20],
]),
);
// A payload disappears out from under the index (leaving it referencing
// a file no longer present); the index must be rebuilt from what is
// actually on disk.
rmSync(join(dir, "200.json"));
const reopened = await MLDataStore.open(dir);
expect(reopened.getIndex().fileIDs).toEqual([100]);
});
it("rebuilds the index when a payload on disk is missing from it", async () => {
const store = await MLDataStore.open(dir);
await store.storeFetched(
new Map([[100, payload([0.5, 0.25, 0.75])]]),
new Map([[100, 10]]),
);
// A crash between storeFetched renaming a payload into place and
// rewriting the index leaves the payload complete on disk but absent
// from clip.json. Write a second payload directly to reproduce that
// torn state without touching the index.
writeFileSync(
join(dir, "200.json"),
JSON.stringify(payload([1, -2, 0.5])),
);
// Reopening self-heals with no manual delete: the index is rebuilt from
// the payloads to include the orphaned embedding.
const reopened = await MLDataStore.open(dir);
const index = reopened.getIndex();
expect(index.fileIDs).toEqual([100, 200]);
expect([...index.embeddings]).toEqual([0.5, 0.25, 0.75, 1, -2, 0.5]);
});
it("appends new payloads and overwrites a refetched file in place", async () => {
const store = await MLDataStore.open(dir);
await store.storeFetched(
new Map([[100, payload([0.5, 0.25, 0.75])]]),
new Map([[100, 10]]),
);
// A later batch adds a new file: appended after the first.
await store.storeFetched(
new Map([[200, payload([1, -2, 0.5])]]),
new Map([[200, 20]]),
);
// Refetching 100 (its embedding changed) updates it in place, not a
// duplicate row.
await store.storeFetched(
new Map([[100, payload([9, 9, 9])]]),
new Map([[100, 30]]),
);
const index = store.getIndex();
expect(index.fileIDs).toEqual([100, 200]);
expect([...index.embeddings]).toEqual([9, 9, 9, 1, -2, 0.5]);
});
it("keeps a payload without a CLIP embedding out of the index", async () => {
const store = await MLDataStore.open(dir);
const res = await store.storeFetched(
new Map<number, MLData>([[100, { face: { faces: [] } }]]),
new Map([[100, 10]]),
);
expect(res.stored).toBe(1);
expect(res.indexed).toBe(0);
// The payload is still cached (present means complete).
expect(existsSync(join(dir, "100.json"))).toBe(true);
expect(store.getIndex().fileIDs).toEqual([]);
});
it("fetches only what is missing or has a newer updationTime", async () => {
const store = await MLDataStore.open(dir);
await store.storeFetched(
new Map([[100, payload([0.5, 0.25, 0.75])]]),
new Map([[100, 10]]),
);
// 100 is cached and current; 200 has never been fetched.
expect(
store.neededFor([
{ id: 100, updationTime: 10 },
{ id: 200, updationTime: 5 },
]),
).toEqual([200]);
// 100's updationTime advanced past what it was fetched at: refetch.
expect(store.neededFor([{ id: 100, updationTime: 15 }])).toEqual([100]);
// Nothing advanced: nothing to fetch.
expect(store.neededFor([{ id: 100, updationTime: 10 }])).toEqual([]);
});
it("survives a corrupt index without losing the payloads", async () => {
const store = await MLDataStore.open(dir);
await store.storeFetched(
new Map([[100, payload([0.5, 0.25, 0.75])]]),
new Map([[100, 10]]),
);
writeFileSync(join(dir, "clip.json"), "not json");
const reopened = await MLDataStore.open(dir);
expect(reopened.getIndex().fileIDs).toEqual([100]);
});
});
// --- Library wiring ---------------------------------------------------------
const USER_ID = 42;
const FAST_INTERVAL = 0.02;
const collection = (id: number, updationTime: number): Collection => ({
id,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
name: `album-${id}`,
type: "album",
updationTime,
isShared: false,
});
const file = (
id: number,
collectionID: number,
updationTime: number,
): EnteFile => ({
id,
collectionID,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
metadata: {
title: `file-${id}.jpg`,
fileType: "image",
creationTime: updationTime,
modificationTime: updationTime,
},
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime,
});
// A mock client that serves scripted collection/file pages and per-file ML
// payloads, recording every ML fetch request so incremental behaviour is
// provable.
class MLMockClient {
userID = USER_ID;
collectionsQueue: CollectionsPage[] = [];
filesByCollection = new Map<number, FilesPage[]>();
mlByFile = new Map<number, MLData>();
mlFetchCalls: number[][] = [];
whoami(): { email: string; userID: number } {
return { email: "user@example.com", userID: this.userID };
}
async collectionsSince(args: {
sinceTime: number;
}): Promise<CollectionsPage> {
return (
this.collectionsQueue.shift() ?? {
collections: [],
deleted: [],
cursor: args.sinceTime,
}
);
}
async filesSince(args: {
collectionID: number;
collectionKey: Uint8Array;
sinceTime: number;
}): Promise<FilesPage> {
const queue = this.filesByCollection.get(args.collectionID);
return (
queue?.shift() ?? {
files: [],
deleted: [],
cursor: args.sinceTime,
}
);
}
async fetchMLData(args: {
fileIDs: number[];
fileKeys: Map<number, Uint8Array>;
}): Promise<Map<number, MLData>> {
this.mlFetchCalls.push([...args.fileIDs]);
const result = new Map<number, MLData>();
for (const id of args.fileIDs) {
const p = this.mlByFile.get(id);
if (p) result.set(id, p);
}
return result;
}
filesFor(collectionID: number, ...pages: FilesPage[]): void {
this.filesByCollection.set(collectionID, pages);
}
}
describe("Library ML-data fetch on refresh", () => {
let dir: string;
let cacheDirectory: string;
beforeEach(() => {
dir = mkdtempSync(join(tmpdir(), "quak-lib-mldata-"));
cacheDirectory = join(dir, "cache");
});
afterEach(() => {
rmSync(dir, { recursive: true, force: true });
});
it("fetches, stores and indexes ML data for known files, then is incremental", async () => {
const client = new MLMockClient();
client.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 90), file(1002, 1, 95)],
deleted: [],
cursor: 95,
});
client.mlByFile.set(1001, payload([0.5, 0.25, 0.75]));
client.mlByFile.set(1002, payload([1, -2, 0.5]));
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: FAST_INTERVAL,
});
try {
// Wait on `lastMLFetchAt`, set only once the pass has persisted the
// index and payloads — not on the in-RAM counts, which advance
// before `storeFetched` writes to disk, so the reopen below reads
// the committed index rather than racing the write.
await vi.waitFor(
() => {
expect(lib.status().lastMLFetchAt).toBeGreaterThan(0);
expect(lib.status().mlIndexed).toBe(2);
expect(lib.status().mlStored).toBe(2);
},
{ timeout: 2000, interval: 5 },
);
// Both files were fetched, in one batch.
expect(client.mlFetchCalls.flat().sort((a, b) => a - b)).toEqual([
1001, 1002,
]);
const callsAfterFirst = client.mlFetchCalls.length;
// The index is on disk and reloads to the same shape.
const reopened = await MLDataStore.open(
join(cacheDirectory, "mldata"),
);
expect(reopened.getIndex().fileIDs).toEqual([1001, 1002]);
// Later refreshes with nothing new must not refetch.
await new Promise((r) => setTimeout(r, FAST_INTERVAL * 1000 * 5));
expect(client.mlFetchCalls.length).toBe(callsAfterFirst);
} finally {
lib.close();
}
});
it("refetches a file whose updationTime advanced", async () => {
const client = new MLMockClient();
client.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 90)],
deleted: [],
cursor: 90,
});
client.mlByFile.set(1001, payload([0.5, 0.25, 0.75]));
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: FAST_INTERVAL,
});
try {
// Wait on `lastMLFetchAt`, set only after the first pass has
// persisted, not on `mlIndexed`, which is bumped in RAM before the
// write lands.
await vi.waitFor(
() => expect(lib.status().lastMLFetchAt).toBeGreaterThan(0),
{ timeout: 2000, interval: 5 },
);
const callsBefore = client.mlFetchCalls.length;
// The file changes on the server (updationTime advances) with a new
// embedding; the next refresh must refetch it.
client.mlByFile.set(1001, payload([9, 9, 9]));
client.collectionsQueue.push({
collections: [collection(1, 200)],
deleted: [],
cursor: 200,
});
client.filesFor(1, {
files: [file(1001, 1, 190)],
deleted: [],
cursor: 190,
});
// Poll the persisted index itself, not the fetch-call log: a call
// is recorded the instant the mock is entered, but `storeFetched`
// rewrites `clip.f32` only after it resolves, so an earlier reopen
// would read the pre-refetch vector. Reopening reads only committed
// (atomically renamed) files, so this sees the new embedding once —
// and only once — the store has written it.
await vi.waitFor(
async () => {
const reopened = await MLDataStore.open(
join(cacheDirectory, "mldata"),
);
expect([...reopened.getIndex().embeddings]).toEqual([
9, 9, 9,
]);
},
{ timeout: 2000, interval: 20 },
);
// The refetch really went back to the server for 1001.
expect(client.mlFetchCalls.length).toBeGreaterThan(callsBefore);
expect(client.mlFetchCalls.flat()).toContain(1001);
} finally {
lib.close();
}
});
});
-108
View File
@@ -1,108 +0,0 @@
/**
* Tests for the content-similarity search surface over the CLIP index
* (issue #50).
*
* The surface is `lib.mldata`: `forFile` reads the full stored payload from
* disk, while `similar` and `searchByEmbedding` rank fileIDs by cosine
* similarity over the in-RAM `Float32Array` index alone (no disk, no network).
* The fixture uses axis-aligned vectors so the correct cosine ranking is
* obvious by inspection; cosine ignores magnitude, so `[2, 0, 0]` ranks above
* `[0.8, 0.6, 0]` for a `[1, 0, 0]` query.
*/
import { describe, it, expect, beforeEach, afterEach } from "vitest";
import { mkdtempSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { MLDataStore } from "../../src/library/mldata.js";
import { makeMLDataAPI, type MLDataAPI } from "../../src/library/mlsearch.js";
import type { MLData } from "../../src/mldata-fetch.js";
// A payload shaped like Ente's: a CLIP embedding plus face data that only the
// on-disk payload carries (never the RAM index).
const payload = (embedding: number[]): MLData => ({
face: { faces: [{ faceID: "f", detection: { box: { x: 0.5 } } }] },
clip: { embedding },
});
// A small fixture index. Directions are chosen so every cosine ranking below
// is unambiguous.
const fixture = (): Map<number, MLData> =>
new Map([
[10, payload([1, 0, 0])],
[20, payload([0.8, 0.6, 0])],
[30, payload([0, 1, 0])],
[40, payload([-1, 0, 0])],
[50, payload([2, 0, 0])],
]);
describe("lib.mldata content-similarity search", () => {
let dir: string;
let store: MLDataStore;
let api: MLDataAPI;
beforeEach(async () => {
dir = mkdtempSync(join(tmpdir(), "quak-mlsearch-"));
store = await MLDataStore.open(dir);
const updation = new Map([...fixture().keys()].map((id) => [id, 1]));
await store.storeFetched(fixture(), updation);
api = makeMLDataAPI(() => store);
});
afterEach(() => {
rmSync(dir, { recursive: true, force: true });
});
it("forFile returns the whole stored payload, or undefined when uncached", async () => {
const full = await api.forFile({ fileID: 20 });
expect(full).toBeDefined();
// Face data lives only in the payload, never in the RAM index.
expect(full?.face).toBeDefined();
expect(full?.clip).toEqual({ embedding: [0.8, 0.6, 0] });
expect(await api.forFile({ fileID: 999 })).toBeUndefined();
});
it("similar ranks other files by cosine and excludes the query itself", () => {
// Query is file 10 = [1, 0, 0]. By cosine: 50 (1.0) > 20 (0.8) >
// 30 (0) > 40 (-1); 10 itself is left out.
const ranked = api.similar({ fileID: 10 });
expect(ranked.map((r) => r.fileID)).toEqual([50, 20, 30, 40]);
// Cosine ignores magnitude: [2,0,0] is a perfect match for [1,0,0].
expect(ranked[0]).toMatchObject({ fileID: 50 });
expect(ranked[0].score).toBeCloseTo(1, 5);
});
it("similar honours limit and returns [] for an unindexed file", () => {
expect(
api.similar({ fileID: 10, limit: 2 }).map((r) => r.fileID),
).toEqual([50, 20]);
expect(api.similar({ fileID: 999 })).toEqual([]);
});
it("searchByEmbedding ranks the index by cosine to the query vector", () => {
// Query [0, 1, 0]: 30 (1.0) > 20 (0.6) > {10, 40, 50} all 0, broken by
// ascending fileID.
const ranked = api.searchByEmbedding({ embedding: [0, 1, 0] });
expect(ranked.map((r) => r.fileID)).toEqual([30, 20, 10, 40, 50]);
expect(ranked[0].score).toBeCloseTo(1, 5);
expect(
api
.searchByEmbedding({ embedding: [0, 1, 0], limit: 2 })
.map((r) => r.fileID),
).toEqual([30, 20]);
});
it("searchByEmbedding returns [] for a wrong-length or zero query", () => {
expect(api.searchByEmbedding({ embedding: [1, 0] })).toEqual([]);
expect(api.searchByEmbedding({ embedding: [0, 0, 0] })).toEqual([]);
});
it("degrades to empty results when no ML store is present", async () => {
const none = makeMLDataAPI(() => undefined);
expect(await none.forFile({ fileID: 10 })).toBeUndefined();
expect(none.similar({ fileID: 10 })).toEqual([]);
expect(none.searchByEmbedding({ embedding: [1, 0, 0] })).toEqual([]);
});
});
-345
View File
@@ -1,345 +0,0 @@
/**
* Tests for `src/library/pools.ts` — the three bounded request pools (issue
* #45).
*
* A `BoundedPool` runs submitted tasks with a fixed concurrency cap. Within a
* pool, on-demand work runs before background work, and a task submitted with a
* key that a still-pending task already carries is not run twice — both callers
* share the one result. `RequestPools` bundles the three the design calls for
* (metadata 10, content 5, thumbnails 25); the pools are independent, so an
* idle pool never lends its slots to a busy one.
*
* ## How the tasks are controlled
*
* Every task here is a gate: it reports when it *starts* and then blocks until
* the test *releases* it, so the test decides exactly how many run at once and
* in what order they finish. A shared tracker counts how many tasks are running
* at any instant and records the peak, which is what the concurrency assertions
* read. No assertion is about wall-clock time.
*
* `drain()` returns a promise that settles on a macrotask, which flushes the
* microtask queue the pool schedules its starts on; the tests await it to let
* the pool react to a submission or a release before they inspect it.
*/
import { describe, it, expect } from "vitest";
import {
BoundedPool,
RequestPools,
DEFAULT_METADATA_CONCURRENCY,
DEFAULT_CONTENT_CONCURRENCY,
DEFAULT_THUMBNAIL_CONCURRENCY,
} from "../../src/library/pools.js";
// Settle on a macrotask so every microtask the pool queued has run.
const drain = (): Promise<void> =>
new Promise((resolve) => setTimeout(resolve, 0));
// A controllable task. `task` blocks until `release()` (resolve) or `fail()`
// (reject) is called; `startOrder` records the sequence in which tasks began.
interface Gate<T> {
task: () => Promise<T>;
release: (value: T) => void;
fail: (err: unknown) => void;
started: () => boolean;
runs: () => number;
}
// Tracks how many gated tasks are running concurrently across a whole test.
class Tracker {
active = 0;
peak = 0;
readonly starts: string[] = [];
gate<T>(label = ""): Gate<T> {
let settleResolve!: (value: T) => void;
let settleReject!: (err: unknown) => void;
const settled = new Promise<T>((resolve, reject) => {
settleResolve = resolve;
settleReject = reject;
});
let started = false;
let runs = 0;
const task = async (): Promise<T> => {
started = true;
runs++;
this.active++;
this.peak = Math.max(this.peak, this.active);
this.starts.push(label);
try {
return await settled;
} finally {
this.active--;
}
};
return {
task,
release: (value: T) => settleResolve(value),
fail: (err: unknown) => settleReject(err),
started: () => started,
runs: () => runs,
};
}
}
describe("BoundedPool concurrency cap", () => {
it("never runs more than `concurrency` tasks at once", async () => {
const pool = new BoundedPool(3);
const t = new Tracker();
const gates = Array.from({ length: 5 }, () => t.gate<void>());
const done = gates.map((g) => pool.run(g.task));
await drain();
// Three started, two queued behind the cap.
expect(t.active).toBe(3);
expect(gates.slice(0, 3).every((g) => g.started())).toBe(true);
expect(gates.slice(3).some((g) => g.started())).toBe(false);
// Finishing one admits exactly one more; the cap holds.
gates[0]!.release();
await drain();
expect(t.active).toBe(3);
expect(gates[3]!.started()).toBe(true);
expect(gates[4]!.started()).toBe(false);
for (const g of gates.slice(1)) g.release();
await Promise.all(done);
expect(t.peak).toBe(3);
});
it("rejects a non-positive or non-integer concurrency", () => {
expect(() => new BoundedPool(0)).toThrow(RangeError);
expect(() => new BoundedPool(-1)).toThrow(RangeError);
expect(() => new BoundedPool(2.5)).toThrow(RangeError);
});
});
describe("BoundedPool priority ordering", () => {
it("runs on-demand work before background, FIFO within a priority", async () => {
const pool = new BoundedPool(1);
const t = new Tracker();
const a = t.gate<void>("a");
const b = t.gate<void>("b");
const c = t.gate<void>("c");
const d = t.gate<void>("d");
// `a` takes the only slot; the rest queue.
void pool.run(a.task, { priority: "background" });
await drain();
void pool.run(b.task, { priority: "background" });
void pool.run(c.task, { priority: "on-demand" });
void pool.run(d.task, { priority: "background" });
await drain();
expect(t.starts).toEqual(["a"]);
// The on-demand `c` jumps ahead of the earlier-queued background `b`.
a.release();
await drain();
expect(t.starts).toEqual(["a", "c"]);
// Then background work drains in submission order: `b` before `d`.
c.release();
await drain();
expect(t.starts).toEqual(["a", "c", "b"]);
b.release();
await drain();
expect(t.starts).toEqual(["a", "c", "b", "d"]);
d.release();
});
it("defaults to background priority", async () => {
const pool = new BoundedPool(1);
const t = new Tracker();
const a = t.gate<void>("a");
const plain = t.gate<void>("plain");
const urgent = t.gate<void>("urgent");
void pool.run(a.task);
await drain();
void pool.run(plain.task); // no options -> background
void pool.run(urgent.task, { priority: "on-demand" });
await drain();
a.release();
await drain();
expect(t.starts).toEqual(["a", "urgent"]);
urgent.release();
plain.release();
});
});
describe("BoundedPool in-flight dedup", () => {
it("fetches a key once and hands both callers the same result", async () => {
const pool = new BoundedPool(5);
const t = new Tracker();
const g = t.gate<number>();
const first = pool.run(g.task, { key: 7 });
const second = pool.run(g.task, { key: 7 });
await drain();
expect(g.runs()).toBe(1);
expect(first).toBe(second);
g.release(99);
expect(await first).toBe(99);
expect(await second).toBe(99);
});
it("dedups only while in flight; a settled key runs again", async () => {
const pool = new BoundedPool(5);
const t = new Tracker();
const g1 = t.gate<number>();
const first = pool.run(g1.task, { key: 7 });
await drain();
g1.release(1);
expect(await first).toBe(1);
// The key is free again once its task settled.
const g2 = t.gate<number>();
const third = pool.run(g2.task, { key: 7 });
await drain();
expect(g2.started()).toBe(true);
g2.release(2);
expect(await third).toBe(2);
});
it("propagates a rejection to every deduped caller", async () => {
const pool = new BoundedPool(5);
const t = new Tracker();
const g = t.gate<number>();
const first = pool.run(g.task, { key: 7 });
const second = pool.run(g.task, { key: 7 });
await drain();
const boom = new Error("boom");
g.fail(boom);
await expect(first).rejects.toBe(boom);
await expect(second).rejects.toBe(boom);
// A failed key is also freed, so it may be retried by a fresh submit.
const g2 = t.gate<number>();
const retry = pool.run(g2.task, { key: 7 });
await drain();
expect(g2.started()).toBe(true);
g2.release(5);
expect(await retry).toBe(5);
});
});
describe("BoundedPool slot lifetime", () => {
it("holds one slot for a task's whole lifetime, retries included", async () => {
const pool = new BoundedPool(1);
const t = new Tracker();
// A task that internally makes two attempts before succeeding — the
// shape of a retrying request. It must occupy exactly one slot for the
// whole of that, so no other task may start until it finally settles.
const attempt1 = t.gate<void>("attempt1");
const attempt2 = t.gate<void>("attempt2");
const retrying = async (): Promise<void> => {
try {
await attempt1.task();
} catch {
await attempt2.task();
}
};
const other = t.gate<void>("other");
const running = pool.run(retrying);
await drain();
void pool.run(other.task);
await drain();
// First attempt is in flight and holds the only slot.
expect(t.starts).toEqual(["attempt1"]);
expect(other.started()).toBe(false);
// The retry is still the same task in the same slot; `other` waits.
attempt1.fail(new Error("transient"));
await drain();
expect(t.starts).toEqual(["attempt1", "attempt2"]);
expect(other.started()).toBe(false);
// Only when the whole task settles does the slot free.
attempt2.release();
await running;
await drain();
expect(other.started()).toBe(true);
other.release();
});
});
describe("RequestPools", () => {
it("exposes three pools at the design's default caps", () => {
expect(DEFAULT_METADATA_CONCURRENCY).toBe(10);
expect(DEFAULT_CONTENT_CONCURRENCY).toBe(5);
expect(DEFAULT_THUMBNAIL_CONCURRENCY).toBe(25);
const pools = new RequestPools();
expect(pools.metadata.concurrency).toBe(10);
expect(pools.content.concurrency).toBe(5);
expect(pools.thumbnails.concurrency).toBe(25);
});
it("takes overridden caps", () => {
const pools = new RequestPools({
metadataConcurrency: 1,
contentConcurrency: 2,
thumbnailConcurrency: 3,
});
expect(pools.metadata.concurrency).toBe(1);
expect(pools.content.concurrency).toBe(2);
expect(pools.thumbnails.concurrency).toBe(3);
});
it("keeps pools independent: an idle pool lends no slots", async () => {
const pools = new RequestPools({ contentConcurrency: 1 });
const t = new Tracker();
const c1 = t.gate<void>();
const c2 = t.gate<void>();
const c3 = t.gate<void>();
// The content pool is capped at 1. The thumbnail pool sits idle with 25
// free slots — none of which may be borrowed to run a second content
// task.
void pools.content.run(c1.task);
void pools.content.run(c2.task);
void pools.content.run(c3.task);
await drain();
expect(t.active).toBe(1);
c1.release();
await drain();
expect(t.active).toBe(1);
c2.release();
await drain();
expect(t.active).toBe(1);
c3.release();
await drain();
expect(t.peak).toBe(1);
});
it("runs different pools concurrently", async () => {
const pools = new RequestPools({
metadataConcurrency: 1,
contentConcurrency: 1,
});
const t = new Tracker();
const m = t.gate<void>();
const c = t.gate<void>();
void pools.metadata.run(m.task);
void pools.content.run(c.task);
await drain();
// One slot each, in two independent pools: both run at once.
expect(t.active).toBe(2);
m.release();
c.release();
});
});
-436
View File
@@ -1,436 +0,0 @@
/**
* The aggressive local precache (issue #48), driven from `Library.open`.
*
* Two background fills start with no caller input: every thumbnail in the
* account newest first, and the originals of the pinned set (the favorites
* album then the latest `precacheOriginalsDays` window ending at the newest
* file). Both run through the shared pools at background priority, so on-demand
* work always preempts them; both report through `status()`. The pinned set is
* the eviction predicate (#47), so a pinned original is never evicted and a
* file that leaves the set becomes an ordinary, evictable original.
*
* The unit tests drive `Precache` against a fake cache that records what it was
* asked to fetch (order and kind) with no pool or network; the integration
* tests drive the real wiring through `Library.open` with a stub content
* source, and the eviction test drives the real `ContentCache`.
*/
import { describe, it, expect, beforeEach, afterEach } from "vitest";
import { mkdtempSync, rmSync, existsSync, utimesSync } from "node:fs";
import { writeFile } from "node:fs/promises";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { Precache, type PrecacheCache } from "../../src/library/precache.js";
import {
deriveRecords,
type DerivedRecords,
} from "../../src/library/records.js";
import {
ContentCache,
type ContentSource,
type EnsureResult,
type StatFsFn,
} from "../../src/library/content.js";
import { RequestPools } from "../../src/library/pools.js";
import { Library } from "../../src/library/index.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
import type { CollectionsPage, FilesPage } from "../../src/client.js";
const DAY_MICROS = 24 * 60 * 60 * 1000 * 1000;
const collection = (
id: number,
type: Collection["type"] = "album",
): Collection => ({
id,
ownerID: 1,
key: new Uint8Array([id & 0xff]),
name: `album-${id}`,
type,
updationTime: 1,
isShared: false,
});
// A file whose creationTime (microseconds) places it `daysAgo` days before a
// fixed reference instant, so the latest-week window is deterministic.
const REFERENCE_MICROS = 1_000 * DAY_MICROS;
const file = (id: number, collectionID: number, daysAgo: number): EnteFile => ({
id,
collectionID,
ownerID: 1,
key: new Uint8Array([id & 0xff]),
metadata: {
title: `file-${id}.jpg`,
fileType: "image",
creationTime: REFERENCE_MICROS - daysAgo * DAY_MICROS,
modificationTime: 0,
},
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime: 1,
});
const records = (
collections: Collection[],
files: EnteFile[],
): DerivedRecords => deriveRecords(collections, files);
// A fake cache: records every fetch (kind + order) and reports the files it has
// stored via `pathsFor`. `presentThumbs`/`presentOriginals` seed already-cached
// files so the precache skips them with a single lookup.
class FakeCache implements PrecacheCache {
readonly thumbFetched: number[] = [];
readonly originalFetched: number[] = [];
readonly presentThumbs = new Set<number>();
readonly presentOriginals = new Set<number>();
pathsFor(fileID: number): {
originalPath?: string;
thumbnailPath?: string;
} {
const out: { originalPath?: string; thumbnailPath?: string } = {};
if (this.presentThumbs.has(fileID))
out.thumbnailPath = `/thumbs/${fileID}`;
if (this.presentOriginals.has(fileID))
out.originalPath = `/originals/${fileID}`;
return out;
}
async ensureThumbnails(args: {
fileIDs: number[];
priority: "background";
}): Promise<EnsureResult[]> {
return args.fileIDs.map((fileID) => {
this.thumbFetched.push(fileID);
this.presentThumbs.add(fileID);
return { fileID, path: `/thumbs/${fileID}` };
});
}
async ensureOriginals(args: {
fileIDs: number[];
}): Promise<EnsureResult[]> {
return args.fileIDs.map((fileID) => {
this.originalFetched.push(fileID);
this.presentOriginals.add(fileID);
return { fileID, path: `/originals/${fileID}` };
});
}
}
// Resolve once a predicate holds, polling the microtask queue; fails fast
// rather than hanging the suite.
const until = async (predicate: () => boolean): Promise<void> => {
for (let i = 0; i < 1000; i++) {
if (predicate()) return;
await new Promise((r) => setTimeout(r, 1));
}
throw new Error("condition not met in time");
};
describe("Precache unit", () => {
it("precaches every thumbnail newest first, skipping present ones", async () => {
const cols = [collection(1)];
const files = [
file(1, 1, 0),
file(2, 1, 1),
file(3, 1, 2),
file(4, 1, 3),
];
const cache = new FakeCache();
cache.presentThumbs.add(3); // already on disk: skipped
const pre = new Precache({ originals: false });
pre.bind(cache);
pre.update(records(cols, files));
pre.start();
await until(() => cache.thumbFetched.length === 3);
// Newest first (file 1 is newest), file 3 skipped by a lookup.
expect(cache.thumbFetched).toEqual([1, 2, 4]);
expect(cache.originalFetched).toEqual([]);
});
it("pins favorites then the latest-week window and precaches their originals in that order", async () => {
const cols = [collection(1), collection(2, "favorites")];
// File 10 is an old favorite (30 days old); files 1..3 are within the
// 7-day window; file 4 is outside it.
const files = [
file(1, 1, 0),
file(2, 1, 2),
file(3, 1, 6),
file(4, 1, 20),
file(10, 2, 30), // favorite, old
];
const cache = new FakeCache();
const pre = new Precache({ originalsDays: 7 });
pre.bind(cache);
pre.update(records(cols, files));
pre.start();
await until(() => cache.originalFetched.length === 4);
// Favorite (10) first, then the window newest-first (1, 2, 3). File 4
// is outside the window and never pinned.
expect(cache.originalFetched).toEqual([10, 1, 2, 3]);
expect(pre.isPinned(10)).toBe(true);
expect(pre.isPinned(1)).toBe(true);
expect(pre.isPinned(4)).toBe(false);
});
it("drops a file from the pinned set when the window moves past it", () => {
const cols = [collection(1)];
const cache = new FakeCache();
const pre = new Precache({ originalsDays: 7 });
pre.bind(cache);
pre.update(records(cols, [file(1, 1, 0), file(2, 1, 3)]));
expect(pre.isPinned(2)).toBe(true);
// A newer file arrives; the window's newest end moves forward so the
// 3-day-old file 2 (now 13 days behind the newest) falls out.
pre.update(
records(cols, [file(3, 1, -10), file(1, 1, 0), file(2, 1, 3)]),
);
expect(pre.isPinned(3)).toBe(true);
expect(pre.isPinned(2)).toBe(false);
});
it("reports progress through status()", async () => {
const cols = [collection(1), collection(2, "favorites")];
const files = [file(1, 1, 0), file(2, 1, 1), file(10, 2, 0)];
const cache = new FakeCache();
const pre = new Precache({ originalsDays: 7 });
pre.bind(cache);
pre.update(records(cols, files));
const before = pre.status();
expect(before.thumbnailsTotal).toBe(3);
expect(before.thumbnailsCached).toBe(0);
expect(before.originalsPinned).toBe(3); // files 1, 2, 10 all in window
expect(before.originalsCached).toBe(0);
pre.start();
await until(
() =>
cache.thumbFetched.length === 3 &&
cache.originalFetched.length === 3,
);
const after = pre.status();
expect(after.thumbnailsCached).toBe(3);
expect(after.originalsCached).toBe(3);
});
it("honours the disable flags", async () => {
const cols = [collection(1)];
const files = [file(1, 1, 0)];
const cache = new FakeCache();
const pre = new Precache({ thumbnails: false, originals: false });
pre.bind(cache);
pre.update(records(cols, files));
pre.start();
await new Promise((r) => setTimeout(r, 20));
expect(cache.thumbFetched).toEqual([]);
expect(cache.originalFetched).toEqual([]);
expect(pre.isPinned(1)).toBe(false);
expect(pre.status().thumbnailsTotal).toBe(0);
});
});
// ---- Integration through the real ContentCache and Library ----
const enteFile = (id: number, collectionID: number): EnteFile =>
file(id, collectionID, 0);
describe("Precache eviction integration", () => {
let root: string;
beforeEach(() => {
root = mkdtempSync(join(tmpdir(), "quak-precache-evict-"));
});
afterEach(() => {
if (root && existsSync(root))
rmSync(root, { recursive: true, force: true });
});
it("never evicts a pinned original the precache put in place", async () => {
const cacheDir = join(root, "cache");
// File 1 is the favorites album's only file (pinned regardless of age);
// files 2 and 3 sit outside the latest-week window, so only file 1 is
// pinned. The by-id map serves the bytes for each fetch.
const cols = [collection(1), collection(2, "favorites")];
const files = [file(1, 2, 30), file(2, 1, 40), file(3, 1, 50)];
const byID = new Map<number, EnteFile>(files.map((f) => [f.id, f]));
const source: ContentSource = {
original: async ({ destination }) => {
await writeFile(destination, Buffer.alloc(10, 1));
return { bytesWritten: 10 };
},
thumbnail: async ({ destination }) => {
await writeFile(destination, Buffer.alloc(10, 1));
return { bytesWritten: 10 };
},
};
const statfs: StatFsFn = async () => ({
bsize: 1,
bavail: 1_000_000_000,
});
const pre = new Precache({ originalsDays: 7 });
const cache = new ContentCache({
pools: new RequestPools(),
source,
cacheDirectory: cacheDir,
getFile: (id) => byID.get(id),
statfs,
cacheOriginalsMaxBytes: 25, // holds two 10-byte originals
freeBelowBytes: 0,
isPinned: (id) => pre.isPinned(id),
});
pre.bind(cache);
pre.update(records(cols, files));
expect(pre.isPinned(1)).toBe(true);
expect(pre.isPinned(2)).toBe(false);
await cache.open();
// Fill three originals; the 25-byte cap forces an eviction on the
// third, and the pinned file 1 must survive it even though it is the
// least-recently-used.
await cache.original(1);
utimesSync(join(cacheDir, "originals", "1.jpg"), 1000, 1000); // oldest
await cache.original(2);
utimesSync(join(cacheDir, "originals", "2.jpg"), 2000, 2000);
await cache.original(3);
expect(existsSync(join(cacheDir, "originals", "1.jpg"))).toBe(true);
expect(existsSync(join(cacheDir, "originals", "2.jpg"))).toBe(false);
expect(existsSync(join(cacheDir, "originals", "3.jpg"))).toBe(true);
});
});
describe("Precache preemption", () => {
let root: string;
beforeEach(() => {
root = mkdtempSync(join(tmpdir(), "quak-precache-preempt-"));
});
afterEach(() => {
if (root && existsSync(root))
rmSync(root, { recursive: true, force: true });
});
it("lets an on-demand original preempt the background originals fill", async () => {
const byID = new Map<number, EnteFile>([
[1, enteFile(1, 1)],
[2, enteFile(2, 1)],
[3, enteFile(3, 1)],
]);
const finished: number[] = [];
let openGate!: () => void;
const gate = new Promise<void>((r) => (openGate = r));
let sawFirst!: () => void;
const firstStarted = new Promise<void>((r) => (sawFirst = r));
let started = 0;
const source: ContentSource = {
original: async ({ file: f, destination }) => {
if (++started === 1) sawFirst();
await gate;
await writeFile(destination, Buffer.alloc(10, 1));
finished.push(f.id);
return { bytesWritten: 10 };
},
thumbnail: async ({ destination }) => {
await writeFile(destination, Buffer.alloc(10, 1));
return { bytesWritten: 10 };
},
};
// One content slot, so file 1 holds it while 2 and 3 wait.
const cache = new ContentCache({
pools: new RequestPools({ contentConcurrency: 1 }),
source,
cacheDirectory: join(root, "cache"),
getFile: (id) => byID.get(id),
statfs: async () => ({ bsize: 1, bavail: 1_000_000_000 }),
freeBelowBytes: 0,
});
await cache.open();
const pA = cache.ensureOriginals({ fileIDs: [1] }); // background
await firstStarted; // file 1 now holds the only slot
const pB = cache.original(2); // on-demand, queued behind file 1
const pC = cache.ensureOriginals({ fileIDs: [3] }); // background, queued
await new Promise((r) => setTimeout(r, 5)); // let both enqueue
openGate();
await Promise.all([pA, pB, pC]);
// On-demand file 2 was served before the background file 3.
expect(finished).toEqual([1, 2, 3]);
});
});
describe("Precache through Library.open", () => {
let root: string;
beforeEach(() => {
root = mkdtempSync(join(tmpdir(), "quak-precache-lib-"));
});
afterEach(() => {
if (root && existsSync(root))
rmSync(root, { recursive: true, force: true });
});
class MockClient {
served = false;
whoami(): { email: string; userID: number } {
return { email: "u@example.com", userID: 7 };
}
async collectionsSince(): Promise<CollectionsPage> {
if (this.served) return { collections: [], deleted: [], cursor: 1 };
this.served = true;
return {
collections: [collection(1), collection(2, "favorites")],
deleted: [],
cursor: 1,
};
}
async filesSince(args: { collectionID: number }): Promise<FilesPage> {
const files =
args.collectionID === 1
? [enteFile(1, 1), enteFile(2, 1)]
: [enteFile(3, 2)];
return { files, deleted: [], cursor: 1 };
}
}
it("starts both precaches from open() and reports them in status()", async () => {
const thumbFetched = new Set<number>();
const origFetched = new Set<number>();
const source: ContentSource = {
original: async ({ file: f, destination }) => {
origFetched.add(f.id);
await writeFile(destination, Buffer.alloc(10, 1));
return { bytesWritten: 10 };
},
thumbnail: async ({ file: f, destination }) => {
thumbFetched.add(f.id);
await writeFile(destination, Buffer.alloc(10, 1));
return { bytesWritten: 10 };
},
};
const lib = await Library.open({
client: new MockClient(),
cacheDirectory: join(root, "cache"),
contentSource: source,
refreshIntervalSeconds: 3600,
});
// Every file's thumbnail is precached; the favorite (file 3) and the
// week's files (1, 2) all have their originals precached.
await until(() => thumbFetched.size === 3 && origFetched.size === 3);
const status = lib.status();
expect(status.thumbnailsTotal).toBe(3);
expect(status.thumbnailsCached).toBe(3);
expect(status.originalsPinned).toBe(3);
expect(status.originalsCached).toBe(3);
lib.close();
});
});
-537
View File
@@ -1,537 +0,0 @@
/**
* Tests for the in-process read surface (issue #44).
*
* Phase 1 (#43) projected the decrypted store into plain `AlbumRecord` /
* `PhotoRecord` values. This phase adds the read API a CLI or in-process script
* uses, all served from RAM with no network:
*
* - `lib.albums` — `list` / `byName` / `byID`, returning thin `Album` wrappers.
* - `lib.photos` — `byID` (a `Photo` wrapper) and `records` (plain records).
* - `lib.timeline.groups` — photos bucketed by local day / week / month.
*
* Every call takes a single named-argument object; there are no positional
* arguments. The wrapper classes are for in-process callers only (they hold
* object identity, not JSON); the plain records remain the IPC-safe surface.
* Content-fetch methods (`Photo.original` / `thumbnail`) are a later unit and
* deliberately absent here — this surface is read-only.
*
* The detailed cases drive the API factories directly over a hand-built
* projection (`deriveRecords`), which keeps them free of disk and timers. A
* final section opens a real `Library` to prove the namespaces are wired to the
* live store and that a read never touches the client.
*/
import { describe, it, expect, beforeAll, afterAll } from "vitest";
import { mkdtempSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import {
deriveRecords,
type DerivedRecords,
} from "../../src/library/records.js";
import {
Album,
Photo,
makeAlbumsAPI,
makePhotosAPI,
makeTimelineAPI,
type TimelineGroup,
} from "../../src/library/read.js";
import { Library } from "../../src/library/index.js";
import { MetadataStore } from "../../src/library/store.js";
import type { CollectionsPage, FilesPage } from "../../src/client.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
const OWNER = 42;
// Ente stores times in microseconds; records expose milliseconds. These
// helpers keep the fixtures readable: `ms(...)` picks an epoch-millisecond
// instant, `micros(...)` is what the fixture stores so the derived record's
// `takenAt` comes back as the same millisecond value.
const ms = (epochMillis: number): number => epochMillis;
const micros = (epochMillis: number): number => epochMillis * 1000;
const collection = (
id: number,
opts: Partial<Collection> = {},
): Collection => ({
id,
ownerID: OWNER,
key: new Uint8Array([id & 0xff, 1, 2, 3]),
name: `album-${id}`,
type: "album",
updationTime: micros(1_700_000_000_000),
isShared: false,
...opts,
});
const file = (
id: number,
collectionID: number,
opts: Partial<EnteFile> & {
creationTime?: number;
title?: string;
fileType?: EnteFile["metadata"]["fileType"];
latitude?: number;
longitude?: number;
} = {},
): EnteFile => {
const { creationTime, title, fileType, latitude, longitude, ...rest } =
opts;
const metadata: EnteFile["metadata"] = {
title: title ?? `file-${id}.jpg`,
fileType: fileType ?? "image",
creationTime: creationTime ?? micros(1_700_000_000_000),
modificationTime: micros(1_700_000_000_000),
};
if (latitude !== undefined) metadata.latitude = latitude;
if (longitude !== undefined) metadata.longitude = longitude;
return {
id,
collectionID,
ownerID: OWNER,
key: new Uint8Array([id & 0xff, 9, 8, 7]),
metadata,
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime: micros(1_700_000_000_000),
...rest,
};
};
// Build the three API objects over one fixed projection, the way `Library`
// wires them over its live store.
const apis = (records: DerivedRecords) => {
const derive = () => records;
return {
albums: makeAlbumsAPI(derive),
photos: makePhotosAPI(derive),
timeline: makeTimelineAPI(derive),
};
};
// Every file id present across all timeline groups, in group-then-member order.
const allFileIDs = (groups: TimelineGroup[]): number[] =>
groups.flatMap((g) => g.fileIDs);
describe("lib.albums", () => {
it("lists albums as Album wrappers, newest updated first", () => {
const records = deriveRecords(
[
collection(1, { updationTime: micros(1_700_000_000_000) }),
collection(2, { updationTime: micros(1_705_000_000_000) }),
],
[file(10, 1), file(20, 2)],
);
const albums = apis(records).albums.list();
expect(albums.every((a) => a instanceof Album)).toBe(true);
// Collection 2 updated later, so it sorts ahead of collection 1.
expect(albums.map((a) => a.collectionID)).toEqual([2, 1]);
});
it("exposes album fields and its photos newest first", () => {
const records = deriveRecords(
[
collection(7, {
name: "Trip",
type: "favorites",
isShared: true,
}),
],
[
file(1, 7, { creationTime: micros(1_600_000_000_000) }),
file(2, 7, { creationTime: micros(1_800_000_000_000) }),
file(3, 7, { creationTime: micros(1_700_000_000_000) }),
],
);
const album = apis(records).albums.byID({ collectionID: 7 })!;
expect(album.name).toBe("Trip");
expect(album.type).toBe("favorites");
expect(album.isShared).toBe(true);
expect(album.fileIDs).toEqual([2, 3, 1]);
const photos = album.photos.list();
expect(photos.every((p) => p instanceof Photo)).toBe(true);
expect(photos.map((p) => p.fileID)).toEqual([2, 3, 1]);
// record() hands back the plain, JSON-safe projection.
expect("key" in album.record()).toBe(false);
});
it("finds an album by exact name and returns undefined when absent", () => {
const records = deriveRecords(
[
collection(1, { name: "Berlin" }),
collection(2, { name: "Paris" }),
],
[file(10, 1), file(20, 2)],
);
const { albums } = apis(records);
expect(albums.byName({ albumName: "Paris" })?.collectionID).toBe(2);
expect(albums.byName({ albumName: "paris" })).toBeUndefined();
expect(albums.byName({ albumName: "Nowhere" })).toBeUndefined();
});
it("returns undefined for an unknown collection id", () => {
const records = deriveRecords([collection(1)], [file(10, 1)]);
expect(
apis(records).albums.byID({ collectionID: 999 }),
).toBeUndefined();
});
});
describe("lib.photos", () => {
it("byID returns a Photo wrapper carrying the mapped fields", () => {
const records = deriveRecords(
[collection(1)],
[
file(1001, 1, {
title: "IMG.jpg",
fileType: "video",
creationTime: micros(1_699_000_000_000),
latitude: 52.52,
longitude: 13.405,
pubMagicMetadata: { caption: "at the lake" },
}),
],
);
const photo = apis(records).photos.byID({ fileID: 1001 })!;
expect(photo).toBeInstanceOf(Photo);
expect(photo.title).toBe("IMG.jpg");
expect(photo.fileType).toBe("video");
expect(photo.takenAt).toBe(ms(1_699_000_000_000));
expect(photo.caption).toBe("at the lake");
expect(photo.latitude).toBeCloseTo(52.52);
expect(photo.isArchived).toBe(false);
expect(photo.isHidden).toBe(false);
// The wrapper hands back the plain record, IPC-safe.
expect("key" in photo.record()).toBe(false);
});
it("byID returns undefined for an unknown file id", () => {
const records = deriveRecords([collection(1)], [file(1, 1)]);
expect(apis(records).photos.byID({ fileID: 999 })).toBeUndefined();
});
it("records() returns plain records in requested order, deduped, skipping unknowns", () => {
const records = deriveRecords(
[collection(1)],
[file(1, 1), file(2, 1), file(3, 1)],
);
const out = apis(records).photos.records({
fileIDs: [3, 1, 3, 999, 2],
});
// Requested order preserved; the repeated 3 appears once; 999 is dropped.
expect(out.map((r) => r.fileID)).toEqual([3, 1, 2]);
// Plain records, not wrappers, and JSON round-trips whole.
expect(out[0]).not.toBeInstanceOf(Photo);
expect(JSON.parse(JSON.stringify(out[0]))).toEqual(out[0]);
});
it("emits one record for a file even when it belongs to several albums", () => {
// File 1001 is a member of collections 1 and 2.
const records = deriveRecords(
[collection(1), collection(2)],
[file(1001, 1), file(1001, 2)],
);
const out = apis(records).photos.records({ fileIDs: [1001, 1001] });
expect(out).toHaveLength(1);
expect(out[0]!.albumIDs).toEqual([1, 2]);
});
});
describe("lib.timeline grouping", () => {
// Group keys and `startsAt` are computed in local time. Pinning the zone to
// UTC makes the expected values exact and lets the fixtures use `Date.UTC`.
const savedTZ = process.env.TZ;
beforeAll(() => {
process.env.TZ = "UTC";
});
afterAll(() => {
if (savedTZ === undefined) delete process.env.TZ;
else process.env.TZ = savedTZ;
});
it("buckets by local day, newest group and newest member first", () => {
const records = deriveRecords(
[collection(1)],
[
file(1, 1, { creationTime: micros(Date.UTC(2024, 0, 15, 9)) }),
file(2, 1, { creationTime: micros(Date.UTC(2024, 0, 15, 18)) }),
file(3, 1, { creationTime: micros(Date.UTC(2024, 0, 16, 12)) }),
file(4, 1, { creationTime: micros(Date.UTC(2024, 1, 1, 12)) }),
],
);
const groups = apis(records).timeline.groups({ groupBy: "day" });
expect(groups.map((g) => g.key)).toEqual([
"2024-02-01",
"2024-01-16",
"2024-01-15",
]);
// Group start is local midnight of the day.
expect(groups[2]!.startsAt).toBe(Date.UTC(2024, 0, 15));
// Within the 2024-01-15 group, the later photo (id 2) is first.
expect(groups[2]!.fileIDs).toEqual([2, 1]);
});
it("buckets by week with weeks starting on Monday", () => {
// 2024-01-15 is a Monday; the week runs through Sunday 2024-01-21.
const records = deriveRecords(
[collection(1)],
[
file(1, 1, { creationTime: micros(Date.UTC(2024, 0, 15, 12)) }), // Mon
file(2, 1, { creationTime: micros(Date.UTC(2024, 0, 17, 12)) }), // Wed
file(3, 1, { creationTime: micros(Date.UTC(2024, 0, 21, 12)) }), // Sun
file(4, 1, { creationTime: micros(Date.UTC(2024, 0, 22, 12)) }), // next Mon
],
);
const groups = apis(records).timeline.groups({ groupBy: "week" });
// ISO week keys: 2024-01-15 is in 2024-W03, the next Monday in 2024-W04.
expect(groups.map((g) => g.key)).toEqual(["2024-W04", "2024-W03"]);
const first = groups.find((g) => g.key === "2024-W03")!;
expect(first.startsAt).toBe(Date.UTC(2024, 0, 15));
// The Sunday belongs to the Monday-started week, not the next one.
expect(first.fileIDs.sort((a, b) => a - b)).toEqual([1, 2, 3]);
});
it("assigns a Sunday to the preceding Monday's week across a month boundary", () => {
// 2024-01-14 is a Sunday; its week started Monday 2024-01-08.
const records = deriveRecords(
[collection(1)],
[file(1, 1, { creationTime: micros(Date.UTC(2024, 0, 14, 12)) })],
);
const groups = apis(records).timeline.groups({ groupBy: "week" });
expect(groups.map((g) => g.key)).toEqual(["2024-W02"]);
expect(groups[0]!.startsAt).toBe(Date.UTC(2024, 0, 8));
});
it("buckets by month", () => {
const records = deriveRecords(
[collection(1)],
[
file(1, 1, { creationTime: micros(Date.UTC(2024, 0, 3, 12)) }),
file(2, 1, { creationTime: micros(Date.UTC(2024, 0, 28, 12)) }),
file(3, 1, { creationTime: micros(Date.UTC(2024, 1, 9, 12)) }),
],
);
const groups = apis(records).timeline.groups({ groupBy: "month" });
expect(groups.map((g) => g.key)).toEqual(["2024-02", "2024-01"]);
expect(groups[1]!.startsAt).toBe(Date.UTC(2024, 0, 1));
expect(groups[1]!.fileIDs).toEqual([2, 1]);
});
it("lists each file once even when it belongs to several albums", () => {
// File 1001 is in collections 1 and 2 but must appear once in a group.
const records = deriveRecords(
[collection(1), collection(2)],
[
file(1001, 1, {
creationTime: micros(Date.UTC(2024, 0, 15, 12)),
}),
file(1001, 2, {
creationTime: micros(Date.UTC(2024, 0, 15, 12)),
}),
],
);
const groups = apis(records).timeline.groups({ groupBy: "day" });
expect(allFileIDs(groups)).toEqual([1001]);
});
});
describe("lib.timeline uses local time, not UTC", () => {
const savedTZ = process.env.TZ;
afterAll(() => {
if (savedTZ === undefined) delete process.env.TZ;
else process.env.TZ = savedTZ;
});
it("buckets by the viewer's local day", () => {
// Kolkata is UTC+5:30 with no DST. An instant at 2024-01-14T20:00Z is
// 2024-01-15 01:30 local, so it belongs to the local day 2024-01-15.
process.env.TZ = "Asia/Kolkata";
const records = deriveRecords(
[collection(1)],
[file(1, 1, { creationTime: micros(Date.UTC(2024, 0, 14, 20)) })],
);
const groups = apis(records).timeline.groups({ groupBy: "day" });
expect(groups[0]!.key).toBe("2024-01-15");
// Local midnight of 2024-01-15, which is 2024-01-14T18:30Z.
expect(groups[0]!.startsAt).toBe(new Date(2024, 0, 15).getTime());
expect(groups[0]!.startsAt).toBe(Date.UTC(2024, 0, 14, 18, 30));
});
});
describe("PhotoFilter", () => {
// A fixture spanning albums, file types, geotags, captions, and the two
// visibility states, all on the same local day so grouping is incidental.
const day = (h: number): number => micros(Date.UTC(2024, 2, 4, h));
const records = (): DerivedRecords =>
deriveRecords(
[
collection(1, { name: "Holidays" }),
collection(2, { name: "Work" }),
],
[
file(1, 1, {
title: "Beach sunset",
creationTime: day(1),
latitude: 1,
longitude: 2,
}),
file(2, 1, {
title: "clip.mov",
fileType: "video",
creationTime: day(2),
}),
file(3, 2, {
title: "invoice scan",
creationTime: day(3),
pubMagicMetadata: { caption: "SUNSET colours" },
}),
file(4, 2, {
title: "archived note",
creationTime: day(4),
magicMetadata: { visibility: 1 },
}),
file(5, 2, {
title: "secret",
creationTime: day(5),
magicMetadata: { visibility: 2 },
}),
],
);
const idsWith = (
filter: Parameters<
ReturnType<typeof apis>["timeline"]["groups"]
>[0]["filter"],
): number[] =>
allFileIDs(
apis(records()).timeline.groups({ groupBy: "day", filter }),
).sort((a, b) => a - b);
it("never includes hidden photos and excludes archived by default", () => {
// No filter: hidden (5) always gone, archived (4) gone unless asked for.
expect(idsWith(undefined)).toEqual([1, 2, 3]);
});
it("includes archived photos when includeArchived is set, hidden still never", () => {
expect(idsWith({ includeArchived: true })).toEqual([1, 2, 3, 4]);
});
it("filters by album membership", () => {
expect(idsWith({ albumID: 1 })).toEqual([1, 2]);
expect(idsWith({ albumID: 2 })).toEqual([3]);
});
it("filters by file type", () => {
expect(idsWith({ fileTypes: ["video"] })).toEqual([2]);
expect(idsWith({ fileTypes: ["image", "video"] })).toEqual([1, 2, 3]);
});
it("filters by presence or absence of location", () => {
expect(idsWith({ hasLocation: true })).toEqual([1]);
expect(idsWith({ hasLocation: false })).toEqual([2, 3]);
});
it("matches text case-insensitively against title, caption, and album name", () => {
// Title match (case-insensitive): "Beach sunset".
expect(idsWith({ text: "SUNSET" })).toEqual([1, 3]);
// Caption-only match: file 3's caption is "SUNSET colours".
expect(idsWith({ text: "colours" })).toEqual([3]);
// Album-name match: everything in "Holidays".
expect(idsWith({ text: "holiday" })).toEqual([1, 2]);
});
it("combines filters", () => {
// Images in album 1 with a location: only file 1.
expect(
idsWith({ albumID: 1, fileTypes: ["image"], hasLocation: true }),
).toEqual([1]);
});
});
/**
* A minimal mock `Client`, enough for `Library.open` to run its refresh loop.
* The queues are empty, so the background refresh over a seeded cache changes
* nothing; the counters prove that a read never calls the client.
*/
class MockClient {
userID = OWNER;
collectionsCalls = 0;
filesCalls = 0;
whoami(): { email: string; userID: number } {
return { email: "user@example.com", userID: this.userID };
}
async collectionsSince(args: {
sinceTime: number;
}): Promise<CollectionsPage> {
this.collectionsCalls++;
return { collections: [], deleted: [], cursor: args.sinceTime };
}
async filesSince(args: {
collectionID: number;
collectionKey: Uint8Array;
sinceTime: number;
}): Promise<FilesPage> {
this.filesCalls++;
return { files: [], deleted: [], cursor: args.sinceTime };
}
}
describe("Library exposes the read surface over its live store", () => {
let dir: string;
beforeAll(() => {
dir = mkdtempSync(join(tmpdir(), "quak-read-"));
});
afterAll(() => {
rmSync(dir, { recursive: true, force: true });
});
it("serves albums, photos, and timeline from RAM without calling the client", async () => {
const cacheDirectory = join(dir, "cache");
const path = join(cacheDirectory, "metadata.json");
// Seed a cache as a prior run left it, so open() serves it at once.
const seed = await MetadataStore.load(path);
seed.userID = OWNER;
seed.collectionsSinceTime = 100;
seed.putCollection(collection(1, { name: "Seeded" }));
seed.putFile(
file(1001, 1, { creationTime: micros(Date.UTC(2024, 5, 1, 12)) }),
);
await seed.save();
const client = new MockClient();
// A long interval keeps the background timer from firing during the test.
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: 3600,
});
try {
const collectionsBefore = client.collectionsCalls;
const filesBefore = client.filesCalls;
expect(lib.albums.list().map((a) => a.name)).toEqual(["Seeded"]);
expect(
lib.albums.byName({ albumName: "Seeded" })?.collectionID,
).toBe(1);
expect(lib.photos.byID({ fileID: 1001 })?.fileID).toBe(1001);
expect(
lib.photos.records({ fileIDs: [1001] }).map((r) => r.fileID),
).toEqual([1001]);
const groups = lib.timeline.groups({ groupBy: "month" });
expect(allFileIDs(groups)).toEqual([1001]);
// Reads are answered from RAM: no read called the client.
expect(client.collectionsCalls).toBe(collectionsBefore);
expect(client.filesCalls).toBe(filesBefore);
} finally {
lib.close();
}
});
});
-279
View File
@@ -1,279 +0,0 @@
/**
* Tests for the plain-record mapping in `src/library/records.ts` (issue #43).
*
* The library keeps decrypted `Collection`/`EnteFile` objects in RAM, but those
* carry binary keys and cannot cross the Electron IPC boundary. `deriveRecords`
* projects them into plain `AlbumRecord`/`PhotoRecord` values — no keys, no
* `Uint8Array`, JSON-safe — that the GUI process consumes. This file pins:
*
* 1. Field mapping from `metadata` and the two magic-metadata layers, using the
* real Ente field names confirmed against the fixtures in
* `test/cli/metadata-backup.test.ts` (`w`/`h`) and `test/library/store.test.ts`
* (`visibility`): title/takenAt precedence, caption, width/height, geo,
* visibility → isArchived/isHidden, fileType.
* 2. `takenAt` is milliseconds; Ente stores creationTime/editedTime in
* microseconds, so the record divides by 1000.
* 3. Deduplication: one `PhotoRecord` per fileID even when the file belongs to
* several collections, with every membership's collection id in `albumIDs`.
* 4. Ordering: photos and album `fileIDs` are newest first.
* 5. No key material survives the projection.
* 6. `diffRecords` reports exactly what changed between two derivations, and
* returns undefined when nothing changed.
*/
import { describe, it, expect } from "vitest";
import {
deriveRecords,
snapshotFrom,
diffRecords,
} from "../../src/library/records.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
const OWNER = 42;
// Microsecond epoch values, as Ente stores times. 1e15 ≈ 2001 in microseconds.
const T = (micros: number): number => micros;
const collection = (
id: number,
opts: Partial<Collection> = {},
): Collection => ({
id,
ownerID: OWNER,
key: new Uint8Array([id & 0xff, 1, 2, 3]),
name: `album-${id}`,
type: "album",
updationTime: T(1_700_000_000_000_000),
isShared: false,
...opts,
});
const file = (
id: number,
collectionID: number,
opts: Partial<EnteFile> & {
creationTime?: number;
title?: string;
} = {},
): EnteFile => {
const { creationTime, title, ...rest } = opts;
return {
id,
collectionID,
ownerID: OWNER,
key: new Uint8Array([id & 0xff, 9, 8, 7]),
metadata: {
title: title ?? `file-${id}.jpg`,
fileType: "image",
creationTime: creationTime ?? T(1_700_000_000_000_000),
modificationTime: T(1_700_000_000_000_000),
},
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime: T(1_700_000_000_000_000),
...rest,
};
};
describe("deriveRecords: photo mapping", () => {
it("projects a file into a plain PhotoRecord with no key material", () => {
const f = file(1001, 1, {
metadata: {
title: "IMG_1.jpg",
fileType: "image",
creationTime: T(1_699_000_000_000_000),
modificationTime: T(1_699_000_000_000_000),
latitude: 52.52,
longitude: 13.405,
},
});
const { photos } = deriveRecords([collection(1)], [f]);
const rec = photos.get(1001)!;
expect(rec.fileID).toBe(1001);
expect(rec.albumIDs).toEqual([1]);
expect(rec.title).toBe("IMG_1.jpg");
expect(rec.fileType).toBe("image");
expect(rec.latitude).toBeCloseTo(52.52);
expect(rec.longitude).toBeCloseTo(13.405);
expect(rec.isArchived).toBe(false);
expect(rec.isHidden).toBe(false);
// Safe to send over IPC: no key, no Uint8Array, JSON round-trips whole.
expect("key" in rec).toBe(false);
expect(JSON.parse(JSON.stringify(rec))).toEqual(rec);
});
it("takenAt is creationTime converted from microseconds to milliseconds", () => {
const f = file(1001, 1, { creationTime: T(1_699_000_000_000_000) });
const { photos } = deriveRecords([collection(1)], [f]);
expect(photos.get(1001)!.takenAt).toBe(1_699_000_000_000);
});
it("prefers pubMagicMetadata.editedName and editedTime over metadata", () => {
const f = file(1001, 1, {
title: "original.jpg",
creationTime: T(1_699_000_000_000_000),
pubMagicMetadata: {
editedName: "Sunset over the bay",
editedTime: T(1_650_000_000_000_000),
},
});
const { photos } = deriveRecords([collection(1)], [f]);
const rec = photos.get(1001)!;
expect(rec.title).toBe("Sunset over the bay");
expect(rec.takenAt).toBe(1_650_000_000_000);
});
it("falls back to metadata when the edited fields are empty or absent", () => {
const f = file(1001, 1, {
title: "original.jpg",
creationTime: T(1_699_000_000_000_000),
pubMagicMetadata: { editedName: "" },
});
const { photos } = deriveRecords([collection(1)], [f]);
const rec = photos.get(1001)!;
expect(rec.title).toBe("original.jpg");
expect(rec.takenAt).toBe(1_699_000_000_000);
});
it("maps caption and width/height from the public magic metadata", () => {
const f = file(1001, 1, {
pubMagicMetadata: {
caption: "at the beach",
w: 3000,
h: 2000,
},
});
const rec = deriveRecords([collection(1)], [f]).photos.get(1001)!;
expect(rec.caption).toBe("at the beach");
expect(rec.width).toBe(3000);
expect(rec.height).toBe(2000);
});
it("omits optional fields that are absent from the metadata", () => {
const rec = deriveRecords([collection(1)], [file(1001, 1)]).photos.get(
1001,
)!;
expect("caption" in rec).toBe(false);
expect("width" in rec).toBe(false);
expect("height" in rec).toBe(false);
expect("latitude" in rec).toBe(false);
});
it("reads archived and hidden from private magicMetadata.visibility", () => {
const archived = file(1, 1, { magicMetadata: { visibility: 1 } });
const hidden = file(2, 1, { magicMetadata: { visibility: 2 } });
const visible = file(3, 1, { magicMetadata: { visibility: 0 } });
const { photos } = deriveRecords(
[collection(1)],
[archived, hidden, visible],
);
expect(photos.get(1)).toMatchObject({
isArchived: true,
isHidden: false,
});
expect(photos.get(2)).toMatchObject({
isArchived: false,
isHidden: true,
});
expect(photos.get(3)).toMatchObject({
isArchived: false,
isHidden: false,
});
});
});
describe("deriveRecords: dedup and ordering", () => {
it("emits one PhotoRecord per fileID across memberships, all albums listed", () => {
// File 1001 belongs to collections 1 and 2; 2002 only to 2.
const files = [
file(1001, 1, { creationTime: T(1_700_000_000_000_000) }),
file(1001, 2, { creationTime: T(1_700_000_000_000_000) }),
file(2002, 2, { creationTime: T(1_710_000_000_000_000) }),
];
const { photos } = deriveRecords([collection(1), collection(2)], files);
expect([...photos.keys()].sort((a, b) => a - b)).toEqual([1001, 2002]);
expect(photos.get(1001)!.albumIDs).toEqual([1, 2]);
expect(photos.get(2002)!.albumIDs).toEqual([2]);
});
it("orders snapshot photos newest first by takenAt", () => {
const files = [
file(1, 1, { creationTime: T(1_600_000_000_000_000) }),
file(2, 1, { creationTime: T(1_800_000_000_000_000) }),
file(3, 1, { creationTime: T(1_700_000_000_000_000) }),
];
const snap = snapshotFrom(deriveRecords([collection(1)], files), 123);
expect(snap.photos.map((p) => p.fileID)).toEqual([2, 3, 1]);
expect(snap.takenAt).toBe(123);
});
it("orders album fileIDs newest first", () => {
const files = [
file(1, 7, { creationTime: T(1_600_000_000_000_000) }),
file(2, 7, { creationTime: T(1_800_000_000_000_000) }),
file(3, 7, { creationTime: T(1_700_000_000_000_000) }),
];
const { albums } = deriveRecords([collection(7)], files);
expect(albums.get(7)!.fileIDs).toEqual([2, 3, 1]);
});
});
describe("deriveRecords: album mapping", () => {
it("carries collection identity, sharing, and the favorites type", () => {
const fav = collection(9, {
name: "Favorites",
type: "favorites",
isShared: true,
updationTime: T(1_705_000_000_000_000),
});
const rec = deriveRecords([fav], [file(1, 9)]).albums.get(9)!;
expect(rec).toMatchObject({
collectionID: 9,
name: "Favorites",
type: "favorites",
isShared: true,
updationTime: T(1_705_000_000_000_000),
});
expect("key" in rec).toBe(false);
expect(JSON.parse(JSON.stringify(rec))).toEqual(rec);
});
});
describe("diffRecords", () => {
const at = 999;
it("returns undefined when nothing changed", () => {
const a = deriveRecords([collection(1)], [file(1, 1)]);
const b = deriveRecords([collection(1)], [file(1, 1)]);
expect(diffRecords(a, b, at)).toBeUndefined();
});
it("reports added and changed albums and photos and removals", () => {
const before = deriveRecords(
[collection(1), collection(2)],
[file(1, 1), file(2, 2)],
);
// Collection 2 is gone (album + its only file removed). Collection 1 is
// renamed (changed album), gains file 3, and file 1 is retitled.
const after = deriveRecords(
[collection(1, { name: "renamed" })],
[
file(1, 1, {
pubMagicMetadata: { editedName: "new title" },
}),
file(3, 1),
],
);
const change = diffRecords(before, after, at)!;
expect(change.refreshedAt).toBe(at);
expect(change.albumIDsRemoved).toEqual([2]);
expect(change.fileIDsRemoved).toEqual([2]);
expect(change.albumsChanged.map((a) => a.collectionID)).toEqual([1]);
expect(change.albumsChanged[0]!.name).toBe("renamed");
expect(
change.photosChanged.map((p) => p.fileID).sort((x, y) => x - y),
).toEqual([1, 3]);
});
});
-303
View File
@@ -1,303 +0,0 @@
/**
* Tests for `Library.snapshot()` and `Library.subscribe()` (issue #43).
*
* These are the surface the GUI consumes across Electron IPC. `snapshot()` is
* synchronous — it reads the in-RAM store and projects it into plain records
* (no keys) — and `subscribe({ onChange })` delivers a `LibraryChange` whenever
* a background refresh actually changes the derived records. The contracts:
*
* 1. `snapshot()` deduplicates a file across memberships into one record with
* every album id, orders photos newest first, and carries no key material.
* 2. `subscribe` fires on a refresh that changes something, with the exact
* changed and removed sets for both albums and photos.
* 3. A refresh that changes nothing (an empty diff) fires no change.
* 4. `unsubscribe()` stops further delivery.
*
* The client is the same scripted mock used by the refresh-loop tests: no
* crypto, no network. Interval tests use a short real interval and `vi.waitFor`.
*/
import { describe, it, expect, beforeEach, afterEach, vi } from "vitest";
import { mkdtempSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { Library } from "../../src/library/index.js";
import type { LibraryChange } from "../../src/library/records.js";
import type { CollectionsPage, FilesPage } from "../../src/client.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
const USER_ID = 42;
const FAST_INTERVAL = 0.02;
const collection = (
id: number,
updationTime: number,
name = `album-${id}`,
): Collection => ({
id,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
name,
type: "album",
updationTime,
isShared: false,
});
const file = (
id: number,
collectionID: number,
creationTime: number,
): EnteFile => ({
id,
collectionID,
ownerID: USER_ID,
key: new Uint8Array([id & 0xff]),
metadata: {
title: `file-${id}.jpg`,
fileType: "image",
creationTime,
modificationTime: creationTime,
},
file: { decryptionHeader: "aGVhZGVy" },
thumbnail: { decryptionHeader: "dGh1bWI=" },
updationTime: creationTime,
});
class MockClient {
userID = USER_ID;
collectionsQueue: CollectionsPage[] = [];
filesByCollection = new Map<number, FilesPage[]>();
whoami(): { email: string; userID: number } {
return { email: "user@example.com", userID: this.userID };
}
async collectionsSince(args: {
sinceTime: number;
}): Promise<CollectionsPage> {
return (
this.collectionsQueue.shift() ?? {
collections: [],
deleted: [],
cursor: args.sinceTime,
}
);
}
async filesSince(args: {
collectionID: number;
collectionKey: Uint8Array;
sinceTime: number;
}): Promise<FilesPage> {
const queue = this.filesByCollection.get(args.collectionID);
return (
queue?.shift() ?? {
files: [],
deleted: [],
cursor: args.sinceTime,
}
);
}
filesFor(collectionID: number, ...pages: FilesPage[]): void {
this.filesByCollection.set(collectionID, pages);
}
}
describe("Library.snapshot and Library.subscribe", () => {
let dir: string;
let cacheDirectory: string;
beforeEach(() => {
dir = mkdtempSync(join(tmpdir(), "quak-snapshot-"));
cacheDirectory = join(dir, "cache");
});
afterEach(() => {
rmSync(dir, { recursive: true, force: true });
});
it("snapshot() dedupes across memberships, orders newest first, holds no keys", async () => {
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100), collection(2, 100)],
deleted: [],
cursor: 100,
});
// File 1001 is in both collections; 2002 only in collection 2 and newer.
client.filesFor(1, {
files: [file(1001, 1, 1_600_000_000_000_000)],
deleted: [],
cursor: 1_600_000_000_000_000,
});
client.filesFor(2, {
files: [
file(1001, 2, 1_600_000_000_000_000),
file(2002, 2, 1_800_000_000_000_000),
],
deleted: [],
cursor: 1_800_000_000_000_000,
});
const lib = await Library.open({ client, cacheDirectory });
try {
const snap = lib.snapshot();
// One record per fileID, newest first, both albums on the shared file.
expect(snap.photos.map((p) => p.fileID)).toEqual([2002, 1001]);
const shared = snap.photos.find((p) => p.fileID === 1001)!;
expect(shared.albumIDs).toEqual([1, 2]);
expect(shared.takenAt).toBe(1_600_000_000_000);
expect(snap.albums.map((a) => a.collectionID).sort()).toEqual([
1, 2,
]);
// Nothing carries key material; the whole snapshot is JSON-safe.
expect(JSON.parse(JSON.stringify(snap))).toEqual(snap);
for (const p of snap.photos) expect("key" in p).toBe(false);
for (const a of snap.albums) expect("key" in a).toBe(false);
} finally {
lib.close();
}
});
it("subscribe fires on a refresh change with the correct changed/removed sets", async () => {
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100), collection(2, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 1_600_000_000_000_000)],
deleted: [],
cursor: 1_600_000_000_000_000,
});
client.filesFor(2, {
files: [file(2002, 2, 1_600_000_000_000_000)],
deleted: [],
cursor: 1_600_000_000_000_000,
});
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: FAST_INTERVAL,
});
const changes: LibraryChange[] = [];
const { unsubscribe } = lib.subscribe({
onChange: (c) => changes.push(c),
});
try {
// Next refresh: collection 2 (and its file) tombstoned; collection 1
// gains file 1003.
client.filesFor(1, {
files: [file(1003, 1, 1_700_000_000_000_000)],
deleted: [],
cursor: 1_700_000_000_000_000,
});
client.collectionsQueue.push({
collections: [collection(1, 200)],
deleted: [2],
cursor: 200,
});
await vi.waitFor(() => expect(changes.length).toBeGreaterThan(0), {
timeout: 2000,
interval: 5,
});
const change = changes[0]!;
expect(change.albumIDsRemoved).toEqual([2]);
expect(change.fileIDsRemoved).toEqual([2002]);
expect(change.photosChanged.map((p) => p.fileID)).toEqual([1003]);
expect(change.albumsChanged.map((a) => a.collectionID)).toEqual([
1,
]);
expect(change.refreshedAt).toBeGreaterThan(0);
} finally {
unsubscribe();
lib.close();
}
});
it("a refresh that changes nothing fires no change", async () => {
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 1_600_000_000_000_000)],
deleted: [],
cursor: 1_600_000_000_000_000,
});
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: FAST_INTERVAL,
});
const changes: LibraryChange[] = [];
const { unsubscribe } = lib.subscribe({
onChange: (c) => changes.push(c),
});
try {
// Let several empty-diff ticks pass; none may deliver a change.
await new Promise((r) => setTimeout(r, FAST_INTERVAL * 1000 * 6));
expect(changes).toEqual([]);
} finally {
unsubscribe();
lib.close();
}
});
it("unsubscribe stops further delivery", async () => {
const client = new MockClient();
client.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
client.filesFor(1, {
files: [file(1001, 1, 1_600_000_000_000_000)],
deleted: [],
cursor: 1_600_000_000_000_000,
});
const lib = await Library.open({
client,
cacheDirectory,
refreshIntervalSeconds: FAST_INTERVAL,
});
const changes: LibraryChange[] = [];
const { unsubscribe } = lib.subscribe({
onChange: (c) => changes.push(c),
});
unsubscribe();
try {
client.collectionsQueue.push({
collections: [collection(3, 300)],
deleted: [],
cursor: 300,
});
client.filesFor(3, {
files: [file(3003, 3, 1_700_000_000_000_000)],
deleted: [],
cursor: 1_700_000_000_000_000,
});
// The change lands in the store, but the cancelled subscriber sees
// nothing.
await vi.waitFor(
() => expect(lib.snapshot().albums.length).toBe(2),
{ timeout: 2000, interval: 5 },
);
expect(changes).toEqual([]);
} finally {
lib.close();
}
});
});
-246
View File
@@ -1,246 +0,0 @@
/**
* Tests for the on-disk JSON metadata store (`MetadataStore`).
*
* The store is the local cache the library keeps of the account's server
* state: one `metadata.json` file holding the user id, a schema version, the
* cursor for the incremental collections listing, and the decrypted
* collection and file records. The whole file is read into RAM on load and
* rewritten as a whole on save. A separate refresh unit (issue #42) is what
* populates it; this unit only stores.
*
* Four contracts are load-bearing and each is exercised below:
*
* 1. **Round-trip fidelity.** Everything put into the store — including the
* binary decryption keys, which JSON cannot hold directly and which the
* store base64-encodes — comes back byte-for-byte after a save and a fresh
* load. A cache that quietly dropped or mangled a field would hand the
* caller wrong keys or stale metadata.
*
* 2. **A missing or corrupt file loads as an empty store, never an error.**
* The file is only a cache: if it is absent (first run) or unreadable
* (interrupted write on an older build, disk corruption, hand-editing),
* the right answer is to start empty and let the refresh unit repopulate,
* not to crash the whole library.
*
* 3. **Writes are atomic and durable.** The store reuses the same
* fsync-before-rename atomic writer the download layer uses, so a reader
* never sees a half-written file and a crash cannot leave a truncated one.
* The observable consequence tested here is that a save leaves exactly the
* destination file behind — no temporary sibling — and that overwriting an
* existing store preserves a complete, re-loadable file.
*
* 4. **Permissions match `session.json`.** The directory is `0700` and the
* file is `0600`, because the records contain decrypted key material and
* must not be readable by other users on a shared machine.
*/
import { describe, it, expect, beforeEach, afterEach } from "vitest";
import {
mkdtempSync,
rmSync,
readdirSync,
statSync,
writeFileSync,
mkdirSync,
} from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import {
MetadataStore,
METADATA_SCHEMA_VERSION,
} from "../../src/library/store.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
// A representative decrypted collection, including a binary key and all three
// magic-metadata layers, so the round-trip test proves every field survives.
const sampleCollection = (): Collection => ({
id: 12345,
ownerID: 42,
key: new Uint8Array([1, 2, 3, 4, 5, 6, 7, 8]),
name: "Holiday 2026",
type: "album",
updationTime: 1_700_000_000_000_000,
isShared: true,
magicMetadata: { visibility: 0 },
pubMagicMetadata: { subType: 0, coverID: 999 },
sharedMagicMetadata: { note: "shared with a friend" },
});
// A representative decrypted file membership: metadata, both blob headers, a
// binary key, a content hash, and file/thumbnail sizes.
const sampleFile = (): EnteFile => ({
id: 67890,
collectionID: 12345,
ownerID: 42,
key: new Uint8Array([9, 8, 7, 6, 5, 4, 3, 2, 1]),
metadata: {
title: "IMG_0001.jpg",
fileType: "image",
creationTime: 1_699_000_000_000_000,
modificationTime: 1_699_000_500_000_000,
latitude: 52.52,
longitude: 13.405,
hash: "sha256:deadbeef",
},
magicMetadata: { editedName: "sunset" },
pubMagicMetadata: { editedTime: 1_699_000_600_000_000 },
file: { decryptionHeader: "ZmlsZUhlYWRlcg==", size: 4_194_304 },
thumbnail: { decryptionHeader: "dGh1bWJIZWFkZXI=", size: 8192 },
updationTime: 1_700_000_100_000_000,
});
describe("MetadataStore", () => {
let dir: string;
let path: string;
beforeEach(() => {
dir = mkdtempSync(join(tmpdir(), "quak-store-"));
// Deliberately nest the store one level below the temp dir so save()
// has to create its own directory and set its mode.
path = join(dir, "cache", "metadata.json");
});
afterEach(() => {
rmSync(dir, { recursive: true, force: true });
});
it("round-trips the whole model, keys and all", async () => {
const store = await MetadataStore.load(path);
store.userID = 42;
store.collectionsSinceTime = 1_700_000_000_000_000;
store.putCollection(sampleCollection());
store.putFile(sampleFile());
await store.save();
const reloaded = await MetadataStore.load(path);
expect(reloaded.userID).toBe(42);
expect(reloaded.collectionsSinceTime).toBe(1_700_000_000_000_000);
// The binary key must come back as the exact bytes, not a base64
// string or a plain object of numbered keys.
const collection = reloaded.getCollection(12345);
expect(collection).toEqual(sampleCollection());
expect(collection?.key).toBeInstanceOf(Uint8Array);
const file = reloaded.getFile(12345, 67890);
expect(file).toEqual(sampleFile());
expect(file?.key).toBeInstanceOf(Uint8Array);
expect(reloaded.listCollections()).toHaveLength(1);
expect(reloaded.listFiles(12345)).toHaveLength(1);
});
it("writes the declared schema version", async () => {
const store = await MetadataStore.load(path);
await store.save();
const reloaded = await MetadataStore.load(path);
expect(reloaded.schemaVersion).toBe(METADATA_SCHEMA_VERSION);
});
it("loads an empty store when the file is missing", async () => {
const store = await MetadataStore.load(path);
expect(store.userID).toBe(0);
expect(store.listCollections()).toEqual([]);
expect(store.getCollection(1)).toBeUndefined();
});
it("loads an empty store when the file is corrupt", async () => {
mkdirSync(join(dir, "cache"), { recursive: true });
writeFileSync(path, "{ this is not valid json ][");
const store = await MetadataStore.load(path);
expect(store.listCollections()).toEqual([]);
expect(store.listFiles(12345)).toEqual([]);
});
it("loads an empty store when the schema version does not match", async () => {
// A cache written by a future build with an incompatible schema is
// discarded rather than misread; the refresh unit repopulates it.
mkdirSync(join(dir, "cache"), { recursive: true });
writeFileSync(
path,
JSON.stringify({
schemaVersion: METADATA_SCHEMA_VERSION + 1,
userID: 42,
collectionsSinceTime: 0,
collections: [],
files: [],
}),
);
const store = await MetadataStore.load(path);
expect(store.userID).toBe(0);
expect(store.listCollections()).toEqual([]);
});
it("creates the directory 0700 and the file 0600", async () => {
const store = await MetadataStore.load(path);
store.putCollection(sampleCollection());
await store.save();
// Directory 0700, file 0600: on a shared machine the decrypted keys
// in this file must be readable only by their owner. Mask to the
// permission bits; the file-type bits are not part of the assertion.
expect(statSync(join(dir, "cache")).mode & 0o777).toBe(0o700);
expect(statSync(path).mode & 0o777).toBe(0o600);
});
it("leaves exactly the destination behind, with no temp sibling", async () => {
const store = await MetadataStore.load(path);
store.putCollection(sampleCollection());
await store.save();
// The atomic writer stages a temporary file and renames it into
// place; on success nothing temporary is left in the directory.
expect(readdirSync(join(dir, "cache"))).toEqual(["metadata.json"]);
});
it("overwrites an existing store atomically and stays re-loadable", async () => {
const first = await MetadataStore.load(path);
first.userID = 1;
first.putCollection(sampleCollection());
await first.save();
const second = await MetadataStore.load(path);
second.userID = 2;
second.deleteCollection(12345);
await second.save();
const reloaded = await MetadataStore.load(path);
expect(reloaded.userID).toBe(2);
expect(reloaded.getCollection(12345)).toBeUndefined();
expect(readdirSync(join(dir, "cache"))).toEqual(["metadata.json"]);
});
it("deletes a collection together with its file memberships", async () => {
const store = await MetadataStore.load(path);
store.putCollection(sampleCollection());
store.putFile(sampleFile());
store.deleteCollection(12345);
expect(store.getCollection(12345)).toBeUndefined();
expect(store.getFile(12345, 67890)).toBeUndefined();
expect(store.listFiles(12345)).toEqual([]);
});
it("scopes file records to their collection membership", async () => {
// The same underlying file can be a member of two collections, each a
// separate record with its own key. Storing one must not touch the
// other, and lookups are per membership.
const store = await MetadataStore.load(path);
const inA = sampleFile();
const inB: EnteFile = {
...sampleFile(),
collectionID: 55555,
key: new Uint8Array([100, 101, 102]),
};
store.putFile(inA);
store.putFile(inB);
expect(store.getFile(12345, 67890)?.key).toEqual(inA.key);
expect(store.getFile(55555, 67890)?.key).toEqual(inB.key);
expect(store.listFiles(12345)).toHaveLength(1);
expect(store.listFiles(55555)).toHaveLength(1);
store.deleteFile(12345, 67890);
expect(store.getFile(12345, 67890)).toBeUndefined();
expect(store.getFile(55555, 67890)?.key).toEqual(inB.key);
});
});
+83 -175
View File
@@ -6,30 +6,17 @@
* working thumbnails, others return 404 or empty bodies. The tests * working thumbnails, others return 404 or empty bodies. The tests
* verify that the detection and repair logic handles each case correctly. * verify that the detection and repair logic handles each case correctly.
* *
* As of issue #52 both helpers take an open `Library` for enumeration and the
* `Client` for the API operations that stay unchanged (the thumbnail existence
* check, and the encrypt-and-upload path). `fixMissingThumbnails` reads each
* original through the library's content cache (`photo.original()`).
*
* `fixMissingThumbnails` is the most complex function in quak: it * `fixMissingThumbnails` is the most complex function in quak: it
* downloads the original file, generates a JPEG thumbnail with jpeg-js, * downloads the original file, generates a JPEG thumbnail with jpeg-js,
* encrypts it with secretstream push, gets a presigned upload URL, * encrypts it with secretstream push, gets a presigned upload URL,
* uploads to S3, and registers the new thumbnail with the API. The * uploads to S3, and registers the new thumbnail with the API. The
* test verifies each step actually happened and the uploaded data is * test verifies each step actually happened and the uploaded data is
* a valid encrypted blob that decrypts to a JPEG. * a valid encrypted blob that decrypts to a JPEG.
*
* It regenerates thumbnails for baseline JPEGs only, because `jpeg-js` decodes
* only JPEG. A non-JPEG image (PNG, HEIC) or a video is reported as "skipped
* (unsupported)" rather than crashing the decoder into an opaque failure
* (issue #17); the mixed test below locks that distinction down.
*/ */
import { existsSync, mkdtempSync, rmSync } from "node:fs";
import { join } from "node:path";
import { tmpdir } from "node:os";
import sodium from "libsodium-wrappers-sumo"; import sodium from "libsodium-wrappers-sumo";
import * as jpegJs from "jpeg-js"; import * as jpegJs from "jpeg-js";
import { afterAll, beforeAll, describe, expect, it } from "vitest"; import { beforeAll, describe, expect, it } from "vitest";
import { import {
init, init,
toBase64, toBase64,
@@ -40,7 +27,6 @@ import {
} from "../../src/crypto/index.js"; } from "../../src/crypto/index.js";
import { SRP, SrpServer } from "fast-srp-hap"; import { SRP, SrpServer } from "fast-srp-hap";
import { Client } from "../../src/client.js"; import { Client } from "../../src/client.js";
import { Library } from "../../src/library/index.js";
import { import {
listMissingThumbnails, listMissingThumbnails,
fixMissingThumbnails, fixMissingThumbnails,
@@ -56,7 +42,6 @@ const TEST_EMAIL = "thumb@example.com";
const TEST_PASSWORD = "thumbpass"; const TEST_PASSWORD = "thumbpass";
const TEST_OPS = 2; const TEST_OPS = 2;
const TEST_MEM = 64 * 1024 * 1024; const TEST_MEM = 64 * 1024 * 1024;
const TEST_TIME = 1700000000000000;
interface ThumbMockState { interface ThumbMockState {
verifier: Buffer; verifier: Buffer;
@@ -78,17 +63,8 @@ interface ThumbMockState {
} }
let mock: ThumbMockState; let mock: ThumbMockState;
let tmpRoot: string;
// PNG signature bytes — enough for `fixMissingThumbnails` to recognise a const buildThumbMock = async (): Promise<ThumbMockState> => {
// non-JPEG image and skip it. It need not be a decodable PNG.
const PNG_BYTES = new Uint8Array([
0x89, 0x50, 0x4e, 0x47, 0x0d, 0x0a, 0x1a, 0x0a, 0x00, 0x00, 0x00, 0x0d,
]);
const buildThumbMock = async (opts?: {
extraFormats?: boolean;
}): Promise<ThumbMockState> => {
const kekSalt = sodium.randombytes_buf(sodium.crypto_pwhash_SALTBYTES); const kekSalt = sodium.randombytes_buf(sodium.crypto_pwhash_SALTBYTES);
const kek = await deriveKEK(TEST_PASSWORD, kekSalt, TEST_OPS, TEST_MEM); const kek = await deriveKEK(TEST_PASSWORD, kekSalt, TEST_OPS, TEST_MEM);
const loginSubKeyBytes = deriveLoginSubkey(kek); const loginSubKeyBytes = deriveLoginSubkey(kek);
@@ -126,6 +102,7 @@ const buildThumbMock = async (opts?: {
opsLimit: TEST_OPS, opsLimit: TEST_OPS,
}; };
// One collection with 3 files: ok thumbnail, empty thumbnail, 404 thumbnail
const collKey = sodium.crypto_secretbox_keygen(); const collKey = sodium.crypto_secretbox_keygen();
const ckN = sodium.randombytes_buf(sodium.crypto_secretbox_NONCEBYTES); const ckN = sodium.randombytes_buf(sodium.crypto_secretbox_NONCEBYTES);
const encCK = sodium.crypto_secretbox_easy(collKey, ckN, masterKey); const encCK = sodium.crypto_secretbox_easy(collKey, ckN, masterKey);
@@ -141,11 +118,10 @@ const buildThumbMock = async (opts?: {
encryptedName: toBase64(encCN), encryptedName: toBase64(encCN),
nameDecryptionNonce: toBase64(cnN), nameDecryptionNonce: toBase64(cnN),
type: "album", type: "album",
updationTime: TEST_TIME, updationTime: 1700000000000000,
}; };
// Generate a real tiny JPEG via jpeg-js, used as the encrypted body of the // Generate a real tiny JPEG via jpeg-js
// JPEG files so a repair actually decodes and re-encodes real pixels.
const w = 100; const w = 100;
const h = 80; const h = 80;
const pixels = new Uint8Array(w * h * 4); const pixels = new Uint8Array(w * h * 4);
@@ -155,32 +131,26 @@ const buildThumbMock = async (opts?: {
pixels[i + 2] = 0; // B pixels[i + 2] = 0; // B
pixels[i + 3] = 255; // A pixels[i + 3] = 255; // A
} }
const tinyJpeg = new Uint8Array( const tinyJpeg = jpegJs.encode(
jpegJs.encode({ data: pixels, width: w, height: h }, 80).data, { data: pixels, width: w, height: h },
); 80,
).data;
const fileKeys: Record<number, Uint8Array> = {}; const fileKeys: Record<number, Uint8Array> = {};
const fileCiphertexts: Record<number, Uint8Array> = {}; const fileCiphertexts: Record<number, Uint8Array> = {};
const rawFiles: Record<string, unknown>[] = [];
// Build one raw file record: encrypt its metadata and its body under a for (const fileID of [100, 101, 102]) {
// fresh per-file key, and record the key and ciphertext for the mock to
// serve and for the test to verify against.
const makeRawFile = (
fileID: number,
fileType: number,
title: string,
body: Uint8Array,
): Record<string, unknown> => {
const fk = sodium.crypto_secretstream_xchacha20poly1305_keygen(); const fk = sodium.crypto_secretstream_xchacha20poly1305_keygen();
fileKeys[fileID] = fk; fileKeys[fileID] = fk;
const fkN = sodium.randombytes_buf(sodium.crypto_secretbox_NONCEBYTES); const fkN = sodium.randombytes_buf(sodium.crypto_secretbox_NONCEBYTES);
const encFK = sodium.crypto_secretbox_easy(fk, fkN, collKey); const encFK = sodium.crypto_secretbox_easy(fk, fkN, collKey);
const meta = JSON.stringify({ const meta = JSON.stringify({
title, title: `file-${fileID}.jpg`,
fileType, fileType: 0,
creationTime: TEST_TIME, creationTime: 1700000000000000,
modificationTime: TEST_TIME, modificationTime: 1700000000000000,
}); });
const metaPush = const metaPush =
sodium.crypto_secretstream_xchacha20poly1305_init_push(fk); sodium.crypto_secretstream_xchacha20poly1305_init_push(fk);
@@ -191,17 +161,18 @@ const buildThumbMock = async (opts?: {
sodium.crypto_secretstream_xchacha20poly1305_TAG_FINAL, sodium.crypto_secretstream_xchacha20poly1305_TAG_FINAL,
); );
// Encrypt the tiny JPEG as the file body
const filePush = const filePush =
sodium.crypto_secretstream_xchacha20poly1305_init_push(fk); sodium.crypto_secretstream_xchacha20poly1305_init_push(fk);
const encFile = sodium.crypto_secretstream_xchacha20poly1305_push( const encFile = sodium.crypto_secretstream_xchacha20poly1305_push(
filePush.state, filePush.state,
body, new Uint8Array(tinyJpeg),
null, null,
sodium.crypto_secretstream_xchacha20poly1305_TAG_FINAL, sodium.crypto_secretstream_xchacha20poly1305_TAG_FINAL,
); );
fileCiphertexts[fileID] = encFile; fileCiphertexts[fileID] = encFile;
return { rawFiles.push({
id: fileID, id: fileID,
collectionID: 1, collectionID: 1,
ownerID: 42, ownerID: 42,
@@ -215,30 +186,8 @@ const buildThumbMock = async (opts?: {
thumbnail: { thumbnail: {
decryptionHeader: toBase64(sodium.randombytes_buf(24)), decryptionHeader: toBase64(sodium.randombytes_buf(24)),
}, },
updationTime: TEST_TIME, updationTime: 1700000000000000,
}; });
};
// Three JPEG files: ok thumbnail, empty thumbnail, 404 thumbnail.
const rawFiles: Record<string, unknown>[] = [];
for (const fileID of [100, 101, 102]) {
rawFiles.push(makeRawFile(fileID, 0, `file-${fileID}.jpg`, tinyJpeg));
}
const thumbnailBehavior: Record<number, "ok" | "empty" | "404" | "500"> = {
100: "ok",
101: "empty",
102: "404",
};
// For the issue #17 mixed test: a non-JPEG image and a video, both with a
// missing (404) thumbnail so they surface in the missing list too.
if (opts?.extraFormats) {
rawFiles.push(makeRawFile(103, 0, "file-103.png", PNG_BYTES));
rawFiles.push(
makeRawFile(104, 1, "file-104.mp4", new Uint8Array([0, 0, 0, 1])),
);
thumbnailBehavior[103] = "404";
thumbnailBehavior[104] = "404";
} }
return { return {
@@ -257,7 +206,11 @@ const buildThumbMock = async (opts?: {
filesByCollection: { 1: rawFiles }, filesByCollection: { 1: rawFiles },
fileCiphertexts, fileCiphertexts,
fileKeys, fileKeys,
thumbnailBehavior, thumbnailBehavior: {
100: "ok",
101: "empty",
102: "404",
},
uploadedThumbnails: [], uploadedThumbnails: [],
}; };
}; };
@@ -428,31 +381,6 @@ const countingFetch = (
return { fetch: fake as typeof globalThis.fetch, matched: () => matched }; return { fetch: fake as typeof globalThis.fetch, matched: () => matched };
}; };
// Open a library over a mock-backed client. As the CLI does for point commands,
// the background precache is off and the refresh interval is long, and the
// library client omits `fetchMLData` so no background ML fetch runs. The real
// `Client` is still used for the API operations the helpers perform directly.
const openLib = (client: Client): Promise<Library> =>
Library.open({
client: {
whoami: () => client.whoami(),
collectionsSince: (args) => client.collectionsSince(args),
filesSince: (args) => client.filesSince(args),
contentSource: () => client.contentSource(),
},
cacheDirectory: mkdtempSync(join(tmpRoot, "cache-")),
refreshIntervalSeconds: 3600,
precacheThumbnails: false,
precacheOriginals: false,
});
const login = (fetch: typeof globalThis.fetch, retry?: RetryOptions) =>
Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: retry ? { fetch, retry } : { fetch },
});
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
// Tests // Tests
// --------------------------------------------------------------------------- // ---------------------------------------------------------------------------
@@ -461,21 +389,17 @@ beforeAll(async () => {
await init(); await init();
await sodium.ready; await sodium.ready;
mock = await buildThumbMock(); mock = await buildThumbMock();
tmpRoot = mkdtempSync(join(tmpdir(), "quak-thumb-test-"));
});
afterAll(() => {
if (tmpRoot && existsSync(tmpRoot))
rmSync(tmpRoot, { recursive: true, force: true });
}); });
describe("listMissingThumbnails", () => { describe("listMissingThumbnails", () => {
it("identifies files with empty and 404 thumbnails, ignores working ones", async () => { it("identifies files with empty and 404 thumbnails, ignores working ones", async () => {
const client = await login(buildThumbFetch(mock)); const client = await Client.login({
const lib = await openLib(client); email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildThumbFetch(mock) },
});
const missing = await listMissingThumbnails(lib, client); const missing = await listMissingThumbnails(client);
lib.close();
// File 100 has a working thumbnail → not reported // File 100 has a working thumbnail → not reported
// File 101 has an empty thumbnail → reported // File 101 has an empty thumbnail → reported
@@ -512,11 +436,13 @@ describe("listMissingThumbnails", () => {
buildThumbFetch(failingMock), buildThumbFetch(failingMock),
(url) => url.includes("thumbnails.ente.io") && url.includes("102"), (url) => url.includes("thumbnails.ente.io") && url.includes("102"),
); );
const client = await login(counted.fetch, { ...noWait }); const client = await Client.login({
const lib = await openLib(client); email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: counted.fetch, retry: { ...noWait } },
});
const missing = await listMissingThumbnails(lib, client); const missing = await listMissingThumbnails(client);
lib.close();
// Only the genuinely empty thumbnail is reported. // Only the genuinely empty thumbnail is reported.
expect(missing.map((m) => m.fileID)).toEqual([101]); expect(missing.map((m) => m.fileID)).toEqual([101]);
@@ -549,11 +475,13 @@ describe("listMissingThumbnails", () => {
return inner(input, init); return inner(input, init);
}) as typeof globalThis.fetch; }) as typeof globalThis.fetch;
const client = await login(fetch, { ...noWait }); const client = await Client.login({
const lib = await openLib(client); email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch, retry: { ...noWait } },
});
const missing = await listMissingThumbnails(lib, client); const missing = await listMissingThumbnails(client);
lib.close();
expect(missing.map((m) => m.fileID)).toEqual([101]); expect(missing.map((m) => m.fileID)).toEqual([101]);
expect(thumbRequests).toBe(4); expect(thumbRequests).toBe(4);
@@ -572,11 +500,13 @@ describe("listMissingThumbnails", () => {
mockWithDupes.filesByCollection[2] = mockWithDupes.filesByCollection[2] =
mockWithDupes.filesByCollection[1]!; mockWithDupes.filesByCollection[1]!;
const client = await login(buildThumbFetch(mockWithDupes)); const client = await Client.login({
const lib = await openLib(client); email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildThumbFetch(mockWithDupes) },
});
const missing = await listMissingThumbnails(lib, client); const missing = await listMissingThumbnails(client);
lib.close();
// Should still be 2, not 4 (each file checked only once) // Should still be 2, not 4 (each file checked only once)
expect(missing.length).toBe(2); expect(missing.length).toBe(2);
@@ -586,14 +516,16 @@ describe("listMissingThumbnails", () => {
describe("fixMissingThumbnails", () => { describe("fixMissingThumbnails", () => {
it("downloads original, generates thumbnail, encrypts, uploads, and registers", async () => { it("downloads original, generates thumbnail, encrypts, uploads, and registers", async () => {
const fixMock = await buildThumbMock(); const fixMock = await buildThumbMock();
const client = await login(buildThumbFetch(fixMock)); const client = await Client.login({
const lib = await openLib(client); email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildThumbFetch(fixMock) },
});
const results = await fixMissingThumbnails(lib, client, [101]); const results = await fixMissingThumbnails(client, [101]);
lib.close();
expect(results.length).toBe(1); expect(results.length).toBe(1);
expect(results[0]!.status).toBe("fixed"); expect(results[0]!.success).toBe(true);
expect(results[0]!.fileID).toBe(101); expect(results[0]!.fileID).toBe(101);
expect(results[0]!.title).toBe("file-101.jpg"); expect(results[0]!.title).toBe("file-101.jpg");
expect(results[0]!.collection).toBe("Photos"); expect(results[0]!.collection).toBe("Photos");
@@ -623,76 +555,48 @@ describe("fixMissingThumbnails", () => {
it("reports failure for nonexistent file IDs without crashing", async () => { it("reports failure for nonexistent file IDs without crashing", async () => {
const fixMock = await buildThumbMock(); const fixMock = await buildThumbMock();
const client = await login(buildThumbFetch(fixMock)); const client = await Client.login({
const lib = await openLib(client); email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildThumbFetch(fixMock) },
});
const results = await fixMissingThumbnails(lib, client, [999]); const results = await fixMissingThumbnails(client, [999]);
lib.close();
expect(results.length).toBe(1); expect(results.length).toBe(1);
expect(results[0]!.status).toBe("failed"); expect(results[0]!.success).toBe(false);
expect(results[0]!.fileID).toBe(999); expect(results[0]!.fileID).toBe(999);
expect(results[0]!.reason).toContain("not found"); expect(results[0]!.error).toContain("not found");
}); });
it("continues after one file fails and reports mixed results", async () => { it("continues after one file fails and reports mixed results", async () => {
const fixMock = await buildThumbMock(); const fixMock = await buildThumbMock();
// Make file 102 fail by removing its ciphertext so the download 404s. // Make file 102 fail by removing its ciphertext so download fails
delete fixMock.fileCiphertexts[102]; delete fixMock.fileCiphertexts[102];
const client = await login(buildThumbFetch(fixMock)); const client = await Client.login({
const lib = await openLib(client); email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildThumbFetch(fixMock) },
});
const results = await fixMissingThumbnails(lib, client, [101, 102]); const results = await fixMissingThumbnails(client, [101, 102]);
lib.close();
expect(results.length).toBe(2); expect(results.length).toBe(2);
const success = results.find((r) => r.fileID === 101)!; const success = results.find((r) => r.fileID === 101)!;
const failure = results.find((r) => r.fileID === 102)!; const failure = results.find((r) => r.fileID === 102)!;
expect(success.status).toBe("fixed"); expect(success.success).toBe(true);
expect(failure.status).toBe("failed"); expect(failure.success).toBe(false);
});
it("skips a non-JPEG image and a video as unsupported, not failed (issue #17)", async () => {
// A PNG and a video both throw inside the JPEG decoder. The helper must
// recognise them up front and report "skipped", distinct from a genuine
// "failed", and must not upload anything for them. The JPEG in the same
// batch is still repaired.
const fixMock = await buildThumbMock({ extraFormats: true });
const client = await login(buildThumbFetch(fixMock));
const lib = await openLib(client);
const results = await fixMissingThumbnails(
lib,
client,
[101, 103, 104],
);
lib.close();
const jpeg = results.find((r) => r.fileID === 101)!;
const png = results.find((r) => r.fileID === 103)!;
const video = results.find((r) => r.fileID === 104)!;
expect(jpeg.status).toBe("fixed");
// The PNG is a still image but not a JPEG: skipped only after its bytes
// are inspected.
expect(png.status).toBe("skipped");
expect(png.reason).toContain("JPEG");
// The video is skipped from its type alone, before any download.
expect(video.status).toBe("skipped");
expect(video.reason).toContain("video");
// Only the JPEG was uploaded; the two skipped files touched no upload.
expect(fixMock.uploadedThumbnails.length).toBe(1);
expect(fixMock.uploadedThumbnails[0]!.fileID).toBe(101);
}); });
}); });
describe("Client.getApiClient", () => { describe("Client.getApiClient", () => {
it("returns the ApiClient when logged in", async () => { it("returns the ApiClient when logged in", async () => {
const client = await login(buildThumbFetch(mock)); const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildThumbFetch(mock) },
});
const api = client.getApiClient(); const api = client.getApiClient();
expect(api).toBeDefined(); expect(api).toBeDefined();
@@ -700,7 +604,11 @@ describe("Client.getApiClient", () => {
}); });
it("throws after logout", async () => { it("throws after logout", async () => {
const client = await login(buildThumbFetch(mock)); const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: { fetch: buildThumbFetch(mock) },
});
client.logout(); client.logout();
expect(() => client.getApiClient()).toThrow(/logged out/); expect(() => client.getApiClient()).toThrow(/logged out/);