Compare commits

..
4 Commits
Author SHA1 Message Date
clawbot 2402ec97f7 README: running quak backup from cron, and its exit codes (closes #170)
check / check (push) Successful in 3m59s
A new README section shows how to run quak backup unattended: log in once
as the job's user, a crontab with a backup every night and --verify on
Sundays, output appended to a log file, and the HOME and XDG_DATA_HOME the
job needs to find the saved session. A table gives each exit code as the
code returns it, and says which failures the next run retries. The
introduction now lists everything the backup keeps for each file. It, the
backup layout tree and the lib.backup() entry say EXIF and XMP are kept for
an image only, and dimensions for a JPEG only.

Judgement call: the --verify run takes Sunday's slot rather than a second
job that night, since an overlapping run would exit 2.

Model: opus-5-5
2026-10-07 00:47:37 +02:00
clawbot ddf58af3cc quak backup --verify re-hashes stored originals and downloads again any that do not match (closes #168)
check / check (push) Failing after 1m22s
`--verify`, or `lib.backup({ verify: true })`, hashes each original already
at its save path as the download check does, streamed, a live photo as
`<imageHash>:<videoHash>`. A mismatch is logged, removed and put back in the
same run, and what is put back is hashed too, since a copy from the content
cache is not checked; one that still does not match, or a failed fetch, goes
into `failures.json`. A file with no recorded hash counts as unchecked. The
result, the summary and `--json` gain `verified`, `mismatched` and
`unchecked`.

Judgement call: a stored original that cannot be read for hashing is recorded as failed and left in place.
Judgement call: a copy put back that still does not match stays at its save path and is not counted as downloaded.
Judgement call: the summary prints the three counts only with `--verify`.

Model: opus-5-5
2026-10-06 22:47:26 +02:00
clawbot bd77422965 quak backup refuses to run while another backup of the same directory runs, exit 2 (closes #169)
check / check (push) Successful in 3m10s
lib.backup() takes a lock, backup.lock in its download directory, made
with proper-lockfile, before its refresh, and removes it when it ends. A
second backup of the directory fails at once with an error naming it.
quak backup takes the lock itself before it opens its library and passes
lockHeld to the backup, so a refused run sends no request; it prints the
error as one line and exits 2. A lock untouched for 10 seconds, left by a
run that could not remove it, is taken over.

Deviation: yarn.lock was regenerated by yarn add in the pinned node image.
Judgement call: a run failing at its refresh leaves the directory, empty.
Judgement call: a lock removed mid-run stops that run with an uncaught
error, the library's default.

Model: opus-5-5
2026-10-06 20:30:51 +02:00
clawbot 31b50a211d quak backup writes each original's EXIF, XMP and dimensions into its JSON (closes #167)
check / check (push) Successful in 3m18s
Each file's JSON gains imageMetadata, what extractImageMetadata finds in
the stored original (for a live photo, its image), or the reason the
read failed in imageMetadataError; a failed read fails neither the file
nor the run. A video is not read. An original is read when the run
stores it or when its JSON has neither field; otherwise the field is
carried over from that JSON, so a run does not read every original
again. The hand-built JPEG fixtures move to test/exif-jpeg.ts so the
backup tests can use them.

Judgement call: an original with nothing to record gets imageMetadata
{} instead of no field, so it is not read again on every run.

Model: opus-5-5
2026-10-06 16:47:31 +02:00
13 changed files with 1019 additions and 197 deletions
+128 -23
View File
@@ -10,9 +10,12 @@ account into a deduplicated local directory tree, skipping files that already
exist on disk and continuing past individual download failures instead of
crashing. For each file it persists the basic metadata fields quak keeps (title,
file type, creation and modification time, latitude, longitude, content hash),
and the private and public magic metadata in full. A helper subcommand can
detect and regenerate missing thumbnails, encrypting and uploading them back to
the server.
its update time, the private and public magic metadata in full, Ente's ML data
for it when there is any, and, for an image (for a live photo, its image), its
original's EXIF and XMP, with its dimensions for a JPEG only; for a video it
keeps none of these three. It runs unattended from cron (see "Running the backup
from cron"). A helper subcommand can detect and regenerate missing thumbnails,
encrypting and uploading them back to the server.
## Getting Started
@@ -80,6 +83,64 @@ await lib.close();
The lower-level `Client` (login, session serialization, and the raw
enumeration/download calls) is exported too and documented under Design below.
## Running the backup from cron
`quak backup` never prompts, so cron can run it. `make install` builds quak as a
single binary and copies it to `~/bin/quak`. Log in once with it, as the user
the cron job will run as; the session is saved in that user's data directory
(see "Session handling"):
```bash
~/bin/quak login
```
Then add the backup to that user's crontab with `crontab -e`. These two lines
back up the account to `~/photos-backup` at 03:30 every night, with `--verify`
on Sundays, and append all output to `~/quak-backup.log`:
```
30 3 * * 1-6 $HOME/bin/quak backup $HOME/photos-backup >> $HOME/quak-backup.log 2>&1
30 3 * * 0 $HOME/bin/quak backup --verify $HOME/photos-backup >> $HOME/quak-backup.log 2>&1
```
A run with `--verify` does all a plain run does, and also checks each original
already in the backup against the content hash Ente records for it, and replaces
any that do not match (see "Backup layout"). Sunday's run is the `--verify` one,
not a second job that night, because a backup that starts while another backup
of the same directory is running exits 2 without backing anything up.
Cron runs the job with a short `PATH`, usually `/usr/bin:/bin`, so the crontab
names quak by its full path. To find the saved session, quak needs the same
`HOME` as when you logged in, which cron sets from the password file, and on
Linux the same `XDG_DATA_HOME`: quak looks for the session in
`$XDG_DATA_HOME/quak`, or in `~/.local/share/quak` when that is not set. Cron
does not set `XDG_DATA_HOME`, so if your login shell does, set it at the top of
the crontab too. Cron does not expand variables in such a line, so give the full
path:
```
XDG_DATA_HOME=/home/you/.data
```
On macOS the session is in `~/Library/Application Support/quak`, and only `HOME`
matters.
### Exit codes
| Code | Meaning |
| ---- | ----------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `0` | The backup is complete: every original is at its save path, and no file failed. |
| `1` | The run finished with files in `failures.json`, or it stopped on another error, printed as one line `quak: <message>`. The next run tries again. |
| `2` | Another backup of the same directory is running. This run sent no request and changed nothing. |
| `3` | There is no usable session: none is saved, the saved one is corrupt, or the server no longer accepts it. Run `quak login` as the user the cron job runs as. |
After a `1`, the next run fetches each missing original again, fetches the ML
data that is not cached, and rebuilds the album links. An original that a
`--verify` run could not read, or put back with bytes that still do not match,
stays at its save path. Only a later `--verify` run checks it again: a run
without `--verify` leaves it as it is and takes the file out of `failures.json`,
so that run can exit 0.
## Examples
`examples/download-albums.ts` downloads every album's photos and their metadata
@@ -549,7 +610,7 @@ quak collections [--json] list all collections
quak files --collection <id> [--json] list files in a collection
quak get <fileID> [--out path] [--collection] download and decrypt a file
quak get-thumb <fileID> [--out] [--collection] download and decrypt a thumbnail
quak backup <dir> [--json] full incremental backup
quak backup <dir> [--json] [--verify] full incremental backup
quak backup-metadata <dir> [--exif] dump the metadata quak keeps as JSON
quak helper list-missing-thumbnails [--json] find files with missing thumbnails
quak helper fix-missing-thumbnails [--file ids] [--json] generate + upload missing thumbnails
@@ -582,6 +643,10 @@ reads no tag from is recorded, base64, as `exifRaw`, with the reason in
`exifError`. `collections`, `files`, `backup`, `helper list-missing-thumbnails`
and `helper fix-missing-thumbnails` take `--json` for machine-readable output.
`backup --verify` also hashes the originals already in the backup and replaces
any that do not match the content hash Ente records; one it cannot replace goes
into `failures.json` (see "Backup layout").
`backup-metadata` fetches ML data in requests of up to 200 files. When a request
fails, the error is logged, each of its files is written with the reason in an
`mlDataError` field instead of `mlData`, and the dump goes on. The exit code is
@@ -611,7 +676,11 @@ the smallest does not.
below)
YYYY-MM-DD.<fileID>.json the file's basic metadata fields quak
keeps, its update time, its private and
public magic metadata, and its ML data
public magic metadata, its ML data, and,
for an image (for a live photo, its
image), its original's EXIF and XMP, with
dimensions for a JPEG only; none of these
three for a video
YYYY-MM-DD.<fileID>.livephoto.json
which of a live photo's two files is which
collections/
@@ -649,6 +718,16 @@ not in the cache gets the reason in `mlDataError` instead and counts as failed,
and the next run fetches it again. The JSON files are rewritten on every run, so
ML data that arrived since the last run appears.
A file's JSON holds, as `imageMetadata`, what `backup-metadata --exif` records
from its original, or from a live photo's image: `format`, `width` and `height`
for a JPEG, `exif` (or `exifRaw` and `exifError`), and `xmp`. An original with
none of these gets `{}`. A video gets no `imageMetadata`: reading a whole video
to look for tags is not worth it. When the original cannot be read, the reason
is in `imageMetadataError` instead; neither the file nor the run fails. The
original is read when the run stores it, or when the JSON beside it has neither
field, as one an earlier version wrote. Otherwise the field is taken from that
JSON when it is rewritten, so a run does not read every stored original again.
`failures.json` records each failed file with the kind of failure, how many
times it has been tried and when it was last tried. A file leaves it once it
succeeds, or once it is no longer in the library or in the backup's scope. The
@@ -658,17 +737,18 @@ library's `lib.backup({ includeThumbnails: true })` also writes
`backup.lock` keeps two backups of the same directory from running at once, such
as a cron run that starts while the previous one is still going. A backup
creates `<dir>` if it is missing and takes the lock before its refresh, and
removes the lock when it ends, whether it succeeds or fails. The lock is a
removes the lock when it ends, whether it succeeds or fails; `quak backup` takes
it before it opens its library, so a refused run sends no request. The lock is a
directory that
[proper-lockfile](https://github.com/moxystudio/node-proper-lockfile) creates
and keeps touching while the backup runs. A second backup of the directory, from
another process or the same one, fails at once: `quak backup` prints
`quak: another backup of <dir> is running` and exits with status 2. It opens its
library before the backup takes the lock, so a refused run still waits for the
refresh that opening the library starts before it exits. A run that is killed
leaves the lock behind; once it has gone 10 seconds untouched, the next run
takes it over, so nobody has to remove it. The lock is outside the date folders
and `collections/`, so it is never taken for an original, and the removal of old
`quak: another backup of <dir> is running` and exits with status 2. A run
stopped with Ctrl-C or `kill` removes the lock as it exits. Only a run that
cannot, such as one killed with SIGKILL or cut off by a crash or power loss,
leaves it behind; once it has gone 10 seconds untouched, the next run takes it
over, so nobody has to remove it. The lock is outside the date folders and
`collections/`, so it is never taken for an original, and the removal of old
album directories never touches it.
A collection's directory and JSON are named after the collection, and a symlink
@@ -703,6 +783,25 @@ if any files failed. `quak backup` opens its library with the thumbnail and
originals precache off, so the only file content it fetches is the originals the
backup stores.
With `--verify`, or `lib.backup({ verify: true })`, a run also hashes each
original already at its save path the way a download is checked (see "On-disk
cache layout" below): its bytes, read in chunks, or a live photo's image and
video, joined as `<imageHash>:<videoHash>`. An original that matches the content
hash its metadata records is left as it is. One that does not is logged on one
line naming the file, deleted (a live photo's image and video both), and put
back in the same run like a missing one: downloaded, or copied from the cache if
the cache holds it. What is put back is hashed too, because a copy from the
cache is not checked as a download is. If it still does not match, it stays at
its save path and the file goes into `failures.json`, as it does when the
download fails. A file whose metadata records no hash is left as it is and
counted as unchecked. A stored original that cannot be read is left as it is and
counts as failed. The summary and `--json` add the counts `verified`,
`mismatched` and `unchecked`, all of originals that were already stored; one
first downloaded in this run is in none of them. A mismatch that was put back
with matching bytes does not make the exit code non-zero. Without `--verify`
nothing is hashed, the summary is unchanged, and the three counts are 0 in
`--json`.
Each original is written to a temporary file in the same directory, synced to
disk, and renamed into place, so an original is either complete or absent, even
after a power cut. A downloaded original's temporary file is named
@@ -920,17 +1019,23 @@ photos newest first). `lib.subscribe({ onChange })` delivers a `LibraryChange`
a query vector the caller produced elsewhere.
- `await lib.backup(opts?)` → `BackupResult`. It takes the lock in the download
directory, and fails at once with an error whose `code` is `ELOCKED` while
another backup of it runs. It waits for a refresh as `fresh()` does, puts
every in-scope original not already at its save path there as
`photo.download()` does (and, with `includeThumbnails`, fetches thumbnails)
through the content cache, waits for an ML data fetch, and rebuilds the
on-disk backup tree, each file's JSON with its ML data, with a durable failure
ledger. A fetched original is written straight to its save path and not into
the cache, which then counts it as present; one the cache already held is
copied from there. `BackupOptions`: `downloadDirectory` (falls back to the
library's), `includeOriginals` (default `true`), `includeThumbnails` (default
`false`), `onlyAlbumNames`, and `onProgress`. See Backup layout above for the
tree it writes.
another backup of it runs. A caller can instead take the lock itself, before
it opens its library, as `quak backup` does: `await lockBackupDirectory(dir)`
takes it, failing the same way, and returns the function that releases it, and
the caller passes `lockHeld: true` to the backup. It waits for a refresh as
`fresh()` does, puts every in-scope original not already at its save path
there as `photo.download()` does (and, with `includeThumbnails`, fetches
thumbnails) through the content cache, waits for an ML data fetch, and
rebuilds the on-disk backup tree with a durable failure ledger. Each file's
JSON holds its ML data and, for an image (for a live photo, its image), its
original's EXIF and XMP, with its dimensions for a JPEG only; a video's JSON
holds none of these three. A fetched original is written straight to its save
path and not into the cache, which then counts it as present; one the cache
already held is copied from there. `BackupOptions`: `downloadDirectory` (falls
back to the library's), `includeOriginals` (default `true`),
`includeThumbnails` (default `false`), `onlyAlbumNames`, `verify` (default
`false`), `onProgress`, and `lockHeld` (default `false`). See Backup layout
above for the tree it writes.
### Request pools
+33 -4
View File
@@ -25,14 +25,43 @@ declares one.
# Completed Steps
- 2026-10-06: The README says how to run `quak backup` from cron (issue 170):
log in once as the job's user, a crontab with a backup every night and
`--verify` on Sundays appending to a log file, and the `HOME` and
`XDG_DATA_HOME` the job needs to find the saved session. A table gives each
exit code: 0, the backup is complete; 1, files are in `failures.json` or
another error stopped the run; 2, another backup of the directory is running;
3, there is no usable session. The introduction lists everything the backup
keeps for each file.
- 2026-10-06: `quak backup --verify` and `lib.backup({ verify: true })` hash
each original already at its save path as the download check does, streamed, a
live photo as `<imageHash>:<videoHash>` (issue 168). One that does not match
the content hash its metadata records is logged, removed (both files of a live
photo) and put back in the same run, downloaded or copied from the cache, and
what is put back is hashed too. One that still does not match, or whose
download fails, goes into `failures.json`. One with no recorded hash is left
alone. The result, `--json` and the summary gain `verified`, `mismatched` and
`unchecked`. Without `--verify` nothing is hashed.
- 2026-10-06: Two backups of the same directory never run at once (issue 169).
`lib.backup()` takes a lock, `backup.lock` in its download directory, made
with `proper-lockfile`, before its refresh, and removes it when it ends,
whether it succeeds or fails. A second backup of the directory, from another
process or the same one, fails at once with an error naming the directory;
`quak backup` prints it as one line and exits 2. A lock that has gone 10
seconds untouched, left by a run that was killed, is taken over by the next
run.
process or the same one, fails at once with an error naming the directory.
`quak backup` takes the lock before it opens its library, so a refused run
sends no request; it prints the error as one line and exits 2. A lock that has
gone 10 seconds untouched, left by a run that could not remove it, is taken
over by the next run.
- 2026-10-06: `quak backup` writes each original's EXIF, XMP and dimensions into
the file's JSON as `imageMetadata`, what `backup-metadata --exif` records
(issue 167): for a live photo from its image, for a video nothing, and `{}`
for an original with none of them. A failed read puts the reason in
`imageMetadataError` and fails neither the file nor the run. An original is
read when the run stores it or when its JSON has neither field; otherwise the
field is taken from that JSON. The hand-built JPEGs moved to
`test/exif-jpeg.ts`, beside the HEIC.
- 2026-10-06: When the server answers HTTP 401 and that ends a command that
loads the saved session, because the server no longer accepts its token, the
+5 -1
View File
@@ -130,9 +130,13 @@ program
)
.argument("<dir>", "Output directory")
.option("--json", "Print result as JSON instead of human-readable summary")
.option(
"--verify",
"Re-hash stored originals and download again any that do not match",
)
// A backup usually runs from cron with nobody watching, so every request
// it makes retries for longer than the other commands' requests do.
.action((dir: string, opts: { json?: boolean }) =>
.action((dir: string, opts: { json?: boolean; verify?: boolean }) =>
run(
backupCommand(
{
+218 -41
View File
@@ -1,18 +1,21 @@
// The backup command, rebuilt on the library API (issue #51).
//
// `lib.backup()` takes the lock in `downloadDirectory`, and fails at once when
// another backup of it holds the lock. It waits for a completed refresh of the
// library (a failed one fails the backup before any file is touched), then, for
// every file in scope, puts its original at its save path under
// `downloadDirectory`, as `Photo.download()` does, waits for an ML data fetch,
// and rebuilds the derived views (per-file sidecars, per-collection symlink
// trees, per-collection JSON) from the model. The on-disk layout:
// `lib.backup()` takes the lock in `downloadDirectory` (unless its caller
// already holds it), and fails at once when another backup of it holds the
// lock. It waits for a completed refresh of the library (a failed one fails the
// backup before any file is touched), then, for every file in scope, puts its
// original at its save path under `downloadDirectory`, as `Photo.download()`
// does, waits for an ML data fetch, and rebuilds the derived views (per-file
// sidecars, per-collection symlink trees, per-collection JSON) from the model.
// The on-disk layout:
//
// <downloadDirectory>/
// YYYY/YYYY-MM/YYYY-MM-DD/
// YYYY-MM-DD.<fileID>.<ext> the decrypted bytes (the save path)
// YYYY-MM-DD.<fileID>.json per-file metadata sidecar, with
// the file's ML data
// the file's ML data and its
// original's EXIF, XMP and
// dimensions
// collections/<name>/<title> symlink to the original
// collections/<name>.json per-collection metadata
// account.json the account's email and user ID
@@ -30,7 +33,15 @@
// no unique state, so they are rebuilt every run; that repairs stale sidecars
// and missing or broken symlinks left by an earlier crash. A rebuild also
// removes the symlinks to originals that no longer belong to an album, and the
// directories of albums that no longer exist.
// directories of albums that no longer exist. The one thing a sidecar takes
// from the sidecar it replaces is its original's EXIF, XMP and dimensions (or
// why they could not be read), so that a run does not read every stored
// original again; a sidecar without them gets them read from the original.
//
// With `verify`, each original already at its save path is hashed as the
// download check hashes it, and one that does not match the content hash its
// metadata records is removed and fetched again in the same run. What is put
// back is hashed too, and recorded as failed if it still does not match.
//
// Resilience (issue #8): no per-file condition aborts the run. A failed
// download, a failed symlink, or ML data missing because the ML data fetch
@@ -44,6 +55,7 @@
// code.
import {
createReadStream,
lstatSync,
mkdirSync,
readdirSync,
@@ -55,9 +67,16 @@ import {
symlinkSync,
writeFileSync,
} from "node:fs";
import { readFile } from "node:fs/promises";
import { dirname, extname, join, relative, resolve } from "node:path";
import lockfile from "proper-lockfile";
import {
chunkHashFinal,
chunkHashInit,
chunkHashUpdate,
init,
} from "./crypto/index.js";
import { removeLeftoverTempFiles } from "./download/index.js";
import { sanitizeFileName, withExtension } from "./filename.js";
import {
@@ -67,6 +86,7 @@ import {
storedAtSavePath,
} from "./library/content.js";
import { representative } from "./library/records.js";
import { extractImageMetadata } from "./metadata-backup.js";
import type { MLData } from "./mldata-fetch.js";
import type { Collection, EnteFile, FileMetadata } from "./model/types.js";
@@ -84,7 +104,14 @@ export interface BackupOptions {
includeThumbnails?: boolean;
// Restrict the backup to albums with these names; others are left untouched.
onlyAlbumNames?: string[];
// Hash each original already at its save path, and fetch again any whose
// bytes do not match the content hash its metadata records. Default false.
verify?: boolean;
onProgress?: ProgressCallback;
// The caller already holds the lock in `downloadDirectory`, taken with
// `lockBackupDirectory`, and releases it itself, so the backup does not
// take it. `quak backup` takes it before it opens its library.
lockHeld?: boolean;
}
export interface BackupError {
@@ -101,6 +128,13 @@ export interface BackupResult {
downloaded: number;
// Originals already at their save path and left untouched.
skipped: number;
// With `verify`, the originals already at their save path whose hash
// matched, those whose hash did not (each removed and fetched again), and
// those whose metadata records no hash (left as they are). All three are
// zero without `verify`.
verified: number;
mismatched: number;
unchecked: number;
// Files with an unresolved failure after this run (the ledger size); the
// CLI exits non-zero while this is above zero. A file can be both
// downloaded and failed if its bytes landed but its symlink did not.
@@ -357,12 +391,73 @@ const saveLedger = (path: string, ledger: Map<number, FailureEntry>): void => {
);
};
// The file's JSON: its basic fields, its magic metadata, and its ML data, or
// the reason the ML data is missing.
// The content hash of the original stored at `stored`, computed as the download
// check computes it: over each file's bytes, read in chunks, and for a live
// photo `<imageHash>:<videoHash>`.
const storedHash = async (stored: {
path: string;
videoPath?: string;
}): Promise<string> => {
await init();
const hashFile = async (path: string): Promise<string> => {
const state = chunkHashInit();
for await (const chunk of createReadStream(path)) {
chunkHashUpdate(state, chunk as Buffer);
}
return chunkHashFinal(state);
};
const hash = await hashFile(stored.path);
if (stored.videoPath === undefined) return hash;
return `${hash}:${await hashFile(stored.videoPath)}`;
};
// A file's EXIF, XMP and dimensions as its JSON holds them: what
// `extractImageMetadata` found in its original, or why the original could not
// be read.
interface ImageMetadata {
imageMetadata?: Record<string, unknown>;
imageMetadataError?: string;
}
// The image metadata for the file whose original is at `originalPath` (for a
// live photo, its image) and whose JSON is at `jsonPath`. A video gets none,
// as `photo.exif()` reads none. An original stored before this run is not read
// again when its JSON already holds image metadata: that is kept. A failed
// read gives the reason, and fails neither the file nor the run. An original
// with no EXIF, XMP or JPEG dimensions gets `{}`, so it is not read again.
const imageMetadataFor = async (
file: EnteFile,
originalPath: string,
jsonPath: string,
storedThisRun: boolean,
): Promise<ImageMetadata> => {
if (file.metadata.fileType === "video") return {};
if (!storedThisRun) {
try {
const { imageMetadata, imageMetadataError } = JSON.parse(
readFileSync(jsonPath, "utf-8"),
) as ImageMetadata;
if (imageMetadata !== undefined || imageMetadataError !== undefined)
return { imageMetadata, imageMetadataError };
} catch {
// No JSON yet, or one that cannot be parsed: read the original.
}
}
try {
const bytes = await readFile(originalPath);
return { imageMetadata: extractImageMetadata(bytes) ?? {} };
} catch (err) {
return { imageMetadataError: errorMessage(err) };
}
};
// The file's JSON: its basic fields, its magic metadata, its ML data or the
// reason the ML data is missing, and its image metadata.
const writeSidecar = (
path: string,
file: EnteFile,
ml: { mlData?: MLData; mlDataError?: string },
image: ImageMetadata,
): void => {
const meta: Record<string, unknown> = {
id: file.id,
@@ -375,6 +470,10 @@ const writeSidecar = (
if (file.pubMagicMetadata) meta.pubMagicMetadata = file.pubMagicMetadata;
if (ml.mlData) meta.mlData = ml.mlData;
if (ml.mlDataError) meta.mlDataError = ml.mlDataError;
if (image.imageMetadata) meta.imageMetadata = image.imageMetadata;
if (image.imageMetadataError) {
meta.imageMetadataError = image.imageMetadataError;
}
writeFileSync(path, JSON.stringify(meta, null, 2));
};
@@ -401,7 +500,7 @@ const writeAlbumJSON = (
writeFileSync(path, JSON.stringify(album, null, 2));
};
// The backup itself, which `runBackup` below runs while it holds the lock.
// The backup itself, which `runBackup` below runs while the lock is held.
const runLockedBackup = async (
lib: BackupLibrary,
opts: BackupOptions,
@@ -411,6 +510,7 @@ const runLockedBackup = async (
const includeThumbnails = opts.includeThumbnails ?? false;
const log = opts.onProgress ?? (() => {});
const only = opts.onlyAlbumNames ? new Set(opts.onlyAlbumNames) : undefined;
const verify = opts.verify ?? false;
log("Refreshing library...");
await lib.refresh();
@@ -466,8 +566,12 @@ const runLockedBackup = async (
const errors: BackupError[] = [];
const failedThisRun = new Set<number>();
const storedThisRun = new Set<number>();
let downloaded = 0;
let skipped = 0;
let verified = 0;
let mismatched = 0;
let unchecked = 0;
const recordFailure = (
file: EnteFile,
@@ -499,10 +603,47 @@ const runLockedBackup = async (
// Phase 1: get the bytes. Put each pending original at its save path
// through the content cache/pools, as `Photo.download()` does, and fetch
// the optional thumbnails; a present file is left as is.
// the optional thumbnails; a present file is left as is. With `verify`, a
// present original is hashed first, and one that does not match the hash
// its metadata records is removed and fetched like a missing one. One
// that cannot be read is recorded as failed and left where it is.
if (includeOriginals) {
for (const [fileID, file] of distinct) {
if (storedAtSavePath(downloadDirectory, file) !== undefined) {
let stored = storedAtSavePath(downloadDirectory, file);
let mismatch = false;
if (stored !== undefined && verify) {
try {
if (file.metadata.hash === undefined) {
unchecked++;
} else if (
(await storedHash(stored)) === file.metadata.hash
) {
verified++;
} else {
log(
`MISMATCH original ${file.metadata.title} (${fileID}): its bytes do not match its content hash`,
);
mismatched++;
mismatch = true;
rmSync(stored.path);
if (stored.videoPath !== undefined) {
rmSync(stored.videoPath);
}
stored = undefined;
}
} catch (err) {
log(
`FAILED verifying original ${file.metadata.title}: ${errorMessage(err)}`,
);
recordFailure(
file,
collectionName.get(file.collectionID) ?? "",
err,
);
continue;
}
}
if (stored !== undefined) {
skipped++;
continue;
}
@@ -511,9 +652,24 @@ const runLockedBackup = async (
// A fetched original is written straight to its save path (a
// live photo beside it); only one that was already cached
// elsewhere is copied.
await placeOriginal(downloadDirectory, file, (dest) =>
lib.original(fileID, dest),
const placed = await placeOriginal(
downloadDirectory,
file,
(dest) => lib.original(fileID, dest),
);
storedThisRun.add(fileID);
// A copy from the cache is not checked as a download is and can
// hold the same bad bytes, so what is put back after a
// mismatch is hashed too. A bad copy stays where it is and the
// file fails.
if (
mismatch &&
(await storedHash(placed)) !== file.metadata.hash
) {
throw new Error(
"the original put back does not match its content hash either",
);
}
downloaded++;
} catch (err) {
log(
@@ -547,9 +703,10 @@ const runLockedBackup = async (
// Phase 2: rebuild the derived views from the model. Sidecars first, for
// every present original (this repairs stale ones), each with the file's
// ML data once an ML data fetch has completed. When the fetch fails, a
// file with no cached ML data gets the reason instead and is recorded as
// failed. The next run fetches its ML data again because none is cached.
// ML data once an ML data fetch has completed, and its image metadata.
// When the fetch fails, a file with no cached ML data gets the reason
// instead and is recorded as failed. The next run fetches its ML data
// again because none is cached.
if (includeOriginals) {
let mlDataError: string | undefined;
try {
@@ -560,23 +717,28 @@ const runLockedBackup = async (
log(`FAILED ML data: ${mlDataError}`);
}
for (const file of distinct.values()) {
if (storedAtSavePath(downloadDirectory, file) === undefined) {
continue;
}
const stored = storedAtSavePath(downloadDirectory, file);
if (stored === undefined) continue;
const path = withExtension(
savePath(downloadDirectory, file),
".json",
);
const image = await imageMetadataFor(
file,
stored.path,
path,
storedThisRun.has(file.id),
);
const mlData = await lib.mlData(file.id);
if (mlData === undefined && mlDataError !== undefined) {
writeSidecar(path, file, { mlDataError });
writeSidecar(path, file, { mlDataError }, image);
recordFailure(
file,
collectionName.get(file.collectionID) ?? "",
new Error(`ML data: ${mlDataError}`),
);
} else {
writeSidecar(path, file, { mlData });
writeSidecar(path, file, { mlData }, image);
}
}
}
@@ -670,16 +832,40 @@ const runLockedBackup = async (
totalFiles: distinct.size,
downloaded,
skipped,
verified,
mismatched,
unchecked,
failed: ledger.size,
errors,
};
};
// Only one backup of a directory runs at a time, in this process or another:
// a second one fails at once with an error whose `code` is `ELOCKED`. The lock
// is the directory `backup.lock`, whose modification time proper-lockfile
// keeps current while the backup runs. One it has not touched for 10 seconds
// was left by a run that was killed, and is taken over.
// Only one backup of a directory runs at a time, in this process or another.
// Creates `downloadDirectory` if it is missing, takes the lock in it and
// returns the function that releases it; while another backup holds the lock,
// fails at once with an error whose `code` is `ELOCKED`. The lock is the
// directory `backup.lock`, whose modification time proper-lockfile keeps
// current while it is held. One it has not touched for 10 seconds was left by
// a run that could not remove it, and is taken over.
export const lockBackupDirectory = async (
downloadDirectory: string,
): Promise<() => Promise<void>> => {
mkdirSync(downloadDirectory, { recursive: true });
try {
return await lockfile.lock(downloadDirectory, {
lockfilePath: join(downloadDirectory, "backup.lock"),
});
} catch (err) {
if ((err as NodeJS.ErrnoException).code !== "ELOCKED") throw err;
throw Object.assign(
new Error(`another backup of ${downloadDirectory} is running`),
{ code: "ELOCKED" },
);
}
};
// Runs the backup holding the lock, which it takes and releases itself unless
// the caller already holds it (`opts.lockHeld`).
export const runBackup = async (
lib: BackupLibrary,
opts: BackupOptions,
@@ -691,19 +877,10 @@ export const runBackup = async (
"open the library with one)",
);
}
mkdirSync(downloadDirectory, { recursive: true });
let release: () => Promise<void>;
try {
release = await lockfile.lock(downloadDirectory, {
lockfilePath: join(downloadDirectory, "backup.lock"),
});
} catch (err) {
if ((err as NodeJS.ErrnoException).code !== "ELOCKED") throw err;
throw Object.assign(
new Error(`another backup of ${downloadDirectory} is running`),
{ code: "ELOCKED" },
);
if (opts.lockHeld) {
return runLockedBackup(lib, opts, downloadDirectory);
}
const release = await lockBackupDirectory(downloadDirectory);
try {
return await runLockedBackup(lib, opts, downloadDirectory);
} finally {
+49 -36
View File
@@ -20,9 +20,9 @@ import {
type ClientSnapshot,
type LoginOptions,
} from "./client.js";
import { lockBackupDirectory } from "./backup.js";
import { init } from "./crypto/index.js";
import {
type BackupResult,
defaultCacheDirectory,
Library,
type LibraryClient,
@@ -402,59 +402,72 @@ export const backupMetadataCommand = async (
export const backupCommand = async (
ctx: CliContext,
dir: string,
opts: { json?: boolean },
opts: { json?: boolean; verify?: boolean },
): Promise<number> => {
await init();
const client = requireSession(ctx);
if (!client) return 3;
ctx.stderr.write("Starting backup...\n");
// The precache is off: the backup fetches what it needs, and must not
// also fill the cache with every thumbnail and the recent originals.
const lib = await Library.open({
client,
downloadDirectory: dir,
cacheDirectory: ctx.cacheDir,
precacheThumbnails: false,
precacheOriginals: false,
});
// The lock is taken before the library opens and starts its refresh, so a
// run refused while another backup of `dir` runs sends no request.
let release: () => Promise<void>;
try {
let result: BackupResult;
release = await lockBackupDirectory(dir);
} catch (err) {
if ((err as NodeJS.ErrnoException).code !== "ELOCKED") throw err;
ctx.stderr.write(`quak: ${(err as Error).message}\n`);
return 2;
}
try {
// The precache is off: the backup fetches what it needs, and must not
// also fill the cache with every thumbnail and the recent originals.
const lib = await Library.open({
client,
downloadDirectory: dir,
cacheDirectory: ctx.cacheDir,
precacheThumbnails: false,
precacheOriginals: false,
});
try {
result = await lib.backup({
const result = await lib.backup({
downloadDirectory: dir,
lockHeld: true,
verify: opts.verify,
onProgress: (msg) => {
if (!opts.json) ctx.stderr.write(msg + "\n");
},
});
} catch (err) {
// Another backup of `dir` is running and holds its lock.
if ((err as NodeJS.ErrnoException).code !== "ELOCKED") throw err;
ctx.stderr.write(`quak: ${(err as Error).message}\n`);
return 2;
}
if (opts.json) {
ctx.stdout.write(JSON.stringify(result, null, 2) + "\n");
} else {
ctx.stderr.write("\n--- Backup complete ---\n");
ctx.stderr.write(` Total files: ${result.totalFiles}\n`);
ctx.stderr.write(` Downloaded: ${result.downloaded}\n`);
ctx.stderr.write(` Skipped: ${result.skipped}\n`);
ctx.stderr.write(` Failed: ${result.failed}\n`);
if (result.errors.length > 0) {
ctx.stderr.write("\nFailed files:\n");
for (const e of result.errors) {
ctx.stderr.write(
` [${e.collection}] ${e.title} (id ${e.fileID}): ${e.error}\n`,
);
if (opts.json) {
ctx.stdout.write(JSON.stringify(result, null, 2) + "\n");
} else {
ctx.stderr.write("\n--- Backup complete ---\n");
ctx.stderr.write(` Total files: ${result.totalFiles}\n`);
ctx.stderr.write(` Downloaded: ${result.downloaded}\n`);
ctx.stderr.write(` Skipped: ${result.skipped}\n`);
if (opts.verify) {
ctx.stderr.write(` Verified: ${result.verified}\n`);
ctx.stderr.write(` Mismatched: ${result.mismatched}\n`);
ctx.stderr.write(` Unchecked: ${result.unchecked}\n`);
}
ctx.stderr.write(` Failed: ${result.failed}\n`);
if (result.errors.length > 0) {
ctx.stderr.write("\nFailed files:\n");
for (const e of result.errors) {
ctx.stderr.write(
` [${e.collection}] ${e.title} (id ${e.fileID}): ${e.error}\n`,
);
}
}
}
}
return result.failed > 0 ? 1 : 0;
return result.failed > 0 ? 1 : 0;
} finally {
await lib.close();
}
} finally {
await lib.close();
await release();
}
};
+1
View File
@@ -67,6 +67,7 @@ export {
type EnsureOptions,
type EnsureResult,
type EnsureEvent,
lockBackupDirectory,
runBackup,
type BackupOptions,
type BackupResult,
+4 -2
View File
@@ -94,6 +94,7 @@ import type { Collection, EnteFile } from "../model/types.js";
import { runBackup, type BackupOptions, type BackupResult } from "../backup.js";
export {
lockBackupDirectory,
runBackup,
type BackupOptions,
type BackupResult,
@@ -566,8 +567,9 @@ export class Library {
// Back up every in-scope file to `opts.downloadDirectory`, or else the
// library's, each original at its save path, with a durable failure
// ledger (issue #51). Takes the lock in that directory first, failing at
// once while another backup of it runs (see `runBackup`). Waits for a
// ledger (issue #51). Takes the lock in that directory first, unless
// `opts.lockHeld` says the caller holds it, failing at once while another
// backup of it runs (see `runBackup`). Waits for a
// completed refresh, as `fresh()` does, joining one already running, and
// rejects before touching any file but the lock when it fails. Then puts
// pending originals at their save paths as `Photo.download()` does (and
+412
View File
@@ -33,6 +33,7 @@
*/
import {
chmodSync,
existsSync,
lstatSync,
mkdirSync,
@@ -54,12 +55,17 @@ import { runBackup, type BackupLibrary } from "../../src/backup.js";
import { Library } from "../../src/library/index.js";
import type { ContentSource } from "../../src/library/content.js";
import type { CollectionsPage, FilesPage } from "../../src/client.js";
import { extractImageMetadata } from "../../src/metadata-backup.js";
import type { MLData } from "../../src/mldata-fetch.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
import { HEIC_WITH_EXIF } from "../exif-heic.js";
import { JPEG_WITH_EXIF } from "../exif-jpeg.js";
import {
asLivePhoto,
blake2b,
cdnSource,
IMAGE,
livePhotoHash,
livePhotoZip,
VIDEO,
} from "../live-photo.js";
@@ -1093,6 +1099,354 @@ describe("account and album records", () => {
});
});
// What a file's JSON holds as `imageMetadata` for an original of `bytes`.
const imageMetadataOf = (bytes: Uint8Array): unknown =>
JSON.parse(JSON.stringify(extractImageMetadata(bytes)));
describe("image metadata in each file's JSON", () => {
// Serves `files` in the Vacation album and nothing in Work.
class FilesClient extends MockClient {
constructor(private readonly files: EnteFile[]) {
super();
}
override async filesSince(args: {
collectionID: number;
}): Promise<FilesPage> {
return {
files: args.collectionID === 1 ? this.files : [],
deleted: [],
cursor: 1,
};
}
}
// Writes `originals.get(fileID)` as each file's original, looked up at
// each fetch, so a test can change it between runs.
const bytesSource = (
originals: Map<number, Uint8Array>,
): ContentSource => ({
original: async ({ file: f, destination }) => {
const bytes = originals.get(f.id)!;
writeFileSync(destination, bytes);
return { bytesWritten: bytes.length };
},
thumbnail: async () => {
throw new Error("no thumbnails in this source");
},
});
// A JPEG, a HEIC, a video, and a file holding no image metadata. The
// video's bytes are the JPEG's, so reading it would find EXIF.
const clip = file(602, 1, "clip.mov");
const files = [
file(600, 1, "photo.jpg"),
file(601, 1, "photo.heic"),
{ ...clip, metadata: { ...clip.metadata, fileType: "video" as const } },
file(603, 1, "notes.png"),
];
const originals = (): Map<number, Uint8Array> =>
new Map([
[600, JPEG_WITH_EXIF],
[601, HEIC_WITH_EXIF],
[602, JPEG_WITH_EXIF],
[603, new TextEncoder().encode("not an image")],
]);
const open = (bytes = originals()): Promise<Library> =>
openLibrary(bytesSource(bytes), new FilesClient(files));
// Rewrite the JSON of `fileID` without its image metadata, as an earlier
// version of quak wrote it.
const dropImageMetadata = (outDir: string, fileID: number): void => {
const json = fileJSON(outDir, fileID);
delete json.imageMetadata;
writeFileSync(saved(outDir, `${fileID}.json`), JSON.stringify(json));
};
it("writes each new original's image metadata, and none for a video", async () => {
const lib = await open();
const outDir = join(root, "backup");
const result = await lib.backup({ downloadDirectory: outDir });
expect(result).toMatchObject({ downloaded: 4, failed: 0 });
expect(fileJSON(outDir, 600).imageMetadata).toEqual(
imageMetadataOf(JPEG_WITH_EXIF),
);
expect(fileJSON(outDir, 601).imageMetadata).toEqual(
imageMetadataOf(HEIC_WITH_EXIF),
);
expect(fileJSON(outDir, 602)).not.toHaveProperty("imageMetadata");
expect(fileJSON(outDir, 602)).not.toHaveProperty("imageMetadataError");
// Empty, so that the next run does not read it again.
expect(fileJSON(outDir, 603).imageMetadata).toEqual({});
await lib.close();
});
it("reads a live photo's image", async () => {
const { file: live, body } = await asLivePhoto(
file(500, 1, "IMG_0500.HEIC"),
livePhotoZip({ "image.heic": HEIC_WITH_EXIF, "video.mov": VIDEO }),
livePhotoHash(HEIC_WITH_EXIF, VIDEO),
);
const lib = await openLibrary(
cdnSource(new Map([[500, body]])),
new FilesClient([live]),
);
const outDir = join(root, "backup");
await lib.backup({ downloadDirectory: outDir });
expect(fileJSON(outDir, 500).imageMetadata).toEqual(
imageMetadataOf(HEIC_WITH_EXIF),
);
await lib.close();
});
it("keeps the image metadata of an original stored by an earlier run without reading it again", async () => {
const lib = await open();
const outDir = join(root, "backup");
await lib.backup({ downloadDirectory: outDir });
// A read of the original now would give the HEIC's.
writeFileSync(saved(outDir, "600.jpg"), HEIC_WITH_EXIF);
const second = await lib.backup({ downloadDirectory: outDir });
expect(second).toMatchObject({ downloaded: 0, failed: 0 });
expect(fileJSON(outDir, 600).imageMetadata).toEqual(
imageMetadataOf(JPEG_WITH_EXIF),
);
await lib.close();
});
it("reads an original stored by an earlier version, whose JSON has no image metadata", async () => {
const lib = await open();
const outDir = join(root, "backup");
await lib.backup({ downloadDirectory: outDir });
dropImageMetadata(outDir, 600);
const second = await lib.backup({ downloadDirectory: outDir });
expect(second).toMatchObject({ downloaded: 0, failed: 0 });
expect(fileJSON(outDir, 600).imageMetadata).toEqual(
imageMetadataOf(JPEG_WITH_EXIF),
);
await lib.close();
});
it("reads an original the run stores again, though its JSON has image metadata", async () => {
const bytes = originals();
const lib = await open(bytes);
const outDir = join(root, "backup");
await lib.backup({ downloadDirectory: outDir });
rmSync(saved(outDir, "600.jpg"));
bytes.set(600, HEIC_WITH_EXIF);
const second = await lib.backup({ downloadDirectory: outDir });
expect(second).toMatchObject({ downloaded: 1, failed: 0 });
expect(fileJSON(outDir, 600).imageMetadata).toEqual(
imageMetadataOf(HEIC_WITH_EXIF),
);
await lib.close();
});
// Root ignores file permissions, so this fails when run as root. The
// `test` phase of the `Dockerfile` runs as the `node` user.
it("gives the reason an original could not be read, failing neither the file nor the run", async () => {
const lib = await open();
const outDir = join(root, "backup");
await lib.backup({ downloadDirectory: outDir });
dropImageMetadata(outDir, 600);
const original = saved(outDir, "600.jpg");
chmodSync(original, 0o000);
const second = await lib
.backup({ downloadDirectory: outDir })
.finally(() => chmodSync(original, 0o600));
expect(second).toMatchObject({ failed: 0, errors: [] });
expect(existsSync(join(outDir, "failures.json"))).toBe(false);
expect(fileJSON(outDir, 600)).not.toHaveProperty("imageMetadata");
expect(fileJSON(outDir, 600).imageMetadataError).toMatch(/EACCES/);
await lib.close();
});
});
// With `verify`, each original already stored is hashed, and one that does not
// match the content hash its metadata records is downloaded again.
describe("backup with verify", () => {
// MockClient's files, each recording the hash of what `stubSource` writes
// for it, except diagram.png (200), which records none.
class HashedClient extends MockClient {
override async filesSince(args: {
collectionID: number;
}): Promise<FilesPage> {
const page = await super.filesSince(args);
const files = page.files.map((f) =>
f.id === 200
? f
: {
...f,
metadata: {
...f.metadata,
hash: blake2b(Buffer.alloc(SIZE_BY_ID[f.id]!)),
},
},
);
return { ...page, files };
}
}
// A backup of the account in `root/backup`, the library that made it, and
// the source it fetched from.
const backedUp = async (): Promise<{
lib: Library;
source: StubSource;
outDir: string;
}> => {
const source = stubSource();
const lib = await openLibrary(source, new HashedClient());
const outDir = join(root, "backup");
await lib.backup({ downloadDirectory: outDir });
return { lib, source, outDir };
};
// What `stubSource` writes for beach.jpg (100), and other bytes of the
// same length.
const good = Buffer.alloc(SIZE_BY_ID[100]!);
const corrupt = Buffer.alloc(SIZE_BY_ID[100]!, 1);
it("leaves an original that matches its hash, and one with no hash, as they are", async () => {
const { lib, source, outDir } = await backedUp();
const calls = source.originalCalls;
const result = await lib.backup({
downloadDirectory: outDir,
verify: true,
});
expect(result).toMatchObject({
downloaded: 0,
skipped: 3,
verified: 2,
mismatched: 0,
unchecked: 1,
failed: 0,
});
expect(source.originalCalls).toBe(calls);
await lib.close();
});
it("downloads again an original that does not match its hash", async () => {
const { lib, outDir } = await backedUp();
writeFileSync(saved(outDir, "100.jpg"), corrupt);
const log: string[] = [];
const result = await lib.backup({
downloadDirectory: outDir,
verify: true,
onProgress: (msg) => log.push(msg),
});
expect(result).toMatchObject({
downloaded: 1,
skipped: 2,
verified: 1,
mismatched: 1,
unchecked: 1,
failed: 0,
});
expect(readFileSync(saved(outDir, "100.jpg"))).toEqual(good);
expect(log.filter((msg) => msg.startsWith("MISMATCH"))).toEqual([
"MISMATCH original beach.jpg (100): its bytes do not match its content hash",
]);
await lib.close();
});
it("records a failed download in failures.json, with the original removed", async () => {
const { lib, source, outDir } = await backedUp();
writeFileSync(saved(outDir, "100.jpg"), corrupt);
source.failID = 100;
const result = await lib.backup({
downloadDirectory: outDir,
verify: true,
});
expect(result).toMatchObject({
downloaded: 0,
mismatched: 1,
failed: 1,
});
expect(Object.keys(readLedger(outDir).files)).toEqual(["100"]);
expect(existsSync(saved(outDir, "100.jpg"))).toBe(false);
await lib.close();
});
it("records in failures.json an original the cache puts back with the same bad bytes", async () => {
const lib = await openLibrary(stubSource(), new HashedClient());
const outDir = join(root, "backup");
// The cache holds a bad copy, and a backup copies an original the
// cache holds to its save path.
const cached = await lib.photos.byID({ fileID: 100 })!.original();
writeFileSync(cached.path, corrupt);
await lib.backup({ downloadDirectory: outDir });
const result = await lib.backup({
downloadDirectory: outDir,
verify: true,
});
expect(result).toMatchObject({
downloaded: 0,
mismatched: 1,
failed: 1,
});
expect(result.errors.map((e) => e.error)).toEqual([
"the original put back does not match its content hash either",
]);
expect(Object.keys(readLedger(outDir).files)).toEqual(["100"]);
expect(readFileSync(saved(outDir, "100.jpg"))).toEqual(corrupt);
await lib.close();
});
it("hashes nothing without verify", async () => {
const { lib, outDir } = await backedUp();
writeFileSync(saved(outDir, "100.jpg"), corrupt);
const result = await lib.backup({ downloadDirectory: outDir });
expect(result).toMatchObject({
downloaded: 0,
skipped: 3,
verified: 0,
mismatched: 0,
unchecked: 0,
failed: 0,
});
expect(readFileSync(saved(outDir, "100.jpg"))).toEqual(corrupt);
await lib.close();
});
// Root ignores file permissions, so this fails when run as root. The
// `test` phase of the `Dockerfile` runs as the `node` user.
it("records an original it cannot read as failed, and leaves it", async () => {
const { lib, outDir } = await backedUp();
const original = saved(outDir, "100.jpg");
chmodSync(original, 0o000);
const result = await lib
.backup({ downloadDirectory: outDir, verify: true })
.finally(() => chmodSync(original, 0o600));
expect(result).toMatchObject({ downloaded: 0, verified: 1, failed: 1 });
expect(result.errors.map((e) => e.fileID)).toEqual([100]);
expect(result.errors[0]!.error).toMatch(/EACCES/);
expect(readFileSync(original)).toEqual(good);
await lib.close();
});
});
// Every entry under collections/, one level of directories deep, with each
// symlink's target.
const tree = (outDir: string): string[] => {
@@ -1588,6 +1942,64 @@ describe("backup of live photos", () => {
},
);
it("verifies a live photo's image and video together, and downloads it again when one does not match", async () => {
const { file: live, body } = await asLivePhoto(
file(500, 10, "IMG_0500.HEIC"),
);
const lib = await open([live], new Map([[500, body]]));
const outDir = join(root, "backup");
await lib.backup({ downloadDirectory: outDir });
const good = await lib.backup({
downloadDirectory: outDir,
verify: true,
});
expect(good).toMatchObject({ skipped: 1, verified: 1, mismatched: 0 });
const video = saved(outDir, "500.mov");
writeFileSync(video, "another few seconds of video");
const bad = await lib.backup({
downloadDirectory: outDir,
verify: true,
});
expect(bad).toMatchObject({
downloaded: 1,
verified: 0,
mismatched: 1,
failed: 0,
});
expect(readFileSync(saved(outDir, "500.heic"))).toEqual(
Buffer.from(IMAGE),
);
expect(readFileSync(video)).toEqual(Buffer.from(VIDEO));
expect(tree(outDir)).toEqual(linked);
await lib.close();
});
it("removes both files of a live photo that does not match its hash when downloading it again fails", async () => {
const { file: live, body } = await asLivePhoto(
file(500, 10, "IMG_0500.HEIC"),
);
const bodies = new Map([[500, body]]);
const lib = await open([live], bodies);
const outDir = join(root, "backup");
await lib.backup({ downloadDirectory: outDir });
writeFileSync(saved(outDir, "500.mov"), "another few seconds of video");
bodies.delete(500);
const result = await lib.backup({
downloadDirectory: outDir,
verify: true,
});
expect(result).toMatchObject({ mismatched: 1, failed: 1 });
expect(Object.keys(readLedger(outDir).files)).toEqual(["500"]);
expect(existsSync(saved(outDir, "500.heic"))).toBe(false);
expect(existsSync(saved(outDir, "500.mov"))).toBe(false);
await lib.close();
});
it("serves a live photo the backup stored to a library reading the backup", async () => {
const { file: live, body } = await asLivePhoto(
file(500, 10, "IMG_0500.HEIC"),
+75 -3
View File
@@ -53,13 +53,14 @@ import {
import { run } from "../../src/cli-run.js";
import { loadSession } from "../../src/cli-session.js";
import type { Client, ClientSnapshot, LoginOptions } from "../../src/client.js";
import type { ContentSource } from "../../src/library/content.js";
import { savePath, type ContentSource } from "../../src/library/content.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
import { init, toBase64 } from "../../src/crypto/index.js";
import { defaultCacheDirectory } from "../../src/library/index.js";
import { HEIC_WITH_EXIF } from "../exif-heic.js";
import {
asLivePhoto,
blake2b,
cdnSource,
IMAGE,
livePhotoHash,
@@ -704,6 +705,7 @@ describe("backup", () => {
" Failed: 0\n",
);
expect(stdout.text).toBe("");
expect(existsSync(join(dir, "backup.lock"))).toBe(false);
});
// The backup opens its library with the precache off: it fetches the
@@ -743,6 +745,61 @@ describe("backup", () => {
expect(stderr.text).toBe("Starting backup...\n");
});
it("--verify downloads again an original that does not match its hash, prints the counts, and exits 0", async () => {
// Each file records the hash of the original the fake writes for it.
const client = {
...fakeClient(),
filesSince: async (args: { collectionID: number }) => ({
files: (FILES[args.collectionID] ?? []).map((f) => ({
...f,
metadata: {
...f.metadata,
hash: blake2b(Buffer.alloc(7, f.id & 0xff)),
},
})),
deleted: [],
cursor: 1,
}),
} as unknown as Client;
const ctx = context(client);
const dir = join(root, "backup");
expect(await backupCommand(ctx, dir, {})).toBe(0);
writeFileSync(savePath(dir, FILES[1]![0]!), "corrupt");
expect(await backupCommand(ctx, dir, { verify: true })).toBe(0);
expect(stderr.text).toContain(
"MISMATCH original beach.jpg (100): its bytes do not match its content hash\n",
);
expect(stderr.text).toContain(
" Downloaded: 1\n" +
" Skipped: 2\n" +
" Verified: 2\n" +
" Mismatched: 1\n" +
" Unchecked: 0\n" +
" Failed: 0\n",
);
});
it("--verify --json adds the verified, mismatched and unchecked counts", async () => {
const dir = join(root, "backup");
expect(await backupCommand(context(), dir, {})).toBe(0);
const code = await backupCommand(context(), dir, {
verify: true,
json: true,
});
expect(code).toBe(0);
expect(JSON.parse(stdout.text)).toMatchObject({
skipped: 3,
verified: 0,
mismatched: 0,
unchecked: 3,
failed: 0,
});
});
it("exits 1 and lists each file when the ML data fetch fails", async () => {
const client = {
...fakeClient(),
@@ -827,7 +884,7 @@ describe("backup", () => {
expect(readdirSync(dir)).toEqual([]);
});
it("exits 2 with one line naming the directory while another backup of it runs", async () => {
it("exits 2 with one line naming the directory, sending no request, while another backup of it runs", async () => {
const dir = join(root, "backup");
// The lock another backup holds. Its modification time is set an hour
// ahead, so it stays current however long this test takes.
@@ -835,12 +892,27 @@ describe("backup", () => {
mkdirSync(lock, { recursive: true });
const hourAhead = new Date(Date.now() + 3_600_000);
utimesSync(lock, hourAhead, hourAhead);
// A client whose refresh never finishes, so a run that started one
// before it exits would never return.
let requests = 0;
const never = (): Promise<never> => {
requests++;
return new Promise(() => {});
};
const client = {
...fakeClient(),
collectionsSince: never,
filesSince: never,
} as unknown as Client;
expect(await backupCommand(context(), dir, {})).toBe(2);
expect(await backupCommand(context(client), dir, {})).toBe(2);
expect(requests).toBe(0);
expect(stderr.text).toBe(
`Starting backup...\nquak: another backup of ${dir} is running\n`,
);
expect(readdirSync(dir)).toEqual(["backup.lock"]);
// Opening the library would have made its cache directory.
expect(existsSync(join(root, "cache"))).toBe(false);
});
});
+2 -2
View File
@@ -1,7 +1,7 @@
/**
* `exif.heic`, beside this file: a real 64x64 HEIC whose EXIF holds the same
* values as the hand-built JPEG in `library/content-library.test.ts`, for the
* tests of `exif()` and `backup-metadata --exif`.
* values as the hand-built JPEG in `exif-jpeg.ts`, for the tests of `exif()`,
* `backup-metadata --exif` and the image metadata `quak backup` records.
*
* It was made once, in a throwaway node:22-alpine container (Alpine 3.23.3),
* with libheif 1.23.0 and exiftool 13.55:
+89
View File
@@ -0,0 +1,89 @@
/**
* Hand-built JPEGs for the tests of `exif()` and of the image metadata
* `quak backup` records: one whose EXIF holds the same values as `exif.heic`
* (see `exif-heic.ts`), and one whose EXIF cannot be parsed.
*/
// Big-endian bytes for the hand-built JPEG below.
const u16 = (n: number): number[] => [n >> 8, n & 0xff];
const u32 = (n: number): number[] => [...u16(n >>> 16), ...u16(n & 0xffff)];
const ascii = (s: string): number[] => [...new TextEncoder().encode(s), 0];
const rational = (num: number, den: number): number[] => [
...u32(num),
...u32(den),
];
// One IFD entry: tag, type (1 BYTE, 2 ASCII, 3 SHORT, 4 LONG, 5 RATIONAL),
// count, then the value when it fits in 4 bytes, else its offset.
const entry = (
tag: number,
type: number,
count: number,
value: number[],
): number[] => [...u16(tag), ...u16(type), ...u32(count), ...value];
// The TIFF block of a JPEG's EXIF segment, holding every field `Photo`'s typed
// methods return: the camera in the first IFD, the exposure in the Exif IFD,
// and a GPS position of 40°26'46" N, 79°58'56" W, 12.5 m below sea level.
// Offsets count from the start of this block.
const TIFF = [
...[0x4d, 0x4d, 0x00, 0x2a], // big-endian TIFF
...u32(8), // the first IFD's offset
// The first IFD, at 8: five entries, then no next IFD.
...u16(5),
...entry(0x010f, 2, 6, u32(74)), // Make
...entry(0x0110, 2, 7, u32(80)), // Model
...entry(0x0112, 3, 1, [...u16(6), 0, 0]), // Orientation
...entry(0x8769, 4, 1, u32(88)), // the Exif IFD's offset
...entry(0x8825, 4, 1, u32(246)), // the GPS IFD's offset
...u32(0),
...ascii("Canon"), // at 74
...ascii("EOS R5"), // at 80
0, // a pad byte
// The Exif IFD, at 88: seven entries, then no next IFD.
...u16(7),
...entry(0x829a, 5, 1, u32(178)), // ExposureTime
...entry(0x829d, 5, 1, u32(186)), // FNumber
...entry(0x8827, 3, 1, [...u16(400), 0, 0]), // ISOSpeedRatings
...entry(0x9003, 2, 20, u32(194)), // DateTimeOriginal
...entry(0x9011, 2, 7, u32(214)), // OffsetTimeOriginal
...entry(0x920a, 5, 1, u32(222)), // FocalLength
...entry(0xa434, 2, 16, u32(230)), // LensModel
...u32(0),
...rational(1, 250), // at 178
...rational(28, 10), // at 186
...ascii("2021:07:15 14:30:00"), // at 194
...ascii("+02:00"), // at 214
0, // a pad byte
...rational(50, 1), // at 222
...ascii("RF50mm F1.8 STM"), // at 230
// The GPS IFD, at 246: six entries, then no next IFD.
...u16(6),
...entry(0x0001, 2, 2, [...ascii("N"), 0, 0]), // GPSLatitudeRef
...entry(0x0002, 5, 3, u32(324)), // GPSLatitude
...entry(0x0003, 2, 2, [...ascii("W"), 0, 0]), // GPSLongitudeRef
...entry(0x0004, 5, 3, u32(348)), // GPSLongitude
...entry(0x0005, 1, 1, [1, 0, 0, 0]), // GPSAltitudeRef: below sea level
...entry(0x0006, 5, 1, u32(372)), // GPSAltitude
...u32(0),
...[...rational(40, 1), ...rational(26, 1), ...rational(46, 1)], // at 324
...[...rational(79, 1), ...rational(58, 1), ...rational(56, 1)], // at 348
...rational(25, 2), // at 372
];
export const JPEG_WITH_EXIF = new Uint8Array([
...[0xff, 0xd8], // start of image
...[0xff, 0xe1, ...u16(2 + 6 + TIFF.length)], // APP1 and its length
...[...ascii("Exif"), 0], // "Exif\0\0"
...TIFF,
...[0xff, 0xda, 0x00, 0x02], // start of scan
]);
// A JPEG whose EXIF segment is laid out correctly but holds "XX" where the TIFF
// byte order belongs, so exifreader cannot parse it.
export const JPEG_WITH_BAD_EXIF = new Uint8Array([
...[0xff, 0xd8], // start of image
...[0xff, 0xe1, ...u16(2 + 6 + 2)], // APP1 and its length
...[...ascii("Exif"), 0], // "Exif\0\0"
...[0x58, 0x58], // "XX"
...[0xff, 0xda, 0x00, 0x02], // start of scan
]);
+1 -84
View File
@@ -29,6 +29,7 @@ import type { CollectionsPage, FilesPage } from "../../src/client.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
import { readPhotoExif, type PhotoExif } from "../../src/exif.js";
import { HEIC_WITH_EXIF } from "../exif-heic.js";
import { JPEG_WITH_BAD_EXIF, JPEG_WITH_EXIF } from "../exif-jpeg.js";
import {
asLivePhoto,
cdnSource,
@@ -256,90 +257,6 @@ describe("Library content wiring", () => {
});
});
// Big-endian bytes for the hand-built JPEG below.
const u16 = (n: number): number[] => [n >> 8, n & 0xff];
const u32 = (n: number): number[] => [...u16(n >>> 16), ...u16(n & 0xffff)];
const ascii = (s: string): number[] => [...new TextEncoder().encode(s), 0];
const rational = (num: number, den: number): number[] => [
...u32(num),
...u32(den),
];
// One IFD entry: tag, type (1 BYTE, 2 ASCII, 3 SHORT, 4 LONG, 5 RATIONAL),
// count, then the value when it fits in 4 bytes, else its offset.
const entry = (
tag: number,
type: number,
count: number,
value: number[],
): number[] => [...u16(tag), ...u16(type), ...u32(count), ...value];
// The TIFF block of a JPEG's EXIF segment, holding every field `Photo`'s typed
// methods return: the camera in the first IFD, the exposure in the Exif IFD,
// and a GPS position of 40°26'46" N, 79°58'56" W, 12.5 m below sea level.
// Offsets count from the start of this block.
const TIFF = [
...[0x4d, 0x4d, 0x00, 0x2a], // big-endian TIFF
...u32(8), // the first IFD's offset
// The first IFD, at 8: five entries, then no next IFD.
...u16(5),
...entry(0x010f, 2, 6, u32(74)), // Make
...entry(0x0110, 2, 7, u32(80)), // Model
...entry(0x0112, 3, 1, [...u16(6), 0, 0]), // Orientation
...entry(0x8769, 4, 1, u32(88)), // the Exif IFD's offset
...entry(0x8825, 4, 1, u32(246)), // the GPS IFD's offset
...u32(0),
...ascii("Canon"), // at 74
...ascii("EOS R5"), // at 80
0, // a pad byte
// The Exif IFD, at 88: seven entries, then no next IFD.
...u16(7),
...entry(0x829a, 5, 1, u32(178)), // ExposureTime
...entry(0x829d, 5, 1, u32(186)), // FNumber
...entry(0x8827, 3, 1, [...u16(400), 0, 0]), // ISOSpeedRatings
...entry(0x9003, 2, 20, u32(194)), // DateTimeOriginal
...entry(0x9011, 2, 7, u32(214)), // OffsetTimeOriginal
...entry(0x920a, 5, 1, u32(222)), // FocalLength
...entry(0xa434, 2, 16, u32(230)), // LensModel
...u32(0),
...rational(1, 250), // at 178
...rational(28, 10), // at 186
...ascii("2021:07:15 14:30:00"), // at 194
...ascii("+02:00"), // at 214
0, // a pad byte
...rational(50, 1), // at 222
...ascii("RF50mm F1.8 STM"), // at 230
// The GPS IFD, at 246: six entries, then no next IFD.
...u16(6),
...entry(0x0001, 2, 2, [...ascii("N"), 0, 0]), // GPSLatitudeRef
...entry(0x0002, 5, 3, u32(324)), // GPSLatitude
...entry(0x0003, 2, 2, [...ascii("W"), 0, 0]), // GPSLongitudeRef
...entry(0x0004, 5, 3, u32(348)), // GPSLongitude
...entry(0x0005, 1, 1, [1, 0, 0, 0]), // GPSAltitudeRef: below sea level
...entry(0x0006, 5, 1, u32(372)), // GPSAltitude
...u32(0),
...[...rational(40, 1), ...rational(26, 1), ...rational(46, 1)], // at 324
...[...rational(79, 1), ...rational(58, 1), ...rational(56, 1)], // at 348
...rational(25, 2), // at 372
];
const JPEG_WITH_EXIF = new Uint8Array([
...[0xff, 0xd8], // start of image
...[0xff, 0xe1, ...u16(2 + 6 + TIFF.length)], // APP1 and its length
...[...ascii("Exif"), 0], // "Exif\0\0"
...TIFF,
...[0xff, 0xda, 0x00, 0x02], // start of scan
]);
// A JPEG whose EXIF segment is laid out correctly but holds "XX" where the TIFF
// byte order belongs, so exifreader cannot parse it.
const JPEG_WITH_BAD_EXIF = new Uint8Array([
...[0xff, 0xd8], // start of image
...[0xff, 0xe1, ...u16(2 + 6 + 2)], // APP1 and its length
...[...ascii("Exif"), 0], // "Exif\0\0"
...[0x58, 0x58], // "XX"
...[0xff, 0xda, 0x00, 0x02], // start of scan
]);
describe("Photo save path, local copy, content and EXIF", () => {
// The same account, with `files` in its album instead.
class FilesClient extends MockClient {
+2 -1
View File
@@ -30,7 +30,8 @@ export const livePhotoZip = (
},
): Uint8Array => zipSync(entries);
const blake2b = (bytes: Uint8Array): string =>
// The content hash Ente's clients record for an original's bytes.
export const blake2b = (bytes: Uint8Array): string =>
createHash("blake2b512").update(bytes).digest("base64");
// The hash Ente's clients record for a live photo: the unkeyed BLAKE2b-512 of