Store live photos as their image and their video (closes #107)
check / check (push) Successful in 41s

A live photo, which Ente stores as one ZIP, is unpacked as it downloads
into its image and its video, each `<fileID>.<ext>` with its extension
from the ZIP, beside `<fileID>.livephoto.json`, which names the two.
Both are checked against the recorded hash and renamed into place only
when both are complete. A ZIP with a second image or video, or whose
parts come to more than 20 times its size plus 16 MiB, is refused. The
backup and the content cache count a live photo as stored only with both
files, album folders link both, `quak get` writes both, and the content
result gives the video as `videoPath`. A ZIP an earlier version stored
is replaced.

Model: opus-5-5
This commit was merged in pull request #128.
This commit is contained in:
2026-09-28 18:06:01 +02:00
parent 04094a8cfb
commit 9e6deb21eb
17 changed files with 1746 additions and 366 deletions
+105 -53
View File
@@ -5,7 +5,7 @@
// gets its original bytes onto disk under `downloadDirectory` and rebuilds the
// derived views (per-file sidecars, per-collection symlink trees,
// per-collection JSON) from the model. The on-disk layout is the historical
// one, unchanged:
// one:
//
// <downloadDirectory>/
// originals/<fileID>.<ext> the decrypted bytes
@@ -14,6 +14,10 @@
// collections/<name>.json per-collection metadata
// failures.json durable ledger of unresolved failures
//
// A live photo's original is its image and its video, `<fileID>.<ext>` each
// with its own extension, and `originals/<fileID>.livephoto.json` naming them;
// its album folders link both.
//
// Crash-safety rests on two properties. Bytes are present-means-complete: an
// original appears under `originals/` only via the content layer's atomic
// temp-then-rename, so a file that exists is whole and is never re-fetched — an
@@ -48,7 +52,12 @@ import { copyFile, rename, rm } from "node:fs/promises";
import { basename, dirname, extname, join, relative } from "node:path";
import { fsyncPath, removeLeftoverTempFiles } from "./download/index.js";
import { safeExtension, sanitizeFileName } from "./filename.js";
import { sanitizeFileName, withExtension } from "./filename.js";
import {
nameInOriginals,
storedOriginal,
writeLivePhotoJSON,
} from "./library/content.js";
import type { Collection, EnteFile } from "./model/types.js";
export type ProgressCallback = (message: string) => void;
@@ -98,8 +107,13 @@ export interface BackupLibrary {
listFiles(collectionID: number): EnteFile[];
// Get an original's bytes onto disk through the content cache/pools,
// returning where they landed: `destination` when they were fetched now,
// otherwise wherever they already were (the cache, or a prior backup).
original(fileID: number, destination: string): Promise<{ path: string }>;
// otherwise wherever they already were (the cache, or a prior backup). A
// live photo lands as its image and its video, fetched now beside
// `destination`.
original(
fileID: number,
destination: string,
): Promise<{ path: string; videoPath?: string }>;
thumbnail(fileID: number): Promise<{ path: string }>;
}
@@ -116,12 +130,6 @@ interface FailureEntry {
const LEDGER_VERSION = 1;
// The originals/ filename for a file: `<id><ext>`, the extension taken from the
// title (or `.bin`). Matches the content cache's own naming so a present check
// lines up with what a fetch would write.
const originalName = (file: EnteFile): string =>
`${file.id}${safeExtension(file.metadata.title)}`;
// A regular file with content is treated as complete. A zero-byte file is not:
// it is the shape an aborted write leaves and must be re-fetched.
const isPresent = (path: string): boolean => {
@@ -187,6 +195,31 @@ const copyAtomic = async (src: string, dest: string): Promise<void> => {
}
};
// Put an original the library returned at `dest` in originals/, where a fresh
// fetch already wrote it. A live photo's image and video go beside `dest`: when
// they came from the cache they are copied, after removing whatever was at
// `dest` (an earlier version's ZIP of the two). Then the JSON file naming them
// is written, which is what makes the live photo count as stored.
const placeOriginal = async (
file: EnteFile,
dest: string,
got: { path: string; videoPath?: string },
): Promise<void> => {
if (got.videoPath === undefined) {
await copyAtomic(got.path, dest);
return;
}
const originalsDir = dirname(dest);
const path = join(originalsDir, basename(got.path));
const videoPath = join(originalsDir, basename(got.videoPath));
if (got.path !== path) {
await rm(dest, { force: true });
await copyAtomic(got.path, path);
await copyAtomic(got.videoPath, videoPath);
}
await writeLivePhotoJSON(originalsDir, file.id, { path, videoPath });
};
// Ensure `linkPath` is a symlink to `target`, rebuilding a missing, wrong, or
// non-symlink entry. Throws on failure (a directory in the way, no permission)
// so the caller records it and moves on rather than aborting the run.
@@ -203,43 +236,61 @@ const rebuildSymlink = (linkPath: string, target: string): void => {
symlinkSync(target, linkPath);
};
// The on-disk names for the entries of one directory, keyed by ID. Each name
// is used as is unless another entry would get the same name, ignoring case
// (two names that differ only in case are one entry on a case-insensitive
// The on-disk names for the entries of one directory, in entry order. Each
// name is used as is unless another entry would get the same name, ignoring
// case (two names that differ only in case are one entry on a case-insensitive
// file system); then every entry sharing it gets ` (<id>)`, before the
// extension when `beforeExtension` is set. A name with an ID added can match
// another entry's own name (`IMG (6).JPG`), so this repeats until no name is
// shared. IDs are stable, so the names are too.
const namesByID = (
const uniqueNames = (
entries: { id: number; name: string }[],
beforeExtension: boolean,
): Map<number, string> => {
): string[] => {
const withID = (id: number, name: string): string => {
const ext = beforeExtension ? extname(name) : "";
const stem = name.slice(0, name.length - ext.length);
return `${stem} (${id})${ext}`;
};
const names = new Map<number, string>();
for (const { id, name } of entries) names.set(id, name);
const names = entries.map((e) => e.name);
const suffixed = new Set<number>();
for (;;) {
const counts = new Map<string, number>();
for (const name of names.values()) {
for (const name of names) {
const key = name.toLowerCase();
counts.set(key, (counts.get(key) ?? 0) + 1);
}
let changed = false;
for (const { id, name } of entries) {
if (suffixed.has(id)) continue;
for (const [i, { id, name }] of entries.entries()) {
if (suffixed.has(i)) continue;
if (counts.get(name.toLowerCase()) === 1) continue;
names.set(id, withID(id, name));
suffixed.add(id);
names[i] = withID(id, name);
suffixed.add(i);
changed = true;
}
if (!changed) return names;
}
};
// The links a file gets in its album's folder: one named after its title, to
// its original if that is stored. A stored live photo gets two, to its image
// and its video, each named after the title with that file's extension.
const linksFor = (
file: EnteFile,
stored: { path: string; videoPath?: string } | undefined,
): { id: number; name: string; file: EnteFile; target?: string }[] => {
const name = sanitizeFileName(file.metadata.title, `file-${file.id}`);
if (stored?.videoPath === undefined) {
return [{ id: file.id, name, file, target: stored?.path }];
}
return [stored.path, stored.videoPath].map((target) => ({
id: file.id,
name: withExtension(name, extname(target)),
file,
target,
}));
};
// Remove the symlinks in the album directory `dir` that point into
// `originalsDir` and are not named in `keep`. Nothing else in the directory
// is touched: anything else there was put there by the user.
@@ -419,17 +470,21 @@ export const runBackup = async (
// tree; a present file is left as is.
if (includeOriginals) {
for (const [fileID, file] of distinct) {
const dest = join(originalsDir, originalName(file));
if (isPresent(dest)) {
if (storedOriginal(originalsDir, file) !== undefined) {
skipped++;
continue;
}
const dest = join(originalsDir, nameInOriginals(file));
try {
log(`Fetching original ${file.metadata.title} (${fileID})...`);
// A fetched original is written straight to `dest`; only one
// that was already cached elsewhere is copied.
const { path } = await lib.original(fileID, dest);
await copyAtomic(path, dest);
// A fetched original is written straight to `dest` (a live
// photo beside it); only one that was already cached elsewhere
// is copied.
await placeOriginal(
file,
dest,
await lib.original(fileID, dest),
);
downloaded++;
} catch (err) {
log(
@@ -465,8 +520,7 @@ export const runBackup = async (
// every present original (this repairs stale ones).
if (includeOriginals) {
for (const [fileID, file] of distinct) {
const orig = join(originalsDir, originalName(file));
if (isPresent(orig)) {
if (storedOriginal(originalsDir, file) !== undefined) {
writeSidecar(join(originalsDir, `${fileID}.json`), file);
}
}
@@ -478,19 +532,18 @@ export const runBackup = async (
// an album it skipped. Stale entries are removed before anything is
// rebuilt, so on a case-insensitive file system removing an old name can
// never remove the new one.
const albumDirNames = namesByID(
const dirNames = uniqueNames(
allCollections.map((c) => ({
id: c.id,
name: sanitizeFileName(c.name, `collection-${c.id}`),
})),
false,
);
const albumDirNames = new Map(
allCollections.map((c, i) => [c.id, dirNames[i]!]),
);
try {
removeStaleAlbumDirs(
collectionsDir,
new Set(albumDirNames.values()),
originalsDir,
);
removeStaleAlbumDirs(collectionsDir, new Set(dirNames), originalsDir);
} catch (err) {
log(`FAILED removing old album directories: ${errorMessage(err)}`);
}
@@ -501,34 +554,33 @@ export const runBackup = async (
mkdirSync(colDir, { recursive: true });
const files = filesByCollection.get(c.id) ?? [];
const linkNames = namesByID(
files.map((f) => ({
id: f.id,
name: sanitizeFileName(f.metadata.title, `file-${f.id}`),
})),
true,
const links = files.flatMap((f) =>
linksFor(f, storedOriginal(originalsDir, f)),
);
const linkNames = uniqueNames(links, true);
try {
removeStaleLinks(colDir, new Set(linkNames.values()), originalsDir);
removeStaleLinks(colDir, new Set(linkNames), originalsDir);
} catch (err) {
log(`FAILED removing old links in ${c.name}: ${errorMessage(err)}`);
}
const metaFiles: { id: number; metadata: EnteFile["metadata"] }[] = [];
for (const file of files) {
metaFiles.push({ id: file.id, metadata: file.metadata });
if (!includeOriginals) continue;
const orig = join(originalsDir, originalName(file));
if (!isPresent(orig)) continue;
const linkName = linkNames.get(file.id)!;
const linkPath = join(colDir, linkName);
const metaFiles = files.map((f) => ({
id: f.id,
metadata: f.metadata,
}));
for (const [i, link] of links.entries()) {
if (!includeOriginals || link.target === undefined) continue;
const linkName = linkNames[i]!;
try {
rebuildSymlink(linkPath, relative(colDir, orig));
rebuildSymlink(
join(colDir, linkName),
relative(colDir, link.target),
);
} catch (err) {
log(
`FAILED symlink ${c.name}/${linkName}: ${errorMessage(err)}`,
);
recordFailure(file, c.name, err);
recordFailure(link.file, c.name, err);
}
}
+26 -3
View File
@@ -10,10 +10,11 @@ import {
copyFileSync,
existsSync,
mkdirSync,
statSync,
unlinkSync,
writeFileSync,
} from "node:fs";
import { join } from "node:path";
import { extname, join } from "node:path";
import {
type Client,
type ClientSnapshot,
@@ -32,6 +33,7 @@ import {
thumbnailName,
} from "./cli-output.js";
import { freshCollections, freshFiles, freshFile } from "./cli-read.js";
import { withExtension } from "./filename.js";
import { runMetadataBackup } from "./metadata-backup.js";
import { listMissingThumbnails, fixMissingThumbnails } from "./thumbnails.js";
@@ -305,8 +307,29 @@ export const getCommand = async (
// Default name is the file's own title, as the pre-library CLI used
// (not the editedName-preferring projection title) (issue #52).
const outPath = opts.out ?? originalName(file);
copyFileSync(result.path, outPath);
ctx.stderr.write(`${result.bytes} bytes -> ${outPath}\n`);
if (result.videoPath === undefined) {
copyFileSync(result.path, outPath);
ctx.stderr.write(`${result.bytes} bytes -> ${outPath}\n`);
return 0;
}
// A live photo is written as its image and its video, each named after
// the title with its own extension, as Ente's clients name them. With
// --out, the image goes there and the video beside it.
const imageOut =
opts.out ?? withExtension(outPath, extname(result.path));
const videoOut = withExtension(outPath, extname(result.videoPath));
if (imageOut.toLowerCase() === videoOut.toLowerCase()) {
ctx.stderr.write(
`File ${fileID} is a live photo, and its video would also be written to ${imageOut}\n`,
);
return 1;
}
copyFileSync(result.path, imageOut);
copyFileSync(result.videoPath, videoOut);
ctx.stderr.write(
`${result.bytes} bytes -> ${imageOut}\n` +
`${statSync(videoOut).size} bytes -> ${videoOut}\n`,
);
return 0;
} finally {
await lib.close();
+248 -133
View File
@@ -16,14 +16,18 @@ import {
streamTagFinal,
} from "../crypto/index.js";
import { TruncatedStreamError } from "../errors.js";
import { sanitizeFileName } from "../filename.js";
import { safeExtension, sanitizeFileName, withExtension } from "../filename.js";
import { withRetry } from "../retry.js";
import type { ApiClient } from "../api/client.js";
import type { EnteFile } from "../model/types.js";
export interface DownloadResult {
// Where the file was written. A live photo is written as two files, its
// image here and its video at `videoPath` (see `decryptLivePhoto`).
path: string;
// The decrypted length; for a live photo, that of the ZIP it arrives as.
bytesWritten: number;
videoPath?: string;
}
// Fired as decrypted plaintext accumulates, with the running total of
@@ -45,9 +49,9 @@ const ENC_CHUNK_SIZE = STREAM_CHUNK_SIZE + STREAM_CHUNK_OVERHEAD;
// new: a body cut short still decrypts and authenticates up to its last whole
// chunk, so the absence of TAG_FINAL is the sole evidence it was cut short, and
// this throws rather than let a caller keep a short file. The sink has already
// seen those chunks by then; the caller (`decryptToTemp`) stages them in a temp
// file that is renamed into place only on a clean return, so a throw leaves
// nothing on disk.
// seen those chunks by then; the callers (`decryptToTemp`, `decryptLivePhoto`)
// stage them in temp files that are renamed into place only on a clean return,
// so a throw leaves nothing on disk.
const streamDecrypt = async (
stream: ReadableStream<Uint8Array>,
header: Uint8Array,
@@ -213,6 +217,13 @@ export const removeLeftoverTempFiles = (dir: string): void => {
}
};
// A new temp file name in `dir`. The random suffix keeps concurrent downloads
// of the same destination from stepping on each other's temporary file; the
// process ID lets `removeLeftoverTempFiles` tell a leftover from a write in
// progress.
const tempPathIn = (dir: string): string =>
join(dir, `.quak-${process.pid}-${randomBytes(16).toString("hex")}.tmp`);
// Stage a write to `destination` atomically and durably, then rename it into
// place. `fill` writes the contents into the open temp file handle — either the
// whole buffer at once (`writeAtomic`) or chunk by chunk as they decrypt
@@ -237,13 +248,7 @@ const stageAtomic = async (
fill: (handle: FileHandle) => Promise<void>,
): Promise<void> => {
const dir = dirname(destination);
// The random suffix keeps concurrent downloads of the same destination
// from stepping on each other's temporary file; the process ID lets
// `removeLeftoverTempFiles` tell a leftover from a write in progress.
const tmpPath = join(
dir,
`.quak-${process.pid}-${randomBytes(16).toString("hex")}.tmp`,
);
const tmpPath = tempPathIn(dir);
try {
const handle = await open(tmpPath, "w");
try {
@@ -278,81 +283,14 @@ export const writeAtomic = async (
): Promise<void> =>
stageAtomic(destination, (handle) => handle.writeFile(plaintext));
// Hashes an original's bytes as they are decrypted, for comparison with the
// hash its uploader recorded.
interface ContentHasher {
update: (plaintext: Uint8Array) => void;
digest: () => string;
}
const fileHasher = (): ContentHasher => {
const state = chunkHashInit();
return {
update: (plaintext) => chunkHashUpdate(state, plaintext),
digest: () => chunkHashFinal(state),
};
};
// A live photo is stored as a ZIP of its image and its video, and its recorded
// hash is `<imageHash>:<videoHash>`, each over that part's own bytes. Like the
// upstream client's decoder, this takes the first entries whose names start
// with `image` and `video`.
//
// The ZIP is chosen by its uploader and may expand enormously, so entries are
// hashed as they decompress and never held. fflate's `Unzip` inflates each
// push in one piece, and deflate expands at most about 1000-fold, so the ZIP
// is pushed in 4 KiB slices to keep each decompressed piece near 4 MiB, one
// plaintext chunk. Every entry is started, even one that is not hashed,
// because fflate keeps an unstarted entry's data in memory.
const livePhotoHasher = (fileID: number): ContentHasher => {
const sliceSize = 4096;
const fail = (message: string, cause?: unknown): Error =>
new Error(`download: file ${fileID}: ${message}`, { cause });
const claimed = new Set<string>();
const hashes = new Map<string, string>();
const unzip = new Unzip((entry) => {
const part = ["image", "video"].find((p) => entry.name.startsWith(p));
const target =
part === undefined || claimed.has(part)
? undefined
: { part, state: chunkHashInit() };
if (target !== undefined) claimed.add(target.part);
entry.ondata = (err, data, final) => {
if (err) throw err;
if (target === undefined) return;
chunkHashUpdate(target.state, data);
if (final) hashes.set(target.part, chunkHashFinal(target.state));
};
entry.start();
});
unzip.register(UnzipInflate);
// fflate reports a bad ZIP by throwing, sometimes a TypeError, which the
// retry would take for a network failure; a bad ZIP is never retried.
const push = (data: Uint8Array, final: boolean): void => {
try {
unzip.push(data, final);
} catch (err) {
throw fail("live photo is not a readable ZIP", err);
}
};
return {
update: (plaintext) => {
for (let i = 0; i < plaintext.length; i += sliceSize) {
push(plaintext.subarray(i, i + sliceSize), false);
}
},
digest: () => {
push(new Uint8Array(0), true);
const image = hashes.get("image");
const video = hashes.get("video");
if (image === undefined || video === undefined) {
throw fail(
"live photo ZIP does not hold both an image and a video",
);
}
return `${image}:${video}`;
},
};
// Refuse an original whose bytes do not hash to what its uploader recorded.
// The error is not retried.
const checkHash = (file: EnteFile, actual: string): void => {
if (actual !== file.metadata.hash) {
throw new Error(
`download: file ${file.id}: content hash ${actual} does not match the hash its uploader recorded, ${file.metadata.hash}`,
);
}
};
// Decrypt `stream` straight to `destination`, one plaintext chunk at a time,
@@ -363,9 +301,8 @@ const livePhotoHasher = (fileID: number): ContentHasher => {
// Returns the plaintext length written.
//
// `original` is the file whose original this is (none for a thumbnail, which
// has no recorded hash). When its metadata has a hash, the decrypted bytes
// must match it or nothing is stored. Both a plain file and a live photo's
// parts are hashed as they stream. The mismatch error is not retried.
// has no recorded hash). When its metadata has a hash, the decrypted bytes are
// hashed as they stream and must match it, or nothing is stored.
const decryptToTemp = async (
destination: string,
stream: ReadableStream<Uint8Array>,
@@ -374,44 +311,199 @@ const decryptToTemp = async (
onProgress?: ProgressCallback,
original?: EnteFile,
): Promise<number> => {
const expected = original?.metadata.hash;
const hasher =
original === undefined || expected === undefined
? undefined
: original.metadata.fileType === "livePhoto"
? livePhotoHasher(original.id)
: fileHasher();
const hash =
original?.metadata.hash === undefined ? undefined : chunkHashInit();
let bytesWritten = 0;
await stageAtomic(destination, async (handle) => {
bytesWritten = await streamDecrypt(
stream,
header,
key,
async (plaintext) => {
if (hash !== undefined) chunkHashUpdate(hash, plaintext);
await handle.write(plaintext);
},
onProgress,
);
if (original !== undefined && hash !== undefined) {
checkHash(original, chunkHashFinal(hash));
}
});
return bytesWritten;
};
// One of the two parts of a live photo being unpacked: the ZIP entry whose
// name starts with `kind`, written to its own temp file.
interface LivePhotoPart {
kind: "image" | "video";
tmpPath: string;
handle: FileHandle;
hash: ReturnType<typeof chunkHashInit>;
// Decompressed bytes not yet written.
pending: Uint8Array[];
// The entry's extension, set once all of the entry has been read.
ext?: string;
}
const openPart = async (
kind: "image" | "video",
dir: string,
): Promise<LivePhotoPart> => {
const tmpPath = tempPathIn(dir);
const handle = await open(tmpPath, "w");
return { kind, tmpPath, handle, hash: chunkHashInit(), pending: [] };
};
// A live photo arrives as a ZIP of its image and its video. Ente's clients
// name the entries `image.<ext>` and `video.<ext>`. This takes the entry whose
// name starts with `image` as the image and the one whose name starts with
// `video` as the video, and refuses a ZIP holding a second of either. It is
// written unpacked: each part is named `destination` with the extension
// replaced by its own entry's, and the two must differ ignoring case. When the
// file records a hash, `<imageHash>:<videoHash>` must match it, each over that
// part's own bytes. Only then is whatever was at `destination` removed and the
// image, then the video, renamed into place; on any failure neither is stored.
//
// The ZIP is chosen by its uploader and may expand enormously, so each part is
// written as it decompresses and never held, and the ZIP is refused once the
// two together come to more than 20 times its size plus 16 MiB, the limit
// Ente's mobile client and CLI set. The ZIP's size is known only at its end,
// so until then the limit is taken over the part of it decrypted so far.
// fflate's `Unzip` inflates each push in one piece before `push` returns, and
// deflate expands at most about 1000-fold, so the ZIP is pushed in 4 KiB
// slices, keeping each decompressed piece near 4 MiB, one plaintext chunk, and
// each piece is checked and written before the next slice is pushed. Every
// entry is started, even one that is not kept, because fflate keeps an
// unstarted entry's data in memory.
const decryptLivePhoto = async (
destination: string,
stream: ReadableStream<Uint8Array>,
header: Uint8Array,
key: Uint8Array,
onProgress: ProgressCallback | undefined,
file: EnteFile,
): Promise<DownloadResult> => {
const sliceSize = 4096;
const dir = dirname(destination);
const fail = (message: string, cause?: unknown): Error =>
new Error(`download: file ${file.id}: ${message}`, { cause });
const parts: LivePhotoPart[] = [];
try {
await stageAtomic(destination, async (handle) => {
bytesWritten = await streamDecrypt(
stream,
header,
key,
async (plaintext) => {
hasher?.update(plaintext);
await handle.write(plaintext);
},
onProgress,
);
if (original === undefined || hasher === undefined) return;
const actual = hasher.digest();
if (actual !== expected) {
throw new Error(
`download: file ${original.id}: content hash ${actual} does not match the hash its uploader recorded, ${expected}`,
const image = await openPart("image", dir);
parts.push(image);
const video = await openPart("video", dir);
parts.push(video);
const claimed = new Set<LivePhotoPart>();
// Set to a part the ZIP holds a second entry for; the ZIP is then
// refused.
let repeated: LivePhotoPart | undefined;
const unzip = new Unzip((entry) => {
const part = parts.find((p) => entry.name.startsWith(p.kind));
if (part !== undefined) {
if (claimed.has(part)) repeated = part;
claimed.add(part);
}
entry.ondata = (err, data, final) => {
if (err) throw err;
if (part === undefined) return;
chunkHashUpdate(part.hash, data);
part.pending.push(data);
if (final) part.ext = safeExtension(entry.name);
};
entry.start();
});
unzip.register(UnzipInflate);
// The bytes of the ZIP pushed so far, and of the two parts they have
// decompressed to.
let zipBytes = 0;
let expanded = 0;
const push = async (
data: Uint8Array,
final: boolean,
): Promise<void> => {
zipBytes += data.length;
// fflate reports a bad ZIP by throwing, sometimes a TypeError,
// which the retry would take for a network failure. A bad ZIP is
// never retried, and nor are the refusals below.
try {
unzip.push(data, final);
} catch (err) {
throw fail("live photo is not a readable ZIP", err);
}
if (repeated !== undefined) {
throw fail(
`live photo ZIP holds more than one ${repeated.kind}`,
);
}
});
for (const part of parts) {
for (const piece of part.pending) expanded += piece.length;
}
if (expanded > 20 * zipBytes + 16 * 1024 * 1024) {
throw fail(
"live photo ZIP expands to more than 20 times its size plus 16 MiB",
);
}
for (const part of parts) {
for (const piece of part.pending)
await part.handle.write(piece);
part.pending = [];
}
};
const bytesWritten = await streamDecrypt(
stream,
header,
key,
async (plaintext) => {
for (let i = 0; i < plaintext.length; i += sliceSize) {
await push(plaintext.subarray(i, i + sliceSize), false);
}
},
onProgress,
);
await push(new Uint8Array(0), true);
if (image.ext === undefined || video.ext === undefined) {
throw fail(
"live photo ZIP does not hold both an image and a video",
);
}
if (file.metadata.hash !== undefined) {
checkHash(
file,
`${chunkHashFinal(image.hash)}:${chunkHashFinal(video.hash)}`,
);
}
if (image.ext.toLowerCase() === video.ext.toLowerCase()) {
throw fail(
`live photo's image and video have the same extension, ${video.ext}`,
);
}
const path = withExtension(destination, image.ext);
const videoPath = withExtension(destination, video.ext);
for (const part of parts) {
await part.handle.sync();
await part.handle.close();
}
await rm(destination, { force: true });
await rename(image.tmpPath, path);
try {
await rename(video.tmpPath, videoPath);
} catch (err) {
await rm(path, { force: true }).catch(() => undefined);
throw err;
}
await fsyncPath(dir);
return { path, bytesWritten, videoPath };
} catch (err) {
// Cancel the body so its connection is closed now rather than held
// until the stream is garbage collected. A backup run carries on past
// a failed file, so without this every failure would hold a socket.
// This covers every failure, including a temp file that cannot be
// opened and a header that is rejected before the body is read.
await stream.cancel(err).catch(() => undefined);
// Best-effort cleanup, as in `stageAtomic`.
for (const part of parts) {
await part.handle.close().catch(() => undefined);
await rm(part.tmpPath, { force: true }).catch(() => undefined);
}
throw err;
}
return bytesWritten;
};
// Fetch a stream and decrypt it to `destination`, retrying the whole sequence.
@@ -442,19 +534,44 @@ const fetchAndDecrypt = async (
destination: string,
onProgress?: ProgressCallback,
original?: EnteFile,
): Promise<number> =>
): Promise<DownloadResult> =>
withRetry(async () => {
const stream = await openStream();
return decryptToTemp(
destination,
stream,
header,
key,
onProgress,
original,
);
try {
if (original?.metadata.fileType === "livePhoto") {
return await decryptLivePhoto(
destination,
stream,
header,
key,
onProgress,
original,
);
}
const bytesWritten = await decryptToTemp(
destination,
stream,
header,
key,
onProgress,
original,
);
return { path: destination, bytesWritten };
} catch (err) {
// Cancel the body so its connection is closed now rather than held
// until the stream is garbage collected. A backup run carries on
// past a failed file, so without this every failure would hold a
// socket. This covers every failure, including a temp file that
// cannot be opened and a header that is rejected before the body
// is read.
await stream.cancel(err).catch(() => undefined);
throw err;
}
}, api.getRetryOptions());
// Write `file`'s original to `outPath`. A live photo is written as its image
// and its video beside `outPath` instead, and whatever was at `outPath` is
// removed (see `decryptLivePhoto`).
export const downloadFile = async (
api: ApiClient,
file: EnteFile,
@@ -466,7 +583,7 @@ export const downloadFile = async (
const resolvedPath =
outPath ?? sanitizeFileName(file.metadata.title, `file-${file.id}`);
const header = fromBase64(file.file.decryptionHeader);
const bytesWritten = await fetchAndDecrypt(
return fetchAndDecrypt(
api,
() => api.getFileStream(file.id, { retry: false }),
header,
@@ -475,7 +592,6 @@ export const downloadFile = async (
onProgress,
file,
);
return { path: resolvedPath, bytesWritten };
};
export const downloadThumbnail = async (
@@ -488,7 +604,7 @@ export const downloadThumbnail = async (
outPath ??
`thumb_${sanitizeFileName(file.metadata.title, `file-${file.id}`)}`;
const header = fromBase64(file.thumbnail.decryptionHeader);
const bytesWritten = await fetchAndDecrypt(
return fetchAndDecrypt(
api,
() => api.getThumbnailStream(file.id, { retry: false }),
header,
@@ -496,5 +612,4 @@ export const downloadThumbnail = async (
resolvedPath,
onProgress,
);
return { path: resolvedPath, bytesWritten };
};
+4
View File
@@ -35,3 +35,7 @@ export const safeExtension = (title: string): string => {
const ext = extname(title);
return /^\.[A-Za-z0-9]+$/.test(ext) ? ext : ".bin";
};
// `name` with its extension, if it has one, replaced by `ext` (".mov").
export const withExtension = (name: string, ext: string): string =>
name.slice(0, name.length - extname(name).length) + ext;
+265 -60
View File
@@ -1,11 +1,13 @@
// The on-disk content and thumbnail cache keyed by fileID (issue #46).
//
// Layout under `cacheDirectory`: `originals/<fileID>.<ext>` and
// `thumbnails/<fileID>.<ext>`, flat directories at 0700 with files at 0600.
// Content appears only by the streaming atomic writer's rename (the download
// layer, #40), so a file that exists is whole — "present means complete". The
// directory listing taken at `open()` is the record of what is cached, and the
// orphan temp files a crashed write may have left are reaped there.
// `thumbnails/<fileID>.<ext>`, flat directories at 0700 with files at 0600. A
// live photo's original is two files, its image and its video, with
// `originals/<fileID>.livephoto.json` naming them. Content appears only by the
// streaming atomic writer's rename (the download layer, #40), so a file that
// exists is whole — "present means complete". The directory listing taken at
// `open()` is the record of what is cached, and the orphan temp files a crashed
// write may have left are reaped there.
//
// A fetch goes through the shared request pools (#45): the content pool for
// originals, the thumbnail pool for thumbnails. The pool limits concurrency,
@@ -22,7 +24,14 @@
// does; thumbnails have none. On top of that this module refuses to record a
// stored file that came out empty.
import { existsSync, statSync } from "node:fs";
import {
closeSync,
existsSync,
openSync,
readFileSync,
readSync,
statSync,
} from "node:fs";
import {
chmod,
mkdir,
@@ -32,7 +41,7 @@ import {
statfs,
utimes,
} from "node:fs/promises";
import { dirname, extname, join } from "node:path";
import { basename, dirname, extname, join } from "node:path";
import type { ApiClient } from "../api/client.js";
import {
@@ -40,6 +49,7 @@ import {
downloadThumbnail,
type ProgressCallback,
removeLeftoverTempFiles,
writeAtomic,
} from "../download/index.js";
import { safeExtension } from "../filename.js";
import type { EnteFile } from "../model/types.js";
@@ -77,6 +87,8 @@ const poolPriorityOf = (priority: ThumbnailPriority): Priority =>
export interface ContentResult {
path: string;
bytes: number;
// A live photo's video. `path` and `bytes` are then its image's.
videoPath?: string;
}
// Progress for a single `original`/`thumbnail` call. A present file emits one
@@ -127,11 +139,14 @@ export interface ThumbnailsAPI {
// stand-in so the cache logic runs with no crypto and no network. Pool routing,
// dedup, present-checks and integrity live in the cache, not here.
export interface ContentSource {
// Writes the original at `destination`. A live photo is written beside it
// as its image and its video instead, and their paths are returned, as
// `downloadFile` does.
original(args: {
file: EnteFile;
destination: string;
onProgress?: ProgressCallback;
}): Promise<{ bytesWritten: number }>;
}): Promise<{ bytesWritten: number; path?: string; videoPath?: string }>;
thumbnail(args: {
file: EnteFile;
destination: string;
@@ -209,7 +224,9 @@ class AbortDrop extends Error {
}
}
const originalName = (file: EnteFile): string =>
// The name of a file's original in originals/: `<fileID><ext>`, the extension
// taken from the title (or `.bin`). A backup names its originals the same way.
export const nameInOriginals = (file: EnteFile): string =>
`${file.id}${safeExtension(file.metadata.title)}`;
// The fileID a cache filename encodes, or undefined when the name is not one
@@ -231,6 +248,92 @@ const fileSize = (path: string): number | undefined => {
}
};
// Whether `path` is a regular file with content. A zero-byte file is the shape
// an aborted write leaves, so it does not count.
const hasContent = (path: string | undefined): boolean =>
path !== undefined && (fileSize(path) ?? 0) > 0;
// Whether the file at `path` begins as a ZIP does, with `PK\x03\x04`. False
// when it cannot be read.
const isZip = (path: string): boolean => {
try {
const fd = openSync(path, "r");
try {
const head = Buffer.alloc(4);
return (
readSync(fd, head, 0, 4, 0) === 4 &&
head.toString("latin1") === "PK\x03\x04"
);
} finally {
closeSync(fd);
}
} catch {
return false;
}
};
// A live photo's image and video are named with the extensions from inside its
// ZIP, so their names alone do not say which is which. Wherever the cache or a
// backup stores one, a JSON file of this name beside them names both.
const livePhotoJSONName = (fileID: number): string =>
`${fileID}.livephoto.json`;
// The image and video that the live photo's JSON file in `dir` names, or
// undefined when there is none. Only names of the form the cache writes are
// taken, so the file cannot point outside `dir`.
const readLivePhotoJSON = (
dir: string,
fileID: number,
): { path: string; videoPath: string } | undefined => {
const valid = (name: unknown): name is string =>
typeof name === "string" && name === `${fileID}${safeExtension(name)}`;
try {
const { image, video } = JSON.parse(
readFileSync(join(dir, livePhotoJSONName(fileID)), "utf-8"),
);
if (valid(image) && valid(video)) {
return { path: join(dir, image), videoPath: join(dir, video) };
}
} catch {
// No such file, or not one the cache wrote.
}
return undefined;
};
// Write the JSON file naming a live photo's image and video, both in `dir`.
export const writeLivePhotoJSON = (
dir: string,
fileID: number,
stored: { path: string; videoPath: string },
): Promise<void> =>
writeAtomic(
join(dir, livePhotoJSONName(fileID)),
new TextEncoder().encode(
JSON.stringify({
image: basename(stored.path),
video: basename(stored.videoPath),
}),
),
);
// The original of `file` as the cache or a backup stored it in `dir`, when all
// of it is there: `<fileID><ext>`, or a live photo's image and video.
export const storedOriginal = (
dir: string,
file: EnteFile,
): { path: string; videoPath?: string } | undefined => {
if (file.metadata.fileType !== "livePhoto") {
const path = join(dir, nameInOriginals(file));
return hasContent(path) ? { path } : undefined;
}
const stored = readLivePhotoJSON(dir, file.id);
return stored !== undefined &&
hasContent(stored.path) &&
hasContent(stored.videoPath)
? stored
: undefined;
};
export class ContentCache implements PhotoContent, ThumbnailsAPI {
private readonly pools: RequestPools;
private readonly source: ContentSource;
@@ -238,10 +341,17 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
private readonly getFile: (fileID: number) => EnteFile | undefined;
private readonly originalsDir: string;
private readonly thumbnailsDir: string;
// fileID -> absolute path of the cached bytes, seeded from the directory
// listing at open() and extended as fetches store new files.
private readonly originals = new Map<number, string>();
private readonly thumbnails = new Map<number, string>();
// fileID -> absolute path of the cached bytes, and for a live photo's
// original its video's, seeded from the directory listing at open() and
// extended as fetches store new files.
private readonly originals = new Map<
number,
{ path: string; videoPath?: string }
>();
private readonly thumbnails = new Map<
number,
{ path: string; videoPath?: string }
>();
private readonly maxOriginalsBytes: number;
private readonly freeBelowBytes: number;
private readonly isPinned: (fileID: number) => boolean;
@@ -277,11 +387,15 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
// Prepare the cache directories, reap orphan temp files, and take the
// record of what is already cached. Called once before the cache serves.
async open(): Promise<void> {
// `isLivePhoto` says which files are live photos, whose original is only
// the image and video their JSON file names.
async open(
isLivePhoto: (fileID: number) => boolean = () => false,
): Promise<void> {
await this.ensureDir(this.originalsDir);
await this.ensureDir(this.thumbnailsDir);
await this.scan(this.originalsDir, this.originals);
await this.scan(this.thumbnailsDir, this.thumbnails);
await this.scan(this.originalsDir, this.originals, isLivePhoto);
await this.scan(this.thumbnailsDir, this.thumbnails, () => false);
// Publish the current usage and limit without evicting; a restart
// reuses whatever survived on disk. Eviction only ever fires on a write.
await this.refreshOriginalsLimit();
@@ -301,9 +415,9 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
pathsFor(fileID: number): CachedPaths {
const out: CachedPaths = {};
const original = this.originals.get(fileID);
if (original !== undefined) out.originalPath = original;
if (original !== undefined) out.originalPath = original.path;
const thumbnail = this.thumbnails.get(fileID);
if (thumbnail !== undefined) out.thumbnailPath = thumbnail;
if (thumbnail !== undefined) out.thumbnailPath = thumbnail.path;
return out;
}
@@ -335,7 +449,11 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
undefined,
{ destination },
);
return { path: result.path, bytes: result.bytes };
return {
path: result.path,
bytes: result.bytes,
videoPath: result.videoPath,
};
}
async ensure(args: EnsureOptions): Promise<EnsureResult[]> {
@@ -434,52 +552,72 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
? { status: "skipped", bytes: result.bytes }
: { status: "done", bytes: result.bytes },
);
return { path: result.path, bytes: result.bytes };
return {
path: result.path,
bytes: result.bytes,
videoPath: result.videoPath,
};
}
// The core: return the cached path if present, else fetch through the pool,
// store, and return it. `cached` distinguishes a present hit (no network,
// no download event) from a fresh fetch. A fetched original is stored at
// `opts.destination` when given, instead of in `originalsDir`.
// `opts.destination` when given, instead of in `originalsDir`. A live
// photo's original is present only with its video, and is returned with
// it.
private async acquire(
fileID: number,
kind: Kind,
priority: Priority,
signal: AbortSignal | undefined,
opts?: { onByte?: ProgressCallback; destination?: string },
): Promise<{ path: string; bytes: number; cached: boolean }> {
): Promise<{
path: string;
bytes: number;
videoPath?: string;
cached: boolean;
}> {
const file = this.getFile(fileID);
if (!file) throw new Error(`content cache: unknown file ${fileID}`);
const isLivePhoto =
kind === "original" && file.metadata.fileType === "livePhoto";
const known = kind === "original" ? this.originals : this.thumbnails;
const cached = known.get(fileID);
if (cached !== undefined) {
const size = fileSize(cached);
if (size !== undefined && size > 0) {
const size = fileSize(cached.path);
if (
size !== undefined &&
size > 0 &&
(!isLivePhoto || hasContent(cached.videoPath))
) {
// Returning an original's path is a use: bump its mtime so LRU
// order reflects it and survives a restart with no ledger.
if (
kind === "original" &&
dirname(cached) === this.originalsDir
dirname(cached.path) === this.originalsDir
)
await this.touch(cached);
return { path: cached, bytes: size, cached: true };
await this.touch(cached.path);
return { ...cached, bytes: size, cached: true };
}
// A recorded file that has since gone re-fetches below.
// A recorded file that has since gone, or a live photo an earlier
// version stored as one ZIP, re-fetches below.
known.delete(fileID);
}
// An original a backup already stored counts as present.
if (kind === "original" && this.downloadDirectory !== undefined) {
const backupPath = join(
this.downloadDirectory,
"originals",
originalName(file),
const stored = storedOriginal(
join(this.downloadDirectory, "originals"),
file,
);
const size = fileSize(backupPath);
if (size !== undefined && size > 0) {
this.originals.set(fileID, backupPath);
return { path: backupPath, bytes: size, cached: true };
if (stored !== undefined) {
this.originals.set(fileID, stored);
return {
...stored,
bytes: fileSize(stored.path) ?? 0,
cached: true,
};
}
}
@@ -487,7 +625,7 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
kind === "original" ? this.originalsDir : this.thumbnailsDir;
const dest =
kind === "original"
? (opts?.destination ?? join(dir, originalName(file)))
? (opts?.destination ?? join(dir, nameInOriginals(file)))
: join(dir, `${fileID}${THUMBNAIL_EXT}`);
const pool =
kind === "original" ? this.pools.content : this.pools.thumbnails;
@@ -508,21 +646,39 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
? this.beginOriginalWrite(fileID)
: null;
try {
await this.download(file, dest, kind, opts?.onByte);
await chmod(dest, FILE_MODE);
const size = (await stat(dest)).size;
if (size === 0) {
throw new Error(
`content cache: ${kind} ${fileID} stored empty`,
);
const stored = await this.download(
file,
dest,
kind,
opts?.onByte,
);
for (const path of [stored.path, stored.videoPath]) {
if (path === undefined) continue;
await chmod(path, FILE_MODE);
if ((await stat(path)).size === 0) {
throw new Error(
`content cache: ${kind} ${fileID} stored empty`,
);
}
}
known.set(fileID, dest);
// A backup records its own live photos.
if (
stored.videoPath !== undefined &&
opts?.destination === undefined
) {
await writeLivePhotoJSON(dir, fileID, {
path: stored.path,
videoPath: stored.videoPath,
});
}
known.set(fileID, stored);
// A fresh original may have crossed the limit; make room by
// evicting least-recently-used originals. An over-budget
// fetch keeps the file it returns, and no overlapping
// sibling is evicted. Thumbnails are never bounded.
if (write) await this.enforceOriginalsLimit(write);
return { path: dest, bytes: size, cached: false };
const size = (await stat(stored.path)).size;
return { ...stored, bytes: size, cached: false };
} finally {
if (write) this.inFlightOriginals.delete(write);
}
@@ -531,18 +687,24 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
);
}
// Fetch into `destination`, returning where the bytes landed: there, or
// for a live photo, its image and video beside it.
private async download(
file: EnteFile,
destination: string,
kind: Kind,
onProgress: ProgressCallback | undefined,
): Promise<number> {
): Promise<{ path: string; videoPath?: string }> {
const args = { file, destination, onProgress };
const result =
kind === "original"
? await this.source.original(args)
: await this.source.thumbnail(args);
return result.bytesWritten;
if (kind === "thumbnail") {
await this.source.thumbnail(args);
return { path: destination };
}
const result = await this.source.original(args);
return {
path: result.path ?? destination,
videoPath: result.videoPath,
};
}
// Best-effort bump of a file's mtime to now; a failed touch must never fail
@@ -553,13 +715,14 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
}
// Every stored original that lives under `originalsDir` (a backup-directory
// hit recorded in the map is excluded), with its size and mtime. Entries
// whose file has vanished are dropped from the map. Backups and thumbnails
// are never counted.
// hit recorded in the map is excluded), with its size and mtime; a live
// photo's size includes its video. Entries whose file has vanished are
// dropped from the map. Backups and thumbnails are never counted.
private async measureOriginals(): Promise<{
entries: {
fileID: number;
path: string;
videoPath?: string;
size: number;
mtimeMs: number;
}[];
@@ -568,21 +731,28 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
const entries: {
fileID: number;
path: string;
videoPath?: string;
size: number;
mtimeMs: number;
}[] = [];
let used = 0;
for (const [fileID, path] of this.originals) {
for (const [fileID, { path, videoPath }] of this.originals) {
if (dirname(path) !== this.originalsDir) continue;
try {
const s = await stat(path);
const size =
s.size +
(videoPath === undefined
? 0
: (await stat(videoPath)).size);
entries.push({
fileID,
path,
size: s.size,
videoPath,
size,
mtimeMs: s.mtimeMs,
});
used += s.size;
used += size;
} catch {
this.originals.delete(fileID);
}
@@ -646,6 +816,18 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
for (const e of evictable) {
if (remaining <= limit) break;
await rm(e.path, { force: true });
// A live photo goes whole: its video and the JSON file
// naming the two go with its image.
if (e.videoPath !== undefined) {
await rm(e.videoPath, { force: true });
await rm(
join(
this.originalsDir,
livePhotoJSONName(e.fileID),
),
{ force: true },
);
}
this.originals.delete(e.fileID);
remaining -= e.size;
}
@@ -670,7 +852,11 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
await chmod(dir, DIR_MODE);
}
private async scan(dir: string, into: Map<number, string>): Promise<void> {
private async scan(
dir: string,
into: Map<number, { path: string; videoPath?: string }>,
isLivePhoto: (fileID: number) => boolean,
): Promise<void> {
// Another process sharing this cache may still be writing its temp
// files, so only those whose process has exited are removed.
removeLeftoverTempFiles(dir);
@@ -680,10 +866,29 @@ export class ContentCache implements PhotoContent, ThumbnailsAPI {
} catch {
return;
}
const names = new Set(entries);
for (const name of entries) {
const id = fileIDFromName(name);
const path = join(dir, name);
if (id !== undefined && existsSync(path)) into.set(id, path);
if (id === undefined || !existsSync(path)) continue;
// A live photo's image and video are one entry, as the JSON file
// beside them names them. A live photo's file with no such JSON
// file is not its original. If it is a ZIP, it is the one an
// earlier version stored under the image's name, and is removed.
// Any other is left alone: another process may have just stored
// it and not yet written the JSON file.
const livePhoto = names.has(livePhotoJSONName(id))
? readLivePhotoJSON(dir, id)
: undefined;
if (livePhoto !== undefined) {
into.set(id, livePhoto);
} else if (isLivePhoto(id)) {
if (isZip(path)) {
await rm(path, { force: true }).catch(() => undefined);
}
} else {
into.set(id, { path });
}
}
}
}
+8 -2
View File
@@ -397,7 +397,8 @@ export class Library {
originalsDays: opts.precacheOriginalsDays,
onEvent: opts.onProgress,
});
precache.update(deriveRecordsFromStore(store));
const records = deriveRecordsFromStore(store);
precache.update(records);
const extraPinned = opts.isOriginalPinned;
cache = new ContentCache({
pools,
@@ -411,7 +412,12 @@ export class Library {
precache!.isPinned(fileID) ||
(extraPinned?.(fileID) ?? false),
});
await cache.open();
// The records say at once which files are live photos, where
// `getFileByID` would search every file for each one cached.
await cache.open(
(fileID) =>
records.photos.get(fileID)?.fileType === "livePhoto",
);
precache.bind(cache);
}
+2 -1
View File
@@ -80,7 +80,8 @@ export class Photo {
}
// Fetch and cache the full-resolution original, returning its on-disk path
// and byte length. Served from the cache (or the backup download directory)
// and byte length; for a live photo, its image's, and its video's path as
// `videoPath`. Served from the cache (or the backup download directory)
// when already present, otherwise fetched through the content pool.
async original(opts?: ContentOptions): Promise<ContentResult> {
return this.contentOrThrow().original(this.rec.fileID, opts);
+1
View File
@@ -42,6 +42,7 @@ export interface PhotoRecord {
isArchived: boolean;
isHidden: boolean;
// Local cache paths, set once a later phase caches the bytes; unset here.
// A live photo's `originalPath` is its image.
thumbnailPath?: string;
originalPath?: string;
}
+2 -1
View File
@@ -128,7 +128,8 @@ export const extractImageMetadata = (
// Read a file's original bytes through the library's content cache and extract
// its embedded image metadata. The bytes come from `photo.original()` — the
// same on-disk cache the rest of the library fills — rather than a fresh
// per-call download to a throwaway temp file.
// per-call download to a throwaway temp file. For a live photo, its `path` is
// the image.
const extractExif = async (
photo: Photo,
): Promise<Record<string, unknown> | undefined> => {