quak backup writes each file's ML data into its JSON (closes #163)
check / check (push) Failing after 3s

lib.backup() now waits for an ML data fetch, joining the one its refresh
started or starting one, before it writes the per-file JSON, and each
file's JSON carries the cached payload as mlData. When the fetch fails,
each file with no cached ML data gets mlDataError and an entry in
failures.json, so the result counts it as failed and quak backup exits 1;
the next run fetches again.

Judgement call: the wait comes after the originals are downloaded, so the fetch runs alongside the downloads.
Judgement call: a failed ML fetch is recorded per file in failures.json, which is how the exit code goes non-zero without changing src/cli-commands.ts.

Model: opus-5-5
This commit is contained in:
2026-10-05 23:26:16 +00:00
parent 14ac7d05ef
commit 2326281860
6 changed files with 208 additions and 31 deletions
+17 -9
View File
@@ -599,8 +599,8 @@ the smallest does not.
per unique file, two for a live photo: see
below)
YYYY-MM-DD.<fileID>.json the file's basic metadata fields quak
keeps, and its private and public magic
metadata
keeps, its private and public magic
metadata, and its ML data
YYYY-MM-DD.<fileID>.livephoto.json
which of a live photo's two files is which
collections/
@@ -619,6 +619,14 @@ name as uploaded, case kept, or `.bin` when it has none or it holds anything but
letters and digits. When the date or the time zone changes, the next run saves
the original at its new path and leaves the old copy where it is.
A file's JSON holds Ente's ML data for it (its faces and its CLIP embedding) as
`mlData`, the same payload `backup-metadata` writes; a file Ente has no ML data
for has no `mlData`. The backup waits for the library's ML data fetch to finish
before it writes the JSON files. If that fetch fails, each file whose ML data is
not in the cache gets the reason in `mlDataError` instead and counts as failed,
and the next run fetches it again. The JSON files are rewritten on every run, so
ML data that arrived since the last run appears.
`failures.json` records each failed file with the kind of failure, how many
times it has been tried and when it was last tried. A file leaves it once it
succeeds, or once it is no longer in the library or in the backup's scope. The
@@ -655,8 +663,7 @@ subsequent runs, existing originals are skipped. If a download fails, the error
is logged and the backup continues with the next file. The exit code is non-zero
if any files failed. `quak backup` opens its library with the thumbnail and
originals precache off, so the only file content it fetches is the originals the
backup stores. The library's ML data fetch still runs and fills the cache's
`mldata/`.
backup stores.
Each original is written to a temporary file in the same directory, synced to
disk, and renamed into place, so an original is either complete or absent, even
@@ -876,11 +883,12 @@ photos newest first). `lib.subscribe({ onChange })` delivers a `LibraryChange`
- `await lib.backup(opts?)` → `BackupResult`. It waits for a refresh as
`fresh()` does, puts every in-scope original not already at its save path
there as `photo.download()` does (and, with `includeThumbnails`, fetches
thumbnails) through the content cache, and rebuilds the on-disk backup tree
with a durable failure ledger. A fetched original is written straight to its
save path and not into the cache, which then counts it as present; one the
cache already held is copied from there. `BackupOptions`: `downloadDirectory`
(falls back to the library's), `includeOriginals` (default `true`),
thumbnails) through the content cache, waits for an ML data fetch, and
rebuilds the on-disk backup tree, each file's JSON with its ML data, with a
durable failure ledger. A fetched original is written straight to its save
path and not into the cache, which then counts it as present; one the cache
already held is copied from there. `BackupOptions`: `downloadDirectory` (falls
back to the library's), `includeOriginals` (default `true`),
`includeThumbnails` (default `false`), `onlyAlbumNames`, and `onProgress`. See
Backup layout above for the tree it writes.
+7
View File
@@ -25,6 +25,13 @@ declares one.
# Completed Steps
- 2026-10-05: `quak backup` writes each file's ML data (its faces and CLIP
embedding) into the file's JSON as `mlData`, the payload
`lib.mldata.forFile()` returns (issue 163). `lib.backup()` waits for an ML
data fetch before it writes the JSON files. When that fetch fails, each file
whose ML data is not cached gets the reason in `mlDataError` and counts as
failed, and the next run fetches it again.
- 2026-10-03: The download-albums example test no longer fails on a disk with
under 50 GiB free (issue 160). It opens its libraries with
`freeBelowBytes: 0`, so the free space of the disk it runs on cannot shrink
+54 -12
View File
@@ -3,14 +3,15 @@
// `lib.backup()` waits for a completed refresh of the library (a failed one
// fails the backup before any file is touched), then, for every file in scope,
// puts its original at its save path under `downloadDirectory`, as
// `Photo.download()` does, and rebuilds the derived views (per-file sidecars,
// per-collection symlink trees, per-collection JSON) from the model. The
// on-disk layout:
// `Photo.download()` does, waits for an ML data fetch, and rebuilds the derived
// views (per-file sidecars, per-collection symlink trees, per-collection JSON)
// from the model. The on-disk layout:
//
// <downloadDirectory>/
// YYYY/YYYY-MM/YYYY-MM-DD/
// YYYY-MM-DD.<fileID>.<ext> the decrypted bytes (the save path)
// YYYY-MM-DD.<fileID>.json per-file metadata sidecar
// YYYY-MM-DD.<fileID>.json per-file metadata sidecar, with
// the file's ML data
// collections/<name>/<title> symlink to the original
// collections/<name>.json per-collection metadata
// failures.json durable ledger of unresolved failures
@@ -29,9 +30,9 @@
// directories of albums that no longer exist.
//
// Resilience (issue #8): no per-file condition aborts the run. A failed
// download or a failed symlink is caught, recorded in `failures.json` with a
// classification, a running attempt count, and the last-tried time, and the run
// continues. `result.failed` — and thus the CLI's exit code — stays non-zero
// download, a failed symlink, or ML data missing because the ML data fetch
// failed is caught, recorded in `failures.json` with a classification, a
// running attempt count, and the last-tried time, and the run continues. `result.failed` — and thus the CLI's exit code — stays non-zero
// while any failure remains unresolved and clears once every one succeeds. Each
// run reconciles the ledger against the files it attempted, so an entry for a
// file that has since left the library (deleted) or this run's scope is dropped
@@ -60,6 +61,7 @@ import {
storedAtSavePath,
} from "./library/content.js";
import { representative } from "./library/records.js";
import type { MLData } from "./mldata-fetch.js";
import type { Collection, EnteFile } from "./model/types.js";
export type ProgressCallback = (message: string) => void;
@@ -117,6 +119,13 @@ export interface BackupLibrary {
destination: string,
): Promise<{ path: string; videoPath?: string }>;
thumbnail(fileID: number): Promise<{ path: string }>;
// Wait for an ML data fetch to complete, joining one already running or
// starting one. Rejects with the reason when it fails; resolves at once
// when the library cannot fetch ML data.
fetchMLData(): Promise<void>;
// A file's cached ML data, as `lib.mldata.forFile()` returns it, or
// undefined when none is cached.
mlData(fileID: number): Promise<MLData | undefined>;
}
type FailureClass = "transient" | "permanent" | "unknown";
@@ -340,7 +349,13 @@ const saveLedger = (path: string, ledger: Map<number, FailureEntry>): void => {
);
};
const writeSidecar = (path: string, file: EnteFile): void => {
// The file's JSON: its basic fields, its magic metadata, and its ML data, or
// the reason the ML data is missing.
const writeSidecar = (
path: string,
file: EnteFile,
ml: { mlData?: MLData; mlDataError?: string },
): void => {
const meta: Record<string, unknown> = {
id: file.id,
collectionID: file.collectionID,
@@ -349,6 +364,8 @@ const writeSidecar = (path: string, file: EnteFile): void => {
};
if (file.magicMetadata) meta.magicMetadata = file.magicMetadata;
if (file.pubMagicMetadata) meta.pubMagicMetadata = file.pubMagicMetadata;
if (ml.mlData) meta.mlData = ml.mlData;
if (ml.mlDataError) meta.mlDataError = ml.mlDataError;
writeFileSync(path, JSON.stringify(meta, null, 2));
};
@@ -497,12 +514,37 @@ export const runBackup = async (
}
// Phase 2: rebuild the derived views from the model. Sidecars first, for
// every present original (this repairs stale ones).
// every present original (this repairs stale ones), each with the file's
// ML data once an ML data fetch has completed. When the fetch fails, a
// file with no cached ML data gets the reason instead and is recorded as
// failed, so the next run fetches it again.
if (includeOriginals) {
let mlDataError: string | undefined;
try {
log("Fetching ML data...");
await lib.fetchMLData();
} catch (err) {
mlDataError = errorMessage(err);
log(`FAILED ML data: ${mlDataError}`);
}
for (const file of distinct.values()) {
if (storedAtSavePath(downloadDirectory, file) !== undefined) {
const path = savePath(downloadDirectory, file);
writeSidecar(withExtension(path, ".json"), file);
if (storedAtSavePath(downloadDirectory, file) === undefined) {
continue;
}
const path = withExtension(
savePath(downloadDirectory, file),
".json",
);
const mlData = await lib.mlData(file.id);
if (mlData === undefined && mlDataError !== undefined) {
writeSidecar(path, file, { mlDataError });
recordFailure(
file,
collectionName.get(file.collectionID) ?? "",
new Error(`ML data: ${mlDataError}`),
);
} else {
writeSidecar(path, file, { mlData });
}
}
}
+23 -10
View File
@@ -279,7 +279,8 @@ export class Library {
private cycle?: Promise<void>;
// Guards the ML fetch pass so a slow backfill never runs twice at once; a
// refresh whose pass is still running kicks nothing new. Holds the running
// pass, so `close()` can wait for it.
// pass, so `close()` and `backup()` can wait for it. It rejects when the
// pass fails.
private mlFetch?: Promise<void>;
private closed = false;
private lastRefreshAt?: number;
@@ -569,8 +570,9 @@ export class Library {
// does, joining one already running, and rejects before touching any file
// when it fails. Then puts pending originals at their save paths as
// `Photo.download()` does (and optional thumbnails) through the content
// cache and pools, and rebuilds the derived symlink/JSON views from the
// model. Throws before any network work when no content cache backs the
// cache and pools, waits for an ML data fetch, and rebuilds the derived
// symlink/JSON views from the model, each file's JSON with its ML data.
// Throws before any network work when no content cache backs the
// originals it must fetch.
backup(opts?: BackupOptions): Promise<BackupResult> {
const downloadDirectory =
@@ -593,6 +595,8 @@ export class Library {
original: (fileID, destination) =>
cache!.backupOriginal(fileID, destination),
thumbnail: (fileID) => cache!.thumbnail(fileID),
fetchMLData: () => this.fetchMLDataNow(),
mlData: (fileID) => this.mldata.forFile({ fileID }),
},
{ ...opts, downloadDirectory },
);
@@ -612,7 +616,7 @@ export class Library {
this.timer = undefined;
}
await this.cycle?.catch(() => {});
await this.mlFetch;
await this.mlFetch?.catch(() => {});
await precacheClosed;
}
@@ -676,10 +680,9 @@ export class Library {
// Backfill ML data for the files this refresh knows about. It runs
// outside the refresh's success/failure so a fetch or disk problem
// there never marks the metadata refresh failed, and it is not
// awaited so it never stalls the refresh interval.
this.mlFetch ??= this.runMLFetch().finally(() => {
this.mlFetch = undefined;
});
// awaited so it never stalls the refresh interval. Its failure is
// reported through `status()` and `onProgress`.
void this.fetchMLDataNow().catch(() => {});
} catch (err) {
const error = err instanceof Error ? err.message : String(err);
this.lastError = error;
@@ -784,10 +787,19 @@ export class Library {
}
}
// Join the running ML fetch pass, or start one when none runs. Resolves at
// once when the client cannot fetch ML data; rejects when the pass fails.
private fetchMLDataNow(): Promise<void> {
this.mlFetch ??= this.runMLFetch().finally(() => {
this.mlFetch = undefined;
});
return this.mlFetch;
}
// One ML fetch pass: fetch, decrypt and store the ML data for every file
// the store knows about that is not cached (or whose `updationTime` has
// advanced), through the metadata pool, and update the CLIP index. Guarded
// so passes never overlap; a failure is reported, not thrown.
// advanced), through the metadata pool, and update the CLIP index. A
// failure is reported through `status()` and `onProgress`, then thrown.
private async runMLFetch(): Promise<void> {
const mldata = this.mlStore;
// Bind so the call keeps the client as its receiver when invoked
@@ -831,6 +843,7 @@ export class Library {
const error = err instanceof Error ? err.message : String(err);
this.lastMLError = error;
this.emit({ operation: "fetchMLData", status: "failed", error });
throw err;
}
}
+91
View File
@@ -52,6 +52,7 @@ import { runBackup, type BackupLibrary } from "../../src/backup.js";
import { Library } from "../../src/library/index.js";
import type { ContentSource } from "../../src/library/content.js";
import type { CollectionsPage, FilesPage } from "../../src/client.js";
import type { MLData } from "../../src/mldata-fetch.js";
import type { Collection, EnteFile } from "../../src/model/types.js";
import {
asLivePhoto,
@@ -828,6 +829,94 @@ describe("the refresh before a backup", () => {
});
});
// Also serves ML data: file 100 has `ML_PAYLOAD`, the other two have none.
// While `mlError` is set, every ML data request fails with it.
const ML_PAYLOAD = { face: { faces: [] }, clip: { embedding: [0.5, 0.25] } };
class MLClient extends MockClient {
mlError?: string;
async fetchMLData(args: {
fileIDs: number[];
}): Promise<Map<number, MLData>> {
if (this.mlError) throw new Error(this.mlError);
const result = new Map<number, MLData>();
if (args.fileIDs.includes(100)) result.set(100, ML_PAYLOAD);
return result;
}
}
// The JSON the backup in `outDir` wrote beside the original of `fileID`.
const fileJSON = (outDir: string, fileID: number): Record<string, unknown> =>
JSON.parse(readFileSync(saved(outDir, `${fileID}.json`), "utf-8"));
describe("ML data in each file's JSON", () => {
it("writes a file's ML data, and no ML field for a file that has none", async () => {
const lib = await openLibrary(stubSource(), new MLClient());
const outDir = join(root, "backup");
const result = await lib.backup({ downloadDirectory: outDir });
expect(result.failed).toBe(0);
expect(fileJSON(outDir, 100).mlData).toEqual(ML_PAYLOAD);
expect(fileJSON(outDir, 101)).not.toHaveProperty("mlData");
expect(fileJSON(outDir, 101)).not.toHaveProperty("mlDataError");
await lib.close();
});
it("gives each file with no cached ML data the reason when the fetch fails, and fails the run", async () => {
const client = new MLClient();
const outDir = join(root, "backup");
const first = await openLibrary(stubSource(), client);
await first.backup({ downloadDirectory: outDir });
await first.close();
client.mlError = "HTTP 503 from server";
const lib = await openLibrary(stubSource(), client);
const result = await lib.backup({ downloadDirectory: outDir });
// File 100's ML data was cached by the first run and is kept.
expect(fileJSON(outDir, 100).mlData).toEqual(ML_PAYLOAD);
expect(fileJSON(outDir, 100)).not.toHaveProperty("mlDataError");
for (const fileID of [101, 200]) {
expect(fileJSON(outDir, fileID).mlDataError).toBe(
"HTTP 503 from server",
);
}
expect(result.failed).toBe(2);
expect(result.errors.map((e) => [e.fileID, e.error])).toEqual([
[101, "ML data: HTTP 503 from server"],
[200, "ML data: HTTP 503 from server"],
]);
expect(Object.keys(readLedger(outDir).files)).toEqual(["101", "200"]);
await lib.close();
});
it("writes the ML data on the run after a failed fetch", async () => {
const client = new MLClient();
client.mlError = "HTTP 503 from server";
const outDir = join(root, "backup");
const first = await openLibrary(stubSource(), client);
expect((await first.backup({ downloadDirectory: outDir })).failed).toBe(
3,
);
expect(fileJSON(outDir, 100)).not.toHaveProperty("mlData");
await first.close();
client.mlError = undefined;
const lib = await openLibrary(stubSource(), client);
const result = await lib.backup({ downloadDirectory: outDir });
expect(result.failed).toBe(0);
expect(result.skipped).toBe(3);
expect(fileJSON(outDir, 100).mlData).toEqual(ML_PAYLOAD);
for (const fileID of [100, 101, 200]) {
expect(fileJSON(outDir, fileID)).not.toHaveProperty("mlDataError");
}
expect(existsSync(join(outDir, "failures.json"))).toBe(false);
await lib.close();
});
});
// Every entry under collections/, one level of directories deep, with each
// symlink's target.
const tree = (outDir: string): string[] => {
@@ -871,6 +960,8 @@ describe("backup album folders", () => {
thumbnail: async () => {
throw new Error("no thumbnails in this stand-in");
},
fetchMLData: async () => {},
mlData: async () => undefined,
});
const albumID = (outDir: string, jsonName: string): number =>
+16
View File
@@ -732,6 +732,22 @@ describe("backup", () => {
expect(stderr.text).toBe("Starting backup...\n");
});
it("exits 1 and lists each file when the ML data fetch fails", async () => {
const client = {
...fakeClient(),
fetchMLData: async () => {
throw new Error("HTTP 503 from server");
},
} as unknown as Client;
expect(
await backupCommand(context(client), join(root, "backup"), {}),
).toBe(1);
expect(stderr.text).toContain(" Failed: 3\n");
expect(stderr.text).toContain(
" [Vacation] beach.jpg (id 100): ML data: HTTP 503 from server\n",
);
});
it("exits 1 with the error on one line when the refresh fails", async () => {
const client = {
...fakeClient(),