Compare commits

..
3 Commits
Author SHA1 Message Date
sneak c6ab2de940 Give backup album folders distinct names and remove stale entries (closes #103)
check / check (push) Successful in 1m59s
Within a collection's folder, files whose sanitized titles match (ignoring
case) each get their file ID added before the extension, and collections
whose sanitized names match get their ID added, so no symlink or JSON
replaces another. Names are chosen across all collections, so a scoped run
names folders the same as a full one.

Each run first removes symlinks into originals/ that no longer belong to a
collection, and the folders quak wrote (a sibling JSON with an album ID) for
collections that are gone or renamed. Anything else is left alone; a folder
still holding user files keeps its JSON.

Model: opus-5-5
2026-09-23 03:56:46 +00:00
clawbot 390401af2c Start empty when the cache directory holds another account's cache (closes #104)
check / check (push) Successful in 35s
When metadata.json in the cache directory was written for a different,
non-zero user ID than the client's, Library.open deletes it and mldata/
and loads an empty store, so the first refresh enumerates from 0 and none
of the other account's collections, files, keys or ML results are served.
Only reachable with --cache-dir or an explicit cacheDirectory.

Model: opus-5-5
2026-09-23 05:56:14 +02:00
clawbot bf3b20df2f backup-metadata: keep going when an ML data request fails (closes #101)
check / check (push) Successful in 27s
Each ML data request of up to 200 files is now tried on its own. A request
that still fails after its retries is logged, its files are written with the
reason in `mlDataError`, and the command exits 1 once the dump is complete.
`fetchMLData`, used only here, is removed in favour of a per-batch loop over
`fetchMLDataBatch`.

Model: opus-5-5
2026-09-23 05:49:38 +02:00
9 changed files with 256 additions and 40 deletions
+10
View File
@@ -471,6 +471,11 @@ accepted for backward compatibility but ignored. `backup-metadata --exif` (alias
metadata. The listing and backup commands support `--json` for machine-readable metadata. The listing and backup commands support `--json` for machine-readable
output. output.
`backup-metadata` fetches ML data in requests of up to 200 files. When a request
still fails after its retries, the error is logged, each of its files is written
with the reason in an `mlDataError` field instead of `mlData`, and the dump goes
on. The exit code is non-zero if any ML data request failed.
`helper fix-missing-thumbnails` regenerates thumbnails for baseline JPEG images `helper fix-missing-thumbnails` regenerates thumbnails for baseline JPEG images
only, because the bundled decoder (`jpeg-js`) decodes only JPEG. A non-JPEG only, because the bundled decoder (`jpeg-js`) decodes only JPEG. A non-JPEG
image (PNG, HEIC) or a video is reported as `skipped` (unsupported format), kept image (PNG, HEIC) or a video is reported as `skipped` (unsupported format), kept
@@ -690,6 +695,11 @@ Under `cacheDirectory`:
fetched.json per-file fetch bookkeeping fetched.json per-file fetch bookkeeping
``` ```
When `metadata.json` belongs to a different account than the client's,
`Library.open` deletes it and `mldata/` and starts from an empty cache. Cached
originals and thumbnails are kept; they are reached only through the files the
current account's records name.
A stored file appears only via an atomic temp-then-rename, so its presence means A stored file appears only via an atomic temp-then-rename, so its presence means
it is complete. The design also calls for a content-hash comparison against it is complete. The design also calls for a content-hash comparison against
`FileMetadata.hash` on each fetched original; that check is deferred (issue `FileMetadata.hash` on each fetched original; that check is deferred (issue
+11
View File
@@ -25,6 +25,17 @@ Tag v1.0.0.
`originals/` for files no longer in the collection, and the folders of deleted `originals/` for files no longer in the collection, and the folders of deleted
or renamed collections, leaving anything else in `collections/` alone. The or renamed collections, leaving anything else in `collections/` alone. The
README backup layout states the naming rule. README backup layout states the naming rule.
- 2026-09-23: Kept one account's cache from mixing with another's (issue 104).
When `metadata.json` in the cache directory was written for a different,
non-zero user ID than the client's, `Library.open` deletes it and `mldata/`
and starts empty, so the first refresh enumerates from 0. This only happens
with `--cache-dir` or an explicit `cacheDirectory`; the default path already
includes the user ID. A test opens one account's cache as another account.
- 2026-09-23: `backup-metadata` no longer stops on one failed ML data request
(issue 101). Each request of up to 200 files is tried on its own; a failed one
is logged, its files are written with the reason in `mlDataError`, and the
command exits 1 once the dump is complete. `fetchMLData`, which only this
command used, is gone; the command calls `fetchMLDataBatch` per batch.
- 2026-09-23: Single-sourced the version string (issue 5). `package.json` is the - 2026-09-23: Single-sourced the version string (issue 5). `package.json` is the
only place it is written: `src/index.ts` imports it for `VERSION` and only place it is written: `src/index.ts` imports it for `VERSION` and
`bin/quak.ts` passes `VERSION` to commander. tsc copies `package.json` to `bin/quak.ts` passes `VERSION` to commander. tsc copies `package.json` to
+2 -2
View File
@@ -323,11 +323,11 @@ export const backupMetadataCommand = async (
if (!client) return 1; if (!client) return 1;
const lib = await openReadLibrary(ctx, client); const lib = await openReadLibrary(ctx, client);
try { try {
await runMetadataBackup(lib, client, dir, { const { failedMLBatches } = await runMetadataBackup(lib, client, dir, {
exif: opts.exif || opts.all, exif: opts.exif || opts.all,
onProgress: (msg) => ctx.stderr.write(msg + "\n"), onProgress: (msg) => ctx.stderr.write(msg + "\n"),
}); });
return 0; return failedMLBatches > 0 ? 1 : 0;
} finally { } finally {
await lib.close(); await lib.close();
} }
+15 -3
View File
@@ -26,6 +26,7 @@
// store marked unsaved until a later save actually lands, so a stuck disk is // store marked unsaved until a later save actually lands, so a stuck disk is
// never masked by a subsequent empty refresh. // never masked by a subsequent empty refresh.
import { rm } from "node:fs/promises";
import { join } from "node:path"; import { join } from "node:path";
import envPaths from "env-paths"; import envPaths from "env-paths";
@@ -345,9 +346,20 @@ export class Library {
const cacheDirectory = const cacheDirectory =
opts.cacheDirectory ?? opts.cacheDirectory ??
join(envPaths("quak", { suffix: "" }).cache, String(userID)); join(envPaths("quak", { suffix: "" }).cache, String(userID));
const store = await MetadataStore.load( const metadataPath = join(cacheDirectory, "metadata.json");
join(cacheDirectory, "metadata.json"), let store = await MetadataStore.load(metadataPath);
); // A cache directory given explicitly can hold another account's cache.
// Its records and cursor are not this account's, so delete it and the
// ML data beside it and start empty. A user ID of 0 means the cache
// was never refreshed and so holds nothing to discard.
if (store.userID !== 0 && store.userID !== userID) {
await rm(metadataPath, { force: true });
await rm(join(cacheDirectory, "mldata"), {
recursive: true,
force: true,
});
store = await MetadataStore.load(metadataPath);
}
const intervalMs = const intervalMs =
(opts.refreshIntervalSeconds ?? DEFAULT_REFRESH_INTERVAL_SECONDS) * (opts.refreshIntervalSeconds ?? DEFAULT_REFRESH_INTERVAL_SECONDS) *
1000; 1000;
+32 -5
View File
@@ -5,7 +5,11 @@ import exifReader from "exif-reader";
import type { Client } from "./client.js"; import type { Client } from "./client.js";
import type { Library, Photo } from "./library/index.js"; import type { Library, Photo } from "./library/index.js";
import { sanitizeFileName } from "./filename.js"; import { sanitizeFileName } from "./filename.js";
import { fetchMLData } from "./mldata-fetch.js"; import {
fetchMLDataBatch,
MLDATA_BATCH_SIZE,
type MLData,
} from "./mldata-fetch.js";
import type { EnteFile } from "./model/types.js"; import type { EnteFile } from "./model/types.js";
export type ProgressCallback = (message: string) => void; export type ProgressCallback = (message: string) => void;
@@ -137,13 +141,14 @@ const extractExif = async (
// of plain JSON: account, per-collection, and per-file records including the // of plain JSON: account, per-collection, and per-file records including the
// private and public magic metadata and (by default) the ML data. Collections // private and public magic metadata and (by default) the ML data. Collections
// and files are enumerated from the library's cache rather than a fresh server // and files are enumerated from the library's cache rather than a fresh server
// scan; the ML fetch and EXIF extraction are unchanged. // scan. Returns how many ML data requests failed; their files are still
// written, with `mlDataError` in place of `mlData`.
export const runMetadataBackup = async ( export const runMetadataBackup = async (
lib: Library, lib: Library,
client: Client, client: Client,
outDir: string, outDir: string,
opts?: MetadataBackupOptions, opts?: MetadataBackupOptions,
): Promise<void> => { ): Promise<{ failedMLBatches: number }> => {
const log = opts?.onProgress ?? (() => {}); const log = opts?.onProgress ?? (() => {});
const wantExif = opts?.exif ?? false; const wantExif = opts?.exif ?? false;
@@ -208,12 +213,31 @@ export const runMetadataBackup = async (
} }
} }
// One failed request (retries exhausted) must not end the dump: its files
// get the reason in `mlDataError` and the other batches go on.
log("Fetching ML data (face detections, CLIP embeddings)..."); log("Fetching ML data (face detections, CLIP embeddings)...");
const mlDataMap = await fetchMLData( const mlDataMap = new Map<number, MLData>();
const mlDataErrors = new Map<number, string>();
let failedMLBatches = 0;
const fileIDs = [...fileKeys.keys()];
for (let i = 0; i < fileIDs.length; i += MLDATA_BATCH_SIZE) {
const batch = fileIDs.slice(i, i + MLDATA_BATCH_SIZE);
try {
const result = await fetchMLDataBatch(
client.getApiClient(), client.getApiClient(),
[...fileKeys.keys()], batch,
fileKeys, fileKeys,
); );
for (const [id, payload] of result) mlDataMap.set(id, payload);
} catch (err) {
const reason = err instanceof Error ? err.message : String(err);
failedMLBatches++;
log(
`ML data request for ${batch.length} file(s) failed: ${reason}`,
);
for (const id of batch) mlDataErrors.set(id, reason);
}
}
log(`Got ML data for ${mlDataMap.size} file(s)`); log(`Got ML data for ${mlDataMap.size} file(s)`);
const writtenFileIDs = new Set<number>(); const writtenFileIDs = new Set<number>();
@@ -233,6 +257,8 @@ export const runMetadataBackup = async (
const ml = mlDataMap.get(file.id); const ml = mlDataMap.get(file.id);
if (ml) fileMeta.mlData = ml; if (ml) fileMeta.mlData = ml;
const mlError = mlDataErrors.get(file.id);
if (mlError) fileMeta.mlDataError = mlError;
if (wantExif && !writtenFileIDs.has(file.id)) { if (wantExif && !writtenFileIDs.has(file.id)) {
log(`[${file.metadata.title}] Extracting EXIF...`); log(`[${file.metadata.title}] Extracting EXIF...`);
@@ -253,4 +279,5 @@ export const runMetadataBackup = async (
} }
log("Metadata backup complete."); log("Metadata backup complete.");
return { failedMLBatches };
}; };
+2 -25
View File
@@ -5,9 +5,8 @@
// comes back encrypted under the file's own key and gzipped; decrypting and // comes back encrypted under the file's own key and gzipped; decrypting and
// gunzipping yields the JSON payload // gunzipping yields the JSON payload
// `{ face: { faces: [...] }, clip: { embedding } }`. Ente caps a request at 200 // `{ face: { faces: [...] }, clip: { embedding } }`. Ente caps a request at 200
// ids, so `fetchMLData` batches for callers that want many at once while // ids, so callers that want many at once split them into batches of
// `fetchMLDataBatch` is the single-request unit the library submits to its // `MLDATA_BATCH_SIZE` and call `fetchMLDataBatch` once per batch.
// request pool.
import { gunzipSync } from "node:zlib"; import { gunzipSync } from "node:zlib";
@@ -69,25 +68,3 @@ export const fetchMLDataBatch = async (
} }
return result; return result;
}; };
// Fetch ML data for arbitrarily many ids, batching at `MLDATA_BATCH_SIZE`. Used
// by the one-shot metadata backup; the library fetches through its request pool
// with `fetchMLDataBatch` instead.
export const fetchMLData = async (
api: ApiClient,
fileIDs: number[],
fileKeys: Map<number, Uint8Array>,
): Promise<Map<number, MLData>> => {
const result = new Map<number, MLData>();
for (let i = 0; i < fileIDs.length; i += MLDATA_BATCH_SIZE) {
const batch = fileIDs.slice(i, i + MLDATA_BATCH_SIZE);
for (const [id, payload] of await fetchMLDataBatch(
api,
batch,
fileKeys,
)) {
result.set(id, payload);
}
}
return result;
};
+47
View File
@@ -735,6 +735,53 @@ describe("backup album folders", () => {
expect(tree(outDir)).toEqual(before); expect(tree(outDir)).toEqual(before);
}); });
it("leaves the albums an onlyAlbumNames run skips as they were", async () => {
const outDir = join(root, "backup");
// "trip" is skipped by the scoped run but its name clashes with the
// in-scope "Trip", so "Trip" must keep its ID suffix.
const lib = libraryOf([
{
collection: collection(10, "Trip"),
files: [file(1, 10, "a.jpg")],
},
{
collection: collection(11, "trip"),
files: [file(2, 11, "b.jpg")],
},
{
collection: collection(12, "Work"),
files: [file(3, 12, "c.jpg")],
},
]);
const json = (name: string): string =>
readFileSync(join(outDir, "collections", name), "utf-8");
await runBackup(lib, { downloadDirectory: outDir });
const before = tree(outDir);
const skippedJSON = [json("trip (11).json"), json("Work.json")];
const scoped = await runBackup(lib, {
downloadDirectory: outDir,
onlyAlbumNames: ["Trip"],
});
expect(scoped.failed).toBe(0);
expect(before).toEqual([
"Trip (10)/",
"Trip (10)/a.jpg -> ../../originals/1.jpg",
"Trip (10).json",
"Work/",
"Work/c.jpg -> ../../originals/3.jpg",
"Work.json",
"trip (11)/",
"trip (11)/b.jpg -> ../../originals/2.jpg",
"trip (11).json",
]);
expect(tree(outDir)).toEqual(before);
expect([json("trip (11).json"), json("Work.json")]).toEqual(
skippedJSON,
);
});
it("removes links and album folders that are gone, and nothing the user added", async () => { it("removes links and album folders that are gone, and nothing the user added", async () => {
const outDir = join(root, "backup"); const outDir = join(root, "backup");
const albums: Album[] = [ const albums: Album[] = [
+73 -2
View File
@@ -38,7 +38,7 @@ import { join } from "node:path";
import { tmpdir } from "node:os"; import { tmpdir } from "node:os";
import sodium from "libsodium-wrappers-sumo"; import sodium from "libsodium-wrappers-sumo";
import { SRP, SrpServer } from "fast-srp-hap"; import { SRP, SrpServer } from "fast-srp-hap";
import { beforeAll, afterAll, describe, expect, it } from "vitest"; import { beforeAll, afterAll, describe, expect, it, vi } from "vitest";
import { import {
init, init,
toBase64, toBase64,
@@ -53,8 +53,16 @@ import {
runMetadataBackup, runMetadataBackup,
type MetadataBackupOptions, type MetadataBackupOptions,
} from "../../src/metadata-backup.js"; } from "../../src/metadata-backup.js";
import { backupMetadataCommand } from "../../src/cli-commands.js";
import type { KeyAttributes } from "../../src/auth/types.js"; import type { KeyAttributes } from "../../src/auth/types.js";
// One file per ML data request, so the two files of the mock account are
// fetched in two requests and one of them can fail on its own.
vi.mock("../../src/mldata-fetch.js", async (importOriginal) => ({
...(await importOriginal<typeof import("../../src/mldata-fetch.js")>()),
MLDATA_BATCH_SIZE: 1,
}));
const TEST_EMAIL = "metabackup@example.com"; const TEST_EMAIL = "metabackup@example.com";
const TEST_PASSWORD = "metapass"; const TEST_PASSWORD = "metapass";
const TEST_OPS = 2; const TEST_OPS = 2;
@@ -347,7 +355,8 @@ const buildMetaMock = async (): Promise<MetaMockState> => {
}; };
}; };
const buildMetaFetch = (m: MetaMockState) => { // `failMLDataFor`: answer 500 to every ML data request that asks for this file.
const buildMetaFetch = (m: MetaMockState, failMLDataFor?: number) => {
let srpServer: SrpServer; let srpServer: SrpServer;
return (async ( return (async (
input: RequestInfo | URL, input: RequestInfo | URL,
@@ -403,6 +412,8 @@ const buildMetaFetch = (m: MetaMockState) => {
} }
if (path === "/files/data/fetch") { if (path === "/files/data/fetch") {
const body = JSON.parse(init?.body as string); const body = JSON.parse(init?.body as string);
if ((body.fileIDs as number[]).includes(failMLDataFor!))
return new Response("server error", { status: 500 });
const data = (body.fileIDs as number[]) const data = (body.fileIDs as number[])
.filter((id: number) => m.encryptedMLData[id]) .filter((id: number) => m.encryptedMLData[id])
.map((id: number) => ({ .map((id: number) => ({
@@ -636,3 +647,63 @@ describe("quak backup-metadata", () => {
expect(failedMeta.imageMetadataError).toEqual(expect.any(String)); expect(failedMeta.imageMetadataError).toEqual(expect.any(String));
}); });
}); });
describe("quak backup-metadata when an ML data request fails", () => {
// Run the CLI command against the mock and return its exit code, stderr
// and output directory.
const runCommand = async (failMLDataFor?: number) => {
const client = await Client.login({
email: TEST_EMAIL,
password: TEST_PASSWORD,
apiOptions: {
fetch: buildMetaFetch(mock, failMLDataFor),
retry: { sleep: async () => {} },
},
});
const outDir = mkdtempSync(join(testDir, "ml-fail-"));
let stderr = "";
const code = await backupMetadataCommand(
{
stdout: { write: () => true },
stderr: { write: (text: string) => (stderr += text) },
sessionDir: testDir,
cacheDir: mkdtempSync(join(testDir, "cache-")),
loadSession: () => client,
},
outDir,
{},
);
return { code, stderr, outDir };
};
it("writes every file, marks the failed batch's files, and exits 1", async () => {
const { code, stderr, outDir } = await runCommand(200);
expect(code).toBe(1);
expect(stderr).toContain("ML data request for 1 file(s) failed");
const ok = JSON.parse(
readFileSync(
join(outDir, "collections", "10-Vacation", "100.json"),
"utf-8",
),
);
expect(ok.mlData.clip.embedding).toEqual([0.5, 0.6, 0.7]);
expect(ok.mlDataError).toBeUndefined();
const failed = JSON.parse(
readFileSync(
join(outDir, "collections", "20-__Work", "200.json"),
"utf-8",
),
);
expect(failed.metadata.title).toBe("diagram.png");
expect(failed.mlData).toBeUndefined();
expect(failed.mlDataError).toContain("500");
});
it("exits 0 when every ML data request succeeds", async () => {
const { code } = await runCommand();
expect(code).toBe(0);
});
});
+61
View File
@@ -13,6 +13,8 @@
* 2. `Library` wiring: after each refresh the library fetches ML data through * 2. `Library` wiring: after each refresh the library fetches ML data through
* the metadata pool for every known file not yet cached, is incremental on * the metadata pool for every known file not yet cached, is incremental on
* later refreshes, and refetches a file whose `updationTime` advanced. * later refreshes, and refetches a file whose `updationTime` advanced.
* Opening a cache directory written by another account starts empty,
* its ML data included (issue #104).
* *
* Embedding values are chosen to be exactly representable as float32 so the * Embedding values are chosen to be exactly representable as float32 so the
* round-trip through `clip.f32` compares equal. * round-trip through `clip.f32` compares equal.
@@ -498,4 +500,63 @@ describe("Library ML-data fetch on refresh", () => {
await lib.close(); await lib.close();
} }
}); });
it("starts empty when the cache directory holds another account's cache", async () => {
// Account A fills the cache directory: metadata and ML data.
const clientA = new MLMockClient();
clientA.collectionsQueue.push({
collections: [collection(1, 100)],
deleted: [],
cursor: 100,
});
clientA.filesFor(1, {
files: [file(1001, 1, 90)],
deleted: [],
cursor: 90,
});
clientA.mlByFile.set(1001, payload([0.5, 0.25, 0.75]));
const libA = await Library.open({
client: clientA,
cacheDirectory,
refreshIntervalSeconds: 3600,
});
try {
await vi.waitFor(
() => expect(libA.status().lastMLFetchAt).toBeGreaterThan(0),
{ timeout: 2000, interval: 5 },
);
} finally {
await libA.close();
}
// Account B opens the same directory.
const clientB = new MLMockClient();
clientB.userID = USER_ID + 1;
const sinceTimes: number[] = [];
const realCollectionsSince = clientB.collectionsSince.bind(clientB);
clientB.collectionsSince = async (args) => {
sinceTimes.push(args.sinceTime);
return realCollectionsSince(args);
};
const libB = await Library.open({
client: clientB,
cacheDirectory,
refreshIntervalSeconds: 3600,
});
try {
expect(sinceTimes[0]).toBe(0);
expect(libB.status().userID).toBe(USER_ID + 1);
expect(libB.listCollections()).toEqual([]);
expect(libB.getFile(1, 1001)).toBeUndefined();
expect(await libB.mldata.forFile({ fileID: 1001 })).toBeUndefined();
expect(
libB.mldata.searchByEmbedding({ embedding: [0.5, 0.25, 0.75] }),
).toEqual([]);
expect(
existsSync(join(cacheDirectory, "mldata", "1001.json")),
).toBe(false);
} finally {
await libB.close();
}
});
}); });