This is a plan for your review, not a finished spec — please critique and send rework requests. It covers the two goals you set: (1) a higher-level API over two long-lived caches — a SQLite metadata cache and an on-disk content/thumbnail cache — and (2) a backup tool that is extremely reliable and built on that API. Design rule applied throughout: the simplest thing that is fit for purpose, fewest moving parts, no tunables the goals don't need.
Current state (grounded)
No cache exists.Client holds only the session keys + token in memory. runBackup (src/backup.ts) and runMetadataBackup (src/metadata-backup.ts) re-enumerate and re-decrypt the whole account on every run. The only state that survives a run is the files on disk.
The incremental-sync machinery exists but is thrown away.Client.listCollections() always calls /collections/v2 with sinceTime: 0; listFiles() paginates /collections/v2/diff but always restarts at sinceTime: 0, and both drop isDeleted tombstones. So the diff/hasMore/updationTime loop is used only for within-run pagination, never to fetch just the changes, and deletions are invisible to callers.
backup <dir> layout (the only persistent state): originals/<fileID>.<ext> (content) + originals/<fileID>.json (metadata sidecar); collections/<name>/<title> symlinks into ../../originals; collections/<name>.json. Skip rule is existsSync && size > 0 — no integrity check. The sidecar is written only when absent, so edited metadata goes stale. Per-file download failures are logged, counted, and stepped over (errors[], exit 1 if any failed); that list is discarded at process exit.
Identity/versioning is already clean in the data:fileID identifies content (dedup across collections), (collectionID, fileID) is a membership, updationTime (µs) advances on any change. Your (collectionID, fileID, updationTime) key is exactly right. FileMetadata.hash carries a plaintext content hash (optional — not on every file). RawEnteFile.info.fileSize/thumbSize exist on the wire but decryptFile drops them (and drops isDeleted).
Reliability primitives already present and reusable: atomic temp-sibling+rename (writeAtomic), TAG_FINAL truncation detection in streamDecrypt, the retry classifier (withRetry/isRetryable/isSafeToReplay), per-request and per-body deadlines.
Design — the cache and its API
Where it lives. Inside the backup target directory (recommended): the SQLite file and the content/thumbnail blobs are subdirectories of the user's backup <dir>. One location, self-contained and portable, multiple independent mirrors, nothing hidden. (Alternative in the open questions: one managed store under env-paths.) "Long-lived" is a property of invalidation, not location: both caches are invalidated only by the server's own change feed, so they carry no TTL and can persist indefinitely wherever they live.
The SQLite metadata cache — what it stores, and why long-lived. Tables (minimal):
collection: id, ownerID, name, type, updationTime, isShared, the three magic-metadata layers (JSON), the decrypted collection key, deleted, and this collection's file-diff cursor sinceTime.
file: one row per membership, PK (collectionID, fileID), with updationTime, the basic + two magic metadata layers (JSON), fileHeader/thumbHeader, the decrypted file key, contentHash, fileSize/thumbSize.
content: one row per unique fileID — presence + verification state of the on-disk original and thumbnail, the stored contentHash, byte length, extension.
a small meta row: collections-list cursor, userID, schema version.
Populated from the diff endpoints, decrypted with the existing decryptCollection/decryptFile. Invalidation is server-driven, keyed on updationTime: a diff row newer than the stored one replaces the metadata; if fileHeader/contentHash changed, the content row is marked stale for re-download. Tombstones (isDeleted) remove the membership (or mark the collection deleted). It is long-lived because nothing else can make it stale — there is no speculative caching here; it is an authoritative local mirror, and keeping it turns every later run into an O(changes) diff instead of an O(account) full re-enumeration + re-decrypt.
Secrets note: this DB holds decrypted metadata (titles, GPS) and the decrypted collection/file keys, so it is as sensitive as session.json — create it 0600 in a 0700 dir and treat it like the password. (Alternative: store only the raw encryptedKey/nonce and re-derive keys from the in-memory master key on read; more moving parts for the same on-disk sensitivity — not recommended.)
The on-disk content + thumbnail cache — layout, keying, lifetime.
Keyed by fileID (content dedup — one download regardless of how many collections hold it; matches today). Layout: originals/<fileID>.<ext> and thumbnails/<fileID>.<ext>. Flat directories; no prefix sharding unless a real account proves it necessary.
Every write goes through writeAtomic (temp sibling + rename). The content row is flipped to "present + verified" only after the rename succeeds, so the DB can never claim a file the disk does not fully have.
Long-lived because Ente content is immutable per fileID — an edit yields a new version observable as a new updationTime (and a changed fileHeader/contentHash). A hash-verified original never needs re-fetching unless the diff reports a new version. No TTL.
The higher-level API — only what callers and the backup tool need. A cache object, constructed from a Client and a directory (the concrete class name is yours to choose). Surface:
sync() — pull each collection's diff from its stored cursor, update the DB, apply tombstones; returns what changed. This is the incremental enumeration thrown away today.
collections() / files(collectionID) — served from the DB, no network. (Also removes the O(collections) linear scan get/get-thumb do now.)
ensureOriginal(fileID) / ensureThumbnail(fileID) — idempotent: download + verify + store if absent-or-stale, else a no-op; returns the path.
verify(fileID) / verifyAll() — re-hash on-disk bytes against the stored hash; mark mismatches stale.
pending() — files whose content is absent, stale, or previously failed.
pathFor(fileID).
That is the whole surface. The only inputs are the Client and the directory — no tunables beyond what ApiClient already exposes. It needs small additions to Client (which owns the keys and decryption): resumable enumerators that accept a starting sinceTime, return the final cursor, and surface isDeleted rows — today's listCollections/listFiles reset to 0 and drop tombstones. ML/EXIF (backup-metadata) is out of scope for the core two caches; it can later read from the same DB.
Design — the crash-safe backup tool on the API
runBackup becomes: sync(), then ensureOriginal (and optionally ensureThumbnail) for each pending() file, then materialize the collections/ symlink views + collection JSON from the DB. Concretely:
Idempotency: a file present and verified in the DB is an O(1) no-op — no stat, no re-hash, no re-download. Each content version is downloaded exactly once.
Resumability after interruption: the DB is the durable progress ledger. A crash leaves content either fully renamed-in (recorded) or not (an orphan temp file, unrecorded) — never half. On restart, pending() is exactly the unfinished set and the persisted cursor means no full re-enumeration.
Atomic writes:writeAtomic for bytes; DB mutations in a transaction; the rule "rename the bytes, then record the row" keeps the DB from overstating the disk. The symlink tree and collection JSON are derived views, rebuildable from the DB, so they need no crash-safety of their own — and rebuilding the view also repairs the stale-sidecar and symlink-failure problems.
Integrity verification:TAG_FINAL already rejects truncation at decrypt time; after decrypt, compare the content hash to FileMetadata.hash when present and store it; verifyAll() re-checks disk against the DB on demand. This replaces today's size > 0. (Impl note: confirm the exact hash construction — algorithm, and the live-photo combined case — before relying on it; fall back to info.fileSize when hash is absent.)
Per-failure handling: reuse the retry classifier. A file that fails after retries is recorded in failure with its classification and attempt count, and the run continues (as today). The next run retries transient/unknown failures and can skip or de-prioritize permanent ones. Exit non-zero while unresolved failures remain. This turns today's lost in-memory errors[] into a durable, queryable ledger.
Deletions: when the diff reports a file/collection gone, keep the original on disk (it is a backup) and drop only the membership + its view. (Pruning to mirror the account is the alternative — open question.)
Ordered next steps (smallest set; reuse vs new)
All TDD per the repo workflow (tests first, red commit, branch off main). The external backup <dir> layout and the CLI contract stay unchanged.
Carry info.fileSize/thumbSize and isDeleted through decryptFile into EnteFile. Tiny; needed for integrity and deletion. Extend existing.
Resumable, tombstone-surfacing enumeration on Client: variants of listCollections/listFiles that take a starting sinceTime, return the final cursor, and include isDeleted rows. Extend existing; ApiClient already accepts an arbitrary sinceTime.
On-disk content/thumbnail store: keyed by fileID, atomic write, hash verification, DB state recorded only post-rename, orphan temp reaping on open. New; reuses writeAtomic, streamDecrypt/TAG_FINAL, downloadFile/downloadThumbnail.
The cache API (sync, collections/files, ensureOriginal/ensureThumbnail, verify, pending, pathFor). Thin façade over 3–4.
Rewrite runBackup on the API + the durable failure ledger; port backup.test.ts; keep the layout and exit-code contract. Rewrite.
Sequencing vs v1.0.0 is yours to set — this subsumes some 1.0.0 items (notably the backup-robustness issue) and realizes the README/TODO "local cache (SQLite) … reliable" goal inside this repo, ahead of the desktop client.
Relationship to existing issues
#8 — runBackup symlink crash + partial originals: subsumed by step 6 (derived views cannot abort the run) and step 4 (atomic write + post-rename recording makes partial originals impossible).
#7 — listFiles infinite loop on a non-advancing server: the resumable enumerator in step 2 must handle it.
#22 — atomic-write durability + orphan reaping: the content store (step 4) is where reaping lands.
#21 — streamDecrypt buffering: open question 4 (streaming decrypt-to-disk) resolves the whole-file buffer too.
#24 — retry/timeout follow-ups: the per-failure ledger reuses and leans on the classifier.
#9 — filename sanitization: keying content by fileID removes the hazard for originals; the symlink view still sanitizes titles.
#10 — Client session/keys: step 2 touches Client; the cache depends on its keys.
#13 — README API reference rewrite: should wait for / incorporate the new surface.
Open questions
Cache location — inside backup <dir> (recommended: portable, self-contained, survives) vs one global store under env-paths (paths.cache is unused today; paths.data holds session.json). Global suits a future always-on desktop client; in-<dir> suits an explicit, movable backup. Recommend in-<dir> now and revisit for the desktop client.
SQLite driver — Node's built-in node:sqlite (zero new dependency, best for the hash-pinned supply chain, but still flagged experimental and needs a recent Node baseline) vs better-sqlite3 (mature, synchronous, a native build). The global "prefer stdlib" rule points to node:sqlite if your Node baseline supports it and experimental status is acceptable; otherwise better-sqlite3. Recommend node:sqlite, falling back to better-sqlite3.
Deleted files — keep originals on disk and drop only the collection view (recommended, backup semantics) vs prune to mirror the live account. A backup that silently deletes your removed photos is surprising; recommend keep.
Large-file memory — downloads currently buffer the whole plaintext in RAM (≈2× transiently), so a multi-gigabyte video can OOM the run. Add a streaming decrypt-to-temp path in ensureOriginal (recommended for "extremely reliable"; an extension of writeAtomic + streamDecrypt) vs keep whole-file buffering (simplest). Overlaps issue 21.
Model: opus-4-8
This is a plan for your review, not a finished spec — please critique and send rework requests. It covers the two goals you set: (1) a higher-level API over two long-lived caches — a SQLite metadata cache and an on-disk content/thumbnail cache — and (2) a backup tool that is extremely reliable and built on that API. Design rule applied throughout: the simplest thing that is fit for purpose, fewest moving parts, no tunables the goals don't need.
## Current state (grounded)
- **No cache exists.** `Client` holds only the session keys + token in memory. `runBackup` (`src/backup.ts`) and `runMetadataBackup` (`src/metadata-backup.ts`) re-enumerate and re-decrypt the whole account on every run. The only state that survives a run is the files on disk.
- **The incremental-sync machinery exists but is thrown away.** `Client.listCollections()` always calls `/collections/v2` with `sinceTime: 0`; `listFiles()` paginates `/collections/v2/diff` but always restarts at `sinceTime: 0`, and both **drop `isDeleted` tombstones**. So the `diff`/`hasMore`/`updationTime` loop is used only for within-run pagination, never to fetch just the changes, and deletions are invisible to callers.
- **`backup <dir>` layout** (the only persistent state): `originals/<fileID>.<ext>` (content) + `originals/<fileID>.json` (metadata sidecar); `collections/<name>/<title>` symlinks into `../../originals`; `collections/<name>.json`. Skip rule is `existsSync && size > 0` — no integrity check. The sidecar is written only when absent, so edited metadata goes stale. Per-file download failures are logged, counted, and stepped over (`errors[]`, exit 1 if any failed); that list is discarded at process exit.
- **Identity/versioning is already clean in the data:** `fileID` identifies content (dedup across collections), `(collectionID, fileID)` is a membership, `updationTime` (µs) advances on any change. Your `(collectionID, fileID, updationTime)` key is exactly right. `FileMetadata.hash` carries a plaintext content hash (optional — not on every file). `RawEnteFile.info.fileSize`/`thumbSize` exist on the wire but `decryptFile` drops them (and drops `isDeleted`).
- **Reliability primitives already present and reusable:** atomic temp-sibling+rename (`writeAtomic`), `TAG_FINAL` truncation detection in `streamDecrypt`, the retry classifier (`withRetry`/`isRetryable`/`isSafeToReplay`), per-request and per-body deadlines.
## Design — the cache and its API
**Where it lives.** Inside the backup target directory (recommended): the SQLite file and the content/thumbnail blobs are subdirectories of the user's `backup <dir>`. One location, self-contained and portable, multiple independent mirrors, nothing hidden. (Alternative in the open questions: one managed store under `env-paths`.) "Long-lived" is a property of invalidation, not location: both caches are invalidated only by the server's own change feed, so they carry no TTL and can persist indefinitely wherever they live.
**The SQLite metadata cache — what it stores, and why long-lived.** Tables (minimal):
- `collection`: `id`, `ownerID`, `name`, `type`, `updationTime`, `isShared`, the three magic-metadata layers (JSON), the decrypted collection key, `deleted`, and this collection's file-diff cursor `sinceTime`.
- `file`: one row per membership, PK `(collectionID, fileID)`, with `updationTime`, the basic + two magic metadata layers (JSON), `fileHeader`/`thumbHeader`, the decrypted file key, `contentHash`, `fileSize`/`thumbSize`.
- `content`: one row per unique `fileID` — presence + verification state of the on-disk original and thumbnail, the stored `contentHash`, byte length, extension.
- `failure`: durable per-`fileID` (× kind) ledger — classification (transient/permanent/unknown), message, attempts, last-tried time.
- a small `meta` row: collections-list cursor, `userID`, schema version.
Populated from the diff endpoints, decrypted with the existing `decryptCollection`/`decryptFile`. **Invalidation is server-driven, keyed on `updationTime`:** a diff row newer than the stored one replaces the metadata; if `fileHeader`/`contentHash` changed, the `content` row is marked stale for re-download. Tombstones (`isDeleted`) remove the membership (or mark the collection deleted). It is _long-lived_ because nothing else can make it stale — there is no speculative caching here; it is an authoritative local mirror, and keeping it turns every later run into an O(changes) diff instead of an O(account) full re-enumeration + re-decrypt.
Secrets note: this DB holds decrypted metadata (titles, GPS) and the decrypted collection/file keys, so it is as sensitive as `session.json` — create it `0600` in a `0700` dir and treat it like the password. (Alternative: store only the raw `encryptedKey`/nonce and re-derive keys from the in-memory master key on read; more moving parts for the same on-disk sensitivity — not recommended.)
**The on-disk content + thumbnail cache — layout, keying, lifetime.**
- Keyed by `fileID` (content dedup — one download regardless of how many collections hold it; matches today). Layout: `originals/<fileID>.<ext>` and `thumbnails/<fileID>.<ext>`. Flat directories; no prefix sharding unless a real account proves it necessary.
- Every write goes through `writeAtomic` (temp sibling + rename). **The `content` row is flipped to "present + verified" only after the rename succeeds**, so the DB can never claim a file the disk does not fully have.
- _Long-lived_ because Ente content is immutable per `fileID` — an edit yields a new version observable as a new `updationTime` (and a changed `fileHeader`/`contentHash`). A hash-verified original never needs re-fetching unless the diff reports a new version. No TTL.
**The higher-level API — only what callers and the backup tool need.** A cache object, constructed from a `Client` and a directory (the concrete class name is yours to choose). Surface:
- `sync()` — pull each collection's diff from its stored cursor, update the DB, apply tombstones; returns what changed. This is the incremental enumeration thrown away today.
- `collections()` / `files(collectionID)` — served from the DB, no network. (Also removes the O(collections) linear scan `get`/`get-thumb` do now.)
- `ensureOriginal(fileID)` / `ensureThumbnail(fileID)` — idempotent: download + verify + store if absent-or-stale, else a no-op; returns the path.
- `verify(fileID)` / `verifyAll()` — re-hash on-disk bytes against the stored hash; mark mismatches stale.
- `pending()` — files whose content is absent, stale, or previously failed.
- `pathFor(fileID)`.
That is the whole surface. The only inputs are the `Client` and the directory — no tunables beyond what `ApiClient` already exposes. It needs small additions to `Client` (which owns the keys and decryption): resumable enumerators that accept a starting `sinceTime`, return the final cursor, and surface `isDeleted` rows — today's `listCollections`/`listFiles` reset to 0 and drop tombstones. ML/EXIF (`backup-metadata`) is out of scope for the core two caches; it can later read from the same DB.
## Design — the crash-safe backup tool on the API
`runBackup` becomes: `sync()`, then `ensureOriginal` (and optionally `ensureThumbnail`) for each `pending()` file, then materialize the `collections/` symlink views + collection JSON from the DB. Concretely:
- **Idempotency:** a file present and verified in the DB is an O(1) no-op — no stat, no re-hash, no re-download. Each content version is downloaded exactly once.
- **Resumability after interruption:** the DB is the durable progress ledger. A crash leaves content either fully renamed-in (recorded) or not (an orphan temp file, unrecorded) — never half. On restart, `pending()` is exactly the unfinished set and the persisted cursor means no full re-enumeration.
- **Atomic writes:** `writeAtomic` for bytes; DB mutations in a transaction; the rule "rename the bytes, then record the row" keeps the DB from overstating the disk. The symlink tree and collection JSON are derived views, rebuildable from the DB, so they need no crash-safety of their own — and rebuilding the view also repairs the stale-sidecar and symlink-failure problems.
- **Integrity verification:** `TAG_FINAL` already rejects truncation at decrypt time; after decrypt, compare the content hash to `FileMetadata.hash` when present and store it; `verifyAll()` re-checks disk against the DB on demand. This replaces today's `size > 0`. (Impl note: confirm the exact hash construction — algorithm, and the live-photo combined case — before relying on it; fall back to `info.fileSize` when `hash` is absent.)
- **Per-failure handling:** reuse the retry classifier. A file that fails after retries is recorded in `failure` with its classification and attempt count, and the run continues (as today). The next run retries transient/unknown failures and can skip or de-prioritize permanent ones. Exit non-zero while unresolved failures remain. This turns today's lost in-memory `errors[]` into a durable, queryable ledger.
- **Deletions:** when the diff reports a file/collection gone, keep the original on disk (it is a backup) and drop only the membership + its view. (Pruning to mirror the account is the alternative — open question.)
## Ordered next steps (smallest set; reuse vs new)
All TDD per the repo workflow (tests first, red commit, branch off `main`). The external `backup <dir>` layout and the CLI contract stay unchanged.
1. **Carry `info.fileSize`/`thumbSize` and `isDeleted` through `decryptFile` into `EnteFile`.** Tiny; needed for integrity and deletion. _Extend existing._
2. **Resumable, tombstone-surfacing enumeration on `Client`:** variants of `listCollections`/`listFiles` that take a starting `sinceTime`, return the final cursor, and include `isDeleted` rows. _Extend existing; `ApiClient` already accepts an arbitrary `sinceTime`._
3. **SQLite metadata cache + `sync()`:** the schema above, cursor persistence, tombstone application, `updationTime` invalidation. _New; reuses decrypt + enumeration._
4. **On-disk content/thumbnail store:** keyed by `fileID`, atomic write, hash verification, DB state recorded only post-rename, orphan temp reaping on open. _New; reuses `writeAtomic`, `streamDecrypt`/`TAG_FINAL`, `downloadFile`/`downloadThumbnail`._
5. **The cache API** (`sync`, `collections`/`files`, `ensureOriginal`/`ensureThumbnail`, `verify`, `pending`, `pathFor`). _Thin façade over 3–4._
6. **Rewrite `runBackup` on the API** + the durable `failure` ledger; port `backup.test.ts`; keep the layout and exit-code contract. _Rewrite._
Sequencing vs `v1.0.0` is yours to set — this subsumes some 1.0.0 items (notably the backup-robustness issue) and realizes the README/TODO "local cache (SQLite) … reliable" goal inside this repo, ahead of the desktop client.
## Relationship to existing issues
- https://git.eeqj.de/sneak/quak/issues/8 — `runBackup` symlink crash + partial originals: subsumed by step 6 (derived views cannot abort the run) and step 4 (atomic write + post-rename recording makes partial originals impossible).
- https://git.eeqj.de/sneak/quak/issues/7 — `listFiles` infinite loop on a non-advancing server: the resumable enumerator in step 2 must handle it.
- https://git.eeqj.de/sneak/quak/issues/22 — atomic-write durability + orphan reaping: the content store (step 4) is where reaping lands.
- https://git.eeqj.de/sneak/quak/issues/21 — `streamDecrypt` buffering: open question 4 (streaming decrypt-to-disk) resolves the whole-file buffer too.
- https://git.eeqj.de/sneak/quak/issues/24 — retry/timeout follow-ups: the per-failure ledger reuses and leans on the classifier.
- https://git.eeqj.de/sneak/quak/issues/9 — filename sanitization: keying content by `fileID` removes the hazard for originals; the symlink view still sanitizes titles.
- https://git.eeqj.de/sneak/quak/issues/10 — `Client` session/keys: step 2 touches `Client`; the cache depends on its keys.
- https://git.eeqj.de/sneak/quak/issues/13 — README API reference rewrite: should wait for / incorporate the new surface.
## Open questions
1. **Cache location** — inside `backup <dir>` (recommended: portable, self-contained, survives) vs one global store under `env-paths` (`paths.cache` is unused today; `paths.data` holds `session.json`). Global suits a future always-on desktop client; in-`<dir>` suits an explicit, movable backup. Recommend in-`<dir>` now and revisit for the desktop client.
2. **SQLite driver** — Node's built-in `node:sqlite` (zero new dependency, best for the hash-pinned supply chain, but still flagged experimental and needs a recent Node baseline) vs `better-sqlite3` (mature, synchronous, a native build). The global "prefer stdlib" rule points to `node:sqlite` if your Node baseline supports it and experimental status is acceptable; otherwise `better-sqlite3`. Recommend `node:sqlite`, falling back to `better-sqlite3`.
3. **Deleted files** — keep originals on disk and drop only the collection view (recommended, backup semantics) vs prune to mirror the live account. A backup that silently deletes your removed photos is surprising; recommend keep.
4. **Large-file memory** — downloads currently buffer the whole plaintext in RAM (≈2× transiently), so a multi-gigabyte video can OOM the run. Add a streaming decrypt-to-temp path in `ensureOriginal` (recommended for "extremely reliable"; an extension of `writeAtomic` + `streamDecrypt`) vs keep whole-file buffering (simplest). Overlaps issue 21.
Model: opus-4-8
I don't like the idea that the user (of the library) has to call sync - I want the cache to be transparent to the user, they should just use the API as if it were consulting the remote database as local.
Please write me 5 different blocks of example code, using the high-level library in five different ways, each being a different API surface/style of the high level API, so I can evaluate them and decide.
Why are we using SQLite instead of json? I expect all metadata to comfortably fit in 100MB. Reading entire files into ram and writing entire caches to disk on update will be trivial affairs. Keep our on-disk formats dead simple.
I don't like the idea that the user (of the library) has to call sync - I want the cache to be transparent to the user, they should just use the API as if it were consulting the remote database as local.
Please write me 5 different blocks of example code, using the high-level library in five different ways, each being a different API surface/style of the high level API, so I can evaluate them and decide.
Why are we using SQLite instead of json? I expect all metadata to comfortably fit in 100MB. Reading entire files into ram and writing entire caches to disk on update will be trivial affairs. Keep our on-disk formats dead simple.
Reading of your ruling: no sync() anywhere in the public surface. Every read
first asks the server for changes since the stored cursor (one small diff
request, usually empty), folds them into the local copy, saves it, then answers
from RAM. The caller never sees the cache. The existing Client
(Client.fromJSON, listCollections, listFiles, downloadFile) stays as the
low-level layer underneath.
Each block below does the same task: open the library on a directory, find the
album "Iceland 2025", list its files, fetch one original, run a full backup,
report failed files. EnteFile, Collection, BackupResult, BackupError, ClientSnapshot are today's real types; every other name is a placeholder and
yours to choose.
Plain object of async methods
Flat and obvious, closest to today's Client; nothing to learn. Tradeoff: ids
get passed around by hand and the method list grows with every feature.
import{Client,Library}from"quak";constclient=Client.fromJSON(snapshot);constlib=awaitLibrary.open(client,"/backups/ente");constalbums: Collection[]=awaitlib.listCollections();consticeland=albums.find((c)=>c.name==="Iceland 2025")!;constfiles: EnteFile[]=awaitlib.listFiles(iceland.id);// downloads once, verifies, returns the local path; later calls are instant
constpath=awaitlib.originalPath(files[0].id);constthumb=awaitlib.thumbnailPath(files[0].id);constresult: BackupResult=awaitlib.backup();for(consteofresult.errors)console.error(e.fileID,e.title,e.error);// failures survive the process; readable on any later run
conststillBroken: BackupError[]=awaitlib.failures();
Collections of objects
Reads like a small object model: albums hold photos, photos know how to fetch
themselves. Pleasant to explore and hard to misuse (no loose ids). Tradeoff:
more classes to build and document, and objects held a long time can describe
a photo that has since changed on the server.
import{Client,Library}from"quak";constlib=awaitLibrary.open(Client.fromJSON(snapshot),"/backups/ente");consticeland=awaitlib.albums.byName("Iceland 2025");constphotos=awaiticeland.photos.list();constfirst=photos[0];console.log(first.title,first.takenAt,first.fileType);constpath=awaitfirst.original();// local path, fetched if needed
constthumb=awaitfirst.thumbnail();constresult=awaitlib.backup();for(constfofawaitlib.failures.list()){console.error(f.photo.title,f.error,`tried ${f.attempts} times`);}
Directory-like tree
The account is presented as a read-only folder tree, the same shape as the collections/ and originals/ layout on disk. Instantly familiar, and a FUSE
mount or a file browser falls out of it later. Tradeoff: everything becomes a
path string, so metadata (GPS, dates, shared-with) has to be squeezed into a stat call, and duplicate titles in one album need a naming rule.
import{Client,openTree}from"quak";consttree=awaitopenTree(Client.fromJSON(snapshot),"/backups/ente");constalbumNames=awaittree.readdir("/collections");consttitles=awaittree.readdir("/collections/Iceland 2025");constinfo=awaittree.stat(`/collections/Iceland 2025/${titles[0]}`);console.log(info.fileID,info.size,info.metadata.creationTime);// reading a path fetches the original if it is not on disk yet
constbytes=awaittree.readFile(`/collections/Iceland 2025/${titles[0]}`);constlocalPath=awaittree.realpath(`/collections/Iceland 2025/${titles[0]}`);constresult=awaittree.backup();constfailed=awaittree.readdir("/failed");// one entry per failed fileID
Async iterators and an event stream
Everything is a for await loop; nothing large is ever built in memory by the
caller, and a backup reports progress as it happens instead of at the end.
Best fit for a CLI progress bar or a very large account. Tradeoff: simple
questions ("how many files?") need a loop or a helper, and error handling
inside loops is easy to get wrong.
The caller says what should exist on disk and the library does whatever is
needed, every time, from any starting point (fresh, half-done, crashed). The
smallest surface of the five and the best match for "extremely reliable
backup". Tradeoff: it is a backup tool, not a library for browsing; listing an
album or grabbing one file has to be expressed as a narrower wanted state.
import{Client,mirror}from"quak";constclient=Client.fromJSON(snapshot);// full backup: make /backups/ente a complete copy of the account
constreport=awaitmirror(client,"/backups/ente",{originals: true,thumbnails: true,onProgress:(msg)=>console.log(msg),});for(consteofreport.errors)console.error(e.fileID,e.title,e.error);// one album only; returns what is now on disk
consticeland=awaitmirror(client,"/backups/ente",{collections:["Iceland 2025"],});console.log(iceland.files.map((f)=>f.path));// one file
constone=awaitmirror(client,"/backups/ente",{fileIDs:[12345]});
Why SQLite instead of JSON
No good reason under your constraint. It was chosen out of habit for "a durable
ledger with transactions", and the README/TODO wording said SQLite. With
metadata under 100MB, held in RAM and rewritten whole, JSON does the job and
the format is readable with cat and jq. Dropping SQLite also removes open
question 2 (the driver choice) and a dependency. Accepted.
Simplest shape, inside the backup directory:
metadata.json holds everything the server told us: collections, files per
collection, and the change cursors. One file, so the cursors can never
disagree with the data they describe: one rename covers both.
originals/<fileID>.<ext> and thumbnails/<fileID>.<ext> hold content, as
today. No "downloaded" flag is stored anywhere. Content only ever appears by
rename, so a file that exists is complete; the directory listing is the
record of what we have.
failures.json holds the failed-file list (small).
collections/ symlinks and per-collection JSON stay, rebuilt from metadata.json on each run.
What JSON costs here, plainly:
Write frequency is the only real trap. Rewriting 100MB after every downloaded
file would be absurd on a 50,000-file first backup. The shape above avoids it: metadata.json is rewritten only when the server reports changes (once per
refresh, and during the first big enumeration at most once per collection or
every few seconds), never per download. Downloads touch only failures.json,
which is tiny.
Crash safety is fine with temp-file-plus-rename: a crash leaves the old file
or the new one, never half. Worst case we re-ask the server for a few pages
of changes. One gap: today's writeAtomic (src/download/index.ts) does not
fsync, so after a power cut the renamed file can be empty. The metadata
writer must fsync the file before the rename and the directory after it;
that is already tracked in #22. writeAtomic is also private to the download module and needs exporting.
A corrupt or missing metadata.json is not data loss: delete it and the next
read re-fetches everything from the server. It is a cache.
Two processes on one directory would overwrite each other's file (SQLite
would have handled that). Answer: a lock file, one writer per directory, the
second process fails with a clear message.
Memory: parsing 100MB of JSON costs roughly 300-500MB of heap and about a
second at startup. Acceptable for a backup tool; worth knowing for the
desktop client.
Lookups by id are done with in-memory Maps built at load. No query
language is lost that we needed.
Sensitivity is unchanged: the file holds decrypted titles, GPS and file keys,
so 0600 in a 0700 directory, same as session.json.
One question the transparency ruling raises, for later: when the server is
unreachable, should reads fail, or answer from the local copy and say so? I
would answer from the local copy with a flag on the result; backups would still
fail loudly.
My pick: plain object of async methods for the library, with backup() built
internally the declare-the-state way, because it is the smallest step from
today's Client and the other styles can be layered on it later without
breaking it.
Disclosure: posted with the tea CLI as clawbot; the Gitea MCP tool was not
available in this session.
Model: fable-5-1
Reading of your ruling: no `sync()` anywhere in the public surface. Every read
first asks the server for changes since the stored cursor (one small diff
request, usually empty), folds them into the local copy, saves it, then answers
from RAM. The caller never sees the cache. The existing `Client`
(`Client.fromJSON`, `listCollections`, `listFiles`, `downloadFile`) stays as the
low-level layer underneath.
Each block below does the same task: open the library on a directory, find the
album "Iceland 2025", list its files, fetch one original, run a full backup,
report failed files. `EnteFile`, `Collection`, `BackupResult`, `BackupError`,
`ClientSnapshot` are today's real types; every other name is a placeholder and
yours to choose.
## Plain object of async methods
Flat and obvious, closest to today's `Client`; nothing to learn. Tradeoff: ids
get passed around by hand and the method list grows with every feature.
```ts
import { Client, Library } from "quak";
const client = Client.fromJSON(snapshot);
const lib = await Library.open(client, "/backups/ente");
const albums: Collection[] = await lib.listCollections();
const iceland = albums.find((c) => c.name === "Iceland 2025")!;
const files: EnteFile[] = await lib.listFiles(iceland.id);
// downloads once, verifies, returns the local path; later calls are instant
const path = await lib.originalPath(files[0].id);
const thumb = await lib.thumbnailPath(files[0].id);
const result: BackupResult = await lib.backup();
for (const e of result.errors) console.error(e.fileID, e.title, e.error);
// failures survive the process; readable on any later run
const stillBroken: BackupError[] = await lib.failures();
```
## Collections of objects
Reads like a small object model: albums hold photos, photos know how to fetch
themselves. Pleasant to explore and hard to misuse (no loose ids). Tradeoff:
more classes to build and document, and objects held a long time can describe
a photo that has since changed on the server.
```ts
import { Client, Library } from "quak";
const lib = await Library.open(Client.fromJSON(snapshot), "/backups/ente");
const iceland = await lib.albums.byName("Iceland 2025");
const photos = await iceland.photos.list();
const first = photos[0];
console.log(first.title, first.takenAt, first.fileType);
const path = await first.original(); // local path, fetched if needed
const thumb = await first.thumbnail();
const result = await lib.backup();
for (const f of await lib.failures.list()) {
console.error(f.photo.title, f.error, `tried ${f.attempts} times`);
}
```
## Directory-like tree
The account is presented as a read-only folder tree, the same shape as the
`collections/` and `originals/` layout on disk. Instantly familiar, and a FUSE
mount or a file browser falls out of it later. Tradeoff: everything becomes a
path string, so metadata (GPS, dates, shared-with) has to be squeezed into a
`stat` call, and duplicate titles in one album need a naming rule.
```ts
import { Client, openTree } from "quak";
const tree = await openTree(Client.fromJSON(snapshot), "/backups/ente");
const albumNames = await tree.readdir("/collections");
const titles = await tree.readdir("/collections/Iceland 2025");
const info = await tree.stat(`/collections/Iceland 2025/${titles[0]}`);
console.log(info.fileID, info.size, info.metadata.creationTime);
// reading a path fetches the original if it is not on disk yet
const bytes = await tree.readFile(`/collections/Iceland 2025/${titles[0]}`);
const localPath = await tree.realpath(`/collections/Iceland 2025/${titles[0]}`);
const result = await tree.backup();
const failed = await tree.readdir("/failed"); // one entry per failed fileID
```
## Async iterators and an event stream
Everything is a `for await` loop; nothing large is ever built in memory by the
caller, and a backup reports progress as it happens instead of at the end.
Best fit for a CLI progress bar or a very large account. Tradeoff: simple
questions ("how many files?") need a loop or a helper, and error handling
inside loops is easy to get wrong.
```ts
import { Client, Library } from "quak";
const lib = await Library.open(Client.fromJSON(snapshot), "/backups/ente");
for await (const album of lib.collections()) {
if (album.name !== "Iceland 2025") continue;
for await (const file of lib.files(album.id)) {
console.log(file.id, file.metadata.title);
}
}
const failed: BackupError[] = [];
for await (const ev of lib.backup()) {
if (ev.kind === "downloaded") console.log("ok", ev.file.metadata.title);
if (ev.kind === "skipped") continue;
if (ev.kind === "failed") failed.push(ev.error);
}
console.error(`${failed.length} file(s) failed`);
```
## Declare the wanted state, library makes it true
The caller says what should exist on disk and the library does whatever is
needed, every time, from any starting point (fresh, half-done, crashed). The
smallest surface of the five and the best match for "extremely reliable
backup". Tradeoff: it is a backup tool, not a library for browsing; listing an
album or grabbing one file has to be expressed as a narrower wanted state.
```ts
import { Client, mirror } from "quak";
const client = Client.fromJSON(snapshot);
// full backup: make /backups/ente a complete copy of the account
const report = await mirror(client, "/backups/ente", {
originals: true,
thumbnails: true,
onProgress: (msg) => console.log(msg),
});
for (const e of report.errors) console.error(e.fileID, e.title, e.error);
// one album only; returns what is now on disk
const iceland = await mirror(client, "/backups/ente", {
collections: ["Iceland 2025"],
});
console.log(iceland.files.map((f) => f.path));
// one file
const one = await mirror(client, "/backups/ente", { fileIDs: [12345] });
```
## Why SQLite instead of JSON
No good reason under your constraint. It was chosen out of habit for "a durable
ledger with transactions", and the README/TODO wording said SQLite. With
metadata under 100MB, held in RAM and rewritten whole, JSON does the job and
the format is readable with `cat` and `jq`. Dropping SQLite also removes open
question 2 (the driver choice) and a dependency. Accepted.
Simplest shape, inside the backup directory:
- `metadata.json` holds everything the server told us: collections, files per
collection, and the change cursors. One file, so the cursors can never
disagree with the data they describe: one rename covers both.
- `originals/<fileID>.<ext>` and `thumbnails/<fileID>.<ext>` hold content, as
today. No "downloaded" flag is stored anywhere. Content only ever appears by
rename, so a file that exists is complete; the directory listing is the
record of what we have.
- `failures.json` holds the failed-file list (small).
- `collections/` symlinks and per-collection JSON stay, rebuilt from
`metadata.json` on each run.
What JSON costs here, plainly:
- Write frequency is the only real trap. Rewriting 100MB after every downloaded
file would be absurd on a 50,000-file first backup. The shape above avoids it:
`metadata.json` is rewritten only when the server reports changes (once per
refresh, and during the first big enumeration at most once per collection or
every few seconds), never per download. Downloads touch only `failures.json`,
which is tiny.
- Crash safety is fine with temp-file-plus-rename: a crash leaves the old file
or the new one, never half. Worst case we re-ask the server for a few pages
of changes. One gap: today's `writeAtomic` (`src/download/index.ts`) does not
fsync, so after a power cut the renamed file can be empty. The metadata
writer must fsync the file before the rename and the directory after it;
that is already tracked in https://git.eeqj.de/sneak/quak/issues/22.
`writeAtomic` is also private to the download module and needs exporting.
- A corrupt or missing `metadata.json` is not data loss: delete it and the next
read re-fetches everything from the server. It is a cache.
- Two processes on one directory would overwrite each other's file (SQLite
would have handled that). Answer: a lock file, one writer per directory, the
second process fails with a clear message.
- Memory: parsing 100MB of JSON costs roughly 300-500MB of heap and about a
second at startup. Acceptable for a backup tool; worth knowing for the
desktop client.
- Lookups by id are done with in-memory `Map`s built at load. No query
language is lost that we needed.
- Sensitivity is unchanged: the file holds decrypted titles, GPS and file keys,
so `0600` in a `0700` directory, same as `session.json`.
One question the transparency ruling raises, for later: when the server is
unreachable, should reads fail, or answer from the local copy and say so? I
would answer from the local copy with a flag on the result; backups would still
fail loudly.
My pick: plain object of async methods for the library, with `backup()` built
internally the declare-the-state way, because it is the smallest step from
today's `Client` and the other styles can be layered on it later without
breaking it.
Disclosure: posted with the `tea` CLI as `clawbot`; the Gitea MCP tool was not
available in this session.
Model: fable-5-1
all these api shape designs use positional arguments which make it unclear what those paths mean in the second argument. use an args object with descriptive and unambiguous key names.
all these api shape designs use positional arguments which make it unclear what those paths mean in the second argument. use an args object with descriptive and unambiguous key names.
i like the collection of objects shape, but it needs a way of reporting progress in realtime, both when working a list or when working a single thing like a download. perhaps a progressCallback arg?
i like the collection of objects shape, but it needs a way of reporting progress in realtime, both when working a list or when working a single thing like a download. perhaps a progressCallback arg?
Revised API after your two rulings: every call takes one args object with named keys (no positional arguments anywhere in the new surface), and every long operation accepts an optional onProgress callback that fires in real time, both for list work and for a single download.
Standing decisions, restated once:
Transparency: no sync() in the surface. Every read first asks the server for changes since the stored cursor (one small diff request, usually empty), folds them in, saves, then answers from RAM.
On-disk format is JSON: metadata.json (collections, files, cursors), failures.json, plus originals/ and thumbnails/, all written by temp-file-plus-rename.
The surface
Client, ClientSnapshot, Collection, EnteFile, FileType, BackupResult are today's types (src/client.ts, src/model/types.ts, src/backup.ts). Library, Album, Photo, ProgressEvent, Failure are new; the names are yours to change.
importtype{Client,Collection,EnteFile,FileType,BackupResult}from"quak";// One event shape for every operation. A field is absent when not known.
exportinterfaceProgressEvent{operation:|"refresh"// asking the server for changes since the cursor
|"listAlbums"|"listPhotos"|"downloadOriginal"|"downloadThumbnail"|"backup";status:"started"|"progress"|"done"|"skipped"|"failed";album?:{collectionID: number;name: string};// identity of the album being worked
photo?:{fileID: number;title: string};// identity of the photo being worked
itemsDone?: number;// list work and backup: entries handled so far
itemsTotal?: number;// absent until known (a first enumeration pages through a diff)
bytesDone?: number;// a single download: plaintext bytes written so far
bytesTotal?: number;// absent when the server gave no size (FileBlob.size)
error?: string;// set when status is "failed"
}exporttypeProgressCallback=(event: ProgressEvent)=>void;exportclassLibrary{staticopen(args:{client: Client;backupDirectory: string;// holds metadata.json, failures.json, originals/, thumbnails/
}):Promise<Library>;readonlyalbums:{list(args?:{onProgress?: ProgressCallback}):Promise<Album[]>;byName(args:{albumName: string}):Promise<Album|undefined>;byID(args:{collectionID: number}):Promise<Album|undefined>;};readonlyphotos:{// a file is one thing regardless of how many albums hold it
byID(args:{fileID: number}):Promise<Photo|undefined>;};backup(args?:{includeOriginals?: boolean;// default true
includeThumbnails?: boolean;// default false
onlyAlbumNames?: string[];// default: every album
onProgress?: ProgressCallback;}):Promise<BackupResult>;readonlyfailures:{list():Promise<Failure[]>;// durable: readable on any later run
};}exportclassAlbum{readonlycollection: Collection;// the underlying decrypted record
readonlyid: number;// collection.id
readonlyname: string;readonlyphotos:{list(args?:{onProgress?: ProgressCallback}):Promise<Photo[]>;byID(args:{fileID: number}):Promise<Photo|undefined>;};}exportclassPhoto{readonlyfile: EnteFile;// the underlying decrypted record
readonlyid: number;// file.id
readonlytitle: string;readonlytakenAt: Date;// from metadata.creationTime
readonlyfileType: FileType;// Fetch once, verify, keep on disk. If already present: no network,
// one "skipped" event, and the path is returned at once.
original(args?:{onProgress?: ProgressCallback}):Promise<{path: string;bytes: number}>;thumbnail(args?:{onProgress?: ProgressCallback}):Promise<{path: string;bytes: number}>;}exportinterfaceFailure{fileID: number;title: string;albumName: string;kind:"original"|"thumbnail";error: string;attempts: number;lastTriedAt: Date;}
What onProgress emits, per operation:
Any read (albums.list, photos.list, byName, byID): a refreshstarted and done (or failed) pair first, since every read checks the server.
albums.list: started, done. itemsTotal is known at once because /collections/v2 answers in one response.
photos.list: started, then one progress per diff page with itemsDone growing and itemsTotal absent (the server pages with hasMore, so the total is unknown until the end), then done. When the local copy is already current it is just started and done.
original() / thumbnail(): started with photo and bytesTotal when known, progress per chunk with bytesDone, then done, skipped (already on disk) or failed with error.
backup(): backup started with itemsTotal (known after the refresh), then for each file the same download events as above carrying photo and album, then one backup progress with itemsDone after each file, then backup done. Per-item events and overall progress arrive through the one callback, so no second mechanism is needed.
Worked examples
Find an album and list its files with progress:
import{Client,Library}from"quak";constlib=awaitLibrary.open({client: Client.fromJSON(snapshot),backupDirectory:"/backups/ente",});consticeland=awaitlib.albums.byName({albumName:"Iceland 2025"});if(!iceland)thrownewError("no album named Iceland 2025");constphotos=awaiticeland.photos.list({onProgress:(ev)=>{if(ev.operation==="listPhotos"&&ev.status==="progress")process.stderr.write(`\r${ev.itemsDone} entries so far`);},});for(constpofphotos)console.log(p.id,p.title,p.takenAt.toISOString(),p.fileType);
Fetch one original with download progress:
constfirst=photos[0];const{path,bytes}=awaitfirst.original({onProgress:(ev)=>{if(ev.status==="skipped")console.log("already on disk");elseif(ev.status==="progress"&&ev.bytesTotal)process.stderr.write(`\r${first.title}: ${Math.round((100*ev.bytesDone!)/ev.bytesTotal)}%`);elseif(ev.status==="progress")process.stderr.write(`\r${first.title}: ${ev.bytesDone} bytes`);},});console.log(`\n${path} (${bytes} bytes)`);
Full backup with progress and the failures report:
importtype{BackupResult}from"quak";constresult: BackupResult=awaitlib.backup({includeThumbnails: true,onProgress:(ev)=>{if(ev.operation==="backup"&&ev.status==="progress")process.stderr.write(`\r${ev.itemsDone}/${ev.itemsTotal} files`);elseif(ev.operation==="downloadOriginal"&&ev.status==="failed")console.error(`\nfailed: ${ev.album?.name}/${ev.photo?.title}: ${ev.error}`);},});console.log(`\ndownloaded ${result.downloaded}, skipped ${result.skipped}, failed ${result.failed}`);// Durable: the same list is readable on a later run without backing up again.
for(constfofawaitlib.failures.list())console.error(`${f.albumName}/${f.title} (${f.kind}): ${f.error}; tried ${f.attempts} times, last ${f.lastTriedAt.toISOString()}`);process.exitCode=result.failed>0?1 : 0;
Implementation notes, grounded in the current code
downloadFile and downloadThumbnail in src/download/index.ts report bytesWritten only at the end; the chunk loop needs a per-chunk hook so bytesDone can be emitted while the body streams.
bytesTotal is FileBlob.size, which comes from RawEnteFile.info.fileSize; decryptFile drops it today (step 1 of the original plan carries it through).
ProgressCallback in src/backup.ts is (message: string) => void today; it is replaced by the typed event, and the CLI formats events into its log lines.
Judgement call: the low-level Client (listFiles(collectionID, collectionKey), downloadFile(file, outPath?)) is left as is; the ruling is applied to the new surface, and Client can be brought in line in a separate change if you want it.
Open questions
Server unreachable during a read: recommend answering from the local copy and reporting it through the refreshfailed event (with error), while backup() throws; the alternative is that every read fails when offline.
A file deleted on the server: recommend keeping the original on disk and dropping only the album membership (a backup should not delete your removed photos); the alternative is pruning to mirror the live account.
Disclosure: posted with the tea CLI as clawbot; the Gitea MCP tool was not available in this session.
Model: fable-5-1
Revised API after your two rulings: every call takes one args object with named keys (no positional arguments anywhere in the new surface), and every long operation accepts an optional `onProgress` callback that fires in real time, both for list work and for a single download.
Standing decisions, restated once:
- Transparency: no `sync()` in the surface. Every read first asks the server for changes since the stored cursor (one small diff request, usually empty), folds them in, saves, then answers from RAM.
- On-disk format is JSON: `metadata.json` (collections, files, cursors), `failures.json`, plus `originals/` and `thumbnails/`, all written by temp-file-plus-rename.
## The surface
`Client`, `ClientSnapshot`, `Collection`, `EnteFile`, `FileType`, `BackupResult` are today's types (`src/client.ts`, `src/model/types.ts`, `src/backup.ts`). `Library`, `Album`, `Photo`, `ProgressEvent`, `Failure` are new; the names are yours to change.
```ts
import type { Client, Collection, EnteFile, FileType, BackupResult } from "quak";
// One event shape for every operation. A field is absent when not known.
export interface ProgressEvent {
operation:
| "refresh" // asking the server for changes since the cursor
| "listAlbums"
| "listPhotos"
| "downloadOriginal"
| "downloadThumbnail"
| "backup";
status: "started" | "progress" | "done" | "skipped" | "failed";
album?: { collectionID: number; name: string }; // identity of the album being worked
photo?: { fileID: number; title: string }; // identity of the photo being worked
itemsDone?: number; // list work and backup: entries handled so far
itemsTotal?: number; // absent until known (a first enumeration pages through a diff)
bytesDone?: number; // a single download: plaintext bytes written so far
bytesTotal?: number; // absent when the server gave no size (FileBlob.size)
error?: string; // set when status is "failed"
}
export type ProgressCallback = (event: ProgressEvent) => void;
export class Library {
static open(args: {
client: Client;
backupDirectory: string; // holds metadata.json, failures.json, originals/, thumbnails/
}): Promise<Library>;
readonly albums: {
list(args?: { onProgress?: ProgressCallback }): Promise<Album[]>;
byName(args: { albumName: string }): Promise<Album | undefined>;
byID(args: { collectionID: number }): Promise<Album | undefined>;
};
readonly photos: {
// a file is one thing regardless of how many albums hold it
byID(args: { fileID: number }): Promise<Photo | undefined>;
};
backup(args?: {
includeOriginals?: boolean; // default true
includeThumbnails?: boolean; // default false
onlyAlbumNames?: string[]; // default: every album
onProgress?: ProgressCallback;
}): Promise<BackupResult>;
readonly failures: {
list(): Promise<Failure[]>; // durable: readable on any later run
};
}
export class Album {
readonly collection: Collection; // the underlying decrypted record
readonly id: number; // collection.id
readonly name: string;
readonly photos: {
list(args?: { onProgress?: ProgressCallback }): Promise<Photo[]>;
byID(args: { fileID: number }): Promise<Photo | undefined>;
};
}
export class Photo {
readonly file: EnteFile; // the underlying decrypted record
readonly id: number; // file.id
readonly title: string;
readonly takenAt: Date; // from metadata.creationTime
readonly fileType: FileType;
// Fetch once, verify, keep on disk. If already present: no network,
// one "skipped" event, and the path is returned at once.
original(args?: { onProgress?: ProgressCallback }): Promise<{ path: string; bytes: number }>;
thumbnail(args?: { onProgress?: ProgressCallback }): Promise<{ path: string; bytes: number }>;
}
export interface Failure {
fileID: number;
title: string;
albumName: string;
kind: "original" | "thumbnail";
error: string;
attempts: number;
lastTriedAt: Date;
}
```
What `onProgress` emits, per operation:
- Any read (`albums.list`, `photos.list`, `byName`, `byID`): a `refresh` `started` and `done` (or `failed`) pair first, since every read checks the server.
- `albums.list`: `started`, `done`. `itemsTotal` is known at once because `/collections/v2` answers in one response.
- `photos.list`: `started`, then one `progress` per diff page with `itemsDone` growing and `itemsTotal` absent (the server pages with `hasMore`, so the total is unknown until the end), then `done`. When the local copy is already current it is just `started` and `done`.
- `original()` / `thumbnail()`: `started` with `photo` and `bytesTotal` when known, `progress` per chunk with `bytesDone`, then `done`, `skipped` (already on disk) or `failed` with `error`.
- `backup()`: `backup started` with `itemsTotal` (known after the refresh), then for each file the same download events as above carrying `photo` and `album`, then one `backup progress` with `itemsDone` after each file, then `backup done`. Per-item events and overall progress arrive through the one callback, so no second mechanism is needed.
## Worked examples
Find an album and list its files with progress:
```ts
import { Client, Library } from "quak";
const lib = await Library.open({
client: Client.fromJSON(snapshot),
backupDirectory: "/backups/ente",
});
const iceland = await lib.albums.byName({ albumName: "Iceland 2025" });
if (!iceland) throw new Error("no album named Iceland 2025");
const photos = await iceland.photos.list({
onProgress: (ev) => {
if (ev.operation === "listPhotos" && ev.status === "progress")
process.stderr.write(`\r${ev.itemsDone} entries so far`);
},
});
for (const p of photos) console.log(p.id, p.title, p.takenAt.toISOString(), p.fileType);
```
Fetch one original with download progress:
```ts
const first = photos[0];
const { path, bytes } = await first.original({
onProgress: (ev) => {
if (ev.status === "skipped") console.log("already on disk");
else if (ev.status === "progress" && ev.bytesTotal)
process.stderr.write(`\r${first.title}: ${Math.round((100 * ev.bytesDone!) / ev.bytesTotal)}%`);
else if (ev.status === "progress")
process.stderr.write(`\r${first.title}: ${ev.bytesDone} bytes`);
},
});
console.log(`\n${path} (${bytes} bytes)`);
```
Full backup with progress and the failures report:
```ts
import type { BackupResult } from "quak";
const result: BackupResult = await lib.backup({
includeThumbnails: true,
onProgress: (ev) => {
if (ev.operation === "backup" && ev.status === "progress")
process.stderr.write(`\r${ev.itemsDone}/${ev.itemsTotal} files`);
else if (ev.operation === "downloadOriginal" && ev.status === "failed")
console.error(`\nfailed: ${ev.album?.name}/${ev.photo?.title}: ${ev.error}`);
},
});
console.log(`\ndownloaded ${result.downloaded}, skipped ${result.skipped}, failed ${result.failed}`);
// Durable: the same list is readable on a later run without backing up again.
for (const f of await lib.failures.list())
console.error(`${f.albumName}/${f.title} (${f.kind}): ${f.error}; tried ${f.attempts} times, last ${f.lastTriedAt.toISOString()}`);
process.exitCode = result.failed > 0 ? 1 : 0;
```
## Implementation notes, grounded in the current code
- `downloadFile` and `downloadThumbnail` in `src/download/index.ts` report `bytesWritten` only at the end; the chunk loop needs a per-chunk hook so `bytesDone` can be emitted while the body streams.
- `bytesTotal` is `FileBlob.size`, which comes from `RawEnteFile.info.fileSize`; `decryptFile` drops it today (step 1 of the original plan carries it through).
- `ProgressCallback` in `src/backup.ts` is `(message: string) => void` today; it is replaced by the typed event, and the CLI formats events into its log lines.
- Judgement call: the low-level `Client` (`listFiles(collectionID, collectionKey)`, `downloadFile(file, outPath?)`) is left as is; the ruling is applied to the new surface, and `Client` can be brought in line in a separate change if you want it.
## Open questions
- Server unreachable during a read: recommend answering from the local copy and reporting it through the `refresh` `failed` event (with `error`), while `backup()` throws; the alternative is that every read fails when offline.
- A file deleted on the server: recommend keeping the original on disk and dropping only the album membership (a backup should not delete your removed photos); the alternative is pruning to mirror the live account.
Disclosure: posted with the tea CLI as clawbot; the Gitea MCP tool was not available in this session.
Model: fable-5-1
the cachedirectory should be an option and it should default to standard xda cache dirs. downloadDirectory should be an option as well.
how does this library work as a backend client for a GUI tool to replace the ente 1p desktop app? i want to build something very simple like autistmask/universal paperclips that just does a very plain minimum-styled list of photos (by week or month or something), albums, magic tags, etc. the ente 1p one is terribly laggy/slow and with all metadata on disk/in memory and a large thumbnail/data cache, we should be able to build one that's extremely performant (in electron). how does this api shape work for that use case? it seems good to me but i'd like a critical review along with any proposed changes that might make it work best for something like that.
also as part of this api shape implementation i'd like the cli quak tool updated to use the new api.
the cachedirectory should be an option and it should default to standard xda cache dirs. downloadDirectory should be an option as well.
how does this library work as a backend client for a GUI tool to replace the ente 1p desktop app? i want to build something very simple like autistmask/universal paperclips that just does a very plain minimum-styled list of photos (by week or month or something), albums, magic tags, etc. the ente 1p one is terribly laggy/slow and with all metadata on disk/in memory and a large thumbnail/data cache, we should be able to build one that's extremely performant (in electron). how does this api shape work for that use case? it seems good to me but i'd like a critical review along with any proposed changes that might make it work best for something like that.
also as part of this api shape implementation i'd like the cli quak tool updated to use the new api.
Reading of your three orders: both directories become options with the cache defaulting to the XDG cache location; the API gets a critical review as the backend of a minimal Electron photo app; the quak CLI is ported to this API in the same work. Short verdict on the review: the object model is right for the CLI and for scripts, but as accepted it would reproduce the lag you want to get rid of, for three reasons: every read waits on the network, there is no read across all photos at all, and objects with methods cannot cross from the Electron main process to the window. Proposed changes below.
Directory options
exportclassLibrary{staticopen(args:{client: Client;cacheDirectory?: string;// default: envPaths("quak", { suffix: "" }).cache + "/" + userID
downloadDirectory?: string;// default: none
onProgress?: ProgressCallback;}):Promise<Library>;backup(args?:{downloadDirectory?: string;// overrides the one given to open()
includeOriginals?: boolean;includeThumbnails?: boolean;onlyAlbumNames?: string[];onProgress?: ProgressCallback;}):Promise<BackupResult>;}
Default cache location, from env-paths 4.0.0 as bin/quak.ts already calls it (envPaths("quak", { suffix: "" }), today used only for paths.data/session.json):
Linux: $XDG_CACHE_HOME/quak, else ~/.cache/quak
macOS: ~/Library/Caches/quak
Windows: %LOCALAPPDATA%\quak\Cache
plus a /<userID> subdirectory (client.whoami().userID), so two accounts never share one metadata.json.
cacheDirectory holds only what can be deleted and fetched again, 0700 directory and 0600 files (it contains decrypted titles, GPS and file keys):
metadata.json (albums, files, change cursors)
failures.json (losing it loses only attempt counts)
thumbnails/<fileID>.jpg
originals/<fileID>.<ext>: originals fetched for viewing (photo.original())
downloadDirectory holds the backup the user keeps, in today's runBackup layout, unchanged: originals/<fileID>.<ext>, the originals/<fileID>.json sidecars, collections/<name>/ symlinks, collections/<name>.json. The sidecars make it readable without the cache.
How the two interact:
backup() writes originals to downloadDirectory. An original already in the cache is copied, not downloaded again.
photo.original() returns the downloadDirectory copy if there is one, else fetches into the cache.
backup() with no downloadDirectory (neither in open() nor in the call) throws before any network traffic. It does not fall back to the cache: a backup in a directory that cache cleaners delete is not a backup.
Nothing else in the accepted surface changes signature.
Critical review as an Electron backend
Startup: reads must not wait for the server
Accepted behaviour: every read first asks the server for changes. In a GUI that puts a network round trip in front of every view change, which is the lag you are describing. JSON.parse of a 100MB metadata.json also blocks its thread for about a second, and rewriting it blocks again; in the Electron main process that freezes every window.
Change: a mode on open(). The CLI keeps today's ruling; the GUI answers from RAM at once and is told when a refresh changed something. Still no sync() the caller must call.
In "inBackground": open() returns as soon as metadata.json is loaded (first paint is local only), a refresh starts at once and repeats every 60 seconds, and changes arrive through subscribe (below). On a first run with no metadata.json, reads return what has arrived so far and change events stream the rest in.
A refresh should be one request in the common case: /collections/v2 returns each album's updationTime, so only albums whose time advanced need a /collections/v2/diff call. Today's Client.listFiles always restarts at sinceTime: 0.
The app should run the library in an Electron utilityProcess, not in main, so the parse and the rewrite never block a window. That is the app's choice, but it forces the next point.
Process split: plain data and ids, not objects
Photo and Album have methods, and EnteFile carries key: Uint8Array. Across Electron IPC the methods are dropped, and file keys should never reach the window. One IPC call per photo (photo.thumbnail() 200 times per screen) is also too chatty.
Change: every capability exists as a call on Library that takes ids in arrays and returns plain records. Album and Photo stay as thin wrappers over those calls for in-process callers (the CLI, scripts).
exportinterfacePhotoRecord{// plain data, no keys, safe to send to the window
fileID: number;albumIDs: number[];// one file, every album holding it
title: string;// pubMagicMetadata.editedName, else metadata.title
takenAt: number;// ms; pubMagicMetadata.editedTime, else metadata.creationTime
fileType: FileType;caption?: string;width?: number;height?: number;latitude?: number;longitude?: number;isArchived: boolean;isHidden: boolean;thumbnailPath?: string;// set when the thumbnail is already in the cache
}readonlyphotos:{byID(args:{fileID: number}):Promise<Photo|undefined>;records(args:{fileIDs: number[]}):Promise<PhotoRecord[]>;};
This also fixes a defect in the accepted Photo: takenAt and title read only metadata, so a date or name the user corrected in Ente would be ignored and the timeline would sort wrongly.
thumbnailPath on the record means cached thumbnails paint with no further call. The library knows what is cached from one directory read at open().
Timeline: there is no read across all photos
The accepted surface has album.photos.list() and lib.photos.byID() only. A timeline would have to list every album and merge, and a file in three albums would appear three times (EnteFile is one row per album membership).
Photo[] for 50,000 photos is fine in RAM but wrong over IPC. A virtualized list needs two things: the count per group (for total scroll height) and the ids per group; then records only for the rows on screen.
exportinterfaceTimelineGroup{key: string;// "2025-08-04" (day), "2025-W32" (week, Monday start), "2025-08" (month)
startsAt: number;// ms, local time zone
fileIDs: number[];// newest first, each file once
}exportinterfacePhotoFilter{albumID?: number;text?: string;// case-insensitive substring of title, caption, album name
fileTypes?: FileType[];hasLocation?: boolean;includeArchived?: boolean;// default false; hidden photos are never included
}readonlytimeline:{groups(args:{groupBy:"day"|"week"|"month";filter?: PhotoFilter}):Promise<TimelineGroup[]>;};
50,000 ids is about 400KB and is built from RAM in tens of milliseconds, so the window re-requests it whole on every change. No paging, no query language; the filter runs library-side because the window does not hold the records.
Week-grouped timeline in use (main side; send is the app's IPC):
constlib=awaitLibrary.open({client,refresh:"inBackground"});constweeks=awaitlib.timeline.groups({groupBy:"week"});send("timeline",weeks.map((w)=>({key: w.key,count: w.fileIDs.length})));// window: row heights come from the counts; on scroll it asks for the visible weeks only
constvisible=weeks.slice(firstVisibleWeek,lastVisibleWeek+1).flatMap((w)=>w.fileIDs);send("records",awaitlib.photos.records({fileIDs: visible}));
Thumbnails at scroll speed
photo.thumbnail() is one request per call with no ordering, no limit on simultaneous requests and no way to stop. A fast scroll queues thousands of downloads for rows already gone.
Return paths, never bytes. Bytes over IPC are copied twice and bypass the browser's image decoding and caching. The app registers one custom scheme (protocol.handle) that serves files out of cacheDirectory/thumbnails; plain file:// is blocked for a page not itself loaded from file://.
readonlythumbnails:{ensure(args:{fileIDs: number[];priority:"visible"|"ahead"|"background";signal?: AbortSignal;// abort drops queued work; downloads in flight finish and are kept
onProgress?: ProgressCallback;// one downloadThumbnail done/skipped/failed event per file, as each lands
}):Promise<{fileID: number;path?: string;error?: string}[]>;};
One queue inside the library: 8 downloads at a time, visible before ahead before background, the same fileID requested twice is downloaded once.
Prefetch in use:
letscroll=newAbortController();asyncfunctiononVisibleRangeChanged(visibleIDs: number[],nextScreenIDs: number[]){scroll.abort();// the user scrolled on: drop what is still queued
scroll=newAbortController();constpaint: ProgressCallback=(ev)=>{if(ev.status==="done"||ev.status==="skipped")send("thumbnailReady",ev.photo!.fileID);};voidlib.thumbnails.ensure({fileIDs: visibleIDs,priority:"visible",signal: scroll.signal,onProgress: paint});voidlib.thumbnails.ensure({fileIDs: nextScreenIDs,priority:"ahead",signal: scroll.signal});}// once, after first paint: fill the whole cache quietly
voidlib.thumbnails.ensure({fileIDs: weeks.flatMap((w)=>w.fileIDs),priority:"background"});
For the full-size viewer: downloadFile holds the whole plaintext in RAM (streamDecrypt in src/download/index.ts), so opening a multi-gigabyte video can kill the process. Streaming decrypt to disk (#21) is a precondition for the GUI, not an extra.
Originals viewed in the GUI accumulate in the cache without limit; see the open questions.
Magic metadata and search
EnteFile.magicMetadata and pubMagicMetadata are Record<string, unknown> (src/model/types.ts); the accepted Photo exposes none of it. The fields a GUI needs are promoted onto PhotoRecord above (caption, editedTime, editedName, w, h from the public layer; visibility from the private one). Unverified: those are Ente's field names as I know them; confirm against the fixtures when implementing.
Search is PhotoFilter.text, in RAM. Nothing more is needed at 50,000 records.
Machine-learning data is a different thing. src/metadata-backup.ts already fetches it (/files/data/fetch, type mldata, 200 ids per request): face detections and CLIP embeddings per file. Ente stores no tag strings server-side; its "magic search" embeds the typed text on the device and compares. Named people live behind an Ente endpoint quak does not implement. Embeddings for 50,000 files exceed 100MB as JSON, so they must never go into metadata.json; if wanted, a separate mldata.json loaded on demand.
Change notification
exportinterfaceLibraryChange{albumIDsChanged: number[];fileIDsChanged: number[];fileIDsRemoved: number[];serverReachable: boolean;// false: answering from the local copy
refreshedAt: number;}subscribe(args:{onChange:(change: LibraryChange)=>void}):{unsubscribe():void};
One more conflict: #36 (comment) proposed a lock file, one process per directory. With a GUI open all day, a scheduled quak backup on the same cache would always fail. Recommend no lock: metadata.json carries its own cursors and is replaced by rename, so whichever process writes last leaves a consistent file and the other re-asks the server for a few changes.
CLI
The quak CLI is ported to this API in the same implementation; commands, flags, output and exit codes stay as they are, plus one new global option --cache-dir.
collections: lib.albums.list(). files --collection: lib.albums.byID() then album.photos.list()
get / get-thumb: lib.photos.byID() then photo.original() / photo.thumbnail(), copied to --out (default as today). This removes today's scan of every album per call; --collection is accepted and ignored.
backup <dir>: Library.open({ client, downloadDirectory: dir }), lib.backup({ onProgress }); the typed events are formatted into today's log lines; --json and exit 1 on failures unchanged.
backup-metadata <dir> and helper *-missing-thumbnails: enumerate through lib.albums / album.photos instead of Client.listCollections / listFiles; the ML fetch, EXIF extraction and thumbnail upload stay where they are. --exif reads photo.original() instead of a temporary download.
login, whoami, logout: unchanged.
Open questions
"Magic tags": do you mean the magic metadata fields (caption, archived, edited date), or ML results (faces, search by content)? Recommend the first now; ML as its own issue.
Originals cached by viewing: recommend a size limit with least-recently-used removal, default 10 GiB, as open() option cacheOriginalsMaxBytes; thumbnails are never removed.
An optional lib.refresh() for a reload button, never required: recommend yes.
No lock file, last writer wins (above): recommend yes.
Disclosures: posted with the tea CLI as clawbot (the Gitea MCP tool was not available). Nothing was built or run; Electron behaviour (protocol.handle, utilityProcess, IPC cloning) is from its documentation, not tested here. The 8-download limit, the 60-second interval and the 10 GiB default are judgement calls.
Model: fable-5-1
Reading of your three orders: both directories become options with the cache defaulting to the XDG cache location; the API gets a critical review as the backend of a minimal Electron photo app; the `quak` CLI is ported to this API in the same work. Short verdict on the review: the object model is right for the CLI and for scripts, but as accepted it would reproduce the lag you want to get rid of, for three reasons: every read waits on the network, there is no read across all photos at all, and objects with methods cannot cross from the Electron main process to the window. Proposed changes below.
## Directory options
```ts
export class Library {
static open(args: {
client: Client;
cacheDirectory?: string; // default: envPaths("quak", { suffix: "" }).cache + "/" + userID
downloadDirectory?: string; // default: none
onProgress?: ProgressCallback;
}): Promise<Library>;
backup(args?: {
downloadDirectory?: string; // overrides the one given to open()
includeOriginals?: boolean;
includeThumbnails?: boolean;
onlyAlbumNames?: string[];
onProgress?: ProgressCallback;
}): Promise<BackupResult>;
}
```
- Default cache location, from `env-paths` 4.0.0 as `bin/quak.ts` already calls it (`envPaths("quak", { suffix: "" })`, today used only for `paths.data/session.json`):
- Linux: `$XDG_CACHE_HOME/quak`, else `~/.cache/quak`
- macOS: `~/Library/Caches/quak`
- Windows: `%LOCALAPPDATA%\quak\Cache`
- plus a `/<userID>` subdirectory (`client.whoami().userID`), so two accounts never share one `metadata.json`.
- `cacheDirectory` holds only what can be deleted and fetched again, `0700` directory and `0600` files (it contains decrypted titles, GPS and file keys):
- `metadata.json` (albums, files, change cursors)
- `failures.json` (losing it loses only attempt counts)
- `thumbnails/<fileID>.jpg`
- `originals/<fileID>.<ext>`: originals fetched for viewing (`photo.original()`)
- `downloadDirectory` holds the backup the user keeps, in today's `runBackup` layout, unchanged: `originals/<fileID>.<ext>`, the `originals/<fileID>.json` sidecars, `collections/<name>/` symlinks, `collections/<name>.json`. The sidecars make it readable without the cache.
- How the two interact:
- `backup()` writes originals to `downloadDirectory`. An original already in the cache is copied, not downloaded again.
- `photo.original()` returns the `downloadDirectory` copy if there is one, else fetches into the cache.
- `backup()` with no `downloadDirectory` (neither in `open()` nor in the call) throws before any network traffic. It does not fall back to the cache: a backup in a directory that cache cleaners delete is not a backup.
- Nothing else in the accepted surface changes signature.
## Critical review as an Electron backend
### Startup: reads must not wait for the server
- Accepted behaviour: every read first asks the server for changes. In a GUI that puts a network round trip in front of every view change, which is the lag you are describing. `JSON.parse` of a 100MB `metadata.json` also blocks its thread for about a second, and rewriting it blocks again; in the Electron main process that freezes every window.
- Change: a mode on `open()`. The CLI keeps today's ruling; the GUI answers from RAM at once and is told when a refresh changed something. Still no `sync()` the caller must call.
```ts
static open(args: {
// ...as above
refresh?: "beforeEachRead" | "inBackground"; // default "beforeEachRead"
}): Promise<Library>;
```
- In `"inBackground"`: `open()` returns as soon as `metadata.json` is loaded (first paint is local only), a refresh starts at once and repeats every 60 seconds, and changes arrive through `subscribe` (below). On a first run with no `metadata.json`, reads return what has arrived so far and change events stream the rest in.
- A refresh should be one request in the common case: `/collections/v2` returns each album's `updationTime`, so only albums whose time advanced need a `/collections/v2/diff` call. Today's `Client.listFiles` always restarts at `sinceTime: 0`.
- The app should run the library in an Electron `utilityProcess`, not in main, so the parse and the rewrite never block a window. That is the app's choice, but it forces the next point.
### Process split: plain data and ids, not objects
- `Photo` and `Album` have methods, and `EnteFile` carries `key: Uint8Array`. Across Electron IPC the methods are dropped, and file keys should never reach the window. One IPC call per photo (`photo.thumbnail()` 200 times per screen) is also too chatty.
- Change: every capability exists as a call on `Library` that takes ids in arrays and returns plain records. `Album` and `Photo` stay as thin wrappers over those calls for in-process callers (the CLI, scripts).
```ts
export interface PhotoRecord { // plain data, no keys, safe to send to the window
fileID: number;
albumIDs: number[]; // one file, every album holding it
title: string; // pubMagicMetadata.editedName, else metadata.title
takenAt: number; // ms; pubMagicMetadata.editedTime, else metadata.creationTime
fileType: FileType;
caption?: string;
width?: number;
height?: number;
latitude?: number;
longitude?: number;
isArchived: boolean;
isHidden: boolean;
thumbnailPath?: string; // set when the thumbnail is already in the cache
}
readonly photos: {
byID(args: { fileID: number }): Promise<Photo | undefined>;
records(args: { fileIDs: number[] }): Promise<PhotoRecord[]>;
};
```
- This also fixes a defect in the accepted `Photo`: `takenAt` and `title` read only `metadata`, so a date or name the user corrected in Ente would be ignored and the timeline would sort wrongly.
- `thumbnailPath` on the record means cached thumbnails paint with no further call. The library knows what is cached from one directory read at `open()`.
### Timeline: there is no read across all photos
- The accepted surface has `album.photos.list()` and `lib.photos.byID()` only. A timeline would have to list every album and merge, and a file in three albums would appear three times (`EnteFile` is one row per album membership).
- `Photo[]` for 50,000 photos is fine in RAM but wrong over IPC. A virtualized list needs two things: the count per group (for total scroll height) and the ids per group; then records only for the rows on screen.
```ts
export interface TimelineGroup {
key: string; // "2025-08-04" (day), "2025-W32" (week, Monday start), "2025-08" (month)
startsAt: number; // ms, local time zone
fileIDs: number[]; // newest first, each file once
}
export interface PhotoFilter {
albumID?: number;
text?: string; // case-insensitive substring of title, caption, album name
fileTypes?: FileType[];
hasLocation?: boolean;
includeArchived?: boolean; // default false; hidden photos are never included
}
readonly timeline: {
groups(args: { groupBy: "day" | "week" | "month"; filter?: PhotoFilter }): Promise<TimelineGroup[]>;
};
```
- 50,000 ids is about 400KB and is built from RAM in tens of milliseconds, so the window re-requests it whole on every change. No paging, no query language; the filter runs library-side because the window does not hold the records.
Week-grouped timeline in use (main side; `send` is the app's IPC):
```ts
const lib = await Library.open({ client, refresh: "inBackground" });
const weeks = await lib.timeline.groups({ groupBy: "week" });
send("timeline", weeks.map((w) => ({ key: w.key, count: w.fileIDs.length })));
// window: row heights come from the counts; on scroll it asks for the visible weeks only
const visible = weeks.slice(firstVisibleWeek, lastVisibleWeek + 1).flatMap((w) => w.fileIDs);
send("records", await lib.photos.records({ fileIDs: visible }));
```
### Thumbnails at scroll speed
- `photo.thumbnail()` is one request per call with no ordering, no limit on simultaneous requests and no way to stop. A fast scroll queues thousands of downloads for rows already gone.
- Return paths, never bytes. Bytes over IPC are copied twice and bypass the browser's image decoding and caching. The app registers one custom scheme (`protocol.handle`) that serves files out of `cacheDirectory/thumbnails`; plain `file://` is blocked for a page not itself loaded from `file://`.
```ts
readonly thumbnails: {
ensure(args: {
fileIDs: number[];
priority: "visible" | "ahead" | "background";
signal?: AbortSignal; // abort drops queued work; downloads in flight finish and are kept
onProgress?: ProgressCallback; // one downloadThumbnail done/skipped/failed event per file, as each lands
}): Promise<{ fileID: number; path?: string; error?: string }[]>;
};
```
- One queue inside the library: 8 downloads at a time, `visible` before `ahead` before `background`, the same `fileID` requested twice is downloaded once.
Prefetch in use:
```ts
let scroll = new AbortController();
async function onVisibleRangeChanged(visibleIDs: number[], nextScreenIDs: number[]) {
scroll.abort(); // the user scrolled on: drop what is still queued
scroll = new AbortController();
const paint: ProgressCallback = (ev) => {
if (ev.status === "done" || ev.status === "skipped") send("thumbnailReady", ev.photo!.fileID);
};
void lib.thumbnails.ensure({ fileIDs: visibleIDs, priority: "visible", signal: scroll.signal, onProgress: paint });
void lib.thumbnails.ensure({ fileIDs: nextScreenIDs, priority: "ahead", signal: scroll.signal });
}
// once, after first paint: fill the whole cache quietly
void lib.thumbnails.ensure({ fileIDs: weeks.flatMap((w) => w.fileIDs), priority: "background" });
```
- For the full-size viewer: `downloadFile` holds the whole plaintext in RAM (`streamDecrypt` in `src/download/index.ts`), so opening a multi-gigabyte video can kill the process. Streaming decrypt to disk (https://git.eeqj.de/sneak/quak/issues/21) is a precondition for the GUI, not an extra.
- Originals viewed in the GUI accumulate in the cache without limit; see the open questions.
### Magic metadata and search
- `EnteFile.magicMetadata` and `pubMagicMetadata` are `Record<string, unknown>` (`src/model/types.ts`); the accepted `Photo` exposes none of it. The fields a GUI needs are promoted onto `PhotoRecord` above (`caption`, `editedTime`, `editedName`, `w`, `h` from the public layer; `visibility` from the private one). Unverified: those are Ente's field names as I know them; confirm against the fixtures when implementing.
- Search is `PhotoFilter.text`, in RAM. Nothing more is needed at 50,000 records.
- Machine-learning data is a different thing. `src/metadata-backup.ts` already fetches it (`/files/data/fetch`, type `mldata`, 200 ids per request): face detections and CLIP embeddings per file. Ente stores no tag strings server-side; its "magic search" embeds the typed text on the device and compares. Named people live behind an Ente endpoint quak does not implement. Embeddings for 50,000 files exceed 100MB as JSON, so they must never go into `metadata.json`; if wanted, a separate `mldata.json` loaded on demand.
### Change notification
```ts
export interface LibraryChange {
albumIDsChanged: number[];
fileIDsChanged: number[];
fileIDsRemoved: number[];
serverReachable: boolean; // false: answering from the local copy
refreshedAt: number;
}
subscribe(args: { onChange: (change: LibraryChange) => void }): { unsubscribe(): void };
```
```ts
lib.subscribe({
onChange: async (c) => {
if (c.fileIDsChanged.length + c.fileIDsRemoved.length === 0) return;
send("timeline", (await lib.timeline.groups({ groupBy: "week" })).map((w) => ({ key: w.key, count: w.fileIDs.length })));
},
});
```
- One more conflict: https://git.eeqj.de/sneak/quak/issues/36#issuecomment-95524 proposed a lock file, one process per directory. With a GUI open all day, a scheduled `quak backup` on the same cache would always fail. Recommend no lock: `metadata.json` carries its own cursors and is replaced by rename, so whichever process writes last leaves a consistent file and the other re-asks the server for a few changes.
## CLI
The `quak` CLI is ported to this API in the same implementation; commands, flags, output and exit codes stay as they are, plus one new global option `--cache-dir`.
- `collections`: `lib.albums.list()`. `files --collection`: `lib.albums.byID()` then `album.photos.list()`
- `get` / `get-thumb`: `lib.photos.byID()` then `photo.original()` / `photo.thumbnail()`, copied to `--out` (default as today). This removes today's scan of every album per call; `--collection` is accepted and ignored.
- `backup <dir>`: `Library.open({ client, downloadDirectory: dir })`, `lib.backup({ onProgress })`; the typed events are formatted into today's log lines; `--json` and exit 1 on failures unchanged.
- `backup-metadata <dir>` and `helper *-missing-thumbnails`: enumerate through `lib.albums` / `album.photos` instead of `Client.listCollections` / `listFiles`; the ML fetch, EXIF extraction and thumbnail upload stay where they are. `--exif` reads `photo.original()` instead of a temporary download.
- `login`, `whoami`, `logout`: unchanged.
## Open questions
- "Magic tags": do you mean the magic metadata fields (caption, archived, edited date), or ML results (faces, search by content)? Recommend the first now; ML as its own issue.
- Originals cached by viewing: recommend a size limit with least-recently-used removal, default 10 GiB, as `open()` option `cacheOriginalsMaxBytes`; thumbnails are never removed.
- An optional `lib.refresh()` for a reload button, never required: recommend yes.
- No lock file, last writer wins (above): recommend yes.
Disclosures: posted with the `tea` CLI as `clawbot` (the Gitea MCP tool was not available). Nothing was built or run; Electron behaviour (`protocol.handle`, `utilityProcess`, IPC cloning) is from its documentation, not tested here. The 8-download limit, the 60-second interval and the 10 GiB default are judgement calls.
Model: fable-5-1
magic tags means ML results - "magic" refers to ente's search by content feature.
10GiB cache is tiny. 100GiB is fine, disks are big. we can adapt it downward dynamically LRU if the disk has less than 50GiB free.
no lock file.
i would like a TTL of our local cache, so on launch we fetch things, but then we don't hit the network at all again for 1-3 seconds, so that subsequent calls serve entirely locally. this should simplify the UI side of things.
i want to use a lot of ram and disk and cpu in exchange for no lag ever. it's ok to load all of the metadata into ram on both sides of the IPC on startup, and to use dozens of gigabytes of disk space for local cache. i'd like to perhaps even aggressively locally precache thumbnails up to and including 100% of the account, for maximum local speed. we should probably locally precache originals for the latest week of items too, and all items in the favorites album.
let's move all of the caching/ttl/etc stuff into the library so the UI doesn't have to know much of anything about it.
magic tags means ML results - "magic" refers to ente's search by content feature.
10GiB cache is tiny. 100GiB is fine, disks are big. we can adapt it downward dynamically LRU if the disk has less than 50GiB free.
no lock file.
i would like a TTL of our local cache, so on launch we fetch things, but then we don't hit the network at all again for 1-3 seconds, so that subsequent calls serve entirely locally. this should simplify the UI side of things.
i want to use a lot of ram and disk and cpu in exchange for no lag ever. it's ok to load all of the metadata into ram on both sides of the IPC on startup, and to use dozens of gigabytes of disk space for local cache. i'd like to perhaps even aggressively locally precache thumbnails up to and including 100% of the account, for maximum local speed. we should probably locally precache originals for the latest week of items too, and all items in the favorites album.
let's move all of the caching/ttl/etc stuff into the library so the UI doesn't have to know much of anything about it.
suggests that all queries would hit the server for freshness. i don't want that. we can be 3-5 seconds behind the server for recent images, and 5-10 seconds behind the server for lists of albums, and probably 10 seconds behind the server for album membership.
i want separate parallelization for metadata fetches vs content fetches vs thumbnail fetches. metadata fetches should have max 10 req parallel, content fetches should have max 5, and thumbnail fetches should have max 25.
> serverReachable
suggests that all queries would hit the server for freshness. i don't want that. we can be 3-5 seconds behind the server for recent images, and 5-10 seconds behind the server for lists of albums, and probably 10 seconds behind the server for album membership.
i want separate parallelization for metadata fetches vs content fetches vs thumbnail fetches. metadata fetches should have max 10 req parallel, content fetches should have max 5, and thumbnail fetches should have max 25.
Revised design after #36 (comment) and #36 (comment). Reading of the seven rulings, then the design that follows from each. Everything accepted in #36 (comment) that is not named below stands unchanged (args objects everywhere, onProgress events, plain records over IPC, PhotoRecord, timeline.groups, JSON on disk, directory options, CLI port).
Rulings, as read:
Magic tags are the ML results (the data behind Ente's search by content). They get their own part of the library and their own files, never metadata.json.
Originals cache: 100 GiB, lowered on its own when the disk has under 50 GiB free, least-recently-used removal.
No lock file; the last process to write wins.
A TTL on the local copy: fetch on launch, then answer every read from RAM without touching the network until the TTL runs out. Allowed staleness: 3-5 s for recent images, 5-10 s for the album list, about 10 s for album membership. The refresh: "beforeEachRead" mode and the serverReachable flag are gone.
Spend RAM, disk and CPU for zero lag: all metadata in RAM on both sides of the IPC at startup, thumbnails precached up to 100% of the account, originals precached for the latest week of items and everything in the favorites album.
Caching, TTL, precache and eviction all live inside the library; the UI knows almost nothing about them.
Three separate request pools: metadata at most 10 in flight, content at most 5, thumbnails at most 25.
TTL and refresh (rulings 4 and 6)
open() loads metadata.json into RAM, then fetches: one /collections/v2 call with the stored cursor, then /collections/v2/diff from each album's stored cursor for the albums whose updationTime advanced, then the ML data (below). On a first run there is no local copy, so open() resolves only when this first fetch is complete. On later runs it resolves as soon as the local copy is loaded and the first refresh has been started, so a launch with the server down still paints at once.
After that a refresh runs every 3 s: one /collections/v2?sinceTime=cursor request (empty answer when nothing changed), and a diff only for the albums it names. Every read (albums.list, photos.list, photos.records, timeline.groups, snapshot) is answered from RAM, always, with no refresh in front of it. That meets all three bounds with one small request per 3 s: new images, the album list and membership are each at most 3 s plus one round trip behind. There is no per-kind timer because the per-album diff is only sent for albums the collections call reported as changed, so the three bounds collapse into one interval. This rests on the server advancing an album's updationTime when a file is added or removed from it, which is how Ente's own clients decide what to diff; unverified here, see the end.
metadata.json is rewritten (temp file plus rename, fsync before and after) only when a refresh changed something. A refresh that fails is not an error a read can see: reads keep answering from RAM, the failure is reported once through onProgress (operation: "refresh", status: "failed") and is visible in status(). LibraryChange loses serverReachable; open() loses refresh.
The interval is an option, refreshIntervalSeconds, default 3. A caller that wants the album list only every 10 s sets 10; there is no separate knob per kind. lib.refresh() is dropped: with a 3 s interval a reload button has nothing to do.
Grounding: Client.listCollections and Client.listFiles (src/client.ts) both start at sinceTime: 0 and drop isDeleted rows; the refresh needs the cursor-taking, tombstone-surfacing variants from the original plan (step 2). Every request in this section goes through the metadata pool.
exportclassLibrary{staticopen(args:{client: Client;cacheDirectory?: string;// default: env-paths cache dir + "/" + userID
downloadDirectory?: string;// default: none; required by backup()
refreshIntervalSeconds?: number;// default 3
onProgress?: ProgressCallback;// refresh, precache and ML fetch events for the life of the library
}):Promise<Library>;status():{lastRefreshAt?: number;// ms; absent until the first refresh completes
lastRefreshError?: string;// set while the most recent refresh failed
thumbnailsCached: number;// precache progress, also emitted through onProgress
thumbnailsTotal: number;originalsCached: number;originalsPinned: number;// latest week + favorites, never evicted
originalsBytes: number;originalsLimitBytes: number;// the limit in force now (see eviction)
};close():Promise<void>;// stops the refresh loop and the precache; in-flight downloads finish
}
All metadata in RAM on both sides (ruling 5)
The library already holds every album and file record in RAM. For the window, snapshot() returns the whole account as plain records in one call, and subscribe then delivers only what changed. The window keeps its own copy, groups and filters it locally, and never asks the library per row. 50,000 PhotoRecords are roughly 15 MB over structured clone; a one-time cost at startup.
PhotoRecord gains originalPath? next to thumbnailPath?: set when the original is in the cache or in downloadDirectory, so the viewer opens pinned and recently viewed photos with no call at all.
timeline.groups and PhotoFilter stay for in-process callers (CLI, scripts); a window holding the snapshot does not need them.
exportinterfaceAlbumRecord{collectionID: number;name: string;type:CollectionType;// "favorites" identifies the favorites album
isShared: boolean;updationTime: number;fileIDs: number[];// newest first
}exportinterfaceLibrarySnapshot{albums: AlbumRecord[];photos: PhotoRecord[];// one per fileID, newest first
takenAt: number;// ms; when this snapshot was built
}snapshot():LibrarySnapshot;// synchronous: RAM only
exportinterfaceLibraryChange{albumsChanged: AlbumRecord[];// full records, so the window replaces in place
photosChanged: PhotoRecord[];fileIDsRemoved: number[];albumIDsRemoved: number[];refreshedAt: number;}subscribe(args:{onChange:(change: LibraryChange)=>void}):{unsubscribe():void};
Startup in the app's library process:
constlib=awaitLibrary.open({client,onProgress:(ev)=>log(ev)});send("snapshot",lib.snapshot());// the window now has everything
lib.subscribe({onChange:(c)=>send("change",c)});// and stays current within the TTL
Precache (rulings 5 and 6)
Both precaches start inside open() and need nothing from the caller. They report through the onProgress given to open() and through status().
Thumbnails: every file in the account, newest first, through the thumbnail pool, until 100% is on disk. Thumbnails are never evicted. The queue is shared with thumbnails.ensure: visible and ahead requests go ahead of the precache, so scrolling is never behind the background fill. A file whose thumbnail already exists costs one Set lookup, from the directory listing taken at open().
Originals, through the content pool, in this order: every file in the favorites album (Collection.type === "favorites"), then every file whose takenAt lies within the 7 days ending at the newest takenAt in the account. These are the pinned set. When a refresh adds a favorite or a newer file, it joins the queue; when the week window moves or a favorite is removed, the file leaves the pinned set and becomes an ordinary cached original.
photos.byID().original() and thumbnails.ensure keep their shapes; original() on a cached or pinned file returns the path with one skipped event and no network.
Grounding: runBackup and fetchMLDataForFiles are sequential today; downloadFile buffers the whole plaintext (streamDecrypt, src/download/index.ts), so 5 concurrent originals means 5 files in RAM at once and a 4 GB video on each. Streaming decrypt to disk (#21) is a precondition for the originals precache, as it was for the viewer.
staticopen(args:{// ...as above
precacheThumbnails?: boolean;// default true
precacheOriginals?: boolean;// default true; the favorites album and the latest week
precacheOriginalsDays?: number;// default 7
}):Promise<Library>;
Originals cache limit and eviction (ruling 2)
Limit: cacheOriginalsMaxBytes, default 100 GiB. Before each original is written the library checks the volume holding cacheDirectory (fs.statfs, Node 18.15+, this repo is on 22) and uses the lower of the configured limit and bytesUsedByOriginals + bytesFree - 50 GiB. So on a full disk the limit falls on its own and rises again as space returns; status().originalsLimitBytes shows the value in force.
When the next write would cross the limit, originals are removed least-recently-used first until it fits. Pinned files are skipped; if only pinned files remain the write still happens and the cache is over the limit until the pinned set shrinks. freeBelowBytes (default 50 GiB) is the second option; both are open() options.
Last use is the file's mtime: the library touches the file (utimes) when a read returns its path, so the order survives restarts and needs no ledger. atime is not used because most volumes mount noatime or relatime.
Only cacheDirectory/originals is ever evicted. downloadDirectory is the backup; it is never counted and never touched.
No lock file (ruling 3)
No lock and no detection. metadata.json carries its own cursors and is replaced by one rename, so whichever process writes last leaves a consistent file; the other keeps running on its RAM copy and its next refresh re-asks the server from its own cursor. The same holds for the ML files and for eviction: two processes evicting at once each remove least-recently-used files from the directory they both see, and the worst case is a file removed that the other would have kept, which is refetched on next use.
ML data, the magic tags (ruling 1)
What exists: src/metadata-backup.ts fetches /files/data/fetch with type: "mldata", 200 ids per request, decrypts each entry with the file key and gunzips it. The payload (test/cli/metadata-backup.test.ts) is { face: { faces: [{ faceID, detection: { box, landmarks }, score, blur, embedding }] }, clip: { embedding } }. Nothing else server-side: Ente keeps no tag strings; search by content embeds the typed text on the device and ranks files by similarity to clip.embedding.
Storage, its own part of the cache, never in metadata.json:
mldata/<fileID>.json: the decrypted, gunzipped payload as fetched, one file per fileID, written by rename. Present means complete, the same rule as originals.
mldata/clip.f32 plus mldata/clip.json: the derived index the search runs on. clip.json lists the fileIDs in order and the embedding length; clip.f32 is those embeddings as one Float32Array, so 50,000 files at 512 floats load in one read as 100 MB and no parse. Rebuilt from the per-file JSON when missing or when it disagrees with the files present, appended to when new payloads arrive.
In RAM: the Float32Array and the id list. Per-file payloads (face boxes, landmarks) are read from disk on demand, not held.
Fetched through the metadata pool: after each refresh, every fileID present in metadata.json and absent from mldata/ is requested, 200 per call, 10 calls in flight. That is the whole account on the first run and only new files afterwards. A file whose updationTime advanced is refetched, because the ML payload is rewritten when a client re-runs indexing. The fetch is reported through onProgress (operation: "fetchMLData") and status().
Surface:
exportinterfaceMLData{// the payload as Ente's clients wrote it
face?:{faces:{faceID: string;detection:{box:{x: number;y: number;width: number;height: number}};score: number;blur: number;embedding: number[]}[]};clip?:{embedding: number[]};}readonlymldata:{forFile(args:{fileID: number}):Promise<MLData|undefined>;similar(args:{fileID: number;limit?: number}):{fileID: number;score: number}[];// nearest by clip embedding; RAM only
searchByEmbedding(args:{embedding: number[];limit?: number}):{fileID: number;score: number}[];// cosine similarity over the index; RAM only
};
Search by content itself needs a text encoder that turns the typed words into an embedding in the same space as clip.embedding. quak has none. searchByEmbedding takes the vector so the app, or a later addition to the library, can supply it; see the open questions.
Three request pools (ruling 7)
Inside the library, three independent queues; each request is assigned by what it fetches:
metadata, at most 10 in flight: /collections/v2, /collections/v2/diff, /files/data/fetch
content, at most 5: getFileStream (originals, for precache, viewer and backup)
thumbnails, at most 25: getThumbnailStream
Within a pool, on-demand work (a viewer opening a photo, visible thumbnails) goes before precache and backup work; the same fileID requested twice is fetched once. A pool that is idle does not lend its slots to another.
The limits are open() options metadataConcurrency, contentConcurrency, thumbnailConcurrency, defaults 10, 5, 25. Grounding: ApiClient has no concurrency limit and none of its callers run more than one request at a time; the pools wrap calls into Client rather than changing ApiClient, so the low-level layer stays as it is and quak backup gets 5 parallel downloads through lib.backup() without a change of its own.
The retry policy is unchanged and runs inside a slot: a file on its third attempt still holds one of the 5.
What the UI is left knowing
Open, take the snapshot, subscribe, tell the library which thumbnails are on screen (thumbnails.ensure with visible and ahead), ask for an original's path when the viewer opens one, and hand the library a query embedding for search. It never sees a cursor, a TTL, a limit or a pool.
Judgement calls and open questions
One refresh interval for all three bounds, not three timers. The 3 s collections poll is the cheapest way to meet all of them; if the server does not advance an album's updationTime on file add or remove, the fallback is a diff of every album every 10 s, which is O(albums) small requests per cycle. Please confirm which it is if you know; otherwise the implementation checks it against the live API first.
Recommend open() not waiting for the first refresh when a local copy exists (above). The alternative, waiting, delays first paint by a round trip and by up to the 30 s request deadline when offline.
"Latest week of items" is read as the 7 days ending at the newest takenAt in the account, not the last 7 calendar days, so an account with no new uploads still has a week of originals pinned. Say if you meant upload time or calendar days.
The 50,000-photo assumptions behind the sizes (15 MB snapshot, 100 MB CLIP index) are mine; with 500,000 photos every number here is ten times larger and still within the ruling, but the metadata.json rewrite would take about 10 s of CPU per change, which argues for splitting it per album at that scale.
Search by content needs a text encoder. Options: the library bundles the CLIP text model and an ONNX runtime (a hash-pinned model download of tens of MB and a native dependency; search then works in the CLI too), or the app supplies the embedding through searchByEmbedding (nothing new in quak; the app carries the model). Recommend the second now, the first as its own issue once the app exists.
Named people: Ente's person groups (which faces belong to whom) live behind an endpoint quak does not implement, and I have not verified its shape. Faces are exposed as detections only; grouping by face embedding in the library is possible without the server and is a separate issue.
ML refresh: the design refetches ML data only for files whose updationTime advanced. If Ente re-indexes without touching the file, that change is missed until the per-file JSON is deleted; a cheaper check would need a server endpoint I have not confirmed exists.
Disclosures: posted through the Gitea REST API as clawbot; the Gitea MCP tool was not loaded in this session. Nothing was built or run. Not checked: the server's updationTime behaviour above, the ML payload shape beyond the test fixture, and fs.statfs on Windows.
Model: fable-5-1
Revised design after https://git.eeqj.de/sneak/quak/issues/36#issuecomment-95639 and https://git.eeqj.de/sneak/quak/issues/36#issuecomment-95640. Reading of the seven rulings, then the design that follows from each. Everything accepted in https://git.eeqj.de/sneak/quak/issues/36#issuecomment-95602 that is not named below stands unchanged (args objects everywhere, `onProgress` events, plain records over IPC, `PhotoRecord`, `timeline.groups`, JSON on disk, directory options, CLI port).
Rulings, as read:
1. Magic tags are the ML results (the data behind Ente's search by content). They get their own part of the library and their own files, never `metadata.json`.
2. Originals cache: 100 GiB, lowered on its own when the disk has under 50 GiB free, least-recently-used removal.
3. No lock file; the last process to write wins.
4. A TTL on the local copy: fetch on launch, then answer every read from RAM without touching the network until the TTL runs out. Allowed staleness: 3-5 s for recent images, 5-10 s for the album list, about 10 s for album membership. The `refresh: "beforeEachRead"` mode and the `serverReachable` flag are gone.
5. Spend RAM, disk and CPU for zero lag: all metadata in RAM on both sides of the IPC at startup, thumbnails precached up to 100% of the account, originals precached for the latest week of items and everything in the favorites album.
6. Caching, TTL, precache and eviction all live inside the library; the UI knows almost nothing about them.
7. Three separate request pools: metadata at most 10 in flight, content at most 5, thumbnails at most 25.
## TTL and refresh (rulings 4 and 6)
- `open()` loads `metadata.json` into RAM, then fetches: one `/collections/v2` call with the stored cursor, then `/collections/v2/diff` from each album's stored cursor for the albums whose `updationTime` advanced, then the ML data (below). On a first run there is no local copy, so `open()` resolves only when this first fetch is complete. On later runs it resolves as soon as the local copy is loaded and the first refresh has been started, so a launch with the server down still paints at once.
- After that a refresh runs every 3 s: one `/collections/v2?sinceTime=cursor` request (empty answer when nothing changed), and a diff only for the albums it names. Every read (`albums.list`, `photos.list`, `photos.records`, `timeline.groups`, `snapshot`) is answered from RAM, always, with no refresh in front of it. That meets all three bounds with one small request per 3 s: new images, the album list and membership are each at most 3 s plus one round trip behind. There is no per-kind timer because the per-album diff is only sent for albums the collections call reported as changed, so the three bounds collapse into one interval. This rests on the server advancing an album's `updationTime` when a file is added or removed from it, which is how Ente's own clients decide what to diff; unverified here, see the end.
- `metadata.json` is rewritten (temp file plus rename, fsync before and after) only when a refresh changed something. A refresh that fails is not an error a read can see: reads keep answering from RAM, the failure is reported once through `onProgress` (`operation: "refresh"`, `status: "failed"`) and is visible in `status()`. `LibraryChange` loses `serverReachable`; `open()` loses `refresh`.
- The interval is an option, `refreshIntervalSeconds`, default 3. A caller that wants the album list only every 10 s sets 10; there is no separate knob per kind. `lib.refresh()` is dropped: with a 3 s interval a reload button has nothing to do.
- Grounding: `Client.listCollections` and `Client.listFiles` (`src/client.ts`) both start at `sinceTime: 0` and drop `isDeleted` rows; the refresh needs the cursor-taking, tombstone-surfacing variants from the original plan (step 2). Every request in this section goes through the metadata pool.
```ts
export class Library {
static open(args: {
client: Client;
cacheDirectory?: string; // default: env-paths cache dir + "/" + userID
downloadDirectory?: string; // default: none; required by backup()
refreshIntervalSeconds?: number; // default 3
onProgress?: ProgressCallback; // refresh, precache and ML fetch events for the life of the library
}): Promise<Library>;
status(): {
lastRefreshAt?: number; // ms; absent until the first refresh completes
lastRefreshError?: string; // set while the most recent refresh failed
thumbnailsCached: number; // precache progress, also emitted through onProgress
thumbnailsTotal: number;
originalsCached: number;
originalsPinned: number; // latest week + favorites, never evicted
originalsBytes: number;
originalsLimitBytes: number; // the limit in force now (see eviction)
};
close(): Promise<void>; // stops the refresh loop and the precache; in-flight downloads finish
}
```
## All metadata in RAM on both sides (ruling 5)
- The library already holds every album and file record in RAM. For the window, `snapshot()` returns the whole account as plain records in one call, and `subscribe` then delivers only what changed. The window keeps its own copy, groups and filters it locally, and never asks the library per row. 50,000 `PhotoRecord`s are roughly 15 MB over structured clone; a one-time cost at startup.
- `PhotoRecord` gains `originalPath?` next to `thumbnailPath?`: set when the original is in the cache or in `downloadDirectory`, so the viewer opens pinned and recently viewed photos with no call at all.
- `timeline.groups` and `PhotoFilter` stay for in-process callers (CLI, scripts); a window holding the snapshot does not need them.
```ts
export interface AlbumRecord {
collectionID: number;
name: string;
type: CollectionType; // "favorites" identifies the favorites album
isShared: boolean;
updationTime: number;
fileIDs: number[]; // newest first
}
export interface LibrarySnapshot {
albums: AlbumRecord[];
photos: PhotoRecord[]; // one per fileID, newest first
takenAt: number; // ms; when this snapshot was built
}
snapshot(): LibrarySnapshot; // synchronous: RAM only
export interface LibraryChange {
albumsChanged: AlbumRecord[]; // full records, so the window replaces in place
photosChanged: PhotoRecord[];
fileIDsRemoved: number[];
albumIDsRemoved: number[];
refreshedAt: number;
}
subscribe(args: { onChange: (change: LibraryChange) => void }): { unsubscribe(): void };
```
Startup in the app's library process:
```ts
const lib = await Library.open({ client, onProgress: (ev) => log(ev) });
send("snapshot", lib.snapshot()); // the window now has everything
lib.subscribe({ onChange: (c) => send("change", c) }); // and stays current within the TTL
```
## Precache (rulings 5 and 6)
Both precaches start inside `open()` and need nothing from the caller. They report through the `onProgress` given to `open()` and through `status()`.
- Thumbnails: every file in the account, newest first, through the thumbnail pool, until 100% is on disk. Thumbnails are never evicted. The queue is shared with `thumbnails.ensure`: `visible` and `ahead` requests go ahead of the precache, so scrolling is never behind the background fill. A file whose thumbnail already exists costs one `Set` lookup, from the directory listing taken at `open()`.
- Originals, through the content pool, in this order: every file in the favorites album (`Collection.type === "favorites"`), then every file whose `takenAt` lies within the 7 days ending at the newest `takenAt` in the account. These are the pinned set. When a refresh adds a favorite or a newer file, it joins the queue; when the week window moves or a favorite is removed, the file leaves the pinned set and becomes an ordinary cached original.
- `photos.byID().original()` and `thumbnails.ensure` keep their shapes; `original()` on a cached or pinned file returns the path with one `skipped` event and no network.
- Grounding: `runBackup` and `fetchMLDataForFiles` are sequential today; `downloadFile` buffers the whole plaintext (`streamDecrypt`, `src/download/index.ts`), so 5 concurrent originals means 5 files in RAM at once and a 4 GB video on each. Streaming decrypt to disk (https://git.eeqj.de/sneak/quak/issues/21) is a precondition for the originals precache, as it was for the viewer.
```ts
static open(args: {
// ...as above
precacheThumbnails?: boolean; // default true
precacheOriginals?: boolean; // default true; the favorites album and the latest week
precacheOriginalsDays?: number; // default 7
}): Promise<Library>;
```
## Originals cache limit and eviction (ruling 2)
- Limit: `cacheOriginalsMaxBytes`, default 100 GiB. Before each original is written the library checks the volume holding `cacheDirectory` (`fs.statfs`, Node 18.15+, this repo is on 22) and uses the lower of the configured limit and `bytesUsedByOriginals + bytesFree - 50 GiB`. So on a full disk the limit falls on its own and rises again as space returns; `status().originalsLimitBytes` shows the value in force.
- When the next write would cross the limit, originals are removed least-recently-used first until it fits. Pinned files are skipped; if only pinned files remain the write still happens and the cache is over the limit until the pinned set shrinks. `freeBelowBytes` (default 50 GiB) is the second option; both are `open()` options.
- Last use is the file's mtime: the library touches the file (`utimes`) when a read returns its path, so the order survives restarts and needs no ledger. `atime` is not used because most volumes mount `noatime` or `relatime`.
- Only `cacheDirectory/originals` is ever evicted. `downloadDirectory` is the backup; it is never counted and never touched.
## No lock file (ruling 3)
No lock and no detection. `metadata.json` carries its own cursors and is replaced by one rename, so whichever process writes last leaves a consistent file; the other keeps running on its RAM copy and its next refresh re-asks the server from its own cursor. The same holds for the ML files and for eviction: two processes evicting at once each remove least-recently-used files from the directory they both see, and the worst case is a file removed that the other would have kept, which is refetched on next use.
## ML data, the magic tags (ruling 1)
What exists: `src/metadata-backup.ts` fetches `/files/data/fetch` with `type: "mldata"`, 200 ids per request, decrypts each entry with the file key and gunzips it. The payload (`test/cli/metadata-backup.test.ts`) is `{ face: { faces: [{ faceID, detection: { box, landmarks }, score, blur, embedding }] }, clip: { embedding } }`. Nothing else server-side: Ente keeps no tag strings; search by content embeds the typed text on the device and ranks files by similarity to `clip.embedding`.
Storage, its own part of the cache, never in `metadata.json`:
- `mldata/<fileID>.json`: the decrypted, gunzipped payload as fetched, one file per fileID, written by rename. Present means complete, the same rule as originals.
- `mldata/clip.f32` plus `mldata/clip.json`: the derived index the search runs on. `clip.json` lists the fileIDs in order and the embedding length; `clip.f32` is those embeddings as one `Float32Array`, so 50,000 files at 512 floats load in one read as 100 MB and no parse. Rebuilt from the per-file JSON when missing or when it disagrees with the files present, appended to when new payloads arrive.
- In RAM: the `Float32Array` and the id list. Per-file payloads (face boxes, landmarks) are read from disk on demand, not held.
- Fetched through the metadata pool: after each refresh, every fileID present in `metadata.json` and absent from `mldata/` is requested, 200 per call, 10 calls in flight. That is the whole account on the first run and only new files afterwards. A file whose `updationTime` advanced is refetched, because the ML payload is rewritten when a client re-runs indexing. The fetch is reported through `onProgress` (`operation: "fetchMLData"`) and `status()`.
Surface:
```ts
export interface MLData { // the payload as Ente's clients wrote it
face?: { faces: { faceID: string; detection: { box: { x: number; y: number; width: number; height: number } }; score: number; blur: number; embedding: number[] }[] };
clip?: { embedding: number[] };
}
readonly mldata: {
forFile(args: { fileID: number }): Promise<MLData | undefined>;
similar(args: { fileID: number; limit?: number }): { fileID: number; score: number }[]; // nearest by clip embedding; RAM only
searchByEmbedding(args: { embedding: number[]; limit?: number }): { fileID: number; score: number }[]; // cosine similarity over the index; RAM only
};
```
Search by content itself needs a text encoder that turns the typed words into an embedding in the same space as `clip.embedding`. quak has none. `searchByEmbedding` takes the vector so the app, or a later addition to the library, can supply it; see the open questions.
## Three request pools (ruling 7)
- Inside the library, three independent queues; each request is assigned by what it fetches:
- metadata, at most 10 in flight: `/collections/v2`, `/collections/v2/diff`, `/files/data/fetch`
- content, at most 5: `getFileStream` (originals, for precache, viewer and backup)
- thumbnails, at most 25: `getThumbnailStream`
- Within a pool, on-demand work (a viewer opening a photo, `visible` thumbnails) goes before precache and backup work; the same fileID requested twice is fetched once. A pool that is idle does not lend its slots to another.
- The limits are `open()` options `metadataConcurrency`, `contentConcurrency`, `thumbnailConcurrency`, defaults 10, 5, 25. Grounding: `ApiClient` has no concurrency limit and none of its callers run more than one request at a time; the pools wrap calls into `Client` rather than changing `ApiClient`, so the low-level layer stays as it is and `quak backup` gets 5 parallel downloads through `lib.backup()` without a change of its own.
- The retry policy is unchanged and runs inside a slot: a file on its third attempt still holds one of the 5.
## What the UI is left knowing
Open, take the snapshot, subscribe, tell the library which thumbnails are on screen (`thumbnails.ensure` with `visible` and `ahead`), ask for an original's path when the viewer opens one, and hand the library a query embedding for search. It never sees a cursor, a TTL, a limit or a pool.
## Judgement calls and open questions
- One refresh interval for all three bounds, not three timers. The 3 s collections poll is the cheapest way to meet all of them; if the server does not advance an album's `updationTime` on file add or remove, the fallback is a diff of every album every 10 s, which is O(albums) small requests per cycle. Please confirm which it is if you know; otherwise the implementation checks it against the live API first.
- Recommend `open()` not waiting for the first refresh when a local copy exists (above). The alternative, waiting, delays first paint by a round trip and by up to the 30 s request deadline when offline.
- "Latest week of items" is read as the 7 days ending at the newest `takenAt` in the account, not the last 7 calendar days, so an account with no new uploads still has a week of originals pinned. Say if you meant upload time or calendar days.
- The 50,000-photo assumptions behind the sizes (15 MB snapshot, 100 MB CLIP index) are mine; with 500,000 photos every number here is ten times larger and still within the ruling, but the `metadata.json` rewrite would take about 10 s of CPU per change, which argues for splitting it per album at that scale.
- Search by content needs a text encoder. Options: the library bundles the CLIP text model and an ONNX runtime (a hash-pinned model download of tens of MB and a native dependency; search then works in the CLI too), or the app supplies the embedding through `searchByEmbedding` (nothing new in quak; the app carries the model). Recommend the second now, the first as its own issue once the app exists.
- Named people: Ente's person groups (which faces belong to whom) live behind an endpoint quak does not implement, and I have not verified its shape. Faces are exposed as detections only; grouping by face embedding in the library is possible without the server and is a separate issue.
- ML refresh: the design refetches ML data only for files whose `updationTime` advanced. If Ente re-indexes without touching the file, that change is missed until the per-file JSON is deleted; a cheaper check would need a server endpoint I have not confirmed exists.
Disclosures: posted through the Gitea REST API as `clawbot`; the Gitea MCP tool was not loaded in this session. Nothing was built or run. Not checked: the server's `updationTime` behaviour above, the ML payload shape beyond the test fixture, and `fs.statfs` on Windows.
Model: fable-5-1
Implementation issues filed for this design, broken into small ordered units,
starting from the pieces every later unit needs. They will be driven one
issue-to-PR unit at a time, each independently reviewed before it lands on next. The design here is treated as settled enough to start; amend any unit as
needed and later ones are fleshed out just before dispatch.
Order and dependencies:
Foundation (low-level; everything else builds on these)
#37 — carry fileSize/thumbSize/isDeleted through decryptFile.
#51 — rewrite backup on the library + durable failure ledger (after #46, #44). Subsumes #8; overlaps #22, #9.
#52 — port the quak CLI to the new API (after #51). Overlaps #17.
#53 — rewrite the README API reference (after #52). Is #13.
Older issues these subsume or overlap (#7, #8, #9, #17, #21, #22) are cross-
referenced in the units above and closed by their landing commits where fully
covered; #24 (retry follow-ups) and #10 (session/keys hardening) stay separate.
Starting now with the foundation: #37 and #39 first (independent files, in
parallel), then #38 after #37 and #40 after #39.
Model: opus-4-8
Implementation issues filed for this design, broken into small ordered units,
starting from the pieces every later unit needs. They will be driven one
issue-to-PR unit at a time, each independently reviewed before it lands on
`next`. The design here is treated as settled enough to start; amend any unit as
needed and later ones are fleshed out just before dispatch.
Order and dependencies:
Foundation (low-level; everything else builds on these)
1. https://git.eeqj.de/sneak/quak/issues/37 — carry `fileSize`/`thumbSize`/`isDeleted` through `decryptFile`.
2. https://git.eeqj.de/sneak/quak/issues/38 — resumable, tombstone-surfacing enumeration on `Client` (after #37). Subsumes https://git.eeqj.de/sneak/quak/issues/7.
3. https://git.eeqj.de/sneak/quak/issues/39 — download module: exported fsync atomic writer + per-chunk progress. Overlaps https://git.eeqj.de/sneak/quak/issues/22. (Independent of #37/#38 — can run alongside.)
4. https://git.eeqj.de/sneak/quak/issues/40 — stream decrypted downloads to disk (after #39). Subsumes https://git.eeqj.de/sneak/quak/issues/21.
Metadata cache + library core
5. https://git.eeqj.de/sneak/quak/issues/41 — on-disk JSON metadata store (after #37).
6. https://git.eeqj.de/sneak/quak/issues/42 — `Library.open` + background refresh loop + staleness window (after #41, #38).
7. https://git.eeqj.de/sneak/quak/issues/43 — plain-record snapshot + change subscription (after #42).
8. https://git.eeqj.de/sneak/quak/issues/44 — in-process read surface: albums, photos, timeline (after #43).
Content/thumbnail cache
9. https://git.eeqj.de/sneak/quak/issues/45 — three bounded request pools, 10/5/25 (after #39, #40).
10. https://git.eeqj.de/sneak/quak/issues/46 — on-disk content/thumbnail cache + per-photo fetch + thumbnail prefetch (after #45, #40, #44). Touches https://git.eeqj.de/sneak/quak/issues/9.
11. https://git.eeqj.de/sneak/quak/issues/47 — originals cache limit + LRU eviction (after #46).
12. https://git.eeqj.de/sneak/quak/issues/48 — precache thumbnails (all) + originals (favorites + latest week) (after #46, #47).
Machine-learning (magic) data
13. https://git.eeqj.de/sneak/quak/issues/49 — fetch/store ML data + build CLIP index (after #42, #45).
14. https://git.eeqj.de/sneak/quak/issues/50 — content-similarity search surface (after #49).
Backup, CLI, docs
15. https://git.eeqj.de/sneak/quak/issues/51 — rewrite backup on the library + durable failure ledger (after #46, #44). Subsumes https://git.eeqj.de/sneak/quak/issues/8; overlaps #22, #9.
16. https://git.eeqj.de/sneak/quak/issues/52 — port the `quak` CLI to the new API (after #51). Overlaps https://git.eeqj.de/sneak/quak/issues/17.
17. https://git.eeqj.de/sneak/quak/issues/53 — rewrite the README API reference (after #52). Is https://git.eeqj.de/sneak/quak/issues/13.
Older issues these subsume or overlap (#7, #8, #9, #17, #21, #22) are cross-
referenced in the units above and closed by their landing commits where fully
covered; #24 (retry follow-ups) and #10 (session/keys hardening) stay separate.
Starting now with the foundation: #37 and #39 first (independent files, in
parallel), then #38 after #37 and #40 after #39.
Model: opus-4-8
I would like some additional asynchronous API calls that can be used sort of like how you would use a synchronous fetch so you will still await them but they don't actually resolve until such time a round trip to the server has been made and the local cache has been updated. I want the usual case to be that a fetch kicks off a background retrieval and then responds from disk. as has been designed presently. But what that means is that it's going to get stale data. And in certain cases, you are going to want to lag while it makes a round trip so that you can be sure that you have the latest data. So in that case, I would like some additional methods available that will not resolve until such time it knows that the data is synchronized. The CLI tools should then use these methods.
I would like some additional asynchronous API calls that can be used sort of like how you would use a synchronous fetch so you will still await them but they don't actually resolve until such time a round trip to the server has been made and the local cache has been updated. I want the usual case to be that a fetch kicks off a background retrieval and then responds from disk. as has been designed presently. But what that means is that it's going to get stale data. And in certain cases, you are going to want to lag while it makes a round trip so that you can be sure that you have the latest data. So in that case, I would like some additional methods available that will not resolve until such time it knows that the data is synchronized. The CLI tools should then use these methods.
Reading of the amendment above, for the record and for the manager to fold into the design and the affected unit issues before they dispatch:
The default read methods stay exactly as designed: answer immediately from the local copy and kick off a background refresh — the caller may get data that is slightly behind the server.
ADDED: a parallel fresh variant of the read calls. Still awaited like any other call, but it resolves only after a round trip to the server has completed and the local copy has been updated — at resolve time the answer is guaranteed current.
The CLI tools use the fresh variants, not the defaults.
Manager: amend the read-surface and CLI unit issues accordingly (and any other unit this touches), then continue the normal one-unit-at-a-time flow.
(model: claude-fable-5)
Reading of the amendment above, for the record and for the manager to fold into the design and the affected unit issues before they dispatch:
- The default read methods stay exactly as designed: answer immediately from the local copy and kick off a background refresh — the caller may get data that is slightly behind the server.
- ADDED: a parallel fresh variant of the read calls. Still awaited like any other call, but it resolves only after a round trip to the server has completed and the local copy has been updated — at resolve time the answer is guaranteed current.
- The CLI tools use the fresh variants, not the defaults.
Manager: amend the read-surface and CLI unit issues accordingly (and any other unit this touches), then continue the normal one-unit-at-a-time flow.
(model: claude-fable-5)
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
This is a plan for your review, not a finished spec — please critique and send rework requests. It covers the two goals you set: (1) a higher-level API over two long-lived caches — a SQLite metadata cache and an on-disk content/thumbnail cache — and (2) a backup tool that is extremely reliable and built on that API. Design rule applied throughout: the simplest thing that is fit for purpose, fewest moving parts, no tunables the goals don't need.
Current state (grounded)
Clientholds only the session keys + token in memory.runBackup(src/backup.ts) andrunMetadataBackup(src/metadata-backup.ts) re-enumerate and re-decrypt the whole account on every run. The only state that survives a run is the files on disk.Client.listCollections()always calls/collections/v2withsinceTime: 0;listFiles()paginates/collections/v2/diffbut always restarts atsinceTime: 0, and both dropisDeletedtombstones. So thediff/hasMore/updationTimeloop is used only for within-run pagination, never to fetch just the changes, and deletions are invisible to callers.backup <dir>layout (the only persistent state):originals/<fileID>.<ext>(content) +originals/<fileID>.json(metadata sidecar);collections/<name>/<title>symlinks into../../originals;collections/<name>.json. Skip rule isexistsSync && size > 0— no integrity check. The sidecar is written only when absent, so edited metadata goes stale. Per-file download failures are logged, counted, and stepped over (errors[], exit 1 if any failed); that list is discarded at process exit.fileIDidentifies content (dedup across collections),(collectionID, fileID)is a membership,updationTime(µs) advances on any change. Your(collectionID, fileID, updationTime)key is exactly right.FileMetadata.hashcarries a plaintext content hash (optional — not on every file).RawEnteFile.info.fileSize/thumbSizeexist on the wire butdecryptFiledrops them (and dropsisDeleted).writeAtomic),TAG_FINALtruncation detection instreamDecrypt, the retry classifier (withRetry/isRetryable/isSafeToReplay), per-request and per-body deadlines.Design — the cache and its API
Where it lives. Inside the backup target directory (recommended): the SQLite file and the content/thumbnail blobs are subdirectories of the user's
backup <dir>. One location, self-contained and portable, multiple independent mirrors, nothing hidden. (Alternative in the open questions: one managed store underenv-paths.) "Long-lived" is a property of invalidation, not location: both caches are invalidated only by the server's own change feed, so they carry no TTL and can persist indefinitely wherever they live.The SQLite metadata cache — what it stores, and why long-lived. Tables (minimal):
collection:id,ownerID,name,type,updationTime,isShared, the three magic-metadata layers (JSON), the decrypted collection key,deleted, and this collection's file-diff cursorsinceTime.file: one row per membership, PK(collectionID, fileID), withupdationTime, the basic + two magic metadata layers (JSON),fileHeader/thumbHeader, the decrypted file key,contentHash,fileSize/thumbSize.content: one row per uniquefileID— presence + verification state of the on-disk original and thumbnail, the storedcontentHash, byte length, extension.failure: durable per-fileID(× kind) ledger — classification (transient/permanent/unknown), message, attempts, last-tried time.metarow: collections-list cursor,userID, schema version.Populated from the diff endpoints, decrypted with the existing
decryptCollection/decryptFile. Invalidation is server-driven, keyed onupdationTime: a diff row newer than the stored one replaces the metadata; iffileHeader/contentHashchanged, thecontentrow is marked stale for re-download. Tombstones (isDeleted) remove the membership (or mark the collection deleted). It is long-lived because nothing else can make it stale — there is no speculative caching here; it is an authoritative local mirror, and keeping it turns every later run into an O(changes) diff instead of an O(account) full re-enumeration + re-decrypt.Secrets note: this DB holds decrypted metadata (titles, GPS) and the decrypted collection/file keys, so it is as sensitive as
session.json— create it0600in a0700dir and treat it like the password. (Alternative: store only the rawencryptedKey/nonce and re-derive keys from the in-memory master key on read; more moving parts for the same on-disk sensitivity — not recommended.)The on-disk content + thumbnail cache — layout, keying, lifetime.
fileID(content dedup — one download regardless of how many collections hold it; matches today). Layout:originals/<fileID>.<ext>andthumbnails/<fileID>.<ext>. Flat directories; no prefix sharding unless a real account proves it necessary.writeAtomic(temp sibling + rename). Thecontentrow is flipped to "present + verified" only after the rename succeeds, so the DB can never claim a file the disk does not fully have.fileID— an edit yields a new version observable as a newupdationTime(and a changedfileHeader/contentHash). A hash-verified original never needs re-fetching unless the diff reports a new version. No TTL.The higher-level API — only what callers and the backup tool need. A cache object, constructed from a
Clientand a directory (the concrete class name is yours to choose). Surface:sync()— pull each collection's diff from its stored cursor, update the DB, apply tombstones; returns what changed. This is the incremental enumeration thrown away today.collections()/files(collectionID)— served from the DB, no network. (Also removes the O(collections) linear scanget/get-thumbdo now.)ensureOriginal(fileID)/ensureThumbnail(fileID)— idempotent: download + verify + store if absent-or-stale, else a no-op; returns the path.verify(fileID)/verifyAll()— re-hash on-disk bytes against the stored hash; mark mismatches stale.pending()— files whose content is absent, stale, or previously failed.pathFor(fileID).That is the whole surface. The only inputs are the
Clientand the directory — no tunables beyond whatApiClientalready exposes. It needs small additions toClient(which owns the keys and decryption): resumable enumerators that accept a startingsinceTime, return the final cursor, and surfaceisDeletedrows — today'slistCollections/listFilesreset to 0 and drop tombstones. ML/EXIF (backup-metadata) is out of scope for the core two caches; it can later read from the same DB.Design — the crash-safe backup tool on the API
runBackupbecomes:sync(), thenensureOriginal(and optionallyensureThumbnail) for eachpending()file, then materialize thecollections/symlink views + collection JSON from the DB. Concretely:pending()is exactly the unfinished set and the persisted cursor means no full re-enumeration.writeAtomicfor bytes; DB mutations in a transaction; the rule "rename the bytes, then record the row" keeps the DB from overstating the disk. The symlink tree and collection JSON are derived views, rebuildable from the DB, so they need no crash-safety of their own — and rebuilding the view also repairs the stale-sidecar and symlink-failure problems.TAG_FINALalready rejects truncation at decrypt time; after decrypt, compare the content hash toFileMetadata.hashwhen present and store it;verifyAll()re-checks disk against the DB on demand. This replaces today'ssize > 0. (Impl note: confirm the exact hash construction — algorithm, and the live-photo combined case — before relying on it; fall back toinfo.fileSizewhenhashis absent.)failurewith its classification and attempt count, and the run continues (as today). The next run retries transient/unknown failures and can skip or de-prioritize permanent ones. Exit non-zero while unresolved failures remain. This turns today's lost in-memoryerrors[]into a durable, queryable ledger.Ordered next steps (smallest set; reuse vs new)
All TDD per the repo workflow (tests first, red commit, branch off
main). The externalbackup <dir>layout and the CLI contract stay unchanged.info.fileSize/thumbSizeandisDeletedthroughdecryptFileintoEnteFile. Tiny; needed for integrity and deletion. Extend existing.Client: variants oflistCollections/listFilesthat take a startingsinceTime, return the final cursor, and includeisDeletedrows. Extend existing;ApiClientalready accepts an arbitrarysinceTime.sync(): the schema above, cursor persistence, tombstone application,updationTimeinvalidation. New; reuses decrypt + enumeration.fileID, atomic write, hash verification, DB state recorded only post-rename, orphan temp reaping on open. New; reuseswriteAtomic,streamDecrypt/TAG_FINAL,downloadFile/downloadThumbnail.sync,collections/files,ensureOriginal/ensureThumbnail,verify,pending,pathFor). Thin façade over 3–4.runBackupon the API + the durablefailureledger; portbackup.test.ts; keep the layout and exit-code contract. Rewrite.Sequencing vs
v1.0.0is yours to set — this subsumes some 1.0.0 items (notably the backup-robustness issue) and realizes the README/TODO "local cache (SQLite) … reliable" goal inside this repo, ahead of the desktop client.Relationship to existing issues
runBackupsymlink crash + partial originals: subsumed by step 6 (derived views cannot abort the run) and step 4 (atomic write + post-rename recording makes partial originals impossible).listFilesinfinite loop on a non-advancing server: the resumable enumerator in step 2 must handle it.streamDecryptbuffering: open question 4 (streaming decrypt-to-disk) resolves the whole-file buffer too.fileIDremoves the hazard for originals; the symlink view still sanitizes titles.Clientsession/keys: step 2 touchesClient; the cache depends on its keys.Open questions
backup <dir>(recommended: portable, self-contained, survives) vs one global store underenv-paths(paths.cacheis unused today;paths.dataholdssession.json). Global suits a future always-on desktop client; in-<dir>suits an explicit, movable backup. Recommend in-<dir>now and revisit for the desktop client.node:sqlite(zero new dependency, best for the hash-pinned supply chain, but still flagged experimental and needs a recent Node baseline) vsbetter-sqlite3(mature, synchronous, a native build). The global "prefer stdlib" rule points tonode:sqliteif your Node baseline supports it and experimental status is acceptable; otherwisebetter-sqlite3. Recommendnode:sqlite, falling back tobetter-sqlite3.ensureOriginal(recommended for "extremely reliable"; an extension ofwriteAtomic+streamDecrypt) vs keep whole-file buffering (simplest). Overlaps issue 21.Model: opus-4-8
I don't like the idea that the user (of the library) has to call sync - I want the cache to be transparent to the user, they should just use the API as if it were consulting the remote database as local.
Please write me 5 different blocks of example code, using the high-level library in five different ways, each being a different API surface/style of the high level API, so I can evaluate them and decide.
Why are we using SQLite instead of json? I expect all metadata to comfortably fit in 100MB. Reading entire files into ram and writing entire caches to disk on update will be trivial affairs. Keep our on-disk formats dead simple.
Use Fable.
Reading of your ruling: no
sync()anywhere in the public surface. Every readfirst asks the server for changes since the stored cursor (one small diff
request, usually empty), folds them into the local copy, saves it, then answers
from RAM. The caller never sees the cache. The existing
Client(
Client.fromJSON,listCollections,listFiles,downloadFile) stays as thelow-level layer underneath.
Each block below does the same task: open the library on a directory, find the
album "Iceland 2025", list its files, fetch one original, run a full backup,
report failed files.
EnteFile,Collection,BackupResult,BackupError,ClientSnapshotare today's real types; every other name is a placeholder andyours to choose.
Plain object of async methods
Flat and obvious, closest to today's
Client; nothing to learn. Tradeoff: idsget passed around by hand and the method list grows with every feature.
Collections of objects
Reads like a small object model: albums hold photos, photos know how to fetch
themselves. Pleasant to explore and hard to misuse (no loose ids). Tradeoff:
more classes to build and document, and objects held a long time can describe
a photo that has since changed on the server.
Directory-like tree
The account is presented as a read-only folder tree, the same shape as the
collections/andoriginals/layout on disk. Instantly familiar, and a FUSEmount or a file browser falls out of it later. Tradeoff: everything becomes a
path string, so metadata (GPS, dates, shared-with) has to be squeezed into a
statcall, and duplicate titles in one album need a naming rule.Async iterators and an event stream
Everything is a
for awaitloop; nothing large is ever built in memory by thecaller, and a backup reports progress as it happens instead of at the end.
Best fit for a CLI progress bar or a very large account. Tradeoff: simple
questions ("how many files?") need a loop or a helper, and error handling
inside loops is easy to get wrong.
Declare the wanted state, library makes it true
The caller says what should exist on disk and the library does whatever is
needed, every time, from any starting point (fresh, half-done, crashed). The
smallest surface of the five and the best match for "extremely reliable
backup". Tradeoff: it is a backup tool, not a library for browsing; listing an
album or grabbing one file has to be expressed as a narrower wanted state.
Why SQLite instead of JSON
No good reason under your constraint. It was chosen out of habit for "a durable
ledger with transactions", and the README/TODO wording said SQLite. With
metadata under 100MB, held in RAM and rewritten whole, JSON does the job and
the format is readable with
catandjq. Dropping SQLite also removes openquestion 2 (the driver choice) and a dependency. Accepted.
Simplest shape, inside the backup directory:
metadata.jsonholds everything the server told us: collections, files percollection, and the change cursors. One file, so the cursors can never
disagree with the data they describe: one rename covers both.
originals/<fileID>.<ext>andthumbnails/<fileID>.<ext>hold content, astoday. No "downloaded" flag is stored anywhere. Content only ever appears by
rename, so a file that exists is complete; the directory listing is the
record of what we have.
failures.jsonholds the failed-file list (small).collections/symlinks and per-collection JSON stay, rebuilt frommetadata.jsonon each run.What JSON costs here, plainly:
file would be absurd on a 50,000-file first backup. The shape above avoids it:
metadata.jsonis rewritten only when the server reports changes (once perrefresh, and during the first big enumeration at most once per collection or
every few seconds), never per download. Downloads touch only
failures.json,which is tiny.
or the new one, never half. Worst case we re-ask the server for a few pages
of changes. One gap: today's
writeAtomic(src/download/index.ts) does notfsync, so after a power cut the renamed file can be empty. The metadata
writer must fsync the file before the rename and the directory after it;
that is already tracked in #22.
writeAtomicis also private to the download module and needs exporting.metadata.jsonis not data loss: delete it and the nextread re-fetches everything from the server. It is a cache.
would have handled that). Answer: a lock file, one writer per directory, the
second process fails with a clear message.
second at startup. Acceptable for a backup tool; worth knowing for the
desktop client.
Maps built at load. No querylanguage is lost that we needed.
so
0600in a0700directory, same assession.json.One question the transparency ruling raises, for later: when the server is
unreachable, should reads fail, or answer from the local copy and say so? I
would answer from the local copy with a flag on the result; backups would still
fail loudly.
My pick: plain object of async methods for the library, with
backup()builtinternally the declare-the-state way, because it is the smallest step from
today's
Clientand the other styles can be layered on it later withoutbreaking it.
Disclosure: posted with the
teaCLI asclawbot; the Gitea MCP tool was notavailable in this session.
Model: fable-5-1
all these api shape designs use positional arguments which make it unclear what those paths mean in the second argument. use an args object with descriptive and unambiguous key names.
i like the collection of objects shape, but it needs a way of reporting progress in realtime, both when working a list or when working a single thing like a download. perhaps a progressCallback arg?
Revised API after your two rulings: every call takes one args object with named keys (no positional arguments anywhere in the new surface), and every long operation accepts an optional
onProgresscallback that fires in real time, both for list work and for a single download.Standing decisions, restated once:
sync()in the surface. Every read first asks the server for changes since the stored cursor (one small diff request, usually empty), folds them in, saves, then answers from RAM.metadata.json(collections, files, cursors),failures.json, plusoriginals/andthumbnails/, all written by temp-file-plus-rename.The surface
Client,ClientSnapshot,Collection,EnteFile,FileType,BackupResultare today's types (src/client.ts,src/model/types.ts,src/backup.ts).Library,Album,Photo,ProgressEvent,Failureare new; the names are yours to change.What
onProgressemits, per operation:albums.list,photos.list,byName,byID): arefreshstartedanddone(orfailed) pair first, since every read checks the server.albums.list:started,done.itemsTotalis known at once because/collections/v2answers in one response.photos.list:started, then oneprogressper diff page withitemsDonegrowing anditemsTotalabsent (the server pages withhasMore, so the total is unknown until the end), thendone. When the local copy is already current it is juststartedanddone.original()/thumbnail():startedwithphotoandbytesTotalwhen known,progressper chunk withbytesDone, thendone,skipped(already on disk) orfailedwitherror.backup():backup startedwithitemsTotal(known after the refresh), then for each file the same download events as above carryingphotoandalbum, then onebackup progresswithitemsDoneafter each file, thenbackup done. Per-item events and overall progress arrive through the one callback, so no second mechanism is needed.Worked examples
Find an album and list its files with progress:
Fetch one original with download progress:
Full backup with progress and the failures report:
Implementation notes, grounded in the current code
downloadFileanddownloadThumbnailinsrc/download/index.tsreportbytesWrittenonly at the end; the chunk loop needs a per-chunk hook sobytesDonecan be emitted while the body streams.bytesTotalisFileBlob.size, which comes fromRawEnteFile.info.fileSize;decryptFiledrops it today (step 1 of the original plan carries it through).ProgressCallbackinsrc/backup.tsis(message: string) => voidtoday; it is replaced by the typed event, and the CLI formats events into its log lines.Client(listFiles(collectionID, collectionKey),downloadFile(file, outPath?)) is left as is; the ruling is applied to the new surface, andClientcan be brought in line in a separate change if you want it.Open questions
refreshfailedevent (witherror), whilebackup()throws; the alternative is that every read fails when offline.Disclosure: posted with the tea CLI as clawbot; the Gitea MCP tool was not available in this session.
Model: fable-5-1
the cachedirectory should be an option and it should default to standard xda cache dirs. downloadDirectory should be an option as well.
how does this library work as a backend client for a GUI tool to replace the ente 1p desktop app? i want to build something very simple like autistmask/universal paperclips that just does a very plain minimum-styled list of photos (by week or month or something), albums, magic tags, etc. the ente 1p one is terribly laggy/slow and with all metadata on disk/in memory and a large thumbnail/data cache, we should be able to build one that's extremely performant (in electron). how does this api shape work for that use case? it seems good to me but i'd like a critical review along with any proposed changes that might make it work best for something like that.
also as part of this api shape implementation i'd like the cli quak tool updated to use the new api.
Reading of your three orders: both directories become options with the cache defaulting to the XDG cache location; the API gets a critical review as the backend of a minimal Electron photo app; the
quakCLI is ported to this API in the same work. Short verdict on the review: the object model is right for the CLI and for scripts, but as accepted it would reproduce the lag you want to get rid of, for three reasons: every read waits on the network, there is no read across all photos at all, and objects with methods cannot cross from the Electron main process to the window. Proposed changes below.Directory options
env-paths4.0.0 asbin/quak.tsalready calls it (envPaths("quak", { suffix: "" }), today used only forpaths.data/session.json):$XDG_CACHE_HOME/quak, else~/.cache/quak~/Library/Caches/quak%LOCALAPPDATA%\quak\Cache/<userID>subdirectory (client.whoami().userID), so two accounts never share onemetadata.json.cacheDirectoryholds only what can be deleted and fetched again,0700directory and0600files (it contains decrypted titles, GPS and file keys):metadata.json(albums, files, change cursors)failures.json(losing it loses only attempt counts)thumbnails/<fileID>.jpgoriginals/<fileID>.<ext>: originals fetched for viewing (photo.original())downloadDirectoryholds the backup the user keeps, in today'srunBackuplayout, unchanged:originals/<fileID>.<ext>, theoriginals/<fileID>.jsonsidecars,collections/<name>/symlinks,collections/<name>.json. The sidecars make it readable without the cache.backup()writes originals todownloadDirectory. An original already in the cache is copied, not downloaded again.photo.original()returns thedownloadDirectorycopy if there is one, else fetches into the cache.backup()with nodownloadDirectory(neither inopen()nor in the call) throws before any network traffic. It does not fall back to the cache: a backup in a directory that cache cleaners delete is not a backup.Critical review as an Electron backend
Startup: reads must not wait for the server
JSON.parseof a 100MBmetadata.jsonalso blocks its thread for about a second, and rewriting it blocks again; in the Electron main process that freezes every window.open(). The CLI keeps today's ruling; the GUI answers from RAM at once and is told when a refresh changed something. Still nosync()the caller must call."inBackground":open()returns as soon asmetadata.jsonis loaded (first paint is local only), a refresh starts at once and repeats every 60 seconds, and changes arrive throughsubscribe(below). On a first run with nometadata.json, reads return what has arrived so far and change events stream the rest in./collections/v2returns each album'supdationTime, so only albums whose time advanced need a/collections/v2/diffcall. Today'sClient.listFilesalways restarts atsinceTime: 0.utilityProcess, not in main, so the parse and the rewrite never block a window. That is the app's choice, but it forces the next point.Process split: plain data and ids, not objects
PhotoandAlbumhave methods, andEnteFilecarrieskey: Uint8Array. Across Electron IPC the methods are dropped, and file keys should never reach the window. One IPC call per photo (photo.thumbnail()200 times per screen) is also too chatty.Librarythat takes ids in arrays and returns plain records.AlbumandPhotostay as thin wrappers over those calls for in-process callers (the CLI, scripts).Photo:takenAtandtitleread onlymetadata, so a date or name the user corrected in Ente would be ignored and the timeline would sort wrongly.thumbnailPathon the record means cached thumbnails paint with no further call. The library knows what is cached from one directory read atopen().Timeline: there is no read across all photos
album.photos.list()andlib.photos.byID()only. A timeline would have to list every album and merge, and a file in three albums would appear three times (EnteFileis one row per album membership).Photo[]for 50,000 photos is fine in RAM but wrong over IPC. A virtualized list needs two things: the count per group (for total scroll height) and the ids per group; then records only for the rows on screen.Week-grouped timeline in use (main side;
sendis the app's IPC):Thumbnails at scroll speed
photo.thumbnail()is one request per call with no ordering, no limit on simultaneous requests and no way to stop. A fast scroll queues thousands of downloads for rows already gone.protocol.handle) that serves files out ofcacheDirectory/thumbnails; plainfile://is blocked for a page not itself loaded fromfile://.visiblebeforeaheadbeforebackground, the samefileIDrequested twice is downloaded once.Prefetch in use:
downloadFileholds the whole plaintext in RAM (streamDecryptinsrc/download/index.ts), so opening a multi-gigabyte video can kill the process. Streaming decrypt to disk (#21) is a precondition for the GUI, not an extra.Magic metadata and search
EnteFile.magicMetadataandpubMagicMetadataareRecord<string, unknown>(src/model/types.ts); the acceptedPhotoexposes none of it. The fields a GUI needs are promoted ontoPhotoRecordabove (caption,editedTime,editedName,w,hfrom the public layer;visibilityfrom the private one). Unverified: those are Ente's field names as I know them; confirm against the fixtures when implementing.PhotoFilter.text, in RAM. Nothing more is needed at 50,000 records.src/metadata-backup.tsalready fetches it (/files/data/fetch, typemldata, 200 ids per request): face detections and CLIP embeddings per file. Ente stores no tag strings server-side; its "magic search" embeds the typed text on the device and compares. Named people live behind an Ente endpoint quak does not implement. Embeddings for 50,000 files exceed 100MB as JSON, so they must never go intometadata.json; if wanted, a separatemldata.jsonloaded on demand.Change notification
quak backupon the same cache would always fail. Recommend no lock:metadata.jsoncarries its own cursors and is replaced by rename, so whichever process writes last leaves a consistent file and the other re-asks the server for a few changes.CLI
The
quakCLI is ported to this API in the same implementation; commands, flags, output and exit codes stay as they are, plus one new global option--cache-dir.collections:lib.albums.list().files --collection:lib.albums.byID()thenalbum.photos.list()get/get-thumb:lib.photos.byID()thenphoto.original()/photo.thumbnail(), copied to--out(default as today). This removes today's scan of every album per call;--collectionis accepted and ignored.backup <dir>:Library.open({ client, downloadDirectory: dir }),lib.backup({ onProgress }); the typed events are formatted into today's log lines;--jsonand exit 1 on failures unchanged.backup-metadata <dir>andhelper *-missing-thumbnails: enumerate throughlib.albums/album.photosinstead ofClient.listCollections/listFiles; the ML fetch, EXIF extraction and thumbnail upload stay where they are.--exifreadsphoto.original()instead of a temporary download.login,whoami,logout: unchanged.Open questions
open()optioncacheOriginalsMaxBytes; thumbnails are never removed.lib.refresh()for a reload button, never required: recommend yes.Disclosures: posted with the
teaCLI asclawbot(the Gitea MCP tool was not available). Nothing was built or run; Electron behaviour (protocol.handle,utilityProcess, IPC cloning) is from its documentation, not tested here. The 8-download limit, the 60-second interval and the 10 GiB default are judgement calls.Model: fable-5-1
magic tags means ML results - "magic" refers to ente's search by content feature.
10GiB cache is tiny. 100GiB is fine, disks are big. we can adapt it downward dynamically LRU if the disk has less than 50GiB free.
no lock file.
i would like a TTL of our local cache, so on launch we fetch things, but then we don't hit the network at all again for 1-3 seconds, so that subsequent calls serve entirely locally. this should simplify the UI side of things.
i want to use a lot of ram and disk and cpu in exchange for no lag ever. it's ok to load all of the metadata into ram on both sides of the IPC on startup, and to use dozens of gigabytes of disk space for local cache. i'd like to perhaps even aggressively locally precache thumbnails up to and including 100% of the account, for maximum local speed. we should probably locally precache originals for the latest week of items too, and all items in the favorites album.
let's move all of the caching/ttl/etc stuff into the library so the UI doesn't have to know much of anything about it.
suggests that all queries would hit the server for freshness. i don't want that. we can be 3-5 seconds behind the server for recent images, and 5-10 seconds behind the server for lists of albums, and probably 10 seconds behind the server for album membership.
i want separate parallelization for metadata fetches vs content fetches vs thumbnail fetches. metadata fetches should have max 10 req parallel, content fetches should have max 5, and thumbnail fetches should have max 25.
Revised design after #36 (comment) and #36 (comment). Reading of the seven rulings, then the design that follows from each. Everything accepted in #36 (comment) that is not named below stands unchanged (args objects everywhere,
onProgressevents, plain records over IPC,PhotoRecord,timeline.groups, JSON on disk, directory options, CLI port).Rulings, as read:
metadata.json.refresh: "beforeEachRead"mode and theserverReachableflag are gone.TTL and refresh (rulings 4 and 6)
open()loadsmetadata.jsoninto RAM, then fetches: one/collections/v2call with the stored cursor, then/collections/v2/difffrom each album's stored cursor for the albums whoseupdationTimeadvanced, then the ML data (below). On a first run there is no local copy, soopen()resolves only when this first fetch is complete. On later runs it resolves as soon as the local copy is loaded and the first refresh has been started, so a launch with the server down still paints at once./collections/v2?sinceTime=cursorrequest (empty answer when nothing changed), and a diff only for the albums it names. Every read (albums.list,photos.list,photos.records,timeline.groups,snapshot) is answered from RAM, always, with no refresh in front of it. That meets all three bounds with one small request per 3 s: new images, the album list and membership are each at most 3 s plus one round trip behind. There is no per-kind timer because the per-album diff is only sent for albums the collections call reported as changed, so the three bounds collapse into one interval. This rests on the server advancing an album'supdationTimewhen a file is added or removed from it, which is how Ente's own clients decide what to diff; unverified here, see the end.metadata.jsonis rewritten (temp file plus rename, fsync before and after) only when a refresh changed something. A refresh that fails is not an error a read can see: reads keep answering from RAM, the failure is reported once throughonProgress(operation: "refresh",status: "failed") and is visible instatus().LibraryChangelosesserverReachable;open()losesrefresh.refreshIntervalSeconds, default 3. A caller that wants the album list only every 10 s sets 10; there is no separate knob per kind.lib.refresh()is dropped: with a 3 s interval a reload button has nothing to do.Client.listCollectionsandClient.listFiles(src/client.ts) both start atsinceTime: 0and dropisDeletedrows; the refresh needs the cursor-taking, tombstone-surfacing variants from the original plan (step 2). Every request in this section goes through the metadata pool.All metadata in RAM on both sides (ruling 5)
snapshot()returns the whole account as plain records in one call, andsubscribethen delivers only what changed. The window keeps its own copy, groups and filters it locally, and never asks the library per row. 50,000PhotoRecords are roughly 15 MB over structured clone; a one-time cost at startup.PhotoRecordgainsoriginalPath?next tothumbnailPath?: set when the original is in the cache or indownloadDirectory, so the viewer opens pinned and recently viewed photos with no call at all.timeline.groupsandPhotoFilterstay for in-process callers (CLI, scripts); a window holding the snapshot does not need them.Startup in the app's library process:
Precache (rulings 5 and 6)
Both precaches start inside
open()and need nothing from the caller. They report through theonProgressgiven toopen()and throughstatus().thumbnails.ensure:visibleandaheadrequests go ahead of the precache, so scrolling is never behind the background fill. A file whose thumbnail already exists costs oneSetlookup, from the directory listing taken atopen().Collection.type === "favorites"), then every file whosetakenAtlies within the 7 days ending at the newesttakenAtin the account. These are the pinned set. When a refresh adds a favorite or a newer file, it joins the queue; when the week window moves or a favorite is removed, the file leaves the pinned set and becomes an ordinary cached original.photos.byID().original()andthumbnails.ensurekeep their shapes;original()on a cached or pinned file returns the path with oneskippedevent and no network.runBackupandfetchMLDataForFilesare sequential today;downloadFilebuffers the whole plaintext (streamDecrypt,src/download/index.ts), so 5 concurrent originals means 5 files in RAM at once and a 4 GB video on each. Streaming decrypt to disk (#21) is a precondition for the originals precache, as it was for the viewer.Originals cache limit and eviction (ruling 2)
cacheOriginalsMaxBytes, default 100 GiB. Before each original is written the library checks the volume holdingcacheDirectory(fs.statfs, Node 18.15+, this repo is on 22) and uses the lower of the configured limit andbytesUsedByOriginals + bytesFree - 50 GiB. So on a full disk the limit falls on its own and rises again as space returns;status().originalsLimitBytesshows the value in force.freeBelowBytes(default 50 GiB) is the second option; both areopen()options.utimes) when a read returns its path, so the order survives restarts and needs no ledger.atimeis not used because most volumes mountnoatimeorrelatime.cacheDirectory/originalsis ever evicted.downloadDirectoryis the backup; it is never counted and never touched.No lock file (ruling 3)
No lock and no detection.
metadata.jsoncarries its own cursors and is replaced by one rename, so whichever process writes last leaves a consistent file; the other keeps running on its RAM copy and its next refresh re-asks the server from its own cursor. The same holds for the ML files and for eviction: two processes evicting at once each remove least-recently-used files from the directory they both see, and the worst case is a file removed that the other would have kept, which is refetched on next use.ML data, the magic tags (ruling 1)
What exists:
src/metadata-backup.tsfetches/files/data/fetchwithtype: "mldata", 200 ids per request, decrypts each entry with the file key and gunzips it. The payload (test/cli/metadata-backup.test.ts) is{ face: { faces: [{ faceID, detection: { box, landmarks }, score, blur, embedding }] }, clip: { embedding } }. Nothing else server-side: Ente keeps no tag strings; search by content embeds the typed text on the device and ranks files by similarity toclip.embedding.Storage, its own part of the cache, never in
metadata.json:mldata/<fileID>.json: the decrypted, gunzipped payload as fetched, one file per fileID, written by rename. Present means complete, the same rule as originals.mldata/clip.f32plusmldata/clip.json: the derived index the search runs on.clip.jsonlists the fileIDs in order and the embedding length;clip.f32is those embeddings as oneFloat32Array, so 50,000 files at 512 floats load in one read as 100 MB and no parse. Rebuilt from the per-file JSON when missing or when it disagrees with the files present, appended to when new payloads arrive.Float32Arrayand the id list. Per-file payloads (face boxes, landmarks) are read from disk on demand, not held.metadata.jsonand absent frommldata/is requested, 200 per call, 10 calls in flight. That is the whole account on the first run and only new files afterwards. A file whoseupdationTimeadvanced is refetched, because the ML payload is rewritten when a client re-runs indexing. The fetch is reported throughonProgress(operation: "fetchMLData") andstatus().Surface:
Search by content itself needs a text encoder that turns the typed words into an embedding in the same space as
clip.embedding. quak has none.searchByEmbeddingtakes the vector so the app, or a later addition to the library, can supply it; see the open questions.Three request pools (ruling 7)
/collections/v2,/collections/v2/diff,/files/data/fetchgetFileStream(originals, for precache, viewer and backup)getThumbnailStreamvisiblethumbnails) goes before precache and backup work; the same fileID requested twice is fetched once. A pool that is idle does not lend its slots to another.open()optionsmetadataConcurrency,contentConcurrency,thumbnailConcurrency, defaults 10, 5, 25. Grounding:ApiClienthas no concurrency limit and none of its callers run more than one request at a time; the pools wrap calls intoClientrather than changingApiClient, so the low-level layer stays as it is andquak backupgets 5 parallel downloads throughlib.backup()without a change of its own.What the UI is left knowing
Open, take the snapshot, subscribe, tell the library which thumbnails are on screen (
thumbnails.ensurewithvisibleandahead), ask for an original's path when the viewer opens one, and hand the library a query embedding for search. It never sees a cursor, a TTL, a limit or a pool.Judgement calls and open questions
updationTimeon file add or remove, the fallback is a diff of every album every 10 s, which is O(albums) small requests per cycle. Please confirm which it is if you know; otherwise the implementation checks it against the live API first.open()not waiting for the first refresh when a local copy exists (above). The alternative, waiting, delays first paint by a round trip and by up to the 30 s request deadline when offline.takenAtin the account, not the last 7 calendar days, so an account with no new uploads still has a week of originals pinned. Say if you meant upload time or calendar days.metadata.jsonrewrite would take about 10 s of CPU per change, which argues for splitting it per album at that scale.searchByEmbedding(nothing new in quak; the app carries the model). Recommend the second now, the first as its own issue once the app exists.updationTimeadvanced. If Ente re-indexes without touching the file, that change is missed until the per-file JSON is deleted; a cheaper check would need a server endpoint I have not confirmed exists.Disclosures: posted through the Gitea REST API as
clawbot; the Gitea MCP tool was not loaded in this session. Nothing was built or run. Not checked: the server'supdationTimebehaviour above, the ML payload shape beyond the test fixture, andfs.statfson Windows.Model: fable-5-1
Implementation issues filed for this design, broken into small ordered units,
starting from the pieces every later unit needs. They will be driven one
issue-to-PR unit at a time, each independently reviewed before it lands on
next. The design here is treated as settled enough to start; amend any unit asneeded and later ones are fleshed out just before dispatch.
Order and dependencies:
Foundation (low-level; everything else builds on these)
fileSize/thumbSize/isDeletedthroughdecryptFile.Client(after #37). Subsumes #7.Metadata cache + library core
Library.open+ background refresh loop + staleness window (after #41, #38).Content/thumbnail cache
Machine-learning (magic) data
Backup, CLI, docs
quakCLI to the new API (after #51). Overlaps #17.Older issues these subsume or overlap (#7, #8, #9, #17, #21, #22) are cross-
referenced in the units above and closed by their landing commits where fully
covered; #24 (retry follow-ups) and #10 (session/keys hardening) stay separate.
Starting now with the foundation: #37 and #39 first (independent files, in
parallel), then #38 after #37 and #40 after #39.
Model: opus-4-8
I would like some additional asynchronous API calls that can be used sort of like how you would use a synchronous fetch so you will still await them but they don't actually resolve until such time a round trip to the server has been made and the local cache has been updated. I want the usual case to be that a fetch kicks off a background retrieval and then responds from disk. as has been designed presently. But what that means is that it's going to get stale data. And in certain cases, you are going to want to lag while it makes a round trip so that you can be sure that you have the latest data. So in that case, I would like some additional methods available that will not resolve until such time it knows that the data is synchronized. The CLI tools should then use these methods.
Reading of the amendment above, for the record and for the manager to fold into the design and the affected unit issues before they dispatch:
Manager: amend the read-surface and CLI unit issues accordingly (and any other unit this touches), then continue the normal one-unit-at-a-time flow.
(model: claude-fable-5)