Stream report and trees instead of loading every record (closes #14)
check / check (push) Successful in 1m49s

report now has SQLite group the records and put the rows in report
order, helped by a new files_signature index on (size, head, tail,
content), and writes each row as it reads it. trees reads the records
in path order, where all the paths under a directory come together, so
it computes each directory's digest as soon as the stream leaves it and
keeps only its path, parent, digest and totals. Output is unchanged.

The tests that called the removed in-memory grouping functions now group
records stored in a database. A new test checks that both commands give
the same output whatever order the records were inserted in.

Model: opus-5-5
This commit is contained in:
2026-10-04 02:04:37 +00:00
parent 2dd1194f33
commit bd8d41b174
9 changed files with 540 additions and 292 deletions
+17 -5
View File
@@ -94,7 +94,14 @@ Goals, in order:
millions of files, ~150 TB filesystem, possibly slow or busy disks
(ZFS pool under resilver). Holding one small record (path, size,
mtime) per file in memory during a scan is acceptable; holding
every file's hashes is not (they stay in the database).
every file's hashes is not (they stay in the database). The
reporting commands hold no file's hashes either: `report` lets
SQLite group and order the records and writes each row as it reads
it, so its memory does not grow with the database, and `trees`
reads the records in path order and keeps each directory's path,
digest and totals, plus the files of the directories holding the
record being read, so its memory grows with the number of
directories, not files.
3. **Scan incrementally, analyze offline.** The expensive filesystem
scan maintains a persistent database; an unchanged file is never
read again on a rescan, except to compute its content hash once a
@@ -201,6 +208,7 @@ All three subcommands operate on a single SQLite database file:
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
) WITHOUT ROWID;
CREATE INDEX files_signature ON files (size, head, tail, content);
```
Paths are stored as BLOBs because Unix paths are raw bytes, not
@@ -215,7 +223,8 @@ All three subcommands operate on a single SQLite database file:
until the content phase of a scan (see "`scan` mode" below) has
read the file. A record with an empty `content` is never part of a
duplicate group, though it still defines the file for tree
reconstruction.
reconstruction. The `files_signature` index lets SQLite group the
records by signature for `report` without sorting the whole table.
### Duplicate detection
@@ -443,9 +452,12 @@ arguments.
**`report` must never touch the filesystem being analyzed.** It does not
stat, open, or otherwise access any path that appears in the records; its
only I/O is reading the database and writing stdout/stderr. It must
produce identical output whether or not the scanned filesystem is still
mounted.
only I/O is reading the database, writing stdout/stderr, and the
temporary file SQLite sorts in when the duplicate rows do not fit in
memory. SQLite puts that file in `$SQLITE_TMPDIR` or `$TMPDIR` when set,
otherwise in `/var/tmp` (or `/tmp`), and deletes it as soon as it has
opened it. `report` must produce identical output whether or not the
scanned filesystem is still mounted.
Processing: