Stream report and trees instead of loading every record (closes #14)
check / check (push) Failing after 46s

report now has SQLite group the records and put the rows in report
order, helped by a new files_signature index on (size, head, tail,
content), and writes each row as it reads it. trees reads the records
in path order, where all the paths under a directory come together, so
it computes each directory's digest as soon as the stream leaves it and
keeps only its path, parent, digest and totals. Output is unchanged.

The tests that called the removed in-memory grouping functions now group
records stored in a database. New tests check that both commands give
the same output whatever order the records were inserted in, and that a
stdout failure partway through a long report is reported as one.

Model: opus-5-5
This commit is contained in:
2026-10-04 03:16:46 +00:00
parent 9abf81535a
commit 7cbb2d7f20
9 changed files with 576 additions and 297 deletions
+22 -6
View File
@@ -94,7 +94,15 @@ Goals, in order:
millions of files, ~150 TB filesystem, possibly slow or busy disks
(ZFS pool under resilver). Holding one small record (path, size,
mtime) per file in memory during a scan is acceptable; holding
every file's hashes is not (they stay in the database).
every file's hashes is not (they stay in the database). The
reporting commands do not hold every file's hashes either: `report`
lets SQLite group and order the records and writes each row as it
reads it, so its memory does not grow with the database, and
`trees` reads the records in path order and keeps each directory's
path, digest and totals, plus the hashes of only the files in the
directories holding the record being read, so its memory grows with
the number of directories and with the size of the largest
directory.
3. **Scan incrementally, analyze offline.** The expensive filesystem
scan maintains a persistent database; an unchanged file is never
read again on a rescan, except to compute its content hash once a
@@ -184,7 +192,10 @@ All three subcommands operate on a single SQLite database file:
starts while a report is still reading waits for it up to the
10-second busy timeout, then fails; a `scan` that ends while a
report has the database open warns and leaves the database in WAL
mode until the next scan.
mode until the next scan. `report` writes each row as it reads it,
so it is still reading while its output is paused (a pager, a
stalled pipe), and a `scan` started then fails after the busy
timeout.
- `report` and `trees` open the database read-only and need only read
access to the database file, and no write access to its directory.
While the database is in WAL mode they also read the `-wal` and
@@ -202,6 +213,7 @@ All three subcommands operate on a single SQLite database file:
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
) WITHOUT ROWID;
CREATE INDEX files_signature ON files (size, head, tail, content);
```
Paths are stored as BLOBs because Unix paths are raw bytes, not
@@ -216,7 +228,8 @@ All three subcommands operate on a single SQLite database file:
until the content phase of a scan (see "`scan` mode" below) has
read the file. A record with an empty `content` is never part of a
duplicate group, though it still defines the file for tree
reconstruction.
reconstruction. The `files_signature` index lets SQLite group the
records by signature for `report` without sorting the whole table.
### Duplicate detection
@@ -444,9 +457,12 @@ arguments.
**`report` must never touch the filesystem being analyzed.** It does not
stat, open, or otherwise access any path that appears in the records; its
only I/O is reading the database and writing stdout/stderr. It must
produce identical output whether or not the scanned filesystem is still
mounted.
only I/O is reading the database, writing stdout/stderr, and the
temporary file SQLite sorts in when the duplicate rows do not fit in
memory. SQLite puts that file in `$SQLITE_TMPDIR` or `$TMPDIR` when set,
otherwise in `/var/tmp` (or `/tmp`), and deletes it as soon as it has
opened it. `report` must produce identical output whether or not the
scanned filesystem is still mounted.
Processing: