Stream report and trees instead of loading every record (closes #14)
check / check (push) Successful in 1m49s
check / check (push) Successful in 1m49s
report now has SQLite group the records and put the rows in report order, helped by a new files_signature index on (size, head, tail, content), and writes each row as it reads it. trees reads the records in path order, where all the paths under a directory come together, so it computes each directory's digest as soon as the stream leaves it and keeps only its path, parent, digest and totals. Output is unchanged. The tests that called the removed in-memory grouping functions now group records stored in a database. A new test checks that both commands give the same output whatever order the records were inserted in. Model: opus-5-5
This commit is contained in:
@@ -94,7 +94,14 @@ Goals, in order:
|
||||
millions of files, ~150 TB filesystem, possibly slow or busy disks
|
||||
(ZFS pool under resilver). Holding one small record (path, size,
|
||||
mtime) per file in memory during a scan is acceptable; holding
|
||||
every file's hashes is not (they stay in the database).
|
||||
every file's hashes is not (they stay in the database). The
|
||||
reporting commands hold no file's hashes either: `report` lets
|
||||
SQLite group and order the records and writes each row as it reads
|
||||
it, so its memory does not grow with the database, and `trees`
|
||||
reads the records in path order and keeps each directory's path,
|
||||
digest and totals, plus the files of the directories holding the
|
||||
record being read, so its memory grows with the number of
|
||||
directories, not files.
|
||||
3. **Scan incrementally, analyze offline.** The expensive filesystem
|
||||
scan maintains a persistent database; an unchanged file is never
|
||||
read again on a rescan, except to compute its content hash once a
|
||||
@@ -201,6 +208,7 @@ All three subcommands operate on a single SQLite database file:
|
||||
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
|
||||
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
|
||||
) WITHOUT ROWID;
|
||||
CREATE INDEX files_signature ON files (size, head, tail, content);
|
||||
```
|
||||
|
||||
Paths are stored as BLOBs because Unix paths are raw bytes, not
|
||||
@@ -215,7 +223,8 @@ All three subcommands operate on a single SQLite database file:
|
||||
until the content phase of a scan (see "`scan` mode" below) has
|
||||
read the file. A record with an empty `content` is never part of a
|
||||
duplicate group, though it still defines the file for tree
|
||||
reconstruction.
|
||||
reconstruction. The `files_signature` index lets SQLite group the
|
||||
records by signature for `report` without sorting the whole table.
|
||||
|
||||
### Duplicate detection
|
||||
|
||||
@@ -443,9 +452,12 @@ arguments.
|
||||
|
||||
**`report` must never touch the filesystem being analyzed.** It does not
|
||||
stat, open, or otherwise access any path that appears in the records; its
|
||||
only I/O is reading the database and writing stdout/stderr. It must
|
||||
produce identical output whether or not the scanned filesystem is still
|
||||
mounted.
|
||||
only I/O is reading the database, writing stdout/stderr, and the
|
||||
temporary file SQLite sorts in when the duplicate rows do not fit in
|
||||
memory. SQLite puts that file in `$SQLITE_TMPDIR` or `$TMPDIR` when set,
|
||||
otherwise in `/var/tmp` (or `/tmp`), and deletes it as soon as it has
|
||||
opened it. `report` must produce identical output whether or not the
|
||||
scanned filesystem is still mounted.
|
||||
|
||||
Processing:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user