Stream report and trees instead of loading every record (closes #14)
check / check (push) Successful in 2m41s
check / check (push) Successful in 2m41s
report now has SQLite group the records and put the rows in report order, helped by a new files_signature index on (size, head, tail, content), and writes each row as it reads it. trees reads the records in path order, where all the paths under a directory come together, so it computes each directory's digest as soon as the stream leaves it and keeps only its path, parent, digest and totals. Output is unchanged. The tests that called the removed in-memory grouping functions now group records stored in a database. New tests check that both commands give the same output whatever order the records were inserted in, and that a stdout failure partway through a long report is reported as one. Model: opus-5-5
This commit is contained in:
@@ -94,7 +94,15 @@ Goals, in order:
|
||||
millions of files, ~150 TB filesystem, possibly slow or busy disks
|
||||
(ZFS pool under resilver). Holding one small record (path, size,
|
||||
mtime) per file in memory during a scan is acceptable; holding
|
||||
every file's hashes is not (they stay in the database).
|
||||
every file's hashes is not (they stay in the database). The
|
||||
reporting commands do not hold every file's hashes either: `report`
|
||||
lets SQLite group and order the records and writes each row as it
|
||||
reads it, so its memory does not grow with the database, and
|
||||
`trees` reads the records in path order and keeps each directory's
|
||||
path, digest and totals, plus the hashes of only the files in the
|
||||
directories holding the record being read, so its memory grows with
|
||||
the number of directories and with the size of the largest
|
||||
directory.
|
||||
3. **Scan incrementally, analyze offline.** The expensive filesystem
|
||||
scan maintains a persistent database; an unchanged file is never
|
||||
read again on a rescan, except to compute its content hash once a
|
||||
@@ -184,7 +192,10 @@ All three subcommands operate on a single SQLite database file:
|
||||
starts while a report is still reading waits for it up to the
|
||||
10-second busy timeout, then fails; a `scan` that ends while a
|
||||
report has the database open warns and leaves the database in WAL
|
||||
mode until the next scan.
|
||||
mode until the next scan. `report` writes each row as it reads it,
|
||||
so it is still reading while its output is paused (a pager, a
|
||||
stalled pipe), and a `scan` started then fails after the busy
|
||||
timeout.
|
||||
- `report` and `trees` open the database read-only and need only read
|
||||
access to the database file, and no write access to its directory.
|
||||
While the database is in WAL mode they also read the `-wal` and
|
||||
@@ -202,6 +213,7 @@ All three subcommands operate on a single SQLite database file:
|
||||
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
|
||||
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
|
||||
) WITHOUT ROWID;
|
||||
CREATE INDEX files_signature ON files (size, head, tail, content);
|
||||
```
|
||||
|
||||
Paths are stored as BLOBs because Unix paths are raw bytes, not
|
||||
@@ -216,7 +228,8 @@ All three subcommands operate on a single SQLite database file:
|
||||
until the content phase of a scan (see "`scan` mode" below) has
|
||||
read the file. A record with an empty `content` is never part of a
|
||||
duplicate group, though it still defines the file for tree
|
||||
reconstruction.
|
||||
reconstruction. The `files_signature` index lets SQLite group the
|
||||
records by signature for `report` without sorting the whole table.
|
||||
|
||||
### Duplicate detection
|
||||
|
||||
@@ -444,9 +457,12 @@ arguments.
|
||||
|
||||
**`report` must never touch the filesystem being analyzed.** It does not
|
||||
stat, open, or otherwise access any path that appears in the records; its
|
||||
only I/O is reading the database and writing stdout/stderr. It must
|
||||
produce identical output whether or not the scanned filesystem is still
|
||||
mounted.
|
||||
only I/O is reading the database, writing stdout/stderr, and the
|
||||
temporary file SQLite sorts in when the duplicate rows do not fit in
|
||||
memory. SQLite puts that file in `$SQLITE_TMPDIR` or `$TMPDIR` when set,
|
||||
otherwise in `/var/tmp` (or `/tmp`), and deletes it as soon as it has
|
||||
opened it. `report` must produce identical output whether or not the
|
||||
scanned filesystem is still mounted.
|
||||
|
||||
Processing:
|
||||
|
||||
|
||||
Reference in New Issue
Block a user