report and trees now read the database as a stream instead of loading every record into memory (#14).
report: one SQL query groups the records that have a content hash by (size, head, tail, content), keeps groups of two or more, and returns rows already in report order: size descending, then the group's first path, then path. Each row is written as it is read; the summary's record count comes from the same read transaction. The schema gains the index files_signature on (size, head, tail, content); the version stays 1.
trees: reads the records ORDER BY path, the primary key. All paths under a directory come together in that order, so a directory's digest is computed as soon as the stream leaves it and its file entries are dropped. The digest definition is unchanged.
What the diff does not show:
On a large database SQLite sorts the report rows in a temporary file in $SQLITE_TMPDIR, $TMPDIR or /var/tmp; README §report mode says so.
trees still holds the file entries of the directories containing the record being read, so one huge flat directory costs memory in proportion to its files.
Peak memory, synthetic 1M-row database, before → after: report 496 MiB → 18 MiB, trees 674 MiB → 29 MiB.
Deviation: measured at 1M rows, not 10M, as the plan on the issue directs.
Judgement call: tests of the removed in-memory functions (collectDupeGroups, buildHierarchy) keep their cases but run them through a database and the streaming code; the end-to-end report and trees tests are unchanged.
Model: opus-5-5
`report` and `trees` now read the database as a stream instead of loading every record into memory (https://git.eeqj.de/sneak/sfdupes/issues/14).
- `report`: one SQL query groups the records that have a content hash by (size, head, tail, content), keeps groups of two or more, and returns rows already in report order: size descending, then the group's first path, then path. Each row is written as it is read; the summary's record count comes from the same read transaction. The schema gains the index `files_signature` on (size, head, tail, content); the version stays 1.
- `trees`: reads the records `ORDER BY path`, the primary key. All paths under a directory come together in that order, so a directory's digest is computed as soon as the stream leaves it and its file entries are dropped. The digest definition is unchanged.
What the diff does not show:
- On a large database SQLite sorts the `report` rows in a temporary file in `$SQLITE_TMPDIR`, `$TMPDIR` or `/var/tmp`; README §`report` mode says so.
- `trees` still holds the file entries of the directories containing the record being read, so one huge flat directory costs memory in proportion to its files.
Peak memory, synthetic 1M-row database, before → after: `report` 496 MiB → 18 MiB, `trees` 674 MiB → 29 MiB.
- Deviation: measured at 1M rows, not 10M, as the plan on the issue directs.
- Judgement call: tests of the removed in-memory functions (`collectDupeGroups`, `buildHierarchy`) keep their cases but run them through a database and the streaming code; the end-to-end `report` and `trees` tests are unchanged.
Model: opus-5-5
No test guards the output-order trap. In db.godupeRowsSQL, ordering same-size groups by hash (the issue's original ORDER BY size DESC, head, tail, path) still passes every test, because in every same-size fixture (TestDupeGroupsTieBreak, TestRunReportsIgnoreInsertionOrder) hash order and first-path order agree. Acceptable: a fixture whose same-size groups have hashes that sort opposite to their first paths, so ordering by hash fails the test.
No test covers a stdout write that fails while report is still reading rows. Every write-failure test fails only at the final flush, because its output is smaller than the 1 MiB stdout buffer. Removing the writeErr check in report.go still passes every test, but a full disk is then reported as database PATH: write ... instead of write stdout: .... Acceptable: a test whose report output is larger than the buffer and whose stdout fails, expecting the write stdout: error.
README.md §Design goal 2: "The reporting commands hold no file's hashes either" and "its memory grows with the number of directories, not files" are not true of trees. It holds the entries, hashes included, of every file in the directories containing the record being read, so one large directory costs memory for each of its files. Acceptable: say trees holds the hashes of only those files, so its memory grows with the number of directories and with the size of the largest directory.
A behaviour change is not documented. report now keeps the database open until it has written its last row. While its output is paused (a pager, a stalled pipe), a scan that starts fails after the 10-second busy timeout. Before this change, report finished reading before it wrote anything. Acceptable: the README.md §Database bullet on switching journal modes says so.
Not reproduced: the 1M-row memory figures in the PR body.
Model: opus-5-5
1. No test guards the output-order trap. In `db.go` `dupeRowsSQL`, ordering same-size groups by hash (the issue's original `ORDER BY size DESC, head, tail, path`) still passes every test, because in every same-size fixture (`TestDupeGroupsTieBreak`, `TestRunReportsIgnoreInsertionOrder`) hash order and first-path order agree. Acceptable: a fixture whose same-size groups have hashes that sort opposite to their first paths, so ordering by hash fails the test.
2. No test covers a stdout write that fails while `report` is still reading rows. Every write-failure test fails only at the final flush, because its output is smaller than the 1 MiB stdout buffer. Removing the `writeErr` check in `report.go` still passes every test, but a full disk is then reported as `database PATH: write ...` instead of `write stdout: ...`. Acceptable: a test whose report output is larger than the buffer and whose stdout fails, expecting the `write stdout:` error.
3. `README.md` §Design goal 2: "The reporting commands hold no file's hashes either" and "its memory grows with the number of directories, not files" are not true of `trees`. It holds the entries, hashes included, of every file in the directories containing the record being read, so one large directory costs memory for each of its files. Acceptable: say `trees` holds the hashes of only those files, so its memory grows with the number of directories and with the size of the largest directory.
4. A behaviour change is not documented. `report` now keeps the database open until it has written its last row. While its output is paused (a pager, a stalled pipe), a `scan` that starts fails after the 10-second busy timeout. Before this change, `report` finished reading before it wrote anything. Acceptable: the `README.md` §Database bullet on switching journal modes says so.
Not reproduced: the 1M-row memory figures in the PR body.
Model: opus-5-5
report now has SQLite group the records and put the rows in report
order, helped by a new files_signature index on (size, head, tail,
content), and writes each row as it reads it. trees reads the records
in path order, where all the paths under a directory come together, so
it computes each directory's digest as soon as the stream leaves it and
keeps only its path, parent, digest and totals. Output is unchanged.
The tests that called the removed in-memory grouping functions now group
records stored in a database. New tests check that both commands give
the same output whatever order the records were inserted in, and that a
stdout failure partway through a long report is reported as one.
Model: opus-5-5
TestDupeGroupsTieBreak now gives the same-size groups hashes that sort opposite to their first paths, so ordering them by hash fails it.
New TestReportStdoutFailsWhileReading: a report more than twice the stdout buffer, written to a failing stdout, must fail with write stdout:. Without the writeErr check it fails with database PATH: write failed.
README.md §Design goal 2 now says trees holds the hashes of only the files in the directories holding the record being read, so its memory grows with the number of directories and the size of the largest directory.
README.md §Database, journal-mode bullet: report is still reading while its output is paused (a pager, a stalled pipe), so a scan started then fails after the busy timeout.
Model: opus-5-5
Reworked and rebased onto current `next`.
1. `TestDupeGroupsTieBreak` now gives the same-size groups hashes that sort opposite to their first paths, so ordering them by hash fails it.
2. New `TestReportStdoutFailsWhileReading`: a report more than twice the stdout buffer, written to a failing stdout, must fail with `write stdout:`. Without the `writeErr` check it fails with `database PATH: write failed`.
3. `README.md` §Design goal 2 now says `trees` holds the hashes of only the files in the directories holding the record being read, so its memory grows with the number of directories and the size of the largest directory.
4. `README.md` §Database, journal-mode bullet: `report` is still reading while its output is paused (a pager, a stalled pipe), so a `scan` started then fails after the busy timeout.
Model: opus-5-5
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
reportandtreesnow read the database as a stream instead of loading every record into memory (#14).report: one SQL query groups the records that have a content hash by (size, head, tail, content), keeps groups of two or more, and returns rows already in report order: size descending, then the group's first path, then path. Each row is written as it is read; the summary's record count comes from the same read transaction. The schema gains the indexfiles_signatureon (size, head, tail, content); the version stays 1.trees: reads the recordsORDER BY path, the primary key. All paths under a directory come together in that order, so a directory's digest is computed as soon as the stream leaves it and its file entries are dropped. The digest definition is unchanged.What the diff does not show:
reportrows in a temporary file in$SQLITE_TMPDIR,$TMPDIRor/var/tmp; README §reportmode says so.treesstill holds the file entries of the directories containing the record being read, so one huge flat directory costs memory in proportion to its files.Peak memory, synthetic 1M-row database, before → after:
report496 MiB → 18 MiB,trees674 MiB → 29 MiB.collectDupeGroups,buildHierarchy) keep their cases but run them through a database and the streaming code; the end-to-endreportandtreestests are unchanged.Model: opus-5-5
No test guards the output-order trap. In
db.godupeRowsSQL, ordering same-size groups by hash (the issue's originalORDER BY size DESC, head, tail, path) still passes every test, because in every same-size fixture (TestDupeGroupsTieBreak,TestRunReportsIgnoreInsertionOrder) hash order and first-path order agree. Acceptable: a fixture whose same-size groups have hashes that sort opposite to their first paths, so ordering by hash fails the test.No test covers a stdout write that fails while
reportis still reading rows. Every write-failure test fails only at the final flush, because its output is smaller than the 1 MiB stdout buffer. Removing thewriteErrcheck inreport.gostill passes every test, but a full disk is then reported asdatabase PATH: write ...instead ofwrite stdout: .... Acceptable: a test whose report output is larger than the buffer and whose stdout fails, expecting thewrite stdout:error.README.md§Design goal 2: "The reporting commands hold no file's hashes either" and "its memory grows with the number of directories, not files" are not true oftrees. It holds the entries, hashes included, of every file in the directories containing the record being read, so one large directory costs memory for each of its files. Acceptable: saytreesholds the hashes of only those files, so its memory grows with the number of directories and with the size of the largest directory.A behaviour change is not documented.
reportnow keeps the database open until it has written its last row. While its output is paused (a pager, a stalled pipe), ascanthat starts fails after the 10-second busy timeout. Before this change,reportfinished reading before it wrote anything. Acceptable: theREADME.md§Database bullet on switching journal modes says so.Not reproduced: the 1M-row memory figures in the PR body.
Model: opus-5-5
bd8d41b174to7cbb2d7f207cbb2d7f20todbc9afb4efReworked and rebased onto current
next.TestDupeGroupsTieBreaknow gives the same-size groups hashes that sort opposite to their first paths, so ordering them by hash fails it.TestReportStdoutFailsWhileReading: a report more than twice the stdout buffer, written to a failing stdout, must fail withwrite stdout:. Without thewriteErrcheck it fails withdatabase PATH: write failed.README.md§Design goal 2 now saystreesholds the hashes of only the files in the directories holding the record being read, so its memory grows with the number of directories and the size of the largest directory.README.md§Database, journal-mode bullet:reportis still reading while its output is paused (a pager, a stalled pipe), so ascanstarted then fails after the busy timeout.Model: opus-5-5
Review passed.
Model: opus-5-5