Stream report and trees instead of loading every record (closes #14) #76

Merged
clawbot merged 1 commits from issue-14-stream-reports into next 2026-10-04 06:47:25 +02:00
Collaborator

report and trees now read the database as a stream instead of loading every record into memory (#14).

  • report: one SQL query groups the records that have a content hash by (size, head, tail, content), keeps groups of two or more, and returns rows already in report order: size descending, then the group's first path, then path. Each row is written as it is read; the summary's record count comes from the same read transaction. The schema gains the index files_signature on (size, head, tail, content); the version stays 1.
  • trees: reads the records ORDER BY path, the primary key. All paths under a directory come together in that order, so a directory's digest is computed as soon as the stream leaves it and its file entries are dropped. The digest definition is unchanged.

What the diff does not show:

  • On a large database SQLite sorts the report rows in a temporary file in $SQLITE_TMPDIR, $TMPDIR or /var/tmp; README §report mode says so.
  • trees still holds the file entries of the directories containing the record being read, so one huge flat directory costs memory in proportion to its files.

Peak memory, synthetic 1M-row database, before → after: report 496 MiB → 18 MiB, trees 674 MiB → 29 MiB.

  • Deviation: measured at 1M rows, not 10M, as the plan on the issue directs.
  • Judgement call: tests of the removed in-memory functions (collectDupeGroups, buildHierarchy) keep their cases but run them through a database and the streaming code; the end-to-end report and trees tests are unchanged.

Model: opus-5-5

`report` and `trees` now read the database as a stream instead of loading every record into memory (https://git.eeqj.de/sneak/sfdupes/issues/14). - `report`: one SQL query groups the records that have a content hash by (size, head, tail, content), keeps groups of two or more, and returns rows already in report order: size descending, then the group's first path, then path. Each row is written as it is read; the summary's record count comes from the same read transaction. The schema gains the index `files_signature` on (size, head, tail, content); the version stays 1. - `trees`: reads the records `ORDER BY path`, the primary key. All paths under a directory come together in that order, so a directory's digest is computed as soon as the stream leaves it and its file entries are dropped. The digest definition is unchanged. What the diff does not show: - On a large database SQLite sorts the `report` rows in a temporary file in `$SQLITE_TMPDIR`, `$TMPDIR` or `/var/tmp`; README §`report` mode says so. - `trees` still holds the file entries of the directories containing the record being read, so one huge flat directory costs memory in proportion to its files. Peak memory, synthetic 1M-row database, before → after: `report` 496 MiB → 18 MiB, `trees` 674 MiB → 29 MiB. - Deviation: measured at 1M rows, not 10M, as the plan on the issue directs. - Judgement call: tests of the removed in-memory functions (`collectDupeGroups`, `buildHierarchy`) keep their cases but run them through a database and the streaming code; the end-to-end `report` and `trees` tests are unchanged. Model: opus-5-5
clawbot added the needs-review label 2026-10-04 04:12:30 +02:00
clawbot self-assigned this 2026-10-04 04:12:30 +02:00
Author
Collaborator
  1. No test guards the output-order trap. In db.go dupeRowsSQL, ordering same-size groups by hash (the issue's original ORDER BY size DESC, head, tail, path) still passes every test, because in every same-size fixture (TestDupeGroupsTieBreak, TestRunReportsIgnoreInsertionOrder) hash order and first-path order agree. Acceptable: a fixture whose same-size groups have hashes that sort opposite to their first paths, so ordering by hash fails the test.

  2. No test covers a stdout write that fails while report is still reading rows. Every write-failure test fails only at the final flush, because its output is smaller than the 1 MiB stdout buffer. Removing the writeErr check in report.go still passes every test, but a full disk is then reported as database PATH: write ... instead of write stdout: .... Acceptable: a test whose report output is larger than the buffer and whose stdout fails, expecting the write stdout: error.

  3. README.md §Design goal 2: "The reporting commands hold no file's hashes either" and "its memory grows with the number of directories, not files" are not true of trees. It holds the entries, hashes included, of every file in the directories containing the record being read, so one large directory costs memory for each of its files. Acceptable: say trees holds the hashes of only those files, so its memory grows with the number of directories and with the size of the largest directory.

  4. A behaviour change is not documented. report now keeps the database open until it has written its last row. While its output is paused (a pager, a stalled pipe), a scan that starts fails after the 10-second busy timeout. Before this change, report finished reading before it wrote anything. Acceptable: the README.md §Database bullet on switching journal modes says so.

Not reproduced: the 1M-row memory figures in the PR body.

Model: opus-5-5

1. No test guards the output-order trap. In `db.go` `dupeRowsSQL`, ordering same-size groups by hash (the issue's original `ORDER BY size DESC, head, tail, path`) still passes every test, because in every same-size fixture (`TestDupeGroupsTieBreak`, `TestRunReportsIgnoreInsertionOrder`) hash order and first-path order agree. Acceptable: a fixture whose same-size groups have hashes that sort opposite to their first paths, so ordering by hash fails the test. 2. No test covers a stdout write that fails while `report` is still reading rows. Every write-failure test fails only at the final flush, because its output is smaller than the 1 MiB stdout buffer. Removing the `writeErr` check in `report.go` still passes every test, but a full disk is then reported as `database PATH: write ...` instead of `write stdout: ...`. Acceptable: a test whose report output is larger than the buffer and whose stdout fails, expecting the `write stdout:` error. 3. `README.md` §Design goal 2: "The reporting commands hold no file's hashes either" and "its memory grows with the number of directories, not files" are not true of `trees`. It holds the entries, hashes included, of every file in the directories containing the record being read, so one large directory costs memory for each of its files. Acceptable: say `trees` holds the hashes of only those files, so its memory grows with the number of directories and with the size of the largest directory. 4. A behaviour change is not documented. `report` now keeps the database open until it has written its last row. While its output is paused (a pager, a stalled pipe), a `scan` that starts fails after the 10-second busy timeout. Before this change, `report` finished reading before it wrote anything. Acceptable: the `README.md` §Database bullet on switching journal modes says so. Not reproduced: the 1M-row memory figures in the PR body. Model: opus-5-5
clawbot added needs-rework and removed needs-review labels 2026-10-04 04:42:26 +02:00
clawbot force-pushed issue-14-stream-reports from bd8d41b174 to 7cbb2d7f20 2026-10-04 05:16:54 +02:00 Compare
clawbot added 1 commit 2026-10-04 05:17:49 +02:00
report now has SQLite group the records and put the rows in report
order, helped by a new files_signature index on (size, head, tail,
content), and writes each row as it reads it. trees reads the records
in path order, where all the paths under a directory come together, so
it computes each directory's digest as soon as the stream leaves it and
keeps only its path, parent, digest and totals. Output is unchanged.

The tests that called the removed in-memory grouping functions now group
records stored in a database. New tests check that both commands give
the same output whatever order the records were inserted in, and that a
stdout failure partway through a long report is reported as one.

Model: opus-5-5
clawbot force-pushed issue-14-stream-reports from 7cbb2d7f20 to dbc9afb4ef 2026-10-04 05:17:49 +02:00 Compare
Author
Collaborator

Reworked and rebased onto current next.

  1. TestDupeGroupsTieBreak now gives the same-size groups hashes that sort opposite to their first paths, so ordering them by hash fails it.
  2. New TestReportStdoutFailsWhileReading: a report more than twice the stdout buffer, written to a failing stdout, must fail with write stdout:. Without the writeErr check it fails with database PATH: write failed.
  3. README.md §Design goal 2 now says trees holds the hashes of only the files in the directories holding the record being read, so its memory grows with the number of directories and the size of the largest directory.
  4. README.md §Database, journal-mode bullet: report is still reading while its output is paused (a pager, a stalled pipe), so a scan started then fails after the busy timeout.

Model: opus-5-5

Reworked and rebased onto current `next`. 1. `TestDupeGroupsTieBreak` now gives the same-size groups hashes that sort opposite to their first paths, so ordering them by hash fails it. 2. New `TestReportStdoutFailsWhileReading`: a report more than twice the stdout buffer, written to a failing stdout, must fail with `write stdout:`. Without the `writeErr` check it fails with `database PATH: write failed`. 3. `README.md` §Design goal 2 now says `trees` holds the hashes of only the files in the directories holding the record being read, so its memory grows with the number of directories and the size of the largest directory. 4. `README.md` §Database, journal-mode bullet: `report` is still reading while its output is paused (a pager, a stalled pipe), so a `scan` started then fails after the busy timeout. Model: opus-5-5
clawbot added needs-review and removed needs-rework labels 2026-10-04 05:19:12 +02:00
Author
Collaborator

Review passed.

Model: opus-5-5

Review passed. Model: opus-5-5
clawbot merged commit 33cf3dd29a into next 2026-10-04 06:47:25 +02:00
clawbot deleted branch issue-14-stream-reports 2026-10-04 06:47:25 +02:00
Sign in to join this conversation.