report.go:32-48 calls loadFileRows (db.go:207-239), which materializes the entire files table as a []scanRec. trees.go:38-43 does the same and then builds the full node graph on top of it.
Each record carries two 64-character hex hash strings plus the path plus slice and string headers — call it 210 bytes at minimum, before Go's allocator overhead. At the stated scale of ~10M files that is over 2 GB resident for report alone, on a machine that may be the same storage server running the scan.
This is the one design constraint the scan phase goes out of its way to honour and the reporting phase ignores. README §Design goal 2 is explicit: "Holding one small record (path, size, mtime) per file in memory during a scan is acceptable; holding every file's hashes is not (they stay in the database)."
Note the two commands have different ceilings. report never needs more than one duplicate group at a time and can stream. trees genuinely needs the whole hierarchy to compute Merkle digests bottom-up, but it does not need to retain the hex hashes once a leaf's signature has been folded into its parent's digest.
Definition of done
report streams: group in SQL with ORDER BY size DESC, head, tail, path over rows where head is non-empty, add the supporting index, and emit each group as its boundary is crossed. It must never hold more than one group. Output must remain byte-identical to today's for a given database — the existing report tests must pass unchanged.
trees stops retaining the record slice and the hex hash strings; leaf signatures are folded into parent digests as rows stream in, so peak memory is proportional to the directory count, not the file count.
Measured, not asserted: build a synthetic database of 10M rows, record peak RSS for both commands before and after, and put the numbers in the PR description. README §Design gets a sentence stating the reporting commands' memory behaviour at scale.
Determinism is preserved exactly — README §report mode and §trees mode both require identical output for identical database contents regardless of insertion order.
make check green, and make test still finishes in under 20 seconds (build the 10M-row database in a throwaway benchmark, not in the test suite).
`report.go:32-48` calls `loadFileRows` (`db.go:207-239`), which materializes the entire `files` table as a `[]scanRec`. `trees.go:38-43` does the same and then builds the full node graph on top of it.
Each record carries two 64-character hex hash strings plus the path plus slice and string headers — call it 210 bytes at minimum, before Go's allocator overhead. At the stated scale of ~10M files that is over 2 GB resident for `report` alone, on a machine that may be the same storage server running the scan.
This is the one design constraint the scan phase goes out of its way to honour and the reporting phase ignores. README §Design goal 2 is explicit: "Holding one small record (path, size, mtime) per file in memory during a scan is acceptable; holding every file's hashes is not (they stay in the database)."
Note the two commands have different ceilings. `report` never needs more than one duplicate group at a time and can stream. `trees` genuinely needs the whole hierarchy to compute Merkle digests bottom-up, but it does not need to retain the hex hashes once a leaf's signature has been folded into its parent's digest.
## Definition of done
1. `report` streams: group in SQL with `ORDER BY size DESC, head, tail, path` over rows where `head` is non-empty, add the supporting index, and emit each group as its boundary is crossed. It must never hold more than one group. Output must remain byte-identical to today's for a given database — the existing report tests must pass unchanged.
2. `trees` stops retaining the record slice and the hex hash strings; leaf signatures are folded into parent digests as rows stream in, so peak memory is proportional to the directory count, not the file count.
3. Measured, not asserted: build a synthetic database of 10M rows, record peak RSS for both commands before and after, and put the numbers in the PR description. README §Design gets a sentence stating the reporting commands' memory behaviour at scale.
4. Determinism is preserved exactly — README §report mode and §trees mode both require identical output for identical database contents regardless of insertion order.
5. `make check` green, and `make test` still finishes in under 20 seconds (build the 10M-row database in a throwaway benchmark, not in the test suite).
clawbot
added this to the 1.0.0 milestone 2026-08-09 03:45:49 +02:00
Plan (implementer's brief). Dispatched after #7 and #30 land, since all three change the same output loops.
report: let SQLite group and order so Go holds one group at a time. The output order stays exactly as today: groups by size descending, ties by the group's smallest path, paths ascending inside a group. Item 1's ORDER BY size DESC, head, tail, path was written before the content column existed and orders same-size groups by hash, which changes the output; do not use it as written. Only records with a content hash are grouped (as today). The report connection is limited to one open connection (SetMaxOpenConns(1)), so a second query while the first is still being read would wait forever; pick a shape that cannot do that. Keep it plain.
The supporting index goes into the version 1 schema in place: no migration, version stays 1.
trees: read rows ORDER BY path (the primary key, so no sort). Every path under one directory is contiguous in byte order, so a directory is complete once the stream leaves its prefix: compute its digest then and drop its file entries. Keep per directory only what grouping and the parent check need. The digest definition, and so the output, is unchanged.
Determinism: the existing report and trees tests pass unchanged; add one that inserts the same records in a different order and gets identical output.
Measurement (item 3), deviation taken: this host limits each session to about 2 GiB, and the current code alone needs more than that at 10M rows. Measure at 1M rows instead, in a throwaway database and program outside the test suite, and delete both afterwards. Put peak memory before and after, for both commands, in one line of the PR body, with a one-line disclosure of the scale used. README.md §Design gets one sentence on how the reporting commands' memory grows.
Model: opus-5-5
Plan (implementer's brief). Dispatched after https://git.eeqj.de/sneak/sfdupes/issues/7 and https://git.eeqj.de/sneak/sfdupes/issues/30 land, since all three change the same output loops.
- `report`: let SQLite group and order so Go holds one group at a time. The output order stays exactly as today: groups by size descending, ties by the group's smallest path, paths ascending inside a group. Item 1's `ORDER BY size DESC, head, tail, path` was written before the `content` column existed and orders same-size groups by hash, which changes the output; do not use it as written. Only records with a `content` hash are grouped (as today). The report connection is limited to one open connection (`SetMaxOpenConns(1)`), so a second query while the first is still being read would wait forever; pick a shape that cannot do that. Keep it plain.
- The supporting index goes into the version 1 schema in place: no migration, version stays 1.
- `trees`: read rows `ORDER BY path` (the primary key, so no sort). Every path under one directory is contiguous in byte order, so a directory is complete once the stream leaves its prefix: compute its digest then and drop its file entries. Keep per directory only what grouping and the parent check need. The digest definition, and so the output, is unchanged.
- Determinism: the existing report and trees tests pass unchanged; add one that inserts the same records in a different order and gets identical output.
- Measurement (item 3), deviation taken: this host limits each session to about 2 GiB, and the current code alone needs more than that at 10M rows. Measure at 1M rows instead, in a throwaway database and program outside the test suite, and delete both afterwards. Put peak memory before and after, for both commands, in one line of the PR body, with a one-line disclosure of the scale used. `README.md` §Design gets one sentence on how the reporting commands' memory grows.
Model: opus-5-5
clawbot
self-assigned this 2026-10-03 14:11:59 +02:00
Built in #76: report writes each duplicate row as SQLite returns it, grouped and ordered by one query over the new files_signature index (schema version still 1); trees reads the records in path order and finishes each directory as the stream leaves it. Output is unchanged.
Peak memory on a synthetic 1M-row database (1M rather than 10M, as the plan directs): report 496 MiB → 18 MiB, trees 674 MiB → 29 MiB.
Model: opus-5-5
Built in https://git.eeqj.de/sneak/sfdupes/pulls/76: `report` writes each duplicate row as SQLite returns it, grouped and ordered by one query over the new `files_signature` index (schema version still 1); `trees` reads the records in path order and finishes each directory as the stream leaves it. Output is unchanged.
Peak memory on a synthetic 1M-row database (1M rather than 10M, as the plan directs): `report` 496 MiB → 18 MiB, `trees` 674 MiB → 29 MiB.
Model: opus-5-5
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
report.go:32-48callsloadFileRows(db.go:207-239), which materializes the entirefilestable as a[]scanRec.trees.go:38-43does the same and then builds the full node graph on top of it.Each record carries two 64-character hex hash strings plus the path plus slice and string headers — call it 210 bytes at minimum, before Go's allocator overhead. At the stated scale of ~10M files that is over 2 GB resident for
reportalone, on a machine that may be the same storage server running the scan.This is the one design constraint the scan phase goes out of its way to honour and the reporting phase ignores. README §Design goal 2 is explicit: "Holding one small record (path, size, mtime) per file in memory during a scan is acceptable; holding every file's hashes is not (they stay in the database)."
Note the two commands have different ceilings.
reportnever needs more than one duplicate group at a time and can stream.treesgenuinely needs the whole hierarchy to compute Merkle digests bottom-up, but it does not need to retain the hex hashes once a leaf's signature has been folded into its parent's digest.Definition of done
reportstreams: group in SQL withORDER BY size DESC, head, tail, pathover rows whereheadis non-empty, add the supporting index, and emit each group as its boundary is crossed. It must never hold more than one group. Output must remain byte-identical to today's for a given database — the existing report tests must pass unchanged.treesstops retaining the record slice and the hex hash strings; leaf signatures are folded into parent digests as rows stream in, so peak memory is proportional to the directory count, not the file count.make checkgreen, andmake teststill finishes in under 20 seconds (build the 10M-row database in a throwaway benchmark, not in the test suite).Plan (implementer's brief). Dispatched after #7 and #30 land, since all three change the same output loops.
report: let SQLite group and order so Go holds one group at a time. The output order stays exactly as today: groups by size descending, ties by the group's smallest path, paths ascending inside a group. Item 1'sORDER BY size DESC, head, tail, pathwas written before thecontentcolumn existed and orders same-size groups by hash, which changes the output; do not use it as written. Only records with acontenthash are grouped (as today). The report connection is limited to one open connection (SetMaxOpenConns(1)), so a second query while the first is still being read would wait forever; pick a shape that cannot do that. Keep it plain.trees: read rowsORDER BY path(the primary key, so no sort). Every path under one directory is contiguous in byte order, so a directory is complete once the stream leaves its prefix: compute its digest then and drop its file entries. Keep per directory only what grouping and the parent check need. The digest definition, and so the output, is unchanged.README.md§Design gets one sentence on how the reporting commands' memory grows.Model: opus-5-5
Built in #76:
reportwrites each duplicate row as SQLite returns it, grouped and ordered by one query over the newfiles_signatureindex (schema version still 1);treesreads the records in path order and finishes each directory as the stream leaves it. Output is unchanged.Peak memory on a synthetic 1M-row database (1M rather than 10M, as the plan directs):
report496 MiB → 18 MiB,trees674 MiB → 29 MiB.Model: opus-5-5