Add 64 KiB head/tail and content-hash duplicate ladder (closes #61)
check / check (push) Successful in 1m0s
check / check (push) Successful in 1m0s
Replace the 1 KiB end sampling with a ladder for same-size candidates: SHA-256 of the first and last 64 KiB, then a content hash that is the whole file below 50 MiB (proof of identity) and gigabyte-spaced 1 MiB samples at or above (deliberately probabilistic). Two files are duplicates only when size, head, tail, and content all agree. The signature gains a content column; schema bumps to version 2, so a version 1 database is rejected and must be rescanned (unavoidable — every stored hash changed). Because report and trees group stored signatures across separate scans, content is computed for every shared-size file, not only within-run head/tail collisions; size remains the sole read gate. README "Duplicate detection" documents each rung; tests cover the window boundaries, the 50 MiB boundary, and a multi-gigabyte sampled case with sparse temp files. Model: opus-4-8
This commit is contained in:
@@ -17,14 +17,15 @@ const ioBufSize = 1 << 20
|
||||
const minGroupSize = 2
|
||||
|
||||
// scanRec is one file record from the database. The signature (size,
|
||||
// head, tail) is the duplicate key; mtime is informational only and
|
||||
// used by scan for change detection.
|
||||
// head, tail, content) is the duplicate key; mtime is informational
|
||||
// only and used by scan for change detection.
|
||||
type scanRec struct {
|
||||
size int64
|
||||
mtime int64
|
||||
head string
|
||||
tail string
|
||||
path string
|
||||
size int64
|
||||
mtime int64
|
||||
head string
|
||||
tail string
|
||||
content string
|
||||
path string
|
||||
}
|
||||
|
||||
// loadRecords opens the database and reads every file record for the
|
||||
@@ -52,8 +53,9 @@ func loadRecords(ctx context.Context) ([]scanRec, error) {
|
||||
}
|
||||
|
||||
// dupeGroup is one set of candidate-duplicate files: identical size,
|
||||
// head hash, and tail hash. paths is sorted lexicographically; the
|
||||
// first entry is the group's "first", the rest are dupes.
|
||||
// head hash, tail hash, and content hash. paths is sorted
|
||||
// lexicographically; the first entry is the group's "first", the rest
|
||||
// are dupes.
|
||||
type dupeGroup struct {
|
||||
size int64
|
||||
paths []string
|
||||
@@ -122,7 +124,9 @@ func collectDupeGroups(recs []scanRec) []dupeGroup {
|
||||
continue
|
||||
}
|
||||
|
||||
k := fileSig{size: r.size, head: r.head, tail: r.tail}
|
||||
k := fileSig{
|
||||
size: r.size, head: r.head, tail: r.tail, content: r.content,
|
||||
}
|
||||
groups[k] = append(groups[k], r.path)
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user