Compare commits
2
Commits
d111b3c665
...
next
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
29a65016d0 | ||
|
|
7ac4f6b723 |
@@ -2,7 +2,6 @@
|
|||||||
.claude
|
.claude
|
||||||
.DS_Store
|
.DS_Store
|
||||||
sfdupes
|
sfdupes
|
||||||
files.dat
|
|
||||||
*.log
|
*.log
|
||||||
*.out
|
*.out
|
||||||
*.test
|
*.test
|
||||||
|
|||||||
@@ -27,7 +27,6 @@ node_modules/
|
|||||||
*.log
|
*.log
|
||||||
|
|
||||||
# Local scan data
|
# Local scan data
|
||||||
files.dat
|
|
||||||
*.sqlite
|
*.sqlite
|
||||||
*.sqlite-shm
|
*.sqlite-shm
|
||||||
*.sqlite-wal
|
*.sqlite-wal
|
||||||
|
|||||||
@@ -46,4 +46,4 @@ hooks:
|
|||||||
@script/install-precommit
|
@script/install-precommit
|
||||||
|
|
||||||
clean:
|
clean:
|
||||||
rm -f $(BINARY) files.dat
|
rm -f $(BINARY)
|
||||||
|
|||||||
@@ -5,14 +5,19 @@
|
|||||||
`sfdupes` is an MIT-licensed Go CLI tool by
|
`sfdupes` is an MIT-licensed Go CLI tool by
|
||||||
[@sneak](https://sneak.berlin) that quickly identifies *candidate*
|
[@sneak](https://sneak.berlin) that quickly identifies *candidate*
|
||||||
duplicate files — and, ultimately, entire duplicate directory trees —
|
duplicate files — and, ultimately, entire duplicate directory trees —
|
||||||
across very large filesystems without reading full file contents. Files
|
across very large filesystems without reading every byte of every file.
|
||||||
are considered duplicates when they have identical size, identical
|
Files are considered duplicates when their sizes are equal and they
|
||||||
SHA-256 of their first 1024 bytes, and identical SHA-256 of their last
|
agree on a short ladder of hashes. A file under 10 MiB is hashed in full
|
||||||
1024 bytes. This is a strong candidate signal, not proof of identical
|
and compared directly. A larger file is gated first on the SHA-256 of
|
||||||
content (the middle of the file is never read); the intended use is
|
its first 64 KiB and of its last 64 KiB, and then compared on a content
|
||||||
|
hash — the SHA-256 of the whole file when it is under 50 MiB, or of
|
||||||
|
gigabyte-spaced 1 MiB samples when it is 50 MiB or larger. Below 50 MiB
|
||||||
|
the content hash is proof of identical content; at or above 50 MiB it is
|
||||||
|
a strong candidate signal rather than proof, because the gaps between
|
||||||
|
samples are never read. The intended use is
|
||||||
finding duplicate downloads and duplicated directory trees on
|
finding duplicate downloads and duplicated directory trees on
|
||||||
multi-terabyte ZFS servers where reading every byte is prohibitively
|
multi-terabyte ZFS servers where reading every byte of every file is
|
||||||
expensive. `scan` maintains a persistent SQLite database of file
|
prohibitively expensive. `scan` maintains a persistent SQLite database of file
|
||||||
signatures that survives between runs, so it can be run from cron and
|
signatures that survives between runs, so it can be run from cron and
|
||||||
the reports can be generated at any time from the most recent scan.
|
the reports can be generated at any time from the most recent scan.
|
||||||
|
|
||||||
@@ -29,7 +34,8 @@ export SFDUPES_DATABASE="$HOME/.local/share/sfdupes/db.sqlite"
|
|||||||
```
|
```
|
||||||
|
|
||||||
`scan` walks one or more filesystem trees and maintains one database
|
`scan` walks one or more filesystem trees and maintains one database
|
||||||
record per regular file (path, size, mtime, head hash, tail hash). The
|
record per regular file (path, size, mtime, head hash, tail hash,
|
||||||
|
content hash). The
|
||||||
database persists between runs; a rescan only hashes files that are new
|
database persists between runs; a rescan only hashes files that are new
|
||||||
or changed, and removes records for files that no longer exist.
|
or changed, and removes records for files that no longer exist.
|
||||||
`report` reads the database and prints the file-level duplicates
|
`report` reads the database and prints the file-level duplicates
|
||||||
@@ -47,10 +53,13 @@ completed scan.
|
|||||||
|
|
||||||
Duplicate finders that hash entire files do not scale to the target
|
Duplicate finders that hash entire files do not scale to the target
|
||||||
environment: ~10 million files and ~150 TB on possibly slow or busy
|
environment: ~10 million files and ~150 TB on possibly slow or busy
|
||||||
disks (a ZFS pool under resilver). Reading at most 2 KiB per file — and
|
disks (a ZFS pool under resilver). sfdupes spends disk I/O only on files
|
||||||
only from files whose size at least one other file shares, since a
|
whose size at least one other file shares, since a size-unique file
|
||||||
size-unique file cannot be a duplicate — makes a full-filesystem sweep
|
cannot be a duplicate; for those it reads the cheap end windows first
|
||||||
tractable, and the signatures are kept in a persistent database, so
|
and a content hash second — the whole file below 50 MiB, but only
|
||||||
|
gigabyte-spaced samples at or above 50 MiB, so the largest files are
|
||||||
|
never read in full. This keeps a full-filesystem sweep tractable, and
|
||||||
|
the signatures are kept in a persistent database, so
|
||||||
the expensive filesystem pass is incremental: a rescan re-hashes only
|
the expensive filesystem pass is incremental: a rescan re-hashes only
|
||||||
files whose recorded mtime or size changed, and all analysis happens
|
files whose recorded mtime or size changed, and all analysis happens
|
||||||
offline from the database alone. The end goal is
|
offline from the database alone. The end goal is
|
||||||
@@ -68,14 +77,18 @@ Goals, in order:
|
|||||||
downloads, copied project trees), so the operator can consider
|
downloads, copied project trees), so the operator can consider
|
||||||
removing an entire subtree at once. File-level duplicate detection is
|
removing an entire subtree at once. File-level duplicate detection is
|
||||||
the foundation; tree-level detection is built on top of it.
|
the foundation; tree-level detection is built on top of it.
|
||||||
2. **Never read full file contents.** At most 2 KiB is read per file
|
2. **Spend I/O in proportion to duplicate likelihood.** Only files
|
||||||
(first and last 1024 bytes), and only files whose size at least
|
whose size at least one other file shares are read at all — a
|
||||||
one other file shares are read at all — a size-unique file cannot
|
size-unique file cannot be a duplicate. Those are compared by the
|
||||||
be a duplicate. Scale target: tens of millions of files, ~150 TB
|
ladder in "Duplicate detection" below: a file under 10 MiB is hashed
|
||||||
filesystem, possibly slow or busy disks (ZFS pool under resilver).
|
in full, while a larger file is gated on cheap 64 KiB end windows
|
||||||
Holding one small record (path, size, mtime) per file in memory
|
first, then a content hash that reads the whole file below 50 MiB
|
||||||
during a scan is acceptable; holding every file's hashes is not
|
but only gigabyte-spaced 1 MiB samples at or above it, so the very
|
||||||
(they stay in the database).
|
largest files are still never read in full. Scale target: tens of
|
||||||
|
millions of files, ~150 TB filesystem, possibly slow or busy disks
|
||||||
|
(ZFS pool under resilver). Holding one small record (path, size,
|
||||||
|
mtime) per file in memory during a scan is acceptable; holding
|
||||||
|
every file's hashes is not (they stay in the database).
|
||||||
3. **Scan incrementally, analyze offline.** The expensive filesystem
|
3. **Scan incrementally, analyze offline.** The expensive filesystem
|
||||||
scan maintains a persistent database; an unchanged file is never
|
scan maintains a persistent database; an unchanged file is never
|
||||||
read again on a rescan. All analysis (`report`, `trees`) works from
|
read again on a rescan. All analysis (`report`, `trees`) works from
|
||||||
@@ -141,26 +154,72 @@ All three subcommands operate on a single SQLite database file:
|
|||||||
that dies partway leaves a valid database holding everything
|
that dies partway leaves a valid database holding everything
|
||||||
hashed so far; the next scan skips those records and converges
|
hashed so far; the next scan skips those records and converges
|
||||||
toward the filesystem.
|
toward the filesystem.
|
||||||
- Schema (`PRAGMA user_version` is the schema version, currently 1; a
|
- Schema (`PRAGMA user_version` is the schema version, currently 2; a
|
||||||
database with any other version is a fatal error):
|
database with any other version is a fatal error). Version 2 added
|
||||||
|
the `content` column and the 64 KiB head/tail signature (replacing
|
||||||
|
the version 1 1 KiB end windows), so a version 1 database cannot be
|
||||||
|
reused: it is rejected and the tree must be rescanned from scratch.
|
||||||
|
|
||||||
```sql
|
```sql
|
||||||
CREATE TABLE files (
|
CREATE TABLE files (
|
||||||
path BLOB PRIMARY KEY, -- absolute path, raw bytes
|
path BLOB PRIMARY KEY, -- absolute path, raw bytes
|
||||||
size INTEGER NOT NULL, -- bytes, from lstat
|
size INTEGER NOT NULL, -- bytes, from lstat
|
||||||
mtime INTEGER NOT NULL, -- Unix seconds, from lstat
|
mtime INTEGER NOT NULL, -- Unix seconds, from lstat
|
||||||
head TEXT NOT NULL, -- lowercase-hex SHA-256, first 1 KiB
|
head TEXT NOT NULL, -- lowercase-hex SHA-256; first 64 KiB, or whole file under 10 MiB
|
||||||
tail TEXT NOT NULL -- lowercase-hex SHA-256, last 1 KiB
|
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
|
||||||
|
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
|
||||||
) WITHOUT ROWID;
|
) WITHOUT ROWID;
|
||||||
```
|
```
|
||||||
|
|
||||||
Paths are stored as BLOBs because Unix paths are raw bytes, not
|
Paths are stored as BLOBs because Unix paths are raw bytes, not
|
||||||
guaranteed UTF-8. `mtime` is used only for change detection; it is
|
guaranteed UTF-8. `mtime` is used only for change detection; it is
|
||||||
not part of the duplicate key. `head` and `tail` are empty strings
|
not part of the duplicate key. For a file under 10 MiB `head`, `tail`,
|
||||||
when the file has never been hashed because its size was unique as
|
and `content` all hold the whole-file hash (that range is hashed in
|
||||||
of the last scan that covered it; such records still define the
|
full, with no end windows); for a larger file `head` and `tail` hold
|
||||||
file for tree reconstruction but never participate in duplicate
|
the first- and last-64 KiB hashes and `content` the whole-file or
|
||||||
groups.
|
sampled hash. All three are empty strings when the file has never
|
||||||
|
been hashed because its size was unique as of the last scan that
|
||||||
|
covered it; such records still define the file for tree
|
||||||
|
reconstruction but never participate in duplicate groups.
|
||||||
|
|
||||||
|
### Duplicate detection
|
||||||
|
|
||||||
|
Two files are duplicates only when they agree on every rung of this
|
||||||
|
ladder; a mismatch at any rung means they are not duplicates. `scan`
|
||||||
|
stores each file's hashes once, and `report` and `trees` group files by
|
||||||
|
the whole signature — size, `head`, `tail`, and `content` — so the
|
||||||
|
grouping is exactly this ladder applied across everything scanned into
|
||||||
|
the database, even across separate scans.
|
||||||
|
|
||||||
|
1. **Size.** Files of different sizes are never compared. Only files
|
||||||
|
whose size at least one other file shares are hashed at all.
|
||||||
|
2. **Under 10 MiB: whole file.** A file smaller than 10 MiB is hashed
|
||||||
|
in full and compared directly, with no separate end-window step —
|
||||||
|
small files are cheap to read to the last byte, and doing so makes
|
||||||
|
the comparison exact. `head`, `tail`, and `content` all hold this
|
||||||
|
whole-file SHA-256, so such a file's signature is decided entirely
|
||||||
|
by its size and its content.
|
||||||
|
3. **10 MiB and above: head and tail.** For a larger file, the SHA-256
|
||||||
|
of the first 64 KiB (`head`) and of the last 64 KiB (`tail`) are a
|
||||||
|
cheap gate that eliminates most same-size pairs before any bulk
|
||||||
|
reading. At 10 MiB and above the two windows never overlap.
|
||||||
|
4. **10 MiB and above, content below 50 MiB.** The SHA-256 of the
|
||||||
|
entire file. Agreement here is proof of identical content (barring a
|
||||||
|
SHA-256 collision).
|
||||||
|
5. **10 MiB and above, content 50 MiB and above.** A sampled SHA-256:
|
||||||
|
the 1 MiB window at each gigabyte-aligned offset (0, 1 GiB, 2 GiB, …
|
||||||
|
while inside the file, the final window truncated at end of file) is
|
||||||
|
fed, in order, into one hash. This is **deliberately probabilistic**
|
||||||
|
— the gaps between samples are never read, so two large files that
|
||||||
|
agree on every sample are reported as duplicates without being read
|
||||||
|
in full. It is the price of never reading a 150 GB file end to end.
|
||||||
|
Because size is already part of the signature, only equal-size files
|
||||||
|
reach this rung, so their sample boundaries always align.
|
||||||
|
|
||||||
|
`head`, `tail`, and `content` are one column each. A file below 10 MiB
|
||||||
|
and one at or above it never share a size, and neither do a file below
|
||||||
|
50 MiB and one at or above it, so a stored value is never ambiguous
|
||||||
|
between the whole-file, end-window, and sampled forms.
|
||||||
|
|
||||||
### `scan` mode
|
### `scan` mode
|
||||||
|
|
||||||
@@ -229,10 +288,9 @@ during the hash phase:
|
|||||||
decides its fate. Size-unique files are never read: new or
|
decides its fate. Size-unique files are never read: new or
|
||||||
changed ones are recorded without hashes in the update phase,
|
changed ones are recorded without hashes in the update phase,
|
||||||
unchanged unhashed ones simply keep their records. Every file
|
unchanged unhashed ones simply keep their records. Every file
|
||||||
with a shared size is hashed by the worker pool: read the first
|
with a shared size is hashed by the worker pool, computing the
|
||||||
`min(1024, size)` bytes and the last `min(1024, size)` bytes
|
full signature — head, tail, and content — described in "Duplicate
|
||||||
(one read when `size <= 1024`, since the two windows coincide)
|
detection" below. Zero-length files have constant
|
||||||
and compute the SHA-256 of each. Zero-length files have constant
|
|
||||||
hashes and are never opened. Files are hashed in **inode order**
|
hashes and are never opened. Files are hashed in **inode order**
|
||||||
(minimizing seeks on spinning disks), and paths that are hard
|
(minimizing seeks on spinning disks), and paths that are hard
|
||||||
links to the same inode are **read once**, all sharing the one
|
links to the same inode are **read once**, all sharing the one
|
||||||
|
|||||||
@@ -29,6 +29,24 @@
|
|||||||
|
|
||||||
# Completed Steps
|
# Completed Steps
|
||||||
|
|
||||||
|
- replace the 1 KiB end-window sampling with the head/tail plus
|
||||||
|
content-hash ladder (2026-09-22, branch `next`, closes
|
||||||
|
https://git.eeqj.de/sneak/sfdupes/issues/61): a file under 10 MiB is
|
||||||
|
hashed in full and compared directly, with no end-window step — its
|
||||||
|
`head`, `tail`, and `content` all hold the whole-file hash. A file at
|
||||||
|
10 MiB or above is gated on the 64 KiB `head` and `tail`, then
|
||||||
|
compared on a `content` hash — the whole file below 50 MiB,
|
||||||
|
gigabyte-spaced 1 MiB samples at or above. Schema bumps to version 2
|
||||||
|
(new `content` column); a version 1 database is rejected and must be
|
||||||
|
rescanned, which is required anyway since every stored hash changed.
|
||||||
|
`report` and `trees` group by the extended signature, so the ladder is
|
||||||
|
applied across the whole database. README "Duplicate detection"
|
||||||
|
documents every rung including the probabilistic large-file path.
|
||||||
|
|
||||||
|
- remove the dead `files.dat` references from `Makefile`, `.gitignore`
|
||||||
|
and `.dockerignore` (2026-09-21, branch `next`, closes
|
||||||
|
https://git.eeqj.de/sneak/sfdupes/issues/22)
|
||||||
|
|
||||||
- fix the lint-image pin comments and `FROM` form in `Dockerfile` and
|
- fix the lint-image pin comments and `FROM` form in `Dockerfile` and
|
||||||
`Dockerfile.lint` (2026-08-10, branch `next`, closes
|
`Dockerfile.lint` (2026-08-10, branch `next`, closes
|
||||||
https://git.eeqj.de/sneak/sfdupes/issues/25): dropped the false
|
https://git.eeqj.de/sneak/sfdupes/issues/25): dropped the false
|
||||||
|
|||||||
@@ -25,8 +25,11 @@ const defaultDatabasePath = "/var/lib/sfdupes/db.sqlite"
|
|||||||
const databaseEnv = "SFDUPES_DATABASE"
|
const databaseEnv = "SFDUPES_DATABASE"
|
||||||
|
|
||||||
// schemaVersion is the database schema version this build reads and
|
// schemaVersion is the database schema version this build reads and
|
||||||
// writes, stored in PRAGMA user_version.
|
// writes, stored in PRAGMA user_version. Version 2 adds the content
|
||||||
const schemaVersion = 1
|
// column and the head/tail/content signature (replacing the version 1
|
||||||
|
// 1 KiB end windows), so a version 1 database is rejected and must be
|
||||||
|
// rescanned.
|
||||||
|
const schemaVersion = 2
|
||||||
|
|
||||||
// dbDirPerm is the mode for a database parent directory created by
|
// dbDirPerm is the mode for a database parent directory created by
|
||||||
// scan.
|
// scan.
|
||||||
@@ -36,22 +39,24 @@ const dbDirPerm = 0o755
|
|||||||
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
|
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
|
||||||
const createTableSQL = `
|
const createTableSQL = `
|
||||||
CREATE TABLE files (
|
CREATE TABLE files (
|
||||||
path BLOB PRIMARY KEY,
|
path BLOB PRIMARY KEY,
|
||||||
size INTEGER NOT NULL,
|
size INTEGER NOT NULL,
|
||||||
mtime INTEGER NOT NULL,
|
mtime INTEGER NOT NULL,
|
||||||
head TEXT NOT NULL,
|
head TEXT NOT NULL,
|
||||||
tail TEXT NOT NULL
|
tail TEXT NOT NULL,
|
||||||
|
content TEXT NOT NULL
|
||||||
) WITHOUT ROWID
|
) WITHOUT ROWID
|
||||||
`
|
`
|
||||||
|
|
||||||
// upsertSQL inserts one file record, replacing any existing record for
|
// upsertSQL inserts one file record, replacing any existing record for
|
||||||
// the same path.
|
// the same path.
|
||||||
const upsertSQL = `
|
const upsertSQL = `
|
||||||
INSERT INTO files (path, size, mtime, head, tail)
|
INSERT INTO files (path, size, mtime, head, tail, content)
|
||||||
VALUES (?, ?, ?, ?, ?)
|
VALUES (?, ?, ?, ?, ?, ?)
|
||||||
ON CONFLICT (path) DO UPDATE SET
|
ON CONFLICT (path) DO UPDATE SET
|
||||||
size = excluded.size, mtime = excluded.mtime,
|
size = excluded.size, mtime = excluded.mtime,
|
||||||
head = excluded.head, tail = excluded.tail
|
head = excluded.head, tail = excluded.tail,
|
||||||
|
content = excluded.content
|
||||||
`
|
`
|
||||||
|
|
||||||
// errNoDatabase reports a missing database file for report/trees.
|
// errNoDatabase reports a missing database file for report/trees.
|
||||||
@@ -205,7 +210,7 @@ func userVersion(ctx context.Context, db *sql.DB) (int, error) {
|
|||||||
// loadFileRows reads every record from the files table.
|
// loadFileRows reads every record from the files table.
|
||||||
func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) {
|
func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) {
|
||||||
rows, err := db.QueryContext(ctx,
|
rows, err := db.QueryContext(ctx,
|
||||||
"SELECT path, size, mtime, head, tail FROM files")
|
"SELECT path, size, mtime, head, tail, content FROM files")
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, fmt.Errorf("read records: %w", err)
|
return nil, fmt.Errorf("read records: %w", err)
|
||||||
}
|
}
|
||||||
@@ -220,7 +225,8 @@ func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) {
|
|||||||
r scanRec
|
r scanRec
|
||||||
)
|
)
|
||||||
|
|
||||||
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail)
|
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail,
|
||||||
|
&r.content)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, fmt.Errorf("read record: %w", err)
|
return nil, fmt.Errorf("read record: %w", err)
|
||||||
}
|
}
|
||||||
@@ -348,7 +354,7 @@ func execUpserts(ctx context.Context, tx *sql.Tx, upserts []scanRec,
|
|||||||
|
|
||||||
for _, r := range upserts {
|
for _, r := range upserts {
|
||||||
_, err = st.ExecContext(ctx,
|
_, err = st.ExecContext(ctx,
|
||||||
[]byte(r.path), r.size, r.mtime, r.head, r.tail)
|
[]byte(r.path), r.size, r.mtime, r.head, r.tail, r.content)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return fmt.Errorf("upsert %s: %w", r.path, err)
|
return fmt.Errorf("upsert %s: %w", r.path, err)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -17,14 +17,15 @@ const ioBufSize = 1 << 20
|
|||||||
const minGroupSize = 2
|
const minGroupSize = 2
|
||||||
|
|
||||||
// scanRec is one file record from the database. The signature (size,
|
// scanRec is one file record from the database. The signature (size,
|
||||||
// head, tail) is the duplicate key; mtime is informational only and
|
// head, tail, content) is the duplicate key; mtime is informational
|
||||||
// used by scan for change detection.
|
// only and used by scan for change detection.
|
||||||
type scanRec struct {
|
type scanRec struct {
|
||||||
size int64
|
size int64
|
||||||
mtime int64
|
mtime int64
|
||||||
head string
|
head string
|
||||||
tail string
|
tail string
|
||||||
path string
|
content string
|
||||||
|
path string
|
||||||
}
|
}
|
||||||
|
|
||||||
// loadRecords opens the database and reads every file record for the
|
// loadRecords opens the database and reads every file record for the
|
||||||
@@ -52,8 +53,9 @@ func loadRecords(ctx context.Context) ([]scanRec, error) {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// dupeGroup is one set of candidate-duplicate files: identical size,
|
// dupeGroup is one set of candidate-duplicate files: identical size,
|
||||||
// head hash, and tail hash. paths is sorted lexicographically; the
|
// head hash, tail hash, and content hash. paths is sorted
|
||||||
// first entry is the group's "first", the rest are dupes.
|
// lexicographically; the first entry is the group's "first", the rest
|
||||||
|
// are dupes.
|
||||||
type dupeGroup struct {
|
type dupeGroup struct {
|
||||||
size int64
|
size int64
|
||||||
paths []string
|
paths []string
|
||||||
@@ -122,7 +124,9 @@ func collectDupeGroups(recs []scanRec) []dupeGroup {
|
|||||||
continue
|
continue
|
||||||
}
|
}
|
||||||
|
|
||||||
k := fileSig{size: r.size, head: r.head, tail: r.tail}
|
k := fileSig{
|
||||||
|
size: r.size, head: r.head, tail: r.tail, content: r.content,
|
||||||
|
}
|
||||||
groups[k] = append(groups[k], r.path)
|
groups[k] = append(groups[k], r.path)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
@@ -38,6 +38,28 @@ func TestCollectDupeGroups(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestCollectDupeGroupsContentSeparates(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
// Same size, head, and tail, but different content hashes: the final
|
||||||
|
// rung keeps them apart, so no group forms. Matching content groups.
|
||||||
|
recs := []scanRec{
|
||||||
|
{size: 100, head: "h", tail: "t", content: "c1", path: "/a"},
|
||||||
|
{size: 100, head: "h", tail: "t", content: "c2", path: "/b"},
|
||||||
|
{size: 100, head: "h", tail: "t", content: "c1", path: "/c"},
|
||||||
|
}
|
||||||
|
|
||||||
|
groups := collectDupeGroups(recs)
|
||||||
|
if len(groups) != 1 {
|
||||||
|
t.Fatalf("len(groups) = %d, want 1 (only the matching content)",
|
||||||
|
len(groups))
|
||||||
|
}
|
||||||
|
|
||||||
|
if !slices.Equal(groups[0].paths, []string{"/a", "/c"}) {
|
||||||
|
t.Errorf("group paths = %q, want /a /c", groups[0].paths)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestCollectDupeGroupsMtimeExcluded(t *testing.T) {
|
func TestCollectDupeGroupsMtimeExcluded(t *testing.T) {
|
||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
|
|||||||
@@ -7,6 +7,7 @@ import (
|
|||||||
"database/sql"
|
"database/sql"
|
||||||
"encoding/hex"
|
"encoding/hex"
|
||||||
"fmt"
|
"fmt"
|
||||||
|
"io"
|
||||||
"io/fs"
|
"io/fs"
|
||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
@@ -16,8 +17,37 @@ import (
|
|||||||
"syscall"
|
"syscall"
|
||||||
)
|
)
|
||||||
|
|
||||||
// chunk is the number of bytes hashed from each end of a file.
|
// The duplicate ladder (see hashSignature and README "Duplicate
|
||||||
const chunk = 1024
|
// detection"). A same-size candidate below headTailMin is hashed in
|
||||||
|
// full and compared directly; a larger one is separated first by the
|
||||||
|
// hashes of its end windows, then by a content hash that is exact below
|
||||||
|
// wholeFileMax and deliberately sampled at or above it.
|
||||||
|
|
||||||
|
// headTailMin is the size threshold for the end-window gate. A file
|
||||||
|
// smaller than this is hashed in full directly, with no separate head
|
||||||
|
// and tail step: its head, tail, and content all carry the whole-file
|
||||||
|
// hash. A file this size or larger is separated first by its end
|
||||||
|
// windows.
|
||||||
|
const headTailMin = 10 * 1024 * 1024
|
||||||
|
|
||||||
|
// headTailWindow is the number of bytes hashed from each end of a file
|
||||||
|
// at or above headTailMin (the head and tail rungs). Because
|
||||||
|
// headTailMin is far larger than two windows, the head and tail windows
|
||||||
|
// never overlap.
|
||||||
|
const headTailWindow = 64 * 1024
|
||||||
|
|
||||||
|
// wholeFileMax is the size boundary between the two content rungs: a
|
||||||
|
// file strictly smaller than this is content-hashed in full; a file
|
||||||
|
// this size or larger is content-hashed by sampling.
|
||||||
|
const wholeFileMax = 50 * 1024 * 1024
|
||||||
|
|
||||||
|
// sampleStride is the spacing between content samples for large files:
|
||||||
|
// one window is read at each gigabyte-aligned offset (0, 1 GiB, ...).
|
||||||
|
const sampleStride = 1024 * 1024 * 1024
|
||||||
|
|
||||||
|
// sampleWindow is the number of bytes read at each large-file sample
|
||||||
|
// offset, truncated at end of file.
|
||||||
|
const sampleWindow = 1024 * 1024
|
||||||
|
|
||||||
// workQueueDepth bounds the job and result channels feeding the walk
|
// workQueueDepth bounds the job and result channels feeding the walk
|
||||||
// and hash worker pools.
|
// and hash worker pools.
|
||||||
@@ -439,11 +469,12 @@ func (s *scanState) recordRun(ctx context.Context, r hashResult) error {
|
|||||||
s.resolve(rec.path)
|
s.resolve(rec.path)
|
||||||
|
|
||||||
s.batch = append(s.batch, scanRec{
|
s.batch = append(s.batch, scanRec{
|
||||||
size: rec.size,
|
size: rec.size,
|
||||||
mtime: rec.mtime,
|
mtime: rec.mtime,
|
||||||
head: r.head,
|
head: r.head,
|
||||||
tail: r.tail,
|
tail: r.tail,
|
||||||
path: rec.path,
|
content: r.content,
|
||||||
|
path: rec.path,
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -834,13 +865,15 @@ func inodeOfInfo(fi fs.FileInfo) (uint64, uint64) {
|
|||||||
return statDev(st), st.Ino
|
return statDev(st), st.Ino
|
||||||
}
|
}
|
||||||
|
|
||||||
// hashResult carries one inode run's head/tail hashes (or the error
|
// hashResult carries one inode run's signature hashes — head, tail, and
|
||||||
// that prevented hashing it) from the hash workers to the hash phase.
|
// content — (or the error that prevented hashing it) from the hash
|
||||||
|
// workers to the hash phase.
|
||||||
type hashResult struct {
|
type hashResult struct {
|
||||||
run []fileRec
|
run []fileRec
|
||||||
head string
|
head string
|
||||||
tail string
|
tail string
|
||||||
err error
|
content string
|
||||||
|
err error
|
||||||
}
|
}
|
||||||
|
|
||||||
// hashPool owns every goroutine of the hash worker pool: the feeder
|
// hashPool owns every goroutine of the hash worker pool: the feeder
|
||||||
@@ -930,11 +963,11 @@ func hashWorker(ctx context.Context, jobs <-chan []fileRec,
|
|||||||
continue
|
continue
|
||||||
}
|
}
|
||||||
|
|
||||||
head, tail, err := hashHeadTail(run[0].path, run[0].size)
|
head, tail, content, err := hashSignature(run[0].path, run[0].size)
|
||||||
|
|
||||||
select {
|
select {
|
||||||
case results <- hashResult{
|
case results <- hashResult{
|
||||||
run: run, head: head, tail: tail, err: err,
|
run: run, head: head, tail: tail, content: content, err: err,
|
||||||
}:
|
}:
|
||||||
case <-ctx.Done():
|
case <-ctx.Done():
|
||||||
return
|
return
|
||||||
@@ -942,55 +975,138 @@ func hashWorker(ctx context.Context, jobs <-chan []fileRec,
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// emptyHash is the lowercase-hex SHA-256 of the empty input: the head
|
// emptyHash is the lowercase-hex SHA-256 of the empty input: the head,
|
||||||
// and tail hash of every zero-length file.
|
// tail, and content hash of every zero-length file.
|
||||||
const emptyHash = "e3b0c44298fc1c149afbf4c8996fb924" +
|
const emptyHash = "e3b0c44298fc1c149afbf4c8996fb924" +
|
||||||
"27ae41e4649b934ca495991b7852b855"
|
"27ae41e4649b934ca495991b7852b855"
|
||||||
|
|
||||||
// hashHeadTail returns the lowercase-hex SHA-256 of the first
|
// hashSignature computes the three content hashes that, with the file
|
||||||
// min(chunk, size) bytes and of the last min(chunk, size) bytes of the
|
// size, form its duplicate signature. A file below headTailMin is
|
||||||
// file at path. The two reads overlap when size < 2*chunk. size is the
|
// hashed in full and its whole-file SHA-256 is returned as head, tail,
|
||||||
// value recorded when the file was statted; a zero-length file's
|
// and content alike — that range takes no separate end-window step. For
|
||||||
// hashes are constant, so it is never even opened.
|
// a file at or above headTailMin the head and tail are the SHA-256 of
|
||||||
func hashHeadTail(path string, size int64) (string, string, error) {
|
// its first and last headTailWindow bytes, and content is the SHA-256
|
||||||
|
// of the whole file below wholeFileMax (the exact rung) or of
|
||||||
|
// gigabyte-spaced samples at or above it (the sampled, deliberately
|
||||||
|
// probabilistic rung). Two files are duplicates only when all four
|
||||||
|
// agree; any mismatch means not a duplicate. size is the value recorded
|
||||||
|
// when the file was statted; a zero-length file has constant hashes and
|
||||||
|
// is never opened.
|
||||||
|
func hashSignature(path string, size int64) (string, string, string, error) {
|
||||||
if size == 0 {
|
if size == 0 {
|
||||||
return emptyHash, emptyHash, nil
|
return emptyHash, emptyHash, emptyHash, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
//nolint:gosec // hashing operator-supplied paths is the tool's purpose
|
//nolint:gosec // hashing operator-supplied paths is the tool's purpose
|
||||||
f, err := os.Open(path)
|
f, err := os.Open(path)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return "", "", err
|
return "", "", "", err
|
||||||
}
|
}
|
||||||
|
|
||||||
defer func() { _ = f.Close() }()
|
defer func() { _ = f.Close() }()
|
||||||
|
|
||||||
n := min(int64(chunk), size)
|
// Below the threshold the whole file is hashed directly, with no
|
||||||
|
// end-window step: head and tail both carry the whole-file hash.
|
||||||
|
if size < int64(headTailMin) {
|
||||||
|
content, err := hashWhole(f, size)
|
||||||
|
if err != nil {
|
||||||
|
return "", "", "", err
|
||||||
|
}
|
||||||
|
|
||||||
buf := make([]byte, n)
|
return content, content, content, nil
|
||||||
|
}
|
||||||
|
|
||||||
_, err = f.ReadAt(buf, 0)
|
head, tail, err := hashEnds(f, size)
|
||||||
|
if err != nil {
|
||||||
|
return "", "", "", err
|
||||||
|
}
|
||||||
|
|
||||||
|
content, err := hashContent(f, size)
|
||||||
|
if err != nil {
|
||||||
|
return "", "", "", err
|
||||||
|
}
|
||||||
|
|
||||||
|
return head, tail, content, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// hashEnds returns the SHA-256 of the first and last headTailWindow
|
||||||
|
// bytes of f. It is called only for files at least headTailMin, which
|
||||||
|
// is far larger than two windows, so the windows never overlap and both
|
||||||
|
// reads are always full.
|
||||||
|
func hashEnds(f *os.File, size int64) (string, string, error) {
|
||||||
|
buf := make([]byte, headTailWindow)
|
||||||
|
|
||||||
|
_, err := f.ReadAt(buf, 0)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return "", "", err
|
return "", "", err
|
||||||
}
|
}
|
||||||
|
|
||||||
h := sha256.Sum256(buf)
|
h := sha256.Sum256(buf)
|
||||||
|
head := hex.EncodeToString(h[:])
|
||||||
|
|
||||||
// When the whole file fits in one chunk the tail window is exactly
|
_, err = f.ReadAt(buf, size-int64(headTailWindow))
|
||||||
// the bytes just read: reuse the head hash instead of issuing a
|
|
||||||
// second read for every small file.
|
|
||||||
if size <= int64(chunk) {
|
|
||||||
hh := hex.EncodeToString(h[:])
|
|
||||||
|
|
||||||
return hh, hh, nil
|
|
||||||
}
|
|
||||||
|
|
||||||
_, err = f.ReadAt(buf, size-n)
|
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return "", "", err
|
return "", "", err
|
||||||
}
|
}
|
||||||
|
|
||||||
t := sha256.Sum256(buf)
|
t := sha256.Sum256(buf)
|
||||||
|
|
||||||
return hex.EncodeToString(h[:]), hex.EncodeToString(t[:]), nil
|
return head, hex.EncodeToString(t[:]), nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// hashContent returns the content-rung hash of f: the SHA-256 of the
|
||||||
|
// whole file when it is smaller than wholeFileMax, or of sampled
|
||||||
|
// windows when it is that size or larger.
|
||||||
|
func hashContent(f *os.File, size int64) (string, error) {
|
||||||
|
if size >= int64(wholeFileMax) {
|
||||||
|
return hashSamples(f, size)
|
||||||
|
}
|
||||||
|
|
||||||
|
return hashWhole(f, size)
|
||||||
|
}
|
||||||
|
|
||||||
|
// hashWhole returns the SHA-256 of the entire file. A SectionReader is
|
||||||
|
// used so the read is independent of the offset left by any end-window
|
||||||
|
// reads. Reading fewer than size bytes means the file shrank between
|
||||||
|
// the stat and the hash; that is an error rather than a hash of content
|
||||||
|
// that no longer matches the recorded size.
|
||||||
|
func hashWhole(f *os.File, size int64) (string, error) {
|
||||||
|
h := sha256.New()
|
||||||
|
|
||||||
|
n, err := io.Copy(h, io.NewSectionReader(f, 0, size))
|
||||||
|
if err != nil {
|
||||||
|
return "", err
|
||||||
|
}
|
||||||
|
|
||||||
|
if n != size {
|
||||||
|
return "", fmt.Errorf("read %d of %d bytes: %w", n, size,
|
||||||
|
io.ErrUnexpectedEOF)
|
||||||
|
}
|
||||||
|
|
||||||
|
return hex.EncodeToString(h.Sum(nil)), nil
|
||||||
|
}
|
||||||
|
|
||||||
|
// hashSamples feeds sampleWindow bytes at each gigabyte-aligned offset
|
||||||
|
// (0, sampleStride, 2*sampleStride, ... while inside the file), in
|
||||||
|
// order, into one hash, each window truncated at end of file. This is
|
||||||
|
// the probabilistic large-file rung: two files of equal size agreeing
|
||||||
|
// on every sample are reported as duplicates without every byte being
|
||||||
|
// read. Because size is part of the signature, files of different sizes
|
||||||
|
// never reach this comparison, so the sample boundaries always align.
|
||||||
|
func hashSamples(f *os.File, size int64) (string, error) {
|
||||||
|
h := sha256.New()
|
||||||
|
buf := make([]byte, sampleWindow)
|
||||||
|
|
||||||
|
for off := int64(0); off < size; off += int64(sampleStride) {
|
||||||
|
n := min(int64(sampleWindow), size-off)
|
||||||
|
|
||||||
|
_, err := f.ReadAt(buf[:n], off)
|
||||||
|
if err != nil {
|
||||||
|
return "", err
|
||||||
|
}
|
||||||
|
|
||||||
|
h.Write(buf[:n])
|
||||||
|
}
|
||||||
|
|
||||||
|
return hex.EncodeToString(h.Sum(nil)), nil
|
||||||
}
|
}
|
||||||
|
|||||||
+229
-25
@@ -53,7 +53,23 @@ func pattern(tag byte, n int) []byte {
|
|||||||
return data
|
return data
|
||||||
}
|
}
|
||||||
|
|
||||||
func TestHashHeadTail(t *testing.T) {
|
// sig returns a file's full signature (head, tail, content), failing the
|
||||||
|
// test on any error.
|
||||||
|
func sig(t *testing.T, path string, size int64) (string, string, string) {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
head, tail, content, err := hashSignature(path, size)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("hashSignature %s: %v", path, err)
|
||||||
|
}
|
||||||
|
|
||||||
|
return head, tail, content
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestHashSignatureBelowThreshold verifies that a file below headTailMin
|
||||||
|
// is hashed in full and compared directly: head, tail, and content all
|
||||||
|
// carry the whole-file SHA-256, with no separate end-window step.
|
||||||
|
func TestHashSignatureBelowThreshold(t *testing.T) {
|
||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
dir := t.TempDir()
|
dir := t.TempDir()
|
||||||
@@ -62,13 +78,10 @@ func TestHashHeadTail(t *testing.T) {
|
|||||||
name string
|
name string
|
||||||
data []byte
|
data []byte
|
||||||
}{
|
}{
|
||||||
{"empty", nil},
|
|
||||||
{"one-byte", []byte("x")},
|
{"one-byte", []byte("x")},
|
||||||
{"under-one-chunk", pattern(1, chunk-1)},
|
{"one-window", pattern(1, headTailWindow)},
|
||||||
{"exactly-one-chunk", pattern(2, chunk)},
|
{"several-windows", pattern(2, 3*headTailWindow)},
|
||||||
{"overlapping-reads", pattern(3, chunk+chunk/2)},
|
{"near-threshold", pattern(3, headTailMin-1)},
|
||||||
{"exactly-two-chunks", pattern(4, 2*chunk)},
|
|
||||||
{"beyond-two-chunks", pattern(5, 3*chunk)},
|
|
||||||
}
|
}
|
||||||
for _, c := range cases {
|
for _, c := range cases {
|
||||||
t.Run(c.name, func(t *testing.T) {
|
t.Run(c.name, func(t *testing.T) {
|
||||||
@@ -76,49 +89,240 @@ func TestHashHeadTail(t *testing.T) {
|
|||||||
|
|
||||||
p := writeFile(t, dir, c.name, c.data)
|
p := writeFile(t, dir, c.name, c.data)
|
||||||
|
|
||||||
head, tail, err := hashHeadTail(p, int64(len(c.data)))
|
head, tail, content := sig(t, p, int64(len(c.data)))
|
||||||
if err != nil {
|
|
||||||
t.Fatalf("hashHeadTail: %v", err)
|
|
||||||
}
|
|
||||||
|
|
||||||
n := min(chunk, len(c.data))
|
whole := hexSum(c.data)
|
||||||
if want := hexSum(c.data[:n]); head != want {
|
if head != whole || tail != whole || content != whole {
|
||||||
t.Errorf("head = %s, want %s", head, want)
|
t.Errorf("head=%s tail=%s content=%s, want all whole-file %s",
|
||||||
}
|
head, tail, content, whole)
|
||||||
|
|
||||||
if want := hexSum(c.data[len(c.data)-n:]); tail != want {
|
|
||||||
t.Errorf("tail = %s, want %s", tail, want)
|
|
||||||
}
|
}
|
||||||
})
|
})
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
func TestHashHeadTailErrors(t *testing.T) {
|
// TestHashSignatureEnds exercises the head and tail rungs, which apply
|
||||||
|
// only to files at least headTailMin. Sparse files keep the fixtures
|
||||||
|
// cheap: a difference in the first window changes only head, a
|
||||||
|
// difference in the last window changes only tail, and a difference
|
||||||
|
// between the windows changes neither end hash but does change the
|
||||||
|
// whole-file content rung (the file is below wholeFileMax).
|
||||||
|
func TestHashSignatureEnds(t *testing.T) {
|
||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
dir := t.TempDir()
|
dir := t.TempDir()
|
||||||
|
|
||||||
_, _, err := hashHeadTail(filepath.Join(dir, "missing"), 1)
|
// Between headTailMin and wholeFileMax: the end-window gate is active
|
||||||
|
// and the content rung is a whole-file hash.
|
||||||
|
const size = int64(headTailMin + 2*1024*1024)
|
||||||
|
|
||||||
|
base := sparseFile(t, dir, "ends-base", size)
|
||||||
|
headDiff := sparseFile(t, dir, "ends-head", size)
|
||||||
|
tailDiff := sparseFile(t, dir, "ends-tail", size)
|
||||||
|
midDiff := sparseFile(t, dir, "ends-mid", size)
|
||||||
|
|
||||||
|
pokeAt(t, headDiff, 0, []byte{1})
|
||||||
|
pokeAt(t, tailDiff, size-1, []byte{1})
|
||||||
|
pokeAt(t, midDiff, size/2, []byte{1})
|
||||||
|
|
||||||
|
bHead, bTail, bContent := sig(t, base, size)
|
||||||
|
|
||||||
|
h, tl, c := sig(t, headDiff, size)
|
||||||
|
if h == bHead {
|
||||||
|
t.Error("a byte in the first window did not change head")
|
||||||
|
}
|
||||||
|
|
||||||
|
if tl != bTail {
|
||||||
|
t.Error("a byte in the first window changed tail")
|
||||||
|
}
|
||||||
|
|
||||||
|
if c == bContent {
|
||||||
|
t.Error("a byte in the first window did not change content")
|
||||||
|
}
|
||||||
|
|
||||||
|
h, tl, c = sig(t, tailDiff, size)
|
||||||
|
if tl == bTail {
|
||||||
|
t.Error("a byte in the last window did not change tail")
|
||||||
|
}
|
||||||
|
|
||||||
|
if h != bHead {
|
||||||
|
t.Error("a byte in the last window changed head")
|
||||||
|
}
|
||||||
|
|
||||||
|
if c == bContent {
|
||||||
|
t.Error("a byte in the last window did not change content")
|
||||||
|
}
|
||||||
|
|
||||||
|
h, tl, c = sig(t, midDiff, size)
|
||||||
|
if h != bHead || tl != bTail {
|
||||||
|
t.Error("a byte between the windows changed an end hash")
|
||||||
|
}
|
||||||
|
|
||||||
|
if c == bContent {
|
||||||
|
t.Error("whole-file content rung ignored a byte between the windows")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestHashSignatureErrors(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
dir := t.TempDir()
|
||||||
|
|
||||||
|
// A missing file: an error, and every hash left empty.
|
||||||
|
head, tail, content, err := hashSignature(filepath.Join(dir, "missing"), 1)
|
||||||
if err == nil {
|
if err == nil {
|
||||||
t.Error("no error for a missing file")
|
t.Error("no error for a missing file")
|
||||||
}
|
}
|
||||||
|
|
||||||
|
if head != "" || tail != "" || content != "" {
|
||||||
|
t.Errorf("missing file returned hashes: %q %q %q", head, tail, content)
|
||||||
|
}
|
||||||
|
|
||||||
// A zero-length file has constant hashes and is never opened: even
|
// A zero-length file has constant hashes and is never opened: even
|
||||||
// a missing path succeeds.
|
// a missing path succeeds.
|
||||||
head, tail, err := hashHeadTail(filepath.Join(dir, "missing"), 0)
|
head, tail, content, err = hashSignature(filepath.Join(dir, "missing"), 0)
|
||||||
if err != nil || head != emptyHash || tail != emptyHash {
|
if err != nil ||
|
||||||
t.Errorf("empty: head=%q tail=%q err=%v, want constant hashes",
|
head != emptyHash || tail != emptyHash || content != emptyHash {
|
||||||
head, tail, err)
|
t.Errorf("empty: head=%q tail=%q content=%q err=%v, "+
|
||||||
|
"want constant hashes", head, tail, content, err)
|
||||||
}
|
}
|
||||||
|
|
||||||
// A file that shrank between the stat and hash passes: reading at
|
// A file that shrank between the stat and hash passes: reading at
|
||||||
// the stat-reported size must fail rather than emit wrong hashes.
|
// the stat-reported size must fail rather than emit wrong hashes.
|
||||||
p := writeFile(t, dir, "shrunk", []byte("tiny"))
|
p := writeFile(t, dir, "shrunk", []byte("tiny"))
|
||||||
|
|
||||||
_, _, err = hashHeadTail(p, int64(2*chunk))
|
head, tail, content, err = hashSignature(p, int64(2*headTailWindow))
|
||||||
if err == nil {
|
if err == nil {
|
||||||
t.Error("no error when the stat size exceeds the file size")
|
t.Error("no error when the stat size exceeds the file size")
|
||||||
}
|
}
|
||||||
|
|
||||||
|
if head != "" || tail != "" || content != "" {
|
||||||
|
t.Errorf("shrunk file returned hashes: %q %q %q", head, tail, content)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// sparseFile creates a file that is logically size bytes long without
|
||||||
|
// allocating blocks for the hole, so multi-gigabyte cases stay cheap.
|
||||||
|
func sparseFile(t *testing.T, dir, name string, size int64) string {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
p := filepath.Join(dir, name)
|
||||||
|
|
||||||
|
f, err := os.Create(p) //nolint:gosec // test-controlled path
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
err = f.Truncate(size)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
err = f.Close()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
// pokeAt writes data into an existing file at off, leaving the rest of
|
||||||
|
// the file (a sparse hole) untouched.
|
||||||
|
func pokeAt(t *testing.T, path string, off int64, data []byte) {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
f, err := os.OpenFile(path, os.O_WRONLY, 0o600) //nolint:gosec // test path
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
_, err = f.WriteAt(data, off)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
err = f.Close()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// contentHash returns just the content rung of a file's signature.
|
||||||
|
func contentHash(t *testing.T, path string, size int64) string {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
_, _, content := sig(t, path, size)
|
||||||
|
|
||||||
|
return content
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestContentRungBoundary checks the 50 MiB boundary between the two
|
||||||
|
// content rungs: just below it the whole file is hashed and any byte
|
||||||
|
// difference shows; at the boundary only the gigabyte-spaced samples are
|
||||||
|
// hashed, so a difference outside a sample window is invisible.
|
||||||
|
func TestContentRungBoundary(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
dir := t.TempDir()
|
||||||
|
|
||||||
|
// A byte that lands outside the single [0, sampleWindow) sample a
|
||||||
|
// sub-gigabyte file has, but well inside the file.
|
||||||
|
const off = 10 * 1024 * 1024
|
||||||
|
|
||||||
|
// Just under the boundary: the whole-file rung sees the poked byte.
|
||||||
|
under := int64(wholeFileMax - 1)
|
||||||
|
underBase := sparseFile(t, dir, "under-base", under)
|
||||||
|
underPoked := sparseFile(t, dir, "under-poked", under)
|
||||||
|
|
||||||
|
pokeAt(t, underPoked, off, []byte{1})
|
||||||
|
|
||||||
|
if contentHash(t, underBase, under) == contentHash(t, underPoked, under) {
|
||||||
|
t.Error("whole-file rung ignored a byte difference below wholeFileMax")
|
||||||
|
}
|
||||||
|
|
||||||
|
// At the boundary: only [0, sampleWindow) is sampled, so the poked
|
||||||
|
// byte at off is invisible and the two content hashes match.
|
||||||
|
at := int64(wholeFileMax)
|
||||||
|
atBase := sparseFile(t, dir, "at-base", at)
|
||||||
|
atPoked := sparseFile(t, dir, "at-poked", at)
|
||||||
|
|
||||||
|
pokeAt(t, atPoked, off, []byte{1})
|
||||||
|
|
||||||
|
if contentHash(t, atBase, at) != contentHash(t, atPoked, at) {
|
||||||
|
t.Error("sampled rung saw a byte outside every sample window")
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestContentRungMultiGigabyte exercises the sampled rung across several
|
||||||
|
// gigabytes using sparse files: a difference inside the third sample
|
||||||
|
// window (at offset 2*sampleStride) changes the hash, while a difference
|
||||||
|
// in the gap after it does not.
|
||||||
|
func TestContentRungMultiGigabyte(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
dir := t.TempDir()
|
||||||
|
|
||||||
|
// Three sample windows (offsets 0, 1 GiB, 2 GiB) plus a trailing gap
|
||||||
|
// that no sample covers.
|
||||||
|
size := int64(2*sampleStride + 2*sampleWindow)
|
||||||
|
thirdSample := int64(2 * sampleStride)
|
||||||
|
gap := thirdSample + int64(sampleWindow)
|
||||||
|
|
||||||
|
base := sparseFile(t, dir, "g-base", size)
|
||||||
|
inSample := sparseFile(t, dir, "g-insample", size)
|
||||||
|
inGap := sparseFile(t, dir, "g-ingap", size)
|
||||||
|
|
||||||
|
pokeAt(t, inSample, thirdSample, []byte{1})
|
||||||
|
pokeAt(t, inGap, gap, []byte{1})
|
||||||
|
|
||||||
|
baseHash := contentHash(t, base, size)
|
||||||
|
|
||||||
|
if contentHash(t, inSample, size) == baseHash {
|
||||||
|
t.Error("sample at 2 GiB was not read: difference there was invisible")
|
||||||
|
}
|
||||||
|
|
||||||
|
if contentHash(t, inGap, size) != baseHash {
|
||||||
|
t.Error("a byte in an unsampled gap changed the content hash")
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// collectWalk runs a walk over roots and returns the emitted records
|
// collectWalk runs a walk over roots and returns the emitted records
|
||||||
|
|||||||
@@ -13,9 +13,10 @@ import (
|
|||||||
|
|
||||||
// fileSig is a file's duplicate signature; mtime is excluded.
|
// fileSig is a file's duplicate signature; mtime is excluded.
|
||||||
type fileSig struct {
|
type fileSig struct {
|
||||||
size int64
|
size int64
|
||||||
head string
|
head string
|
||||||
tail string
|
tail string
|
||||||
|
content string
|
||||||
}
|
}
|
||||||
|
|
||||||
// treeNode is one directory reconstructed from the scan stream.
|
// treeNode is one directory reconstructed from the scan stream.
|
||||||
@@ -122,7 +123,9 @@ func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
|
|||||||
node.files = make(map[string]fileSig)
|
node.files = make(map[string]fileSig)
|
||||||
}
|
}
|
||||||
|
|
||||||
sig := fileSig{size: r.size, head: r.head, tail: r.tail}
|
sig := fileSig{
|
||||||
|
size: r.size, head: r.head, tail: r.tail, content: r.content,
|
||||||
|
}
|
||||||
|
|
||||||
// An unhashed record (its size was unique when last scanned)
|
// An unhashed record (its size was unique when last scanned)
|
||||||
// has unknown content: give it a signature no other file can
|
// has unknown content: give it a signature no other file can
|
||||||
@@ -188,7 +191,7 @@ func (n *treeNode) compute() {
|
|||||||
for name, sig := range n.files {
|
for name, sig := range n.files {
|
||||||
entries = append(entries,
|
entries = append(entries,
|
||||||
"f\x00"+name+"\x00"+strconv.FormatInt(sig.size, 10)+
|
"f\x00"+name+"\x00"+strconv.FormatInt(sig.size, 10)+
|
||||||
"\x00"+sig.head+"\x00"+sig.tail)
|
"\x00"+sig.head+"\x00"+sig.tail+"\x00"+sig.content)
|
||||||
n.fileCount++
|
n.fileCount++
|
||||||
n.totalSize += sig.size
|
n.totalSize += sig.size
|
||||||
}
|
}
|
||||||
|
|||||||
Reference in New Issue
Block a user