Author SHA1 Message Date
sneak bd8d41b174 Stream report and trees instead of loading every record (closes #14)
check / check (push) Successful in 1m49s
report now has SQLite group the records and put the rows in report
order, helped by a new files_signature index on (size, head, tail,
content), and writes each row as it reads it. trees reads the records
in path order, where all the paths under a directory come together, so
it computes each directory's digest as soon as the stream leaves it and
keeps only its path, parent, digest and totals. Output is unchanged.

The tests that called the removed in-memory grouping functions now group
records stored in a database. A new test checks that both commands give
the same output whatever order the records were inserted in.

Model: opus-5-5
2026-10-04 02:04:37 +00:00
clawbot 2dd1194f33 Warn about and skip non-regular and .zfs operands, keeping their records (closes #9)
check / check (push) Successful in 2m48s
A symlink, socket, FIFO or device-node operand, or a directory operand
named .zfs, was silently ignored yet stayed in the scanned operands, so
the update phase deleted every record stored beneath it. Such an
operand now gets a one-line warning, counts as skipped, and is dropped
before overlapping operands are pruned and the database index is
loaded: another operand beneath it is still scanned, and the records
beneath it count as outside the scanned operands and are not deleted,
unless it lies under another operand. The exit status stays 0. An
operand that turns into one of these after that check is warned about
and skipped by the walk instead. README "scan mode" and "Rules for the
walk" say so.

Model: opus-5-5
2026-10-04 04:01:43 +02:00
clawbot 705c8729ca Hold a lock so a second scan fails at once (closes #53)
check / check (push) Successful in 1m43s
scan takes an exclusive flock(2) on a lock file beside the database
(its path with .lock appended) before it walks anything or opens the
database, and holds it until it returns. A second scan against the
same database fails at once with a one-line error naming the lock
file and exits 1. report and trees never take the lock. The lock ends
with the process, so a fatal error or an interrupt releases it; the
file is never deleted. golang.org/x/sys becomes a direct dependency.

The README smoke test now keeps the database outside the scanned
tree, where its empty lock file would have joined the empty-file
group.

Model: opus-5-5
2026-10-04 02:30:21 +02:00
clawbot 01ff3bb5f0 Test report and trees stdout write failures (closes #30)
check / check (push) Successful in 1m34s
report and trees already checked every stdout write and the final
flush. run now takes the stdout it hands to them, so tests pass a
closed file or a failing writer instead of swapping os.Stdout: a
closed stdout exits 1 with a one-line diagnostic, and the writer's
error reaches the caller.

README "Error handling" now states the two cases that never reach
sfdupes as a failed write: a pipe reader that exits early ends the
process with SIGPIPE, as with cat; and stdout closed with >&- is
replaced by /dev/null by the Go runtime before main runs, so the run
succeeds.

Model: opus-5-5
2026-10-03 18:01:27 +02:00
clawbot d63d3cc7fc Open the database read-only for report and trees (closes #8)
check / check (push) Successful in 1m17s
report and trees now connect read-only (mode=ro, query_only, the same
busy timeout) and no longer set the journal mode, which is a write. A
read-only connection to a WAL database still needs its -wal and -shm
files, or write access to the directory to create them, so scan now
switches the database back to rollback-journal mode whenever it closes
it: between scans the file alone holds the database. If a report has
the database open at that moment the switch is refused; scan warns and
the database stays in WAL mode, with its -wal and -shm files, until the
next scan. README §Database states what readers need.

Model: opus-5-5
2026-10-03 16:30:19 +02:00
clawbot c887f80f57 Escape tab, newline, CR and backslash in report paths (closes #7)
check / check (push) Successful in 1m29s
A path holding a tab or newline split a row of the report or trees
output. The path columns of both now write a backslash, tab, newline
and carriage return as \\, \t, \n and \r; every other byte is written
unchanged. Grouping and sorting still use the stored path. Warnings
on stderr are escaped the same way in warnf, so each stays one line.

In trees, the root directory's node now has the path "/" instead of
an empty string, and its children's paths start with a single slash.

README states the rule under "Report output format".

Model: opus-5-5
2026-10-03 15:30:37 +02:00
clawbot c9bf22d483 Stamp the git tag or short commit in a plain docker build (closes #67)
check / check (push) Successful in 1m1s
.dockerignore now sends .git, without .git/config, which can hold a
credential. The build stage takes the VERSION build argument when one
is given, otherwise git describe --tags --always of that .git, and
fails if the context carries .git and still yields no version. A plain
docker build . used to stamp dev. The CI checkout fetches full history
so CI sees the tag and stamps the same value as make build.

Model: opus-5-5
2026-10-02 08:49:03 +02:00
clawbot c737490a53 Compute the content hash only when head and tail match (closes #61)
check / check (push) Successful in 49s
A file of 10 MiB or more now gets only its 64 KiB head and tail in the
hash phase, so its content is read only when it can be a duplicate. A
new content phase after the update phase finds every group of records,
anywhere in the database, that share size, head and tail and include
one without a content hash. It checks every member with lstat and, when
at least two pass, reads those without a content hash through the
existing worker pool; a stale file does not count as a match. report
and trees leave out records without a content hash. The README, help
text and TODO entry describe the gate; the schema stays at version 1.

Lint suppressed: gosec on the file open in hashContentOnly, as in
hashSignature, and on one chmod in a test.

Model: opus-5-5
2026-09-23 16:06:09 +02:00
clawbot 09a39ddf37 Keep the database schema at version 1 (closes #61)
check / check (push) Successful in 42s
sfdupes is pre-1.0, with no installed base and no databases anywhere,
so the schema is changed in place and its version stays 1.
schemaVersion goes back to 1; the six-column files table, content
included, is the version 1 schema. The check that stops on a database
with any other version stays. README.md and TODO.md no longer describe
a version 2 or rejecting and rescanning version 1 databases. The
main.go package comment still described 1024-byte end windows and said
full file contents are never read; it now describes the hashes the
code computes.

Model: opus-5-5
2026-09-23 13:38:06 +02:00
clawbot 29a65016d0 Add 64 KiB head/tail and content-hash duplicate ladder (closes #61) (#62)
check / check (push) Successful in 57s
2026-09-22 16:40:43 +02:00
clawbot 7ac4f6b723 Remove dead files.dat references from build config (closes #22)
check / check (push) Failing after 0s
files.dat was the scan format before the SQLite database; nothing has
produced it since. Drop the stale references from the Makefile clean
target, .gitignore and .dockerignore. make clean still removes the
binary and .gitignore still covers the database files. The only
remaining mention is the historical entry in TODO.md.

Model: opus-4-8 (implementation); fable-5-1 (merge)
2026-09-21 15:01:57 +02:00
21 changed files with 2862 additions and 656 deletions
+6 -2
View File
@@ -1,8 +1,12 @@
.git # .git is sent without its config. Without a VERSION build argument the
# stage that compiles runs `git describe --tags --always` on .git, which
# does not need .git/config; that file can hold a credential, such as a
# password in a remote URL or the token the CI checkout step stores there.
.git/config
.claude .claude
.DS_Store .DS_Store
sfdupes sfdupes
files.dat
*.log *.log
*.out *.out
*.test *.test
+2
View File
@@ -6,4 +6,6 @@ jobs:
steps: steps:
# actions/checkout v4.2.2, 2026-02-22 # actions/checkout v4.2.2, 2026-02-22
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 - uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
with:
fetch-depth: 0
- run: script/cibuild - run: script/cibuild
-1
View File
@@ -27,7 +27,6 @@ node_modules/
*.log *.log
# Local scan data # Local scan data
files.dat
*.sqlite *.sqlite
*.sqlite-shm *.sqlite-shm
*.sqlite-wal *.sqlite-wal
+13 -1
View File
@@ -108,7 +108,19 @@ ARG CHECK_EPOCH
RUN echo "gate test, epoch ${CHECK_EPOCH}" && make test RUN echo "gate test, epoch ${CHECK_EPOCH}" && make test
RUN echo "gate fmt-check, epoch ${CHECK_EPOCH}" && make fmt-check RUN echo "gate fmt-check, epoch ${CHECK_EPOCH}" && make fmt-check
RUN make build # The version stamped into the binary: the VERSION build argument when
# one is given, otherwise `git describe --tags --always` of the .git in
# the build context (git is installed by script/bootstrap above). A
# context that carries .git and still yields no version fails the build;
# with neither, as from a source tarball, it is "dev".
ARG VERSION
RUN version="${VERSION:-$(git describe --tags --always || echo dev)}"; \
if [ -e .git ] && { [ -z "$version" ] || [ "$version" = dev ] || \
[ "$version" = unknown ]; }; then \
echo "no version could be derived although the build context carries .git" >&2; \
exit 1; \
fi; \
make build VERSION="$version"
# Runtime stage # Runtime stage
# alpine:3.22, 2026-07-23 # alpine:3.22, 2026-07-23
+1 -1
View File
@@ -46,4 +46,4 @@ hooks:
@script/install-precommit @script/install-precommit
clean: clean:
rm -f $(BINARY) files.dat rm -f $(BINARY)
+286 -95
View File
@@ -5,14 +5,20 @@
`sfdupes` is an MIT-licensed Go CLI tool by `sfdupes` is an MIT-licensed Go CLI tool by
[@sneak](https://sneak.berlin) that quickly identifies *candidate* [@sneak](https://sneak.berlin) that quickly identifies *candidate*
duplicate files — and, ultimately, entire duplicate directory trees — duplicate files — and, ultimately, entire duplicate directory trees —
across very large filesystems without reading full file contents. Files across very large filesystems without reading every byte of every file.
are considered duplicates when they have identical size, identical Files are considered duplicates when their sizes are equal and they
SHA-256 of their first 1024 bytes, and identical SHA-256 of their last agree on a short ladder of hashes. A file under 10 MiB is hashed in full
1024 bytes. This is a strong candidate signal, not proof of identical and compared directly. A larger file is gated first on the SHA-256 of
content (the middle of the file is never read); the intended use is its first 64 KiB and of its last 64 KiB, and only when its size and
both of those match another file's is it read for a content hash to
compare — the SHA-256 of the whole file when it is under 50 MiB, or of
gigabyte-spaced 1 MiB samples when it is 50 MiB or larger. Below 50 MiB
the content hash is proof of identical content; at or above 50 MiB it is
a strong candidate signal rather than proof, because the gaps between
samples are never read. The intended use is
finding duplicate downloads and duplicated directory trees on finding duplicate downloads and duplicated directory trees on
multi-terabyte ZFS servers where reading every byte is prohibitively multi-terabyte ZFS servers where reading every byte of every file is
expensive. `scan` maintains a persistent SQLite database of file prohibitively expensive. `scan` maintains a persistent SQLite database of file
signatures that survives between runs, so it can be run from cron and signatures that survives between runs, so it can be run from cron and
the reports can be generated at any time from the most recent scan. the reports can be generated at any time from the most recent scan.
@@ -29,9 +35,11 @@ export SFDUPES_DATABASE="$HOME/.local/share/sfdupes/db.sqlite"
``` ```
`scan` walks one or more filesystem trees and maintains one database `scan` walks one or more filesystem trees and maintains one database
record per regular file (path, size, mtime, head hash, tail hash). The record per regular file (path, size, mtime, head hash, tail hash,
content hash). The
database persists between runs; a rescan only hashes files that are new database persists between runs; a rescan only hashes files that are new
or changed, and removes records for files that no longer exist. or changed, or that may have gained a duplicate since the last scan,
and removes records for files that no longer exist.
`report` reads the database and prints the file-level duplicates `report` reads the database and prints the file-level duplicates
report. `trees` reads the same database and prints the duplicate-tree report. `trees` reads the same database and prints the duplicate-tree
report. A missing/invalid subcommand — or a `scan` invocation with no report. A missing/invalid subcommand — or a `scan` invocation with no
@@ -47,13 +55,19 @@ completed scan.
Duplicate finders that hash entire files do not scale to the target Duplicate finders that hash entire files do not scale to the target
environment: ~10 million files and ~150 TB on possibly slow or busy environment: ~10 million files and ~150 TB on possibly slow or busy
disks (a ZFS pool under resilver). Reading at most 2 KiB per file — and disks (a ZFS pool under resilver). sfdupes spends disk I/O only on files
only from files whose size at least one other file shares, since a whose size at least one other file shares, since a size-unique file
size-unique file cannot be a duplicate — makes a full-filesystem sweep cannot be a duplicate. Of those, a file under 10 MiB is read in full; a
tractable, and the signatures are kept in a persistent database, so larger one has its cheap end windows read first, and is read for a
the expensive filesystem pass is incremental: a rescan re-hashes only content hash only when its size and both end windows match another
files whose recorded mtime or size changed, and all analysis happens file's — the whole file below 50 MiB, but only gigabyte-spaced samples
offline from the database alone. The end goal is at or above 50 MiB, so the largest files are never read in full. This
keeps a full-filesystem sweep tractable, and the signatures are kept in
a persistent database, so the expensive filesystem pass is incremental:
a rescan re-hashes only files whose recorded mtime or size changed, plus
— for its content hash — a file of 10 MiB or more whose size and end
windows have come to match another file's. All analysis happens offline
from the database alone. The end goal is
not individual files but whole duplicated trees — duplicate not individual files but whole duplicated trees — duplicate
extractions, duplicate downloads, copied project trees — which an extractions, duplicate downloads, copied project trees — which an
operator can consider removing as a unit. operator can consider removing as a unit.
@@ -68,17 +82,31 @@ Goals, in order:
downloads, copied project trees), so the operator can consider downloads, copied project trees), so the operator can consider
removing an entire subtree at once. File-level duplicate detection is removing an entire subtree at once. File-level duplicate detection is
the foundation; tree-level detection is built on top of it. the foundation; tree-level detection is built on top of it.
2. **Never read full file contents.** At most 2 KiB is read per file 2. **Spend I/O in proportion to duplicate likelihood.** Only files
(first and last 1024 bytes), and only files whose size at least whose size at least one other file shares are read at all — a
one other file shares are read at all — a size-unique file cannot size-unique file cannot be a duplicate. Those are compared by the
be a duplicate. Scale target: tens of millions of files, ~150 TB ladder in "Duplicate detection" below: a file under 10 MiB is hashed
filesystem, possibly slow or busy disks (ZFS pool under resilver). in full, while a larger file is gated on cheap 64 KiB end windows
Holding one small record (path, size, mtime) per file in memory first, and gets a content hash only when its size and both end
during a scan is acceptable; holding every file's hashes is not windows match another file's. That hash reads the whole file below
(they stay in the database). 50 MiB but only gigabyte-spaced 1 MiB samples at or above it, so the
very largest files are still never read in full. Scale target: tens of
millions of files, ~150 TB filesystem, possibly slow or busy disks
(ZFS pool under resilver). Holding one small record (path, size,
mtime) per file in memory during a scan is acceptable; holding
every file's hashes is not (they stay in the database). The
reporting commands hold no file's hashes either: `report` lets
SQLite group and order the records and writes each row as it reads
it, so its memory does not grow with the database, and `trees`
reads the records in path order and keeps each directory's path,
digest and totals, plus the files of the directories holding the
record being read, so its memory grows with the number of
directories, not files.
3. **Scan incrementally, analyze offline.** The expensive filesystem 3. **Scan incrementally, analyze offline.** The expensive filesystem
scan maintains a persistent database; an unchanged file is never scan maintains a persistent database; an unchanged file is never
read again on a rescan. All analysis (`report`, `trees`) works from read again on a rescan, except to compute its content hash once a
file of 10 MiB or more comes to match another on size and both end
windows. All analysis (`report`, `trees`) works from
the database alone and must never touch the scanned filesystem the database alone and must never touch the scanned filesystem
again. `scan` is designed to be cronned; the reports run at any again. `scan` is designed to be cronned; the reports run at any
time against the last completed scan. time against the last completed scan.
@@ -92,8 +120,9 @@ Goals, in order:
`sfdupes`. `sfdupes`.
- Dependencies: standard library, `github.com/spf13/cobra` for the - Dependencies: standard library, `github.com/spf13/cobra` for the
CLI, **one progress-bar library** CLI, **one progress-bar library**
(`github.com/schollz/progressbar/v3`), and **one SQLite driver** (`github.com/schollz/progressbar/v3`), **one SQLite driver**
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled). (`modernc.org/sqlite`, pure Go, so builds keep cgo disabled), and
`golang.org/x/sys` for `flock(2)` (the scan lock, see "Database").
`github.com/spf13/viper` is permitted if configuration-file support `github.com/spf13/viper` is permitted if configuration-file support
is ever needed, but is not currently used. No other third-party is ever needed, but is not currently used. No other third-party
deps. deps.
@@ -131,36 +160,116 @@ All three subcommands operate on a single SQLite database file:
use. `report` and `trees` require an existing database; a missing use. `report` and `trees` require an existing database; a missing
database file is a fatal error (exit 1) telling the user to run database file is a fatal error (exit 1) telling the user to run
`scan` first. `scan` first.
- The database uses WAL journal mode and a busy timeout, so running a - Only one `scan` runs against a database at a time. For its whole
report while a cron `scan` is in progress is safe. The filesystem run, `scan` holds an exclusive `flock(2)` lock on a lock file
is authoritative; the database is an eventually-consistent beside the database, named by appending `.lock` to the database
reflection of it. Hashed records are committed in batched path (`/var/lib/sfdupes/db.sqlite.lock` by default), taken before
transactions while the scan is still running (keeping the WAL it walks the filesystem or opens the database. A second `scan`
small and letting concurrent reports observe progress), so a against the same database does not wait: it fails at once with a
report may see a scan's changes partially applied, and a scan one-line error naming the lock file and exits 1, without walking
that dies partway leaves a valid database holding everything anything or opening the database, and the running scan carries on.
hashed so far; the next scan skips those records and converges The lock file is created on first use, open to its owner only, and
toward the filesystem. left in place: a leftover file blocks nothing, because the lock
ends with the process holding it however it ends, a fatal error or
an interrupt included, and deleting the file while a scan runs
would let a second scan start. `report` and `trees` never take the
lock, so they run during a scan.
- While `scan` runs, the database is in WAL journal mode with a busy
timeout, so running a report while a cron `scan` is in progress is
safe. The filesystem is authoritative; the database is an
eventually-consistent reflection of it. Hashed records are
committed in batched transactions while the scan is still running
(keeping the WAL small and letting concurrent reports observe
progress), so a report may see a scan's changes partially applied,
and a scan that dies partway leaves a valid database holding
everything hashed so far; the next scan skips those records and
converges toward the filesystem.
- `scan` switches the database back to rollback-journal mode when it
closes it, so between scans the database file alone holds the whole
database. Each switch needs the database to itself: a `scan` that
starts while a report is still reading waits for it up to the
10-second busy timeout, then fails; a `scan` that ends while a
report has the database open warns and leaves the database in WAL
mode until the next scan.
- `report` and `trees` open the database read-only and need only read
access to the database file, and no write access to its directory.
While the database is in WAL mode they also read the `-wal` and
`-shm` files beside it, which SQLite creates with the database
file's permissions.
- Schema (`PRAGMA user_version` is the schema version, currently 1; a - Schema (`PRAGMA user_version` is the schema version, currently 1; a
database with any other version is a fatal error): database with any other version is a fatal error):
```sql ```sql
CREATE TABLE files ( CREATE TABLE files (
path BLOB PRIMARY KEY, -- absolute path, raw bytes path BLOB PRIMARY KEY, -- absolute path, raw bytes
size INTEGER NOT NULL, -- bytes, from lstat size INTEGER NOT NULL, -- bytes, from lstat
mtime INTEGER NOT NULL, -- Unix seconds, from lstat mtime INTEGER NOT NULL, -- Unix seconds, from lstat
head TEXT NOT NULL, -- lowercase-hex SHA-256, first 1 KiB head TEXT NOT NULL, -- lowercase-hex SHA-256; first 64 KiB, or whole file under 10 MiB
tail TEXT NOT NULL -- lowercase-hex SHA-256, last 1 KiB tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
) WITHOUT ROWID; ) WITHOUT ROWID;
CREATE INDEX files_signature ON files (size, head, tail, content);
``` ```
Paths are stored as BLOBs because Unix paths are raw bytes, not Paths are stored as BLOBs because Unix paths are raw bytes, not
guaranteed UTF-8. `mtime` is used only for change detection; it is guaranteed UTF-8. `mtime` is used only for change detection; it is
not part of the duplicate key. `head` and `tail` are empty strings not part of the duplicate key. For a file under 10 MiB `head`, `tail`,
when the file has never been hashed because its size was unique as and `content` all hold the whole-file hash (that range is hashed in
of the last scan that covered it; such records still define the full, with no end windows); for a larger file `head` and `tail` hold
file for tree reconstruction but never participate in duplicate the first- and last-64 KiB hashes and `content` the whole-file or
groups. sampled hash. All three are empty strings when the file has never
been hashed because its size was unique as of the last scan that
covered it. For a file of 10 MiB or more, `content` stays empty
until the content phase of a scan (see "`scan` mode" below) has
read the file. A record with an empty `content` is never part of a
duplicate group, though it still defines the file for tree
reconstruction. The `files_signature` index lets SQLite group the
records by signature for `report` without sorting the whole table.
### Duplicate detection
Two files are duplicates only when they agree on every rung of this
ladder; a mismatch at any rung means they are not duplicates. `scan`
stores each file's hashes, and `report` and `trees` group files by the
whole signature — size, `head`, `tail`, and `content` — so the grouping
is exactly this ladder applied across everything scanned into the
database, even across separate scans.
1. **Size.** Files of different sizes are never compared. Only files
whose size at least one other file shares are hashed at all.
2. **Under 10 MiB: whole file.** A file smaller than 10 MiB is hashed
in full and compared directly, with no separate end-window step —
small files are cheap to read to the last byte, and doing so makes
the comparison exact. `head`, `tail`, and `content` all hold this
whole-file SHA-256, so such a file's signature is decided entirely
by its size and its content.
3. **10 MiB and above: head and tail.** For a larger file, the SHA-256
of the first 64 KiB (`head`) and of the last 64 KiB (`tail`) are a
cheap gate that eliminates most same-size pairs before any bulk
reading: the content hash of the next two rungs is computed only for
a file whose size, `head`, and `tail` match another file's, whether
that file is scanned in the same run or stored by an earlier scan.
A stored file that first gains such a match in a later scan gets its
content hash then; until it has one, its `content` is empty and it
is not a duplicate. At 10 MiB and above the two windows never
overlap.
4. **10 MiB and above, content below 50 MiB.** The SHA-256 of the
entire file. Agreement here is proof of identical content (barring a
SHA-256 collision).
5. **10 MiB and above, content 50 MiB and above.** A sampled SHA-256:
the 1 MiB window at each gigabyte-aligned offset (0, 1 GiB, 2 GiB, …
while inside the file, the final window truncated at end of file) is
fed, in order, into one hash. This is **deliberately probabilistic**
— the gaps between samples are never read, so two large files that
agree on every sample are reported as duplicates without being read
in full. It is the price of never reading a 150 GB file end to end.
Because size is already part of the signature, only equal-size files
reach this rung, so their sample boundaries always align.
`head`, `tail`, and `content` are one column each. A file below 10 MiB
and one at or above it never share a size, and neither do a file below
50 MiB and one at or above it, so a stored value is never ambiguous
between the whole-file, end-window, and sampled forms.
### `scan` mode ### `scan` mode
@@ -178,15 +287,27 @@ duplicates another or lies under another is dropped before walking,
so every file is reached exactly once and produces one database so every file is reached exactly once and produces one database
record. record.
An operand that is a symlink (never followed, not even as an operand),
socket, FIFO, or device node, or a directory named `.zfs`, is not
scanned. `scan` prints a one-line warning naming the path and what it
is, counts it as skipped, and drops it from the scanned operands before
reading the database. Another operand beneath it is still scanned. The
records stored beneath it are not deleted: they are treated like any
other record outside the scanned operands, including the content-phase
exception below. If it lies under another operand, they are under that
operand instead, and are deleted like any other record there that this
scan did not verify. This is not an error: a scan whose every operand
is dropped walks nothing and exits 0.
`scan` synchronizes the database with the filesystem state under the `scan` synchronizes the database with the filesystem state under the
scanned operands: scanned operands:
- Only a file whose size at least one other file shares is ever - Only a file whose size at least one other file shares is ever
read: a size-unique file cannot be a duplicate, so it is recorded read: a size-unique file cannot be a duplicate, so it is recorded
without hashes (`head` and `tail` empty). The size census covers without hashes (`head`, `tail`, and `content` empty). The size
every file walked this scan plus every database record outside census covers every file walked this scan plus every database
the scanned operands, so a possible duplicate of a separately record outside the scanned operands, so a possible duplicate of a
scanned tree is still recognized. separately scanned tree is still recognized.
- A file not yet in the database is inserted: hashed when its size - A file not yet in the database is inserted: hashed when its size
is shared, without hashes otherwise. is shared, without hashes otherwise.
- A file already in the database is **skipped without reading its - A file already in the database is **skipped without reading its
@@ -195,7 +316,10 @@ scanned operands:
makes a daily rescan cheap. Exception: an unchanged file whose makes a daily rescan cheap. Exception: an unchanged file whose
record lacks hashes is hashed — and its record updated — once its record lacks hashes is hashed — and its record updated — once its
size becomes shared, so hashing deferred by size-uniqueness size becomes shared, so hashing deferred by size-uniqueness
happens as soon as it could matter. happens as soon as it could matter. Likewise, an unchanged file of
10 MiB or more whose record has no `content` hash is read for one
by the content phase below once its size, `head`, and `tail` match
another record's.
- A file whose mtime is newer than recorded, or whose size differs, - A file whose mtime is newer than recorded, or whose size differs,
is processed as if new: re-hashed, or recorded without hashes, is processed as if new: re-hashed, or recorded without hashes,
per the shared-size rule. per the shared-size rule.
@@ -204,12 +328,18 @@ scanned operands:
removes records for deleted files. It also removes records for removes records for deleted files. It also removes records for
paths that failed to stat or hash this run: the database only ever paths that failed to stat or hash this run: the database only ever
contains signatures verified by the most recent scan that covered contains signatures verified by the most recent scan that covered
them (a subsequent successful scan re-adds such files). them (a subsequent successful scan re-adds such files). A failure
in the content phase below removes nothing: the record is left as
it is.
- Database records outside the scanned operands are untouched, so - Database records outside the scanned operands are untouched, so
disjoint trees can be scanned on different schedules into the same disjoint trees can be scanned on different schedules into the same
database. database. The one exception is the content phase below: a stored
file of 10 MiB or more without a `content` hash is read for one,
wherever it lies, once its size, `head`, and `tail` match another
record's. If that file is gone or has changed since its record was
written, the record is left as it is.
`scan` runs **three sequential phases over the whole scan**. `scan` runs **four sequential phases over the whole scan**.
Parallelism lives inside each phase; batched database writes begin Parallelism lives inside each phase; batched database writes begin
during the hash phase: during the hash phase:
@@ -229,12 +359,13 @@ during the hash phase:
decides its fate. Size-unique files are never read: new or decides its fate. Size-unique files are never read: new or
changed ones are recorded without hashes in the update phase, changed ones are recorded without hashes in the update phase,
unchanged unhashed ones simply keep their records. Every file unchanged unhashed ones simply keep their records. Every file
with a shared size is hashed by the worker pool: read the first with a shared size is hashed by the worker pool as described in
`min(1024, size)` bytes and the last `min(1024, size)` bytes "Duplicate detection" above: a file under 10 MiB in full, which
(one read when `size <= 1024`, since the two windows coincide) gives its `head`, `tail`, and `content` alike, and a larger file
and compute the SHA-256 of each. Zero-length files have constant only in its end windows, which give its `head` and `tail`; its
hashes and are never opened. Files are hashed in **inode order** content hash is left to the content phase. Zero-length files have
(minimizing seeks on spinning disks), and paths that are hard constant hashes and are never opened. Files are hashed in **inode
order** (minimizing seeks on spinning disks), and paths that are hard
links to the same inode are **read once**, all sharing the one links to the same inode are **read once**, all sharing the one
result — a hard-link backup farm costs one read per inode, not result — a hard-link backup farm costs one read per inode, not
per path. The phase total counts actual reads, so progress and per path. The phase total counts actual reads, so progress and
@@ -246,13 +377,39 @@ during the hash phase:
records for size-unique new and changed files, and the deletions records for size-unique new and changed files, and the deletions
for records the scan did not verify (vanished files, plus paths for records the scan did not verify (vanished files, plus paths
that failed to stat or hash). that failed to stat or hash).
4. **content** — find every record of 10 MiB or more without a
`content` hash whose size, `head`, and `tail` equal another
record's, anywhere in the database: records from this scan and
records stored by earlier scans, inside or outside the scanned
operands. SQLite finds them, so only the records to be read are
kept in memory, never every file's hashes. Every record sharing
their size, `head`, and `tail`, including one that already has a
`content` hash, has its file checked with `lstat` first. A file
that is gone, is no longer a regular file, or has changed (a
different size, or an mtime newer than recorded) keeps its record
as it is and does not count as a match for the others. Any other
`lstat` error is warned about and counted as skipped, with the same
result. If such a record has no `content` hash, it stays out of
duplicate groups; if it has one, it is still reported until a scan
covering its own tree updates or removes it. The files that pass
and have no `content` hash are read only if at least two of those
records pass, so a file whose only matches are stale costs no read;
a file that already has a `content` hash is never read again.
They are read by a worker pool as in the hash phase, in inode order
and once per inode, and their content hashes are committed in
batches. A failed read is warned about and counted as skipped; its
record keeps an empty `content`, so it is not a duplicate, and a
later scan tries again.
Rules for the walk: Rules for the walk:
- Only regular files. Skip directories, symlinks (do not follow, - Only regular files. Skip directories, symlinks (do not follow,
including symlink operands), sockets, FIFOs, and device nodes. including symlink operands), sockets, FIFOs, and device nodes. An
operand that is a symlink, socket, FIFO, or device node is dropped
as described in "`scan` mode" above.
- Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs; - Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs;
walking them would list every file once per snapshot). walking them would list every file once per snapshot), not even
when it is an operand; such an operand is dropped the same way.
- Filesystem boundaries are crossed by default. With `-x` - Filesystem boundaries are crossed by default. With `-x`
(long form `--one-file-system`, following the GNU `du`/`rsync` (long form `--one-file-system`, following the GNU `du`/`rsync`
convention), never descend into a directory on a different convention), never descend into a directory on a different
@@ -263,18 +420,20 @@ Rules for the walk:
path, and continue. Per-file errors never abort the run; the final path, and continue. Per-file errors never abort the run; the final
summary reports how many were skipped. As specified above, a summary reports how many were skipped. As specified above, a
skipped path that has a database record from an earlier scan loses skipped path that has a database record from an earlier scan loses
that record; an unreadable directory subtree likewise loses its that record, unless it failed only in the content phase, or is an
operand dropped before the database was read that lies under no
other operand; an unreadable directory subtree likewise loses its
records (accepted: the database mirrors what the latest scan could records (accepted: the database mirrors what the latest scan could
actually verify). actually verify).
Concurrency: the walk phase (which also stats files) and the hash Concurrency: the walk phase (which also stats files), the hash phase,
phase each use a worker pool of `--workers` workers (default and the content phase each use a worker pool of `--workers` workers
`runtime.NumCPU()`); the walk parallelizes across directories, (default `runtime.NumCPU()`); the walk parallelizes across
hashing across files. Both phases are seek-bound on spinning disks, directories, hashing across files. All three phases are seek-bound on
so raising `--workers` well past the core count can help on pools spinning disks, so raising `--workers` well past the core count can
with many spindles. The main goroutine owns partitioning, database help on pools with many spindles. The main goroutine owns
writes, and progress rendering; progress display must never block partitioning, database writes, and progress rendering; progress
the workers. display must never block the workers.
`scan` writes nothing to stdout. The summary line on stderr reports the `scan` writes nothing to stdout. The summary line on stderr reports the
files seen this run broken down by disposition, plus skips: files seen this run broken down by disposition, plus skips:
@@ -293,17 +452,20 @@ arguments.
**`report` must never touch the filesystem being analyzed.** It does not **`report` must never touch the filesystem being analyzed.** It does not
stat, open, or otherwise access any path that appears in the records; its stat, open, or otherwise access any path that appears in the records; its
only I/O is reading the database and writing stdout/stderr. It must only I/O is reading the database, writing stdout/stderr, and the
produce identical output whether or not the scanned filesystem is still temporary file SQLite sorts in when the duplicate rows do not fit in
mounted. memory. SQLite puts that file in `$SQLITE_TMPDIR` or `$TMPDIR` when set,
otherwise in `/var/tmp` (or `/tmp`), and deletes it as soon as it has
opened it. `report` must produce identical output whether or not the
scanned filesystem is still mounted.
Processing: Processing:
- Records without hashes (size-unique when last scanned) are - Records without a `content` hash (see "Database" above) are
excluded: their content is unknown, so they are never reported as excluded: their content is unknown, so they are never reported as
duplicates. duplicates.
- Group the remaining records by the key - Group the remaining records by the key
`(size, head_hash, tail_hash)`. `(size, head, tail, content)`.
- Every group with two or more paths is a duplicate group. - Every group with two or more paths is a duplicate group.
- Within each group, sort paths lexicographically (byte order). The - Within each group, sort paths lexicographically (byte order). The
first path is the group's `first`; every other path is a `dupe`. first path is the group's `first`; every other path is a `dupe`.
@@ -322,6 +484,15 @@ first dupe size
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296 /srv/a/big.iso /srv/c/big-copy2.iso 4294967296
``` ```
Paths are raw bytes and may hold any byte except NUL, so the path
columns (`first` and `dupe`) are escaped to keep every row one line of
tab-separated fields: a backslash is written as `\\`, a tab as `\t`, a
newline as `\n`, and a carriage return as `\r`. Every other byte is
written unchanged, including bytes that are not valid UTF-8. Undoing
those four escapes gives back the stored path. Grouping and ordering
use the stored path, not the escaped one. The warnings `scan` prints on
stderr are escaped the same way, so each warning is one line.
Summary to stderr: records read, number of duplicate groups, number of Summary to stderr: records read, number of duplicate groups, number of
dupe files, and total reclaimable bytes (sum of `size` over all dupe dupe files, and total reclaimable bytes (sum of `size` over all dupe
rows) in human units. rows) in human units.
@@ -339,11 +510,11 @@ the paths in the records, split on `/`.
Definitions: Definitions:
- A file's **signature** is `(size, head_hash, tail_hash)` — mtime is - A file's **signature** is `(size, head, tail, content)` — mtime is
informational and excluded. An unhashed record (empty hashes) has informational and excluded. A record without a `content` hash has
unknown content: its signature is treated as unique to that file, unknown content: its signature is treated as unique to that file,
so a tree containing an unhashed file never compares equal to any so a tree containing such a file never compares equal to any other
other tree. tree.
- A directory's **digest** is a SHA-256 Merkle digest computed - A directory's **digest** is a SHA-256 Merkle digest computed
bottom-up: serialize the directory's child entries — for a file bottom-up: serialize the directory's child entries — for a file
child, its name and signature; for a subdirectory child, its name child, its name and signature; for a subdirectory child, its name
@@ -393,6 +564,9 @@ first dupe files size
/srv/a/project /srv/backup/project 3417 104857600 /srv/a/project /srv/backup/project 3417 104857600
``` ```
The `first` and `dupe` paths are escaped as described under "Report
output format". The root directory's path is `/`.
Summary to stderr: records read, number of duplicate-tree groups, Summary to stderr: records read, number of duplicate-tree groups,
number of dupe trees, and total reclaimable bytes (sum of `size` over number of dupe trees, and total reclaimable bytes (sum of `size` over
all dupe rows) in human units. all dupe rows) in human units.
@@ -406,10 +580,13 @@ Each phase gets its own display, rendered the moment the phase
starts — a scan must never look hung. Loading the existing-record starts — a scan must never look hung. Loading the existing-record
index (`load`) and the walk have no known totals while running: show index (`load`) and the walk have no known totals while running: show
a live count, rate, and elapsed time (spinner-style, no percentage or a live count, rate, and elapsed time (spinner-style, no percentage or
ETA). The hash and update phases ETA). The content phase's display (`content`) starts the same way,
have exact totals — only files that actually need hashing appear in counting the records checked while SQLite finds the files to read and
the hash total, so its ETA is meaningful. Required elements for the `lstat` checks them, then shows a bar once reading starts. The hash
bars with known totals: and update phases, and the content phase's reads, have exact totals —
only files that actually need hashing appear in the hash and content
totals, so their ETAs are meaningful. Required elements for the bars
with known totals:
- elapsed time - elapsed time
- estimated time remaining - estimated time remaining
@@ -434,12 +611,25 @@ Additional requirements:
### Error handling and exit codes ### Error handling and exit codes
- `0`: success, even if individual files were skipped with warnings. - `0`: success, even if individual files were skipped with warnings.
- `1`: fatal error (e.g., a `PATH` operand does not exist, the - `1`: fatal error (e.g., a `PATH` operand does not exist, another
database cannot be created/opened/read/written, a missing database `scan` is already running against the same database, the database
for `report`/`trees`, stdout write failure). cannot be created/opened/read/written, a missing database for
`report`/`trees`, stdout write failure).
- `2`: usage error (including `scan` with no `PATH` operand and - `2`: usage error (including `scan` with no `PATH` operand and
`report`/`trees` with any positional argument). `report`/`trees` with any positional argument).
A stdout write failure, such as a full disk, is reported in one line on
stderr and exits 1. Two cases never reach sfdupes as a failed write:
- When the reader of a stdout pipe exits early, as in
`sfdupes report | head`, the next write ends sfdupes with `SIGPIPE`,
quietly and without a summary, the way `cat` or `sort` end. The
shell reports the signal (status 141 in most shells), not exit 1.
- When stdout is closed outright (`sfdupes report >&-`), the Go
runtime opens `/dev/null` in its place before sfdupes starts, so
the output is discarded and the run succeeds, as with
`> /dev/null`.
## Entrypoints ## Entrypoints
This repository adheres to the This repository adheres to the
@@ -577,7 +767,7 @@ All of the following, run in this directory, must pass:
```sh ```sh
d=$(mktemp -d) d=$(mktemp -d)
export SFDUPES_DATABASE="$d/db.sqlite" export SFDUPES_DATABASE="$(mktemp -d)/db.sqlite"
mkdir -p "$d/a" "$d/b" mkdir -p "$d/a" "$d/b"
head -c 2000 /dev/urandom > "$d/a/one.bin" head -c 2000 /dev/urandom > "$d/a/one.bin"
cp "$d/a/one.bin" "$d/b/copy.bin" cp "$d/a/one.bin" "$d/b/copy.bin"
@@ -605,9 +795,9 @@ All of the following, run in this directory, must pass:
./sfdupes report ./sfdupes report
``` ```
(The scan database lives inside `$d` here purely for test hygiene; (The database lives in a temp directory of its own: inside `$d`,
scanning `$d` therefore also records the SQLite file itself, which the scan would record it, and its empty lock file would join the
is harmless.) `empty1`/`empty2` group.)
Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin` Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin`
form one group (two dupe rows, `first` is the lexicographically form one group (two dupe rows, `first` is the lexicographically
@@ -636,9 +826,10 @@ Tracked in [TODO.md](TODO.md).
## Non-goals ## Non-goals
- No full-content verification, no byte-for-byte compare, no deletion - No byte-for-byte compare, and no deletion or linking of
or linking of duplicates. The reports are advisory; acting on them is duplicates. Files that match are compared by a SHA-256 of the whole
the user's job. file below 50 MiB, and only by samples at 50 MiB and over. The
reports are advisory; acting on them is the user's job.
- No persistence beyond the SQLite database described above; no - No persistence beyond the SQLite database described above; no
export/import formats. export/import formats.
- No daemon or filesystem watcher; scheduling rescans is cron's job. - No daemon or filesystem watcher; scheduling rescans is cron's job.
+57
View File
@@ -29,6 +29,63 @@
# Completed Steps # Completed Steps
- `report` and `trees` stream the records instead of holding them all in
memory; the schema gains the `files_signature` index (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/14)
- warn about and skip symlink, socket, FIFO, device and `.zfs`
operands, keeping the records beneath them (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/9)
- `scan` holds a lock on a lock file beside the database for its whole run,
so a second `scan` fails at once with exit 1 (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/53)
- test stdout write failures in `report` and `trees`; README states that
`| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null`
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/30)
- `report` and `trees` open the database read-only, and `scan` leaves it
out of WAL mode, so reading needs only read access (2026-10-03, closes
https://git.eeqj.de/sneak/sfdupes/issues/8)
- escape tabs, newlines, carriage returns and backslashes in report,
trees and warning paths; the root directory's path is `/`
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/7)
- stamp the git tag or short commit in a plain `docker build .`
instead of `dev` (2026-10-02, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/67): `.dockerignore` now
sends `.git`, without `.git/config`, and the `Dockerfile` build
stage takes the `VERSION` build argument when one is given,
otherwise `git describe --tags --always` of that `.git`. The build
fails if the context carries `.git` and the version still comes out
empty, `dev` or `unknown`. The CI checkout step fetches the full
history (`fetch-depth: 0`) so CI sees the tag and stamps the same
value as `make build`.
- replace the 1 KiB end-window sampling with the head/tail plus
content-hash ladder (2026-09-22, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/61): a file under 10 MiB is
hashed in full and compared directly, with no end-window step — its
`head`, `tail`, and `content` all hold the whole-file hash. A file at
10 MiB or above gets only the 64 KiB `head` and `tail` in the hash
phase; a new content phase, after the update phase, reads it for its
`content` hash — the whole file below 50 MiB, gigabyte-spaced 1 MiB
samples at or above — only when its size, `head`, and `tail` match
another record's, from the same scan or stored by an earlier one, so
a stored file gains its content hash when it gains a match. A file
that is gone or has changed since its record was written is not
read. The `content` column is part of the version 1 schema. `report`
and `trees` group by the extended signature and leave out any record
without a `content` hash, so the ladder is applied across the whole
database. README "Duplicate detection" documents every rung including
the probabilistic large-file path.
- remove the dead `files.dat` references from `Makefile`, `.gitignore`
and `.dockerignore` (2026-09-21, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/22)
- fix the lint-image pin comments and `FROM` form in `Dockerfile` and - fix the lint-image pin comments and `FROM` form in `Dockerfile` and
`Dockerfile.lint` (2026-08-10, branch `next`, closes `Dockerfile.lint` (2026-08-10, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/25): dropped the false https://git.eeqj.de/sneak/sfdupes/issues/25): dropped the false
+1 -1
View File
@@ -456,7 +456,7 @@ func TestHashWorkerDropsQueuedRuns(t *testing.T) {
go func() { go func() {
defer close(done) defer close(done)
hashWorker(cancelledContext(t), jobs, results) hashWorker(cancelledContext(t), jobs, results, hashSignature)
}() }()
awaitReturn(t, done, "hashWorker") awaitReturn(t, done, "hashWorker")
+247 -31
View File
@@ -11,6 +11,7 @@ import (
"slices" "slices"
"strconv" "strconv"
"golang.org/x/sys/unix"
// The pure-Go SQLite driver, registered as "sqlite"; keeps cgo // The pure-Go SQLite driver, registered as "sqlite"; keeps cgo
// disabled. // disabled.
_ "modernc.org/sqlite" _ "modernc.org/sqlite"
@@ -32,26 +33,39 @@ const schemaVersion = 1
// scan. // scan.
const dbDirPerm = 0o755 const dbDirPerm = 0o755
// lockFilePerm is the mode for the scan lock file. Anyone who can open
// the file can hold the lock and keep every scan from running, so it
// is open to its owner only.
const lockFilePerm = 0o600
// createTableSQL is the schema applied to a fresh database. Paths are // createTableSQL is the schema applied to a fresh database. Paths are
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8. // BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
const createTableSQL = ` const createTableSQL = `
CREATE TABLE files ( CREATE TABLE files (
path BLOB PRIMARY KEY, path BLOB PRIMARY KEY,
size INTEGER NOT NULL, size INTEGER NOT NULL,
mtime INTEGER NOT NULL, mtime INTEGER NOT NULL,
head TEXT NOT NULL, head TEXT NOT NULL,
tail TEXT NOT NULL tail TEXT NOT NULL,
content TEXT NOT NULL
) WITHOUT ROWID ) WITHOUT ROWID
` `
// createIndexSQL indexes the records by signature, so report can have
// SQLite group them without sorting the whole table.
const createIndexSQL = `
CREATE INDEX files_signature ON files (size, head, tail, content)
`
// upsertSQL inserts one file record, replacing any existing record for // upsertSQL inserts one file record, replacing any existing record for
// the same path. // the same path.
const upsertSQL = ` const upsertSQL = `
INSERT INTO files (path, size, mtime, head, tail) INSERT INTO files (path, size, mtime, head, tail, content)
VALUES (?, ?, ?, ?, ?) VALUES (?, ?, ?, ?, ?, ?)
ON CONFLICT (path) DO UPDATE SET ON CONFLICT (path) DO UPDATE SET
size = excluded.size, mtime = excluded.mtime, size = excluded.size, mtime = excluded.mtime,
head = excluded.head, tail = excluded.tail head = excluded.head, tail = excluded.tail,
content = excluded.content
` `
// errNoDatabase reports a missing database file for report/trees. // errNoDatabase reports a missing database file for report/trees.
@@ -62,6 +76,10 @@ var errNoDatabase = errors.New(
// does not understand. // does not understand.
var errSchemaVersion = errors.New("unsupported database schema version") var errSchemaVersion = errors.New("unsupported database schema version")
// errScanRunning reports that another scan holds the lock on the
// database.
var errScanRunning = errors.New("another scan is running")
// databasePath resolves the database location: SFDUPES_DATABASE when // databasePath resolves the database location: SFDUPES_DATABASE when
// set and non-empty, the compiled-in default otherwise. // set and non-empty, the compiled-in default otherwise.
func databasePath() string { func databasePath() string {
@@ -72,16 +90,24 @@ func databasePath() string {
return defaultDatabasePath return defaultDatabasePath
} }
// openDB opens the SQLite database at path with WAL journaling and a // scanParams are the connection parameters for scan: read-write, with
// busy timeout, so a report can run while a cron scan is in progress. // WAL journaling and a busy timeout, so a report can run while a cron
// It does not create or verify the schema. // scan is in progress. closeScanDatabase leaves WAL mode again.
func openDB(path string) (*sql.DB, error) { const scanParams = "_pragma=busy_timeout(10000)" +
dsn := "file:" + path + "&_pragma=journal_mode(WAL)" +
"?_pragma=busy_timeout(10000)" + "&_pragma=synchronous(NORMAL)"
"&_pragma=journal_mode(WAL)" +
"&_pragma=synchronous(NORMAL)"
db, err := sql.Open("sqlite", dsn) // reportParams are the connection parameters for report and trees:
// read-only, with the same busy timeout. They set no journal mode,
// because setting one is a write.
const reportParams = "mode=ro" +
"&_pragma=busy_timeout(10000)" +
"&_pragma=query_only(1)"
// openDB opens the SQLite database at path with the connection
// parameters params. It does not create or verify the schema.
func openDB(path, params string) (*sql.DB, error) {
db, err := sql.Open("sqlite", "file:"+path+"?"+params)
if err != nil { if err != nil {
return nil, fmt.Errorf("open database %s: %w", path, err) return nil, fmt.Errorf("open database %s: %w", path, err)
} }
@@ -94,6 +120,43 @@ func openDB(path string) (*sql.DB, error) {
return db, nil return db, nil
} }
// lockScanDatabase takes the lock that keeps a second scan off the
// database at path: an exclusive flock(2) on the file beside it named
// path with ".lock" appended, created along with the database's parent
// directory if missing. A lock held by another scan fails at once
// instead of waiting. The lock lasts until the returned file is closed
// or the process ends. The file is never deleted: a scan that deleted
// it would let the next scan lock a new file while another still holds
// the old one.
func lockScanDatabase(path string) (*os.File, error) {
err := os.MkdirAll(filepath.Dir(path), dbDirPerm)
if err != nil {
return nil, fmt.Errorf("create database directory: %w", err)
}
lockPath := path + ".lock"
//nolint:gosec // the operator chooses the database path
f, err := os.OpenFile(lockPath, os.O_RDWR|os.O_CREATE, lockFilePerm)
if err != nil {
return nil, err
}
err = unix.Flock(int(f.Fd()), unix.LOCK_EX|unix.LOCK_NB)
if err != nil {
_ = f.Close()
if errors.Is(err, unix.EWOULDBLOCK) {
return nil, fmt.Errorf("%w (lock held on %s)",
errScanRunning, lockPath)
}
return nil, fmt.Errorf("lock %s: %w", lockPath, err)
}
return f, nil
}
// openScanDatabase opens the database for the scan subcommand, creating // openScanDatabase opens the database for the scan subcommand, creating
// the file, its parent directory, and the schema as needed. // the file, its parent directory, and the schema as needed.
func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) { func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
@@ -102,7 +165,7 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return nil, fmt.Errorf("create database directory: %w", err) return nil, fmt.Errorf("create database directory: %w", err)
} }
db, err := openDB(path) db, err := openDB(path, scanParams)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -117,6 +180,24 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return db, nil return db, nil
} }
// closeScanDatabase switches the database at path from WAL back to
// rollback-journal mode and closes it. Out of WAL mode the database
// file alone holds the whole database, so a reader needs no -wal or
// -shm file beside it, nor write access to create them. The switch
// fails while a report has the database open; the database then stays
// in WAL mode, still readable, until a later scan closes it.
func closeScanDatabase(ctx context.Context, db *sql.DB, path string) {
// Runs on the way out of a cancelled scan too.
_, err := db.ExecContext(context.WithoutCancel(ctx),
"PRAGMA journal_mode = DELETE")
if err != nil {
fmt.Fprintf(os.Stderr, "scan: database %s left in WAL mode: %v\n",
path, err)
}
_ = db.Close()
}
// openReportDatabase opens an existing database for the report and // openReportDatabase opens an existing database for the report and
// trees subcommands. A missing database file is an error directing the // trees subcommands. A missing database file is an error directing the
// user to run scan first; the schema version must match exactly. // user to run scan first; the schema version must match exactly.
@@ -132,7 +213,7 @@ func openReportDatabase(ctx context.Context,
return nil, fmt.Errorf("database: %w", err) return nil, fmt.Errorf("database: %w", err)
} }
db, err := openDB(path) db, err := openDB(path, reportParams)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -181,6 +262,11 @@ func createSchema(ctx context.Context, db *sql.DB) error {
return fmt.Errorf("create schema: %w", err) return fmt.Errorf("create schema: %w", err)
} }
_, err = db.ExecContext(ctx, createIndexSQL)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
_, err = db.ExecContext(ctx, _, err = db.ExecContext(ctx,
"PRAGMA user_version = "+strconv.Itoa(schemaVersion)) "PRAGMA user_version = "+strconv.Itoa(schemaVersion))
if err != nil { if err != nil {
@@ -202,39 +288,113 @@ func userVersion(ctx context.Context, db *sql.DB) (int, error) {
return v, nil return v, nil
} }
// loadFileRows reads every record from the files table. // loadFileRows streams every record to fn in path order: byte order,
func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) { // which is the order of the primary key, so SQLite does not sort.
func loadFileRows(ctx context.Context, db *sql.DB, fn func(r scanRec)) error {
rows, err := db.QueryContext(ctx, rows, err := db.QueryContext(ctx,
"SELECT path, size, mtime, head, tail FROM files") "SELECT path, size, mtime, head, tail, content FROM files "+
"ORDER BY path")
if err != nil { if err != nil {
return nil, fmt.Errorf("read records: %w", err) return fmt.Errorf("read records: %w", err)
} }
defer func() { _ = rows.Close() }() defer func() { _ = rows.Close() }()
var recs []scanRec
for rows.Next() { for rows.Next() {
var ( var (
path []byte path []byte
r scanRec r scanRec
) )
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail) err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail,
&r.content)
if err != nil { if err != nil {
return nil, fmt.Errorf("read record: %w", err) return fmt.Errorf("read record: %w", err)
} }
r.path = string(path) r.path = string(path)
recs = append(recs, r) fn(r)
} }
err = rows.Err() err = rows.Err()
if err != nil { if err != nil {
return nil, fmt.Errorf("read records: %w", err) return fmt.Errorf("read records: %w", err)
} }
return recs, nil return nil
}
// dupeRowsSQL selects every record in a duplicate group, with the
// group's first path. A group is the records with a content hash that
// share a size, head, tail, and content, when there are two or more of
// them. The rows come in report order: groups by size descending, then
// by first path, and each group's paths ascending.
const dupeRowsSQL = `
SELECT g.first, f.path, f.size
FROM files AS f
JOIN (
SELECT size, head, tail, content, MIN(path) AS first
FROM files
WHERE content <> ''
GROUP BY size, head, tail, content
HAVING COUNT(*) > 1
) AS g USING (size, head, tail, content)
ORDER BY f.size DESC, g.first, f.path
`
// loadDupeRows streams the rows of dupeRowsSQL to fn and returns the
// number of records in the database. The count and the rows are read
// in one transaction, so they agree while a scan is committing. An
// error from fn stops the reading and is returned as it is.
func loadDupeRows(ctx context.Context, db *sql.DB,
fn func(first, path string, size int64) error,
) (int, error) {
// Everything goes through tx: the report connection is the only
// one, so a query on db would wait for tx forever.
tx, err := db.BeginTx(ctx, &sql.TxOptions{ReadOnly: true})
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
defer func() { _ = tx.Rollback() }()
var records int
err = tx.QueryRowContext(ctx, "SELECT COUNT(*) FROM files").Scan(&records)
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
rows, err := tx.QueryContext(ctx, dupeRowsSQL)
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
defer func() { _ = rows.Close() }()
for rows.Next() {
var (
first, path []byte
size int64
)
err = rows.Scan(&first, &path, &size)
if err != nil {
return 0, fmt.Errorf("read record: %w", err)
}
err = fn(string(first), string(path), size)
if err != nil {
return 0, err
}
}
err = rows.Err()
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
return records, nil
} }
// loadFileMeta streams every record's path, size, mtime, and whether // loadFileMeta streams every record's path, size, mtime, and whether
@@ -275,6 +435,62 @@ func loadFileMeta(ctx context.Context, db *sql.DB,
return nil return nil
} }
// contentCandidatesSQL selects every record of at least headTailMin
// bytes whose size, head, and tail equal another record's, in each
// group (the records sharing a size, head, and tail) where at least one
// record has no content hash, with whether each record has one. SQLite
// does the grouping, so no other record's hashes are loaded into
// memory; the rows come ordered by size, head, and tail, so each
// group's rows arrive together.
const contentCandidatesSQL = `
SELECT f.path, f.size, f.mtime, f.head, f.tail, f.content <> ''
FROM files AS f
JOIN (
SELECT size, head, tail
FROM files
WHERE size >= ? AND head <> ''
GROUP BY size, head, tail
HAVING COUNT(*) > 1 AND SUM(content = '') > 0
) AS g USING (size, head, tail)
ORDER BY size, head, tail
`
// loadContentCandidates streams the rows of contentCandidatesSQL to fn:
// each record, without its content hash, and whether it has one.
func loadContentCandidates(ctx context.Context, db *sql.DB,
fn func(r scanRec, hashed bool),
) error {
rows, err := db.QueryContext(ctx, contentCandidatesSQL, headTailMin)
if err != nil {
return fmt.Errorf("read records: %w", err)
}
defer func() { _ = rows.Close() }()
for rows.Next() {
var (
path []byte
r scanRec
hashed int64
)
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail, &hashed)
if err != nil {
return fmt.Errorf("read record: %w", err)
}
r.path = string(path)
fn(r, hashed != 0)
}
err = rows.Err()
if err != nil {
return fmt.Errorf("read records: %w", err)
}
return nil
}
// updateBatchSize is the number of record changes committed per // updateBatchSize is the number of record changes committed per
// transaction during the update pass. The filesystem is authoritative // transaction during the update pass. The filesystem is authoritative
// and the database an eventually-consistent reflection of it, so // and the database an eventually-consistent reflection of it, so
@@ -348,7 +564,7 @@ func execUpserts(ctx context.Context, tx *sql.Tx, upserts []scanRec,
for _, r := range upserts { for _, r := range upserts {
_, err = st.ExecContext(ctx, _, err = st.ExecContext(ctx,
[]byte(r.path), r.size, r.mtime, r.head, r.tail) []byte(r.path), r.size, r.mtime, r.head, r.tail, r.content)
if err != nil { if err != nil {
return fmt.Errorf("upsert %s: %w", r.path, err) return fmt.Errorf("upsert %s: %w", r.path, err)
} }
+63 -28
View File
@@ -5,9 +5,9 @@ import (
"database/sql" "database/sql"
"errors" "errors"
"fmt" "fmt"
"os"
"path/filepath" "path/filepath"
"slices" "slices"
"strings"
"testing" "testing"
) )
@@ -72,9 +72,8 @@ func TestOpenScanDatabaseCreates(t *testing.T) {
defer func() { _ = db.Close() }() defer func() { _ = db.Close() }()
recs, err := loadFileRows(t.Context(), db) if recs := dbRecords(t, db); len(recs) != 0 {
if err != nil || len(recs) != 0 { t.Fatalf("records = %v, want none", recs)
t.Fatalf("loadFileRows = %v, %v; want empty, nil", recs, err)
} }
} }
@@ -130,16 +129,63 @@ func TestOpenReportDatabaseOK(t *testing.T) {
_ = db.Close() _ = db.Close()
} }
func TestCloseScanDatabaseWhileReportOpen(t *testing.T) {
t.Parallel()
// A report holding the database open stops scan from taking it out
// of WAL mode. The -wal and -shm files must then stay beside it, so
// that a later report still needs only read access.
path := testDBPath(t)
scanDB, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
reportDB, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
closeScanDatabase(t.Context(), scanDB, path)
_ = reportDB.Close()
_, err = os.Stat(path + "-wal")
if err != nil {
t.Fatalf("no -wal left: the switch out of WAL mode was not "+
"stopped: %v", err)
}
makeReadOnly(t, path)
reportDB, err = openReportDatabase(t.Context(), path)
if err != nil {
t.Fatalf("openReportDatabase: %v", err)
}
defer func() { _ = reportDB.Close() }()
err = loadFileRows(t.Context(), reportDB, func(scanRec) {})
if err != nil {
t.Fatalf("loadFileRows: %v", err)
}
}
func TestApplyChangesRoundTrip(t *testing.T) { func TestApplyChangesRoundTrip(t *testing.T) {
t.Parallel() t.Parallel()
db := openTestDB(t) db := openTestDB(t)
// Paths may contain tabs and newlines; the database must store // Paths may contain tabs and newlines; the database must store
// them byte-exactly. // them byte-exactly. Every hash, content included, comes back as
// written.
recs := []scanRec{ recs := []scanRec{
{size: 2, mtime: 20, head: "h2", tail: "t2", path: "/a/tab\tnew\nline"}, {
{size: 1, mtime: 10, head: "h1", tail: "t1", path: "/a/x"}, size: 2, mtime: 20, head: "h2", tail: "t2", content: "c2",
path: "/a/tab\tnew\nline",
},
{size: 1, mtime: 10, head: "h1", tail: "t1", content: "c1", path: "/a/x"},
} }
err := applyChanges(t.Context(), db, recs, nil, err := applyChanges(t.Context(), db, recs, nil,
@@ -148,22 +194,17 @@ func TestApplyChangesRoundTrip(t *testing.T) {
t.Fatalf("applyChanges: %v", err) t.Fatalf("applyChanges: %v", err)
} }
got, err := loadFileRows(t.Context(), db) // The records come back in path order, which is the order of recs.
if err != nil { got := dbRecords(t, db)
t.Fatal(err)
}
slices.SortFunc(got, func(a, b scanRec) int {
return strings.Compare(a.path, b.path)
})
if !slices.Equal(got, recs) { if !slices.Equal(got, recs) {
t.Fatalf("rows = %+v, want %+v", got, recs) t.Fatalf("rows = %+v, want %+v", got, recs)
} }
// An upsert for an existing path updates in place; a delete // An upsert for an existing path updates in place; a delete
// removes exactly its path. // removes exactly its path.
upd := scanRec{size: 3, mtime: 30, head: "h3", tail: "t3", path: "/a/x"} upd := scanRec{
size: 3, mtime: 30, head: "h3", tail: "t3", content: "c3", path: "/a/x",
}
err = applyChanges(t.Context(), db, []scanRec{upd}, err = applyChanges(t.Context(), db, []scanRec{upd},
[]string{"/a/tab\tnew\nline"}, newProgress("update", 2)) []string{"/a/tab\tnew\nline"}, newProgress("update", 2))
@@ -171,11 +212,7 @@ func TestApplyChangesRoundTrip(t *testing.T) {
t.Fatalf("applyChanges: %v", err) t.Fatalf("applyChanges: %v", err)
} }
got, err = loadFileRows(t.Context(), db) got = dbRecords(t, db)
if err != nil {
t.Fatal(err)
}
if len(got) != 1 || got[0] != upd { if len(got) != 1 || got[0] != upd {
t.Fatalf("rows = %+v, want just %+v", got, upd) t.Fatalf("rows = %+v, want just %+v", got, upd)
} }
@@ -204,9 +241,8 @@ func TestApplyChangesBatching(t *testing.T) {
t.Fatalf("applyChanges: %v", err) t.Fatalf("applyChanges: %v", err)
} }
got, err := loadFileRows(t.Context(), db) if got := dbRecords(t, db); len(got) != n {
if err != nil || len(got) != n { t.Fatalf("records = %d, want %d", len(got), n)
t.Fatalf("loadFileRows = %d rows, %v; want %d", len(got), err, n)
} }
deletes := make([]string, 0, n) deletes := make([]string, 0, n)
@@ -220,8 +256,7 @@ func TestApplyChangesBatching(t *testing.T) {
t.Fatalf("applyChanges deletes: %v", err) t.Fatalf("applyChanges deletes: %v", err)
} }
got, err = loadFileRows(t.Context(), db) if got := dbRecords(t, db); len(got) != 0 {
if err != nil || len(got) != 0 { t.Fatalf("records = %d, want 0", len(got))
t.Fatalf("loadFileRows = %d rows, %v; want 0", len(got), err)
} }
} }
+1 -1
View File
@@ -5,6 +5,7 @@ go 1.25.7
require ( require (
github.com/schollz/progressbar/v3 v3.19.1 github.com/schollz/progressbar/v3 v3.19.1
github.com/spf13/cobra v1.10.2 github.com/spf13/cobra v1.10.2
golang.org/x/sys v0.46.0
modernc.org/sqlite v1.54.0 modernc.org/sqlite v1.54.0
) )
@@ -18,7 +19,6 @@ require (
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/rivo/uniseg v0.4.7 // indirect github.com/rivo/uniseg v0.4.7 // indirect
github.com/spf13/pflag v1.0.9 // indirect github.com/spf13/pflag v1.0.9 // indirect
golang.org/x/sys v0.46.0 // indirect
golang.org/x/term v0.44.0 // indirect golang.org/x/term v0.44.0 // indirect
modernc.org/libc v1.74.1 // indirect modernc.org/libc v1.74.1 // indirect
modernc.org/mathutil v1.7.1 // indirect modernc.org/mathutil v1.7.1 // indirect
+24 -15
View File
@@ -1,10 +1,14 @@
// Command sfdupes quickly identifies candidate duplicate files across // Command sfdupes quickly identifies candidate duplicate files across
// very large filesystems without reading full file contents. Files are // very large filesystems without reading every byte of every file.
// considered duplicates when they have identical size, identical SHA-256 // Files are considered duplicates when their sizes are equal and they
// of their first 1024 bytes, and identical SHA-256 of their last 1024 // agree on a short ladder of SHA-256 hashes. A file under 10 MiB is
// bytes. scan maintains a persistent SQLite database of file signatures // hashed in full. A larger file is compared on the hashes of its first
// (SFDUPES_DATABASE, default /var/lib/sfdupes/db.sqlite) that the // and last 64 KiB, and only when those match another file's is its
// reporting subcommands read. // content hash computed and compared: of the whole file when it is
// under 50 MiB, or of gigabyte-spaced 1 MiB samples when it is 50 MiB
// or larger. scan maintains a persistent SQLite database of file
// signatures (SFDUPES_DATABASE, default /var/lib/sfdupes/db.sqlite)
// that the reporting subcommands read.
// //
// Usage: // Usage:
// //
@@ -53,22 +57,27 @@ var errNoSubcommand = errors.New("no subcommand")
var Version = "dev" var Version = "dev"
func main() { func main() {
os.Exit(run(os.Args[1:], os.Stderr)) // Once the reader of a stdout pipe has gone, as in "sfdupes report |
// head", the Go runtime ends the process with SIGPIPE on the next
// write instead of returning an error (README "Error handling").
// Registering for SIGPIPE with os/signal would change that.
os.Exit(run(os.Args[1:], os.Stdout, os.Stderr))
} }
// run executes args against the command tree and returns the process // run executes args against the command tree and returns the process
// exit code. It is the program's single exit point: the subcommands // exit code. It is the program's single exit point: the subcommands
// return their errors instead of exiting, so every deferred cleanup — // return their errors instead of exiting, so every deferred cleanup —
// above all closing the database, which checkpoints the SQLite WAL — // above all closing the database, which checkpoints the SQLite WAL —
// runs before the process ends. // runs before the process ends. The report and trees subcommands write
func run(args []string, stderr io.Writer) int { // their data to stdout.
func run(args []string, stdout, stderr io.Writer) int {
// A nil slice makes cobra fall back to os.Args, which would let a // A nil slice makes cobra fall back to os.Args, which would let a
// test binary's own flags reach the command tree. // test binary's own flags reach the command tree.
if args == nil { if args == nil {
args = []string{} args = []string{}
} }
root := newRootCommand(stderr) root := newRootCommand(stdout, stderr)
root.SetArgs(args) root.SetArgs(args)
err := root.Execute() err := root.Execute()
@@ -94,10 +103,10 @@ func run(args []string, stderr io.Writer) int {
// newRootCommand builds the command tree. Everything on stdout is // newRootCommand builds the command tree. Everything on stdout is
// machine-readable data; all human-facing output (help, usage, errors) // machine-readable data; all human-facing output (help, usage, errors)
// goes to stderr. // goes to stderr.
func newRootCommand(stderr io.Writer) *cobra.Command { func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
root := &cobra.Command{ root := &cobra.Command{
Use: "sfdupes", Use: "sfdupes",
Short: "Find candidate duplicate files by size and head/tail SHA-256", Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
Version: Version, Version: Version,
Args: cobra.NoArgs, Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, _ []string) error { RunE: func(cmd *cobra.Command, _ []string) error {
@@ -127,7 +136,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
}), }),
} }
scanCmd.Flags().IntVar(&scanWorkers, "workers", runtime.NumCPU(), scanCmd.Flags().IntVar(&scanWorkers, "workers", runtime.NumCPU(),
"concurrent workers for the walk and hash phases") "concurrent workers for the walk, hash, and content phases")
scanCmd.Flags().BoolVarP(&scanOneFS, "one-file-system", "x", false, scanCmd.Flags().BoolVarP(&scanOneFS, "one-file-system", "x", false,
"do not cross filesystem boundaries") "do not cross filesystem boundaries")
@@ -136,7 +145,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the file-level duplicates report", Short: "Read the scan database and print the file-level duplicates report",
Args: cobra.NoArgs, Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error { RunE: runE(func(ctx context.Context, _ []string) error {
return runReport(ctx) return runReport(ctx, stdout)
}), }),
} }
@@ -145,7 +154,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the duplicate-tree report", Short: "Read the scan database and print the duplicate-tree report",
Args: cobra.NoArgs, Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error { RunE: runE(func(ctx context.Context, _ []string) error {
return runTrees(ctx) return runTrees(ctx, stdout)
}), }),
} }
+376 -55
View File
@@ -44,23 +44,61 @@ func assertNoSidecars(t *testing.T, path string) {
} }
} }
// captureStdout redirects os.Stdout to a file for the rest of the test // makeReadOnly takes write permission away from the database at path,
// and returns a function reading back everything written to it. Only // from any WAL sidecar beside it, and from their directory, as for a
// machine-readable data belongs on stdout (README design goal 4), so // user reading a database that a root cron scan keeps. Root ignores
// the tests assert on it directly. // file permissions, so it skips the test when run as root.
func captureStdout(t *testing.T) func() string { func makeReadOnly(t *testing.T, path string) {
t.Helper() t.Helper()
f, err := os.Create(filepath.Join(t.TempDir(), "stdout")) if os.Geteuid() == 0 {
t.Skip("root ignores file permissions")
}
err := os.Chmod(path, 0o400)
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
saved := os.Stdout for _, suffix := range walSuffixes {
os.Stdout = f err = os.Chmod(path+suffix, 0o400)
if err != nil && !errors.Is(err, fs.ErrNotExist) {
t.Fatal(err)
}
}
dir := filepath.Dir(path)
//nolint:gosec // reaching the database needs the search bit
err = os.Chmod(dir, 0o500)
if err != nil {
t.Fatal(err)
}
// Runs before t.TempDir's own cleanup, which must delete the files.
t.Cleanup(func() {
//nolint:gosec // removing the directory needs its search bit back
_ = os.Chmod(dir, 0o700)
})
}
// captureStderr redirects os.Stderr to a file for the rest of the test
// and returns a function reading back everything written to it. scan
// writes its warnings and summary straight to os.Stderr, not to the
// stderr writer run is given.
func captureStderr(t *testing.T) func() string {
t.Helper()
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
os.Stderr = f
t.Cleanup(func() { t.Cleanup(func() {
os.Stdout = saved os.Stderr = saved
_ = f.Close() _ = f.Close()
}) })
@@ -91,13 +129,13 @@ func captureStdout(t *testing.T) func() string {
// brokenDatabase writes a database that opens cleanly and passes the // brokenDatabase writes a database that opens cleanly and passes the
// schema-version check but has no files table, so the first query // schema-version check but has no files table, so the first query
// fails with the database already open: a fatal error on a path that // fails with the database already open: a fatal error on a path that
// owns an open database. // owns an open database. It closes the database the way scan does.
func brokenDatabase(t *testing.T) string { func brokenDatabase(t *testing.T) string {
t.Helper() t.Helper()
path := testDBPath(t) path := testDBPath(t)
db, err := openDB(path) db, err := openDB(path, scanParams)
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -108,10 +146,7 @@ func brokenDatabase(t *testing.T) string {
t.Fatal(err) t.Fatal(err)
} }
err = db.Close() closeScanDatabase(t.Context(), db, path)
if err != nil {
t.Fatal(err)
}
return path return path
} }
@@ -144,7 +179,10 @@ func TestOpenDatabaseKeepsWALWhileOpen(t *testing.T) {
func TestRunFatalAfterOpenClosesDatabase(t *testing.T) { func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
// Every subcommand that owns an open database must close it when // Every subcommand that owns an open database must close it when
// it fails: no os.Exit between the open and the return. // it fails: no os.Exit between the open and the return. The
// sidecar check is evidence of the close only for scan: report and
// trees only read a database that is out of WAL mode, which leaves
// nothing on disk whether they close it or not.
cases := map[string][]string{ cases := map[string][]string{
cmdScan: {cmdScan}, cmdScan: {cmdScan},
cmdReport: {cmdReport}, cmdReport: {cmdReport},
@@ -160,17 +198,15 @@ func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
args = append(args, t.TempDir()) args = append(args, t.TempDir())
} }
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run(args, &stdout, &stderr)
code := run(args, &stderr)
if code != exitFatal { if code != exitFatal {
t.Errorf("run(%v) = %d, want %d", args, code, exitFatal) t.Errorf("run(%v) = %d, want %d", args, code, exitFatal)
} }
assertNoSidecars(t, path) assertNoSidecars(t, path)
assertFatalOutput(t, stderr.String(), stdout()) assertFatalOutput(t, stderr.String(), stdout.String())
// Proof that the failure happened after the open: only a // Proof that the failure happened after the open: only a
// query against the opened database can report this. // query against the opened database can report this.
@@ -188,18 +224,16 @@ func TestRunMissingOperandIsFatalNotUsage(t *testing.T) {
// must not dump the usage text. // must not dump the usage text.
t.Setenv(databaseEnv, testDBPath(t)) t.Setenv(databaseEnv, testDBPath(t))
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t)
missing := filepath.Join(t.TempDir(), "nope") missing := filepath.Join(t.TempDir(), "nope")
code := run([]string{cmdScan, missing}, &stderr) code := run([]string{cmdScan, missing}, &stdout, &stderr)
if code != exitFatal { if code != exitFatal {
t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal) t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal)
} }
assertFatalOutput(t, stderr.String(), stdout()) assertFatalOutput(t, stderr.String(), stdout.String())
} }
// assertFatalOutput checks that a fatal error was reported the way // assertFatalOutput checks that a fatal error was reported the way
@@ -243,11 +277,9 @@ func TestRunUsageErrors(t *testing.T) {
// path that does not exist. // path that does not exist.
t.Setenv(databaseEnv, testDBPath(t)) t.Setenv(databaseEnv, testDBPath(t))
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run(tc.args, &stdout, &stderr)
code := run(tc.args, &stderr)
if code != exitUsage { if code != exitUsage {
t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage) t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage)
} }
@@ -256,7 +288,7 @@ func TestRunUsageErrors(t *testing.T) {
t.Errorf("stderr = %q, want %q", stderr.String(), tc.want) t.Errorf("stderr = %q, want %q", stderr.String(), tc.want)
} }
if got := stdout(); got != "" { if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got) t.Errorf("stdout = %q, want nothing (data only)", got)
} }
}) })
@@ -265,9 +297,9 @@ func TestRunUsageErrors(t *testing.T) {
// TestRunHelpAndVersionSucceed checks that the two informational flags // TestRunHelpAndVersionSucceed checks that the two informational flags
// exit 0 and keep their human-facing output on stderr. // exit 0 and keep their human-facing output on stderr.
//
//nolint:paralleltest // captureStdout replaces the process-wide os.Stdout
func TestRunHelpAndVersionSucceed(t *testing.T) { func TestRunHelpAndVersionSucceed(t *testing.T) {
t.Parallel()
assertHumanOutput(t, "--help") assertHumanOutput(t, "--help")
assertHumanOutput(t, "--version") assertHumanOutput(t, "--version")
} }
@@ -278,11 +310,9 @@ func TestRunHelpAndVersionSucceed(t *testing.T) {
func assertHumanOutput(t *testing.T, arg string) { func assertHumanOutput(t *testing.T, arg string) {
t.Helper() t.Helper()
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run([]string{arg}, &stdout, &stderr)
code := run([]string{arg}, &stderr)
if code != exitOK { if code != exitOK {
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK) t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
} }
@@ -291,7 +321,7 @@ func assertHumanOutput(t *testing.T, arg string) {
t.Errorf("run(%s) wrote nothing to stderr", arg) t.Errorf("run(%s) wrote nothing to stderr", arg)
} }
if got := stdout(); got != "" { if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got) t.Errorf("stdout = %q, want nothing (data only)", got)
} }
} }
@@ -320,21 +350,31 @@ func scanFixture(t *testing.T) []string {
t.Fatal(err) t.Fatal(err)
} }
var stderr bytes.Buffer scanOK(t, dir)
stdout := captureStdout(t) return dupes
}
code := run([]string{cmdScan, dir}, &stderr) // scanOK runs scan over operands, fails the test unless it exits 0 with
// nothing on stdout, and returns everything it printed to stderr.
func scanOK(t *testing.T, operands ...string) string {
t.Helper()
var stdout bytes.Buffer
stderr := captureStderr(t)
code := run(append([]string{cmdScan}, operands...), &stdout, os.Stderr)
if code != exitOK { if code != exitOK {
t.Fatalf("run(scan) = %d, want %d; stderr: %s", t.Fatalf("run(scan %q) = %d, want %d; stderr: %s",
code, exitOK, stderr.String()) operands, code, exitOK, stderr())
} }
if got := stdout(); got != "" { if got := stdout.String(); got != "" {
t.Errorf("scan stdout = %q, want nothing (data only)", got) t.Errorf("scan stdout = %q, want nothing (data only)", got)
} }
return dupes return stderr()
} }
func TestRunScanSucceedsDespiteWarnings(t *testing.T) { func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
@@ -345,24 +385,119 @@ func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
assertNoSidecars(t, path) assertNoSidecars(t, path)
} }
func TestRunScanSkipsSymlinkOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dir := t.TempDir()
writeFile(t, dir, "target/sub/f", pattern(1, 10))
link := filepath.Join(dir, "link")
err := os.Symlink(filepath.Join(dir, "target"), link)
if err != nil {
t.Fatal(err)
}
// Scanning a directory through the symlink stores a record beneath
// the symlink's own path for a file beneath its target.
scanOK(t, filepath.Join(link, "sub"))
assertOperandSkipped(t, path, link, "symlink",
filepath.Join(link, "sub", "f"))
}
func TestRunScanWalksOperandUnderSymlinkOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dir := t.TempDir()
writeFile(t, dir, "target/sub/f", pattern(1, 10))
link := filepath.Join(dir, "link")
err := os.Symlink(filepath.Join(dir, "target"), link)
if err != nil {
t.Fatal(err)
}
// link is dropped as a symlink, but link/sub must still be scanned,
// not dropped as lying under link.
scanOK(t, link, filepath.Join(link, "sub"))
db, err := openDB(path, reportParams)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
recordByPath(t, dbRecords(t, db), filepath.Join(link, "sub", "f"))
}
func TestRunScanSkipsZFSOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
zfs := filepath.Join(t.TempDir(), ".zfs")
snapshot := filepath.Join(zfs, "snapshot", "hourly")
f := writeFile(t, snapshot, "f", pattern(1, 10))
// An operand beneath a .zfs directory is walked, because it is not
// itself named .zfs.
scanOK(t, snapshot)
assertOperandSkipped(t, path, zfs, ".zfs directory", f)
}
// assertOperandSkipped scans operand alone and checks that it is skipped
// as kind: a warning naming it, one skip in the summary, exit 0, and the
// record for kept, which an earlier scan stored beneath operand, still
// in the database at dbPath.
func assertOperandSkipped(t *testing.T, dbPath, operand, kind,
kept string,
) {
t.Helper()
stderr := scanOK(t, operand)
warning := "walk " + operand + ": skipping " + kind + " operand\n"
if !strings.Contains(stderr, warning) {
t.Errorf("stderr = %q, want %q", stderr, warning)
}
summary := "scan: 0 files seen (0 added, 0 updated, 0 removed, " +
"0 unchanged), 1 skipped\n"
if !strings.Contains(stderr, summary) {
t.Errorf("stderr = %q, want %q", stderr, summary)
}
db, err := openDB(dbPath, reportParams)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
recordByPath(t, dbRecords(t, db), kept)
}
func TestRunReportSucceeds(t *testing.T) { func TestRunReportSucceeds(t *testing.T) {
path := testDBPath(t) path := testDBPath(t)
t.Setenv(databaseEnv, path) t.Setenv(databaseEnv, path)
dupes := scanFixture(t) dupes := scanFixture(t)
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run([]string{cmdReport}, &stdout, &stderr)
code := run([]string{cmdReport}, &stderr)
if code != exitOK { if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s", t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String()) code, exitOK, stderr.String())
} }
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n" want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
if got := stdout(); got != want { if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want) t.Errorf("stdout = %q, want %q", got, want)
} }
@@ -375,11 +510,9 @@ func TestRunTreesSucceeds(t *testing.T) {
dupes := scanFixture(t) dupes := scanFixture(t)
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run([]string{cmdTrees}, &stdout, &stderr)
code := run([]string{cmdTrees}, &stderr)
if code != exitOK { if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s", t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String()) code, exitOK, stderr.String())
@@ -389,9 +522,197 @@ func TestRunTreesSucceeds(t *testing.T) {
// trees of each other. // trees of each other.
want := "first\tdupe\tfiles\tsize\n" + want := "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n" filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
if got := stdout(); got != want { if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want) t.Errorf("stdout = %q, want %q", got, want)
} }
assertNoSidecars(t, path) assertNoSidecars(t, path)
} }
func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
// README §Database: report and trees need only read access to the
// database file. With its directory read-only as well, SQLite
// cannot create any file beside it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dupes := scanFixture(t)
assertNoSidecars(t, path)
makeReadOnly(t, path)
cases := map[string]string{
cmdReport: "first\tdupe\tsize\n" +
dupes[0] + "\t" + dupes[1] + "\t300\n",
cmdTrees: "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) +
"\t1\t300\n",
}
for name, want := range cases {
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
continue
}
if got := stdout.String(); got != want {
t.Errorf("%s stdout = %q, want %q", name, got, want)
}
}
}
// holdScanLock takes the lock on the database at path, as a running
// scan does, and holds it until the test ends. It fails the test when
// the lock is already held.
func holdScanLock(t *testing.T, path string) {
t.Helper()
lock, err := lockScanDatabase(path)
if err != nil {
t.Fatalf("lock %s: %v", path, err)
}
t.Cleanup(func() { _ = lock.Close() })
}
func TestRunSecondScanFails(t *testing.T) {
// README §Database: while one scan holds the lock, a second scan
// fails at once, naming the lock file, without creating the
// database.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
holdScanLock(t, path)
var stdout, stderr bytes.Buffer
code := run([]string{cmdScan, t.TempDir()}, &stdout, &stderr)
if code != exitFatal {
t.Errorf("run(scan) = %d, want %d", code, exitFatal)
}
want := "sfdupes: another scan is running (lock held on " +
path + ".lock)\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
_, err := os.Stat(path)
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("stat %s = %v, want the database not created", path, err)
}
}
func TestRunScanReleasesLock(t *testing.T) {
// README §Database: a scan releases the lock however it ends.
t.Run("success", func(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
})
t.Run("fatal error", func(t *testing.T) {
path := brokenDatabase(t)
t.Setenv(databaseEnv, path)
code := run([]string{cmdScan, t.TempDir()}, io.Discard, io.Discard)
if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d", code, exitFatal)
}
holdScanLock(t, path)
})
}
func TestRunReportsDuringScan(t *testing.T) {
// README §Database: report and trees never take the lock, so they
// run while a scan holds it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
for _, name := range []string{cmdReport, cmdTrees} {
var stderr bytes.Buffer
code := run([]string{name}, io.Discard, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
}
}
}
func TestRunStdoutClosedIsFatal(t *testing.T) {
// README §Error handling: a stdout write failure exits 1, reported
// in one line on stderr.
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
stdout, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
if err != nil {
t.Fatal(err)
}
err = stdout.Close()
if err != nil {
t.Fatal(err)
}
var stderr bytes.Buffer
code := run([]string{name}, stdout, &stderr)
if code != exitFatal {
t.Errorf("run(%s) = %d, want %d", name, code, exitFatal)
}
got := stderr.String()
if !strings.HasPrefix(got, "sfdupes: write stdout: ") ||
!strings.Contains(got, os.ErrClosed.Error()) ||
strings.Count(got, "\n") != 1 {
t.Errorf("stderr = %q, want one line reporting the "+
"failed stdout write", got)
}
})
}
}
// errWriteFailed is the error failingWriter returns.
var errWriteFailed = errors.New("write failed")
// failingWriter is a stdout that fails every write.
type failingWriter struct{}
func (failingWriter) Write([]byte) (int, error) { return 0, errWriteFailed }
func TestStdoutWriteErrorPropagates(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
cases := map[string]func(context.Context, io.Writer) error{
cmdReport: runReport,
cmdTrees: runTrees,
}
for name, fn := range cases {
err := fn(t.Context(), failingWriter{})
if !errors.Is(err, errWriteFailed) {
t.Errorf("%s: error = %v, want %v", name, err, errWriteFailed)
}
}
}
+3 -1
View File
@@ -106,6 +106,8 @@ func (p *progress) increment() {
} }
// warnf prints a one-line warning to stderr without corrupting the bar. // warnf prints a one-line warning to stderr without corrupting the bar.
// The whole message is escaped like a report's path columns, so a path
// holding a newline cannot split the warning.
func (p *progress) warnf(format string, args ...any) { func (p *progress) warnf(format string, args ...any) {
if p == nil { if p == nil {
return return
@@ -115,7 +117,7 @@ func (p *progress) warnf(format string, args ...any) {
_ = p.bar.Clear() _ = p.bar.Clear()
} }
fmt.Fprintf(os.Stderr, format+"\n", args...) fmt.Fprintln(os.Stderr, escapePath(fmt.Sprintf(format, args...)))
} }
// finish terminates the pass's display. // finish terminates the pass's display.
+55 -97
View File
@@ -4,8 +4,8 @@ import (
"bufio" "bufio"
"context" "context"
"fmt" "fmt"
"io"
"os" "os"
"slices"
"strings" "strings"
) )
@@ -17,82 +17,70 @@ const ioBufSize = 1 << 20
const minGroupSize = 2 const minGroupSize = 2
// scanRec is one file record from the database. The signature (size, // scanRec is one file record from the database. The signature (size,
// head, tail) is the duplicate key; mtime is informational only and // head, tail, content) is the duplicate key; mtime is informational
// used by scan for change detection. // only and used by scan for change detection.
type scanRec struct { type scanRec struct {
size int64 size int64
mtime int64 mtime int64
head string head string
tail string tail string
path string content string
path string
} }
// loadRecords opens the database and reads every file record for the // runReport implements the report subcommand: it prints the file-level
// report and trees subcommands. Any database problem — including a // duplicates report as TSV on stdout. SQLite groups and orders the
// missing database — is fatal. The error is returned rather than // records, and each row is written as it is read, so no group is held
// exiting, so that the deferred close — which checkpoints the SQLite // in memory. It never touches the scanned filesystem; its only I/O is
// WAL — always runs; the database is closed before the caller formats // the database (with SQLite's temporary sort file), stdout, and stderr.
// its output, so it stays closed even if that output fails. // Any database problem, including a missing database, is fatal.
func loadRecords(ctx context.Context) ([]scanRec, error) { func runReport(ctx context.Context, stdout io.Writer) error {
dbPath := databasePath() dbPath := databasePath()
db, err := openReportDatabase(ctx, dbPath) db, err := openReportDatabase(ctx, dbPath)
if err != nil { if err != nil {
return nil, err return err
} }
defer func() { _ = db.Close() }() defer func() { _ = db.Close() }()
recs, err := loadFileRows(ctx, db) out := bufio.NewWriterSize(stdout, ioBufSize)
if err != nil {
return nil, fmt.Errorf("database %s: %w", dbPath, err)
}
return recs, nil
}
// dupeGroup is one set of candidate-duplicate files: identical size,
// head hash, and tail hash. paths is sorted lexicographically; the
// first entry is the group's "first", the rest are dupes.
type dupeGroup struct {
size int64
paths []string
}
// runReport implements the report subcommand: it reads every record
// from the database and prints the file-level duplicates report as TSV
// on stdout. It never touches the scanned filesystem; its only I/O is
// the database, stdout, and stderr.
func runReport(ctx context.Context) error {
recs, err := loadRecords(ctx)
if err != nil {
return err
}
dupes := collectDupeGroups(recs)
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tsize") _, err = fmt.Fprintln(out, "first\tdupe\tsize")
if err != nil { if err != nil {
return fmt.Errorf("write stdout: %w", err) return fmt.Errorf("write stdout: %w", err)
} }
dupeFiles := 0 var (
groups, dupeFiles int
reclaimable int64
writeErr error
)
var reclaimable int64 records, err := loadDupeRows(ctx, db,
func(first, path string, size int64) error {
// A group's first path is its first row; every other
// path is a dupe.
if path == first {
groups++
for _, g := range dupes { return nil
for _, p := range g.paths[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\n",
g.paths[0], p, g.size)
if err != nil {
return fmt.Errorf("write stdout: %w", err)
} }
_, writeErr = fmt.Fprintf(out, "%s\t%s\t%d\n",
escapePath(first), escapePath(path), size)
dupeFiles++ dupeFiles++
reclaimable += g.size reclaimable += size
}
return writeErr
})
if writeErr != nil {
return fmt.Errorf("write stdout: %w", writeErr)
}
if err != nil {
return fmt.Errorf("database %s: %w", dbPath, err)
} }
err = out.Flush() err = out.Flush()
@@ -103,54 +91,24 @@ func runReport(ctx context.Context) error {
fmt.Fprintf(os.Stderr, fmt.Fprintf(os.Stderr,
"report: %d records read, %d duplicate groups, %d dupe files, "+ "report: %d records read, %d duplicate groups, %d dupe files, "+
"%s reclaimable\n", "%s reclaimable\n",
len(recs), len(dupes), dupeFiles, humanBytes(reclaimable)) records, groups, dupeFiles, humanBytes(reclaimable))
return nil return nil
} }
// collectDupeGroups groups records by signature and returns every group // escapePath returns a path as it is written in a report column (README
// with two or more paths, each group's paths sorted lexicographically, // "Report output format"): a backslash, tab, newline or carriage return
// groups ordered by size descending then by first path ascending. // becomes \\, \t, \n or \r, and every other byte is kept as it is.
func collectDupeGroups(recs []scanRec) []dupeGroup { // Grouping and sorting use the raw path, never this form.
groups := make(map[fileSig][]string) func escapePath(p string) string {
// Most paths need no escaping; skip building a replacer for them.
for _, r := range recs { if !strings.ContainsAny(p, "\\\t\n\r") {
// A record without hashes (its size was unique when last return p
// scanned) has unknown content and is never reported as a
// duplicate.
if r.head == "" {
continue
}
k := fileSig{size: r.size, head: r.head, tail: r.tail}
groups[k] = append(groups[k], r.path)
} }
var dupes []dupeGroup return strings.NewReplacer(
`\`, `\\`, "\t", `\t`, "\n", `\n`, "\r", `\r`,
for k, paths := range groups { ).Replace(p)
if len(paths) < minGroupSize {
continue
}
slices.Sort(paths)
dupes = append(dupes, dupeGroup{size: k.size, paths: paths})
}
// Biggest reclaimable space first; ties broken by first path.
slices.SortFunc(dupes, func(a, b dupeGroup) int {
if a.size != b.size {
if a.size > b.size {
return -1
}
return 1
}
return strings.Compare(a.paths[0], b.paths[0])
})
return dupes
} }
// humanBytes formats a byte count in human units (binary prefixes). // humanBytes formats a byte count in human units (binary prefixes).
+270 -26
View File
@@ -1,26 +1,244 @@
package main package main
import ( import (
"bytes"
"database/sql"
"io"
"os"
"path/filepath"
"slices" "slices"
"strings"
"testing" "testing"
) )
func TestCollectDupeGroups(t *testing.T) { // awkwardDir is a directory name holding every byte the reports escape.
const awkwardDir = "/d/\tone\ntwo\rthree\\four"
// awkwardPairRecs is a duplicate pair in sibling directories /d/A and
// awkwardDir. A raw tab sorts before "A" but its escaped form `\t`
// sorts after it, so awkwardDir coming first shows that sorting uses
// the raw path.
func awkwardPairRecs() []scanRec {
return []scanRec{
{size: 5, head: "h", tail: "t", content: "c", path: "/d/A/f"},
{size: 5, head: "h", tail: "t", content: "c", path: awkwardDir + "/f"},
}
}
// seedDatabase writes recs into a fresh database and returns its path.
func seedDatabase(t *testing.T, recs []scanRec) string {
t.Helper()
path := testDBPath(t)
db, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
err = applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
err = db.Close()
if err != nil {
t.Fatal(err)
}
return path
}
// dupeGroup is one duplicate group as report reads it: the size, and
// the paths in report order, first path first.
type dupeGroup struct {
size int64
paths []string
}
// dupeGroups returns the duplicate groups report reads from db, in
// report order.
func dupeGroups(t *testing.T, db *sql.DB) []dupeGroup {
t.Helper()
var groups []dupeGroup
_, err := loadDupeRows(t.Context(), db,
func(first, path string, size int64) error {
if path == first {
groups = append(groups, dupeGroup{size: size})
}
g := &groups[len(groups)-1]
g.paths = append(g.paths, path)
return nil
})
if err != nil {
t.Fatal(err)
}
return groups
}
// dupeGroupsOf writes recs into a fresh database and returns the
// duplicate groups report reads from it.
func dupeGroupsOf(t *testing.T, recs []scanRec) []dupeGroup {
t.Helper()
db := openTestDB(t)
err := applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
return dupeGroups(t, db)
}
func TestRunReportEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
code := run([]string{cmdReport}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tsize\n" +
`/d/\tone\ntwo\rthree\\four/f` + "\t/d/A/f\t5\n"
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
func TestRunReportsIgnoreInsertionOrder(t *testing.T) {
// README §Constraints: identical database contents give identical
// output, whatever order the records were inserted in.
recs := append(smokeTreeRecs(), awkwardPairRecs()...)
recs = append(recs,
scanRec{size: 50, head: "b", tail: "b", content: "b", path: "/y/2"},
scanRec{size: 50, head: "b", tail: "b", content: "b", path: "/y/1"},
scanRec{size: 50, head: "a", tail: "a", content: "a", path: "/x/2"},
scanRec{size: 50, head: "a", tail: "a", content: "a", path: "/x/1"},
scanRec{size: 50, path: "/x/unhashed"},
)
reversed := slices.Clone(recs)
slices.Reverse(reversed)
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, recs))
forward := runStdout(t, name)
t.Setenv(databaseEnv, seedDatabase(t, reversed))
backward := runStdout(t, name)
if strings.Count(forward, "\n") < 3 {
t.Errorf("stdout = %q, want at least two rows", forward)
}
if forward != backward {
t.Errorf("stdout depends on insertion order: %q vs %q",
forward, backward)
}
})
}
}
// runStdout runs the subcommand name and returns its stdout, failing
// the test unless it succeeds.
func runStdout(t *testing.T, name string) string {
t.Helper()
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
}
return stdout.String()
}
func TestEscapePath(t *testing.T) {
t.Parallel()
cases := map[string]string{
"/srv/plain": "/srv/plain",
"/a\tb": `/a\tb`,
"/a\nb": `/a\nb`,
"/a\rb": `/a\rb`,
`/a\b`: `/a\\b`,
`/a\tb`: `/a\\tb`,
"/not-utf8\xff": "/not-utf8\xff",
}
for in, want := range cases {
if got := escapePath(in); got != want {
t.Errorf("escapePath(%q) = %q, want %q", in, got, want)
}
}
}
// TestWarnfEscapes checks that a warning naming a path that holds a
// newline is still one line.
//
//nolint:paralleltest // replaces the process-wide os.Stderr
func TestWarnfEscapes(t *testing.T) {
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
os.Stderr = f
t.Cleanup(func() {
os.Stderr = saved
_ = f.Close()
})
(&progress{}).warnf("stat %s: %s", "/d/a\nb", "gone")
_, err = f.Seek(0, io.SeekStart)
if err != nil {
t.Fatal(err)
}
got, err := io.ReadAll(f)
if err != nil {
t.Fatal(err)
}
want := `stat /d/a\nb: gone` + "\n"
if string(got) != want {
t.Errorf("warning = %q, want %q", got, want)
}
}
func TestDupeGroups(t *testing.T) {
t.Parallel() t.Parallel()
recs := []scanRec{ recs := []scanRec{
{size: 100, head: "h", tail: "t", path: "/z/b"}, {size: 100, head: "h", tail: "t", content: "c", path: "/z/b"},
{size: 100, head: "h", tail: "t", path: "/z/a"}, {size: 100, head: "h", tail: "t", content: "c", path: "/z/a"},
{size: 100, head: "h", tail: "t", path: "/z/c"}, {size: 100, head: "h", tail: "t", content: "c", path: "/z/c"},
{size: 4000, head: "H", tail: "T", path: "/big/2"}, {size: 4000, head: "H", tail: "T", content: "C", path: "/big/2"},
{size: 4000, head: "H", tail: "T", path: "/big/1"}, {size: 4000, head: "H", tail: "T", content: "C", path: "/big/1"},
// Same size as the /z group but a different head hash. // Same size as the /z group but a different head hash.
{size: 100, head: "other", tail: "t", path: "/z/d"}, {size: 100, head: "other", tail: "t", content: "c", path: "/z/d"},
// A singleton signature must not form a group. // A singleton signature must not form a group.
{size: 7, head: "u", tail: "u", path: "/lonely"}, {size: 7, head: "u", tail: "u", content: "u", path: "/lonely"},
} }
groups := collectDupeGroups(recs) groups := dupeGroupsOf(t, recs)
if len(groups) != 2 { if len(groups) != 2 {
t.Fatalf("len(groups) = %d, want 2", len(groups)) t.Fatalf("len(groups) = %d, want 2", len(groups))
} }
@@ -38,33 +256,59 @@ func TestCollectDupeGroups(t *testing.T) {
} }
} }
func TestCollectDupeGroupsMtimeExcluded(t *testing.T) { func TestDupeGroupsContentSeparates(t *testing.T) {
t.Parallel()
// Same size, head, and tail, but different content hashes: the final
// rung keeps them apart, so no group forms. Matching content groups.
// Records without a content hash never group, not even with each
// other.
recs := []scanRec{
{size: 100, head: "h", tail: "t", content: "c1", path: "/a"},
{size: 100, head: "h", tail: "t", content: "c2", path: "/b"},
{size: 100, head: "h", tail: "t", content: "c1", path: "/c"},
{size: 100, head: "h", tail: "t", path: "/d"},
{size: 100, head: "h", tail: "t", path: "/e"},
}
groups := dupeGroupsOf(t, recs)
if len(groups) != 1 {
t.Fatalf("len(groups) = %d, want 1 (only the matching content)",
len(groups))
}
if !slices.Equal(groups[0].paths, []string{"/a", "/c"}) {
t.Errorf("group paths = %q, want /a /c", groups[0].paths)
}
}
func TestDupeGroupsMtimeExcluded(t *testing.T) {
t.Parallel() t.Parallel()
// mtime is informational only; records differing only in mtime // mtime is informational only; records differing only in mtime
// still group together. // still group together.
recs := []scanRec{ recs := []scanRec{
{size: 9, mtime: 100, head: "h", tail: "t", path: "/m/1"}, {size: 9, mtime: 100, head: "h", tail: "t", content: "c", path: "/m/1"},
{size: 9, mtime: 200, head: "h", tail: "t", path: "/m/2"}, {size: 9, mtime: 200, head: "h", tail: "t", content: "c", path: "/m/2"},
} }
groups := collectDupeGroups(recs) groups := dupeGroupsOf(t, recs)
if len(groups) != 1 { if len(groups) != 1 {
t.Fatalf("len(groups) = %d, want 1", len(groups)) t.Fatalf("len(groups) = %d, want 1", len(groups))
} }
} }
func TestCollectDupeGroupsTieBreak(t *testing.T) { func TestDupeGroupsTieBreak(t *testing.T) {
t.Parallel() t.Parallel()
recs := []scanRec{ recs := []scanRec{
{size: 50, head: "b", tail: "b", path: "/beta/2"}, {size: 50, head: "b", tail: "b", content: "b", path: "/beta/2"},
{size: 50, head: "b", tail: "b", path: "/beta/1"}, {size: 50, head: "b", tail: "b", content: "b", path: "/beta/1"},
{size: 50, head: "a", tail: "a", path: "/alpha/2"}, {size: 50, head: "a", tail: "a", content: "a", path: "/alpha/2"},
{size: 50, head: "a", tail: "a", path: "/alpha/1"}, {size: 50, head: "a", tail: "a", content: "a", path: "/alpha/1"},
} }
groups := collectDupeGroups(recs) groups := dupeGroupsOf(t, recs)
if len(groups) != 2 { if len(groups) != 2 {
t.Fatalf("len(groups) = %d, want 2", len(groups)) t.Fatalf("len(groups) = %d, want 2", len(groups))
} }
@@ -76,22 +320,22 @@ func TestCollectDupeGroupsTieBreak(t *testing.T) {
} }
} }
func TestCollectDupeGroupsDeterministic(t *testing.T) { func TestDupeGroupsDeterministic(t *testing.T) {
t.Parallel() t.Parallel()
recs := []scanRec{ recs := []scanRec{
{size: 1, head: "a", tail: "a", path: "/p/1"}, {size: 1, head: "a", tail: "a", content: "a", path: "/p/1"},
{size: 1, head: "a", tail: "a", path: "/p/2"}, {size: 1, head: "a", tail: "a", content: "a", path: "/p/2"},
{size: 2, head: "b", tail: "b", path: "/q/1"}, {size: 2, head: "b", tail: "b", content: "b", path: "/q/1"},
{size: 2, head: "b", tail: "b", path: "/q/2"}, {size: 2, head: "b", tail: "b", content: "b", path: "/q/2"},
} }
forward := collectDupeGroups(recs) forward := dupeGroupsOf(t, recs)
reversed := slices.Clone(recs) reversed := slices.Clone(recs)
slices.Reverse(reversed) slices.Reverse(reversed)
backward := collectDupeGroups(reversed) backward := dupeGroupsOf(t, reversed)
if !slices.EqualFunc(forward, backward, func(a, b dupeGroup) bool { if !slices.EqualFunc(forward, backward, func(a, b dupeGroup) bool {
return a.size == b.size && slices.Equal(a.paths, b.paths) return a.size == b.size && slices.Equal(a.paths, b.paths)
}) { }) {
+487 -107
View File
@@ -6,7 +6,9 @@ import (
"crypto/sha256" "crypto/sha256"
"database/sql" "database/sql"
"encoding/hex" "encoding/hex"
"errors"
"fmt" "fmt"
"io"
"io/fs" "io/fs"
"os" "os"
"path/filepath" "path/filepath"
@@ -16,8 +18,40 @@ import (
"syscall" "syscall"
) )
// chunk is the number of bytes hashed from each end of a file. // The duplicate ladder (see hashSignature and README "Duplicate
const chunk = 1024 // detection"). A same-size candidate below headTailMin is hashed in
// full and compared directly; a larger one is separated first by the
// hashes of its end windows, then by a content hash that is exact below
// wholeFileMax and deliberately sampled at or above it. The hash phase
// reads only the end windows of a larger file; the content phase reads
// it for its content hash only once its size, head, and tail match
// another file's.
// headTailMin is the size threshold for the end-window gate. A file
// smaller than this is hashed in full directly, with no separate head
// and tail step: its head, tail, and content all carry the whole-file
// hash. A file this size or larger is separated first by its end
// windows.
const headTailMin = 10 * 1024 * 1024
// headTailWindow is the number of bytes hashed from each end of a file
// at or above headTailMin (the head and tail rungs). Because
// headTailMin is far larger than two windows, the head and tail windows
// never overlap.
const headTailWindow = 64 * 1024
// wholeFileMax is the size boundary between the two content rungs: a
// file strictly smaller than this is content-hashed in full; a file
// this size or larger is content-hashed by sampling.
const wholeFileMax = 50 * 1024 * 1024
// sampleStride is the spacing between content samples for large files:
// one window is read at each gigabyte-aligned offset (0, 1 GiB, ...).
const sampleStride = 1024 * 1024 * 1024
// sampleWindow is the number of bytes read at each large-file sample
// offset, truncated at end of file.
const sampleWindow = 1024 * 1024
// workQueueDepth bounds the job and result channels feeding the walk // workQueueDepth bounds the job and result channels feeding the walk
// and hash worker pools. // and hash worker pools.
@@ -44,16 +78,20 @@ type fileMeta struct {
hashed bool hashed bool
} }
// runScan implements the scan subcommand: three sequential phases — // runScan implements the scan subcommand: four sequential phases —
// walk (which stats each file as it is discovered), hash, update — // walk (which stats each file as it is discovered), hash, update,
// that synchronize the persistent database with the filesystem state // content — that synchronize the persistent database with the
// under the PATH operands. Only files whose size at least one other // filesystem state under the PATH operands. Only files whose size at
// file shares are ever hashed: a size-unique file cannot be a // least one other file shares are ever hashed: a size-unique file
// duplicate. Flag parsing and the at-least-one-operand check are done // cannot be a duplicate. A file of headTailMin or more gets its content
// by cobra. Errors are returned rather than exiting, so that the // hash only when its size, head, and tail match another file's. Flag
// deferred close — which checkpoints the SQLite WAL — always runs. // parsing and the at-least-one-operand check are done by cobra. The
// Cancelling ctx unwinds the worker pools and aborts the scan with the // scan holds the lock on the database for its whole run, so a second
// context's error. // scan fails before it walks the filesystem or opens the database.
// Errors are returned rather than exiting, so that the deferred close —
// which takes the database out of WAL mode — always runs, and the lock
// is released after it. Cancelling ctx unwinds the worker pools and
// aborts the scan with the context's error.
func runScan(ctx context.Context, roots []string, workers int, func runScan(ctx context.Context, roots []string, workers int,
oneFS bool, oneFS bool,
) error { ) error {
@@ -68,12 +106,19 @@ func runScan(ctx context.Context, roots []string, workers int,
dbPath := databasePath() dbPath := databasePath()
lock, err := lockScanDatabase(dbPath)
if err != nil {
return err
}
defer func() { _ = lock.Close() }()
db, err := openScanDatabase(ctx, dbPath) db, err := openScanDatabase(ctx, dbPath)
if err != nil { if err != nil {
return err return err
} }
defer func() { _ = db.Close() }() defer closeScanDatabase(ctx, db, dbPath)
st, err := syncScan(ctx, db, roots, workers, oneFS) st, err := syncScan(ctx, db, roots, workers, oneFS)
if err != nil { if err != nil {
@@ -166,20 +211,27 @@ type scanState struct {
} }
// syncScan synchronizes the database with the filesystem under roots // syncScan synchronizes the database with the filesystem under roots
// in three sequential phases: walk (enumerate and stat every file, // in four sequential phases: walk (enumerate and stat every file,
// building a complete size census), hash (read only the new or // building a complete size census), hash (read only the new or
// changed — or previously unhashed — files whose size at least one // changed — or previously unhashed — files whose size at least one
// other file shares, committing results in batches as they arrive), // other file shares, committing results in batches as they arrive),
// and update (record the size-unique files without reading them, and // update (record the size-unique files without reading them, and
// delete the records the scan no longer verifies). Records outside // delete the records the scan no longer verifies), and content (fill
// the roots are never touched. // in the content hash of every record of headTailMin or more whose
// size, head, and tail match another record's). Records outside the
// roots are never touched, except that the content phase fills in
// their content hash. Operands the walk cannot start from are dropped
// first, so the records beneath them count as outside the roots unless
// they lie under another root.
func syncScan(ctx context.Context, db *sql.DB, roots []string, func syncScan(ctx context.Context, db *sql.DB, roots []string,
workers int, oneFS bool, workers int, oneFS bool,
) (scanStats, error) { ) (scanStats, error) {
roots = pruneRoots(roots)
s := &scanState{db: db} s := &scanState{db: db}
// Types are checked before pruning so that an operand under a
// dropped one is still scanned, not dropped as lying under it.
roots = pruneRoots(s.walkableRoots(roots))
err := s.loadIndex(ctx, roots) err := s.loadIndex(ctx, roots)
if err != nil { if err != nil {
return s.st, err return s.st, err
@@ -207,7 +259,41 @@ func syncScan(ctx context.Context, db *sql.DB, roots []string,
return s.st, err return s.st, err
} }
return s.st, s.updatePhase(ctx) err = s.updatePhase(ctx)
if err != nil {
return s.st, err
}
return s.st, s.contentPhase(ctx, workers)
}
// walkableRoots returns the operands the walk can start from: regular
// files, and directories not named .zfs. Every other operand is warned
// about, counted as skipped, and dropped. A dropped operand is no
// longer a root, so the records stored beneath it count as outside the
// roots and are not deleted as unverified, unless it lies under another
// root. An operand that fails lstat here is kept, and the walk warns
// about it.
func (s *scanState) walkableRoots(roots []string) []string {
kept := make([]string, 0, len(roots))
for _, root := range roots {
fi, err := os.Lstat(root)
if err == nil {
warn := operandWarning(root, fi)
if warn != "" {
s.st.skipped++
fmt.Fprintln(os.Stderr, escapePath(warn))
continue
}
}
kept = append(kept, root)
}
return kept
} }
// loadIndex indexes the database records under the scan roots for // loadIndex indexes the database records under the scan roots for
@@ -241,7 +327,8 @@ func (s *scanState) loadIndex(ctx context.Context, roots []string) error {
// walkPhase drains the walk, appending every walked file's size to // walkPhase drains the walk, appending every walked file's size to
// the census and resolving what it can immediately: an unchanged file // the census and resolving what it can immediately: an unchanged file
// whose record already has hashes needs nothing further. It returns // whose record already has hashes needs nothing from the hash phase
// (the content phase may still fill in its content hash). It returns
// the new-or-changed files and the unchanged files whose records lack // the new-or-changed files and the unchanged files whose records lack
// hashes; both remain candidates until the census decides whether // hashes; both remain candidates until the census decides whether
// their sizes are shared. // their sizes are shared.
@@ -380,27 +467,40 @@ func sameInode(a, b fileRec) bool {
return (a.dev != 0 || a.ino != 0) && a.dev == b.dev && a.ino == b.ino return (a.dev != 0 || a.ino != 0) && a.dev == b.dev && a.ino == b.ino
} }
// hashPhase hashes every queued file with the worker pool — one read // hashPhase hashes every queued file with hashSignature — the head and
// per inode run, in inode order — committing completed records to the // tail of a file of headTailMin or more, the whole file below that —
// database in batches as results arrive, so a long scan persists its // committing completed records to the database in batches as results
// progress as it goes (an interrupted scan resumes cheaply: the next // arrive, so a long scan persists its progress as it goes (an
// run skips everything already recorded). The total counts actual // interrupted scan resumes cheaply: the next run skips everything
// reads, so the bar shows a real ETA. A run that fails to hash is // already recorded). A run that fails to hash is warned about and
// warned about and skipped; stale records for its paths, if any, are // skipped; stale records for its paths, if any, are deleted by the
// deleted by the update phase. // update phase.
func (s *scanState) hashPhase(ctx context.Context, workers int) error {
runs := hashRuns(s.toHash)
s.toHash = nil
return s.readRuns(ctx, workers, "hash", runs, hashSignature, s.recordRun)
}
// readRuns reads runs with the worker pool, one read per inode run, in
// the order given, under a progress display named label. The workers
// compute each run's hashes with hash, and each result goes to record;
// a run that fails to read is warned about and counted as skipped
// instead. The total counts actual reads, so the bar shows a real ETA.
// //
// Returning early — a failed database write, or a cancelled scan — must // Returning early — a failed database write, or a cancelled scan — must
// not strand the pool: the feeder would park forever on a full jobs // not strand the pool: the feeder would park forever on a full jobs
// channel and every worker on a full results channel. The deferred stop // channel and every worker on a full results channel. The deferred stop
// is what prevents that. // is what prevents that.
func (s *scanState) hashPhase(ctx context.Context, workers int) error { func (s *scanState) readRuns(ctx context.Context, workers int,
runs := hashRuns(s.toHash) label string, runs [][]fileRec,
s.toHash = nil hash func(path string, size int64) (string, string, string, error),
record func(ctx context.Context, r hashResult) error,
pool := startHashPool(ctx, runs, workers) ) error {
pool := startHashPool(ctx, runs, workers, hash)
defer pool.stop() defer pool.stop()
prog := newProgress("hash", int64(len(runs))) prog := newProgress(label, int64(len(runs)))
defer prog.finish() defer prog.finish()
for range runs { for range runs {
@@ -417,12 +517,12 @@ func (s *scanState) hashPhase(ctx context.Context, workers int) error {
if r.err != nil { if r.err != nil {
s.st.skipped += len(r.run) s.st.skipped += len(r.run)
prog.warnf("hash %s: %v", r.run[0].path, r.err) prog.warnf("%s %s: %v", label, r.run[0].path, r.err)
continue continue
} }
err := s.recordRun(ctx, r) err := record(ctx, r)
if err != nil { if err != nil {
return err return err
} }
@@ -439,14 +539,21 @@ func (s *scanState) recordRun(ctx context.Context, r hashResult) error {
s.resolve(rec.path) s.resolve(rec.path)
s.batch = append(s.batch, scanRec{ s.batch = append(s.batch, scanRec{
size: rec.size, size: rec.size,
mtime: rec.mtime, mtime: rec.mtime,
head: r.head, head: r.head,
tail: r.tail, tail: r.tail,
path: rec.path, content: r.content,
path: rec.path,
}) })
} }
return s.commitFullBatch(ctx)
}
// commitFullBatch commits the running batch once it holds
// updateBatchSize records.
func (s *scanState) commitFullBatch(ctx context.Context) error {
if len(s.batch) < updateBatchSize { if len(s.batch) < updateBatchSize {
return nil return nil
} }
@@ -504,6 +611,147 @@ func (s *scanState) updatePhase(ctx context.Context) error {
return applyChanges(ctx, s.db, nil, deletes, prog) return applyChanges(ctx, s.db, nil, deletes, prog)
} }
// contentPhase fills in the content hash of every record of headTailMin
// or more that lacks one and whose size, head, and tail equal another
// record's, anywhere in the database: records from this scan and
// records stored by earlier scans, inside or outside the roots. Only
// such a file can still be a duplicate, so no other file of headTailMin
// or more is read beyond its end windows. The files are read with the
// hash phase's worker pool and their records written back in batches. A
// failed read is warned about and counted as skipped; the record keeps
// its empty content, so it is never grouped, and a later scan tries
// again.
func (s *scanState) contentPhase(ctx context.Context, workers int) error {
toRead, recs, err := s.contentCandidates(ctx)
if err != nil {
return err
}
err = s.readRuns(ctx, workers, "content", hashRuns(toRead),
hashContentOnly, func(ctx context.Context, r hashResult) error {
// Every path in the run keeps its record's head and tail
// and gains the one content hash read for the run.
for _, f := range r.run {
rec := recs[f.path]
rec.content = r.content
s.batch = append(s.batch, rec)
}
return s.commitFullBatch(ctx)
})
if err != nil {
return err
}
return applyChanges(ctx, s.db, s.batch, nil, nil)
}
// contentCandidates returns the files the content phase reads, and
// their records by path. Every record contentCandidatesSQL returns has
// its file checked with lstat, whether or not it already has a content
// hash: a file that is gone, is no longer a regular file, or has
// changed by the walk's rule keeps its record as it is and does not
// count as a match for the others, and any other lstat error is warned
// about and counted as skipped, with the same result. If such a record
// has no content hash, it stays out of duplicate groups; if it has one,
// it is still reported until a scan covering its own tree updates or
// removes it. The files of a group that pass and have no content hash
// are read only if at least minGroupSize of the group's files pass, so
// a group whose other members are all stale costs no reads. Only the
// records to be read are kept.
func (s *scanState) contentCandidates(
ctx context.Context,
) ([]fileRec, map[string]scanRec, error) {
// The query and the checks take real time on a large database;
// without a display the scan looks hung before the reads begin.
prog := newProgress("content", -1)
defer prog.finish()
var (
toRead []fileRec
first scanRec // the current group's first record
passed int // the current group's files that passed the check
unread []fileRec // those of them without a content hash
)
recs := make(map[string]scanRec)
// endGroup queues the current group's files to read if at least
// minGroupSize of its files passed, and drops their records if not.
endGroup := func() {
if passed >= minGroupSize {
toRead = append(toRead, unread...)
} else {
for _, f := range unread {
delete(recs, f.path)
}
}
passed, unread = 0, nil
}
err := loadContentCandidates(ctx, s.db, func(r scanRec, hashed bool) {
prog.increment()
if r.size != first.size || r.head != first.head || r.tail != first.tail {
endGroup()
first = r
}
f, ok, err := unchangedFile(r)
if err != nil {
s.st.skipped++
prog.warnf("content %s: %v", r.path, err)
}
if !ok {
return
}
passed++
if !hashed {
unread = append(unread, f)
recs[r.path] = r
}
})
if err != nil {
return nil, nil, err
}
endGroup()
return toRead, recs, nil
}
// unchangedFile lstats the file r names and returns it for reading if
// it is still the regular file r records: the same size, and an mtime
// no newer than recorded (the walk's change rule). A file that is gone
// or has changed reports false; any other lstat error is returned.
func unchangedFile(r scanRec) (fileRec, bool, error) {
fi, err := os.Lstat(r.path)
if errors.Is(err, fs.ErrNotExist) {
return fileRec{}, false, nil
}
if err != nil {
return fileRec{}, false, err
}
if !fi.Mode().IsRegular() || fi.Size() != r.size ||
fi.ModTime().Unix() > r.mtime {
return fileRec{}, false, nil
}
dev, ino := inodeOfInfo(fi)
return fileRec{
path: r.path, size: r.size, mtime: r.mtime, dev: dev, ino: ino,
}, true, nil
}
// underAnyRoot reports whether path is any of the roots or lies under // underAnyRoot reports whether path is any of the roots or lies under
// one of them. // one of them.
func underAnyRoot(path string, roots []string) bool { func underAnyRoot(path string, roots []string) bool {
@@ -584,11 +832,43 @@ func sendEvent(ctx context.Context, events chan<- walkEvent,
} }
} }
// operandWarning returns the one-line warning for an operand the walk
// does not start from, naming the path and what it is, or "" for one it
// does: a regular file, or a directory not named .zfs. Symlinks are
// never followed, including as operands.
func operandWarning(root string, fi fs.FileInfo) string {
var kind string
switch mode := fi.Mode(); {
case mode.IsRegular():
return ""
case mode.IsDir():
if filepath.Base(root) != ".zfs" {
return ""
}
kind = ".zfs directory"
case mode&fs.ModeSymlink != 0:
kind = "symlink"
case mode&fs.ModeSocket != 0:
kind = "socket"
case mode&fs.ModeNamedPipe != 0:
kind = "FIFO"
case mode&fs.ModeDevice != 0:
kind = "device node"
default:
kind = "non-regular file"
}
return fmt.Sprintf("walk %s: skipping %s operand", root, kind)
}
// seedRoot turns one PATH operand into the walk's starting state: a // seedRoot turns one PATH operand into the walk's starting state: a
// regular-file operand is statted and emitted directly, a directory // regular-file operand is statted and emitted directly, and a directory
// operand becomes an initial job, and a symlink or other non-regular // operand becomes an initial job. walkableRoots has already dropped
// operand yields nothing (symlinks are never followed, including as // every other operand. One that has changed into something else since
// operands). // is warned about and skipped here; it is still a root, so the records
// stored beneath it are deleted as unverified.
func seedRoot(ctx context.Context, root string, func seedRoot(ctx context.Context, root string,
events chan<- walkEvent, events chan<- walkEvent,
) []dirJob { ) []dirJob {
@@ -602,30 +882,30 @@ func seedRoot(ctx context.Context, root string,
return nil return nil
} }
switch { warn := operandWarning(root, fi)
case fi.IsDir(): if warn != "" {
if filepath.Base(root) == ".zfs" { sendEvent(ctx, events, walkEvent{warn: warn, fail: true})
return nil
}
return nil
}
if fi.IsDir() {
dev, ok := deviceOfInfo(fi) dev, ok := deviceOfInfo(fi)
return []dirJob{{path: root, rootDev: dev, rootDevOK: ok}} return []dirJob{{path: root, rootDev: dev, rootDevOK: ok}}
case fi.Mode().IsRegular():
dev, ino := inodeOfInfo(fi)
sendEvent(ctx, events, walkEvent{rec: fileRec{
path: root,
size: fi.Size(),
mtime: fi.ModTime().Unix(),
dev: dev,
ino: ino,
}})
return nil
default:
return nil
} }
dev, ino := inodeOfInfo(fi)
sendEvent(ctx, events, walkEvent{rec: fileRec{
path: root,
size: fi.Size(),
mtime: fi.ModTime().Unix(),
dev: dev,
ino: ino,
}})
return nil
} }
// startWalkWorkers starts the walk worker pool. Each worker processes // startWalkWorkers starts the walk worker pool. Each worker processes
@@ -834,13 +1114,16 @@ func inodeOfInfo(fi fs.FileInfo) (uint64, uint64) {
return statDev(st), st.Ino return statDev(st), st.Ino
} }
// hashResult carries one inode run's head/tail hashes (or the error // hashResult carries the hashes computed for one inode run (or the
// that prevented hashing it) from the hash workers to the hash phase. // error that prevented computing them) from the pool's workers to the
// phase that started the pool: head, tail, and content from
// hashSignature, content alone from hashContentOnly.
type hashResult struct { type hashResult struct {
run []fileRec run []fileRec
head string head string
tail string tail string
err error content string
err error
} }
// hashPool owns every goroutine of the hash worker pool: the feeder // hashPool owns every goroutine of the hash worker pool: the feeder
@@ -856,10 +1139,10 @@ type hashPool struct {
} }
// startHashPool starts the feeder and the workers over runs. Workers // startHashPool starts the feeder and the workers over runs. Workers
// hash each run's first path (all paths in a run are hard links to the // hash each run's first path with hash (all paths in a run are hard
// same inode) and write one result per run. // links to the same inode) and write one result per run.
func startHashPool(ctx context.Context, runs [][]fileRec, func startHashPool(ctx context.Context, runs [][]fileRec, workers int,
workers int, hash func(path string, size int64) (string, string, string, error),
) *hashPool { ) *hashPool {
ctx, cancel := context.WithCancel(ctx) ctx, cancel := context.WithCancel(ctx)
@@ -871,7 +1154,7 @@ func startHashPool(ctx context.Context, runs [][]fileRec,
wg.Go(func() { feedHashJobs(ctx, runs, jobs) }) wg.Go(func() { feedHashJobs(ctx, runs, jobs) })
for range workers { for range workers {
wg.Go(func() { hashWorker(ctx, jobs, results) }) wg.Go(func() { hashWorker(ctx, jobs, results, hash) })
} }
done := make(chan struct{}) done := make(chan struct{})
@@ -917,24 +1200,25 @@ func feedHashJobs(ctx context.Context, runs [][]fileRec,
} }
} }
// hashWorker hashes one inode run at a time until jobs is closed or the // hashWorker hashes one inode run at a time with hash until jobs is
// scan is cancelled. A cancelled worker drops the runs still queued // closed or the scan is cancelled. A cancelled worker drops the runs
// instead of stopping its reads of jobs: the range must run out for the // still queued instead of stopping its reads of jobs: the range must
// pool to tear down, and reading a file nobody wants the hash of only // run out for the pool to tear down, and reading a file nobody wants
// delays that. // the hash of only delays that.
func hashWorker(ctx context.Context, jobs <-chan []fileRec, func hashWorker(ctx context.Context, jobs <-chan []fileRec,
results chan<- hashResult, results chan<- hashResult,
hash func(path string, size int64) (string, string, string, error),
) { ) {
for run := range jobs { for run := range jobs {
if ctx.Err() != nil { if ctx.Err() != nil {
continue continue
} }
head, tail, err := hashHeadTail(run[0].path, run[0].size) head, tail, content, err := hash(run[0].path, run[0].size)
select { select {
case results <- hashResult{ case results <- hashResult{
run: run, head: head, tail: tail, err: err, run: run, head: head, tail: tail, content: content, err: err,
}: }:
case <-ctx.Done(): case <-ctx.Done():
return return
@@ -942,55 +1226,151 @@ func hashWorker(ctx context.Context, jobs <-chan []fileRec,
} }
} }
// emptyHash is the lowercase-hex SHA-256 of the empty input: the head // emptyHash is the lowercase-hex SHA-256 of the empty input: the head,
// and tail hash of every zero-length file. // tail, and content hash of every zero-length file.
const emptyHash = "e3b0c44298fc1c149afbf4c8996fb924" + const emptyHash = "e3b0c44298fc1c149afbf4c8996fb924" +
"27ae41e4649b934ca495991b7852b855" "27ae41e4649b934ca495991b7852b855"
// hashHeadTail returns the lowercase-hex SHA-256 of the first // hashSignature computes the hashes the hash phase records for a file
// min(chunk, size) bytes and of the last min(chunk, size) bytes of the // whose size is shared; with the file size they form its duplicate
// file at path. The two reads overlap when size < 2*chunk. size is the // signature. A file below headTailMin is hashed in full and its
// value recorded when the file was statted; a zero-length file's // whole-file SHA-256 is returned as head, tail, and content alike —
// hashes are constant, so it is never even opened. // that range takes no separate end-window step. For a file at or above
func hashHeadTail(path string, size int64) (string, string, error) { // headTailMin only the head and tail are computed, the SHA-256 of its
// first and last headTailWindow bytes, and content is returned empty:
// the content phase computes it with hashContentOnly once the file's
// size, head, and tail match another file's. Two files are duplicates
// only when all four agree; any mismatch means not a duplicate. size
// is the value recorded when the file was statted; a zero-length file
// has constant hashes and is never opened.
func hashSignature(path string, size int64) (string, string, string, error) {
if size == 0 { if size == 0 {
return emptyHash, emptyHash, nil return emptyHash, emptyHash, emptyHash, nil
} }
//nolint:gosec // hashing operator-supplied paths is the tool's purpose //nolint:gosec // hashing operator-supplied paths is the tool's purpose
f, err := os.Open(path) f, err := os.Open(path)
if err != nil { if err != nil {
return "", "", err return "", "", "", err
} }
defer func() { _ = f.Close() }() defer func() { _ = f.Close() }()
n := min(int64(chunk), size) // Below the threshold the whole file is hashed directly, with no
// end-window step: head and tail both carry the whole-file hash.
if size < int64(headTailMin) {
content, err := hashWhole(f, size)
if err != nil {
return "", "", "", err
}
buf := make([]byte, n) return content, content, content, nil
}
_, err = f.ReadAt(buf, 0) head, tail, err := hashEnds(f, size)
if err != nil {
return "", "", "", err
}
return head, tail, "", nil
}
// hashContentOnly returns the content hash of the file at path, which
// is at least headTailMin bytes: the content phase's read. head and
// tail are returned empty, because the content phase keeps the ones its
// records already hold.
func hashContentOnly(path string, size int64) (string, string, string, error) {
//nolint:gosec // hashing operator-supplied paths is the tool's purpose
f, err := os.Open(path)
if err != nil {
return "", "", "", err
}
defer func() { _ = f.Close() }()
content, err := hashContent(f, size)
return "", "", content, err
}
// hashEnds returns the SHA-256 of the first and last headTailWindow
// bytes of f. It is called only for files at least headTailMin, which
// is far larger than two windows, so the windows never overlap and both
// reads are always full.
func hashEnds(f *os.File, size int64) (string, string, error) {
buf := make([]byte, headTailWindow)
_, err := f.ReadAt(buf, 0)
if err != nil { if err != nil {
return "", "", err return "", "", err
} }
h := sha256.Sum256(buf) h := sha256.Sum256(buf)
head := hex.EncodeToString(h[:])
// When the whole file fits in one chunk the tail window is exactly _, err = f.ReadAt(buf, size-int64(headTailWindow))
// the bytes just read: reuse the head hash instead of issuing a
// second read for every small file.
if size <= int64(chunk) {
hh := hex.EncodeToString(h[:])
return hh, hh, nil
}
_, err = f.ReadAt(buf, size-n)
if err != nil { if err != nil {
return "", "", err return "", "", err
} }
t := sha256.Sum256(buf) t := sha256.Sum256(buf)
return hex.EncodeToString(h[:]), hex.EncodeToString(t[:]), nil return head, hex.EncodeToString(t[:]), nil
}
// hashContent returns the content-rung hash of f: the SHA-256 of the
// whole file when it is smaller than wholeFileMax, or of sampled
// windows when it is that size or larger.
func hashContent(f *os.File, size int64) (string, error) {
if size >= int64(wholeFileMax) {
return hashSamples(f, size)
}
return hashWhole(f, size)
}
// hashWhole returns the SHA-256 of the entire file. A SectionReader is
// used so the read is independent of the offset left by any end-window
// reads. Reading fewer than size bytes means the file shrank between
// the stat and the hash; that is an error rather than a hash of content
// that no longer matches the recorded size.
func hashWhole(f *os.File, size int64) (string, error) {
h := sha256.New()
n, err := io.Copy(h, io.NewSectionReader(f, 0, size))
if err != nil {
return "", err
}
if n != size {
return "", fmt.Errorf("read %d of %d bytes: %w", n, size,
io.ErrUnexpectedEOF)
}
return hex.EncodeToString(h.Sum(nil)), nil
}
// hashSamples feeds sampleWindow bytes at each gigabyte-aligned offset
// (0, sampleStride, 2*sampleStride, ... while inside the file), in
// order, into one hash, each window truncated at end of file. This is
// the probabilistic large-file rung: two files of equal size agreeing
// on every sample are reported as duplicates without every byte being
// read. Because size is part of the signature, files of different sizes
// never reach this comparison, so the sample boundaries always align.
func hashSamples(f *os.File, size int64) (string, error) {
h := sha256.New()
buf := make([]byte, sampleWindow)
for off := int64(0); off < size; off += int64(sampleStride) {
n := min(int64(sampleWindow), size-off)
_, err := f.ReadAt(buf[:n], off)
if err != nil {
return "", err
}
h.Write(buf[:n])
}
return hex.EncodeToString(h.Sum(nil)), nil
} }
+661 -61
View File
@@ -7,6 +7,7 @@ import (
"database/sql" "database/sql"
"encoding/hex" "encoding/hex"
"fmt" "fmt"
"io"
"os" "os"
"path/filepath" "path/filepath"
"runtime" "runtime"
@@ -53,7 +54,23 @@ func pattern(tag byte, n int) []byte {
return data return data
} }
func TestHashHeadTail(t *testing.T) { // sig returns a file's full signature (head, tail, content), failing the
// test on any error.
func sig(t *testing.T, path string, size int64) (string, string, string) {
t.Helper()
head, tail, content, err := hashSignature(path, size)
if err != nil {
t.Fatalf("hashSignature %s: %v", path, err)
}
return head, tail, content
}
// TestHashSignatureBelowThreshold verifies that a file below headTailMin
// is hashed in full and compared directly: head, tail, and content all
// carry the whole-file SHA-256, with no separate end-window step.
func TestHashSignatureBelowThreshold(t *testing.T) {
t.Parallel() t.Parallel()
dir := t.TempDir() dir := t.TempDir()
@@ -62,13 +79,10 @@ func TestHashHeadTail(t *testing.T) {
name string name string
data []byte data []byte
}{ }{
{"empty", nil},
{"one-byte", []byte("x")}, {"one-byte", []byte("x")},
{"under-one-chunk", pattern(1, chunk-1)}, {"one-window", pattern(1, headTailWindow)},
{"exactly-one-chunk", pattern(2, chunk)}, {"several-windows", pattern(2, 3*headTailWindow)},
{"overlapping-reads", pattern(3, chunk+chunk/2)}, {"near-threshold", pattern(3, headTailMin-1)},
{"exactly-two-chunks", pattern(4, 2*chunk)},
{"beyond-two-chunks", pattern(5, 3*chunk)},
} }
for _, c := range cases { for _, c := range cases {
t.Run(c.name, func(t *testing.T) { t.Run(c.name, func(t *testing.T) {
@@ -76,49 +90,619 @@ func TestHashHeadTail(t *testing.T) {
p := writeFile(t, dir, c.name, c.data) p := writeFile(t, dir, c.name, c.data)
head, tail, err := hashHeadTail(p, int64(len(c.data))) head, tail, content := sig(t, p, int64(len(c.data)))
if err != nil {
t.Fatalf("hashHeadTail: %v", err)
}
n := min(chunk, len(c.data)) whole := hexSum(c.data)
if want := hexSum(c.data[:n]); head != want { if head != whole || tail != whole || content != whole {
t.Errorf("head = %s, want %s", head, want) t.Errorf("head=%s tail=%s content=%s, want all whole-file %s",
} head, tail, content, whole)
if want := hexSum(c.data[len(c.data)-n:]); tail != want {
t.Errorf("tail = %s, want %s", tail, want)
} }
}) })
} }
} }
func TestHashHeadTailErrors(t *testing.T) { // TestHashSignatureEnds exercises the head and tail rungs, which apply
// only to files at least headTailMin. Sparse files keep the fixtures
// cheap: a difference in the first window changes only head, a
// difference in the last window changes only tail, and a difference
// between the windows changes neither end hash but does change the
// whole-file content rung (the file is below wholeFileMax).
// hashSignature leaves the content hash of a file this size to the
// content phase, so that rung is checked through a scan.
func TestHashSignatureEnds(t *testing.T) {
t.Parallel() t.Parallel()
dir := t.TempDir() dir := t.TempDir()
_, _, err := hashHeadTail(filepath.Join(dir, "missing"), 1) // Between headTailMin and wholeFileMax: the end-window gate is active
// and the content rung is a whole-file hash.
const size = int64(headTailMin + 2*1024*1024)
base := sparseFile(t, dir, "ends-base", size)
headDiff := sparseFile(t, dir, "ends-head", size)
tailDiff := sparseFile(t, dir, "ends-tail", size)
midDiff := sparseFile(t, dir, "ends-mid", size)
pokeAt(t, headDiff, 0, []byte{1})
pokeAt(t, tailDiff, size-1, []byte{1})
pokeAt(t, midDiff, size/2, []byte{1})
bHead, bTail, bContent := sig(t, base, size)
if bContent != "" {
t.Errorf("content = %q, want none from the hash phase", bContent)
}
h, tl, _ := sig(t, headDiff, size)
if h == bHead {
t.Error("a byte in the first window did not change head")
}
if tl != bTail {
t.Error("a byte in the first window changed tail")
}
h, tl, _ = sig(t, tailDiff, size)
if tl == bTail {
t.Error("a byte in the last window did not change tail")
}
if h != bHead {
t.Error("a byte in the last window changed head")
}
h, tl, _ = sig(t, midDiff, size)
if h != bHead || tl != bTail {
t.Error("a byte between the windows changed an end hash")
}
// base and midDiff match on size, head, and tail, so the scan reads
// both for their content hashes.
c := scanContents(t, dir, base, midDiff)
if c[midDiff] == c[base] {
t.Error("whole-file content rung ignored a byte between the windows")
}
}
func TestHashSignatureErrors(t *testing.T) {
t.Parallel()
dir := t.TempDir()
// A missing file: an error, and every hash left empty.
head, tail, content, err := hashSignature(filepath.Join(dir, "missing"), 1)
if err == nil { if err == nil {
t.Error("no error for a missing file") t.Error("no error for a missing file")
} }
if head != "" || tail != "" || content != "" {
t.Errorf("missing file returned hashes: %q %q %q", head, tail, content)
}
// A zero-length file has constant hashes and is never opened: even // A zero-length file has constant hashes and is never opened: even
// a missing path succeeds. // a missing path succeeds.
head, tail, err := hashHeadTail(filepath.Join(dir, "missing"), 0) head, tail, content, err = hashSignature(filepath.Join(dir, "missing"), 0)
if err != nil || head != emptyHash || tail != emptyHash { if err != nil ||
t.Errorf("empty: head=%q tail=%q err=%v, want constant hashes", head != emptyHash || tail != emptyHash || content != emptyHash {
head, tail, err) t.Errorf("empty: head=%q tail=%q content=%q err=%v, "+
"want constant hashes", head, tail, content, err)
} }
// A file that shrank between the stat and hash passes: reading at // A file that shrank between the stat and hash passes: reading at
// the stat-reported size must fail rather than emit wrong hashes. // the stat-reported size must fail rather than emit wrong hashes.
p := writeFile(t, dir, "shrunk", []byte("tiny")) p := writeFile(t, dir, "shrunk", []byte("tiny"))
_, _, err = hashHeadTail(p, int64(2*chunk)) head, tail, content, err = hashSignature(p, int64(2*headTailWindow))
if err == nil { if err == nil {
t.Error("no error when the stat size exceeds the file size") t.Error("no error when the stat size exceeds the file size")
} }
if head != "" || tail != "" || content != "" {
t.Errorf("shrunk file returned hashes: %q %q %q", head, tail, content)
}
}
// sparseFile creates a file that is logically size bytes long without
// allocating blocks for the hole, so multi-gigabyte cases stay cheap.
func sparseFile(t *testing.T, dir, name string, size int64) string {
t.Helper()
p := filepath.Join(dir, name)
f, err := os.Create(p) //nolint:gosec // test-controlled path
if err != nil {
t.Fatal(err)
}
err = f.Truncate(size)
if err != nil {
t.Fatal(err)
}
err = f.Close()
if err != nil {
t.Fatal(err)
}
return p
}
// pokeAt writes data into an existing file at off, leaving the rest of
// the file (a sparse hole) untouched.
func pokeAt(t *testing.T, path string, off int64, data []byte) {
t.Helper()
f, err := os.OpenFile(path, os.O_WRONLY, 0o600) //nolint:gosec // test path
if err != nil {
t.Fatal(err)
}
_, err = f.WriteAt(data, off)
if err != nil {
t.Fatal(err)
}
err = f.Close()
if err != nil {
t.Fatal(err)
}
}
// scanContents scans dir into a fresh database and returns the content
// hash recorded for each file, by path, failing the test if one of want
// has none. A file of headTailMin or more gets a content hash only when
// it is scanned with a file of the same size, head, and tail.
func scanContents(t *testing.T, dir string,
want ...string,
) map[string]string {
t.Helper()
db := openTestDB(t)
syncTree(t, db, dir)
contents := make(map[string]string)
for _, r := range dbRecords(t, db) {
contents[r.path] = r.content
}
for _, p := range want {
if contents[p] == "" {
t.Fatalf("%s: no content hash", p)
}
}
return contents
}
// TestContentRungBoundary checks the 50 MiB boundary between the two
// content rungs: just below it the whole file is hashed and any byte
// difference shows; at the boundary only the gigabyte-spaced samples are
// hashed, so a difference outside a sample window is invisible. The
// files of each pair match on size, head, and tail, so the scan reads
// both for their content hashes.
func TestContentRungBoundary(t *testing.T) {
t.Parallel()
dir := t.TempDir()
// A byte that lands outside the single [0, sampleWindow) sample a
// sub-gigabyte file has, but well inside the file.
const off = 10 * 1024 * 1024
// Just under the boundary: the whole-file rung sees the poked byte.
under := int64(wholeFileMax - 1)
underBase := sparseFile(t, dir, "under-base", under)
underPoked := sparseFile(t, dir, "under-poked", under)
pokeAt(t, underPoked, off, []byte{1})
// At the boundary: only [0, sampleWindow) is sampled, so the poked
// byte at off is invisible and the two content hashes match.
at := int64(wholeFileMax)
atBase := sparseFile(t, dir, "at-base", at)
atPoked := sparseFile(t, dir, "at-poked", at)
pokeAt(t, atPoked, off, []byte{1})
c := scanContents(t, dir, underBase, underPoked, atBase, atPoked)
if c[underBase] == c[underPoked] {
t.Error("whole-file rung ignored a byte difference below wholeFileMax")
}
if c[atBase] != c[atPoked] {
t.Error("sampled rung saw a byte outside every sample window")
}
}
// TestContentRungMultiGigabyte exercises the sampled rung across several
// gigabytes using sparse files: a difference inside the third sample
// window (at offset 2*sampleStride) changes the hash, while a difference
// in the gap after it does not. The three files match on size, head,
// and tail, so the scan reads each for its content hash.
func TestContentRungMultiGigabyte(t *testing.T) {
t.Parallel()
dir := t.TempDir()
// Three sample windows (offsets 0, 1 GiB, 2 GiB) plus a trailing gap
// that no sample covers.
size := int64(2*sampleStride + 2*sampleWindow)
thirdSample := int64(2 * sampleStride)
gap := thirdSample + int64(sampleWindow)
base := sparseFile(t, dir, "g-base", size)
inSample := sparseFile(t, dir, "g-insample", size)
inGap := sparseFile(t, dir, "g-ingap", size)
pokeAt(t, inSample, thirdSample, []byte{1})
pokeAt(t, inGap, gap, []byte{1})
c := scanContents(t, dir, base, inSample, inGap)
if c[inSample] == c[base] {
t.Error("sample at 2 GiB was not read: difference there was invisible")
}
if c[inGap] != c[base] {
t.Error("a byte in an unsampled gap changed the content hash")
}
}
// sparseFileWithoutMatch writes name in dir as a sparse file of size
// bytes, next to another file of that size whose first byte differs. A
// scan then reads the file's head and tail, since its size is shared,
// but finds no file matching them, so it gets no content hash.
func sparseFileWithoutMatch(t *testing.T, dir, name string,
size int64,
) string {
t.Helper()
p := sparseFile(t, dir, name, size)
other := sparseFile(t, dir, name+"-other-head", size)
pokeAt(t, other, 0, []byte{1})
return p
}
// TestScanContentGate checks that a file of headTailMin or more is read
// for its content hash only when its size, head, and tail match another
// file's: a same-size pair whose heads differ and one whose tails differ
// get no content hash and are not reported, while an identical pair is
// read and reported.
func TestScanContentGate(t *testing.T) {
t.Parallel()
dir := t.TempDir()
db := openTestDB(t)
// Three sizes, so that no pair meets another.
headA := sparseFile(t, dir, "head-a", headTailMin)
headB := sparseFile(t, dir, "head-b", headTailMin)
tailA := sparseFile(t, dir, "tail-a", headTailMin+1)
tailB := sparseFile(t, dir, "tail-b", headTailMin+1)
same := []string{
sparseFile(t, dir, "same-a", headTailMin+2),
sparseFile(t, dir, "same-b", headTailMin+2),
}
pokeAt(t, headB, 0, []byte{1})
pokeAt(t, tailB, headTailMin, []byte{1}) // its last byte
syncTree(t, db, dir)
recs := dbRecords(t, db)
for _, p := range []string{headA, headB, tailA, tailB} {
r := recordByPath(t, recs, p)
if r.head == "" || r.tail == "" || r.content != "" {
t.Errorf("%s: head = %q tail = %q content = %q, "+
"want head and tail only", p, r.head, r.tail, r.content)
}
}
groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, same) {
t.Fatalf("groups = %+v, want only the identical pair %q",
groups, same)
}
}
// TestScanContentAcrossOperands checks that a stored file gets its
// content hash when a later scan of a separate operand brings its
// match: tree A's file has a head and tail but no content hash until
// tree B, holding an identical file, is scanned.
func TestScanContentAcrossOperands(t *testing.T) {
t.Parallel()
db := openTestDB(t)
a := sparseFileWithoutMatch(t, t.TempDir(), "a", headTailMin)
syncTree(t, db, filepath.Dir(a))
if r := recordByPath(t, dbRecords(t, db), a); r.head == "" || r.content != "" {
t.Fatalf("after scanning A: %+v, want head and tail only", r)
}
b := sparseFile(t, t.TempDir(), "b", headTailMin)
syncTree(t, db, filepath.Dir(b))
recs := dbRecords(t, db)
if r := recordByPath(t, recs, a); r.content == "" {
t.Fatalf("after scanning B: %+v, want A's file content-hashed", r)
}
want := []string{a, b}
slices.Sort(want)
groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
t.Fatalf("groups = %+v, want the pair %q", groups, want)
}
}
// TestScanContentWithinOperand checks that a rescan adding a match next
// to an unchanged stored file gives the stored file its content hash,
// though the hash phase leaves it alone as unchanged.
func TestScanContentWithinOperand(t *testing.T) {
t.Parallel()
dir := t.TempDir()
db := openTestDB(t)
stored := sparseFileWithoutMatch(t, dir, "d1", headTailMin)
syncTree(t, db, dir)
added := sparseFile(t, dir, "d2", headTailMin)
st := syncTree(t, db, dir)
if st != (scanStats{added: 1, unchanged: 2}) {
t.Fatalf("rescan stats = %+v, want 1 added 2 unchanged", st)
}
want := []string{stored, added}
groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
t.Fatalf("groups = %+v, want the pair %q", groups, want)
}
}
// TestScanContentStalePartners checks that a stored file outside the
// operand that has vanished, or changed, since it was recorded is not
// read, and that its match inside the operand is not read either: the
// match has no other partner left, so neither gets a content hash and
// no duplicate is reported.
func TestScanContentStalePartners(t *testing.T) {
t.Parallel()
db := openTestDB(t)
dirA := t.TempDir()
gone := sparseFileWithoutMatch(t, dirA, "gone", headTailMin)
changed := sparseFileWithoutMatch(t, dirA, "changed", headTailMin+1)
syncTree(t, db, dirA)
before := dbRecords(t, db)
err := os.Remove(gone)
if err != nil {
t.Fatal(err)
}
future := time.Now().Add(time.Hour)
err = os.Chtimes(changed, future, future)
if err != nil {
t.Fatal(err)
}
dirB := t.TempDir()
sparseFile(t, dirB, "gone-copy", headTailMin)
sparseFile(t, dirB, "changed-copy", headTailMin+1)
st := syncTree(t, db, dirB)
if st != (scanStats{added: 2}) {
t.Errorf("stats = %+v, want 2 added and nothing skipped", st)
}
recs := dbRecords(t, db)
for _, r := range recs {
if r.content != "" {
t.Errorf("%s: content = %q, want none: its only match is stale",
r.path, r.content)
}
}
for _, old := range before {
if r := recordByPath(t, recs, old.path); r != old {
t.Errorf("record = %+v, want it left as %+v", r, old)
}
}
if groups := dupeGroups(t, db); len(groups) != 0 {
t.Errorf("groups = %+v, want none", groups)
}
}
// TestScanContentHashedStalePartners checks that stored matches outside
// the operand that already have a content hash are checked like any
// other: once one has vanished and the other has changed, a copy of
// them scanned in another tree has no match left, so it is not read and
// is not reported as their duplicate.
func TestScanContentHashedStalePartners(t *testing.T) {
t.Parallel()
db := openTestDB(t)
dirA := t.TempDir()
stored := []string{
sparseFile(t, dirA, "changed", headTailMin),
sparseFile(t, dirA, "gone", headTailMin),
}
// The two stored files match, so this scan gives both a content
// hash.
syncTree(t, db, dirA)
err := os.Remove(stored[1])
if err != nil {
t.Fatal(err)
}
future := time.Now().Add(time.Hour)
err = os.Chtimes(stored[0], future, future)
if err != nil {
t.Fatal(err)
}
b := sparseFile(t, t.TempDir(), "copy", headTailMin)
st := syncTree(t, db, filepath.Dir(b))
if st != (scanStats{added: 1}) {
t.Errorf("stats = %+v, want 1 added and nothing skipped", st)
}
recs := dbRecords(t, db)
if r := recordByPath(t, recs, b); r.content != "" {
t.Errorf("copy: content = %q, want none: its only matches are stale",
r.content)
}
// The stored records lie outside the operand and are left as they
// are, so they still group with each other, but not with the copy.
groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, stored) {
t.Errorf("groups = %+v, want only the stored pair %q", groups, stored)
}
}
// TestScanContentReadFailure checks that a failed content read is
// counted as skipped and leaves the record without a content hash, and
// that a later scan tries the read again.
func TestScanContentReadFailure(t *testing.T) {
t.Parallel()
db := openTestDB(t)
a := sparseFileWithoutMatch(t, t.TempDir(), "a", headTailMin)
syncTree(t, db, filepath.Dir(a))
// lstat still works on the unreadable file, so it passes the check
// and fails only when it is read.
err := os.Chmod(a, 0)
if err != nil {
t.Fatal(err)
}
dirB := t.TempDir()
b := sparseFile(t, dirB, "b", headTailMin)
st := syncTree(t, db, dirB)
if st != (scanStats{added: 1, skipped: 1}) {
t.Fatalf("stats = %+v, want 1 added 1 skipped", st)
}
if r := recordByPath(t, dbRecords(t, db), a); r.content != "" {
t.Fatalf("unreadable file: %+v, want no content hash", r)
}
err = os.Chmod(a, 0o600)
if err != nil {
t.Fatal(err)
}
st = syncTree(t, db, dirB)
if st != (scanStats{unchanged: 1}) {
t.Fatalf("rescan stats = %+v, want 1 unchanged", st)
}
want := []string{a, b}
slices.Sort(want)
groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
t.Fatalf("groups = %+v, want the pair %q after the retry",
groups, want)
}
}
// TestScanContentCheckError checks that a stored file the content phase
// cannot lstat, for a reason other than its being gone, is counted as
// skipped and does not count as a match.
func TestScanContentCheckError(t *testing.T) {
t.Parallel()
db := openTestDB(t)
sub := filepath.Join(t.TempDir(), "sub")
err := os.Mkdir(sub, 0o700)
if err != nil {
t.Fatal(err)
}
sparseFileWithoutMatch(t, sub, "a", headTailMin)
syncTree(t, db, sub)
// Without search permission on its directory, the stored file's
// lstat fails with permission denied.
err = os.Chmod(sub, 0)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() {
//nolint:gosec // removing the directory needs its search bit back
_ = os.Chmod(sub, 0o700)
})
b := sparseFile(t, t.TempDir(), "b", headTailMin)
st := syncTree(t, db, filepath.Dir(b))
if st != (scanStats{added: 1, skipped: 1}) {
t.Fatalf("stats = %+v, want 1 added 1 skipped", st)
}
if r := recordByPath(t, dbRecords(t, db), b); r.content != "" {
t.Errorf("b: content = %q, want none: its only match could not be "+
"checked", r.content)
}
}
// TestScanContentHardlinks checks that the content phase stores the
// content hash of a hard-linked file on every one of its links.
func TestScanContentHardlinks(t *testing.T) {
t.Parallel()
dir := t.TempDir()
db := openTestDB(t)
a := sparseFile(t, dir, "a", headTailMin)
b := filepath.Join(dir, "b")
err := os.Link(a, b)
if err != nil {
t.Fatal(err)
}
c := sparseFile(t, dir, "copy", headTailMin)
st := syncTree(t, db, dir)
if st != (scanStats{added: 3}) {
t.Fatalf("stats = %+v, want 3 added", st)
}
recs := dbRecords(t, db)
want := recordByPath(t, recs, c).content
if want == "" {
t.Fatal("the copy has no content hash")
}
for _, p := range []string{a, b} {
if got := recordByPath(t, recs, p).content; got != want {
t.Errorf("%s: content = %q, want %q", p, got, want)
}
}
} }
// collectWalk runs a walk over roots and returns the emitted records // collectWalk runs a walk over roots and returns the emitted records
@@ -273,9 +857,11 @@ func TestWalkFileAndSymlinkOperands(t *testing.T) {
t.Fatalf("file operand: recs = %+v, errs = %d", recs, errs) t.Fatalf("file operand: recs = %+v, errs = %d", recs, errs)
} }
// A symlink operand is not followed and yields nothing. // A symlink operand that reaches the walk (it became one after
// walkableRoots checked it) is not followed: it yields a warning and
// no records.
recs, errs = collectWalk(t, []string{link}, false, 2) recs, errs = collectWalk(t, []string{link}, false, 2)
if errs != 0 || len(recs) != 0 { if errs != 1 || len(recs) != 0 {
t.Fatalf("symlink operand: recs = %+v, errs = %d", recs, errs) t.Fatalf("symlink operand: recs = %+v, errs = %d", recs, errs)
} }
} }
@@ -370,11 +956,16 @@ func syncTree(t *testing.T, db *sql.DB, roots ...string) scanStats {
return st return st
} }
// dbRecords returns every record currently in the database. // dbRecords returns every record currently in the database, in path
// order.
func dbRecords(t *testing.T, db *sql.DB) []scanRec { func dbRecords(t *testing.T, db *sql.DB) []scanRec {
t.Helper() t.Helper()
recs, err := loadFileRows(t.Context(), db) var recs []scanRec
err := loadFileRows(t.Context(), db, func(r scanRec) {
recs = append(recs, r)
})
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -411,10 +1002,10 @@ func recordPaths(recs []scanRec) []string {
// assertSmokeDupeGroups checks the file-level duplicate groups for the // assertSmokeDupeGroups checks the file-level duplicate groups for the
// smoke tree rooted at dir. // smoke tree rooted at dir.
func assertSmokeDupeGroups(t *testing.T, dir string, parsed []scanRec) { func assertSmokeDupeGroups(t *testing.T, dir string, db *sql.DB) {
t.Helper() t.Helper()
groups := collectDupeGroups(parsed) groups := dupeGroups(t, db)
if len(groups) != 5 { if len(groups) != 5 {
t.Fatalf("len(groups) = %d, want 5", len(groups)) t.Fatalf("len(groups) = %d, want 5", len(groups))
} }
@@ -439,11 +1030,10 @@ func assertSmokeDupeGroups(t *testing.T, dir string, parsed []scanRec) {
// assertSmokeTreeGroups checks the duplicate-tree groups for the smoke // assertSmokeTreeGroups checks the duplicate-tree groups for the smoke
// tree rooted at dir. // tree rooted at dir.
func assertSmokeTreeGroups(t *testing.T, dir string, parsed []scanRec) { func assertSmokeTreeGroups(t *testing.T, dir string, db *sql.DB) {
t.Helper() t.Helper()
super, dirs := buildHierarchy(parsed) super, dirs := dbTree(t, db)
super.compute()
tg := collectTreeGroups(dirs, super) tg := collectTreeGroups(dirs, super)
if len(tg) != 1 { if len(tg) != 1 {
@@ -477,8 +1067,8 @@ func TestScanPipeline(t *testing.T) {
t.Fatalf("len(records) = %d, want %d", len(parsed), smokeTreeFiles) t.Fatalf("len(records) = %d, want %d", len(parsed), smokeTreeFiles)
} }
assertSmokeDupeGroups(t, dir, parsed) assertSmokeDupeGroups(t, dir, db)
assertSmokeTreeGroups(t, dir, parsed) assertSmokeTreeGroups(t, dir, db)
} }
func TestSyncScanUnchangedReuse(t *testing.T) { func TestSyncScanUnchangedReuse(t *testing.T) {
@@ -706,13 +1296,13 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
recs := dbRecords(t, db) recs := dbRecords(t, db)
for _, r := range recs { for _, r := range recs {
if r.head != "" || r.tail != "" { if r.head != "" || r.tail != "" || r.content != "" {
t.Errorf("%s: head = %q tail = %q, want unhashed", t.Errorf("%s: head = %q tail = %q content = %q, want unhashed",
r.path, r.head, r.tail) r.path, r.head, r.tail, r.content)
} }
} }
if groups := collectDupeGroups(recs); len(groups) != 0 { if groups := dupeGroups(t, db); len(groups) != 0 {
t.Fatalf("groups = %+v, want none from unhashed records", groups) t.Fatalf("groups = %+v, want none from unhashed records", groups)
} }
@@ -727,7 +1317,7 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
st) st)
} }
groups := collectDupeGroups(dbRecords(t, db)) groups := dupeGroups(t, db)
if len(groups) != 1 { if len(groups) != 1 {
t.Fatalf("groups = %+v, want the a/c pair", groups) t.Fatalf("groups = %+v, want the a/c pair", groups)
} }
@@ -740,23 +1330,35 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
func TestTreesUnhashedNeverEqual(t *testing.T) { func TestTreesUnhashedNeverEqual(t *testing.T) {
t.Parallel() t.Parallel()
// Two trees identical except for unhashed same-name, same-size // Two trees identical except for same-name, same-size files without
// files (possible when the trees were scanned separately) must not // a content hash must not compare equal: their content is unknown.
// compare equal: unhashed content is unknown. // That holds for unhashed files (possible when the trees were
shared := pattern(1, 100) // scanned separately) and for files of headTailMin or more that
recs := []scanRec{ // have only a head and tail.
{path: "/x/t1/f1", size: 100, head: hexSum(shared), tail: hexSum(shared)}, sum := hexSum(pattern(1, 100))
{path: "/x/t2/f1", size: 100, head: hexSum(shared), tail: hexSum(shared)}, shared := []scanRec{
{path: "/x/t1/u", size: 50}, {path: "/x/t1/f1", size: 100, head: sum, tail: sum, content: sum},
{path: "/x/t2/u", size: 50}, {path: "/x/t2/f1", size: 100, head: sum, tail: sum, content: sum},
} }
super, dirs := buildHierarchy(recs) cases := map[string][]scanRec{
super.compute() "unhashed": {
{path: "/x/t1/u", size: 50},
{path: "/x/t2/u", size: 50},
},
"head and tail only": {
{path: "/x/t1/u", size: headTailMin, head: "h", tail: "t"},
{path: "/x/t2/u", size: headTailMin, head: "h", tail: "t"},
},
}
if tg := collectTreeGroups(dirs, super); len(tg) != 0 { for name, unknown := range cases {
t.Fatalf("tree groups = %d, want 0 (unhashed files differ)", super, dirs := treeOf(t, append(slices.Clone(shared), unknown...))
len(tg))
if tg := collectTreeGroups(dirs, super); len(tg) != 0 {
t.Errorf("%s: tree groups = %d, want 0 (the files may differ)",
name, len(tg))
}
} }
} }
@@ -788,7 +1390,7 @@ func TestScanHardlinksReadOnce(t *testing.T) {
t.Fatalf("hardlink hashes differ: %+v vs %+v", ra, rb) t.Fatalf("hardlink hashes differ: %+v vs %+v", ra, rb)
} }
if groups := collectDupeGroups(recs); len(groups) != 1 { if groups := dupeGroups(t, db); len(groups) != 1 {
t.Fatalf("groups = %+v, want the hardlink pair", groups) t.Fatalf("groups = %+v, want the hardlink pair", groups)
} }
} }
@@ -952,7 +1554,7 @@ func TestScanHashWriteFailureUnwindsPool(t *testing.T) {
code := run([]string{ code := run([]string{
cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir, cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir,
}, &stderr) }, io.Discard, &stderr)
if code != exitFatal { if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d; stderr: %s", t.Fatalf("run(scan) = %d, want %d; stderr: %s",
code, exitFatal, stderr.String()) code, exitFatal, stderr.String())
@@ -1029,10 +1631,8 @@ func TestReportsNeverTouchFilesystem(t *testing.T) {
t.Fatal(err) t.Fatal(err)
} }
recs := dbRecords(t, db) assertSmokeDupeGroups(t, dir, db)
assertSmokeTreeGroups(t, dir, db)
assertSmokeDupeGroups(t, dir, recs)
assertSmokeTreeGroups(t, dir, recs)
} }
func TestUnderRoot(t *testing.T) { func TestUnderRoot(t *testing.T) {
+4 -4
View File
@@ -16,10 +16,10 @@
# implies the repo is green. # implies the repo is green.
# #
# That implication holds only because of CHECK_EPOCH. A COPY layer is # That implication holds only because of CHECK_EPOCH. A COPY layer is
# invalidated by changed content, and a merge commit's tree is # invalidated only by changed content, and a rebuild of an unchanged
# byte-identical to the branch head it merges, so without a fresh value # checkout sends the same content, so without a fresh value here Docker
# here Docker serves the gate layers from cache and the build reports a # serves the gate layers from cache and the build reports a green it
# green it never earned. Passing the current epoch invalidates the gate # never earned. Passing the current epoch invalidates the gate
# layers on every run while leaving the pinned base images and # layers on every run while leaving the pinned base images and
# go mod download cached; see the Dockerfile for the placement. # go mod download cached; see the Dockerfile for the placement.
set -eu set -eu
+156 -95
View File
@@ -5,48 +5,59 @@ import (
"context" "context"
"crypto/sha256" "crypto/sha256"
"fmt" "fmt"
"io"
"os" "os"
"slices" "slices"
"strconv" "strconv"
"strings" "strings"
) )
// fileSig is a file's duplicate signature; mtime is excluded. // treeNode is one directory reconstructed from the record paths.
type fileSig struct {
size int64
head string
tail string
}
// treeNode is one directory reconstructed from the scan stream.
type treeNode struct { type treeNode struct {
path string path string
parent *treeNode parent *treeNode
dirs map[string]*treeNode // entries holds the serialized child entries until the digest is
files map[string]fileSig // computed from them, and is then dropped.
entries []string
digest [sha256.Size]byte digest [sha256.Size]byte
fileCount int64 fileCount int64
totalSize int64 totalSize int64
} }
// runTrees implements the trees subcommand: it reads every record from // runTrees implements the trees subcommand: it reads every record from
// the database, reconstructs the directory hierarchy from the record // the database in path order, reconstructs the directory hierarchy from
// paths, computes a Merkle-style digest per directory, and prints // the record paths, computes a Merkle-style digest per directory, and
// maximal duplicate-tree groups as TSV on stdout. It never touches the // prints maximal duplicate-tree groups as TSV on stdout. It never
// scanned filesystem; its only I/O is the database, stdout, and // touches the scanned filesystem; its only I/O is the database, stdout,
// stderr. // and stderr. Any database problem, including a missing database, is
func runTrees(ctx context.Context) error { // fatal.
recs, err := loadRecords(ctx) func runTrees(ctx context.Context, stdout io.Writer) error {
dbPath := databasePath()
db, err := openReportDatabase(ctx, dbPath)
if err != nil { if err != nil {
return err return err
} }
super, allDirs := buildHierarchy(recs) defer func() { _ = db.Close() }()
super.compute()
records := 0
tree := newTreeBuilder()
err = loadFileRows(ctx, db, func(r scanRec) {
records++
tree.add(r)
})
if err != nil {
return fmt.Errorf("database %s: %w", dbPath, err)
}
super, allDirs := tree.finish()
dupes := collectTreeGroups(allDirs, super) dupes := collectTreeGroups(allDirs, super)
out := bufio.NewWriterSize(os.Stdout, ioBufSize) out := bufio.NewWriterSize(stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize") _, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
if err != nil { if err != nil {
@@ -61,7 +72,8 @@ func runTrees(ctx context.Context) error {
first := g[0] first := g[0]
for _, n := range g[1:] { for _, n := range g[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n", _, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n",
first.path, n.path, first.fileCount, first.totalSize) escapePath(first.path), escapePath(n.path),
first.fileCount, first.totalSize)
if err != nil { if err != nil {
return fmt.Errorf("write stdout: %w", err) return fmt.Errorf("write stdout: %w", err)
} }
@@ -79,63 +91,128 @@ func runTrees(ctx context.Context) error {
fmt.Fprintf(os.Stderr, fmt.Fprintf(os.Stderr,
"trees: %d records read, %d duplicate tree groups, %d dupe trees, "+ "trees: %d records read, %d duplicate tree groups, %d dupe trees, "+
"%s reclaimable\n", "%s reclaimable\n",
len(recs), len(dupes), dupeTrees, humanBytes(reclaimable)) records, len(dupes), dupeTrees, humanBytes(reclaimable))
return nil return nil
} }
// buildHierarchy reconstructs the directory hierarchy from the record // treeBuilder reconstructs the directory hierarchy from records added
// paths under a synthetic super-root. Paths are split on "/"; for // in path order, under a synthetic super-root. Paths are split on "/";
// absolute paths the first component is empty, which simply becomes a // for absolute paths the first component is empty, which becomes the
// top-level node representing "/". It returns the super-root and every // top-level directory with path "/". In path order all the paths under
// directory node created. // one directory come together, so a directory is complete once a path
func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) { // outside it is added: its digest is computed then and its entries are
// dropped. Only the directories holding the latest path keep entries.
type treeBuilder struct {
super *treeNode
// open lists the directories holding the latest path, outermost
// first, starting with the super-root; names[i] is open[i]'s name.
open []*treeNode
names []string
// dirs lists every completed directory.
dirs []*treeNode
}
func newTreeBuilder() *treeBuilder {
super := &treeNode{} super := &treeNode{}
var allDirs []*treeNode return &treeBuilder{
super: super,
open: []*treeNode{super},
names: []string{""},
}
}
for _, r := range recs { // add adds one record. Each record must come after the previous one in
comps := strings.Split(r.path, "/") // path order (byte order); otherwise a completed directory would be
// started again as a second directory with the same path.
func (b *treeBuilder) add(r scanRec) {
comps := strings.Split(r.path, "/")
dirNames, name := comps[:len(comps)-1], comps[len(comps)-1]
node := super // Keep the open directories that hold this path; complete the rest.
for _, c := range comps[:len(comps)-1] { depth := 1
child := node.dirs[c] for depth < len(b.open) && depth <= len(dirNames) &&
if child == nil { b.names[depth] == dirNames[depth-1] {
childPath := c depth++
if node != super {
childPath = node.path + "/" + c
}
child = &treeNode{path: childPath, parent: node}
if node.dirs == nil {
node.dirs = make(map[string]*treeNode)
}
node.dirs[c] = child
allDirs = append(allDirs, child)
}
node = child
}
if node.files == nil {
node.files = make(map[string]fileSig)
}
sig := fileSig{size: r.size, head: r.head, tail: r.tail}
// An unhashed record (its size was unique when last scanned)
// has unknown content: give it a signature no other file can
// share, so trees containing it never compare equal. Real
// heads are hex, so the NUL-prefixed form cannot collide.
if sig.head == "" {
sig.head = "unhashed\x00" + r.path
}
node.files[comps[len(comps)-1]] = sig
} }
return super, allDirs b.closeTo(depth)
for _, c := range dirNames[depth-1:] {
b.openDir(c)
}
dir := b.open[len(b.open)-1]
dir.entries = append(dir.entries, fileEntry(name, r))
dir.fileCount++
dir.totalSize += r.size
}
// openDir starts the directory called name inside the innermost open
// one.
func (b *treeBuilder) openDir(name string) {
parent := b.open[len(b.open)-1]
path := parent.path + "/" + name
// The root directory's path is "/", not empty, and its children's
// paths start with one slash, not two.
switch {
case parent == b.super && name == "":
path = "/"
case parent == b.super:
path = name
case parent.path == "/":
path = "/" + name
}
b.open = append(b.open, &treeNode{path: path, parent: parent})
b.names = append(b.names, name)
}
// closeTo completes the open directories after the first n, innermost
// first: each one's digest is computed and entered in its parent along
// with its totals.
func (b *treeBuilder) closeTo(n int) {
for len(b.open) > n {
last := len(b.open) - 1
dir, name := b.open[last], b.names[last]
b.open, b.names = b.open[:last], b.names[:last]
dir.computeDigest()
dir.parent.entries = append(dir.parent.entries,
"d\x00"+name+"\x00"+string(dir.digest[:]))
dir.parent.fileCount += dir.fileCount
dir.parent.totalSize += dir.totalSize
b.dirs = append(b.dirs, dir)
}
}
// finish completes every open directory and returns the super-root and
// every directory.
func (b *treeBuilder) finish() (*treeNode, []*treeNode) {
b.closeTo(1)
return b.super, b.dirs
}
// fileEntry serializes a file child for its directory's digest: its
// name and its signature (size, head, tail, content); mtime is
// excluded.
func fileEntry(name string, r scanRec) string {
content := r.content
// A record without a content hash has unknown content (README
// "Database"): give it a signature no other file can share, so
// trees containing it never compare equal. Real hashes are hex, so
// the NUL-prefixed form cannot collide.
if content == "" {
content = "unhashed\x00" + r.path
}
return "f\x00" + name + "\x00" + strconv.FormatInt(r.size, 10) +
"\x00" + r.head + "\x00" + r.tail + "\x00" + content
} }
// collectTreeGroups groups directories by digest and returns every // collectTreeGroups groups directories by digest and returns every
@@ -177,38 +254,22 @@ func collectTreeGroups(allDirs []*treeNode, super *treeNode) [][]*treeNode {
return dupes return dupes
} }
// compute fills in digest, fileCount, and totalSize for n and all of // computeDigest sets n's digest and drops its entries. A directory's
// its descendants. A directory's digest is the SHA-256 of its child // digest is the SHA-256 of its child entries — files serialized with
// entries — files serialized with name and signature, subdirectories // name and signature, subdirectories with name and recursive digest —
// with name and recursive digest — sorted byte-lexicographically. // sorted byte-lexicographically. Filenames cannot contain NUL or "/",
// Filenames cannot contain NUL or "/", so NUL delimiters are // so NUL delimiters are unambiguous.
// unambiguous. func (n *treeNode) computeDigest() {
func (n *treeNode) compute() { slices.Sort(n.entries)
entries := make([]string, 0, len(n.dirs)+len(n.files))
for name, sig := range n.files {
entries = append(entries,
"f\x00"+name+"\x00"+strconv.FormatInt(sig.size, 10)+
"\x00"+sig.head+"\x00"+sig.tail)
n.fileCount++
n.totalSize += sig.size
}
for name, child := range n.dirs {
child.compute()
entries = append(entries, "d\x00"+name+"\x00"+string(child.digest[:]))
n.fileCount += child.fileCount
n.totalSize += child.totalSize
}
slices.Sort(entries)
h := sha256.New() h := sha256.New()
for _, e := range entries { for _, e := range n.entries {
h.Write([]byte(e)) h.Write([]byte(e))
h.Write([]byte{0}) h.Write([]byte{0})
} }
copy(n.digest[:], h.Sum(nil)) copy(n.digest[:], h.Sum(nil))
n.entries = nil
} }
// suppressed reports whether a duplicate-tree group is non-maximal: its // suppressed reports whether a duplicate-tree group is non-maximal: its
+149 -34
View File
@@ -1,31 +1,66 @@
package main package main
import ( import (
"bytes"
"database/sql"
"slices" "slices"
"testing" "testing"
) )
// Signature hashes shared by the smoke-test records. // Signature hashes shared by the smoke-test records.
const ( const (
f1Head = "f1h" f1Head = "f1h"
f1Tail = "f1t" f1Tail = "f1t"
f2Head = "f2h" f1Content = "f1c"
f2Tail = "f2t" f2Head = "f2h"
f2Tail = "f2t"
f2Content = "f2c"
) )
// smokeTreeRecs mirrors the README smoke-test tree layout: /d/t1 and // smokeTreeRecs mirrors the README smoke-test tree layout: /d/t1 and
// /d/t2 are identical, /d/t3 differs from them only by one filename. // /d/t2 are identical, /d/t3 differs from them only by one filename.
func smokeTreeRecs() []scanRec { func smokeTreeRecs() []scanRec {
return []scanRec{ return []scanRec{
{size: 3000, head: f1Head, tail: f1Tail, path: "/d/t1/f1"}, {size: 3000, head: f1Head, tail: f1Tail, content: f1Content, path: "/d/t1/f1"},
{size: 100, head: f2Head, tail: f2Tail, path: "/d/t1/sub/f2"}, {size: 100, head: f2Head, tail: f2Tail, content: f2Content, path: "/d/t1/sub/f2"},
{size: 3000, head: f1Head, tail: f1Tail, path: "/d/t2/f1"}, {size: 3000, head: f1Head, tail: f1Tail, content: f1Content, path: "/d/t2/f1"},
{size: 100, head: f2Head, tail: f2Tail, path: "/d/t2/sub/f2"}, {size: 100, head: f2Head, tail: f2Tail, content: f2Content, path: "/d/t2/sub/f2"},
{size: 3000, head: f1Head, tail: f1Tail, path: "/d/t3/f1"}, {size: 3000, head: f1Head, tail: f1Tail, content: f1Content, path: "/d/t3/f1"},
{size: 100, head: f2Head, tail: f2Tail, path: "/d/t3/sub/f2renamed"}, {size: 100, head: f2Head, tail: f2Tail, content: f2Content,
path: "/d/t3/sub/f2renamed"},
} }
} }
// dbTree builds the directory hierarchy from the records in db the way
// trees does, and returns the super-root and every directory.
func dbTree(t *testing.T, db *sql.DB) (*treeNode, []*treeNode) {
t.Helper()
tree := newTreeBuilder()
err := loadFileRows(t.Context(), db, tree.add)
if err != nil {
t.Fatal(err)
}
return tree.finish()
}
// treeOf writes recs into a fresh database and builds the directory
// hierarchy from it the way trees does.
func treeOf(t *testing.T, recs []scanRec) (*treeNode, []*treeNode) {
t.Helper()
db := openTestDB(t)
err := applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
return dbTree(t, db)
}
// nodeByPath finds the directory node with the given path. // nodeByPath finds the directory node with the given path.
func nodeByPath(t *testing.T, dirs []*treeNode, path string) *treeNode { func nodeByPath(t *testing.T, dirs []*treeNode, path string) *treeNode {
t.Helper() t.Helper()
@@ -56,11 +91,10 @@ func groupPaths(groups [][]*treeNode) [][]string {
return out return out
} }
func TestBuildHierarchyCounts(t *testing.T) { func TestTreeCounts(t *testing.T) {
t.Parallel() t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs()) _, dirs := treeOf(t, smokeTreeRecs())
super.compute()
d := nodeByPath(t, dirs, "/d") d := nodeByPath(t, dirs, "/d")
if d.fileCount != 6 || d.totalSize != 9300 { if d.fileCount != 6 || d.totalSize != 9300 {
@@ -81,11 +115,98 @@ func TestBuildHierarchyCounts(t *testing.T) {
} }
} }
func TestTreeRootPath(t *testing.T) {
t.Parallel()
// The root directory's path is "/", never empty, and its
// children's paths start with a single slash.
_, dirs := treeOf(t, []scanRec{{path: "/f"}, {path: "/srv/g"}})
got := make([]string, 0, len(dirs))
for _, d := range dirs {
got = append(got, d.path)
}
slices.Sort(got)
want := []string{"/", "/srv"}
if !slices.Equal(got, want) {
t.Fatalf("directory paths = %q, want %q", got, want)
}
}
func TestTreeNamesSortingBeforeSlash(t *testing.T) {
t.Parallel()
// In path order "/a/b-x/f" and "/a/b.txt" come between the file
// "/a/b" and "/a/b/f", because "-" and "." sort before "/". Each
// directory must still be built once, whole, so /a matches /c.
recs := make([]scanRec, 0, 8)
for _, top := range []string{"/a", "/c"} {
for _, p := range []string{"/b", "/b-x/f", "/b.txt", "/b/f"} {
content := "c"
if p == "/b-x/f" {
content = "other"
}
recs = append(recs, scanRec{
size: 1, head: "h", tail: "t", content: content, path: top + p,
})
}
}
super, dirs := treeOf(t, recs)
got := make([]string, 0, len(dirs))
for _, d := range dirs {
got = append(got, d.path)
}
slices.Sort(got)
want := []string{"/", "/a", "/a/b", "/a/b-x", "/c", "/c/b", "/c/b-x"}
if !slices.Equal(got, want) {
t.Fatalf("directory paths = %q, want %q", got, want)
}
groups := collectTreeGroups(dirs, super)
gotGroups := groupPaths(groups)
wantGroups := [][]string{{"/a", "/c"}}
if !slices.EqualFunc(gotGroups, wantGroups, slices.Equal) {
t.Fatalf("groups = %v, want %v", gotGroups, wantGroups)
}
if groups[0][0].fileCount != 4 || groups[0][0].totalSize != 4 {
t.Errorf("group totals: %d files %d bytes, want 4 4",
groups[0][0].fileCount, groups[0][0].totalSize)
}
}
func TestRunTreesEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
code := run([]string{cmdTrees}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tfiles\tsize\n" +
`/d/\tone\ntwo\rthree\\four` + "\t/d/A\t1\t5\n"
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
func TestTreeDigests(t *testing.T) { func TestTreeDigests(t *testing.T) {
t.Parallel() t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs()) _, dirs := treeOf(t, smokeTreeRecs())
super.compute()
t1 := nodeByPath(t, dirs, "/d/t1") t1 := nodeByPath(t, dirs, "/d/t1")
t2 := nodeByPath(t, dirs, "/d/t2") t2 := nodeByPath(t, dirs, "/d/t2")
@@ -114,12 +235,11 @@ func TestTreeDigestContentSensitivity(t *testing.T) {
const sharedTail = "same" const sharedTail = "same"
recs := []scanRec{ recs := []scanRec{
{size: 10, head: sharedTail, tail: sharedTail, path: "/r/a/f"}, {size: 10, head: sharedTail, tail: sharedTail, content: "c", path: "/r/a/f"},
{size: 10, head: "DIFF", tail: sharedTail, path: "/r/b/f"}, {size: 10, head: "DIFF", tail: sharedTail, content: "c", path: "/r/b/f"},
} }
super, dirs := buildHierarchy(recs) _, dirs := treeOf(t, recs)
super.compute()
a := nodeByPath(t, dirs, "/r/a") a := nodeByPath(t, dirs, "/r/a")
b := nodeByPath(t, dirs, "/r/b") b := nodeByPath(t, dirs, "/r/b")
@@ -132,8 +252,7 @@ func TestTreeDigestContentSensitivity(t *testing.T) {
func TestCollectTreeGroupsMaximal(t *testing.T) { func TestCollectTreeGroupsMaximal(t *testing.T) {
t.Parallel() t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs()) super, dirs := treeOf(t, smokeTreeRecs())
super.compute()
groups := collectTreeGroups(dirs, super) groups := collectTreeGroups(dirs, super)
@@ -157,16 +276,14 @@ func TestCollectTreeGroupsDeterministic(t *testing.T) {
recs := smokeTreeRecs() recs := smokeTreeRecs()
super, dirs := buildHierarchy(recs) super, dirs := treeOf(t, recs)
super.compute()
forward := groupPaths(collectTreeGroups(dirs, super)) forward := groupPaths(collectTreeGroups(dirs, super))
reversed := slices.Clone(recs) reversed := slices.Clone(recs)
slices.Reverse(reversed) slices.Reverse(reversed)
superR, dirsR := buildHierarchy(reversed) superR, dirsR := treeOf(t, reversed)
superR.compute()
backward := groupPaths(collectTreeGroups(dirsR, superR)) backward := groupPaths(collectTreeGroups(dirsR, superR))
if !slices.EqualFunc(forward, backward, slices.Equal) { if !slices.EqualFunc(forward, backward, slices.Equal) {
@@ -181,12 +298,11 @@ func TestCollectTreeGroupsSiblings(t *testing.T) {
// Identical sibling dirs share a parent, so their group cannot be // Identical sibling dirs share a parent, so their group cannot be
// implied by a parent group and must be reported. // implied by a parent group and must be reported.
recs := []scanRec{ recs := []scanRec{
{size: 10, head: "h", tail: "t", path: "/p/x1/f"}, {size: 10, head: "h", tail: "t", content: "c", path: "/p/x1/f"},
{size: 10, head: "h", tail: "t", path: "/p/x2/f"}, {size: 10, head: "h", tail: "t", content: "c", path: "/p/x2/f"},
} }
super, dirs := buildHierarchy(recs) super, dirs := treeOf(t, recs)
super.compute()
got := groupPaths(collectTreeGroups(dirs, super)) got := groupPaths(collectTreeGroups(dirs, super))
@@ -203,13 +319,12 @@ func TestCollectTreeGroupsDifferingParents(t *testing.T) {
// extra file, so the parents' digests differ and the x group must // extra file, so the parents' digests differ and the x group must
// be reported. // be reported.
recs := []scanRec{ recs := []scanRec{
{size: 10, head: "h", tail: "t", path: "/p/a/x/f"}, {size: 10, head: "h", tail: "t", content: "c", path: "/p/a/x/f"},
{size: 99, head: "e", tail: "e", path: "/p/a/extra"}, {size: 99, head: "e", tail: "e", content: "e", path: "/p/a/extra"},
{size: 10, head: "h", tail: "t", path: "/q/b/x/f"}, {size: 10, head: "h", tail: "t", content: "c", path: "/q/b/x/f"},
} }
super, dirs := buildHierarchy(recs) super, dirs := treeOf(t, recs)
super.compute()
got := groupPaths(collectTreeGroups(dirs, super)) got := groupPaths(collectTreeGroups(dirs, super))