check / check (push) Failing after 2s
script/test-race runs go test -race in a digest-pinned Debian golang image that has gcc, since the detector needs cgo and the build keeps it off. The checkout is mounted read-only and the container is removed afterwards. The tests run as the calling user, or as nobody when that is root, so the tests that make a file unreadable still see the read fail. It is not part of make check. The detector found no races. Model: opus-5-5
950 lines
52 KiB
Markdown
950 lines
52 KiB
Markdown
# sfdupes
|
|
|
|
## Description
|
|
|
|
`sfdupes` is an MIT-licensed Go CLI tool by [@sneak](https://sneak.berlin) that
|
|
quickly identifies _candidate_ duplicate files — and, ultimately, entire
|
|
duplicate directory trees — across very large filesystems without reading every
|
|
byte of every file. Files are considered duplicates when their sizes are equal
|
|
and they agree on a short ladder of hashes. A file under 10 MiB is hashed in
|
|
full and compared directly. A larger file is gated first on the SHA-256 of its
|
|
first 64 KiB and of its last 64 KiB, and only when its size and both of those
|
|
match another file's is it read for a content hash to compare — the SHA-256 of
|
|
the whole file when it is under 50 MiB, or of gigabyte-spaced 1 MiB samples when
|
|
it is 50 MiB or larger. Below 50 MiB the content hash is proof of identical
|
|
content; at or above 50 MiB it is a strong candidate signal rather than proof,
|
|
because the gaps between samples are never read. The intended use is finding
|
|
duplicate downloads and duplicated directory trees on multi-terabyte ZFS servers
|
|
where reading every byte of every file is prohibitively expensive. `scan`
|
|
maintains a persistent SQLite database of file signatures that survives between
|
|
runs, so it can be run from cron and the reports can be generated at any time
|
|
from the most recent scan.
|
|
|
|
This README is the complete and authoritative specification.
|
|
|
|
## Getting Started
|
|
|
|
```sh
|
|
make build
|
|
export SFDUPES_DATABASE="$HOME/.local/share/sfdupes/db.sqlite"
|
|
./sfdupes scan /srv
|
|
./sfdupes report > dupes.tsv
|
|
./sfdupes trees > dupetrees.tsv
|
|
```
|
|
|
|
`scan` walks one or more filesystem trees and maintains one database record per
|
|
regular file (path, size, mtime, head hash, tail hash, content hash). The
|
|
database persists between runs; a rescan only hashes files that are new or
|
|
changed, or that may have gained a duplicate since the last scan, and removes
|
|
records for files that no longer exist. `report` reads the database and prints
|
|
the file-level duplicates report. `trees` reads the same database and prints the
|
|
duplicate-tree report. A missing/invalid subcommand — or a `scan` invocation
|
|
with no `PATH` operand — prints a usage message and exits 2.
|
|
|
|
The database defaults to `/var/lib/sfdupes/db.sqlite` and can be placed anywhere
|
|
by setting `SFDUPES_DATABASE`. The intended deployment is a daily `sfdupes scan`
|
|
cron job, with the reporting commands run interactively whenever needed; their
|
|
results are as fresh as the last completed scan.
|
|
|
|
### Install
|
|
|
|
With Go installed, this builds and installs the current `main` branch:
|
|
|
|
```sh
|
|
go install sneak.berlin/go/sfdupes@main
|
|
```
|
|
|
|
The binary goes to `$(go env GOPATH)/bin`, or to `$GOBIN` when that is set. A
|
|
binary installed this way reports its version as `dev`; one built from a clone
|
|
or into the Docker image carries the git tag or commit it was built from.
|
|
|
|
From a clone, `make build` writes the binary to `./sfdupes`:
|
|
|
|
```sh
|
|
git clone https://git.eeqj.de/sneak/sfdupes.git
|
|
cd sfdupes
|
|
make build
|
|
```
|
|
|
|
Copy the binary to `/usr/local/bin` for the cron job below.
|
|
|
|
`make docker` builds the Docker image, tagged `sfdupes`, after running the tests
|
|
and the linter (see "Build"). The image runs `sfdupes` as root with the database
|
|
at its default path, so a bind mount of `/var/lib/sfdupes` keeps the database
|
|
between runs. Mount the scanned tree at the same path inside the container as on
|
|
the host; read-only is enough. The database records paths as the container sees
|
|
them, so the reports then name the host's paths.
|
|
|
|
```sh
|
|
make docker
|
|
docker run --rm -v /srv:/srv:ro -v /var/lib/sfdupes:/var/lib/sfdupes \
|
|
sfdupes scan /srv
|
|
docker run --rm -v /var/lib/sfdupes:/var/lib/sfdupes sfdupes report > dupes.tsv
|
|
```
|
|
|
|
### Daily scan from cron
|
|
|
|
Run `scan` as root, so that it can read every file: a path it cannot read is
|
|
skipped with a warning and loses its database record (see "Rules for the walk").
|
|
As a file `/etc/cron.d/sfdupes`:
|
|
|
|
```
|
|
30 3 * * * root /usr/local/bin/sfdupes scan /srv 2>>/var/log/sfdupes.log || tail -n 3 /var/log/sfdupes.log
|
|
```
|
|
|
|
- The database is `/var/lib/sfdupes/db.sqlite`, created with its directory by
|
|
the first scan. To keep it elsewhere, set
|
|
`SFDUPES_DATABASE=/path/to/db.sqlite` before the command on the same line.
|
|
- `scan` writes nothing to stdout. Its stderr, appended here to
|
|
`/var/log/sfdupes.log`, holds a plain progress line as each phase starts and
|
|
then at most every 5 seconds, a warning for each path it skips, and the
|
|
summary line (see "Progress" and "`scan` mode"). The log grows with every
|
|
scan; rotate it like any other.
|
|
- Skipped paths do not fail a scan: it still exits 0, and cron sends nothing. A
|
|
scan that fails, or is stopped by `SIGINT` or `SIGTERM`, exits 1 with the
|
|
reason among the last lines of the log; `tail` prints them, and cron mails
|
|
them to root if the host can send mail.
|
|
- A scan still running when the next one starts carries on. The new one fails at
|
|
once, and the lines cron mails include
|
|
`sfdupes: another scan is running (lock held on /var/lib/sfdupes/db.sqlite.lock)`.
|
|
- `report` and `trees` need only read access to the database (see "Database").
|
|
Under the usual umask of `022` the first scan creates it readable by every
|
|
user, so an unprivileged user can run them against root's database.
|
|
|
|
### Reading the reports
|
|
|
|
Each row of `report` names two copies of one file, and each row of `trees` two
|
|
copies of one directory tree (see "Report output format" and "Trees output
|
|
format"). In a group of copies, the path that sorts first byte by byte is
|
|
`first` and every other path is a `dupe` of it. `first` says nothing about which
|
|
copy is the original or the oldest; which copy to keep is your choice.
|
|
|
|
A row is a candidate, not proof:
|
|
|
|
- The reports read only the database, so they show the files as of the last
|
|
scan; a file may have changed or gone since.
|
|
- A file of 50 MiB or more is compared only on samples of its content (see
|
|
"Duplicate detection").
|
|
- Paths that are hard links to one file are listed as duplicates, but they share
|
|
their data, so removing one frees nothing.
|
|
|
|
Compare a pair byte for byte before removing either copy. For the row
|
|
`/srv/a/big.iso`, `/srv/b/big-copy.iso`, `4294967296`:
|
|
|
|
```sh
|
|
cmp /srv/a/big.iso /srv/b/big-copy.iso && echo identical
|
|
[ /srv/a/big.iso -ef /srv/b/big-copy.iso ] && echo "hard links"
|
|
```
|
|
|
|
`cmp` prints nothing and exits 0 only when every byte matches, and otherwise
|
|
reports where the files differ. The second line prints `hard links` when the two
|
|
paths are the same file, so removing either frees nothing.
|
|
|
|
A path holding a backslash, tab, newline or carriage return is escaped in the
|
|
reports (see "Report output format"). Undo the escapes before using it.
|
|
`printf '%b'` does exactly that, because every backslash in an escaped path
|
|
starts one of the four escapes. Command substitution drops trailing newlines, so
|
|
print an `x` after the path and remove it afterwards, or a path that ends in a
|
|
newline names a different file:
|
|
|
|
```sh
|
|
p="$(printf '%bx' '/srv/a/tab\tname.txt')"; p="${p%x}"
|
|
cmp "$p" /srv/b/tab-copy.txt
|
|
```
|
|
|
|
Check a `trees` row with `diff -r`, which compares the two trees file by file
|
|
and also names anything present in only one of them, such as an empty directory
|
|
or a symlink, which `trees` does not see.
|
|
|
|
## Rationale
|
|
|
|
Duplicate finders that hash entire files do not scale to the target environment:
|
|
~10 million files and ~150 TB on possibly slow or busy disks (a ZFS pool under
|
|
resilver). sfdupes spends disk I/O only on files whose size at least one other
|
|
file shares, since a size-unique file cannot be a duplicate. Of those, a file
|
|
under 10 MiB is read in full; a larger one has its cheap end windows read first,
|
|
and is read for a content hash only when its size and both end windows match
|
|
another file's — the whole file below 50 MiB, but only gigabyte-spaced samples
|
|
at or above 50 MiB, so the largest files are never read in full. This keeps a
|
|
full-filesystem sweep tractable, and the signatures are kept in a persistent
|
|
database, so the expensive filesystem pass is incremental: a rescan re-hashes
|
|
only files whose recorded mtime or size changed, plus — for its content hash — a
|
|
file of 10 MiB or more whose size and end windows have come to match another
|
|
file's. All analysis happens offline from the database alone. The end goal is
|
|
not individual files but whole duplicated trees — duplicate extractions,
|
|
duplicate downloads, copied project trees — which an operator can consider
|
|
removing as a unit.
|
|
|
|
## Design
|
|
|
|
Goals, in order:
|
|
|
|
1. **Find whole duplicate trees, not just files.** The end goal is to identify
|
|
places where the exact same set of files and directories exists at two or
|
|
more paths (duplicate extractions, duplicate downloads, copied project
|
|
trees), so the operator can consider removing an entire subtree at once.
|
|
File-level duplicate detection is the foundation; tree-level detection is
|
|
built on top of it.
|
|
2. **Spend I/O in proportion to duplicate likelihood.** Only files whose size
|
|
at least one other file shares are read at all — a size-unique file cannot
|
|
be a duplicate. Those are compared by the ladder in "Duplicate detection"
|
|
below: a file under 10 MiB is hashed in full, while a larger file is gated
|
|
on cheap 64 KiB end windows first, and gets a content hash only when its
|
|
size and both end windows match another file's. That hash reads the whole
|
|
file below 50 MiB but only gigabyte-spaced 1 MiB samples at or above it, so
|
|
the very largest files are still never read in full. Scale target: tens of
|
|
millions of files, ~150 TB filesystem, possibly slow or busy disks (ZFS pool
|
|
under resilver). Holding one small record (path, size, mtime) per file in
|
|
memory during a scan is acceptable; holding every file's hashes is not (they
|
|
stay in the database). The reporting commands do not hold every file's
|
|
hashes either: `report` lets SQLite group and order the records and writes
|
|
each row as it reads it, so its memory does not grow with the database, and
|
|
`trees` reads the records in path order and keeps each directory's path,
|
|
digest and totals, plus the hashes of only the files in the directories
|
|
holding the record being read, so its memory grows with the number of
|
|
directories and with the size of the largest directory.
|
|
3. **Scan incrementally, analyze offline.** The expensive filesystem scan
|
|
maintains a persistent database; an unchanged file is never read again on a
|
|
rescan, except to compute its content hash once a file of 10 MiB or more
|
|
comes to match another on size and both end windows. All analysis (`report`,
|
|
`trees`) works from the database alone and must never touch the scanned
|
|
filesystem again. `scan` is designed to be cronned; the reports run at any
|
|
time against the last completed scan.
|
|
4. **Clean stream separation.** Everything on stdout is machine-readable data.
|
|
All progress, warnings, summaries, and help and usage text go to stderr.
|
|
Never mix them.
|
|
|
|
### Constraints
|
|
|
|
- Language: Go (module `sneak.berlin/go/sfdupes`). Binary name: `sfdupes`.
|
|
- Dependencies: standard library, `github.com/spf13/cobra` for the CLI, **one
|
|
progress-bar library** (`github.com/schollz/progressbar/v3`),
|
|
`golang.org/x/term` to tell whether stderr is a terminal, **one SQLite
|
|
driver** (`modernc.org/sqlite`, pure Go, so builds keep cgo disabled), and
|
|
`golang.org/x/sys` for `flock(2)` (the scan lock, see "Database").
|
|
`github.com/spf13/viper` is permitted if configuration-file support is ever
|
|
needed, but is not currently used. No other third-party deps.
|
|
- Cross-compilation is not a concern. Builds run with cgo disabled (the
|
|
`Makefile` exports `CGO_ENABLED=0`); the code must remain pure Go.
|
|
- Analysis modes (`report`, `trees`) must be deterministic: identical database
|
|
contents, identical output, regardless of the order in which records were
|
|
inserted.
|
|
|
|
### Subcommands
|
|
|
|
Three subcommands, all implemented:
|
|
|
|
1. `scan` — walk the filesystem and synchronize the database: one signature
|
|
record per regular file.
|
|
2. `report` — file-level duplicate report from the database.
|
|
3. `trees` — tree-level duplicate report: reconstruct the directory hierarchy
|
|
from the database records, compute a Merkle-style digest per directory, and
|
|
report maximal groups of identical trees.
|
|
|
|
```
|
|
sfdupes scan [--workers N] [-x] PATH...
|
|
sfdupes report > dupes.tsv
|
|
sfdupes trees > dupetrees.tsv
|
|
sfdupes --version
|
|
sfdupes [command] --help
|
|
```
|
|
|
|
`--workers N` sets the size of each `scan` worker pool (default: the number of
|
|
CPUs), and `-x` (`--one-file-system`) keeps the walk of each operand on that
|
|
operand's filesystem; both are described under "`scan` mode".
|
|
`sfdupes --version` (or `-v`) prints one line, `sfdupes VERSION`, to stdout and
|
|
exits 0, writing nothing to stderr. `-h` or `--help`, alone or after a
|
|
subcommand, prints the help text to stderr and exits 0, writing nothing to
|
|
stdout.
|
|
|
|
### Database
|
|
|
|
All three subcommands operate on a single SQLite database file:
|
|
|
|
- Location: the value of the `SFDUPES_DATABASE` environment variable when set
|
|
and non-empty, otherwise `/var/lib/sfdupes/db.sqlite`. There is no
|
|
command-line flag. The path names the file exactly, whatever characters it
|
|
holds (`?`, `#` and `%` included); a relative path is relative to the working
|
|
directory.
|
|
- `scan` creates the database (and its parent directory) on first use. `report`
|
|
and `trees` require an existing database; a missing database file is a fatal
|
|
error (exit 1) telling the user to run `scan` first.
|
|
- Only one `scan` runs against a database at a time. For its whole run, `scan`
|
|
holds an exclusive `flock(2)` lock on a lock file beside the database, named
|
|
by appending `.lock` to the database path (`/var/lib/sfdupes/db.sqlite.lock`
|
|
by default), taken before it walks the filesystem or opens the database. A
|
|
second `scan` against the same database does not wait: it fails at once with a
|
|
one-line error naming the lock file and exits 1, without walking anything or
|
|
opening the database, and the running scan carries on. The lock file is
|
|
created on first use, open to its owner only, and left in place: a leftover
|
|
file blocks nothing, because the lock ends with the process holding it however
|
|
it ends, a fatal error or an interrupt included, and deleting the file while a
|
|
scan runs would let a second scan start. `report` and `trees` never take the
|
|
lock, so they run during a scan.
|
|
- While `scan` runs, the database is in WAL journal mode with a busy timeout, so
|
|
running a report while a cron `scan` is in progress is safe. The filesystem is
|
|
authoritative; the database is an eventually-consistent reflection of it.
|
|
Hashed records are committed in batched transactions while the scan is still
|
|
running (keeping the WAL small and letting concurrent reports observe
|
|
progress), so a report may see a scan's changes partially applied, and a scan
|
|
that dies partway leaves a valid database holding every batch committed so far
|
|
(an interrupted scan also commits the batch in progress, see "Error handling
|
|
and exit codes"); the next scan skips those records and converges toward the
|
|
filesystem.
|
|
- `scan` switches the database back to rollback-journal mode when it closes it,
|
|
so between scans the database file alone holds the whole database. Each switch
|
|
needs the database to itself: a `scan` that starts while a report is still
|
|
reading waits for it up to the 10-second busy timeout, then fails; a `scan`
|
|
that ends while a report has the database open warns and leaves the database
|
|
in WAL mode until the next scan. `report` writes each row as it reads it, so
|
|
it is still reading while its output is paused (a pager, a stalled pipe), and
|
|
a `scan` started then fails after the busy timeout.
|
|
- `report` and `trees` open the database read-only and need only read access to
|
|
the database file, and no write access to its directory. While the database is
|
|
in WAL mode they also read the `-wal` and `-shm` files beside it, which SQLite
|
|
creates with the database file's permissions.
|
|
- Schema (`PRAGMA user_version` is the schema version, currently 1; a database
|
|
with any other version is a fatal error. `scan` creates the schema and sets
|
|
the version in one transaction, so a first scan stopped while doing so leaves
|
|
an empty database the next scan sets up. A database at version 0 that already
|
|
has a `files` table was therefore not made by sfdupes; every subcommand
|
|
refuses it with an error telling the user to remove the file and rescan):
|
|
|
|
```sql
|
|
CREATE TABLE files (
|
|
path BLOB PRIMARY KEY, -- absolute path, raw bytes
|
|
size INTEGER NOT NULL, -- bytes, from lstat
|
|
mtime INTEGER NOT NULL, -- Unix seconds, from lstat
|
|
head TEXT NOT NULL, -- lowercase-hex SHA-256; first 64 KiB, or whole file under 10 MiB
|
|
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
|
|
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
|
|
) WITHOUT ROWID;
|
|
CREATE INDEX files_signature ON files (size, head, tail, content);
|
|
```
|
|
|
|
Paths are stored as BLOBs because Unix paths are raw bytes, not guaranteed
|
|
UTF-8. `mtime` is used only for change detection; it is not part of the
|
|
duplicate key. For a file under 10 MiB `head`, `tail`, and `content` all
|
|
hold the whole-file hash (that range is hashed in full, with no end
|
|
windows); for a larger file `head` and `tail` hold the first- and last-64
|
|
KiB hashes and `content` the whole-file or sampled hash. All three are empty
|
|
strings when the file has never been hashed because its size was unique as
|
|
of the last scan that covered it. For a file of 10 MiB or more, `content`
|
|
stays empty until the content phase of a scan (see "`scan` mode" below) has
|
|
read the file. A record with an empty `content` is never part of a duplicate
|
|
group, though it still defines the file for tree reconstruction. The
|
|
`files_signature` index lets SQLite group the records by signature for
|
|
`report` without sorting the whole table.
|
|
|
|
### Duplicate detection
|
|
|
|
Two files are duplicates only when they agree on every rung of this ladder; a
|
|
mismatch at any rung means they are not duplicates. `scan` stores each file's
|
|
hashes, and `report` and `trees` group files by the whole signature — size,
|
|
`head`, `tail`, and `content` — so the grouping is exactly this ladder applied
|
|
across everything scanned into the database, even across separate scans.
|
|
|
|
1. **Size.** Files of different sizes are never compared. Only files whose size
|
|
at least one other file shares are hashed at all.
|
|
2. **Under 10 MiB: whole file.** A file smaller than 10 MiB is hashed in full
|
|
and compared directly, with no separate end-window step — small files are
|
|
cheap to read to the last byte, and doing so makes the comparison exact.
|
|
`head`, `tail`, and `content` all hold this whole-file SHA-256, so such a
|
|
file's signature is decided entirely by its size and its content.
|
|
3. **10 MiB and above: head and tail.** For a larger file, the SHA-256 of the
|
|
first 64 KiB (`head`) and of the last 64 KiB (`tail`) are a cheap gate that
|
|
eliminates most same-size pairs before any bulk reading: the content hash of
|
|
the next two rungs is computed only for a file whose size, `head`, and
|
|
`tail` match another file's, whether that file is scanned in the same run or
|
|
stored by an earlier scan. A stored file that first gains such a match in a
|
|
later scan gets its content hash then; until it has one, its `content` is
|
|
empty and it is not a duplicate. At 10 MiB and above the two windows never
|
|
overlap.
|
|
4. **10 MiB and above, content below 50 MiB.** The SHA-256 of the entire file.
|
|
Agreement here is proof of identical content (barring a SHA-256 collision).
|
|
5. **10 MiB and above, content 50 MiB and above.** A sampled SHA-256: the 1 MiB
|
|
window at each gigabyte-aligned offset (0, 1 GiB, 2 GiB, … while inside the
|
|
file, the final window truncated at end of file) is fed, in order, into one
|
|
hash. This is **deliberately probabilistic** — the gaps between samples are
|
|
never read, so two large files that agree on every sample are reported as
|
|
duplicates without being read in full. It is the price of never reading a
|
|
150 GB file end to end. Because size is already part of the signature, only
|
|
equal-size files reach this rung, so their sample boundaries always align.
|
|
|
|
`head`, `tail`, and `content` are one column each. A file below 10 MiB and one
|
|
at or above it never share a size, and neither do a file below 50 MiB and one at
|
|
or above it, so a stored value is never ambiguous between the whole-file,
|
|
end-window, and sampled forms.
|
|
|
|
### `scan` mode
|
|
|
|
`scan` requires one or more `PATH` operands naming the trees to scan. There is
|
|
no default path; invoking `scan` with no operand is a usage error (usage message
|
|
on stderr, exit 2). An operand may be a directory or a regular file; an operand
|
|
that does not exist is a fatal error (exit 1). Because database records persist
|
|
between runs and are keyed by absolute path, each operand is resolved to an
|
|
absolute, lexically cleaned path (symlinks are not resolved) before walking, so
|
|
results do not depend on the working directory. All operands belong to a single
|
|
scan and are enumerated concurrently: every operand seeds the shared walk worker
|
|
pool. Overlapping operands are harmless — an operand that duplicates another or
|
|
lies under another is dropped before walking, so every file is reached exactly
|
|
once and produces one database record.
|
|
|
|
An operand that is a symlink (never followed, not even as an operand), socket,
|
|
FIFO, or device node, or a directory named `.zfs`, is not scanned. `scan` prints
|
|
a one-line warning naming the path and what it is, counts it as skipped, and
|
|
drops it from the scanned operands before reading the database. Another operand
|
|
beneath it is still scanned. The records stored beneath it are not deleted: they
|
|
are treated like any other record outside the scanned operands, including the
|
|
content-phase exception below. If it lies under another operand, they are under
|
|
that operand instead, and are deleted like any other record there that this scan
|
|
did not verify. This is not an error: a scan whose every operand is dropped
|
|
walks nothing and exits 0.
|
|
|
|
`scan` synchronizes the database with the filesystem state under the scanned
|
|
operands:
|
|
|
|
- Only a file whose size at least one other file shares is ever read: a
|
|
size-unique file cannot be a duplicate, so it is recorded without hashes
|
|
(`head`, `tail`, and `content` empty). The size census covers every file
|
|
walked this scan plus every database record outside the scanned operands, so a
|
|
possible duplicate of a separately scanned tree is still recognized.
|
|
- A file not yet in the database is inserted: hashed when its size is shared,
|
|
without hashes otherwise.
|
|
- A file already in the database is **skipped without reading its contents**
|
|
when its lstat size equals the recorded size and its lstat mtime is not newer
|
|
than the recorded mtime. This is what makes a daily rescan cheap. Exception:
|
|
an unchanged file whose record lacks hashes is hashed — and its record updated
|
|
— once its size becomes shared, so hashing deferred by size-uniqueness happens
|
|
as soon as it could matter. Likewise, an unchanged file of 10 MiB or more
|
|
whose record has no `content` hash is read for one by the content phase below
|
|
once its size, `head`, and `tail` match another record's.
|
|
- A file whose mtime is newer than recorded, or whose size differs, is processed
|
|
as if new: re-hashed, or recorded without hashes, per the shared-size rule.
|
|
- A database record whose path lies under one of the scanned operands but was
|
|
not successfully processed this run is deleted. This removes records for
|
|
deleted files. It also removes records for paths that failed to stat or hash
|
|
this run: the database only ever contains signatures verified by the most
|
|
recent scan that covered them (a subsequent successful scan re-adds such
|
|
files). A failure in the content phase below removes nothing: the record is
|
|
left as it is.
|
|
- Database records outside the scanned operands are untouched, so disjoint trees
|
|
can be scanned on different schedules into the same database. The one
|
|
exception is the content phase below: a stored file of 10 MiB or more without
|
|
a `content` hash is read for one, wherever it lies, once its size, `head`, and
|
|
`tail` match another record's. If that file is gone or has changed since its
|
|
record was written, the record is left as it is.
|
|
|
|
`scan` runs **four sequential phases over the whole scan**. Parallelism lives
|
|
inside each phase; batched database writes begin during the hash phase:
|
|
|
|
1. **walk + stat** — enumerate the trees under all `PATH` operands concurrently
|
|
with the walk worker pool: every operand seeds the shared queue, and each
|
|
worker reads one directory at a time, handing discovered subdirectories back
|
|
to the queue and running `lstat` on each regular file as it is discovered
|
|
(while the directory's metadata is still hot). Sequential directory
|
|
enumeration is metadata-latency-bound and takes hours at tens of millions of
|
|
files; per-directory parallelism is what makes the walk tractable on large
|
|
or busy pools. The walk builds the size census and resolves unchanged
|
|
already-hashed files on the fly; every other file is carried to the hash
|
|
phase as a (path, size, mtime) record.
|
|
2. **hash** — with the census complete, each carried file's size decides its
|
|
fate. Size-unique files are never read: new or changed ones are recorded
|
|
without hashes in the update phase, unchanged unhashed ones simply keep
|
|
their records. Every file with a shared size is hashed by the worker pool as
|
|
described in "Duplicate detection" above: a file under 10 MiB in full, which
|
|
gives its `head`, `tail`, and `content` alike, and a larger file only in its
|
|
end windows, which give its `head` and `tail`; its content hash is left to
|
|
the content phase. Zero-length files have constant hashes and are never
|
|
opened. Files are hashed in **inode order** (minimizing seeks on spinning
|
|
disks), and paths that are hard links to the same inode are **read once**,
|
|
all sharing the one result — a hard-link backup farm costs one read per
|
|
inode, not per path. The phase total counts actual reads, so progress and
|
|
ETA are meaningful. Completed records are committed in batched transactions
|
|
**while hashing runs**, so a scan interrupted after hours keeps everything
|
|
hashed so far and the next scan resumes cheaply, skipping records already
|
|
written.
|
|
3. **update** — commit the final partial batch, the hash-less records for
|
|
size-unique new and changed files, and the deletions for records the scan
|
|
did not verify (vanished files, plus paths that failed to stat or hash).
|
|
4. **content** — find every record of 10 MiB or more without a `content` hash
|
|
whose size, `head`, and `tail` equal another record's, anywhere in the
|
|
database: records from this scan and records stored by earlier scans, inside
|
|
or outside the scanned operands. SQLite finds them, so only the records to
|
|
be read are kept in memory, never every file's hashes. Every record sharing
|
|
their size, `head`, and `tail`, including one that already has a `content`
|
|
hash, has its file checked with `lstat` first. A file that is gone, is no
|
|
longer a regular file, or has changed (a different size, or an mtime newer
|
|
than recorded) keeps its record as it is and does not count as a match for
|
|
the others. Any other `lstat` error is warned about and counted as skipped,
|
|
with the same result. If such a record has no `content` hash, it stays out
|
|
of duplicate groups; if it has one, it is still reported until a scan
|
|
covering its own tree updates or removes it. The files that pass and have no
|
|
`content` hash are read only if at least two of those records pass, so a
|
|
file whose only matches are stale costs no read; a file that already has a
|
|
`content` hash is never read again. They are read by a worker pool as in the
|
|
hash phase, in inode order and once per inode, and their content hashes are
|
|
committed in batches. A failed read is warned about and counted as skipped;
|
|
its record keeps an empty `content`, so it is not a duplicate, and a later
|
|
scan tries again.
|
|
|
|
Rules for the walk:
|
|
|
|
- Only regular files. Skip directories, symlinks (do not follow, including
|
|
symlink operands), sockets, FIFOs, and device nodes. An operand that is a
|
|
symlink, socket, FIFO, or device node is dropped as described in "`scan` mode"
|
|
above.
|
|
- Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs; walking
|
|
them would list every file once per snapshot), not even when it is an operand;
|
|
such an operand is dropped the same way.
|
|
- Filesystem boundaries are crossed by default. With `-x` (long form
|
|
`--one-file-system`, following the GNU `du`/`rsync` convention), never descend
|
|
into a directory on a different filesystem than its `PATH` operand; each
|
|
operand is bounded by its own filesystem.
|
|
- On any per-path error (permission denied, file vanished between passes,
|
|
unreadable): print a one-line warning to stderr, skip the path, and continue.
|
|
Per-file errors never abort the run; the final summary reports how many were
|
|
skipped. As specified above, a skipped path that has a database record from an
|
|
earlier scan loses that record, unless it failed only in the content phase, or
|
|
is an operand dropped before the database was read that lies under no other
|
|
operand; an unreadable directory subtree likewise loses its records (accepted:
|
|
the database mirrors what the latest scan could actually verify).
|
|
|
|
Concurrency: the walk phase (which also stats files), the hash phase, and the
|
|
content phase each use a worker pool of `--workers` workers (default
|
|
`runtime.NumCPU()`); the walk parallelizes across directories, hashing across
|
|
files. `--workers` must be at least 1: a smaller value is a usage error,
|
|
reported in one line on stderr with exit 2 before anything is scanned. All three
|
|
phases are seek-bound on spinning disks, so raising `--workers` well past the
|
|
core count can help on pools with many spindles. The main goroutine owns
|
|
partitioning, database writes, and progress rendering; progress display must
|
|
never block the workers.
|
|
|
|
`scan` writes nothing to stdout. The summary line on stderr reports the files
|
|
seen this run broken down by disposition, plus skips:
|
|
|
|
```
|
|
scan: 123400 files seen (1200 added, 34 updated, 56 removed, 122166 unchanged), 3 skipped
|
|
```
|
|
|
|
(`removed` counts deleted database records, which are not part of the files-seen
|
|
total.)
|
|
|
|
### `report` mode
|
|
|
|
`report` reads every record from the database and takes no positional arguments.
|
|
|
|
**`report` must never touch the filesystem being analyzed.** It does not stat,
|
|
open, or otherwise access any path that appears in the records; its only I/O is
|
|
reading the database, writing stdout/stderr, and the temporary file SQLite sorts
|
|
in when the duplicate rows do not fit in memory. SQLite puts that file in
|
|
`$SQLITE_TMPDIR` or `$TMPDIR` when set, otherwise in `/var/tmp` (or `/tmp`), and
|
|
deletes it as soon as it has opened it. `report` must produce identical output
|
|
whether or not the scanned filesystem is still mounted.
|
|
|
|
Processing:
|
|
|
|
- Records without a `content` hash (see "Database" above) are excluded: their
|
|
content is unknown, so they are never reported as duplicates.
|
|
- Group the remaining records by the key `(size, head, tail, content)`.
|
|
- Every group with two or more paths is a duplicate group.
|
|
- Within each group, sort paths lexicographically (byte order). The first path
|
|
is the group's `first`; every other path is a `dupe`.
|
|
- Order groups by size descending (biggest reclaimable space first), tie-broken
|
|
by `first` path ascending. Output must be fully deterministic for a given
|
|
database state.
|
|
|
|
#### Report output format
|
|
|
|
TSV on stdout: a header line, then one row per duplicate file (N-1 rows for a
|
|
group of N):
|
|
|
|
```
|
|
first dupe size
|
|
/srv/a/big.iso /srv/b/big-copy.iso 4294967296
|
|
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296
|
|
```
|
|
|
|
Paths are raw bytes and may hold any byte except NUL, so the path columns
|
|
(`first` and `dupe`) are escaped to keep every row one line of tab-separated
|
|
fields: a backslash is written as `\\`, a tab as `\t`, a newline as `\n`, and a
|
|
carriage return as `\r`. Every other byte is written unchanged, including bytes
|
|
that are not valid UTF-8. Undoing those four escapes gives back the stored path.
|
|
Grouping and ordering use the stored path, not the escaped one. The warnings
|
|
`scan` prints on stderr are escaped the same way, so each warning is one line.
|
|
|
|
Summary to stderr: records read, number of duplicate groups, number of dupe
|
|
files, and total reclaimable bytes (sum of `size` over all dupe rows) in human
|
|
units.
|
|
|
|
### `trees` mode
|
|
|
|
`trees` reads the same database as `report` (no positional arguments) and
|
|
reports **entire duplicate directory trees**: directories under which the exact
|
|
same set of relative paths exists with the exact same file signatures.
|
|
|
|
**`trees` must never touch the filesystem being analyzed** — the same rule as
|
|
`report`. The directory hierarchy is reconstructed purely from the paths in the
|
|
records, split on `/`.
|
|
|
|
Definitions:
|
|
|
|
- A file's **signature** is `(size, head, tail, content)` — mtime is
|
|
informational and excluded. A record without a `content` hash has unknown
|
|
content: its signature is treated as unique to that file, so a tree containing
|
|
such a file never compares equal to any other tree.
|
|
- A directory's **digest** is a SHA-256 Merkle digest computed bottom-up:
|
|
serialize the directory's child entries — for a file child, its name and
|
|
signature; for a subdirectory child, its name and that subdirectory's digest —
|
|
sort the serialized entries byte-lexicographically, and hash the
|
|
concatenation. Names are part of the digest: two trees whose files differ only
|
|
in name are _not_ duplicates.
|
|
- Two directories are **duplicate trees** when their digests are equal. Equal
|
|
digests imply equal recursive file count and equal total byte size.
|
|
|
|
Known limitation (accepted): hard-linked paths are reported as duplicates by
|
|
`report` and count toward duplicate trees — their content is genuinely identical
|
|
— even though they share storage, so removing one reclaims no space. Inode
|
|
identity is used during the scan to avoid redundant reads but is not persisted
|
|
in the database.
|
|
|
|
Known limitation (accepted): only regular files that appear in the database
|
|
define a tree. Empty directories are invisible, and a file skipped during the
|
|
scan (e.g. permission error) in one copy but not the other will make
|
|
otherwise-identical trees compare as different.
|
|
|
|
Processing:
|
|
|
|
- Build the hierarchy, compute every directory's digest, and group directories
|
|
by digest. Every group with two or more directories is a duplicate-tree group.
|
|
- **Report only maximal trees.** A group is suppressed when its members' parents
|
|
are pairwise distinct directories that all share a single digest — such a
|
|
group is wholly implied by its parents' (or a further ancestor's) group.
|
|
Groups containing sibling directories, or members whose parents differ, are
|
|
always reported.
|
|
- Within each group, sort paths lexicographically (byte order); the first path
|
|
is `first`, every other path is a `dupe`.
|
|
- Order groups by total tree size descending, tie-broken by `first` path
|
|
ascending. Output must be fully deterministic for a given input.
|
|
|
|
#### Trees output format
|
|
|
|
TSV on stdout: a header line, then one row per duplicate tree (N-1 rows for a
|
|
group of N). `files` is the recursive regular-file count of one copy of the
|
|
tree; `size` is the recursive total byte size of one copy:
|
|
|
|
```
|
|
first dupe files size
|
|
/srv/a/project /srv/backup/project 3417 104857600
|
|
```
|
|
|
|
The `first` and `dupe` paths are escaped as described under "Report output
|
|
format". The root directory's path is `/`.
|
|
|
|
Summary to stderr: records read, number of duplicate-tree groups, number of dupe
|
|
trees, and total reclaimable bytes (sum of `size` over all dupe rows) in human
|
|
units.
|
|
|
|
### Progress
|
|
|
|
Use the progress-bar library for all scan progress; rendering in the style of
|
|
`pv` is the model. All progress goes to stderr.
|
|
|
|
Each phase gets its own display, rendered the moment the phase starts — a scan
|
|
must never look hung. Loading the existing-record index (`load`) and the walk
|
|
have no known totals while running: show a live count, rate, and elapsed time
|
|
(spinner-style, no percentage or ETA). The content phase's display (`content`)
|
|
starts the same way, counting the records checked while SQLite finds the files
|
|
to read and `lstat` checks them, then shows a bar once reading starts. The hash
|
|
and update phases, and the content phase's reads, have exact totals — only files
|
|
that actually need hashing appear in the hash and content totals, so their ETAs
|
|
are meaningful. Required elements for the bars with known totals:
|
|
|
|
- elapsed time
|
|
- estimated time remaining
|
|
- a `[m/n] x%` display (items processed / total items, percent)
|
|
- current rate (items/s)
|
|
|
|
Example shape (exact layout is flexible, content is not):
|
|
|
|
```
|
|
hash: [12345/98765] 12% |████ | 92 files/s elapsed 2:32 eta 17:54
|
|
```
|
|
|
|
Additional requirements:
|
|
|
|
- When stderr is not a terminal (a pipe, a file, `/dev/null`), do not emit ANSI
|
|
redraws: print a plain one-line progress update the moment each phase starts,
|
|
then no more often than every 5 seconds.
|
|
- Progress updates are driven from the main goroutine and must be non-blocking
|
|
with respect to the worker pool. On a terminal the spinner-style displays also
|
|
redraw on their own several times a second, so their count and elapsed time
|
|
stay current while a phase waits for its next item.
|
|
- A warning printed during a phase always lands on a line of its own, never
|
|
inside the progress display.
|
|
- A bar whose phase stops short of its total, as an interrupted one does, is
|
|
left as last drawn rather than filled up.
|
|
- `report` and `trees` modes need no progress display, only their stderr
|
|
summaries.
|
|
|
|
### Error handling and exit codes
|
|
|
|
- `0`: success, even if individual files were skipped with warnings.
|
|
- `1`: fatal error (e.g., a `PATH` operand does not exist, another `scan` is
|
|
already running against the same database, the database cannot be
|
|
created/opened/read/written, a missing database for `report`/`trees`, stdout
|
|
write failure), or a `scan` stopped by `SIGINT` or `SIGTERM` (see below).
|
|
- `2`: usage error (including `scan` with no `PATH` operand, `scan` with
|
|
`--workers` below 1, and `report`/`trees` with any positional argument).
|
|
|
|
A stdout write failure, such as a full disk, is reported in one line on stderr
|
|
and exits 1. Two cases never reach sfdupes as a failed write:
|
|
|
|
- When the reader of a stdout pipe exits early, as in `sfdupes report | head`,
|
|
the next write ends sfdupes with `SIGPIPE`, quietly and without a summary, the
|
|
way `cat` or `sort` end. The shell reports the signal (status 141 in most
|
|
shells), not exit 1.
|
|
- When stdout is closed outright (`sfdupes report >&-`), the Go runtime opens
|
|
`/dev/null` in its place before sfdupes starts, so the output is discarded and
|
|
the run succeeds, as with `> /dev/null`.
|
|
|
|
`scan` stops cleanly on `SIGINT` (Ctrl-C) or `SIGTERM`. Its workers stop taking
|
|
work, each finishing at most the directory listing or file it is reading; the
|
|
progress display is finished; and the records it has hashed but not yet
|
|
committed are committed, so the next scan does not hash them again. Apart from
|
|
that commit it starts no further writes or deletions: records are deleted only
|
|
after a complete walk, so those under paths an interrupted walk never reached
|
|
are kept. The database is closed and the lock released as on any other exit, the
|
|
line `scan: interrupted after N files` goes to stderr, N being the number of
|
|
files the walk reached, and the exit code is 1. The next scan skips the records
|
|
already written and converges as usual.
|
|
|
|
After the first signal `scan` stops catching them, so a second one ends it at
|
|
once, as an uncaught signal does: the records not yet committed are lost, and
|
|
the database is left valid, as when any scan dies (see "Database"). A `SIGINT`
|
|
that `scan` inherits as ignored, as a script's background job does, stays
|
|
ignored.
|
|
|
|
## Entrypoints
|
|
|
|
This repository adheres to the
|
|
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
|
|
standard: the normalized executables in `script/` are the entrypoints for the
|
|
development workflow, and the `Makefile` targets are thin shims that call them.
|
|
Every script is POSIX `sh`, resolves the repository root itself so it can be run
|
|
from any working directory, and may be invoked directly. The provided
|
|
entrypoints are:
|
|
|
|
- `script/bootstrap` — install everything needed to build and develop this
|
|
repository, idempotently, assuming nothing is present. `git`, `make`, and `go`
|
|
come from the first of nix, apt, brew, or apk found on the host, and are
|
|
presence-checked only. `golangci-lint` and prettier are deliberately **not**
|
|
installed: they run in Docker (see `script/lint` and `script/fmt`) and never
|
|
from a host install, so there is no host copy to drift from the pin. A missing
|
|
`docker` is warned about rather than installed or treated as fatal —
|
|
everything except linting, formatting and `make test-race` works without it.
|
|
Ends with `go mod download`.
|
|
- `script/setup` — make a fresh clone ready for development: runs
|
|
`script/bootstrap`, then `script/install-precommit`.
|
|
- `script/projectname` — print this project's name (`sfdupes`). Scripts that
|
|
need the name call it, so they stay identical across repositories.
|
|
- `script/test` — run the test suite with a 30-second timeout and coverage
|
|
enabled, rerunning verbosely on failure so the logs show which test failed.
|
|
- `script/test-race` — run the test suite under the race detector with a
|
|
60-second timeout. The detector needs cgo and a C compiler, which the build
|
|
never uses, so the tests run in a digest-pinned Debian `golang` image that has
|
|
`gcc`, with the checkout mounted read-only; the container is removed when it
|
|
exits. They run as the calling user, or as `nobody` when that is root, because
|
|
several tests make a file unreadable and root reads it anyway; only then must
|
|
the checkout be readable by other users. Not part of `script/check`. Every run
|
|
starts with empty caches, so it needs the network and takes minutes, and the
|
|
mount needs a local docker daemon.
|
|
- `script/lint` — run the linter. It builds `Dockerfile.lint`, which copies the
|
|
repository into the digest-pinned `golangci/golangci-lint` image and runs
|
|
`golangci-lint config verify` and `golangci-lint run` as build steps, so a
|
|
successful build is a clean lint. That exit status is all it produces, so it
|
|
runs with `--output=type=cacheonly` and writes no image; a run leaves only
|
|
build cache. The linter is never run on the host, which makes a working
|
|
`docker` the one prerequisite for linting — and therefore for `make check` and
|
|
the pre-commit hook. Offline machines: the gate steps themselves make no
|
|
network calls. `golangci-lint run` does not, and neither does
|
|
`golangci-lint config verify` — it validates against a schema the pinned
|
|
binary embeds, measured under `--network none` to both pass a valid config and
|
|
reject an invalid one. The build around them does. `Dockerfile.lint` runs
|
|
`go mod download` before the gates and this module has external dependencies,
|
|
so a first lint on a machine with a cold BuildKit cache reaches the network
|
|
there (as well as pulling the pinned image); under `--network none` it fails
|
|
at that step, before any gate. That layer sits above the gates and stays
|
|
cached, so once it is warm `script/lint` — and with it `make check` — runs
|
|
entirely offline, until `go.mod` or `go.sum` changes and the download layer
|
|
goes cold again. Because the daemon only ever sees a build context, this works
|
|
when the docker daemon is remote and bind mounts are impossible.
|
|
- `script/fmt` — format in place: the Go sources with `gofmt -s -w`, and every
|
|
Markdown file with prettier, at the settings in `.prettierrc` (4-space
|
|
indents, prose wrapped at 80 columns). prettier is pinned by hash through
|
|
`package.json` and `yarn.lock` and never installed on the host: this builds
|
|
the `Dockerfile`'s `prettier` stage, a digest-pinned node image into which
|
|
`yarn install --frozen-lockfile` installs it, tagged `sfdupes-prettier`, and
|
|
runs that with the repository mounted, as the calling user. Needs `docker`,
|
|
and because of the mount, unlike `script/lint`, a local docker daemon.
|
|
- `script/fmt-check` — the read-only counterpart of `script/fmt`, with the
|
|
repository mounted read-only: prints any unformatted file and exits non-zero
|
|
instead of writing. gofmt and prettier both run every time, and each names
|
|
itself when it fails. The `Dockerfile` runs the same two checks as gates: the
|
|
gofmt check in its lint stage, prettier in its `markdown` stage.
|
|
- `script/check` — run `script/test`, `script/lint`, and `script/fmt-check`, in
|
|
that order. Modifies nothing. Needs `docker`, because `script/lint` and
|
|
`script/fmt-check` do.
|
|
- `script/docker` — build the Docker image, tagged with the name from
|
|
`script/projectname`. The `Dockerfile` runs the gates as build steps, so this
|
|
is also the check a developer or reviewer runs by hand.
|
|
- `script/cibuild` — build the Docker image untagged. This is what the Gitea
|
|
workflow runs on push; because the gates run as build steps, a successful
|
|
build implies the repository is green.
|
|
- `script/precommit` — run by the git pre-commit hook: `go mod tidy` must be a
|
|
no-op (a resulting change to `go.mod` or `go.sum` fails the commit), then
|
|
`script/check`.
|
|
- `script/install-precommit` — install the git pre-commit hook that runs
|
|
`script/precommit`. The hook is written to the common git directory, so the
|
|
main checkout and every worktree share it.
|
|
- `script/verify-lint-image-pin` — fail unless the `golangci/golangci-lint`
|
|
reference in `Dockerfile.lint` and the one in the `Dockerfile` lint stage are
|
|
the same image at the same digest, naming both if not. The linter is pinned in
|
|
those two files and nothing else keeps them in sync, so a bump applied to one
|
|
alone would leave `make lint` and the `Dockerfile`'s fail-fast lint stage
|
|
checking the same tree against different rulesets, both green. The guard
|
|
restates neither pin — a third copy would be the same drift one file further
|
|
out — and runs as a gate in both files, so `make lint`, `make check` and
|
|
`make docker` all catch it.
|
|
|
|
`script/verify-linter-pin` used to live here. It compared a linter binary
|
|
against a version pin in `script/bootstrap`, and both of its subjects are gone:
|
|
no linter binary is copied between build stages any more, and bootstrap pins no
|
|
version because it installs no linter. The drift it existed to catch has moved
|
|
from binary-versus-pin to pin-versus-pin, which is what
|
|
`script/verify-lint-image-pin` above checks.
|
|
|
|
`script/lint`, `script/docker` and `script/cibuild` all pass a freshly computed
|
|
`CHECK_EPOCH` build argument, and the gate steps in `Dockerfile.lint` and
|
|
`Dockerfile` reference it. Without that, an unchanged tree lets Docker serve the
|
|
gate layers from cache and the build exits 0 having executed no tests and no
|
|
lint — a green it never earned, and one this repository has produced twice.
|
|
`CHECK_EPOCH` invalidates the gate layers on every run while leaving the pinned
|
|
base images and the dependency layers cached. Each script's value carries the
|
|
process id as well as the epoch, because two runs land inside the same second
|
|
easily and a bare epoch would cache the second one. Each `Dockerfile` stage with
|
|
gates fails when the value is empty, so a bare `docker build .` stops with
|
|
`CHECK_EPOCH is unset; build via script/cibuild or script/docker` instead of
|
|
serving the gates from cache.
|
|
|
|
## Build
|
|
|
|
The `script/` entrypoints above are where the implementations live; the
|
|
`Makefile` targets are shims onto them, except `build`, which carries the
|
|
compile recipe:
|
|
|
|
- `make` / `make build` — build the `sfdupes` binary (cgo disabled); building is
|
|
the default target.
|
|
- `make bootstrap` — install the build and development dependencies.
|
|
- `make setup` — prepare a fresh clone: `bootstrap` plus the pre-commit hook.
|
|
- `make test` — run the test suite (30-second timeout; reruns with `-v` on
|
|
failure).
|
|
- `make test-race` — run the test suite under the race detector, in Docker (see
|
|
`script/test-race`); requires `docker`. Not part of `make check`.
|
|
- `make lint` — run `golangci-lint` with the repo config, in Docker (see
|
|
`script/lint`); requires `docker`.
|
|
- `make fmt` / `make fmt-check` — format the Go sources and the Markdown /
|
|
verify formatting without writing; requires `docker`, for prettier (see
|
|
`script/fmt`).
|
|
- `make check` — `test`, `lint`, and `fmt-check`; modifies nothing. Requires
|
|
`docker`, via `lint` and `fmt-check`.
|
|
- `make docker` — build the Docker image, which runs the gates as build stages.
|
|
- `make hooks` — install the pre-commit hook.
|
|
- `make clean` — remove the binary.
|
|
|
|
### Definition of done
|
|
|
|
All of the following, run in this directory, must pass:
|
|
|
|
1. `make check` passes (tests, lint, `gofmt`, prettier).
|
|
2. `make docker` succeeds.
|
|
3. Smoke test — create a throwaway tree in a temp dir (never test against real
|
|
data):
|
|
|
|
```sh
|
|
d=$(mktemp -d)
|
|
export SFDUPES_DATABASE="$(mktemp -d)/db.sqlite"
|
|
mkdir -p "$d/a" "$d/b"
|
|
head -c 2000 /dev/urandom > "$d/a/one.bin"
|
|
cp "$d/a/one.bin" "$d/b/copy.bin"
|
|
cp "$d/a/one.bin" "$d/b/copy2.bin"
|
|
head -c 2000 /dev/urandom > "$d/a/unique.bin" # same size, different content
|
|
printf 'x' > "$d/tiny1"; printf 'x' > "$d/tiny2" # 1-byte duplicates
|
|
printf 'y' > "$d/tiny3" # 1-byte non-duplicate
|
|
: > "$d/empty1"; : > "$d/empty2" # empty duplicates
|
|
# duplicate trees: t1 and t2 are identical; t3 differs by one filename
|
|
mkdir -p "$d/t1/sub" "$d/t2/sub" "$d/t3/sub"
|
|
head -c 3000 /dev/urandom > "$d/t1/f1"
|
|
head -c 100 /dev/urandom > "$d/t1/sub/f2"
|
|
cp "$d/t1/f1" "$d/t2/f1"
|
|
cp "$d/t1/sub/f2" "$d/t2/sub/f2"
|
|
cp "$d/t1/f1" "$d/t3/f1"
|
|
cp "$d/t1/sub/f2" "$d/t3/sub/f2renamed"
|
|
./sfdupes scan "$d"
|
|
./sfdupes report
|
|
./sfdupes trees
|
|
# incremental behavior (scan a subtree; records elsewhere persist):
|
|
./sfdupes scan "$d/a" # everything unchanged, nothing hashed
|
|
printf 'z' >> "$d/a/one.bin" # modify: next scan re-hashes it
|
|
rm "$d/a/unique.bin" # delete: next scan removes its record
|
|
./sfdupes scan "$d/a" # 1 updated, 1 removed
|
|
./sfdupes report
|
|
```
|
|
|
|
(The database lives in a temp directory of its own: inside `$d`, the scan
|
|
would record it, and its empty lock file would join the `empty1`/`empty2`
|
|
group.)
|
|
|
|
Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin` form one
|
|
group (two dupe rows, `first` is the lexicographically smallest path);
|
|
`t1/f1`/`t2/f1`/`t3/f1` form one group; `t1/sub/f2`/
|
|
`t2/sub/f2`/`t3/sub/f2renamed` form one group; `tiny1`/`tiny2` pair;
|
|
`empty1`/`empty2` pair; `unique.bin` and `tiny3` appear nowhere; groups
|
|
ordered by size descending.
|
|
|
|
Expected from `trees`: exactly one row — `first` `$d/t1`, `dupe` `$d/t2`, 2
|
|
files, 3100 bytes. `$d/t1/sub` vs `$d/t2/sub` is suppressed as non-maximal
|
|
(implied by the `t1`/`t2` group), and `t3` appears nowhere (its file set
|
|
differs by name).
|
|
|
|
Expected from the second `report` (after the modify/delete rescan):
|
|
`one.bin` has left its group (its content changed), so
|
|
`copy.bin`/`copy2.bin` remain as one pair, and `unique.bin` is gone from the
|
|
database.
|
|
|
|
The test suite automates this scenario (see `scan_test.go`), plus a negative
|
|
check: `report` and `trees` operate on the database alone and never touch
|
|
the scanned filesystem.
|
|
|
|
## TODO
|
|
|
|
Tracked in [TODO.md](TODO.md).
|
|
|
|
## Non-goals
|
|
|
|
- No byte-for-byte compare, and no deletion or linking of duplicates. Files that
|
|
match are compared by a SHA-256 of the whole file below 50 MiB, and only by
|
|
samples at 50 MiB and over. The reports are advisory; acting on them is the
|
|
user's job.
|
|
- No persistence beyond the SQLite database described above; no export/import
|
|
formats.
|
|
- No daemon or filesystem watcher; scheduling rescans is cron's job.
|
|
|
|
## License
|
|
|
|
MIT. See [LICENSE](LICENSE).
|
|
|
|
## Author
|
|
|
|
[@sneak](https://sneak.berlin)
|