sfdupes
Description
sfdupes is an MIT-licensed Go CLI tool by
@sneak that quickly identifies candidate
duplicate files — and, ultimately, entire duplicate directory trees —
across very large filesystems without reading full file contents. Files
are considered duplicates when they have identical size, identical
SHA-256 of their first 1024 bytes, and identical SHA-256 of their last
1024 bytes. This is a strong candidate signal, not proof of identical
content (the middle of the file is never read); the intended use is
finding duplicate downloads and duplicated directory trees on
multi-terabyte ZFS servers where reading every byte is prohibitively
expensive. scan maintains a persistent SQLite database of file
signatures that survives between runs, so it can be run from cron and
the reports can be generated at any time from the most recent scan.
This tool was created by @sneak to scratch an itch, using Claude Code/Fable.
This README is the complete and authoritative specification.
Getting Started
make build
export SFDUPES_DATABASE="$HOME/.local/share/sfdupes/db.sqlite"
./sfdupes scan /srv
./sfdupes report > dupes.tsv
./sfdupes trees > dupetrees.tsv
scan walks one or more filesystem trees and maintains one database
record per regular file (path, size, mtime, head hash, tail hash). The
database persists between runs; a rescan only hashes files that are new
or changed, and removes records for files that no longer exist.
report reads the database and prints the file-level duplicates
report. trees reads the same database and prints the duplicate-tree
report. A missing/invalid subcommand — or a scan invocation with no
PATH operand — prints a usage message and exits 2.
The database defaults to /var/lib/sfdupes/db.sqlite and can be placed
anywhere by setting SFDUPES_DATABASE. The intended deployment is a
daily sfdupes scan cron job, with the reporting commands run
interactively whenever needed; their results are as fresh as the last
completed scan.
Rationale
Duplicate finders that hash entire files do not scale to the target environment: ~10 million files and ~150 TB on possibly slow or busy disks (a ZFS pool under resilver). Reading at most 2 KiB per file — and only from files whose size at least one other file shares, since a size-unique file cannot be a duplicate — makes a full-filesystem sweep tractable, and the signatures are kept in a persistent database, so the expensive filesystem pass is incremental: a rescan re-hashes only files whose recorded mtime or size changed, and all analysis happens offline from the database alone. The end goal is not individual files but whole duplicated trees — duplicate extractions, duplicate downloads, copied project trees — which an operator can consider removing as a unit.
Design
Goals, in order:
- Find whole duplicate trees, not just files. The end goal is to identify places where the exact same set of files and directories exists at two or more paths (duplicate extractions, duplicate downloads, copied project trees), so the operator can consider removing an entire subtree at once. File-level duplicate detection is the foundation; tree-level detection is built on top of it.
- Never read full file contents. At most 2 KiB is read per file (first and last 1024 bytes), and only files whose size at least one other file shares are read at all — a size-unique file cannot be a duplicate. Scale target: tens of millions of files, ~150 TB filesystem, possibly slow or busy disks (ZFS pool under resilver). Holding one small record (path, size, mtime) per file in memory during a scan is acceptable; holding every file's hashes is not (they stay in the database).
- Scan incrementally, analyze offline. The expensive filesystem
scan maintains a persistent database; an unchanged file is never
read again on a rescan. All analysis (
report,trees) works from the database alone and must never touch the scanned filesystem again.scanis designed to be cronned; the reports run at any time against the last completed scan. - Clean stream separation. Everything on stdout is machine-readable data. All progress, warnings, and summaries go to stderr. Never mix them.
Constraints
- Language: Go (module
sneak.berlin/go/sfdupes). Binary name:sfdupes. - Dependencies: standard library,
github.com/spf13/cobrafor the CLI, one progress-bar library (github.com/schollz/progressbar/v3), and one SQLite driver (modernc.org/sqlite, pure Go, so builds keep cgo disabled).github.com/spf13/viperis permitted if configuration-file support is ever needed, but is not currently used. No other third-party deps. - Cross-compilation is not a concern. Builds run with cgo disabled (the
MakefileexportsCGO_ENABLED=0); the code must remain pure Go. - Analysis modes (
report,trees) must be deterministic: identical database contents, identical output, regardless of the order in which records were inserted.
Subcommands
Three subcommands, all implemented:
scan— walk the filesystem and synchronize the database: one signature record per regular file.report— file-level duplicate report from the database.trees— tree-level duplicate report: reconstruct the directory hierarchy from the database records, compute a Merkle-style digest per directory, and report maximal groups of identical trees.
sfdupes scan [--workers N] [-x] PATH...
sfdupes report > dupes.tsv
sfdupes trees > dupetrees.tsv
Database
All three subcommands operate on a single SQLite database file:
-
Location: the value of the
SFDUPES_DATABASEenvironment variable when set and non-empty, otherwise/var/lib/sfdupes/db.sqlite. There is no command-line flag. -
scancreates the database (and its parent directory) on first use.reportandtreesrequire an existing database; a missing database file is a fatal error (exit 1) telling the user to runscanfirst. -
The database uses WAL journal mode and a busy timeout, so running a report while a cron
scanis in progress is safe. The filesystem is authoritative; the database is an eventually-consistent reflection of it. Hashed records are committed in batched transactions while the scan is still running (keeping the WAL small and letting concurrent reports observe progress), so a report may see a scan's changes partially applied, and a scan that dies partway leaves a valid database holding everything hashed so far; the next scan skips those records and converges toward the filesystem. -
Schema (
PRAGMA user_versionis the schema version, currently 1; a database with any other version is a fatal error):CREATE TABLE files ( path BLOB PRIMARY KEY, -- absolute path, raw bytes size INTEGER NOT NULL, -- bytes, from lstat mtime INTEGER NOT NULL, -- Unix seconds, from lstat head TEXT NOT NULL, -- lowercase-hex SHA-256, first 1 KiB tail TEXT NOT NULL -- lowercase-hex SHA-256, last 1 KiB ) WITHOUT ROWID;Paths are stored as BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
mtimeis used only for change detection; it is not part of the duplicate key.headandtailare empty strings when the file has never been hashed because its size was unique as of the last scan that covered it; such records still define the file for tree reconstruction but never participate in duplicate groups.
scan mode
scan requires one or more PATH operands naming the trees to scan.
There is no default path; invoking scan with no operand is a usage
error (usage message on stderr, exit 2). An operand may be a directory
or a regular file; an operand that does not exist is a fatal error
(exit 1). Because database records persist between runs and are keyed
by absolute path, each operand is resolved to an absolute, lexically
cleaned path (symlinks are not resolved) before walking, so results do
not depend on the working directory. All operands belong to a single
scan and are enumerated concurrently: every operand seeds the shared
walk worker pool. Overlapping operands are harmless — an operand that
duplicates another or lies under another is dropped before walking,
so every file is reached exactly once and produces one database
record.
scan synchronizes the database with the filesystem state under the
scanned operands:
- Only a file whose size at least one other file shares is ever
read: a size-unique file cannot be a duplicate, so it is recorded
without hashes (
headandtailempty). The size census covers every file walked this scan plus every database record outside the scanned operands, so a possible duplicate of a separately scanned tree is still recognized. - A file not yet in the database is inserted: hashed when its size is shared, without hashes otherwise.
- A file already in the database is skipped without reading its contents when its lstat size equals the recorded size and its lstat mtime is not newer than the recorded mtime. This is what makes a daily rescan cheap. Exception: an unchanged file whose record lacks hashes is hashed — and its record updated — once its size becomes shared, so hashing deferred by size-uniqueness happens as soon as it could matter.
- A file whose mtime is newer than recorded, or whose size differs, is processed as if new: re-hashed, or recorded without hashes, per the shared-size rule.
- A database record whose path lies under one of the scanned operands but was not successfully processed this run is deleted. This removes records for deleted files. It also removes records for paths that failed to stat or hash this run: the database only ever contains signatures verified by the most recent scan that covered them (a subsequent successful scan re-adds such files).
- Database records outside the scanned operands are untouched, so disjoint trees can be scanned on different schedules into the same database.
scan runs three sequential phases over the whole scan.
Parallelism lives inside each phase; batched database writes begin
during the hash phase:
- walk + stat — enumerate the trees under all
PATHoperands concurrently with the walk worker pool: every operand seeds the shared queue, and each worker reads one directory at a time, handing discovered subdirectories back to the queue and runninglstaton each regular file as it is discovered (while the directory's metadata is still hot). Sequential directory enumeration is metadata-latency-bound and takes hours at tens of millions of files; per-directory parallelism is what makes the walk tractable on large or busy pools. The walk builds the size census and resolves unchanged already-hashed files on the fly; every other file is carried to the hash phase as a (path, size, mtime) record. - hash — with the census complete, each carried file's size
decides its fate. Size-unique files are never read: new or
changed ones are recorded without hashes in the update phase,
unchanged unhashed ones simply keep their records. Every file
with a shared size is hashed by the worker pool: read the first
min(1024, size)bytes and the lastmin(1024, size)bytes (one read whensize <= 1024, since the two windows coincide; forsize == 0hash the empty input) and compute the SHA-256 of each. The phase total is exact, so progress and ETA are meaningful. Completed records are committed in batched transactions while hashing runs, so a scan interrupted after hours keeps everything hashed so far and the next scan resumes cheaply, skipping records already written. - update — commit the final partial batch, the hash-less records for size-unique new and changed files, and the deletions for records the scan did not verify (vanished files, plus paths that failed to stat or hash).
Rules for the walk:
- Only regular files. Skip directories, symlinks (do not follow, including symlink operands), sockets, FIFOs, and device nodes.
- Never descend into a directory named
.zfs(ZFS snapshot pseudo-dirs; walking them would list every file once per snapshot). - Filesystem boundaries are crossed by default. With
-x(long form--one-file-system, following the GNUdu/rsyncconvention), never descend into a directory on a different filesystem than itsPATHoperand; each operand is bounded by its own filesystem. - On any per-path error (permission denied, file vanished between passes, unreadable): print a one-line warning to stderr, skip the path, and continue. Per-file errors never abort the run; the final summary reports how many were skipped. As specified above, a skipped path that has a database record from an earlier scan loses that record; an unreadable directory subtree likewise loses its records (accepted: the database mirrors what the latest scan could actually verify).
Concurrency: the walk phase (which also stats files) and the hash
phase each use a worker pool of --workers workers (default
runtime.NumCPU()); the walk parallelizes across directories,
hashing across files. Both phases are seek-bound on spinning disks,
so raising --workers well past the core count can help on pools
with many spindles. The main goroutine owns partitioning, database
writes, and progress rendering; progress display must never block
the workers.
scan writes nothing to stdout. The summary line on stderr reports the
files seen this run broken down by disposition, plus skips:
scan: 123456 files seen (1200 added, 34 updated, 56 removed, 122166 unchanged), 3 skipped
(removed counts deleted database records, which are not part of the
files-seen total.)
report mode
report reads every record from the database and takes no positional
arguments.
report must never touch the filesystem being analyzed. It does not
stat, open, or otherwise access any path that appears in the records; its
only I/O is reading the database and writing stdout/stderr. It must
produce identical output whether or not the scanned filesystem is still
mounted.
Processing:
- Records without hashes (size-unique when last scanned) are excluded: their content is unknown, so they are never reported as duplicates.
- Group the remaining records by the key
(size, head_hash, tail_hash). - Every group with two or more paths is a duplicate group.
- Within each group, sort paths lexicographically (byte order). The
first path is the group's
first; every other path is adupe. - Order groups by size descending (biggest reclaimable space first),
tie-broken by
firstpath ascending. Output must be fully deterministic for a given database state.
Report output format
TSV on stdout: a header line, then one row per duplicate file (N-1 rows for a group of N):
first dupe size
/srv/a/big.iso /srv/b/big-copy.iso 4294967296
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296
Summary to stderr: records read, number of duplicate groups, number of
dupe files, and total reclaimable bytes (sum of size over all dupe
rows) in human units.
trees mode
trees reads the same database as report (no positional arguments)
and reports entire duplicate directory trees: directories under
which the exact same set of relative paths exists with the exact same
file signatures.
trees must never touch the filesystem being analyzed — the same
rule as report. The directory hierarchy is reconstructed purely from
the paths in the records, split on /.
Definitions:
- A file's signature is
(size, head_hash, tail_hash)— mtime is informational and excluded. An unhashed record (empty hashes) has unknown content: its signature is treated as unique to that file, so a tree containing an unhashed file never compares equal to any other tree. - A directory's digest is a SHA-256 Merkle digest computed bottom-up: serialize the directory's child entries — for a file child, its name and signature; for a subdirectory child, its name and that subdirectory's digest — sort the serialized entries byte-lexicographically, and hash the concatenation. Names are part of the digest: two trees whose files differ only in name are not duplicates.
- Two directories are duplicate trees when their digests are equal. Equal digests imply equal recursive file count and equal total byte size.
Known limitation (accepted): only regular files that appear in the database define a tree. Empty directories are invisible, and a file skipped during the scan (e.g. permission error) in one copy but not the other will make otherwise-identical trees compare as different.
Processing:
- Build the hierarchy, compute every directory's digest, and group directories by digest. Every group with two or more directories is a duplicate-tree group.
- Report only maximal trees. A group is suppressed when its members' parents are pairwise distinct directories that all share a single digest — such a group is wholly implied by its parents' (or a further ancestor's) group. Groups containing sibling directories, or members whose parents differ, are always reported.
- Within each group, sort paths lexicographically (byte order); the
first path is
first, every other path is adupe. - Order groups by total tree size descending, tie-broken by
firstpath ascending. Output must be fully deterministic for a given input.
Trees output format
TSV on stdout: a header line, then one row per duplicate tree (N-1 rows
for a group of N). files is the recursive regular-file count of one
copy of the tree; size is the recursive total byte size of one copy:
first dupe files size
/srv/a/project /srv/backup/project 3417 104857600
Summary to stderr: records read, number of duplicate-tree groups,
number of dupe trees, and total reclaimable bytes (sum of size over
all dupe rows) in human units.
Progress
Use the progress-bar library for all scan progress; rendering in the
style of pv is the model. All progress goes to stderr.
Each phase gets its own display. The walk has no known total while running: show a live file count, rate, and elapsed time (spinner-style, no percentage or ETA). The hash and update phases have exact totals — only files that actually need hashing appear in the hash total, so its ETA is meaningful. Required elements for the bars with known totals:
- elapsed time
- estimated time remaining
- a
[m/n] x%display (items processed / total items, percent) - current rate (items/s)
Example shape (exact layout is flexible, content is not):
hash: [12345/98765] 12% |████ | 92 files/s elapsed 2:32 eta 17:54
Additional requirements:
- When stderr is not a TTY, do not emit ANSI redraws: print a plain one-line progress update no more often than every 5 seconds instead.
- Progress updates are driven from the main goroutine and must be non-blocking with respect to the worker pool.
reportandtreesmodes need no progress display, only their stderr summaries.
Error handling and exit codes
0: success, even if individual files were skipped with warnings.1: fatal error (e.g., aPATHoperand does not exist, the database cannot be created/opened/read/written, a missing database forreport/trees, stdout write failure).2: usage error (includingscanwith noPATHoperand andreport/treeswith any positional argument).
Build
The Makefile is the single source of truth for all operations:
make/make build— build thesfdupesbinary (cgo disabled); building is the default target.make test— run the test suite (30-second timeout; reruns with-von failure).make lint— rungolangci-lintwith the repo config.make fmt/make fmt-check— format Go sources / verify formatting without writing.make check—test,lint, andfmt-check; modifies nothing.make docker— build the Docker image, which runsmake checkas a build stage.make hooks— install the pre-commit hook.make clean— remove the binary and any legacy localfiles.dat.
Definition of done
All of the following, run in this directory, must pass:
-
make checkpasses (tests, lint,gofmt). -
make dockersucceeds. -
Smoke test — create a throwaway tree in a temp dir (never test against real data):
d=$(mktemp -d) export SFDUPES_DATABASE="$d/db.sqlite" mkdir -p "$d/a" "$d/b" head -c 2000 /dev/urandom > "$d/a/one.bin" cp "$d/a/one.bin" "$d/b/copy.bin" cp "$d/a/one.bin" "$d/b/copy2.bin" head -c 2000 /dev/urandom > "$d/a/unique.bin" # same size, different content printf 'x' > "$d/tiny1"; printf 'x' > "$d/tiny2" # 1-byte duplicates printf 'y' > "$d/tiny3" # 1-byte non-duplicate : > "$d/empty1"; : > "$d/empty2" # empty duplicates # duplicate trees: t1 and t2 are identical; t3 differs by one filename mkdir -p "$d/t1/sub" "$d/t2/sub" "$d/t3/sub" head -c 3000 /dev/urandom > "$d/t1/f1" head -c 100 /dev/urandom > "$d/t1/sub/f2" cp "$d/t1/f1" "$d/t2/f1" cp "$d/t1/sub/f2" "$d/t2/sub/f2" cp "$d/t1/f1" "$d/t3/f1" cp "$d/t1/sub/f2" "$d/t3/sub/f2renamed" ./sfdupes scan "$d" ./sfdupes report ./sfdupes trees # incremental behavior (scan a subtree; records elsewhere persist): ./sfdupes scan "$d/a" # everything unchanged, nothing hashed printf 'z' >> "$d/a/one.bin" # modify: next scan re-hashes it rm "$d/a/unique.bin" # delete: next scan removes its record ./sfdupes scan "$d/a" # 1 updated, 1 removed ./sfdupes report(The scan database lives inside
$dhere purely for test hygiene; scanning$dtherefore also records the SQLite file itself, which is harmless.)Expected from the first
report:one.bin/copy.bin/copy2.binform one group (two dupe rows,firstis the lexicographically smallest path);t1/f1/t2/f1/t3/f1form one group;t1/sub/f2/t2/sub/f2/t3/sub/f2renamedform one group;tiny1/tiny2pair;empty1/empty2pair;unique.binandtiny3appear nowhere; groups ordered by size descending.Expected from
trees: exactly one row —first$d/t1,dupe$d/t2, 2 files, 3100 bytes.$d/t1/subvs$d/t2/subis suppressed as non-maximal (implied by thet1/t2group), andt3appears nowhere (its file set differs by name).Expected from the second
report(after the modify/delete rescan):one.binhas left its group (its content changed), socopy.bin/copy2.binremain as one pair, andunique.binis gone from the database.The test suite automates this scenario (see
scan_test.go), plus a negative check:reportandtreesoperate on the database alone and never touch the scanned filesystem.
TODO
Tracked in TODO.md.
Non-goals
- No full-content verification, no byte-for-byte compare, no deletion or linking of duplicates. The reports are advisory; acting on them is the user's job.
- No persistence beyond the SQLite database described above; no export/import formats.
- No daemon or filesystem watcher; scheduling rescans is cron's job.
License
MIT. See LICENSE.