script/fmt and script/fmt-check run prettier over every Markdown file again, next to gofmt. prettier is pinned by hash through package.json and yarn.lock, copied from the prompts repo with .prettierrc and .prettierignore, and is never installed on a host: a new prettier stage of the Dockerfile installs it into a digest-pinned node image, and both scripts build that stage and run it with the repository mounted. CI checks the Markdown in a markdown stage that the build stage waits on. Because make fmt-check now runs docker, the Dockerfile runs gofmt directly in its lint stage instead. All Markdown is reformatted. Model: opus-5-5
sfdupes
Description
sfdupes is an MIT-licensed Go CLI tool by @sneak that
quickly identifies candidate duplicate files — and, ultimately, entire
duplicate directory trees — across very large filesystems without reading every
byte of every file. Files are considered duplicates when their sizes are equal
and they agree on a short ladder of hashes. A file under 10 MiB is hashed in
full and compared directly. A larger file is gated first on the SHA-256 of its
first 64 KiB and of its last 64 KiB, and only when its size and both of those
match another file's is it read for a content hash to compare — the SHA-256 of
the whole file when it is under 50 MiB, or of gigabyte-spaced 1 MiB samples when
it is 50 MiB or larger. Below 50 MiB the content hash is proof of identical
content; at or above 50 MiB it is a strong candidate signal rather than proof,
because the gaps between samples are never read. The intended use is finding
duplicate downloads and duplicated directory trees on multi-terabyte ZFS servers
where reading every byte of every file is prohibitively expensive. scan
maintains a persistent SQLite database of file signatures that survives between
runs, so it can be run from cron and the reports can be generated at any time
from the most recent scan.
This README is the complete and authoritative specification.
Getting Started
make build
export SFDUPES_DATABASE="$HOME/.local/share/sfdupes/db.sqlite"
./sfdupes scan /srv
./sfdupes report > dupes.tsv
./sfdupes trees > dupetrees.tsv
scan walks one or more filesystem trees and maintains one database record per
regular file (path, size, mtime, head hash, tail hash, content hash). The
database persists between runs; a rescan only hashes files that are new or
changed, or that may have gained a duplicate since the last scan, and removes
records for files that no longer exist. report reads the database and prints
the file-level duplicates report. trees reads the same database and prints the
duplicate-tree report. A missing/invalid subcommand — or a scan invocation
with no PATH operand — prints a usage message and exits 2.
The database defaults to /var/lib/sfdupes/db.sqlite and can be placed anywhere
by setting SFDUPES_DATABASE. The intended deployment is a daily sfdupes scan
cron job, with the reporting commands run interactively whenever needed; their
results are as fresh as the last completed scan.
Install
With Go installed, this builds and installs the current main branch:
go install sneak.berlin/go/sfdupes@main
The binary goes to $(go env GOPATH)/bin, or to $GOBIN when that is set. A
binary installed this way reports its version as dev; one built from a clone
or into the Docker image carries the git tag or commit it was built from.
From a clone, make build writes the binary to ./sfdupes:
git clone https://git.eeqj.de/sneak/sfdupes.git
cd sfdupes
make build
Copy the binary to /usr/local/bin for the cron job below.
make docker builds the Docker image, tagged sfdupes, after running the tests
and the linter (see "Build"). The image runs sfdupes as root with the database
at its default path, so a bind mount of /var/lib/sfdupes keeps the database
between runs. Mount the scanned tree at the same path inside the container as on
the host; read-only is enough. The database records paths as the container sees
them, so the reports then name the host's paths.
make docker
docker run --rm -v /srv:/srv:ro -v /var/lib/sfdupes:/var/lib/sfdupes \
sfdupes scan /srv
docker run --rm -v /var/lib/sfdupes:/var/lib/sfdupes sfdupes report > dupes.tsv
Daily scan from cron
Run scan as root, so that it can read every file: a path it cannot read is
skipped with a warning and loses its database record (see "Rules for the walk").
As a file /etc/cron.d/sfdupes:
30 3 * * * root /usr/local/bin/sfdupes scan /srv 2>>/var/log/sfdupes.log || tail -n 3 /var/log/sfdupes.log
- The database is
/var/lib/sfdupes/db.sqlite, created with its directory by the first scan. To keep it elsewhere, setSFDUPES_DATABASE=/path/to/db.sqlitebefore the command on the same line. scanwrites nothing to stdout. Its stderr, appended here to/var/log/sfdupes.log, holds a plain progress line as each phase starts and then at most every 5 seconds, a warning for each path it skips, and the summary line (see "Progress" and "scanmode"). The log grows with every scan; rotate it like any other.- Skipped paths do not fail a scan: it still exits 0, and cron sends nothing. A
scan that fails, or is stopped by
SIGINTorSIGTERM, exits 1 with the reason among the last lines of the log;tailprints them, and cron mails them to root if the host can send mail. - A scan still running when the next one starts carries on. The new one fails at
once, and the lines cron mails include
sfdupes: another scan is running (lock held on /var/lib/sfdupes/db.sqlite.lock). reportandtreesneed only read access to the database (see "Database"). Under the usual umask of022the first scan creates it readable by every user, so an unprivileged user can run them against root's database.
Reading the reports
Each row of report names two copies of one file, and each row of trees two
copies of one directory tree (see "Report output format" and "Trees output
format"). In a group of copies, the path that sorts first byte by byte is
first and every other path is a dupe of it. first says nothing about which
copy is the original or the oldest; which copy to keep is your choice.
A row is a candidate, not proof:
- The reports read only the database, so they show the files as of the last scan; a file may have changed or gone since.
- A file of 50 MiB or more is compared only on samples of its content (see "Duplicate detection").
- Paths that are hard links to one file are listed as duplicates, but they share their data, so removing one frees nothing.
Compare a pair byte for byte before removing either copy. For the row
/srv/a/big.iso, /srv/b/big-copy.iso, 4294967296:
cmp /srv/a/big.iso /srv/b/big-copy.iso && echo identical
[ /srv/a/big.iso -ef /srv/b/big-copy.iso ] && echo "hard links"
cmp prints nothing and exits 0 only when every byte matches, and otherwise
reports where the files differ. The second line prints hard links when the two
paths are the same file, so removing either frees nothing.
A path holding a backslash, tab, newline or carriage return is escaped in the
reports (see "Report output format"). Undo the escapes before using it.
printf '%b' does exactly that, because every backslash in an escaped path
starts one of the four escapes. Command substitution drops trailing newlines, so
print an x after the path and remove it afterwards, or a path that ends in a
newline names a different file:
p="$(printf '%bx' '/srv/a/tab\tname.txt')"; p="${p%x}"
cmp "$p" /srv/b/tab-copy.txt
Check a trees row with diff -r, which compares the two trees file by file
and also names anything present in only one of them, such as an empty directory
or a symlink, which trees does not see.
Rationale
Duplicate finders that hash entire files do not scale to the target environment: ~10 million files and ~150 TB on possibly slow or busy disks (a ZFS pool under resilver). sfdupes spends disk I/O only on files whose size at least one other file shares, since a size-unique file cannot be a duplicate. Of those, a file under 10 MiB is read in full; a larger one has its cheap end windows read first, and is read for a content hash only when its size and both end windows match another file's — the whole file below 50 MiB, but only gigabyte-spaced samples at or above 50 MiB, so the largest files are never read in full. This keeps a full-filesystem sweep tractable, and the signatures are kept in a persistent database, so the expensive filesystem pass is incremental: a rescan re-hashes only files whose recorded mtime or size changed, plus — for its content hash — a file of 10 MiB or more whose size and end windows have come to match another file's. All analysis happens offline from the database alone. The end goal is not individual files but whole duplicated trees — duplicate extractions, duplicate downloads, copied project trees — which an operator can consider removing as a unit.
Design
Goals, in order:
- Find whole duplicate trees, not just files. The end goal is to identify places where the exact same set of files and directories exists at two or more paths (duplicate extractions, duplicate downloads, copied project trees), so the operator can consider removing an entire subtree at once. File-level duplicate detection is the foundation; tree-level detection is built on top of it.
- Spend I/O in proportion to duplicate likelihood. Only files whose size
at least one other file shares are read at all — a size-unique file cannot
be a duplicate. Those are compared by the ladder in "Duplicate detection"
below: a file under 10 MiB is hashed in full, while a larger file is gated
on cheap 64 KiB end windows first, and gets a content hash only when its
size and both end windows match another file's. That hash reads the whole
file below 50 MiB but only gigabyte-spaced 1 MiB samples at or above it, so
the very largest files are still never read in full. Scale target: tens of
millions of files, ~150 TB filesystem, possibly slow or busy disks (ZFS pool
under resilver). Holding one small record (path, size, mtime) per file in
memory during a scan is acceptable; holding every file's hashes is not (they
stay in the database). The reporting commands do not hold every file's
hashes either:
reportlets SQLite group and order the records and writes each row as it reads it, so its memory does not grow with the database, andtreesreads the records in path order and keeps each directory's path, digest and totals, plus the hashes of only the files in the directories holding the record being read, so its memory grows with the number of directories and with the size of the largest directory. - Scan incrementally, analyze offline. The expensive filesystem scan
maintains a persistent database; an unchanged file is never read again on a
rescan, except to compute its content hash once a file of 10 MiB or more
comes to match another on size and both end windows. All analysis (
report,trees) works from the database alone and must never touch the scanned filesystem again.scanis designed to be cronned; the reports run at any time against the last completed scan. - Clean stream separation. Everything on stdout is machine-readable data. All progress, warnings, summaries, and help and usage text go to stderr. Never mix them.
Constraints
- Language: Go (module
sneak.berlin/go/sfdupes). Binary name:sfdupes. - Dependencies: standard library,
github.com/spf13/cobrafor the CLI, one progress-bar library (github.com/schollz/progressbar/v3),golang.org/x/termto tell whether stderr is a terminal, one SQLite driver (modernc.org/sqlite, pure Go, so builds keep cgo disabled), andgolang.org/x/sysforflock(2)(the scan lock, see "Database").github.com/spf13/viperis permitted if configuration-file support is ever needed, but is not currently used. No other third-party deps. - Cross-compilation is not a concern. Builds run with cgo disabled (the
MakefileexportsCGO_ENABLED=0); the code must remain pure Go. - Analysis modes (
report,trees) must be deterministic: identical database contents, identical output, regardless of the order in which records were inserted.
Subcommands
Three subcommands, all implemented:
scan— walk the filesystem and synchronize the database: one signature record per regular file.report— file-level duplicate report from the database.trees— tree-level duplicate report: reconstruct the directory hierarchy from the database records, compute a Merkle-style digest per directory, and report maximal groups of identical trees.
sfdupes scan [--workers N] [-x] PATH...
sfdupes report > dupes.tsv
sfdupes trees > dupetrees.tsv
sfdupes --version
sfdupes [command] --help
--workers N sets the size of each scan worker pool (default: the number of
CPUs), and -x (--one-file-system) keeps the walk of each operand on that
operand's filesystem; both are described under "scan mode".
sfdupes --version (or -v) prints one line, sfdupes VERSION, to stdout and
exits 0, writing nothing to stderr. -h or --help, alone or after a
subcommand, prints the help text to stderr and exits 0, writing nothing to
stdout.
Database
All three subcommands operate on a single SQLite database file:
-
Location: the value of the
SFDUPES_DATABASEenvironment variable when set and non-empty, otherwise/var/lib/sfdupes/db.sqlite. There is no command-line flag. The path names the file exactly, whatever characters it holds (?,#and%included); a relative path is relative to the working directory. -
scancreates the database (and its parent directory) on first use.reportandtreesrequire an existing database; a missing database file is a fatal error (exit 1) telling the user to runscanfirst. -
Only one
scanruns against a database at a time. For its whole run,scanholds an exclusiveflock(2)lock on a lock file beside the database, named by appending.lockto the database path (/var/lib/sfdupes/db.sqlite.lockby default), taken before it walks the filesystem or opens the database. A secondscanagainst the same database does not wait: it fails at once with a one-line error naming the lock file and exits 1, without walking anything or opening the database, and the running scan carries on. The lock file is created on first use, open to its owner only, and left in place: a leftover file blocks nothing, because the lock ends with the process holding it however it ends, a fatal error or an interrupt included, and deleting the file while a scan runs would let a second scan start.reportandtreesnever take the lock, so they run during a scan. -
While
scanruns, the database is in WAL journal mode with a busy timeout, so running a report while a cronscanis in progress is safe. The filesystem is authoritative; the database is an eventually-consistent reflection of it. Hashed records are committed in batched transactions while the scan is still running (keeping the WAL small and letting concurrent reports observe progress), so a report may see a scan's changes partially applied, and a scan that dies partway leaves a valid database holding every batch committed so far (an interrupted scan also commits the batch in progress, see "Error handling and exit codes"); the next scan skips those records and converges toward the filesystem. -
scanswitches the database back to rollback-journal mode when it closes it, so between scans the database file alone holds the whole database. Each switch needs the database to itself: ascanthat starts while a report is still reading waits for it up to the 10-second busy timeout, then fails; ascanthat ends while a report has the database open warns and leaves the database in WAL mode until the next scan.reportwrites each row as it reads it, so it is still reading while its output is paused (a pager, a stalled pipe), and ascanstarted then fails after the busy timeout. -
reportandtreesopen the database read-only and need only read access to the database file, and no write access to its directory. While the database is in WAL mode they also read the-waland-shmfiles beside it, which SQLite creates with the database file's permissions. -
Schema (
PRAGMA user_versionis the schema version, currently 1; a database with any other version is a fatal error.scancreates the schema and sets the version in one transaction, so a first scan stopped while doing so leaves an empty database the next scan sets up. A database at version 0 that already has afilestable was therefore not made by sfdupes; every subcommand refuses it with an error telling the user to remove the file and rescan):CREATE TABLE files ( path BLOB PRIMARY KEY, -- absolute path, raw bytes size INTEGER NOT NULL, -- bytes, from lstat mtime INTEGER NOT NULL, -- Unix seconds, from lstat head TEXT NOT NULL, -- lowercase-hex SHA-256; first 64 KiB, or whole file under 10 MiB tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples ) WITHOUT ROWID; CREATE INDEX files_signature ON files (size, head, tail, content);Paths are stored as BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
mtimeis used only for change detection; it is not part of the duplicate key. For a file under 10 MiBhead,tail, andcontentall hold the whole-file hash (that range is hashed in full, with no end windows); for a larger fileheadandtailhold the first- and last-64 KiB hashes andcontentthe whole-file or sampled hash. All three are empty strings when the file has never been hashed because its size was unique as of the last scan that covered it. For a file of 10 MiB or more,contentstays empty until the content phase of a scan (see "scanmode" below) has read the file. A record with an emptycontentis never part of a duplicate group, though it still defines the file for tree reconstruction. Thefiles_signatureindex lets SQLite group the records by signature forreportwithout sorting the whole table.
Duplicate detection
Two files are duplicates only when they agree on every rung of this ladder; a
mismatch at any rung means they are not duplicates. scan stores each file's
hashes, and report and trees group files by the whole signature — size,
head, tail, and content — so the grouping is exactly this ladder applied
across everything scanned into the database, even across separate scans.
- Size. Files of different sizes are never compared. Only files whose size at least one other file shares are hashed at all.
- Under 10 MiB: whole file. A file smaller than 10 MiB is hashed in full
and compared directly, with no separate end-window step — small files are
cheap to read to the last byte, and doing so makes the comparison exact.
head,tail, andcontentall hold this whole-file SHA-256, so such a file's signature is decided entirely by its size and its content. - 10 MiB and above: head and tail. For a larger file, the SHA-256 of the
first 64 KiB (
head) and of the last 64 KiB (tail) are a cheap gate that eliminates most same-size pairs before any bulk reading: the content hash of the next two rungs is computed only for a file whose size,head, andtailmatch another file's, whether that file is scanned in the same run or stored by an earlier scan. A stored file that first gains such a match in a later scan gets its content hash then; until it has one, itscontentis empty and it is not a duplicate. At 10 MiB and above the two windows never overlap. - 10 MiB and above, content below 50 MiB. The SHA-256 of the entire file. Agreement here is proof of identical content (barring a SHA-256 collision).
- 10 MiB and above, content 50 MiB and above. A sampled SHA-256: the 1 MiB window at each gigabyte-aligned offset (0, 1 GiB, 2 GiB, … while inside the file, the final window truncated at end of file) is fed, in order, into one hash. This is deliberately probabilistic — the gaps between samples are never read, so two large files that agree on every sample are reported as duplicates without being read in full. It is the price of never reading a 150 GB file end to end. Because size is already part of the signature, only equal-size files reach this rung, so their sample boundaries always align.
head, tail, and content are one column each. A file below 10 MiB and one
at or above it never share a size, and neither do a file below 50 MiB and one at
or above it, so a stored value is never ambiguous between the whole-file,
end-window, and sampled forms.
scan mode
scan requires one or more PATH operands naming the trees to scan. There is
no default path; invoking scan with no operand is a usage error (usage message
on stderr, exit 2). An operand may be a directory or a regular file; an operand
that does not exist is a fatal error (exit 1). Because database records persist
between runs and are keyed by absolute path, each operand is resolved to an
absolute, lexically cleaned path (symlinks are not resolved) before walking, so
results do not depend on the working directory. All operands belong to a single
scan and are enumerated concurrently: every operand seeds the shared walk worker
pool. Overlapping operands are harmless — an operand that duplicates another or
lies under another is dropped before walking, so every file is reached exactly
once and produces one database record.
An operand that is a symlink (never followed, not even as an operand), socket,
FIFO, or device node, or a directory named .zfs, is not scanned. scan prints
a one-line warning naming the path and what it is, counts it as skipped, and
drops it from the scanned operands before reading the database. Another operand
beneath it is still scanned. The records stored beneath it are not deleted: they
are treated like any other record outside the scanned operands, including the
content-phase exception below. If it lies under another operand, they are under
that operand instead, and are deleted like any other record there that this scan
did not verify. This is not an error: a scan whose every operand is dropped
walks nothing and exits 0.
scan synchronizes the database with the filesystem state under the scanned
operands:
- Only a file whose size at least one other file shares is ever read: a
size-unique file cannot be a duplicate, so it is recorded without hashes
(
head,tail, andcontentempty). The size census covers every file walked this scan plus every database record outside the scanned operands, so a possible duplicate of a separately scanned tree is still recognized. - A file not yet in the database is inserted: hashed when its size is shared, without hashes otherwise.
- A file already in the database is skipped without reading its contents
when its lstat size equals the recorded size and its lstat mtime is not newer
than the recorded mtime. This is what makes a daily rescan cheap. Exception:
an unchanged file whose record lacks hashes is hashed — and its record updated
— once its size becomes shared, so hashing deferred by size-uniqueness happens
as soon as it could matter. Likewise, an unchanged file of 10 MiB or more
whose record has no
contenthash is read for one by the content phase below once its size,head, andtailmatch another record's. - A file whose mtime is newer than recorded, or whose size differs, is processed as if new: re-hashed, or recorded without hashes, per the shared-size rule.
- A database record whose path lies under one of the scanned operands but was not successfully processed this run is deleted. This removes records for deleted files. It also removes records for paths that failed to stat or hash this run: the database only ever contains signatures verified by the most recent scan that covered them (a subsequent successful scan re-adds such files). A failure in the content phase below removes nothing: the record is left as it is.
- Database records outside the scanned operands are untouched, so disjoint trees
can be scanned on different schedules into the same database. The one
exception is the content phase below: a stored file of 10 MiB or more without
a
contenthash is read for one, wherever it lies, once its size,head, andtailmatch another record's. If that file is gone or has changed since its record was written, the record is left as it is.
scan runs four sequential phases over the whole scan. Parallelism lives
inside each phase; batched database writes begin during the hash phase:
- walk + stat — enumerate the trees under all
PATHoperands concurrently with the walk worker pool: every operand seeds the shared queue, and each worker reads one directory at a time, handing discovered subdirectories back to the queue and runninglstaton each regular file as it is discovered (while the directory's metadata is still hot). Sequential directory enumeration is metadata-latency-bound and takes hours at tens of millions of files; per-directory parallelism is what makes the walk tractable on large or busy pools. The walk builds the size census and resolves unchanged already-hashed files on the fly; every other file is carried to the hash phase as a (path, size, mtime) record. - hash — with the census complete, each carried file's size decides its
fate. Size-unique files are never read: new or changed ones are recorded
without hashes in the update phase, unchanged unhashed ones simply keep
their records. Every file with a shared size is hashed by the worker pool as
described in "Duplicate detection" above: a file under 10 MiB in full, which
gives its
head,tail, andcontentalike, and a larger file only in its end windows, which give itsheadandtail; its content hash is left to the content phase. Zero-length files have constant hashes and are never opened. Files are hashed in inode order (minimizing seeks on spinning disks), and paths that are hard links to the same inode are read once, all sharing the one result — a hard-link backup farm costs one read per inode, not per path. The phase total counts actual reads, so progress and ETA are meaningful. Completed records are committed in batched transactions while hashing runs, so a scan interrupted after hours keeps everything hashed so far and the next scan resumes cheaply, skipping records already written. - update — commit the final partial batch, the hash-less records for size-unique new and changed files, and the deletions for records the scan did not verify (vanished files, plus paths that failed to stat or hash).
- content — find every record of 10 MiB or more without a
contenthash whose size,head, andtailequal another record's, anywhere in the database: records from this scan and records stored by earlier scans, inside or outside the scanned operands. SQLite finds them, so only the records to be read are kept in memory, never every file's hashes. Every record sharing their size,head, andtail, including one that already has acontenthash, has its file checked withlstatfirst. A file that is gone, is no longer a regular file, or has changed (a different size, or an mtime newer than recorded) keeps its record as it is and does not count as a match for the others. Any otherlstaterror is warned about and counted as skipped, with the same result. If such a record has nocontenthash, it stays out of duplicate groups; if it has one, it is still reported until a scan covering its own tree updates or removes it. The files that pass and have nocontenthash are read only if at least two of those records pass, so a file whose only matches are stale costs no read; a file that already has acontenthash is never read again. They are read by a worker pool as in the hash phase, in inode order and once per inode, and their content hashes are committed in batches. A failed read is warned about and counted as skipped; its record keeps an emptycontent, so it is not a duplicate, and a later scan tries again.
Rules for the walk:
- Only regular files. Skip directories, symlinks (do not follow, including
symlink operands), sockets, FIFOs, and device nodes. An operand that is a
symlink, socket, FIFO, or device node is dropped as described in "
scanmode" above. - Never descend into a directory named
.zfs(ZFS snapshot pseudo-dirs; walking them would list every file once per snapshot), not even when it is an operand; such an operand is dropped the same way. - Filesystem boundaries are crossed by default. With
-x(long form--one-file-system, following the GNUdu/rsyncconvention), never descend into a directory on a different filesystem than itsPATHoperand; each operand is bounded by its own filesystem. - On any per-path error (permission denied, file vanished between passes, unreadable): print a one-line warning to stderr, skip the path, and continue. Per-file errors never abort the run; the final summary reports how many were skipped. As specified above, a skipped path that has a database record from an earlier scan loses that record, unless it failed only in the content phase, or is an operand dropped before the database was read that lies under no other operand; an unreadable directory subtree likewise loses its records (accepted: the database mirrors what the latest scan could actually verify).
Concurrency: the walk phase (which also stats files), the hash phase, and the
content phase each use a worker pool of --workers workers (default
runtime.NumCPU()); the walk parallelizes across directories, hashing across
files. --workers must be at least 1: a smaller value is a usage error,
reported in one line on stderr with exit 2 before anything is scanned. All three
phases are seek-bound on spinning disks, so raising --workers well past the
core count can help on pools with many spindles. The main goroutine owns
partitioning, database writes, and progress rendering; progress display must
never block the workers.
scan writes nothing to stdout. The summary line on stderr reports the files
seen this run broken down by disposition, plus skips:
scan: 123400 files seen (1200 added, 34 updated, 56 removed, 122166 unchanged), 3 skipped
(removed counts deleted database records, which are not part of the files-seen
total.)
report mode
report reads every record from the database and takes no positional arguments.
report must never touch the filesystem being analyzed. It does not stat,
open, or otherwise access any path that appears in the records; its only I/O is
reading the database, writing stdout/stderr, and the temporary file SQLite sorts
in when the duplicate rows do not fit in memory. SQLite puts that file in
$SQLITE_TMPDIR or $TMPDIR when set, otherwise in /var/tmp (or /tmp), and
deletes it as soon as it has opened it. report must produce identical output
whether or not the scanned filesystem is still mounted.
Processing:
- Records without a
contenthash (see "Database" above) are excluded: their content is unknown, so they are never reported as duplicates. - Group the remaining records by the key
(size, head, tail, content). - Every group with two or more paths is a duplicate group.
- Within each group, sort paths lexicographically (byte order). The first path
is the group's
first; every other path is adupe. - Order groups by size descending (biggest reclaimable space first), tie-broken
by
firstpath ascending. Output must be fully deterministic for a given database state.
Report output format
TSV on stdout: a header line, then one row per duplicate file (N-1 rows for a group of N):
first dupe size
/srv/a/big.iso /srv/b/big-copy.iso 4294967296
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296
Paths are raw bytes and may hold any byte except NUL, so the path columns
(first and dupe) are escaped to keep every row one line of tab-separated
fields: a backslash is written as \\, a tab as \t, a newline as \n, and a
carriage return as \r. Every other byte is written unchanged, including bytes
that are not valid UTF-8. Undoing those four escapes gives back the stored path.
Grouping and ordering use the stored path, not the escaped one. The warnings
scan prints on stderr are escaped the same way, so each warning is one line.
Summary to stderr: records read, number of duplicate groups, number of dupe
files, and total reclaimable bytes (sum of size over all dupe rows) in human
units.
trees mode
trees reads the same database as report (no positional arguments) and
reports entire duplicate directory trees: directories under which the exact
same set of relative paths exists with the exact same file signatures.
trees must never touch the filesystem being analyzed — the same rule as
report. The directory hierarchy is reconstructed purely from the paths in the
records, split on /.
Definitions:
- A file's signature is
(size, head, tail, content)— mtime is informational and excluded. A record without acontenthash has unknown content: its signature is treated as unique to that file, so a tree containing such a file never compares equal to any other tree. - A directory's digest is a SHA-256 Merkle digest computed bottom-up: serialize the directory's child entries — for a file child, its name and signature; for a subdirectory child, its name and that subdirectory's digest — sort the serialized entries byte-lexicographically, and hash the concatenation. Names are part of the digest: two trees whose files differ only in name are not duplicates.
- Two directories are duplicate trees when their digests are equal. Equal digests imply equal recursive file count and equal total byte size.
Known limitation (accepted): hard-linked paths are reported as duplicates by
report and count toward duplicate trees — their content is genuinely identical
— even though they share storage, so removing one reclaims no space. Inode
identity is used during the scan to avoid redundant reads but is not persisted
in the database.
Known limitation (accepted): only regular files that appear in the database define a tree. Empty directories are invisible, and a file skipped during the scan (e.g. permission error) in one copy but not the other will make otherwise-identical trees compare as different.
Processing:
- Build the hierarchy, compute every directory's digest, and group directories by digest. Every group with two or more directories is a duplicate-tree group.
- Report only maximal trees. A group is suppressed when its members' parents are pairwise distinct directories that all share a single digest — such a group is wholly implied by its parents' (or a further ancestor's) group. Groups containing sibling directories, or members whose parents differ, are always reported.
- Within each group, sort paths lexicographically (byte order); the first path
is
first, every other path is adupe. - Order groups by total tree size descending, tie-broken by
firstpath ascending. Output must be fully deterministic for a given input.
Trees output format
TSV on stdout: a header line, then one row per duplicate tree (N-1 rows for a
group of N). files is the recursive regular-file count of one copy of the
tree; size is the recursive total byte size of one copy:
first dupe files size
/srv/a/project /srv/backup/project 3417 104857600
The first and dupe paths are escaped as described under "Report output
format". The root directory's path is /.
Summary to stderr: records read, number of duplicate-tree groups, number of dupe
trees, and total reclaimable bytes (sum of size over all dupe rows) in human
units.
Progress
Use the progress-bar library for all scan progress; rendering in the style of
pv is the model. All progress goes to stderr.
Each phase gets its own display, rendered the moment the phase starts — a scan
must never look hung. Loading the existing-record index (load) and the walk
have no known totals while running: show a live count, rate, and elapsed time
(spinner-style, no percentage or ETA). The content phase's display (content)
starts the same way, counting the records checked while SQLite finds the files
to read and lstat checks them, then shows a bar once reading starts. The hash
and update phases, and the content phase's reads, have exact totals — only files
that actually need hashing appear in the hash and content totals, so their ETAs
are meaningful. Required elements for the bars with known totals:
- elapsed time
- estimated time remaining
- a
[m/n] x%display (items processed / total items, percent) - current rate (items/s)
Example shape (exact layout is flexible, content is not):
hash: [12345/98765] 12% |████ | 92 files/s elapsed 2:32 eta 17:54
Additional requirements:
- When stderr is not a terminal (a pipe, a file,
/dev/null), do not emit ANSI redraws: print a plain one-line progress update the moment each phase starts, then no more often than every 5 seconds. - Progress updates are driven from the main goroutine and must be non-blocking with respect to the worker pool. On a terminal the spinner-style displays also redraw on their own several times a second, so their count and elapsed time stay current while a phase waits for its next item.
- A warning printed during a phase always lands on a line of its own, never inside the progress display.
- A bar whose phase stops short of its total, as an interrupted one does, is left as last drawn rather than filled up.
reportandtreesmodes need no progress display, only their stderr summaries.
Error handling and exit codes
0: success, even if individual files were skipped with warnings.1: fatal error (e.g., aPATHoperand does not exist, anotherscanis already running against the same database, the database cannot be created/opened/read/written, a missing database forreport/trees, stdout write failure), or ascanstopped bySIGINTorSIGTERM(see below).2: usage error (includingscanwith noPATHoperand,scanwith--workersbelow 1, andreport/treeswith any positional argument).
A stdout write failure, such as a full disk, is reported in one line on stderr and exits 1. Two cases never reach sfdupes as a failed write:
- When the reader of a stdout pipe exits early, as in
sfdupes report | head, the next write ends sfdupes withSIGPIPE, quietly and without a summary, the waycatorsortend. The shell reports the signal (status 141 in most shells), not exit 1. - When stdout is closed outright (
sfdupes report >&-), the Go runtime opens/dev/nullin its place before sfdupes starts, so the output is discarded and the run succeeds, as with> /dev/null.
scan stops cleanly on SIGINT (Ctrl-C) or SIGTERM. Its workers stop taking
work, each finishing at most the directory listing or file it is reading; the
progress display is finished; and the records it has hashed but not yet
committed are committed, so the next scan does not hash them again. Apart from
that commit it starts no further writes or deletions: records are deleted only
after a complete walk, so those under paths an interrupted walk never reached
are kept. The database is closed and the lock released as on any other exit, the
line scan: interrupted after N files goes to stderr, N being the number of
files the walk reached, and the exit code is 1. The next scan skips the records
already written and converges as usual.
After the first signal scan stops catching them, so a second one ends it at
once, as an uncaught signal does: the records not yet committed are lost, and
the database is left valid, as when any scan dies (see "Database"). A SIGINT
that scan inherits as ignored, as a script's background job does, stays
ignored.
Entrypoints
This repository adheres to the
Scripts to Rule Them All
standard: the normalized executables in script/ are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call them.
Every script is POSIX sh, resolves the repository root itself so it can be run
from any working directory, and may be invoked directly. The provided
entrypoints are:
script/bootstrap— install everything needed to build and develop this repository, idempotently, assuming nothing is present.git,make, andgocome from the first of nix, apt, brew, or apk found on the host, and are presence-checked only.golangci-lintand prettier are deliberately not installed: they run in Docker (seescript/lintandscript/fmt) and never from a host install, so there is no host copy to drift from the pin. A missingdockeris warned about rather than installed or treated as fatal — everything except linting and formatting works without it. Ends withgo mod download.script/setup— make a fresh clone ready for development: runsscript/bootstrap, thenscript/install-precommit.script/projectname— print this project's name (sfdupes). Scripts that need the name call it, so they stay identical across repositories.script/test— run the test suite with a 30-second timeout and coverage enabled, rerunning verbosely on failure so the logs show which test failed.script/lint— run the linter. It buildsDockerfile.lint, which copies the repository into the digest-pinnedgolangci/golangci-lintimage and runsgolangci-lint config verifyandgolangci-lint runas build steps, so a successful build is a clean lint. That exit status is all it produces, so it runs with--output=type=cacheonlyand writes no image; a run leaves only build cache. The linter is never run on the host, which makes a workingdockerthe one prerequisite for linting — and therefore formake checkand the pre-commit hook. Offline machines: the gate steps themselves make no network calls.golangci-lint rundoes not, and neither doesgolangci-lint config verify— it validates against a schema the pinned binary embeds, measured under--network noneto both pass a valid config and reject an invalid one. The build around them does.Dockerfile.lintrunsgo mod downloadbefore the gates and this module has external dependencies, so a first lint on a machine with a cold BuildKit cache reaches the network there (as well as pulling the pinned image); under--network noneit fails at that step, before any gate. That layer sits above the gates and stays cached, so once it is warmscript/lint— and with itmake check— runs entirely offline, untilgo.modorgo.sumchanges and the download layer goes cold again. Because the daemon only ever sees a build context, this works when the docker daemon is remote and bind mounts are impossible.script/fmt— format in place: the Go sources withgofmt -s -w, and every Markdown file with prettier, at the settings in.prettierrc(4-space indents, prose wrapped at 80 columns). prettier is pinned by hash throughpackage.jsonandyarn.lockand never installed on the host: this builds theDockerfile'sprettierstage, a digest-pinned node image into whichyarn install --frozen-lockfileinstalls it, taggedsfdupes-prettier, and runs that with the repository mounted, as the calling user. Needsdocker, and because of the mount, unlikescript/lint, a local docker daemon.script/fmt-check— the read-only counterpart ofscript/fmt, with the repository mounted read-only: prints any unformatted file and exits non-zero instead of writing. gofmt and prettier both run every time, and each names itself when it fails. TheDockerfileruns the same two checks as gates: the gofmt check in its lint stage, prettier in itsmarkdownstage.script/check— runscript/test,script/lint, andscript/fmt-check, in that order. Modifies nothing. Needsdocker, becausescript/lintandscript/fmt-checkdo.script/docker— build the Docker image, tagged with the name fromscript/projectname. TheDockerfileruns the gates as build steps, so this is also the check a developer or reviewer runs by hand.script/cibuild— build the Docker image untagged. This is what the Gitea workflow runs on push; because the gates run as build steps, a successful build implies the repository is green.script/precommit— run by the git pre-commit hook:go mod tidymust be a no-op (a resulting change togo.modorgo.sumfails the commit), thenscript/check.script/install-precommit— install the git pre-commit hook that runsscript/precommit. The hook is written to the common git directory, so the main checkout and every worktree share it.script/verify-lint-image-pin— fail unless thegolangci/golangci-lintreference inDockerfile.lintand the one in theDockerfilelint stage are the same image at the same digest, naming both if not. The linter is pinned in those two files and nothing else keeps them in sync, so a bump applied to one alone would leavemake lintand theDockerfile's fail-fast lint stage checking the same tree against different rulesets, both green. The guard restates neither pin — a third copy would be the same drift one file further out — and runs as a gate in both files, somake lint,make checkandmake dockerall catch it.
script/verify-linter-pin used to live here. It compared a linter binary
against a version pin in script/bootstrap, and both of its subjects are gone:
no linter binary is copied between build stages any more, and bootstrap pins no
version because it installs no linter. The drift it existed to catch has moved
from binary-versus-pin to pin-versus-pin, which is what
script/verify-lint-image-pin above checks.
script/lint, script/docker and script/cibuild all pass a freshly computed
CHECK_EPOCH build argument, and the gate steps in Dockerfile.lint and
Dockerfile reference it. Without that, an unchanged tree lets Docker serve the
gate layers from cache and the build exits 0 having executed no tests and no
lint — a green it never earned, and one this repository has produced twice.
CHECK_EPOCH invalidates the gate layers on every run while leaving the pinned
base images and the dependency layers cached. script/lint's value carries the
process id as well as the epoch, because two lint runs land inside the same
second easily and a bare epoch would cache the second one.
Build
The script/ entrypoints above are where the implementations live; the
Makefile targets are shims onto them, except build, which carries the
compile recipe:
make/make build— build thesfdupesbinary (cgo disabled); building is the default target.make bootstrap— install the build and development dependencies.make setup— prepare a fresh clone:bootstrapplus the pre-commit hook.make test— run the test suite (30-second timeout; reruns with-von failure).make lint— rungolangci-lintwith the repo config, in Docker (seescript/lint); requiresdocker.make fmt/make fmt-check— format the Go sources and the Markdown / verify formatting without writing; requiresdocker, for prettier (seescript/fmt).make check—test,lint, andfmt-check; modifies nothing. Requiresdocker, vialintandfmt-check.make docker— build the Docker image, which runs the gates as build stages.make hooks— install the pre-commit hook.make clean— remove the binary.
Definition of done
All of the following, run in this directory, must pass:
-
make checkpasses (tests, lint,gofmt, prettier). -
make dockersucceeds. -
Smoke test — create a throwaway tree in a temp dir (never test against real data):
d=$(mktemp -d) export SFDUPES_DATABASE="$(mktemp -d)/db.sqlite" mkdir -p "$d/a" "$d/b" head -c 2000 /dev/urandom > "$d/a/one.bin" cp "$d/a/one.bin" "$d/b/copy.bin" cp "$d/a/one.bin" "$d/b/copy2.bin" head -c 2000 /dev/urandom > "$d/a/unique.bin" # same size, different content printf 'x' > "$d/tiny1"; printf 'x' > "$d/tiny2" # 1-byte duplicates printf 'y' > "$d/tiny3" # 1-byte non-duplicate : > "$d/empty1"; : > "$d/empty2" # empty duplicates # duplicate trees: t1 and t2 are identical; t3 differs by one filename mkdir -p "$d/t1/sub" "$d/t2/sub" "$d/t3/sub" head -c 3000 /dev/urandom > "$d/t1/f1" head -c 100 /dev/urandom > "$d/t1/sub/f2" cp "$d/t1/f1" "$d/t2/f1" cp "$d/t1/sub/f2" "$d/t2/sub/f2" cp "$d/t1/f1" "$d/t3/f1" cp "$d/t1/sub/f2" "$d/t3/sub/f2renamed" ./sfdupes scan "$d" ./sfdupes report ./sfdupes trees # incremental behavior (scan a subtree; records elsewhere persist): ./sfdupes scan "$d/a" # everything unchanged, nothing hashed printf 'z' >> "$d/a/one.bin" # modify: next scan re-hashes it rm "$d/a/unique.bin" # delete: next scan removes its record ./sfdupes scan "$d/a" # 1 updated, 1 removed ./sfdupes report(The database lives in a temp directory of its own: inside
$d, the scan would record it, and its empty lock file would join theempty1/empty2group.)Expected from the first
report:one.bin/copy.bin/copy2.binform one group (two dupe rows,firstis the lexicographically smallest path);t1/f1/t2/f1/t3/f1form one group;t1/sub/f2/t2/sub/f2/t3/sub/f2renamedform one group;tiny1/tiny2pair;empty1/empty2pair;unique.binandtiny3appear nowhere; groups ordered by size descending.Expected from
trees: exactly one row —first$d/t1,dupe$d/t2, 2 files, 3100 bytes.$d/t1/subvs$d/t2/subis suppressed as non-maximal (implied by thet1/t2group), andt3appears nowhere (its file set differs by name).Expected from the second
report(after the modify/delete rescan):one.binhas left its group (its content changed), socopy.bin/copy2.binremain as one pair, andunique.binis gone from the database.The test suite automates this scenario (see
scan_test.go), plus a negative check:reportandtreesoperate on the database alone and never touch the scanned filesystem.
TODO
Tracked in TODO.md.
Non-goals
- No byte-for-byte compare, and no deletion or linking of duplicates. Files that match are compared by a SHA-256 of the whole file below 50 MiB, and only by samples at 50 MiB and over. The reports are advisory; acting on them is the user's job.
- No persistence beyond the SQLite database described above; no export/import formats.
- No daemon or filesystem watcher; scheduling rescans is cron's job.
License
MIT. See LICENSE.