Author SHA1 Message Date
clawbot 9776984e19 Test that scan refuses another schema version (closes #64)
check / check (push) Waiting to run
The version-mismatch test only opened its database through
openReportDatabase, so nothing exercised the branch of initSchema that
stops scan on a database stamped with an unknown schema version. The
test now opens the same database through openScanDatabase too and
requires errSchemaVersion, and is renamed to match the unversioned-file
test beside it, which also covers both paths.

Model: opus-5-5
2026-10-04 10:14:59 +00:00
clawbot 7ff0dbb6e1 Test the -x filesystem-boundary rejection branch (closes #17)
check / check (push) Waiting to run
No test ever ran the part of subdirJob that refuses to descend across
a filesystem boundary under -x. The new tests call subdirJob directly
with a made-up device for the operand, so no second real filesystem is
needed. They cover refusal across a boundary, skipping the check when
the operand's device is unknown, the warning when the subdirectory
cannot be statted, and crossing by default without -x. When descent is
accepted, the whole returned job is compared, so losing the operand's
device on the way down fails a test. scan.go is unchanged.

Model: opus-4-8 (implementation); opus-5-5 (rebase)
2026-10-04 12:13:23 +02:00
clawbot bccc14ffef Refuse an unversioned database that already has a files table (closes #11)
check / check (push) Waiting to run
A database at user_version 0 that already has a files table was made
by something else: scan used to run its CREATE TABLE on it and fail
with a raw SQLite error, and report and trees gave only a bare version
mismatch. All three now refuse such a database with the schema-version
error telling the operator to remove the file and rescan. scan creates
the table and index and sets the version in one transaction, so a first
scan stopped partway leaves an empty database the next scan sets up,
never a files table at version 0. A genuinely empty database is
unchanged.

Model: opus-4-8 (implementation); opus-5-5 (rebase)
2026-10-04 12:01:36 +02:00
clawbot 33f8607e17 Document install, a daily cron scan and reading the reports (closes #54)
check / check (push) Waiting to run
Getting Started gains three parts: installing with go install, from a
clone or as the Docker image; a crontab line for a daily root scan,
with where its stderr and failures go; and how to read the two reports,
why a row is a candidate rather than proof, and how to check a pair
with cmp before removing anything. The usage block names --help, and
the text below it --workers and -x.

The install line uses @main, not @latest: @latest resolves to the
v0.0.1 tag, which predates the database.

Model: opus-5-5
2026-10-04 11:47:20 +02:00
clawbot 2eeba3df6f Keep the module cache out of the build stage's chown (closes #43)
check / check (push) Waiting to run
The build stage handed /src and the whole Go module cache to the
unprivileged user with chown -R, in a layer that re-ran on every source
change. On this host that step took from about 80 s to over ten minutes,
depending on load.

The module cache now sits at /go/pkg/mod and stays root's: bootstrap
fills it as root. The sources are copied with --chown. A small chown,
cached with bootstrap, hands builder the /src directory itself, the
module cache's cache/download directory (where make build saves its
lookup of this module's own version), and the telemetry files root's go
commands left in its home. Tests and the build still run as builder.

Model: opus-5-5
2026-10-04 10:01:28 +02:00
clawbot cb5dda4f45 Print --version to stdout (closes #15)
check / check (push) Waiting to run
Cobra's built-in version flag prints through the writer that carries
help and usage, which is stderr here. The root command now defines
its own -v/--version flag and prints one line, "sfdupes VERSION", to
stdout; a failed write is a fatal error (exit 1). Help and usage stay
on stderr. README documents --version and --help, their streams and
exit codes. Tests cover both flags and the failed write.

Model: opus-5-5
2026-10-04 09:47:28 +02:00
clawbot 4847882e46 Stop scan cleanly on SIGINT or SIGTERM (closes #5)
check / check (push) Waiting to run
A first SIGINT or SIGTERM cancels the scan. It commits the hashed
records still in its batch, with a context that is not cancelled for
that one write, and starts no other write or deletion; deletions need
a complete walk, so records under paths an interrupted walk never
reached are kept. A batch whose commit failed is now kept for that
final commit instead of dropped. The progress display is finished (a
bar stopped short is no longer filled up), `scan: interrupted after N
files` goes to stderr, and the exit code is 1. A second signal ends the
process at once. A SIGINT inherited as ignored stays ignored.

Model: opus-5-5
2026-10-04 09:30:38 +02:00
clawbot 33cf3dd29a Stream report and trees instead of loading every record (closes #14)
check / check (push) Waiting to run
report now has SQLite group the records and put the rows in report
order, helped by a new files_signature index on (size, head, tail,
content), and writes each row as it reads it. trees reads the records
in path order, where all the paths under a directory come together, so
it computes each directory's digest as soon as the stream leaves it and
keeps only its path, parent, digest and totals. Output is unchanged.

The tests that called the removed in-memory grouping functions now group
records stored in a database. New tests check that both commands give
the same output whatever order the records were inserted in, and that a
stdout failure partway through a long report is reported as one.

Model: opus-5-5
2026-10-04 06:47:25 +02:00
clawbot 9abf81535a Print progress at once off a terminal, keep warnings out of redraws (closes #13)
check / check (push) Successful in 1m56s
When stderr is not a terminal, each phase prints its zero-state line
as it starts instead of after its first item. stderrIsTTY uses
term.IsTerminal from golang.org/x/term, now a direct dependency, so
/dev/null is no longer taken for a terminal.

A spinner keeps the library's background redraw, so its count and
elapsed time stay current while a phase waits for its next item. A
warning printed during a spinner phase goes through the bar
(progressbar.Bprintln), which prints it before its next redraw instead
of racing it. Bars with a total have no background redraw and still
print warnings directly.

Model: opus-5-5
2026-10-04 04:47:35 +02:00
clawbot 2dd1194f33 Warn about and skip non-regular and .zfs operands, keeping their records (closes #9)
check / check (push) Successful in 2m48s
A symlink, socket, FIFO or device-node operand, or a directory operand
named .zfs, was silently ignored yet stayed in the scanned operands, so
the update phase deleted every record stored beneath it. Such an
operand now gets a one-line warning, counts as skipped, and is dropped
before overlapping operands are pruned and the database index is
loaded: another operand beneath it is still scanned, and the records
beneath it count as outside the scanned operands and are not deleted,
unless it lies under another operand. The exit status stays 0. An
operand that turns into one of these after that check is warned about
and skipped by the walk instead. README "scan mode" and "Rules for the
walk" say so.

Model: opus-5-5
2026-10-04 04:01:43 +02:00
clawbot 705c8729ca Hold a lock so a second scan fails at once (closes #53)
check / check (push) Successful in 1m43s
scan takes an exclusive flock(2) on a lock file beside the database
(its path with .lock appended) before it walks anything or opens the
database, and holds it until it returns. A second scan against the
same database fails at once with a one-line error naming the lock
file and exits 1. report and trees never take the lock. The lock ends
with the process, so a fatal error or an interrupt releases it; the
file is never deleted. golang.org/x/sys becomes a direct dependency.

The README smoke test now keeps the database outside the scanned
tree, where its empty lock file would have joined the empty-file
group.

Model: opus-5-5
2026-10-04 02:30:21 +02:00
clawbot 01ff3bb5f0 Test report and trees stdout write failures (closes #30)
check / check (push) Successful in 1m34s
report and trees already checked every stdout write and the final
flush. run now takes the stdout it hands to them, so tests pass a
closed file or a failing writer instead of swapping os.Stdout: a
closed stdout exits 1 with a one-line diagnostic, and the writer's
error reaches the caller.

README "Error handling" now states the two cases that never reach
sfdupes as a failed write: a pipe reader that exits early ends the
process with SIGPIPE, as with cat; and stdout closed with >&- is
replaced by /dev/null by the Go runtime before main runs, so the run
succeeds.

Model: opus-5-5
2026-10-03 18:01:27 +02:00
clawbot d63d3cc7fc Open the database read-only for report and trees (closes #8)
check / check (push) Successful in 1m17s
report and trees now connect read-only (mode=ro, query_only, the same
busy timeout) and no longer set the journal mode, which is a write. A
read-only connection to a WAL database still needs its -wal and -shm
files, or write access to the directory to create them, so scan now
switches the database back to rollback-journal mode whenever it closes
it: between scans the file alone holds the database. If a report has
the database open at that moment the switch is refused; scan warns and
the database stays in WAL mode, with its -wal and -shm files, until the
next scan. README §Database states what readers need.

Model: opus-5-5
2026-10-03 16:30:19 +02:00
clawbot c887f80f57 Escape tab, newline, CR and backslash in report paths (closes #7)
check / check (push) Successful in 1m29s
A path holding a tab or newline split a row of the report or trees
output. The path columns of both now write a backslash, tab, newline
and carriage return as \\, \t, \n and \r; every other byte is written
unchanged. Grouping and sorting still use the stored path. Warnings
on stderr are escaped the same way in warnf, so each stays one line.

In trees, the root directory's node now has the path "/" instead of
an empty string, and its children's paths start with a single slash.

README states the rule under "Report output format".

Model: opus-5-5
2026-10-03 15:30:37 +02:00
clawbot c9bf22d483 Stamp the git tag or short commit in a plain docker build (closes #67)
check / check (push) Successful in 1m1s
.dockerignore now sends .git, without .git/config, which can hold a
credential. The build stage takes the VERSION build argument when one
is given, otherwise git describe --tags --always of that .git, and
fails if the context carries .git and still yields no version. A plain
docker build . used to stamp dev. The CI checkout fetches full history
so CI sees the tag and stamps the same value as make build.

Model: opus-5-5
2026-10-02 08:49:03 +02:00
17 changed files with 2694 additions and 533 deletions
+26 -8
View File
@@ -51,14 +51,21 @@ RUN echo "gate lint, epoch ${CHECK_EPOCH}" && \
FROM golang@sha256:56961d79ea8129efddcc0b8643fd8a5416b4e6228cfd477e3fd61deb2672c587 AS builder
# We never build or run as root. Create an unprivileged user and point
# HOME and the Go caches at its home so go build and go test can write
# their caches when we drop to it below. $GOPATH/bin is deliberately not
# on PATH: script/bootstrap no longer `go install`s anything (the linter
# runs from a pinned image, never from a host install), so nothing lands
# HOME and the build cache at its home so go build and go test can write
# it when we drop to it below. $GOPATH/bin is deliberately not on PATH:
# script/bootstrap no longer `go install`s anything (the linter runs
# from a pinned image, never from a host install), so nothing lands
# there and adding it would only widen what this image resolves.
#
# The module cache is kept outside that home, at the base image's
# default /go/pkg/mod, and belongs to root: script/bootstrap fills it as
# root. Do not move it into the home and hand it over with `chown -R`:
# that walks every file in it, which took from about 80 s to over ten
# minutes on a shared host, depending on load.
RUN adduser -D -u 1000 builder
ENV HOME=/home/builder
ENV GOPATH=/home/builder/go
ENV GOMODCACHE=/go/pkg/mod
ENV GOCACHE=/home/builder/.cache/go-build
WORKDIR /src
@@ -83,11 +90,22 @@ COPY script/ script/
COPY go.mod go.sum ./
RUN script/bootstrap
COPY . .
# Hand builder only what it writes to, without walking the module cache.
# This layer stays cached with bootstrap.
# - /src itself: make build writes the binary into it, and git refuses
# a repository whose top directory belongs to another user.
# - the module cache's cache/download directory itself, not what is in
# it: Go only reads the downloaded modules, but make build saves its
# lookup of this module's own version from git there, in a new
# directory named after the module path.
# - builder's home: the go commands bootstrap ran as root left Go's
# telemetry files there, a few small files.
RUN chown builder:builder /src /go/pkg/mod/cache/download && \
chown -R builder:builder /home/builder
# Hand the sources and caches to the unprivileged user, then drop root
# before running any checks or builds.
RUN chown -R builder:builder /src /home/builder
# The sources are handed to builder as they are copied, so no layer has
# to walk them. Then drop root before running any checks or builds.
COPY --chown=builder:builder . .
USER builder
# Fail the build unless the branch is green. Runs as non-root so the
+280 -35
View File
@@ -51,6 +51,122 @@ daily `sfdupes scan` cron job, with the reporting commands run
interactively whenever needed; their results are as fresh as the last
completed scan.
### Install
With Go installed, this builds and installs the current `main` branch:
```sh
go install sneak.berlin/go/sfdupes@main
```
The binary goes to `$(go env GOPATH)/bin`, or to `$GOBIN` when that is
set. A binary installed this way reports its version as `dev`; one
built from a clone or into the Docker image carries the git tag or
commit it was built from.
From a clone, `make build` writes the binary to `./sfdupes`:
```sh
git clone https://git.eeqj.de/sneak/sfdupes.git
cd sfdupes
make build
```
Copy the binary to `/usr/local/bin` for the cron job below.
`make docker` builds the Docker image, tagged `sfdupes`, after running
the tests and the linter (see "Build"). The image runs `sfdupes` as
root with the database at its default path, so a bind mount of
`/var/lib/sfdupes` keeps the database between runs. Mount the scanned
tree at the same path inside the container as on the host; read-only
is enough. The database records paths as the container sees them, so
the reports then name the host's paths.
```sh
make docker
docker run --rm -v /srv:/srv:ro -v /var/lib/sfdupes:/var/lib/sfdupes \
sfdupes scan /srv
docker run --rm -v /var/lib/sfdupes:/var/lib/sfdupes sfdupes report > dupes.tsv
```
### Daily scan from cron
Run `scan` as root, so that it can read every file: a path it cannot
read is skipped with a warning and loses its database record (see
"Rules for the walk"). As a file `/etc/cron.d/sfdupes`:
```
30 3 * * * root /usr/local/bin/sfdupes scan /srv 2>>/var/log/sfdupes.log || tail -n 3 /var/log/sfdupes.log
```
- The database is `/var/lib/sfdupes/db.sqlite`, created with its
directory by the first scan. To keep it elsewhere, set
`SFDUPES_DATABASE=/path/to/db.sqlite` before the command on the same
line.
- `scan` writes nothing to stdout. Its stderr, appended here to
`/var/log/sfdupes.log`, holds a plain progress line as each phase
starts and then at most every 5 seconds, a warning for each path it
skips, and the summary line (see "Progress" and "`scan` mode"). The
log grows with every scan; rotate it like any other.
- Skipped paths do not fail a scan: it still exits 0, and cron sends
nothing. A scan that fails, or is stopped by `SIGINT` or `SIGTERM`,
exits 1 with the reason among the last lines of the log; `tail`
prints them, and cron mails them to root if the host can send mail.
- A scan still running when the next one starts carries on. The new
one fails at once, and the lines cron mails include
`sfdupes: another scan is running (lock held on /var/lib/sfdupes/db.sqlite.lock)`.
- `report` and `trees` need only read access to the database (see
"Database"). Under the usual umask of `022` the first scan creates
it readable by every user, so an unprivileged user can run them
against root's database.
### Reading the reports
Each row of `report` names two copies of one file, and each row of
`trees` two copies of one directory tree (see "Report output format"
and "Trees output format"). In a group of copies, the path that sorts
first byte by byte is `first` and every other path is a `dupe` of it.
`first` says nothing about which copy is the original or the oldest;
which copy to keep is your choice.
A row is a candidate, not proof:
- The reports read only the database, so they show the files as of
the last scan; a file may have changed or gone since.
- A file of 50 MiB or more is compared only on samples of its content
(see "Duplicate detection").
- Paths that are hard links to one file are listed as duplicates, but
they share their data, so removing one frees nothing.
Compare a pair byte for byte before removing either copy. For the row
`/srv/a/big.iso`, `/srv/b/big-copy.iso`, `4294967296`:
```sh
cmp /srv/a/big.iso /srv/b/big-copy.iso && echo identical
[ /srv/a/big.iso -ef /srv/b/big-copy.iso ] && echo "hard links"
```
`cmp` prints nothing and exits 0 only when every byte matches, and
otherwise reports where the files differ. The second line prints
`hard links` when the two paths are the same file, so removing either
frees nothing.
A path holding a backslash, tab, newline or carriage return is escaped
in the reports (see "Report output format"). Undo the escapes before
using it. `printf '%b'` does exactly that, because every backslash in
an escaped path starts one of the four escapes. Command substitution
drops trailing newlines, so print an `x` after the path and remove it
afterwards, or a path that ends in a newline names a different file:
```sh
p="$(printf '%bx' '/srv/a/tab\tname.txt')"; p="${p%x}"
cmp "$p" /srv/b/tab-copy.txt
```
Check a `trees` row with `diff -r`, which compares the two trees file
by file and also names anything present in only one of them, such as
an empty directory or a symlink, which `trees` does not see.
## Rationale
Duplicate finders that hash entire files do not scale to the target
@@ -94,7 +210,15 @@ Goals, in order:
millions of files, ~150 TB filesystem, possibly slow or busy disks
(ZFS pool under resilver). Holding one small record (path, size,
mtime) per file in memory during a scan is acceptable; holding
every file's hashes is not (they stay in the database).
every file's hashes is not (they stay in the database). The
reporting commands do not hold every file's hashes either: `report`
lets SQLite group and order the records and writes each row as it
reads it, so its memory does not grow with the database, and
`trees` reads the records in path order and keeps each directory's
path, digest and totals, plus the hashes of only the files in the
directories holding the record being read, so its memory grows with
the number of directories and with the size of the largest
directory.
3. **Scan incrementally, analyze offline.** The expensive filesystem
scan maintains a persistent database; an unchanged file is never
read again on a rescan, except to compute its content hash once a
@@ -104,8 +228,8 @@ Goals, in order:
again. `scan` is designed to be cronned; the reports run at any
time against the last completed scan.
4. **Clean stream separation.** Everything on stdout is machine-readable
data. All progress, warnings, and summaries go to stderr. Never mix
them.
data. All progress, warnings, summaries, and help and usage text go
to stderr. Never mix them.
### Constraints
@@ -113,8 +237,10 @@ Goals, in order:
`sfdupes`.
- Dependencies: standard library, `github.com/spf13/cobra` for the
CLI, **one progress-bar library**
(`github.com/schollz/progressbar/v3`), and **one SQLite driver**
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled).
(`github.com/schollz/progressbar/v3`), `golang.org/x/term` to tell
whether stderr is a terminal, **one SQLite driver**
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled), and
`golang.org/x/sys` for `flock(2)` (the scan lock, see "Database").
`github.com/spf13/viper` is permitted if configuration-file support
is ever needed, but is not currently used. No other third-party
deps.
@@ -139,8 +265,18 @@ Three subcommands, all implemented:
sfdupes scan [--workers N] [-x] PATH...
sfdupes report > dupes.tsv
sfdupes trees > dupetrees.tsv
sfdupes --version
sfdupes [command] --help
```
`--workers N` sets the size of each `scan` worker pool (default: the
number of CPUs), and `-x` (`--one-file-system`) keeps the walk of each
operand on that operand's filesystem; both are described under
"`scan` mode". `sfdupes --version` (or `-v`) prints one line,
`sfdupes VERSION`, to stdout and exits 0, writing nothing to stderr.
`-h` or `--help`, alone or after a subcommand, prints the help text to
stderr and exits 0, writing nothing to stdout.
### Database
All three subcommands operate on a single SQLite database file:
@@ -152,18 +288,53 @@ All three subcommands operate on a single SQLite database file:
use. `report` and `trees` require an existing database; a missing
database file is a fatal error (exit 1) telling the user to run
`scan` first.
- The database uses WAL journal mode and a busy timeout, so running a
report while a cron `scan` is in progress is safe. The filesystem
is authoritative; the database is an eventually-consistent
reflection of it. Hashed records are committed in batched
transactions while the scan is still running (keeping the WAL
small and letting concurrent reports observe progress), so a
report may see a scan's changes partially applied, and a scan
that dies partway leaves a valid database holding everything
hashed so far; the next scan skips those records and converges
toward the filesystem.
- Only one `scan` runs against a database at a time. For its whole
run, `scan` holds an exclusive `flock(2)` lock on a lock file
beside the database, named by appending `.lock` to the database
path (`/var/lib/sfdupes/db.sqlite.lock` by default), taken before
it walks the filesystem or opens the database. A second `scan`
against the same database does not wait: it fails at once with a
one-line error naming the lock file and exits 1, without walking
anything or opening the database, and the running scan carries on.
The lock file is created on first use, open to its owner only, and
left in place: a leftover file blocks nothing, because the lock
ends with the process holding it however it ends, a fatal error or
an interrupt included, and deleting the file while a scan runs
would let a second scan start. `report` and `trees` never take the
lock, so they run during a scan.
- While `scan` runs, the database is in WAL journal mode with a busy
timeout, so running a report while a cron `scan` is in progress is
safe. The filesystem is authoritative; the database is an
eventually-consistent reflection of it. Hashed records are
committed in batched transactions while the scan is still running
(keeping the WAL small and letting concurrent reports observe
progress), so a report may see a scan's changes partially applied,
and a scan that dies partway leaves a valid database holding
every batch committed so far (an interrupted scan also commits the
batch in progress, see "Error handling and exit codes"); the next
scan skips those records and converges toward the filesystem.
- `scan` switches the database back to rollback-journal mode when it
closes it, so between scans the database file alone holds the whole
database. Each switch needs the database to itself: a `scan` that
starts while a report is still reading waits for it up to the
10-second busy timeout, then fails; a `scan` that ends while a
report has the database open warns and leaves the database in WAL
mode until the next scan. `report` writes each row as it reads it,
so it is still reading while its output is paused (a pager, a
stalled pipe), and a `scan` started then fails after the busy
timeout.
- `report` and `trees` open the database read-only and need only read
access to the database file, and no write access to its directory.
While the database is in WAL mode they also read the `-wal` and
`-shm` files beside it, which SQLite creates with the database
file's permissions.
- Schema (`PRAGMA user_version` is the schema version, currently 1; a
database with any other version is a fatal error):
database with any other version is a fatal error. `scan` creates
the schema and sets the version in one transaction, so a first scan
stopped while doing so leaves an empty database the next scan sets
up. A database at version 0 that already has a `files` table was
therefore not made by sfdupes; every subcommand refuses it with an
error telling the user to remove the file and rescan):
```sql
CREATE TABLE files (
@@ -174,6 +345,7 @@ All three subcommands operate on a single SQLite database file:
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
) WITHOUT ROWID;
CREATE INDEX files_signature ON files (size, head, tail, content);
```
Paths are stored as BLOBs because Unix paths are raw bytes, not
@@ -188,7 +360,8 @@ All three subcommands operate on a single SQLite database file:
until the content phase of a scan (see "`scan` mode" below) has
read the file. A record with an empty `content` is never part of a
duplicate group, though it still defines the file for tree
reconstruction.
reconstruction. The `files_signature` index lets SQLite group the
records by signature for `report` without sorting the whole table.
### Duplicate detection
@@ -251,6 +424,18 @@ duplicates another or lies under another is dropped before walking,
so every file is reached exactly once and produces one database
record.
An operand that is a symlink (never followed, not even as an operand),
socket, FIFO, or device node, or a directory named `.zfs`, is not
scanned. `scan` prints a one-line warning naming the path and what it
is, counts it as skipped, and drops it from the scanned operands before
reading the database. Another operand beneath it is still scanned. The
records stored beneath it are not deleted: they are treated like any
other record outside the scanned operands, including the content-phase
exception below. If it lies under another operand, they are under that
operand instead, and are deleted like any other record there that this
scan did not verify. This is not an error: a scan whose every operand
is dropped walks nothing and exits 0.
`scan` synchronizes the database with the filesystem state under the
scanned operands:
@@ -356,9 +541,12 @@ during the hash phase:
Rules for the walk:
- Only regular files. Skip directories, symlinks (do not follow,
including symlink operands), sockets, FIFOs, and device nodes.
including symlink operands), sockets, FIFOs, and device nodes. An
operand that is a symlink, socket, FIFO, or device node is dropped
as described in "`scan` mode" above.
- Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs;
walking them would list every file once per snapshot).
walking them would list every file once per snapshot), not even
when it is an operand; such an operand is dropped the same way.
- Filesystem boundaries are crossed by default. With `-x`
(long form `--one-file-system`, following the GNU `du`/`rsync`
convention), never descend into a directory on a different
@@ -369,9 +557,11 @@ Rules for the walk:
path, and continue. Per-file errors never abort the run; the final
summary reports how many were skipped. As specified above, a
skipped path that has a database record from an earlier scan loses
that record, unless it failed only in the content phase; an
unreadable directory subtree likewise loses its records (accepted:
the database mirrors what the latest scan could actually verify).
that record, unless it failed only in the content phase, or is an
operand dropped before the database was read that lies under no
other operand; an unreadable directory subtree likewise loses its
records (accepted: the database mirrors what the latest scan could
actually verify).
Concurrency: the walk phase (which also stats files), the hash phase,
and the content phase each use a worker pool of `--workers` workers
@@ -399,9 +589,12 @@ arguments.
**`report` must never touch the filesystem being analyzed.** It does not
stat, open, or otherwise access any path that appears in the records; its
only I/O is reading the database and writing stdout/stderr. It must
produce identical output whether or not the scanned filesystem is still
mounted.
only I/O is reading the database, writing stdout/stderr, and the
temporary file SQLite sorts in when the duplicate rows do not fit in
memory. SQLite puts that file in `$SQLITE_TMPDIR` or `$TMPDIR` when set,
otherwise in `/var/tmp` (or `/tmp`), and deletes it as soon as it has
opened it. `report` must produce identical output whether or not the
scanned filesystem is still mounted.
Processing:
@@ -428,6 +621,15 @@ first dupe size
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296
```
Paths are raw bytes and may hold any byte except NUL, so the path
columns (`first` and `dupe`) are escaped to keep every row one line of
tab-separated fields: a backslash is written as `\\`, a tab as `\t`, a
newline as `\n`, and a carriage return as `\r`. Every other byte is
written unchanged, including bytes that are not valid UTF-8. Undoing
those four escapes gives back the stored path. Grouping and ordering
use the stored path, not the escaped one. The warnings `scan` prints on
stderr are escaped the same way, so each warning is one line.
Summary to stderr: records read, number of duplicate groups, number of
dupe files, and total reclaimable bytes (sum of `size` over all dupe
rows) in human units.
@@ -499,6 +701,9 @@ first dupe files size
/srv/a/project /srv/backup/project 3417 104857600
```
The `first` and `dupe` paths are escaped as described under "Report
output format". The root directory's path is `/`.
Summary to stderr: records read, number of duplicate-tree groups,
number of dupe trees, and total reclaimable bytes (sum of `size` over
all dupe rows) in human units.
@@ -533,22 +738,62 @@ hash: [12345/98765] 12% |████ | 92 files/s elapsed 2:32 eta 17:54
Additional requirements:
- When stderr is not a TTY, do not emit ANSI redraws: print a plain
one-line progress update no more often than every 5 seconds instead.
- When stderr is not a terminal (a pipe, a file, `/dev/null`), do not
emit ANSI redraws: print a plain one-line progress update the moment
each phase starts, then no more often than every 5 seconds.
- Progress updates are driven from the main goroutine and must be
non-blocking with respect to the worker pool.
non-blocking with respect to the worker pool. On a terminal the
spinner-style displays also redraw on their own several times a
second, so their count and elapsed time stay current while a phase
waits for its next item.
- A warning printed during a phase always lands on a line of its own,
never inside the progress display.
- A bar whose phase stops short of its total, as an interrupted one
does, is left as last drawn rather than filled up.
- `report` and `trees` modes need no progress display, only their
stderr summaries.
### Error handling and exit codes
- `0`: success, even if individual files were skipped with warnings.
- `1`: fatal error (e.g., a `PATH` operand does not exist, the
database cannot be created/opened/read/written, a missing database
for `report`/`trees`, stdout write failure).
- `1`: fatal error (e.g., a `PATH` operand does not exist, another
`scan` is already running against the same database, the database
cannot be created/opened/read/written, a missing database for
`report`/`trees`, stdout write failure), or a `scan` stopped by
`SIGINT` or `SIGTERM` (see below).
- `2`: usage error (including `scan` with no `PATH` operand and
`report`/`trees` with any positional argument).
A stdout write failure, such as a full disk, is reported in one line on
stderr and exits 1. Two cases never reach sfdupes as a failed write:
- When the reader of a stdout pipe exits early, as in
`sfdupes report | head`, the next write ends sfdupes with `SIGPIPE`,
quietly and without a summary, the way `cat` or `sort` end. The
shell reports the signal (status 141 in most shells), not exit 1.
- When stdout is closed outright (`sfdupes report >&-`), the Go
runtime opens `/dev/null` in its place before sfdupes starts, so
the output is discarded and the run succeeds, as with
`> /dev/null`.
`scan` stops cleanly on `SIGINT` (Ctrl-C) or `SIGTERM`. Its workers
stop taking work, each finishing at most the directory listing or file
it is reading; the progress display is finished; and the records it has
hashed but not yet committed are committed, so the next scan does not
hash them again. Apart from that commit it starts no further writes or
deletions: records are deleted only after a complete walk, so those
under paths an interrupted walk never reached are kept. The database is
closed and the lock released as on any other exit, the line
`scan: interrupted after N files` goes to stderr, N being the number of
files the walk reached, and the exit code is 1. The next scan skips the
records already written and converges as usual.
After the first signal `scan` stops catching them, so a second one ends
it at once, as an uncaught signal does: the records not yet committed
are lost, and the database is left valid, as when any scan dies (see
"Database"). A `SIGINT` that `scan` inherits as ignored, as a script's
background job does, stays ignored.
## Entrypoints
This repository adheres to the
@@ -686,7 +931,7 @@ All of the following, run in this directory, must pass:
```sh
d=$(mktemp -d)
export SFDUPES_DATABASE="$d/db.sqlite"
export SFDUPES_DATABASE="$(mktemp -d)/db.sqlite"
mkdir -p "$d/a" "$d/b"
head -c 2000 /dev/urandom > "$d/a/one.bin"
cp "$d/a/one.bin" "$d/b/copy.bin"
@@ -714,9 +959,9 @@ All of the following, run in this directory, must pass:
./sfdupes report
```
(The scan database lives inside `$d` here purely for test hygiene;
scanning `$d` therefore also records the SQLite file itself, which
is harmless.)
(The database lives in a temp directory of its own: inside `$d`,
the scan would record it, and its empty lock file would join the
`empty1`/`empty2` group.)
Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin`
form one group (two dupe rows, `first` is the lexicographically
+53
View File
@@ -29,6 +29,59 @@
# Completed Steps
- test that `scan` refuses a database with another schema version
(2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/64)
- test the `-x` filesystem-boundary rules in `subdirJob` (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/17)
- `scan` creates the schema in one transaction; a version-0 database with a
`files` table is refused with a clear schema-version error (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/11)
- README documents install, Docker, a daily cron scan and how to read
and check the reports (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/54)
- the `Dockerfile` build stage keeps the Go module cache out of `builder`'s
home and copies the sources with `--chown`, so no `chown -R` walks them
(2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/43)
- `--version` prints `sfdupes VERSION` to stdout; README documents it and
`--help` (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/15)
- `scan` stops cleanly on `SIGINT` or `SIGTERM`: commits what it has
hashed, deletes nothing more, exits 1 (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/5)
- `report` and `trees` stream the records instead of holding them all in
memory; the schema gains the `files_signature` index (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/14)
- progress prints at once on a non-terminal, uses a real terminal test,
and prints warnings through a spinner instead of racing its redraw
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/13)
- warn about and skip symlink, socket, FIFO, device and `.zfs`
operands, keeping the records beneath them (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/9)
- `scan` holds a lock on a lock file beside the database for its whole run,
so a second `scan` fails at once with exit 1 (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/53)
- test stdout write failures in `report` and `trees`; README states that
`| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null`
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/30)
- `report` and `trees` open the database read-only, and `scan` leaves it
out of WAL mode, so reading needs only read access (2026-10-03, closes
https://git.eeqj.de/sneak/sfdupes/issues/8)
- escape tabs, newlines, carriage returns and backslashes in report,
trees and warning paths; the root directory's path is `/`
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/7)
- stamp the git tag or short commit in a plain `docker build .`
instead of `dev` (2026-10-02, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/67): `.dockerignore` now
+247 -13
View File
@@ -4,11 +4,16 @@ import (
"context"
"database/sql"
"errors"
"fmt"
"os"
"os/signal"
"path/filepath"
"slices"
"strconv"
"strings"
"sync"
"sync/atomic"
"syscall"
"testing"
"time"
)
@@ -150,9 +155,8 @@ func assertRecordsIntact(t *testing.T, db *sql.DB, before []string) {
// Every one of those records would look vanished to the update phase.
// The guard is what stops the scan there, and this test is what
// notices if it stops doing so: deleting the guard, or making it
// unreachable, makes the scan carry its truncated view into a later
// phase and fail there instead, with a wrapped error rather than the
// bare cancellation.
// unreachable, makes the scan carry its truncated view into the update
// phase, which counts every record the walk never reached for removal.
//
//nolint:paralleltest // counts goroutines: must not run beside others
func TestSyncScanCancelledMidWalkKeepsRecords(t *testing.T) {
@@ -182,10 +186,10 @@ func TestSyncScanCancelledMidWalkKeepsRecords(t *testing.T) {
// assertWalkGuardAborted checks that the scan stopped at the post-walk
// guard: with a census that is neither empty (the walk really ran)
// nor complete (it really was cut short), and with the guard's own
// bare cancellation as the error. A wrapped error means the partial
// census was carried past the guard into the hash or update phase,
// which is the failure this test exists to catch.
// nor complete (it really was cut short), and with no record counted
// for removal. A removal count means the partial census was carried
// past the guard into the update phase, which is the failure this test
// exists to catch.
func assertWalkGuardAborted(t *testing.T, st scanStats, err error) {
t.Helper()
@@ -194,12 +198,6 @@ func assertWalkGuardAborted(t *testing.T, st scanStats, err error) {
err, context.Canceled)
}
if errors.Unwrap(err) != nil {
t.Errorf("syncScan reported %q, want the guard's bare "+
"cancellation: a wrapped error means the truncated census "+
"reached a later phase", err)
}
if st.unchanged == 0 {
t.Fatalf("stats = %+v: the census is empty, so the walk never "+
"ran and the guard was reached for the wrong reason", st)
@@ -260,6 +258,242 @@ func TestSyncScanCancelledBeforeLoadIndex(t *testing.T) {
assertRecordsIntact(t, db, before)
}
// hashCancelAtDone is the consultation on which the mid-hash test's
// context cancels itself. The walk of buildWalkCancelTree spends about
// one per file and three per directory, and the hash phase then one per
// file hashed, so this lands about half way through the hash phase.
const hashCancelAtDone = walkCancelFiles + 3*walkCancelDirs +
walkCancelFiles/2
// TestSyncScanCancelledMidHashKeepsHashedRecords cancels a first scan
// part-way through its hash phase. The fixture holds fewer files than a
// batch, so every file hashed is still waiting to be committed: the scan
// must commit them all before it returns, and the next scan must hash
// only the rest.
func TestSyncScanCancelledMidHashKeepsHashedRecords(t *testing.T) {
t.Parallel()
dir := buildWalkCancelTree(t)
db := openTestDB(t)
st, err := syncScan(newWalkClock(hashCancelAtDone), db,
[]string{dir}, walkCancelWorkers, false)
if !errors.Is(err, context.Canceled) {
t.Fatalf("syncScan cancelled mid-hash = %v, want %v",
err, context.Canceled)
}
if st.walked != walkCancelFiles || st.added == 0 ||
st.added >= walkCancelFiles {
t.Fatalf("stats = %+v: want the walk complete and the hash phase "+
"cut short", st)
}
if got := len(dbRecords(t, db)); got != st.added {
t.Errorf("%d records after the cancelled scan, want the %d it hashed",
got, st.added)
}
hashed := st.added
st = syncTree(t, db, dir)
if st.added != walkCancelFiles-hashed || st.unchanged != hashed {
t.Errorf("next scan stats = %+v, want %d added %d unchanged",
st, walkCancelFiles-hashed, hashed)
}
}
// storedPaths opens the database at path as report does, which fails
// unless it is a valid database, and returns its records' paths.
func storedPaths(t *testing.T, path string) []string {
t.Helper()
db, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
defer func() { _ = db.Close() }()
return recordPaths(dbRecords(t, db))
}
// TestRunScanInterrupted calls the scan entrypoint with a context that
// is already cancelled, as when a signal arrives at once. It must return
// errInterrupted promptly with its one line on stderr, leave the
// database valid and as it was, and leave nothing in the way of the
// next scan, which must bring the database up to date.
func TestRunScanInterrupted(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
stderr := captureStderr(t)
dir := buildSmokeTree(t)
err := runScan(t.Context(), []string{dir}, walkCancelWorkers, false)
if err != nil {
t.Fatal(err)
}
before := storedPaths(t, path)
// A vanished file and a new one: the interrupted scan records
// neither.
gone := filepath.Join(dir, "a", "unique.bin")
err = os.Remove(gone)
if err != nil {
t.Fatal(err)
}
added := writeFile(t, dir, "a/new.bin", pattern(50, 10))
shown := len(stderr())
done := make(chan struct{})
go func() {
defer close(done)
err = runScan(cancelledContext(t), []string{dir}, walkCancelWorkers,
false)
}()
awaitReturn(t, done, "runScan")
if !errors.Is(err, errInterrupted) {
t.Fatalf("runScan on a cancelled context = %v, want %v",
err, errInterrupted)
}
want := "scan: interrupted after 0 files\n"
if got := stderr()[shown:]; got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
assertNoSidecars(t, path)
if got := storedPaths(t, path); !slices.Equal(got, before) {
t.Errorf("records = %q after the interrupted scan, want %q",
got, before)
}
err = runScan(t.Context(), []string{dir}, walkCancelWorkers, false)
if err != nil {
t.Fatal(err)
}
got := storedPaths(t, path)
if slices.Contains(got, gone) || !slices.Contains(got, added) {
t.Errorf("records = %q after the next scan, want %q gone and %q "+
"added", got, gone, added)
}
}
// TestRunScanInterruptedMidHash interrupts the scan entrypoint part-way
// through its hash phase, after the database is open. It must return
// errInterrupted, release the lock, end stderr with its line counting
// every file the walk reached, close the database out of WAL mode, and
// keep the records it hashed.
func TestRunScanInterruptedMidHash(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
stderr := captureStderr(t)
dir := buildWalkCancelTree(t)
err := runScan(newWalkClock(hashCancelAtDone), []string{dir},
walkCancelWorkers, false)
if !errors.Is(err, errInterrupted) {
t.Fatalf("runScan interrupted mid-hash = %v, want %v",
err, errInterrupted)
}
holdScanLock(t, path)
want := fmt.Sprintf("scan: interrupted after %d files\n", walkCancelFiles)
if got := stderr(); !strings.HasSuffix(got, want) {
t.Errorf("stderr = %q, want it to end with %q", got, want)
}
assertNoSidecars(t, path)
db, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
defer func() { _ = db.Close() }()
// A plain close also removes the sidecars, but leaves WAL mode on.
var mode string
err = db.QueryRowContext(t.Context(), "PRAGMA journal_mode").Scan(&mode)
if err != nil {
t.Fatal(err)
}
if mode != "delete" {
t.Errorf("journal mode = %q after the interrupted scan, want %q",
mode, "delete")
}
kept := len(dbRecords(t, db))
if kept == 0 || kept >= walkCancelFiles {
t.Errorf("%d records after the interrupted scan, want those it "+
"hashed: some but not all of the %d files", kept, walkCancelFiles)
}
}
// TestInterruptContextCatchesSIGTERM sends SIGTERM to the test process
// while the scan's handler is installed, and checks that it cancels the
// scan's context.
//
//nolint:paralleltest // signals the whole process: must not run beside a scan
func TestInterruptContextCatchesSIGTERM(t *testing.T) {
// Caught here as well, so that a handler that misses SIGTERM fails
// this test instead of ending the test process.
caught := make(chan os.Signal, 1)
signal.Notify(caught, syscall.SIGTERM)
defer signal.Stop(caught)
ctx, stop := interruptContext(t.Context())
defer stop()
err := syscall.Kill(os.Getpid(), syscall.SIGTERM)
if err != nil {
t.Fatal(err)
}
select {
case <-ctx.Done():
case <-time.After(poolUnwind):
t.Fatal("SIGTERM did not cancel the scan's context")
}
}
// TestCommitFullBatchKeepsFailedBatch checks that a full batch whose
// commit fails, as it does once the scan is interrupted, stays in the
// batch, so that syncScan's final commit saves it.
func TestCommitFullBatchKeepsFailedBatch(t *testing.T) {
t.Parallel()
s := &scanState{db: openTestDB(t)}
for i := range updateBatchSize {
s.batch = append(s.batch, scanRec{path: "/f" + strconv.Itoa(i)})
}
err := s.commitFullBatch(cancelledContext(t))
if !errors.Is(err, context.Canceled) {
t.Fatalf("commitFullBatch on a cancelled context = %v, want %v",
err, context.Canceled)
}
if len(s.batch) != updateBatchSize {
t.Errorf("batch holds %d records after the failed commit, want %d",
len(s.batch), updateBatchSize)
}
}
// drainClosed counts the values received from ch until it closes,
// failing the test if it does not close within poolUnwind. A pool that
// ignored its cancellation leaves its channel open with its goroutines
+232 -24
View File
@@ -11,6 +11,7 @@ import (
"slices"
"strconv"
"golang.org/x/sys/unix"
// The pure-Go SQLite driver, registered as "sqlite"; keeps cgo
// disabled.
_ "modernc.org/sqlite"
@@ -32,6 +33,11 @@ const schemaVersion = 1
// scan.
const dbDirPerm = 0o755
// lockFilePerm is the mode for the scan lock file. Anyone who can open
// the file can hold the lock and keep every scan from running, so it
// is open to its owner only.
const lockFilePerm = 0o600
// createTableSQL is the schema applied to a fresh database. Paths are
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
const createTableSQL = `
@@ -45,6 +51,12 @@ CREATE TABLE files (
) WITHOUT ROWID
`
// createIndexSQL indexes the records by signature, so report can have
// SQLite group them without sorting the whole table.
const createIndexSQL = `
CREATE INDEX files_signature ON files (size, head, tail, content)
`
// upsertSQL inserts one file record, replacing any existing record for
// the same path.
const upsertSQL = `
@@ -64,6 +76,10 @@ var errNoDatabase = errors.New(
// does not understand.
var errSchemaVersion = errors.New("unsupported database schema version")
// errScanRunning reports that another scan holds the lock on the
// database.
var errScanRunning = errors.New("another scan is running")
// databasePath resolves the database location: SFDUPES_DATABASE when
// set and non-empty, the compiled-in default otherwise.
func databasePath() string {
@@ -74,16 +90,24 @@ func databasePath() string {
return defaultDatabasePath
}
// openDB opens the SQLite database at path with WAL journaling and a
// busy timeout, so a report can run while a cron scan is in progress.
// It does not create or verify the schema.
func openDB(path string) (*sql.DB, error) {
dsn := "file:" + path +
"?_pragma=busy_timeout(10000)" +
"&_pragma=journal_mode(WAL)" +
"&_pragma=synchronous(NORMAL)"
// scanParams are the connection parameters for scan: read-write, with
// WAL journaling and a busy timeout, so a report can run while a cron
// scan is in progress. closeScanDatabase leaves WAL mode again.
const scanParams = "_pragma=busy_timeout(10000)" +
"&_pragma=journal_mode(WAL)" +
"&_pragma=synchronous(NORMAL)"
db, err := sql.Open("sqlite", dsn)
// reportParams are the connection parameters for report and trees:
// read-only, with the same busy timeout. They set no journal mode,
// because setting one is a write.
const reportParams = "mode=ro" +
"&_pragma=busy_timeout(10000)" +
"&_pragma=query_only(1)"
// openDB opens the SQLite database at path with the connection
// parameters params. It does not create or verify the schema.
func openDB(path, params string) (*sql.DB, error) {
db, err := sql.Open("sqlite", "file:"+path+"?"+params)
if err != nil {
return nil, fmt.Errorf("open database %s: %w", path, err)
}
@@ -96,6 +120,43 @@ func openDB(path string) (*sql.DB, error) {
return db, nil
}
// lockScanDatabase takes the lock that keeps a second scan off the
// database at path: an exclusive flock(2) on the file beside it named
// path with ".lock" appended, created along with the database's parent
// directory if missing. A lock held by another scan fails at once
// instead of waiting. The lock lasts until the returned file is closed
// or the process ends. The file is never deleted: a scan that deleted
// it would let the next scan lock a new file while another still holds
// the old one.
func lockScanDatabase(path string) (*os.File, error) {
err := os.MkdirAll(filepath.Dir(path), dbDirPerm)
if err != nil {
return nil, fmt.Errorf("create database directory: %w", err)
}
lockPath := path + ".lock"
//nolint:gosec // the operator chooses the database path
f, err := os.OpenFile(lockPath, os.O_RDWR|os.O_CREATE, lockFilePerm)
if err != nil {
return nil, err
}
err = unix.Flock(int(f.Fd()), unix.LOCK_EX|unix.LOCK_NB)
if err != nil {
_ = f.Close()
if errors.Is(err, unix.EWOULDBLOCK) {
return nil, fmt.Errorf("%w (lock held on %s)",
errScanRunning, lockPath)
}
return nil, fmt.Errorf("lock %s: %w", lockPath, err)
}
return f, nil
}
// openScanDatabase opens the database for the scan subcommand, creating
// the file, its parent directory, and the schema as needed.
func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
@@ -104,7 +165,7 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return nil, fmt.Errorf("create database directory: %w", err)
}
db, err := openDB(path)
db, err := openDB(path, scanParams)
if err != nil {
return nil, err
}
@@ -119,6 +180,24 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return db, nil
}
// closeScanDatabase switches the database at path from WAL back to
// rollback-journal mode and closes it. Out of WAL mode the database
// file alone holds the whole database, so a reader needs no -wal or
// -shm file beside it, nor write access to create them. The switch
// fails while a report has the database open; the database then stays
// in WAL mode, still readable, until a later scan closes it.
func closeScanDatabase(ctx context.Context, db *sql.DB, path string) {
// Runs on the way out of a cancelled scan too.
_, err := db.ExecContext(context.WithoutCancel(ctx),
"PRAGMA journal_mode = DELETE")
if err != nil {
fmt.Fprintf(os.Stderr, "scan: database %s left in WAL mode: %v\n",
path, err)
}
_ = db.Close()
}
// openReportDatabase opens an existing database for the report and
// trees subcommands. A missing database file is an error directing the
// user to run scan first; the schema version must match exactly.
@@ -134,12 +213,18 @@ func openReportDatabase(ctx context.Context,
return nil, fmt.Errorf("database: %w", err)
}
db, err := openDB(path)
db, err := openDB(path, reportParams)
if err != nil {
return nil, err
}
v, err := userVersion(ctx, db)
if err == nil && v == 0 {
// An empty database passes this check and fails the version
// check below.
err = checkUnversioned(ctx, db)
}
if err != nil {
_ = db.Close()
@@ -166,6 +251,11 @@ func initSchema(ctx context.Context, db *sql.DB) error {
switch v {
case 0:
err = checkUnversioned(ctx, db)
if err != nil {
return err
}
return createSchema(ctx, db)
case schemaVersion:
return nil
@@ -175,20 +265,65 @@ func initSchema(ctx context.Context, db *sql.DB) error {
}
}
// checkUnversioned checks a database at user_version 0 before it is
// taken for an empty one. createSchema creates the files table and
// sets the version together, so a files table at version 0 was made by
// something else. Adopting it could corrupt unrelated data, so that is
// a schema-version error telling the operator to remove the file and
// rescan.
func checkUnversioned(ctx context.Context, db *sql.DB) error {
var name string
err := db.QueryRowContext(ctx,
"SELECT name FROM sqlite_master "+
"WHERE type = 'table' AND name = 'files'").Scan(&name)
switch {
case err == nil:
return fmt.Errorf(
"has a files table but no schema version; "+
"remove the file and rescan: %w", errSchemaVersion)
case errors.Is(err, sql.ErrNoRows):
return nil
default:
return fmt.Errorf("check for files table: %w", err)
}
}
// createSchema applies the schema to a fresh database and stamps the
// schema version.
// schema version in one transaction, so a creation stopped partway, by
// an interrupt or an error, leaves an empty database the next scan
// sets up, never a files table at version 0, which checkUnversioned
// refuses.
func createSchema(ctx context.Context, db *sql.DB) error {
_, err := db.ExecContext(ctx, createTableSQL)
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
_, err = db.ExecContext(ctx,
defer func() { _ = tx.Rollback() }()
_, err = tx.ExecContext(ctx, createTableSQL)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
_, err = tx.ExecContext(ctx, createIndexSQL)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
_, err = tx.ExecContext(ctx,
"PRAGMA user_version = "+strconv.Itoa(schemaVersion))
if err != nil {
return fmt.Errorf("set schema version: %w", err)
}
err = tx.Commit()
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
return nil
}
@@ -204,18 +339,18 @@ func userVersion(ctx context.Context, db *sql.DB) (int, error) {
return v, nil
}
// loadFileRows reads every record from the files table.
func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) {
// loadFileRows streams every record to fn in path order: byte order,
// which is the order of the primary key, so SQLite does not sort.
func loadFileRows(ctx context.Context, db *sql.DB, fn func(r scanRec)) error {
rows, err := db.QueryContext(ctx,
"SELECT path, size, mtime, head, tail, content FROM files")
"SELECT path, size, mtime, head, tail, content FROM files "+
"ORDER BY path")
if err != nil {
return nil, fmt.Errorf("read records: %w", err)
return fmt.Errorf("read records: %w", err)
}
defer func() { _ = rows.Close() }()
var recs []scanRec
for rows.Next() {
var (
path []byte
@@ -225,19 +360,92 @@ func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) {
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail,
&r.content)
if err != nil {
return nil, fmt.Errorf("read record: %w", err)
return fmt.Errorf("read record: %w", err)
}
r.path = string(path)
recs = append(recs, r)
fn(r)
}
err = rows.Err()
if err != nil {
return nil, fmt.Errorf("read records: %w", err)
return fmt.Errorf("read records: %w", err)
}
return recs, nil
return nil
}
// dupeRowsSQL selects every record in a duplicate group, with the
// group's first path. A group is the records with a content hash that
// share a size, head, tail, and content, when there are two or more of
// them. The rows come in report order: groups by size descending, then
// by first path, and each group's paths ascending.
const dupeRowsSQL = `
SELECT g.first, f.path, f.size
FROM files AS f
JOIN (
SELECT size, head, tail, content, MIN(path) AS first
FROM files
WHERE content <> ''
GROUP BY size, head, tail, content
HAVING COUNT(*) > 1
) AS g USING (size, head, tail, content)
ORDER BY f.size DESC, g.first, f.path
`
// loadDupeRows streams the rows of dupeRowsSQL to fn and returns the
// number of records in the database. The count and the rows are read
// in one transaction, so they agree while a scan is committing. An
// error from fn stops the reading and is returned as it is.
func loadDupeRows(ctx context.Context, db *sql.DB,
fn func(first, path string, size int64) error,
) (int, error) {
// Everything goes through tx: the report connection is the only
// one, so a query on db would wait for tx forever.
tx, err := db.BeginTx(ctx, &sql.TxOptions{ReadOnly: true})
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
defer func() { _ = tx.Rollback() }()
var records int
err = tx.QueryRowContext(ctx, "SELECT COUNT(*) FROM files").Scan(&records)
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
rows, err := tx.QueryContext(ctx, dupeRowsSQL)
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
defer func() { _ = rows.Close() }()
for rows.Next() {
var (
first, path []byte
size int64
)
err = rows.Scan(&first, &path, &size)
if err != nil {
return 0, fmt.Errorf("read record: %w", err)
}
err = fn(string(first), string(path), size)
if err != nil {
return 0, err
}
}
err = rows.Err()
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
return records, nil
}
// loadFileMeta streams every record's path, size, mtime, and whether
+132 -25
View File
@@ -5,6 +5,7 @@ import (
"database/sql"
"errors"
"fmt"
"os"
"path/filepath"
"slices"
"strings"
@@ -72,9 +73,78 @@ func TestOpenScanDatabaseCreates(t *testing.T) {
defer func() { _ = db.Close() }()
recs, err := loadFileRows(t.Context(), db)
if err != nil || len(recs) != 0 {
t.Fatalf("loadFileRows = %v, %v; want empty, nil", recs, err)
if recs := dbRecords(t, db); len(recs) != 0 {
t.Fatalf("records = %v, want none", recs)
}
}
func TestOpenDatabaseUnversionedForeign(t *testing.T) {
t.Parallel()
// A database that has a files table but user_version 0, written by
// some other tool. report, trees and scan must refuse it with the
// schema-version error, not adopt it and not emit a raw SQLite
// "table files already exists".
path := testDBPath(t)
db, err := sql.Open("sqlite", path)
if err != nil {
t.Fatal(err)
}
_, err = db.ExecContext(t.Context(), "CREATE TABLE files (x INTEGER)")
if err != nil {
t.Fatal(err)
}
_ = db.Close()
_, err = openReportDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) ||
!strings.Contains(err.Error(), "remove the file and rescan") {
t.Fatalf("report: err = %v, want errSchemaVersion telling the "+
"operator to remove the file and rescan", err)
}
_, err = openScanDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) ||
!strings.Contains(err.Error(), "remove the file and rescan") {
t.Fatalf("scan: err = %v, want errSchemaVersion telling the "+
"operator to remove the file and rescan", err)
}
}
func TestSchemaCreationStoppedPartway(t *testing.T) {
t.Parallel()
// A first scan stopped while creating the schema must leave a
// database the next scan accepts. max_page_count(2) leaves room for
// the files table but not its index, so schema creation fails right
// after CREATE TABLE, a point an interrupt could also stop it at.
path := testDBPath(t)
db, err := openDB(path, scanParams+"&_pragma=max_page_count(2)")
if err != nil {
t.Fatal(err)
}
err = initSchema(t.Context(), db)
_ = db.Close()
if err == nil {
t.Fatal("initSchema with no room for the index succeeded")
}
db, err = openScanDatabase(t.Context(), path)
if err != nil {
t.Fatalf("next scan: %v", err)
}
defer func() { _ = db.Close() }()
v, err := userVersion(t.Context(), db)
if err != nil || v != schemaVersion {
t.Fatalf("userVersion = %d, %v; want %d, nil", v, err, schemaVersion)
}
}
@@ -87,9 +157,11 @@ func TestOpenReportDatabaseMissing(t *testing.T) {
}
}
func TestOpenReportDatabaseVersionMismatch(t *testing.T) {
func TestOpenDatabaseVersionMismatch(t *testing.T) {
t.Parallel()
// A database stamped with a schema version other than 0 and
// schemaVersion. report, trees and scan must all refuse it.
path := testDBPath(t)
db, err := openScanDatabase(t.Context(), path)
@@ -106,7 +178,12 @@ func TestOpenReportDatabaseVersionMismatch(t *testing.T) {
_, err = openReportDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) {
t.Fatalf("err = %v, want errSchemaVersion", err)
t.Fatalf("report: err = %v, want errSchemaVersion", err)
}
_, err = openScanDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) {
t.Fatalf("scan: err = %v, want errSchemaVersion", err)
}
}
@@ -130,6 +207,49 @@ func TestOpenReportDatabaseOK(t *testing.T) {
_ = db.Close()
}
func TestCloseScanDatabaseWhileReportOpen(t *testing.T) {
t.Parallel()
// A report holding the database open stops scan from taking it out
// of WAL mode. The -wal and -shm files must then stay beside it, so
// that a later report still needs only read access.
path := testDBPath(t)
scanDB, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
reportDB, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
closeScanDatabase(t.Context(), scanDB, path)
_ = reportDB.Close()
_, err = os.Stat(path + "-wal")
if err != nil {
t.Fatalf("no -wal left: the switch out of WAL mode was not "+
"stopped: %v", err)
}
makeReadOnly(t, path)
reportDB, err = openReportDatabase(t.Context(), path)
if err != nil {
t.Fatalf("openReportDatabase: %v", err)
}
defer func() { _ = reportDB.Close() }()
err = loadFileRows(t.Context(), reportDB, func(scanRec) {})
if err != nil {
t.Fatalf("loadFileRows: %v", err)
}
}
func TestApplyChangesRoundTrip(t *testing.T) {
t.Parallel()
@@ -152,15 +272,8 @@ func TestApplyChangesRoundTrip(t *testing.T) {
t.Fatalf("applyChanges: %v", err)
}
got, err := loadFileRows(t.Context(), db)
if err != nil {
t.Fatal(err)
}
slices.SortFunc(got, func(a, b scanRec) int {
return strings.Compare(a.path, b.path)
})
// The records come back in path order, which is the order of recs.
got := dbRecords(t, db)
if !slices.Equal(got, recs) {
t.Fatalf("rows = %+v, want %+v", got, recs)
}
@@ -177,11 +290,7 @@ func TestApplyChangesRoundTrip(t *testing.T) {
t.Fatalf("applyChanges: %v", err)
}
got, err = loadFileRows(t.Context(), db)
if err != nil {
t.Fatal(err)
}
got = dbRecords(t, db)
if len(got) != 1 || got[0] != upd {
t.Fatalf("rows = %+v, want just %+v", got, upd)
}
@@ -210,9 +319,8 @@ func TestApplyChangesBatching(t *testing.T) {
t.Fatalf("applyChanges: %v", err)
}
got, err := loadFileRows(t.Context(), db)
if err != nil || len(got) != n {
t.Fatalf("loadFileRows = %d rows, %v; want %d", len(got), err, n)
if got := dbRecords(t, db); len(got) != n {
t.Fatalf("records = %d, want %d", len(got), n)
}
deletes := make([]string, 0, n)
@@ -226,8 +334,7 @@ func TestApplyChangesBatching(t *testing.T) {
t.Fatalf("applyChanges deletes: %v", err)
}
got, err = loadFileRows(t.Context(), db)
if err != nil || len(got) != 0 {
t.Fatalf("loadFileRows = %d rows, %v; want 0", len(got), err)
if got := dbRecords(t, db); len(got) != 0 {
t.Fatalf("records = %d, want 0", len(got))
}
}
+2 -2
View File
@@ -5,6 +5,8 @@ go 1.25.7
require (
github.com/schollz/progressbar/v3 v3.19.1
github.com/spf13/cobra v1.10.2
golang.org/x/sys v0.46.0
golang.org/x/term v0.44.0
modernc.org/sqlite v1.54.0
)
@@ -18,8 +20,6 @@ require (
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/rivo/uniseg v0.4.7 // indirect
github.com/spf13/pflag v1.0.9 // indirect
golang.org/x/sys v0.46.0 // indirect
golang.org/x/term v0.44.0 // indirect
modernc.org/libc v1.74.1 // indirect
modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect
+52 -21
View File
@@ -15,6 +15,7 @@
// sfdupes scan [--workers N] [-x] PATH...
// sfdupes report > dupes.tsv
// sfdupes trees > dupetrees.tsv
// sfdupes --version
//
// See README.md for the complete specification.
package main
@@ -57,22 +58,27 @@ var errNoSubcommand = errors.New("no subcommand")
var Version = "dev"
func main() {
os.Exit(run(os.Args[1:], os.Stderr))
// Once the reader of a stdout pipe has gone, as in "sfdupes report |
// head", the Go runtime ends the process with SIGPIPE on the next
// write instead of returning an error (README "Error handling").
// Registering for SIGPIPE with os/signal would change that.
os.Exit(run(os.Args[1:], os.Stdout, os.Stderr))
}
// run executes args against the command tree and returns the process
// exit code. It is the program's single exit point: the subcommands
// return their errors instead of exiting, so every deferred cleanup —
// above all closing the database, which checkpoints the SQLite WAL —
// runs before the process ends.
func run(args []string, stderr io.Writer) int {
// runs before the process ends. The report and trees subcommands write
// their data to stdout.
func run(args []string, stdout, stderr io.Writer) int {
// A nil slice makes cobra fall back to os.Args, which would let a
// test binary's own flags reach the command tree.
if args == nil {
args = []string{}
}
root := newRootCommand(stderr)
root := newRootCommand(stdout, stderr)
root.SetArgs(args)
err := root.Execute()
@@ -82,6 +88,9 @@ func run(args []string, stderr io.Writer) int {
switch {
case err == nil:
return exitOK
case errors.Is(err, errInterrupted):
// The interrupted scan has printed its own line.
return exitFatal
case errors.As(err, &fatal):
// The command ran and failed: a runtime error, reported
// without the usage text that a usage error gets.
@@ -96,15 +105,29 @@ func run(args []string, stderr io.Writer) int {
}
// newRootCommand builds the command tree. Everything on stdout is
// machine-readable data; all human-facing output (help, usage, errors)
// goes to stderr.
func newRootCommand(stderr io.Writer) *cobra.Command {
// machine-readable data, the version line included; all human-facing
// output (help, usage, errors) goes to stderr.
func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
var showVersion bool
printVersion := runE(func(context.Context, []string) error {
_, err := fmt.Fprintf(stdout, "sfdupes %s\n", Version)
if err != nil {
return fmt.Errorf("write stdout: %w", err)
}
return nil
})
root := &cobra.Command{
Use: "sfdupes",
Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
Version: Version,
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, _ []string) error {
Use: "sfdupes",
Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, args []string) error {
if showVersion {
return printVersion(cmd, args)
}
// A missing subcommand prints usage and exits 2: cobra
// prints the usage text for the returned error, and run
// maps everything that is not a fatal error to exit 2.
@@ -117,6 +140,11 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
root.SetErr(stderr)
root.CompletionOptions.DisableDefaultCmd = true
// Cobra's built-in version flag prints through the help writer,
// stderr; this one prints to stdout.
root.Flags().BoolVarP(&showVersion, "version", "v", false,
"print the version to stdout")
var (
scanWorkers int
scanOneFS bool
@@ -127,6 +155,9 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Walk trees and synchronize the scan database",
Args: cobra.MinimumNArgs(1),
RunE: runE(func(ctx context.Context, args []string) error {
ctx, stop := interruptContext(ctx)
defer stop()
return runScan(ctx, args, scanWorkers, scanOneFS)
}),
}
@@ -140,7 +171,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the file-level duplicates report",
Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error {
return runReport(ctx)
return runReport(ctx, stdout)
}),
}
@@ -149,7 +180,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the duplicate-tree report",
Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error {
return runTrees(ctx)
return runTrees(ctx, stdout)
}),
}
@@ -158,13 +189,13 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
return root
}
// runE adapts a subcommand implementation to cobra's RunE. Cobra
// prints the error and the command's usage text for every error RunE
// returns, but a subcommand that ran and failed has no usage problem
// to report: both are silenced here, and the error is marked fatal so
// that run reports it on stderr and exits 1 rather than 2. The command's
// context is handed to the implementation: cancelling it unwinds the
// scan's worker pools.
// runE adapts a subcommand implementation, or the version print, to
// cobra's RunE. Cobra prints the error and the command's usage text for
// every error RunE returns, but a subcommand that ran and failed has no
// usage problem to report: both are silenced here, and the error is
// marked fatal so that run reports it on stderr and exits 1 rather than
// 2. The command's context is handed to the implementation: cancelling
// it unwinds the scan's worker pools.
func runE(
fn func(ctx context.Context, args []string) error,
) func(*cobra.Command, []string) error {
+430 -71
View File
@@ -44,23 +44,61 @@ func assertNoSidecars(t *testing.T, path string) {
}
}
// captureStdout redirects os.Stdout to a file for the rest of the test
// and returns a function reading back everything written to it. Only
// machine-readable data belongs on stdout (README design goal 4), so
// the tests assert on it directly.
func captureStdout(t *testing.T) func() string {
// makeReadOnly takes write permission away from the database at path,
// from any WAL sidecar beside it, and from their directory, as for a
// user reading a database that a root cron scan keeps. Root ignores
// file permissions, so it skips the test when run as root.
func makeReadOnly(t *testing.T, path string) {
t.Helper()
f, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
if os.Geteuid() == 0 {
t.Skip("root ignores file permissions")
}
err := os.Chmod(path, 0o400)
if err != nil {
t.Fatal(err)
}
saved := os.Stdout
os.Stdout = f
for _, suffix := range walSuffixes {
err = os.Chmod(path+suffix, 0o400)
if err != nil && !errors.Is(err, fs.ErrNotExist) {
t.Fatal(err)
}
}
dir := filepath.Dir(path)
//nolint:gosec // reaching the database needs the search bit
err = os.Chmod(dir, 0o500)
if err != nil {
t.Fatal(err)
}
// Runs before t.TempDir's own cleanup, which must delete the files.
t.Cleanup(func() {
//nolint:gosec // removing the directory needs its search bit back
_ = os.Chmod(dir, 0o700)
})
}
// captureStderr redirects os.Stderr to a file for the rest of the test
// and returns a function reading back everything written to it. scan
// writes its warnings and summary straight to os.Stderr, not to the
// stderr writer run is given.
func captureStderr(t *testing.T) func() string {
t.Helper()
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
os.Stderr = f
t.Cleanup(func() {
os.Stdout = saved
os.Stderr = saved
_ = f.Close()
})
@@ -91,13 +129,13 @@ func captureStdout(t *testing.T) func() string {
// brokenDatabase writes a database that opens cleanly and passes the
// schema-version check but has no files table, so the first query
// fails with the database already open: a fatal error on a path that
// owns an open database.
// owns an open database. It closes the database the way scan does.
func brokenDatabase(t *testing.T) string {
t.Helper()
path := testDBPath(t)
db, err := openDB(path)
db, err := openDB(path, scanParams)
if err != nil {
t.Fatal(err)
}
@@ -108,10 +146,7 @@ func brokenDatabase(t *testing.T) string {
t.Fatal(err)
}
err = db.Close()
if err != nil {
t.Fatal(err)
}
closeScanDatabase(t.Context(), db, path)
return path
}
@@ -144,7 +179,10 @@ func TestOpenDatabaseKeepsWALWhileOpen(t *testing.T) {
func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
// Every subcommand that owns an open database must close it when
// it fails: no os.Exit between the open and the return.
// it fails: no os.Exit between the open and the return. The
// sidecar check is evidence of the close only for scan: report and
// trees only read a database that is out of WAL mode, which leaves
// nothing on disk whether they close it or not.
cases := map[string][]string{
cmdScan: {cmdScan},
cmdReport: {cmdReport},
@@ -160,17 +198,15 @@ func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
args = append(args, t.TempDir())
}
var stderr bytes.Buffer
var stdout, stderr bytes.Buffer
stdout := captureStdout(t)
code := run(args, &stderr)
code := run(args, &stdout, &stderr)
if code != exitFatal {
t.Errorf("run(%v) = %d, want %d", args, code, exitFatal)
}
assertNoSidecars(t, path)
assertFatalOutput(t, stderr.String(), stdout())
assertFatalOutput(t, stderr.String(), stdout.String())
// Proof that the failure happened after the open: only a
// query against the opened database can report this.
@@ -188,18 +224,16 @@ func TestRunMissingOperandIsFatalNotUsage(t *testing.T) {
// must not dump the usage text.
t.Setenv(databaseEnv, testDBPath(t))
var stderr bytes.Buffer
stdout := captureStdout(t)
var stdout, stderr bytes.Buffer
missing := filepath.Join(t.TempDir(), "nope")
code := run([]string{cmdScan, missing}, &stderr)
code := run([]string{cmdScan, missing}, &stdout, &stderr)
if code != exitFatal {
t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal)
}
assertFatalOutput(t, stderr.String(), stdout())
assertFatalOutput(t, stderr.String(), stdout.String())
}
// assertFatalOutput checks that a fatal error was reported the way
@@ -243,11 +277,9 @@ func TestRunUsageErrors(t *testing.T) {
// path that does not exist.
t.Setenv(databaseEnv, testDBPath(t))
var stderr bytes.Buffer
var stdout, stderr bytes.Buffer
stdout := captureStdout(t)
code := run(tc.args, &stderr)
code := run(tc.args, &stdout, &stderr)
if code != exitUsage {
t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage)
}
@@ -256,43 +288,79 @@ func TestRunUsageErrors(t *testing.T) {
t.Errorf("stderr = %q, want %q", stderr.String(), tc.want)
}
if got := stdout(); got != "" {
if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
})
}
}
// TestRunHelpAndVersionSucceed checks that the two informational flags
// exit 0 and keep their human-facing output on stderr.
//
//nolint:paralleltest // captureStdout replaces the process-wide os.Stdout
func TestRunHelpAndVersionSucceed(t *testing.T) {
assertHumanOutput(t, "--help")
assertHumanOutput(t, "--version")
func TestRunHelp(t *testing.T) {
t.Parallel()
// README §Subcommands: help goes to stderr, exits 0, and leaves
// stdout empty.
cases := [][]string{{"--help"}, {"-h"}, {cmdScan, "--help"}}
for _, args := range cases {
var stdout, stderr bytes.Buffer
code := run(args, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%v) = %d, want %d", args, code, exitOK)
}
if !strings.Contains(stderr.String(), usageMarker) {
t.Errorf("run(%v) stderr = %q, want the help text",
args, stderr.String())
}
if got := stdout.String(); got != "" {
t.Errorf("run(%v) stdout = %q, want nothing (data only)",
args, got)
}
}
}
// assertHumanOutput runs sfdupes with one informational flag and checks
// that it succeeds with its output on stderr and stdout untouched
// (README design goal 4).
func assertHumanOutput(t *testing.T, arg string) {
t.Helper()
func TestRunVersion(t *testing.T) {
t.Parallel()
// README §Subcommands: the version is one line on stdout, with
// nothing on stderr, and exits 0.
for _, arg := range []string{"--version", "-v"} {
var stdout, stderr bytes.Buffer
code := run([]string{arg}, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
}
want := "sfdupes " + Version + "\n"
if got := stdout.String(); got != want {
t.Errorf("run(%s) stdout = %q, want %q", arg, got, want)
}
if got := stderr.String(); got != "" {
t.Errorf("run(%s) stderr = %q, want nothing", arg, got)
}
}
}
func TestRunVersionWriteFailureIsFatal(t *testing.T) {
t.Parallel()
// README §Error handling: a stdout write failure exits 1, reported
// in one line on stderr.
var stderr bytes.Buffer
stdout := captureStdout(t)
code := run([]string{arg}, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
code := run([]string{"--version"}, failingWriter{}, &stderr)
if code != exitFatal {
t.Errorf("run(--version) = %d, want %d", code, exitFatal)
}
if stderr.Len() == 0 {
t.Errorf("run(%s) wrote nothing to stderr", arg)
}
if got := stdout(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
want := "sfdupes: write stdout: " + errWriteFailed.Error() + "\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
}
@@ -320,21 +388,31 @@ func scanFixture(t *testing.T) []string {
t.Fatal(err)
}
var stderr bytes.Buffer
scanOK(t, dir)
stdout := captureStdout(t)
return dupes
}
code := run([]string{cmdScan, dir}, &stderr)
// scanOK runs scan over operands, fails the test unless it exits 0 with
// nothing on stdout, and returns everything it printed to stderr.
func scanOK(t *testing.T, operands ...string) string {
t.Helper()
var stdout bytes.Buffer
stderr := captureStderr(t)
code := run(append([]string{cmdScan}, operands...), &stdout, os.Stderr)
if code != exitOK {
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
t.Fatalf("run(scan %q) = %d, want %d; stderr: %s",
operands, code, exitOK, stderr())
}
if got := stdout(); got != "" {
if got := stdout.String(); got != "" {
t.Errorf("scan stdout = %q, want nothing (data only)", got)
}
return dupes
return stderr()
}
func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
@@ -345,24 +423,119 @@ func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
assertNoSidecars(t, path)
}
func TestRunScanSkipsSymlinkOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dir := t.TempDir()
writeFile(t, dir, "target/sub/f", pattern(1, 10))
link := filepath.Join(dir, "link")
err := os.Symlink(filepath.Join(dir, "target"), link)
if err != nil {
t.Fatal(err)
}
// Scanning a directory through the symlink stores a record beneath
// the symlink's own path for a file beneath its target.
scanOK(t, filepath.Join(link, "sub"))
assertOperandSkipped(t, path, link, "symlink",
filepath.Join(link, "sub", "f"))
}
func TestRunScanWalksOperandUnderSymlinkOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dir := t.TempDir()
writeFile(t, dir, "target/sub/f", pattern(1, 10))
link := filepath.Join(dir, "link")
err := os.Symlink(filepath.Join(dir, "target"), link)
if err != nil {
t.Fatal(err)
}
// link is dropped as a symlink, but link/sub must still be scanned,
// not dropped as lying under link.
scanOK(t, link, filepath.Join(link, "sub"))
db, err := openDB(path, reportParams)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
recordByPath(t, dbRecords(t, db), filepath.Join(link, "sub", "f"))
}
func TestRunScanSkipsZFSOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
zfs := filepath.Join(t.TempDir(), ".zfs")
snapshot := filepath.Join(zfs, "snapshot", "hourly")
f := writeFile(t, snapshot, "f", pattern(1, 10))
// An operand beneath a .zfs directory is walked, because it is not
// itself named .zfs.
scanOK(t, snapshot)
assertOperandSkipped(t, path, zfs, ".zfs directory", f)
}
// assertOperandSkipped scans operand alone and checks that it is skipped
// as kind: a warning naming it, one skip in the summary, exit 0, and the
// record for kept, which an earlier scan stored beneath operand, still
// in the database at dbPath.
func assertOperandSkipped(t *testing.T, dbPath, operand, kind,
kept string,
) {
t.Helper()
stderr := scanOK(t, operand)
warning := "walk " + operand + ": skipping " + kind + " operand\n"
if !strings.Contains(stderr, warning) {
t.Errorf("stderr = %q, want %q", stderr, warning)
}
summary := "scan: 0 files seen (0 added, 0 updated, 0 removed, " +
"0 unchanged), 1 skipped\n"
if !strings.Contains(stderr, summary) {
t.Errorf("stderr = %q, want %q", stderr, summary)
}
db, err := openDB(dbPath, reportParams)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
recordByPath(t, dbRecords(t, db), kept)
}
func TestRunReportSucceeds(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dupes := scanFixture(t)
var stderr bytes.Buffer
var stdout, stderr bytes.Buffer
stdout := captureStdout(t)
code := run([]string{cmdReport}, &stderr)
code := run([]string{cmdReport}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
if got := stdout(); got != want {
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
@@ -375,11 +548,9 @@ func TestRunTreesSucceeds(t *testing.T) {
dupes := scanFixture(t)
var stderr bytes.Buffer
var stdout, stderr bytes.Buffer
stdout := captureStdout(t)
code := run([]string{cmdTrees}, &stderr)
code := run([]string{cmdTrees}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
@@ -389,9 +560,197 @@ func TestRunTreesSucceeds(t *testing.T) {
// trees of each other.
want := "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
if got := stdout(); got != want {
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
assertNoSidecars(t, path)
}
func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
// README §Database: report and trees need only read access to the
// database file. With its directory read-only as well, SQLite
// cannot create any file beside it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dupes := scanFixture(t)
assertNoSidecars(t, path)
makeReadOnly(t, path)
cases := map[string]string{
cmdReport: "first\tdupe\tsize\n" +
dupes[0] + "\t" + dupes[1] + "\t300\n",
cmdTrees: "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) +
"\t1\t300\n",
}
for name, want := range cases {
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
continue
}
if got := stdout.String(); got != want {
t.Errorf("%s stdout = %q, want %q", name, got, want)
}
}
}
// holdScanLock takes the lock on the database at path, as a running
// scan does, and holds it until the test ends. It fails the test when
// the lock is already held.
func holdScanLock(t *testing.T, path string) {
t.Helper()
lock, err := lockScanDatabase(path)
if err != nil {
t.Fatalf("lock %s: %v", path, err)
}
t.Cleanup(func() { _ = lock.Close() })
}
func TestRunSecondScanFails(t *testing.T) {
// README §Database: while one scan holds the lock, a second scan
// fails at once, naming the lock file, without creating the
// database.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
holdScanLock(t, path)
var stdout, stderr bytes.Buffer
code := run([]string{cmdScan, t.TempDir()}, &stdout, &stderr)
if code != exitFatal {
t.Errorf("run(scan) = %d, want %d", code, exitFatal)
}
want := "sfdupes: another scan is running (lock held on " +
path + ".lock)\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
_, err := os.Stat(path)
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("stat %s = %v, want the database not created", path, err)
}
}
func TestRunScanReleasesLock(t *testing.T) {
// README §Database: a scan releases the lock however it ends.
t.Run("success", func(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
})
t.Run("fatal error", func(t *testing.T) {
path := brokenDatabase(t)
t.Setenv(databaseEnv, path)
code := run([]string{cmdScan, t.TempDir()}, io.Discard, io.Discard)
if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d", code, exitFatal)
}
holdScanLock(t, path)
})
}
func TestRunReportsDuringScan(t *testing.T) {
// README §Database: report and trees never take the lock, so they
// run while a scan holds it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
for _, name := range []string{cmdReport, cmdTrees} {
var stderr bytes.Buffer
code := run([]string{name}, io.Discard, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
}
}
}
func TestRunStdoutClosedIsFatal(t *testing.T) {
// README §Error handling: a stdout write failure exits 1, reported
// in one line on stderr.
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
stdout, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
if err != nil {
t.Fatal(err)
}
err = stdout.Close()
if err != nil {
t.Fatal(err)
}
var stderr bytes.Buffer
code := run([]string{name}, stdout, &stderr)
if code != exitFatal {
t.Errorf("run(%s) = %d, want %d", name, code, exitFatal)
}
got := stderr.String()
if !strings.HasPrefix(got, "sfdupes: write stdout: ") ||
!strings.Contains(got, os.ErrClosed.Error()) ||
strings.Count(got, "\n") != 1 {
t.Errorf("stderr = %q, want one line reporting the "+
"failed stdout write", got)
}
})
}
}
// errWriteFailed is the error failingWriter returns.
var errWriteFailed = errors.New("write failed")
// failingWriter is a stdout that fails every write.
type failingWriter struct{}
func (failingWriter) Write([]byte) (int, error) { return 0, errWriteFailed }
func TestStdoutWriteErrorPropagates(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
cases := map[string]func(context.Context, io.Writer) error{
cmdReport: runReport,
cmdTrees: runTrees,
}
for name, fn := range cases {
err := fn(t.Context(), failingWriter{})
if !errors.Is(err, errWriteFailed) {
t.Errorf("%s: error = %v, want %v", name, err, errWriteFailed)
}
}
}
+47 -18
View File
@@ -6,6 +6,7 @@ import (
"time"
"github.com/schollz/progressbar/v3"
"golang.org/x/term"
)
// plainInterval is the minimum time between progress lines when stderr
@@ -24,24 +25,23 @@ const percentScale = 100
// stderrIsTTY reports whether stderr is attached to a terminal.
func stderrIsTTY() bool {
fi, err := os.Stderr.Stat()
if err != nil {
return false
}
return fi.Mode()&os.ModeCharDevice != 0
return term.IsTerminal(int(os.Stderr.Fd()))
}
// progress renders one scan pass's progress on stderr. On a TTY it
// delegates to the progressbar library (spinner style when the total is
// unknown, full bar with count/percent/rate/elapsed/ETA otherwise). When
// stderr is not a TTY it emits no ANSI redraws: it prints a plain
// one-line update no more often than every plainInterval.
// one-line update as the pass starts, then no more often than every
// plainInterval.
//
// All methods must be called from the main goroutine only. A nil
// *progress is a valid no-display receiver: every method is a no-op,
// so batched database flushes during the streaming pass can reuse the
// update-pass helpers without rendering anything.
// All methods must be called from the main goroutine only. On a TTY
// the library also redraws a spinner from its own goroutine, several
// times a second, so its count and elapsed time stay current while a
// pass waits for its next item. A nil *progress is a valid
// no-display receiver: every method is a no-op, so batched database
// flushes during the streaming pass can reuse the update-pass helpers
// without rendering anything.
type progress struct {
label string
total int64 // -1 when unknown (walk pass)
@@ -53,10 +53,22 @@ type progress struct {
func newProgress(label string, total int64) *progress {
p := &progress{label: label, total: total, start: time.Now()}
if !stderrIsTTY() {
if stderrIsTTY() {
p.bar = newBar(label, total)
return p
}
// Print the zero state at once: the first item may take minutes,
// and a pass must never look hung.
p.last = p.start
fmt.Fprintln(os.Stderr, p.plainLine())
return p
}
// newBar builds the TTY display for newProgress.
func newBar(label string, total int64) *progressbar.ProgressBar {
opts := []progressbar.Option{
progressbar.OptionSetWriter(os.Stderr),
progressbar.OptionSetDescription(label),
@@ -81,9 +93,7 @@ func newProgress(label string, total int64) *progress {
)
}
p.bar = progressbar.NewOptions64(total, opts...)
return p
return progressbar.NewOptions64(total, opts...)
}
// increment records one completed item and refreshes the display.
@@ -106,26 +116,45 @@ func (p *progress) increment() {
}
// warnf prints a one-line warning to stderr without corrupting the bar.
// The whole message is escaped like a report's path columns, so a path
// holding a newline cannot split the warning.
func (p *progress) warnf(format string, args ...any) {
if p == nil {
return
}
msg := escapePath(fmt.Sprintf(format, args...))
if p.bar != nil && p.total < 0 {
// The library also redraws a spinner from its own goroutine, so
// a direct write could land inside a redraw. The bar prints the
// warning itself, just before its next redraw.
_, _ = progressbar.Bprintln(p.bar, msg)
return
}
if p.bar != nil {
_ = p.bar.Clear()
}
fmt.Fprintf(os.Stderr, format+"\n", args...)
fmt.Fprintln(os.Stderr, msg)
}
// finish terminates the pass's display.
// finish terminates the pass's display. A bar whose pass stopped short
// of its total, as an interrupted one does, is left as last drawn; the
// library's Finish would fill it up.
func (p *progress) finish() {
if p == nil {
return
}
if p.bar != nil {
_ = p.bar.Finish()
if p.total >= 0 && p.count < p.total {
_ = p.bar.Exit()
} else {
_ = p.bar.Finish()
}
fmt.Fprintln(os.Stderr)
+189
View File
@@ -0,0 +1,189 @@
package main
import (
"os"
"path/filepath"
"strings"
"testing"
"time"
)
// spinnerIdle comfortably outlasts the 100ms interval at which the
// progressbar library redraws a spinner from its own goroutine.
const spinnerIdle = 500 * time.Millisecond
//nolint:paralleltest // replaces the process-wide os.Stderr
func TestStderrIsTTYFalseForNonTerminals(t *testing.T) {
r, pipe, err := os.Pipe()
if err != nil {
t.Fatal(err)
}
regular, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
devNull, err := os.OpenFile(os.DevNull, os.O_WRONLY, 0)
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
t.Cleanup(func() {
os.Stderr = saved
for _, f := range []*os.File{r, pipe, regular, devNull} {
_ = f.Close()
}
})
cases := map[string]*os.File{
"a pipe": pipe,
"a regular file": regular,
os.DevNull: devNull,
}
for name, f := range cases {
os.Stderr = f
if stderrIsTTY() {
t.Errorf("stderrIsTTY() = true with stderr on %s", name)
}
}
}
// TestNewProgressPrintsBeforeFirstItem checks that each pass shows its
// zero state the moment it starts when stderr is not a terminal, and
// that the next line still waits for plainInterval.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestNewProgressPrintsBeforeFirstItem(t *testing.T) {
stderr := captureStderr(t)
newProgress("walk", -1).increment()
newProgress("hash", 10).increment()
want := "walk: 0 files, elapsed 0s\n" +
"hash: [0/10] 0% 0 files/s elapsed 0s eta ?\n"
if got := stderr(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
}
// newWalkSpinner returns the walk pass's terminal display, writing to
// os.Stderr whether or not it is a terminal, and stops the library's
// redraws when the test ends.
func newWalkSpinner(t *testing.T) *progress {
t.Helper()
p := &progress{
label: "walk", total: -1, start: time.Now(),
bar: newBar("walk", -1),
}
t.Cleanup(p.finish)
return p
}
// TestProgressWarningsOnOwnLines drives the terminal display of the walk
// pass through a run of warnings with no items between them, as when the
// walk meets many unreadable paths, for several of the spinner's
// redraws: every warning must land on a line of its own, never inside a
// redraw.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestProgressWarningsOnOwnLines(t *testing.T) {
stderr := captureStderr(t)
p := newWalkSpinner(t)
// No pause between warnings: one written straight to stderr is
// garbled only if a redraw lands while it is being written.
issued := 0
for start := time.Now(); time.Since(start) < spinnerIdle; issued++ {
p.warnf("warning")
}
// The spinner prints the warnings at its next redraw.
time.Sleep(spinnerIdle)
// A terminal shows each line as the text after its last carriage
// return.
shown := 0
for line := range strings.SplitSeq(stderr(), "\n") {
if !strings.Contains(line, "warning") {
continue
}
shown++
if text := line[strings.LastIndex(line, "\r")+1:]; text != "warning" {
t.Errorf("terminal shows %q, want %q", text, "warning")
}
}
if shown != issued {
t.Errorf("%d warning lines, want %d", shown, issued)
}
}
// TestSpinnerShowsCountAfterBurst checks that once a burst of items
// faster than the redraw limit is over, the walk display shows every
// item completed while it waits for the next one.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestSpinnerShowsCountAfterBurst(t *testing.T) {
stderr := captureStderr(t)
p := newWalkSpinner(t)
for range 50 {
p.increment()
}
time.Sleep(spinnerIdle)
if shown := lastFrame(stderr()); !strings.Contains(shown, "(50/-,") {
t.Errorf("terminal shows %q, want a count of 50", shown)
}
}
// lastFrame returns what a terminal shows of the frames a bar drew: the
// last one. The library starts each frame with a carriage return and
// erases the previous one with spaces first.
func lastFrame(out string) string {
var shown string
for frame := range strings.SplitSeq(out, "\r") {
if strings.TrimSpace(frame) != "" {
shown = frame
}
}
return shown
}
// TestBarStoppedShortKeepsCount checks that the terminal display of a
// pass that stops before its total, as an interrupted one does, is left
// as last drawn instead of being filled up.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestBarStoppedShortKeepsCount(t *testing.T) {
stderr := captureStderr(t)
p := &progress{
label: "hash", total: 10, start: time.Now(),
bar: newBar("hash", 10),
}
p.increment()
// Past the redraw limit, so the bar draws the next count.
time.Sleep(2 * barThrottle)
p.increment()
p.finish()
if shown := lastFrame(stderr()); !strings.Contains(shown, "(2/10,") {
t.Errorf("terminal shows %q, want a count of 2 of 10", shown)
}
}
+47 -92
View File
@@ -4,8 +4,8 @@ import (
"bufio"
"context"
"fmt"
"io"
"os"
"slices"
"strings"
)
@@ -28,73 +28,59 @@ type scanRec struct {
path string
}
// loadRecords opens the database and reads every file record for the
// report and trees subcommands. Any database problem — including a
// missing database — is fatal. The error is returned rather than
// exiting, so that the deferred close — which checkpoints the SQLite
// WAL — always runs; the database is closed before the caller formats
// its output, so it stays closed even if that output fails.
func loadRecords(ctx context.Context) ([]scanRec, error) {
// runReport implements the report subcommand: it prints the file-level
// duplicates report as TSV on stdout. SQLite groups and orders the
// records, and each row is written as it is read, so no group is held
// in memory. It never touches the scanned filesystem; its only I/O is
// the database (with SQLite's temporary sort file), stdout, and stderr.
// Any database problem, including a missing database, is fatal.
func runReport(ctx context.Context, stdout io.Writer) error {
dbPath := databasePath()
db, err := openReportDatabase(ctx, dbPath)
if err != nil {
return nil, err
return err
}
defer func() { _ = db.Close() }()
recs, err := loadFileRows(ctx, db)
if err != nil {
return nil, fmt.Errorf("database %s: %w", dbPath, err)
}
return recs, nil
}
// dupeGroup is one set of candidate-duplicate files: identical size,
// head hash, tail hash, and content hash. paths is sorted
// lexicographically; the first entry is the group's "first", the rest
// are dupes.
type dupeGroup struct {
size int64
paths []string
}
// runReport implements the report subcommand: it reads every record
// from the database and prints the file-level duplicates report as TSV
// on stdout. It never touches the scanned filesystem; its only I/O is
// the database, stdout, and stderr.
func runReport(ctx context.Context) error {
recs, err := loadRecords(ctx)
if err != nil {
return err
}
dupes := collectDupeGroups(recs)
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
out := bufio.NewWriterSize(stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tsize")
if err != nil {
return fmt.Errorf("write stdout: %w", err)
}
dupeFiles := 0
var (
groups, dupeFiles int
reclaimable int64
writeErr error
)
var reclaimable int64
records, err := loadDupeRows(ctx, db,
func(first, path string, size int64) error {
// A group's first path is its first row; every other
// path is a dupe.
if path == first {
groups++
for _, g := range dupes {
for _, p := range g.paths[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\n",
g.paths[0], p, g.size)
if err != nil {
return fmt.Errorf("write stdout: %w", err)
return nil
}
_, writeErr = fmt.Fprintf(out, "%s\t%s\t%d\n",
escapePath(first), escapePath(path), size)
dupeFiles++
reclaimable += g.size
}
reclaimable += size
return writeErr
})
if writeErr != nil {
return fmt.Errorf("write stdout: %w", writeErr)
}
if err != nil {
return fmt.Errorf("database %s: %w", dbPath, err)
}
err = out.Flush()
@@ -105,55 +91,24 @@ func runReport(ctx context.Context) error {
fmt.Fprintf(os.Stderr,
"report: %d records read, %d duplicate groups, %d dupe files, "+
"%s reclaimable\n",
len(recs), len(dupes), dupeFiles, humanBytes(reclaimable))
records, groups, dupeFiles, humanBytes(reclaimable))
return nil
}
// collectDupeGroups groups records by signature and returns every group
// with two or more paths, each group's paths sorted lexicographically,
// groups ordered by size descending then by first path ascending.
func collectDupeGroups(recs []scanRec) []dupeGroup {
groups := make(map[fileSig][]string)
for _, r := range recs {
// A record without a content hash has unknown content and is
// never reported as a duplicate (README "Database").
if r.content == "" {
continue
}
k := fileSig{
size: r.size, head: r.head, tail: r.tail, content: r.content,
}
groups[k] = append(groups[k], r.path)
// escapePath returns a path as it is written in a report column (README
// "Report output format"): a backslash, tab, newline or carriage return
// becomes \\, \t, \n or \r, and every other byte is kept as it is.
// Grouping and sorting use the raw path, never this form.
func escapePath(p string) string {
// Most paths need no escaping; skip building a replacer for them.
if !strings.ContainsAny(p, "\\\t\n\r") {
return p
}
var dupes []dupeGroup
for k, paths := range groups {
if len(paths) < minGroupSize {
continue
}
slices.Sort(paths)
dupes = append(dupes, dupeGroup{size: k.size, paths: paths})
}
// Biggest reclaimable space first; ties broken by first path.
slices.SortFunc(dupes, func(a, b dupeGroup) int {
if a.size != b.size {
if a.size > b.size {
return -1
}
return 1
}
return strings.Compare(a.paths[0], b.paths[0])
})
return dupes
return strings.NewReplacer(
`\`, `\\`, "\t", `\t`, "\n", `\n`, "\r", `\r`,
).Replace(p)
}
// humanBytes formats a byte count in human units (binary prefixes).
+260 -15
View File
@@ -1,11 +1,254 @@
package main
import (
"bytes"
"database/sql"
"errors"
"fmt"
"io"
"os"
"path/filepath"
"slices"
"strings"
"testing"
)
func TestCollectDupeGroups(t *testing.T) {
// awkwardDir is a directory name holding every byte the reports escape.
const awkwardDir = "/d/\tone\ntwo\rthree\\four"
// awkwardPairRecs is a duplicate pair in sibling directories /d/A and
// awkwardDir. A raw tab sorts before "A" but its escaped form `\t`
// sorts after it, so awkwardDir coming first shows that sorting uses
// the raw path.
func awkwardPairRecs() []scanRec {
return []scanRec{
{size: 5, head: "h", tail: "t", content: "c", path: "/d/A/f"},
{size: 5, head: "h", tail: "t", content: "c", path: awkwardDir + "/f"},
}
}
// seedDatabase writes recs into a fresh database and returns its path.
func seedDatabase(t *testing.T, recs []scanRec) string {
t.Helper()
path := testDBPath(t)
db, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
err = applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
err = db.Close()
if err != nil {
t.Fatal(err)
}
return path
}
// dupeGroup is one duplicate group as report reads it: the size, and
// the paths in report order, first path first.
type dupeGroup struct {
size int64
paths []string
}
// dupeGroups returns the duplicate groups report reads from db, in
// report order.
func dupeGroups(t *testing.T, db *sql.DB) []dupeGroup {
t.Helper()
var groups []dupeGroup
_, err := loadDupeRows(t.Context(), db,
func(first, path string, size int64) error {
if path == first {
groups = append(groups, dupeGroup{size: size})
}
g := &groups[len(groups)-1]
g.paths = append(g.paths, path)
return nil
})
if err != nil {
t.Fatal(err)
}
return groups
}
// dupeGroupsOf writes recs into a fresh database and returns the
// duplicate groups report reads from it.
func dupeGroupsOf(t *testing.T, recs []scanRec) []dupeGroup {
t.Helper()
db := openTestDB(t)
err := applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
return dupeGroups(t, db)
}
func TestRunReportEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
code := run([]string{cmdReport}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tsize\n" +
`/d/\tone\ntwo\rthree\\four/f` + "\t/d/A/f\t5\n"
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
func TestReportStdoutFailsWhileReading(t *testing.T) {
// Each row holds two paths longer than dir, so the report is more
// than twice the stdout buffer and stdout fails while rows are
// still being read, not at the final flush.
dir := "/" + strings.Repeat("d", 4096)
recs := make([]scanRec, ioBufSize/len(dir))
for i := range recs {
recs[i] = scanRec{
size: 1, head: "h", tail: "t", content: "c",
path: fmt.Sprintf("%s/%d", dir, i),
}
}
t.Setenv(databaseEnv, seedDatabase(t, recs))
err := runReport(t.Context(), failingWriter{})
if !errors.Is(err, errWriteFailed) ||
!strings.HasPrefix(err.Error(), "write stdout: ") {
t.Errorf("error = %v, want write stdout: %v", err, errWriteFailed)
}
}
func TestRunReportsIgnoreInsertionOrder(t *testing.T) {
// README §Constraints: identical database contents give identical
// output, whatever order the records were inserted in.
recs := append(smokeTreeRecs(), awkwardPairRecs()...)
recs = append(recs,
scanRec{size: 50, head: "b", tail: "b", content: "b", path: "/y/2"},
scanRec{size: 50, head: "b", tail: "b", content: "b", path: "/y/1"},
scanRec{size: 50, head: "a", tail: "a", content: "a", path: "/x/2"},
scanRec{size: 50, head: "a", tail: "a", content: "a", path: "/x/1"},
scanRec{size: 50, path: "/x/unhashed"},
)
reversed := slices.Clone(recs)
slices.Reverse(reversed)
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, recs))
forward := runStdout(t, name)
t.Setenv(databaseEnv, seedDatabase(t, reversed))
backward := runStdout(t, name)
if strings.Count(forward, "\n") < 3 {
t.Errorf("stdout = %q, want at least two rows", forward)
}
if forward != backward {
t.Errorf("stdout depends on insertion order: %q vs %q",
forward, backward)
}
})
}
}
// runStdout runs the subcommand name and returns its stdout, failing
// the test unless it succeeds.
func runStdout(t *testing.T, name string) string {
t.Helper()
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
}
return stdout.String()
}
func TestEscapePath(t *testing.T) {
t.Parallel()
cases := map[string]string{
"/srv/plain": "/srv/plain",
"/a\tb": `/a\tb`,
"/a\nb": `/a\nb`,
"/a\rb": `/a\rb`,
`/a\b`: `/a\\b`,
`/a\tb`: `/a\\tb`,
"/not-utf8\xff": "/not-utf8\xff",
}
for in, want := range cases {
if got := escapePath(in); got != want {
t.Errorf("escapePath(%q) = %q, want %q", in, got, want)
}
}
}
// TestWarnfEscapes checks that a warning naming a path that holds a
// newline is still one line.
//
//nolint:paralleltest // replaces the process-wide os.Stderr
func TestWarnfEscapes(t *testing.T) {
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
os.Stderr = f
t.Cleanup(func() {
os.Stderr = saved
_ = f.Close()
})
(&progress{}).warnf("stat %s: %s", "/d/a\nb", "gone")
_, err = f.Seek(0, io.SeekStart)
if err != nil {
t.Fatal(err)
}
got, err := io.ReadAll(f)
if err != nil {
t.Fatal(err)
}
want := `stat /d/a\nb: gone` + "\n"
if string(got) != want {
t.Errorf("warning = %q, want %q", got, want)
}
}
func TestDupeGroups(t *testing.T) {
t.Parallel()
recs := []scanRec{
@@ -20,7 +263,7 @@ func TestCollectDupeGroups(t *testing.T) {
{size: 7, head: "u", tail: "u", content: "u", path: "/lonely"},
}
groups := collectDupeGroups(recs)
groups := dupeGroupsOf(t, recs)
if len(groups) != 2 {
t.Fatalf("len(groups) = %d, want 2", len(groups))
}
@@ -38,7 +281,7 @@ func TestCollectDupeGroups(t *testing.T) {
}
}
func TestCollectDupeGroupsContentSeparates(t *testing.T) {
func TestDupeGroupsContentSeparates(t *testing.T) {
t.Parallel()
// Same size, head, and tail, but different content hashes: the final
@@ -53,7 +296,7 @@ func TestCollectDupeGroupsContentSeparates(t *testing.T) {
{size: 100, head: "h", tail: "t", path: "/e"},
}
groups := collectDupeGroups(recs)
groups := dupeGroupsOf(t, recs)
if len(groups) != 1 {
t.Fatalf("len(groups) = %d, want 1 (only the matching content)",
len(groups))
@@ -64,7 +307,7 @@ func TestCollectDupeGroupsContentSeparates(t *testing.T) {
}
}
func TestCollectDupeGroupsMtimeExcluded(t *testing.T) {
func TestDupeGroupsMtimeExcluded(t *testing.T) {
t.Parallel()
// mtime is informational only; records differing only in mtime
@@ -74,23 +317,25 @@ func TestCollectDupeGroupsMtimeExcluded(t *testing.T) {
{size: 9, mtime: 200, head: "h", tail: "t", content: "c", path: "/m/2"},
}
groups := collectDupeGroups(recs)
groups := dupeGroupsOf(t, recs)
if len(groups) != 1 {
t.Fatalf("len(groups) = %d, want 1", len(groups))
}
}
func TestCollectDupeGroupsTieBreak(t *testing.T) {
func TestDupeGroupsTieBreak(t *testing.T) {
t.Parallel()
// The hashes sort opposite to the first paths, so ordering the
// groups by hash instead of by first path fails this test.
recs := []scanRec{
{size: 50, head: "b", tail: "b", content: "b", path: "/beta/2"},
{size: 50, head: "b", tail: "b", content: "b", path: "/beta/1"},
{size: 50, head: "a", tail: "a", content: "a", path: "/alpha/2"},
{size: 50, head: "a", tail: "a", content: "a", path: "/alpha/1"},
{size: 50, head: "a", tail: "a", content: "a", path: "/beta/2"},
{size: 50, head: "a", tail: "a", content: "a", path: "/beta/1"},
{size: 50, head: "b", tail: "b", content: "b", path: "/alpha/2"},
{size: 50, head: "b", tail: "b", content: "b", path: "/alpha/1"},
}
groups := collectDupeGroups(recs)
groups := dupeGroupsOf(t, recs)
if len(groups) != 2 {
t.Fatalf("len(groups) = %d, want 2", len(groups))
}
@@ -102,7 +347,7 @@ func TestCollectDupeGroupsTieBreak(t *testing.T) {
}
}
func TestCollectDupeGroupsDeterministic(t *testing.T) {
func TestDupeGroupsDeterministic(t *testing.T) {
t.Parallel()
recs := []scanRec{
@@ -112,12 +357,12 @@ func TestCollectDupeGroupsDeterministic(t *testing.T) {
{size: 2, head: "b", tail: "b", content: "b", path: "/q/2"},
}
forward := collectDupeGroups(recs)
forward := dupeGroupsOf(t, recs)
reversed := slices.Clone(recs)
slices.Reverse(reversed)
backward := collectDupeGroups(reversed)
backward := dupeGroupsOf(t, reversed)
if !slices.EqualFunc(forward, backward, func(a, b dupeGroup) bool {
return a.size == b.size && slices.Equal(a.paths, b.paths)
}) {
+201 -51
View File
@@ -11,6 +11,7 @@ import (
"io"
"io/fs"
"os"
"os/signal"
"path/filepath"
"slices"
"strings"
@@ -57,6 +58,10 @@ const sampleWindow = 1024 * 1024
// and hash worker pools.
const workQueueDepth = 1024
// errInterrupted reports a scan stopped by SIGINT or SIGTERM. runScan
// has already printed its line, so run prints nothing more.
var errInterrupted = errors.New("scan interrupted")
// fileRec carries one statted file between the scan phases. dev and
// ino identify the underlying inode so hard-linked paths can share
// one read; both are zero when the platform exposes no inode.
@@ -85,10 +90,15 @@ type fileMeta struct {
// least one other file shares are ever hashed: a size-unique file
// cannot be a duplicate. A file of headTailMin or more gets its content
// hash only when its size, head, and tail match another file's. Flag
// parsing and the at-least-one-operand check are done by cobra. Errors
// are returned rather than exiting, so that the deferred close — which
// checkpoints the SQLite WAL — always runs. Cancelling ctx unwinds the
// worker pools and aborts the scan with the context's error.
// parsing and the at-least-one-operand check are done by cobra. The
// scan holds the lock on the database for its whole run, so a second
// scan fails before it walks the filesystem or opens the database.
// Errors are returned rather than exiting, so that the deferred close —
// which takes the database out of WAL mode — always runs, and the lock
// is released after it. When ctx is cancelled, as by the SIGINT or
// SIGTERM that interruptContext catches, the scan keeps what it has
// hashed (see syncScan), prints how many files its walk reached, and
// returns errInterrupted.
func runScan(ctx context.Context, roots []string, workers int,
oneFS bool,
) error {
@@ -103,14 +113,31 @@ func runScan(ctx context.Context, roots []string, workers int,
dbPath := databasePath()
db, err := openScanDatabase(ctx, dbPath)
lock, err := lockScanDatabase(dbPath)
if err != nil {
return err
}
defer func() { _ = db.Close() }()
defer func() { _ = lock.Close() }()
db, err := openScanDatabase(ctx, dbPath)
if err != nil && ctx.Err() != nil {
// Interrupted while opening; SQLite may report that with an
// error of its own rather than the context's.
return interrupted(0)
}
if err != nil {
return err
}
defer closeScanDatabase(ctx, db, dbPath)
st, err := syncScan(ctx, db, roots, workers, oneFS)
if errors.Is(err, context.Canceled) {
return interrupted(st.walked)
}
if err != nil {
return fmt.Errorf("update database %s: %w", dbPath, err)
}
@@ -124,6 +151,34 @@ func runScan(ctx context.Context, roots []string, workers int,
return nil
}
// interruptContext returns a copy of ctx that the first SIGINT or
// SIGTERM cancels; the scan command runs the scan under it. stop
// releases the signals.
func interruptContext(ctx context.Context) (context.Context, func()) {
// A SIGINT ignored from the start, as by a script's background job,
// stays ignored.
signals := []os.Signal{syscall.SIGTERM}
if !signal.Ignored(syscall.SIGINT) {
signals = append(signals, syscall.SIGINT)
}
ctx, stop := signal.NotifyContext(ctx, signals...)
// Stopping restores the default handling, so a second signal ends
// the process at once.
context.AfterFunc(ctx, stop)
return ctx, stop
}
// interrupted prints the line for a scan stopped by a signal after its
// walk reached walked files, and returns errInterrupted.
func interrupted(walked int) error {
fmt.Fprintf(os.Stderr, "scan: interrupted after %d files\n", walked)
return errInterrupted
}
// resolveRoots converts each PATH operand to an absolute, lexically
// cleaned path (symlinks are not resolved) and verifies that it
// exists. Database records are keyed by absolute path, so scan results
@@ -177,8 +232,9 @@ func pruneRoots(roots []string) []string {
}
// scanStats summarizes one scan's database synchronization for the
// final stderr summary.
// final stderr summary, or for the line an interrupted scan prints.
type scanStats struct {
walked int // files the walk reached
added int
updated int
removed int
@@ -200,7 +256,30 @@ type scanState struct {
st scanStats
}
// syncScan synchronizes the database with the filesystem under roots
// syncScan synchronizes the database with the filesystem under roots;
// see runPhases. When ctx is cancelled, as by an interrupt, it commits
// the hashed records still waiting in the batch, starts no other write
// or deletion, and returns the cancellation.
func syncScan(ctx context.Context, db *sql.DB, roots []string,
workers int, oneFS bool,
) (scanStats, error) {
s := &scanState{db: db}
err := s.runPhases(ctx, roots, workers, oneFS)
if err == nil || ctx.Err() == nil {
return s.st, err
}
// The one write made after the cancellation, so it cannot use ctx.
err = applyChanges(context.WithoutCancel(ctx), db, s.batch, nil, nil)
if err != nil {
return s.st, err
}
return s.st, ctx.Err()
}
// runPhases synchronizes the database with the filesystem under roots
// in four sequential phases: walk (enumerate and stat every file,
// building a complete size census), hash (read only the new or
// changed — or previously unhashed — files whose size at least one
@@ -210,47 +289,73 @@ type scanState struct {
// in the content hash of every record of headTailMin or more whose
// size, head, and tail match another record's). Records outside the
// roots are never touched, except that the content phase fills in
// their content hash.
func syncScan(ctx context.Context, db *sql.DB, roots []string,
// their content hash. Operands the walk cannot start from are dropped
// first, so the records beneath them count as outside the roots unless
// they lie under another root.
func (s *scanState) runPhases(ctx context.Context, roots []string,
workers int, oneFS bool,
) (scanStats, error) {
roots = pruneRoots(roots)
s := &scanState{db: db}
) error {
// Types are checked before pruning so that an operand under a
// dropped one is still scanned, not dropped as lying under it.
roots = pruneRoots(s.walkableRoots(roots))
err := s.loadIndex(ctx, roots)
if err != nil {
return s.st, err
return err
}
changed, unhashed := s.walkPhase(startWalk(ctx, roots, oneFS, workers))
// A cancelled walk stops early, so its size census covers only part
// of the roots, and every file it never reached looks vanished to
// the update phase. Defence in depth rather than the only barrier:
// that phase would today fail on its first BeginTx with the same
// cancelled context before deleting anything. But it is the barrier
// that survives a later decision to let an interrupted scan commit
// what it has, and it turns a confusing failure deep in the update
// phase into a clean abort at the phase boundary.
// of the roots, and every file it never reached would look vanished
// to the update phase. Stop before anything is written or deleted.
err = ctx.Err()
if err != nil {
return s.st, err
return err
}
s.partition(changed, unhashed)
err = s.hashPhase(ctx, workers)
if err != nil {
return s.st, err
return err
}
err = s.updatePhase(ctx)
if err != nil {
return s.st, err
return err
}
return s.st, s.contentPhase(ctx, workers)
return s.contentPhase(ctx, workers)
}
// walkableRoots returns the operands the walk can start from: regular
// files, and directories not named .zfs. Every other operand is warned
// about, counted as skipped, and dropped. A dropped operand is no
// longer a root, so the records stored beneath it count as outside the
// roots and are not deleted as unverified, unless it lies under another
// root. An operand that fails lstat here is kept, and the walk warns
// about it.
func (s *scanState) walkableRoots(roots []string) []string {
kept := make([]string, 0, len(roots))
for _, root := range roots {
fi, err := os.Lstat(root)
if err == nil {
warn := operandWarning(root, fi)
if warn != "" {
s.st.skipped++
fmt.Fprintln(os.Stderr, escapePath(warn))
continue
}
}
kept = append(kept, root)
}
return kept
}
// loadIndex indexes the database records under the scan roots for
@@ -306,6 +411,7 @@ func (s *scanState) walkPhase(
}
s.sizes = append(s.sizes, ev.rec.size)
s.st.walked++
prog.increment()
@@ -509,16 +615,22 @@ func (s *scanState) recordRun(ctx context.Context, r hashResult) error {
}
// commitFullBatch commits the running batch once it holds
// updateBatchSize records.
// updateBatchSize records. A batch that fails to commit is kept: the
// commit fails when the scan is interrupted, and syncScan then commits
// the batch itself.
func (s *scanState) commitFullBatch(ctx context.Context) error {
if len(s.batch) < updateBatchSize {
return nil
}
err := applyBatch(ctx, s.db, s.batch, nil, nil)
if err != nil {
return err
}
s.batch = s.batch[:0]
return err
return nil
}
// updatePhase writes the scan's tail under one progress display: the
@@ -789,11 +901,43 @@ func sendEvent(ctx context.Context, events chan<- walkEvent,
}
}
// operandWarning returns the one-line warning for an operand the walk
// does not start from, naming the path and what it is, or "" for one it
// does: a regular file, or a directory not named .zfs. Symlinks are
// never followed, including as operands.
func operandWarning(root string, fi fs.FileInfo) string {
var kind string
switch mode := fi.Mode(); {
case mode.IsRegular():
return ""
case mode.IsDir():
if filepath.Base(root) != ".zfs" {
return ""
}
kind = ".zfs directory"
case mode&fs.ModeSymlink != 0:
kind = "symlink"
case mode&fs.ModeSocket != 0:
kind = "socket"
case mode&fs.ModeNamedPipe != 0:
kind = "FIFO"
case mode&fs.ModeDevice != 0:
kind = "device node"
default:
kind = "non-regular file"
}
return fmt.Sprintf("walk %s: skipping %s operand", root, kind)
}
// seedRoot turns one PATH operand into the walk's starting state: a
// regular-file operand is statted and emitted directly, a directory
// operand becomes an initial job, and a symlink or other non-regular
// operand yields nothing (symlinks are never followed, including as
// operands).
// regular-file operand is statted and emitted directly, and a directory
// operand becomes an initial job. walkableRoots has already dropped
// every other operand. One that has changed into something else since
// is warned about and skipped here; it is still a root, so the records
// stored beneath it are deleted as unverified.
func seedRoot(ctx context.Context, root string,
events chan<- walkEvent,
) []dirJob {
@@ -807,30 +951,30 @@ func seedRoot(ctx context.Context, root string,
return nil
}
switch {
case fi.IsDir():
if filepath.Base(root) == ".zfs" {
return nil
}
warn := operandWarning(root, fi)
if warn != "" {
sendEvent(ctx, events, walkEvent{warn: warn, fail: true})
return nil
}
if fi.IsDir() {
dev, ok := deviceOfInfo(fi)
return []dirJob{{path: root, rootDev: dev, rootDevOK: ok}}
case fi.Mode().IsRegular():
dev, ino := inodeOfInfo(fi)
sendEvent(ctx, events, walkEvent{rec: fileRec{
path: root,
size: fi.Size(),
mtime: fi.ModTime().Unix(),
dev: dev,
ino: ino,
}})
return nil
default:
return nil
}
dev, ino := inodeOfInfo(fi)
sendEvent(ctx, events, walkEvent{rec: fileRec{
path: root,
size: fi.Size(),
mtime: fi.ModTime().Unix(),
dev: dev,
ino: ino,
}})
return nil
}
// startWalkWorkers starts the walk worker pool. Each worker processes
@@ -929,6 +1073,12 @@ func walkOneDir(ctx context.Context, job dirJob, oneFS bool,
var subs []dirJob
for _, e := range entries {
// A cancelled scan wants nothing more from this directory: stop
// rather than lstat the rest of a large one.
if ctx.Err() != nil {
return nil
}
p := filepath.Join(job.path, e.Name())
if e.IsDir() {
+211 -43
View File
@@ -6,13 +6,18 @@ import (
"crypto/sha256"
"database/sql"
"encoding/hex"
"errors"
"fmt"
"io"
"io/fs"
"os"
"os/signal"
"path/filepath"
"runtime"
"slices"
"strconv"
"strings"
"syscall"
"testing"
"time"
)
@@ -399,7 +404,7 @@ func TestScanContentGate(t *testing.T) {
}
}
groups := collectDupeGroups(recs)
groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, same) {
t.Fatalf("groups = %+v, want only the identical pair %q",
groups, same)
@@ -434,7 +439,7 @@ func TestScanContentAcrossOperands(t *testing.T) {
want := []string{a, b}
slices.Sort(want)
groups := collectDupeGroups(recs)
groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
t.Fatalf("groups = %+v, want the pair %q", groups, want)
}
@@ -455,13 +460,13 @@ func TestScanContentWithinOperand(t *testing.T) {
added := sparseFile(t, dir, "d2", headTailMin)
st := syncTree(t, db, dir)
if st != (scanStats{added: 1, unchanged: 2}) {
if st != (scanStats{walked: 3, added: 1, unchanged: 2}) {
t.Fatalf("rescan stats = %+v, want 1 added 2 unchanged", st)
}
want := []string{stored, added}
groups := collectDupeGroups(dbRecords(t, db))
groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
t.Fatalf("groups = %+v, want the pair %q", groups, want)
}
@@ -501,7 +506,7 @@ func TestScanContentStalePartners(t *testing.T) {
sparseFile(t, dirB, "changed-copy", headTailMin+1)
st := syncTree(t, db, dirB)
if st != (scanStats{added: 2}) {
if st != (scanStats{walked: 2, added: 2}) {
t.Errorf("stats = %+v, want 2 added and nothing skipped", st)
}
@@ -519,7 +524,7 @@ func TestScanContentStalePartners(t *testing.T) {
}
}
if groups := collectDupeGroups(recs); len(groups) != 0 {
if groups := dupeGroups(t, db); len(groups) != 0 {
t.Errorf("groups = %+v, want none", groups)
}
}
@@ -558,7 +563,7 @@ func TestScanContentHashedStalePartners(t *testing.T) {
b := sparseFile(t, t.TempDir(), "copy", headTailMin)
st := syncTree(t, db, filepath.Dir(b))
if st != (scanStats{added: 1}) {
if st != (scanStats{walked: 1, added: 1}) {
t.Errorf("stats = %+v, want 1 added and nothing skipped", st)
}
@@ -570,7 +575,7 @@ func TestScanContentHashedStalePartners(t *testing.T) {
// The stored records lie outside the operand and are left as they
// are, so they still group with each other, but not with the copy.
groups := collectDupeGroups(recs)
groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, stored) {
t.Errorf("groups = %+v, want only the stored pair %q", groups, stored)
}
@@ -598,7 +603,7 @@ func TestScanContentReadFailure(t *testing.T) {
b := sparseFile(t, dirB, "b", headTailMin)
st := syncTree(t, db, dirB)
if st != (scanStats{added: 1, skipped: 1}) {
if st != (scanStats{walked: 1, added: 1, skipped: 1}) {
t.Fatalf("stats = %+v, want 1 added 1 skipped", st)
}
@@ -612,14 +617,14 @@ func TestScanContentReadFailure(t *testing.T) {
}
st = syncTree(t, db, dirB)
if st != (scanStats{unchanged: 1}) {
if st != (scanStats{walked: 1, unchanged: 1}) {
t.Fatalf("rescan stats = %+v, want 1 unchanged", st)
}
want := []string{a, b}
slices.Sort(want)
groups := collectDupeGroups(dbRecords(t, db))
groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
t.Fatalf("groups = %+v, want the pair %q after the retry",
groups, want)
@@ -658,7 +663,7 @@ func TestScanContentCheckError(t *testing.T) {
b := sparseFile(t, t.TempDir(), "b", headTailMin)
st := syncTree(t, db, filepath.Dir(b))
if st != (scanStats{added: 1, skipped: 1}) {
if st != (scanStats{walked: 1, added: 1, skipped: 1}) {
t.Fatalf("stats = %+v, want 1 added 1 skipped", st)
}
@@ -686,7 +691,7 @@ func TestScanContentHardlinks(t *testing.T) {
c := sparseFile(t, dir, "copy", headTailMin)
st := syncTree(t, db, dir)
if st != (scanStats{added: 3}) {
if st != (scanStats{walked: 3, added: 3}) {
t.Fatalf("stats = %+v, want 3 added", st)
}
@@ -856,9 +861,11 @@ func TestWalkFileAndSymlinkOperands(t *testing.T) {
t.Fatalf("file operand: recs = %+v, errs = %d", recs, errs)
}
// A symlink operand is not followed and yields nothing.
// A symlink operand that reaches the walk (it became one after
// walkableRoots checked it) is not followed: it yields a warning and
// no records.
recs, errs = collectWalk(t, []string{link}, false, 2)
if errs != 0 || len(recs) != 0 {
if errs != 1 || len(recs) != 0 {
t.Fatalf("symlink operand: recs = %+v, errs = %d", recs, errs)
}
}
@@ -907,6 +914,159 @@ func TestDeviceOfInfo(t *testing.T) {
}
}
// dirEntryFor returns the fs.DirEntry for name within dir, obtained via
// the same os.ReadDir the walk uses, so it carries a real Info().
func dirEntryFor(t *testing.T, dir, name string) fs.DirEntry {
t.Helper()
entries, err := os.ReadDir(dir)
if err != nil {
t.Fatal(err)
}
for _, e := range entries {
if e.Name() == name {
return e
}
}
t.Fatalf("entry %q not found in %q", name, dir)
return nil
}
// callSubdirJob runs subdirJob against e under parent, collecting any
// warning events it emits (subdirJob emits at most one).
func callSubdirJob(t *testing.T, p string, e fs.DirEntry,
parent dirJob, oneFS bool,
) (dirJob, bool, []walkEvent) {
t.Helper()
events := make(chan walkEvent, 1)
job, ok := subdirJob(t.Context(), p, e, parent, oneFS, events)
close(events)
var evs []walkEvent
for ev := range events {
evs = append(evs, ev)
}
return job, ok, evs
}
// TestSubdirJobOneFilesystem exercises the -x boundary check in
// subdirJob directly, so no second real filesystem is needed. The
// subdirectory's real device is compared against a fabricated operand
// device.
func TestSubdirJobOneFilesystem(t *testing.T) {
t.Parallel()
dir := t.TempDir()
sub := filepath.Join(dir, "sub")
err := os.Mkdir(sub, 0o750)
if err != nil {
t.Fatal(err)
}
info, err := os.Lstat(sub)
if err != nil {
t.Fatal(err)
}
dev, ok := deviceOfInfo(info)
if !ok {
t.Skip("platform exposes no device id")
}
// A device the subdirectory is not on, standing in for an operand
// rooted on a different filesystem.
otherDev := dev + 1
e := dirEntryFor(t, dir, "sub")
cases := []struct {
name string
oneFS bool
parent dirJob
wantOK bool
}{
// -x on, subdirectory on a different device than its operand:
// descent is refused.
{"reject across boundary", true,
dirJob{rootDev: otherDev, rootDevOK: true}, false},
// -x on but the operand's own device is unknown: the boundary
// check is bypassed and descent proceeds.
{"bypass when root device unknown", true,
dirJob{rootDev: otherDev, rootDevOK: false}, true},
// Default (no -x): boundaries are crossed even onto a different
// device.
{"cross by default", false,
dirJob{rootDev: otherDev, rootDevOK: true}, true},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
job, ok, evs := callSubdirJob(t, sub, e, tc.parent, tc.oneFS)
if ok != tc.wantOK {
t.Fatalf("accepted = %v, want %v", ok, tc.wantOK)
}
if len(evs) != 0 {
t.Fatalf("unexpected events: %+v", evs)
}
// The accepted job must carry the operand's device down, or -x
// stops checking below the first level.
want := dirJob{
path: sub,
rootDev: tc.parent.rootDev,
rootDevOK: tc.parent.rootDevOK,
}
if ok && job != want {
t.Fatalf("job = %+v, want %+v", job, want)
}
})
}
}
// errInfoUnavailable is returned by errDirEntry.Info().
var errInfoUnavailable = errors.New("info unavailable")
// errDirEntry is a directory entry whose Info() always fails, driving
// subdirJob's stat-error branch deterministically.
type errDirEntry struct{ name string }
func (e errDirEntry) Name() string { return e.name }
func (errDirEntry) IsDir() bool { return true }
func (errDirEntry) Type() fs.FileMode {
return fs.ModeDir
}
func (errDirEntry) Info() (fs.FileInfo, error) {
return nil, errInfoUnavailable
}
// TestSubdirJobStatError asserts that when a subdirectory's Info()
// fails under -x, subdirJob warns and refuses descent.
func TestSubdirJobStatError(t *testing.T) {
t.Parallel()
p := "/does/not/matter/sub"
_, ok, evs := callSubdirJob(t, p, errDirEntry{name: "sub"},
dirJob{rootDev: 1, rootDevOK: true}, true)
if ok {
t.Fatal("descent accepted after stat error, want refused")
}
if len(evs) != 1 || !evs[0].fail || !strings.Contains(evs[0].warn, p) {
t.Fatalf("want one warning naming the path, got %+v", evs)
}
}
// buildSmokeTree recreates the README smoke-test filesystem layout
// with deterministic content and returns the tree root.
func buildSmokeTree(t *testing.T) string {
@@ -953,11 +1113,16 @@ func syncTree(t *testing.T, db *sql.DB, roots ...string) scanStats {
return st
}
// dbRecords returns every record currently in the database.
// dbRecords returns every record currently in the database, in path
// order.
func dbRecords(t *testing.T, db *sql.DB) []scanRec {
t.Helper()
recs, err := loadFileRows(t.Context(), db)
var recs []scanRec
err := loadFileRows(t.Context(), db, func(r scanRec) {
recs = append(recs, r)
})
if err != nil {
t.Fatal(err)
}
@@ -994,10 +1159,10 @@ func recordPaths(recs []scanRec) []string {
// assertSmokeDupeGroups checks the file-level duplicate groups for the
// smoke tree rooted at dir.
func assertSmokeDupeGroups(t *testing.T, dir string, parsed []scanRec) {
func assertSmokeDupeGroups(t *testing.T, dir string, db *sql.DB) {
t.Helper()
groups := collectDupeGroups(parsed)
groups := dupeGroups(t, db)
if len(groups) != 5 {
t.Fatalf("len(groups) = %d, want 5", len(groups))
}
@@ -1022,11 +1187,10 @@ func assertSmokeDupeGroups(t *testing.T, dir string, parsed []scanRec) {
// assertSmokeTreeGroups checks the duplicate-tree groups for the smoke
// tree rooted at dir.
func assertSmokeTreeGroups(t *testing.T, dir string, parsed []scanRec) {
func assertSmokeTreeGroups(t *testing.T, dir string, db *sql.DB) {
t.Helper()
super, dirs := buildHierarchy(parsed)
super.compute()
super, dirs := dbTree(t, db)
tg := collectTreeGroups(dirs, super)
if len(tg) != 1 {
@@ -1051,7 +1215,7 @@ func TestScanPipeline(t *testing.T) {
db := openTestDB(t)
st := syncTree(t, db, dir)
if st != (scanStats{added: smokeTreeFiles}) {
if st != (scanStats{walked: smokeTreeFiles, added: smokeTreeFiles}) {
t.Fatalf("stats = %+v, want %d added only", st, smokeTreeFiles)
}
@@ -1060,8 +1224,8 @@ func TestScanPipeline(t *testing.T) {
t.Fatalf("len(records) = %d, want %d", len(parsed), smokeTreeFiles)
}
assertSmokeDupeGroups(t, dir, parsed)
assertSmokeTreeGroups(t, dir, parsed)
assertSmokeDupeGroups(t, dir, db)
assertSmokeTreeGroups(t, dir, db)
}
func TestSyncScanUnchangedReuse(t *testing.T) {
@@ -1074,7 +1238,7 @@ func TestSyncScanUnchangedReuse(t *testing.T) {
writeFile(t, dir, "b.bin", pattern(2, 600))
st := syncTree(t, db, dir)
if st != (scanStats{added: 2}) {
if st != (scanStats{walked: 2, added: 2}) {
t.Fatalf("first scan stats = %+v, want 2 added", st)
}
@@ -1088,7 +1252,7 @@ func TestSyncScanUnchangedReuse(t *testing.T) {
}
st = syncTree(t, db, dir)
if st != (scanStats{unchanged: 2}) {
if st != (scanStats{walked: 2, unchanged: 2}) {
t.Fatalf("rescan stats = %+v, want 2 unchanged", st)
}
@@ -1117,7 +1281,7 @@ func TestSyncScanMtimeBump(t *testing.T) {
}
st := syncTree(t, db, dir)
if st != (scanStats{updated: 1}) {
if st != (scanStats{walked: 1, updated: 1}) {
t.Fatalf("mtime-bump stats = %+v, want 1 updated", st)
}
@@ -1145,7 +1309,7 @@ func TestSyncScanAddRemove(t *testing.T) {
}
st := syncTree(t, db, dir)
if st != (scanStats{added: 1, removed: 1, unchanged: 1}) {
if st != (scanStats{walked: 2, added: 1, removed: 1, unchanged: 1}) {
t.Fatalf("add/remove stats = %+v, want 1 added 1 removed 1 unchanged",
st)
}
@@ -1262,7 +1426,7 @@ func TestSyncScanOverlappingRoots(t *testing.T) {
// A file reachable via two overlapping operands is deduplicated
// by path in the shared walk and processed once.
st := syncTree(t, db, dir, filepath.Join(dir, "sub"))
if st != (scanStats{added: 1}) {
if st != (scanStats{walked: 1, added: 1}) {
t.Fatalf("stats = %+v, want 1 added", st)
}
@@ -1283,7 +1447,7 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
// Neither size is shared, so neither file is read: both records
// are written without hashes and no duplicates are reported.
st := syncTree(t, db, dir)
if st != (scanStats{added: 2}) {
if st != (scanStats{walked: 2, added: 2}) {
t.Fatalf("stats = %+v, want 2 added", st)
}
@@ -1295,7 +1459,7 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
}
}
if groups := collectDupeGroups(recs); len(groups) != 0 {
if groups := dupeGroups(t, db); len(groups) != 0 {
t.Fatalf("groups = %+v, want none from unhashed records", groups)
}
@@ -1305,12 +1469,12 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
c := writeFile(t, dir, "c.bin", pattern(1, 500))
st = syncTree(t, db, dir)
if st != (scanStats{added: 1, updated: 1, unchanged: 1}) {
if st != (scanStats{walked: 3, added: 1, updated: 1, unchanged: 1}) {
t.Fatalf("rescan stats = %+v, want 1 added 1 updated 1 unchanged",
st)
}
groups := collectDupeGroups(dbRecords(t, db))
groups := dupeGroups(t, db)
if len(groups) != 1 {
t.Fatalf("groups = %+v, want the a/c pair", groups)
}
@@ -1346,8 +1510,7 @@ func TestTreesUnhashedNeverEqual(t *testing.T) {
}
for name, unknown := range cases {
super, dirs := buildHierarchy(append(slices.Clone(shared), unknown...))
super.compute()
super, dirs := treeOf(t, append(slices.Clone(shared), unknown...))
if tg := collectTreeGroups(dirs, super); len(tg) != 0 {
t.Errorf("%s: tree groups = %d, want 0 (the files may differ)",
@@ -1370,7 +1533,7 @@ func TestScanHardlinksReadOnce(t *testing.T) {
}
st := syncTree(t, db, dir)
if st != (scanStats{added: 2}) {
if st != (scanStats{walked: 2, added: 2}) {
t.Fatalf("stats = %+v, want 2 added", st)
}
@@ -1384,7 +1547,7 @@ func TestScanHardlinksReadOnce(t *testing.T) {
t.Fatalf("hardlink hashes differ: %+v vs %+v", ra, rb)
}
if groups := collectDupeGroups(recs); len(groups) != 1 {
if groups := dupeGroups(t, db); len(groups) != 1 {
t.Fatalf("groups = %+v, want the hardlink pair", groups)
}
}
@@ -1492,6 +1655,13 @@ func injectWriteFailure(t *testing.T, path string) {
func baselineGoroutines(t *testing.T) int {
t.Helper()
// The first scan command in a process starts os/signal's goroutine,
// which never exits. Start it now, so the baseline counts it
// instead of the scan seeming to leave it behind.
ch := make(chan os.Signal, 1)
signal.Notify(ch, syscall.SIGINT)
signal.Stop(ch)
deadline := time.Now().Add(goroutineSettle)
last := runtime.NumGoroutine()
@@ -1548,7 +1718,7 @@ func TestScanHashWriteFailureUnwindsPool(t *testing.T) {
code := run([]string{
cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir,
}, &stderr)
}, io.Discard, &stderr)
if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
code, exitFatal, stderr.String())
@@ -1625,10 +1795,8 @@ func TestReportsNeverTouchFilesystem(t *testing.T) {
t.Fatal(err)
}
recs := dbRecords(t, db)
assertSmokeDupeGroups(t, dir, recs)
assertSmokeTreeGroups(t, dir, recs)
assertSmokeDupeGroups(t, dir, db)
assertSmokeTreeGroups(t, dir, db)
}
func TestUnderRoot(t *testing.T) {
+156 -98
View File
@@ -5,49 +5,59 @@ import (
"context"
"crypto/sha256"
"fmt"
"io"
"os"
"slices"
"strconv"
"strings"
)
// fileSig is a file's duplicate signature; mtime is excluded.
type fileSig struct {
size int64
head string
tail string
content string
}
// treeNode is one directory reconstructed from the scan stream.
// treeNode is one directory reconstructed from the record paths.
type treeNode struct {
path string
parent *treeNode
dirs map[string]*treeNode
files map[string]fileSig
path string
parent *treeNode
// entries holds the serialized child entries until the digest is
// computed from them, and is then dropped.
entries []string
digest [sha256.Size]byte
fileCount int64
totalSize int64
}
// runTrees implements the trees subcommand: it reads every record from
// the database, reconstructs the directory hierarchy from the record
// paths, computes a Merkle-style digest per directory, and prints
// maximal duplicate-tree groups as TSV on stdout. It never touches the
// scanned filesystem; its only I/O is the database, stdout, and
// stderr.
func runTrees(ctx context.Context) error {
recs, err := loadRecords(ctx)
// the database in path order, reconstructs the directory hierarchy from
// the record paths, computes a Merkle-style digest per directory, and
// prints maximal duplicate-tree groups as TSV on stdout. It never
// touches the scanned filesystem; its only I/O is the database, stdout,
// and stderr. Any database problem, including a missing database, is
// fatal.
func runTrees(ctx context.Context, stdout io.Writer) error {
dbPath := databasePath()
db, err := openReportDatabase(ctx, dbPath)
if err != nil {
return err
}
super, allDirs := buildHierarchy(recs)
super.compute()
defer func() { _ = db.Close() }()
records := 0
tree := newTreeBuilder()
err = loadFileRows(ctx, db, func(r scanRec) {
records++
tree.add(r)
})
if err != nil {
return fmt.Errorf("database %s: %w", dbPath, err)
}
super, allDirs := tree.finish()
dupes := collectTreeGroups(allDirs, super)
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
out := bufio.NewWriterSize(stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
if err != nil {
@@ -62,7 +72,8 @@ func runTrees(ctx context.Context) error {
first := g[0]
for _, n := range g[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n",
first.path, n.path, first.fileCount, first.totalSize)
escapePath(first.path), escapePath(n.path),
first.fileCount, first.totalSize)
if err != nil {
return fmt.Errorf("write stdout: %w", err)
}
@@ -80,65 +91,128 @@ func runTrees(ctx context.Context) error {
fmt.Fprintf(os.Stderr,
"trees: %d records read, %d duplicate tree groups, %d dupe trees, "+
"%s reclaimable\n",
len(recs), len(dupes), dupeTrees, humanBytes(reclaimable))
records, len(dupes), dupeTrees, humanBytes(reclaimable))
return nil
}
// buildHierarchy reconstructs the directory hierarchy from the record
// paths under a synthetic super-root. Paths are split on "/"; for
// absolute paths the first component is empty, which simply becomes a
// top-level node representing "/". It returns the super-root and every
// directory node created.
func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
// treeBuilder reconstructs the directory hierarchy from records added
// in path order, under a synthetic super-root. Paths are split on "/";
// for absolute paths the first component is empty, which becomes the
// top-level directory with path "/". In path order all the paths under
// one directory come together, so a directory is complete once a path
// outside it is added: its digest is computed then and its entries are
// dropped. Only the directories holding the latest path keep entries.
type treeBuilder struct {
super *treeNode
// open lists the directories holding the latest path, outermost
// first, starting with the super-root; names[i] is open[i]'s name.
open []*treeNode
names []string
// dirs lists every completed directory.
dirs []*treeNode
}
func newTreeBuilder() *treeBuilder {
super := &treeNode{}
var allDirs []*treeNode
return &treeBuilder{
super: super,
open: []*treeNode{super},
names: []string{""},
}
}
for _, r := range recs {
comps := strings.Split(r.path, "/")
// add adds one record. Each record must come after the previous one in
// path order (byte order); otherwise a completed directory would be
// started again as a second directory with the same path.
func (b *treeBuilder) add(r scanRec) {
comps := strings.Split(r.path, "/")
dirNames, name := comps[:len(comps)-1], comps[len(comps)-1]
node := super
for _, c := range comps[:len(comps)-1] {
child := node.dirs[c]
if child == nil {
childPath := c
if node != super {
childPath = node.path + "/" + c
}
child = &treeNode{path: childPath, parent: node}
if node.dirs == nil {
node.dirs = make(map[string]*treeNode)
}
node.dirs[c] = child
allDirs = append(allDirs, child)
}
node = child
}
if node.files == nil {
node.files = make(map[string]fileSig)
}
sig := fileSig{
size: r.size, head: r.head, tail: r.tail, content: r.content,
}
// A record without a content hash has unknown content (README
// "Database"): give it a signature no other file can share, so
// trees containing it never compare equal. Real hashes are
// hex, so the NUL-prefixed form cannot collide.
if sig.content == "" {
sig.content = "unhashed\x00" + r.path
}
node.files[comps[len(comps)-1]] = sig
// Keep the open directories that hold this path; complete the rest.
depth := 1
for depth < len(b.open) && depth <= len(dirNames) &&
b.names[depth] == dirNames[depth-1] {
depth++
}
return super, allDirs
b.closeTo(depth)
for _, c := range dirNames[depth-1:] {
b.openDir(c)
}
dir := b.open[len(b.open)-1]
dir.entries = append(dir.entries, fileEntry(name, r))
dir.fileCount++
dir.totalSize += r.size
}
// openDir starts the directory called name inside the innermost open
// one.
func (b *treeBuilder) openDir(name string) {
parent := b.open[len(b.open)-1]
path := parent.path + "/" + name
// The root directory's path is "/", not empty, and its children's
// paths start with one slash, not two.
switch {
case parent == b.super && name == "":
path = "/"
case parent == b.super:
path = name
case parent.path == "/":
path = "/" + name
}
b.open = append(b.open, &treeNode{path: path, parent: parent})
b.names = append(b.names, name)
}
// closeTo completes the open directories after the first n, innermost
// first: each one's digest is computed and entered in its parent along
// with its totals.
func (b *treeBuilder) closeTo(n int) {
for len(b.open) > n {
last := len(b.open) - 1
dir, name := b.open[last], b.names[last]
b.open, b.names = b.open[:last], b.names[:last]
dir.computeDigest()
dir.parent.entries = append(dir.parent.entries,
"d\x00"+name+"\x00"+string(dir.digest[:]))
dir.parent.fileCount += dir.fileCount
dir.parent.totalSize += dir.totalSize
b.dirs = append(b.dirs, dir)
}
}
// finish completes every open directory and returns the super-root and
// every directory.
func (b *treeBuilder) finish() (*treeNode, []*treeNode) {
b.closeTo(1)
return b.super, b.dirs
}
// fileEntry serializes a file child for its directory's digest: its
// name and its signature (size, head, tail, content); mtime is
// excluded.
func fileEntry(name string, r scanRec) string {
content := r.content
// A record without a content hash has unknown content (README
// "Database"): give it a signature no other file can share, so
// trees containing it never compare equal. Real hashes are hex, so
// the NUL-prefixed form cannot collide.
if content == "" {
content = "unhashed\x00" + r.path
}
return "f\x00" + name + "\x00" + strconv.FormatInt(r.size, 10) +
"\x00" + r.head + "\x00" + r.tail + "\x00" + content
}
// collectTreeGroups groups directories by digest and returns every
@@ -180,38 +254,22 @@ func collectTreeGroups(allDirs []*treeNode, super *treeNode) [][]*treeNode {
return dupes
}
// compute fills in digest, fileCount, and totalSize for n and all of
// its descendants. A directory's digest is the SHA-256 of its child
// entries — files serialized with name and signature, subdirectories
// with name and recursive digest — sorted byte-lexicographically.
// Filenames cannot contain NUL or "/", so NUL delimiters are
// unambiguous.
func (n *treeNode) compute() {
entries := make([]string, 0, len(n.dirs)+len(n.files))
for name, sig := range n.files {
entries = append(entries,
"f\x00"+name+"\x00"+strconv.FormatInt(sig.size, 10)+
"\x00"+sig.head+"\x00"+sig.tail+"\x00"+sig.content)
n.fileCount++
n.totalSize += sig.size
}
for name, child := range n.dirs {
child.compute()
entries = append(entries, "d\x00"+name+"\x00"+string(child.digest[:]))
n.fileCount += child.fileCount
n.totalSize += child.totalSize
}
slices.Sort(entries)
// computeDigest sets n's digest and drops its entries. A directory's
// digest is the SHA-256 of its child entries — files serialized with
// name and signature, subdirectories with name and recursive digest —
// sorted byte-lexicographically. Filenames cannot contain NUL or "/",
// so NUL delimiters are unambiguous.
func (n *treeNode) computeDigest() {
slices.Sort(n.entries)
h := sha256.New()
for _, e := range entries {
for _, e := range n.entries {
h.Write([]byte(e))
h.Write([]byte{0})
}
copy(n.digest[:], h.Sum(nil))
n.entries = nil
}
// suppressed reports whether a duplicate-tree group is non-maximal: its
+129 -17
View File
@@ -1,6 +1,8 @@
package main
import (
"bytes"
"database/sql"
"slices"
"testing"
)
@@ -29,6 +31,36 @@ func smokeTreeRecs() []scanRec {
}
}
// dbTree builds the directory hierarchy from the records in db the way
// trees does, and returns the super-root and every directory.
func dbTree(t *testing.T, db *sql.DB) (*treeNode, []*treeNode) {
t.Helper()
tree := newTreeBuilder()
err := loadFileRows(t.Context(), db, tree.add)
if err != nil {
t.Fatal(err)
}
return tree.finish()
}
// treeOf writes recs into a fresh database and builds the directory
// hierarchy from it the way trees does.
func treeOf(t *testing.T, recs []scanRec) (*treeNode, []*treeNode) {
t.Helper()
db := openTestDB(t)
err := applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
return dbTree(t, db)
}
// nodeByPath finds the directory node with the given path.
func nodeByPath(t *testing.T, dirs []*treeNode, path string) *treeNode {
t.Helper()
@@ -59,11 +91,10 @@ func groupPaths(groups [][]*treeNode) [][]string {
return out
}
func TestBuildHierarchyCounts(t *testing.T) {
func TestTreeCounts(t *testing.T) {
t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs())
super.compute()
_, dirs := treeOf(t, smokeTreeRecs())
d := nodeByPath(t, dirs, "/d")
if d.fileCount != 6 || d.totalSize != 9300 {
@@ -84,11 +115,98 @@ func TestBuildHierarchyCounts(t *testing.T) {
}
}
func TestTreeRootPath(t *testing.T) {
t.Parallel()
// The root directory's path is "/", never empty, and its
// children's paths start with a single slash.
_, dirs := treeOf(t, []scanRec{{path: "/f"}, {path: "/srv/g"}})
got := make([]string, 0, len(dirs))
for _, d := range dirs {
got = append(got, d.path)
}
slices.Sort(got)
want := []string{"/", "/srv"}
if !slices.Equal(got, want) {
t.Fatalf("directory paths = %q, want %q", got, want)
}
}
func TestTreeNamesSortingBeforeSlash(t *testing.T) {
t.Parallel()
// In path order "/a/b-x/f" and "/a/b.txt" come between the file
// "/a/b" and "/a/b/f", because "-" and "." sort before "/". Each
// directory must still be built once, whole, so /a matches /c.
recs := make([]scanRec, 0, 8)
for _, top := range []string{"/a", "/c"} {
for _, p := range []string{"/b", "/b-x/f", "/b.txt", "/b/f"} {
content := "c"
if p == "/b-x/f" {
content = "other"
}
recs = append(recs, scanRec{
size: 1, head: "h", tail: "t", content: content, path: top + p,
})
}
}
super, dirs := treeOf(t, recs)
got := make([]string, 0, len(dirs))
for _, d := range dirs {
got = append(got, d.path)
}
slices.Sort(got)
want := []string{"/", "/a", "/a/b", "/a/b-x", "/c", "/c/b", "/c/b-x"}
if !slices.Equal(got, want) {
t.Fatalf("directory paths = %q, want %q", got, want)
}
groups := collectTreeGroups(dirs, super)
gotGroups := groupPaths(groups)
wantGroups := [][]string{{"/a", "/c"}}
if !slices.EqualFunc(gotGroups, wantGroups, slices.Equal) {
t.Fatalf("groups = %v, want %v", gotGroups, wantGroups)
}
if groups[0][0].fileCount != 4 || groups[0][0].totalSize != 4 {
t.Errorf("group totals: %d files %d bytes, want 4 4",
groups[0][0].fileCount, groups[0][0].totalSize)
}
}
func TestRunTreesEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
code := run([]string{cmdTrees}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tfiles\tsize\n" +
`/d/\tone\ntwo\rthree\\four` + "\t/d/A\t1\t5\n"
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
func TestTreeDigests(t *testing.T) {
t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs())
super.compute()
_, dirs := treeOf(t, smokeTreeRecs())
t1 := nodeByPath(t, dirs, "/d/t1")
t2 := nodeByPath(t, dirs, "/d/t2")
@@ -121,8 +239,7 @@ func TestTreeDigestContentSensitivity(t *testing.T) {
{size: 10, head: "DIFF", tail: sharedTail, content: "c", path: "/r/b/f"},
}
super, dirs := buildHierarchy(recs)
super.compute()
_, dirs := treeOf(t, recs)
a := nodeByPath(t, dirs, "/r/a")
b := nodeByPath(t, dirs, "/r/b")
@@ -135,8 +252,7 @@ func TestTreeDigestContentSensitivity(t *testing.T) {
func TestCollectTreeGroupsMaximal(t *testing.T) {
t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs())
super.compute()
super, dirs := treeOf(t, smokeTreeRecs())
groups := collectTreeGroups(dirs, super)
@@ -160,16 +276,14 @@ func TestCollectTreeGroupsDeterministic(t *testing.T) {
recs := smokeTreeRecs()
super, dirs := buildHierarchy(recs)
super.compute()
super, dirs := treeOf(t, recs)
forward := groupPaths(collectTreeGroups(dirs, super))
reversed := slices.Clone(recs)
slices.Reverse(reversed)
superR, dirsR := buildHierarchy(reversed)
superR.compute()
superR, dirsR := treeOf(t, reversed)
backward := groupPaths(collectTreeGroups(dirsR, superR))
if !slices.EqualFunc(forward, backward, slices.Equal) {
@@ -188,8 +302,7 @@ func TestCollectTreeGroupsSiblings(t *testing.T) {
{size: 10, head: "h", tail: "t", content: "c", path: "/p/x2/f"},
}
super, dirs := buildHierarchy(recs)
super.compute()
super, dirs := treeOf(t, recs)
got := groupPaths(collectTreeGroups(dirs, super))
@@ -211,8 +324,7 @@ func TestCollectTreeGroupsDifferingParents(t *testing.T) {
{size: 10, head: "h", tail: "t", content: "c", path: "/q/b/x/f"},
}
super, dirs := buildHierarchy(recs)
super.compute()
super, dirs := treeOf(t, recs)
got := groupPaths(collectTreeGroups(dirs, super))