Compare commits
7
Commits
main
..
bd8d41b174
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
bd8d41b174 | ||
|
|
2dd1194f33 | ||
|
|
705c8729ca | ||
|
|
01ff3bb5f0 | ||
|
|
d63d3cc7fc | ||
|
|
c887f80f57 | ||
|
|
c9bf22d483 |
+6
-1
@@ -1,4 +1,9 @@
|
||||
.git
|
||||
# .git is sent without its config. Without a VERSION build argument the
|
||||
# stage that compiles runs `git describe --tags --always` on .git, which
|
||||
# does not need .git/config; that file can hold a credential, such as a
|
||||
# password in a remote URL or the token the CI checkout step stores there.
|
||||
.git/config
|
||||
|
||||
.claude
|
||||
.DS_Store
|
||||
sfdupes
|
||||
|
||||
@@ -6,4 +6,6 @@ jobs:
|
||||
steps:
|
||||
# actions/checkout v4.2.2, 2026-02-22
|
||||
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
|
||||
with:
|
||||
fetch-depth: 0
|
||||
- run: script/cibuild
|
||||
|
||||
+13
-1
@@ -108,7 +108,19 @@ ARG CHECK_EPOCH
|
||||
RUN echo "gate test, epoch ${CHECK_EPOCH}" && make test
|
||||
RUN echo "gate fmt-check, epoch ${CHECK_EPOCH}" && make fmt-check
|
||||
|
||||
RUN make build
|
||||
# The version stamped into the binary: the VERSION build argument when
|
||||
# one is given, otherwise `git describe --tags --always` of the .git in
|
||||
# the build context (git is installed by script/bootstrap above). A
|
||||
# context that carries .git and still yields no version fails the build;
|
||||
# with neither, as from a source tarball, it is "dev".
|
||||
ARG VERSION
|
||||
RUN version="${VERSION:-$(git describe --tags --always || echo dev)}"; \
|
||||
if [ -e .git ] && { [ -z "$version" ] || [ "$version" = dev ] || \
|
||||
[ "$version" = unknown ]; }; then \
|
||||
echo "no version could be derived although the build context carries .git" >&2; \
|
||||
exit 1; \
|
||||
fi; \
|
||||
make build VERSION="$version"
|
||||
|
||||
# Runtime stage
|
||||
# alpine:3.22, 2026-07-23
|
||||
|
||||
@@ -94,7 +94,14 @@ Goals, in order:
|
||||
millions of files, ~150 TB filesystem, possibly slow or busy disks
|
||||
(ZFS pool under resilver). Holding one small record (path, size,
|
||||
mtime) per file in memory during a scan is acceptable; holding
|
||||
every file's hashes is not (they stay in the database).
|
||||
every file's hashes is not (they stay in the database). The
|
||||
reporting commands hold no file's hashes either: `report` lets
|
||||
SQLite group and order the records and writes each row as it reads
|
||||
it, so its memory does not grow with the database, and `trees`
|
||||
reads the records in path order and keeps each directory's path,
|
||||
digest and totals, plus the files of the directories holding the
|
||||
record being read, so its memory grows with the number of
|
||||
directories, not files.
|
||||
3. **Scan incrementally, analyze offline.** The expensive filesystem
|
||||
scan maintains a persistent database; an unchanged file is never
|
||||
read again on a rescan, except to compute its content hash once a
|
||||
@@ -113,8 +120,9 @@ Goals, in order:
|
||||
`sfdupes`.
|
||||
- Dependencies: standard library, `github.com/spf13/cobra` for the
|
||||
CLI, **one progress-bar library**
|
||||
(`github.com/schollz/progressbar/v3`), and **one SQLite driver**
|
||||
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled).
|
||||
(`github.com/schollz/progressbar/v3`), **one SQLite driver**
|
||||
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled), and
|
||||
`golang.org/x/sys` for `flock(2)` (the scan lock, see "Database").
|
||||
`github.com/spf13/viper` is permitted if configuration-file support
|
||||
is ever needed, but is not currently used. No other third-party
|
||||
deps.
|
||||
@@ -152,16 +160,42 @@ All three subcommands operate on a single SQLite database file:
|
||||
use. `report` and `trees` require an existing database; a missing
|
||||
database file is a fatal error (exit 1) telling the user to run
|
||||
`scan` first.
|
||||
- The database uses WAL journal mode and a busy timeout, so running a
|
||||
report while a cron `scan` is in progress is safe. The filesystem
|
||||
is authoritative; the database is an eventually-consistent
|
||||
reflection of it. Hashed records are committed in batched
|
||||
transactions while the scan is still running (keeping the WAL
|
||||
small and letting concurrent reports observe progress), so a
|
||||
report may see a scan's changes partially applied, and a scan
|
||||
that dies partway leaves a valid database holding everything
|
||||
hashed so far; the next scan skips those records and converges
|
||||
toward the filesystem.
|
||||
- Only one `scan` runs against a database at a time. For its whole
|
||||
run, `scan` holds an exclusive `flock(2)` lock on a lock file
|
||||
beside the database, named by appending `.lock` to the database
|
||||
path (`/var/lib/sfdupes/db.sqlite.lock` by default), taken before
|
||||
it walks the filesystem or opens the database. A second `scan`
|
||||
against the same database does not wait: it fails at once with a
|
||||
one-line error naming the lock file and exits 1, without walking
|
||||
anything or opening the database, and the running scan carries on.
|
||||
The lock file is created on first use, open to its owner only, and
|
||||
left in place: a leftover file blocks nothing, because the lock
|
||||
ends with the process holding it however it ends, a fatal error or
|
||||
an interrupt included, and deleting the file while a scan runs
|
||||
would let a second scan start. `report` and `trees` never take the
|
||||
lock, so they run during a scan.
|
||||
- While `scan` runs, the database is in WAL journal mode with a busy
|
||||
timeout, so running a report while a cron `scan` is in progress is
|
||||
safe. The filesystem is authoritative; the database is an
|
||||
eventually-consistent reflection of it. Hashed records are
|
||||
committed in batched transactions while the scan is still running
|
||||
(keeping the WAL small and letting concurrent reports observe
|
||||
progress), so a report may see a scan's changes partially applied,
|
||||
and a scan that dies partway leaves a valid database holding
|
||||
everything hashed so far; the next scan skips those records and
|
||||
converges toward the filesystem.
|
||||
- `scan` switches the database back to rollback-journal mode when it
|
||||
closes it, so between scans the database file alone holds the whole
|
||||
database. Each switch needs the database to itself: a `scan` that
|
||||
starts while a report is still reading waits for it up to the
|
||||
10-second busy timeout, then fails; a `scan` that ends while a
|
||||
report has the database open warns and leaves the database in WAL
|
||||
mode until the next scan.
|
||||
- `report` and `trees` open the database read-only and need only read
|
||||
access to the database file, and no write access to its directory.
|
||||
While the database is in WAL mode they also read the `-wal` and
|
||||
`-shm` files beside it, which SQLite creates with the database
|
||||
file's permissions.
|
||||
- Schema (`PRAGMA user_version` is the schema version, currently 1; a
|
||||
database with any other version is a fatal error):
|
||||
|
||||
@@ -174,6 +208,7 @@ All three subcommands operate on a single SQLite database file:
|
||||
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
|
||||
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
|
||||
) WITHOUT ROWID;
|
||||
CREATE INDEX files_signature ON files (size, head, tail, content);
|
||||
```
|
||||
|
||||
Paths are stored as BLOBs because Unix paths are raw bytes, not
|
||||
@@ -188,7 +223,8 @@ All three subcommands operate on a single SQLite database file:
|
||||
until the content phase of a scan (see "`scan` mode" below) has
|
||||
read the file. A record with an empty `content` is never part of a
|
||||
duplicate group, though it still defines the file for tree
|
||||
reconstruction.
|
||||
reconstruction. The `files_signature` index lets SQLite group the
|
||||
records by signature for `report` without sorting the whole table.
|
||||
|
||||
### Duplicate detection
|
||||
|
||||
@@ -251,6 +287,18 @@ duplicates another or lies under another is dropped before walking,
|
||||
so every file is reached exactly once and produces one database
|
||||
record.
|
||||
|
||||
An operand that is a symlink (never followed, not even as an operand),
|
||||
socket, FIFO, or device node, or a directory named `.zfs`, is not
|
||||
scanned. `scan` prints a one-line warning naming the path and what it
|
||||
is, counts it as skipped, and drops it from the scanned operands before
|
||||
reading the database. Another operand beneath it is still scanned. The
|
||||
records stored beneath it are not deleted: they are treated like any
|
||||
other record outside the scanned operands, including the content-phase
|
||||
exception below. If it lies under another operand, they are under that
|
||||
operand instead, and are deleted like any other record there that this
|
||||
scan did not verify. This is not an error: a scan whose every operand
|
||||
is dropped walks nothing and exits 0.
|
||||
|
||||
`scan` synchronizes the database with the filesystem state under the
|
||||
scanned operands:
|
||||
|
||||
@@ -356,9 +404,12 @@ during the hash phase:
|
||||
Rules for the walk:
|
||||
|
||||
- Only regular files. Skip directories, symlinks (do not follow,
|
||||
including symlink operands), sockets, FIFOs, and device nodes.
|
||||
including symlink operands), sockets, FIFOs, and device nodes. An
|
||||
operand that is a symlink, socket, FIFO, or device node is dropped
|
||||
as described in "`scan` mode" above.
|
||||
- Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs;
|
||||
walking them would list every file once per snapshot).
|
||||
walking them would list every file once per snapshot), not even
|
||||
when it is an operand; such an operand is dropped the same way.
|
||||
- Filesystem boundaries are crossed by default. With `-x`
|
||||
(long form `--one-file-system`, following the GNU `du`/`rsync`
|
||||
convention), never descend into a directory on a different
|
||||
@@ -369,9 +420,11 @@ Rules for the walk:
|
||||
path, and continue. Per-file errors never abort the run; the final
|
||||
summary reports how many were skipped. As specified above, a
|
||||
skipped path that has a database record from an earlier scan loses
|
||||
that record, unless it failed only in the content phase; an
|
||||
unreadable directory subtree likewise loses its records (accepted:
|
||||
the database mirrors what the latest scan could actually verify).
|
||||
that record, unless it failed only in the content phase, or is an
|
||||
operand dropped before the database was read that lies under no
|
||||
other operand; an unreadable directory subtree likewise loses its
|
||||
records (accepted: the database mirrors what the latest scan could
|
||||
actually verify).
|
||||
|
||||
Concurrency: the walk phase (which also stats files), the hash phase,
|
||||
and the content phase each use a worker pool of `--workers` workers
|
||||
@@ -399,9 +452,12 @@ arguments.
|
||||
|
||||
**`report` must never touch the filesystem being analyzed.** It does not
|
||||
stat, open, or otherwise access any path that appears in the records; its
|
||||
only I/O is reading the database and writing stdout/stderr. It must
|
||||
produce identical output whether or not the scanned filesystem is still
|
||||
mounted.
|
||||
only I/O is reading the database, writing stdout/stderr, and the
|
||||
temporary file SQLite sorts in when the duplicate rows do not fit in
|
||||
memory. SQLite puts that file in `$SQLITE_TMPDIR` or `$TMPDIR` when set,
|
||||
otherwise in `/var/tmp` (or `/tmp`), and deletes it as soon as it has
|
||||
opened it. `report` must produce identical output whether or not the
|
||||
scanned filesystem is still mounted.
|
||||
|
||||
Processing:
|
||||
|
||||
@@ -428,6 +484,15 @@ first dupe size
|
||||
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296
|
||||
```
|
||||
|
||||
Paths are raw bytes and may hold any byte except NUL, so the path
|
||||
columns (`first` and `dupe`) are escaped to keep every row one line of
|
||||
tab-separated fields: a backslash is written as `\\`, a tab as `\t`, a
|
||||
newline as `\n`, and a carriage return as `\r`. Every other byte is
|
||||
written unchanged, including bytes that are not valid UTF-8. Undoing
|
||||
those four escapes gives back the stored path. Grouping and ordering
|
||||
use the stored path, not the escaped one. The warnings `scan` prints on
|
||||
stderr are escaped the same way, so each warning is one line.
|
||||
|
||||
Summary to stderr: records read, number of duplicate groups, number of
|
||||
dupe files, and total reclaimable bytes (sum of `size` over all dupe
|
||||
rows) in human units.
|
||||
@@ -499,6 +564,9 @@ first dupe files size
|
||||
/srv/a/project /srv/backup/project 3417 104857600
|
||||
```
|
||||
|
||||
The `first` and `dupe` paths are escaped as described under "Report
|
||||
output format". The root directory's path is `/`.
|
||||
|
||||
Summary to stderr: records read, number of duplicate-tree groups,
|
||||
number of dupe trees, and total reclaimable bytes (sum of `size` over
|
||||
all dupe rows) in human units.
|
||||
@@ -543,12 +611,25 @@ Additional requirements:
|
||||
### Error handling and exit codes
|
||||
|
||||
- `0`: success, even if individual files were skipped with warnings.
|
||||
- `1`: fatal error (e.g., a `PATH` operand does not exist, the
|
||||
database cannot be created/opened/read/written, a missing database
|
||||
for `report`/`trees`, stdout write failure).
|
||||
- `1`: fatal error (e.g., a `PATH` operand does not exist, another
|
||||
`scan` is already running against the same database, the database
|
||||
cannot be created/opened/read/written, a missing database for
|
||||
`report`/`trees`, stdout write failure).
|
||||
- `2`: usage error (including `scan` with no `PATH` operand and
|
||||
`report`/`trees` with any positional argument).
|
||||
|
||||
A stdout write failure, such as a full disk, is reported in one line on
|
||||
stderr and exits 1. Two cases never reach sfdupes as a failed write:
|
||||
|
||||
- When the reader of a stdout pipe exits early, as in
|
||||
`sfdupes report | head`, the next write ends sfdupes with `SIGPIPE`,
|
||||
quietly and without a summary, the way `cat` or `sort` end. The
|
||||
shell reports the signal (status 141 in most shells), not exit 1.
|
||||
- When stdout is closed outright (`sfdupes report >&-`), the Go
|
||||
runtime opens `/dev/null` in its place before sfdupes starts, so
|
||||
the output is discarded and the run succeeds, as with
|
||||
`> /dev/null`.
|
||||
|
||||
## Entrypoints
|
||||
|
||||
This repository adheres to the
|
||||
@@ -686,7 +767,7 @@ All of the following, run in this directory, must pass:
|
||||
|
||||
```sh
|
||||
d=$(mktemp -d)
|
||||
export SFDUPES_DATABASE="$d/db.sqlite"
|
||||
export SFDUPES_DATABASE="$(mktemp -d)/db.sqlite"
|
||||
mkdir -p "$d/a" "$d/b"
|
||||
head -c 2000 /dev/urandom > "$d/a/one.bin"
|
||||
cp "$d/a/one.bin" "$d/b/copy.bin"
|
||||
@@ -714,9 +795,9 @@ All of the following, run in this directory, must pass:
|
||||
./sfdupes report
|
||||
```
|
||||
|
||||
(The scan database lives inside `$d` here purely for test hygiene;
|
||||
scanning `$d` therefore also records the SQLite file itself, which
|
||||
is harmless.)
|
||||
(The database lives in a temp directory of its own: inside `$d`,
|
||||
the scan would record it, and its empty lock file would join the
|
||||
`empty1`/`empty2` group.)
|
||||
|
||||
Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin`
|
||||
form one group (two dupe rows, `first` is the lexicographically
|
||||
|
||||
@@ -29,6 +29,41 @@
|
||||
|
||||
# Completed Steps
|
||||
|
||||
- `report` and `trees` stream the records instead of holding them all in
|
||||
memory; the schema gains the `files_signature` index (2026-10-04,
|
||||
https://git.eeqj.de/sneak/sfdupes/issues/14)
|
||||
|
||||
- warn about and skip symlink, socket, FIFO, device and `.zfs`
|
||||
operands, keeping the records beneath them (2026-10-03,
|
||||
https://git.eeqj.de/sneak/sfdupes/issues/9)
|
||||
|
||||
- `scan` holds a lock on a lock file beside the database for its whole run,
|
||||
so a second `scan` fails at once with exit 1 (2026-10-03,
|
||||
https://git.eeqj.de/sneak/sfdupes/issues/53)
|
||||
|
||||
- test stdout write failures in `report` and `trees`; README states that
|
||||
`| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null`
|
||||
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/30)
|
||||
|
||||
- `report` and `trees` open the database read-only, and `scan` leaves it
|
||||
out of WAL mode, so reading needs only read access (2026-10-03, closes
|
||||
https://git.eeqj.de/sneak/sfdupes/issues/8)
|
||||
|
||||
- escape tabs, newlines, carriage returns and backslashes in report,
|
||||
trees and warning paths; the root directory's path is `/`
|
||||
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/7)
|
||||
|
||||
- stamp the git tag or short commit in a plain `docker build .`
|
||||
instead of `dev` (2026-10-02, branch `next`, closes
|
||||
https://git.eeqj.de/sneak/sfdupes/issues/67): `.dockerignore` now
|
||||
sends `.git`, without `.git/config`, and the `Dockerfile` build
|
||||
stage takes the `VERSION` build argument when one is given,
|
||||
otherwise `git describe --tags --always` of that `.git`. The build
|
||||
fails if the context carries `.git` and the version still comes out
|
||||
empty, `dev` or `unknown`. The CI checkout step fetches the full
|
||||
history (`fetch-depth: 0`) so CI sees the tag and stamps the same
|
||||
value as `make build`.
|
||||
|
||||
- replace the 1 KiB end-window sampling with the head/tail plus
|
||||
content-hash ladder (2026-09-22, branch `next`, closes
|
||||
https://git.eeqj.de/sneak/sfdupes/issues/61): a file under 10 MiB is
|
||||
|
||||
@@ -11,6 +11,7 @@ import (
|
||||
"slices"
|
||||
"strconv"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
// The pure-Go SQLite driver, registered as "sqlite"; keeps cgo
|
||||
// disabled.
|
||||
_ "modernc.org/sqlite"
|
||||
@@ -32,6 +33,11 @@ const schemaVersion = 1
|
||||
// scan.
|
||||
const dbDirPerm = 0o755
|
||||
|
||||
// lockFilePerm is the mode for the scan lock file. Anyone who can open
|
||||
// the file can hold the lock and keep every scan from running, so it
|
||||
// is open to its owner only.
|
||||
const lockFilePerm = 0o600
|
||||
|
||||
// createTableSQL is the schema applied to a fresh database. Paths are
|
||||
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
|
||||
const createTableSQL = `
|
||||
@@ -45,6 +51,12 @@ CREATE TABLE files (
|
||||
) WITHOUT ROWID
|
||||
`
|
||||
|
||||
// createIndexSQL indexes the records by signature, so report can have
|
||||
// SQLite group them without sorting the whole table.
|
||||
const createIndexSQL = `
|
||||
CREATE INDEX files_signature ON files (size, head, tail, content)
|
||||
`
|
||||
|
||||
// upsertSQL inserts one file record, replacing any existing record for
|
||||
// the same path.
|
||||
const upsertSQL = `
|
||||
@@ -64,6 +76,10 @@ var errNoDatabase = errors.New(
|
||||
// does not understand.
|
||||
var errSchemaVersion = errors.New("unsupported database schema version")
|
||||
|
||||
// errScanRunning reports that another scan holds the lock on the
|
||||
// database.
|
||||
var errScanRunning = errors.New("another scan is running")
|
||||
|
||||
// databasePath resolves the database location: SFDUPES_DATABASE when
|
||||
// set and non-empty, the compiled-in default otherwise.
|
||||
func databasePath() string {
|
||||
@@ -74,16 +90,24 @@ func databasePath() string {
|
||||
return defaultDatabasePath
|
||||
}
|
||||
|
||||
// openDB opens the SQLite database at path with WAL journaling and a
|
||||
// busy timeout, so a report can run while a cron scan is in progress.
|
||||
// It does not create or verify the schema.
|
||||
func openDB(path string) (*sql.DB, error) {
|
||||
dsn := "file:" + path +
|
||||
"?_pragma=busy_timeout(10000)" +
|
||||
"&_pragma=journal_mode(WAL)" +
|
||||
"&_pragma=synchronous(NORMAL)"
|
||||
// scanParams are the connection parameters for scan: read-write, with
|
||||
// WAL journaling and a busy timeout, so a report can run while a cron
|
||||
// scan is in progress. closeScanDatabase leaves WAL mode again.
|
||||
const scanParams = "_pragma=busy_timeout(10000)" +
|
||||
"&_pragma=journal_mode(WAL)" +
|
||||
"&_pragma=synchronous(NORMAL)"
|
||||
|
||||
db, err := sql.Open("sqlite", dsn)
|
||||
// reportParams are the connection parameters for report and trees:
|
||||
// read-only, with the same busy timeout. They set no journal mode,
|
||||
// because setting one is a write.
|
||||
const reportParams = "mode=ro" +
|
||||
"&_pragma=busy_timeout(10000)" +
|
||||
"&_pragma=query_only(1)"
|
||||
|
||||
// openDB opens the SQLite database at path with the connection
|
||||
// parameters params. It does not create or verify the schema.
|
||||
func openDB(path, params string) (*sql.DB, error) {
|
||||
db, err := sql.Open("sqlite", "file:"+path+"?"+params)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("open database %s: %w", path, err)
|
||||
}
|
||||
@@ -96,6 +120,43 @@ func openDB(path string) (*sql.DB, error) {
|
||||
return db, nil
|
||||
}
|
||||
|
||||
// lockScanDatabase takes the lock that keeps a second scan off the
|
||||
// database at path: an exclusive flock(2) on the file beside it named
|
||||
// path with ".lock" appended, created along with the database's parent
|
||||
// directory if missing. A lock held by another scan fails at once
|
||||
// instead of waiting. The lock lasts until the returned file is closed
|
||||
// or the process ends. The file is never deleted: a scan that deleted
|
||||
// it would let the next scan lock a new file while another still holds
|
||||
// the old one.
|
||||
func lockScanDatabase(path string) (*os.File, error) {
|
||||
err := os.MkdirAll(filepath.Dir(path), dbDirPerm)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("create database directory: %w", err)
|
||||
}
|
||||
|
||||
lockPath := path + ".lock"
|
||||
|
||||
//nolint:gosec // the operator chooses the database path
|
||||
f, err := os.OpenFile(lockPath, os.O_RDWR|os.O_CREATE, lockFilePerm)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
err = unix.Flock(int(f.Fd()), unix.LOCK_EX|unix.LOCK_NB)
|
||||
if err != nil {
|
||||
_ = f.Close()
|
||||
|
||||
if errors.Is(err, unix.EWOULDBLOCK) {
|
||||
return nil, fmt.Errorf("%w (lock held on %s)",
|
||||
errScanRunning, lockPath)
|
||||
}
|
||||
|
||||
return nil, fmt.Errorf("lock %s: %w", lockPath, err)
|
||||
}
|
||||
|
||||
return f, nil
|
||||
}
|
||||
|
||||
// openScanDatabase opens the database for the scan subcommand, creating
|
||||
// the file, its parent directory, and the schema as needed.
|
||||
func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
|
||||
@@ -104,7 +165,7 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
|
||||
return nil, fmt.Errorf("create database directory: %w", err)
|
||||
}
|
||||
|
||||
db, err := openDB(path)
|
||||
db, err := openDB(path, scanParams)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
@@ -119,6 +180,24 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
|
||||
return db, nil
|
||||
}
|
||||
|
||||
// closeScanDatabase switches the database at path from WAL back to
|
||||
// rollback-journal mode and closes it. Out of WAL mode the database
|
||||
// file alone holds the whole database, so a reader needs no -wal or
|
||||
// -shm file beside it, nor write access to create them. The switch
|
||||
// fails while a report has the database open; the database then stays
|
||||
// in WAL mode, still readable, until a later scan closes it.
|
||||
func closeScanDatabase(ctx context.Context, db *sql.DB, path string) {
|
||||
// Runs on the way out of a cancelled scan too.
|
||||
_, err := db.ExecContext(context.WithoutCancel(ctx),
|
||||
"PRAGMA journal_mode = DELETE")
|
||||
if err != nil {
|
||||
fmt.Fprintf(os.Stderr, "scan: database %s left in WAL mode: %v\n",
|
||||
path, err)
|
||||
}
|
||||
|
||||
_ = db.Close()
|
||||
}
|
||||
|
||||
// openReportDatabase opens an existing database for the report and
|
||||
// trees subcommands. A missing database file is an error directing the
|
||||
// user to run scan first; the schema version must match exactly.
|
||||
@@ -134,7 +213,7 @@ func openReportDatabase(ctx context.Context,
|
||||
return nil, fmt.Errorf("database: %w", err)
|
||||
}
|
||||
|
||||
db, err := openDB(path)
|
||||
db, err := openDB(path, reportParams)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
@@ -183,6 +262,11 @@ func createSchema(ctx context.Context, db *sql.DB) error {
|
||||
return fmt.Errorf("create schema: %w", err)
|
||||
}
|
||||
|
||||
_, err = db.ExecContext(ctx, createIndexSQL)
|
||||
if err != nil {
|
||||
return fmt.Errorf("create schema: %w", err)
|
||||
}
|
||||
|
||||
_, err = db.ExecContext(ctx,
|
||||
"PRAGMA user_version = "+strconv.Itoa(schemaVersion))
|
||||
if err != nil {
|
||||
@@ -204,18 +288,18 @@ func userVersion(ctx context.Context, db *sql.DB) (int, error) {
|
||||
return v, nil
|
||||
}
|
||||
|
||||
// loadFileRows reads every record from the files table.
|
||||
func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) {
|
||||
// loadFileRows streams every record to fn in path order: byte order,
|
||||
// which is the order of the primary key, so SQLite does not sort.
|
||||
func loadFileRows(ctx context.Context, db *sql.DB, fn func(r scanRec)) error {
|
||||
rows, err := db.QueryContext(ctx,
|
||||
"SELECT path, size, mtime, head, tail, content FROM files")
|
||||
"SELECT path, size, mtime, head, tail, content FROM files "+
|
||||
"ORDER BY path")
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read records: %w", err)
|
||||
return fmt.Errorf("read records: %w", err)
|
||||
}
|
||||
|
||||
defer func() { _ = rows.Close() }()
|
||||
|
||||
var recs []scanRec
|
||||
|
||||
for rows.Next() {
|
||||
var (
|
||||
path []byte
|
||||
@@ -225,19 +309,92 @@ func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) {
|
||||
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail,
|
||||
&r.content)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read record: %w", err)
|
||||
return fmt.Errorf("read record: %w", err)
|
||||
}
|
||||
|
||||
r.path = string(path)
|
||||
recs = append(recs, r)
|
||||
fn(r)
|
||||
}
|
||||
|
||||
err = rows.Err()
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("read records: %w", err)
|
||||
return fmt.Errorf("read records: %w", err)
|
||||
}
|
||||
|
||||
return recs, nil
|
||||
return nil
|
||||
}
|
||||
|
||||
// dupeRowsSQL selects every record in a duplicate group, with the
|
||||
// group's first path. A group is the records with a content hash that
|
||||
// share a size, head, tail, and content, when there are two or more of
|
||||
// them. The rows come in report order: groups by size descending, then
|
||||
// by first path, and each group's paths ascending.
|
||||
const dupeRowsSQL = `
|
||||
SELECT g.first, f.path, f.size
|
||||
FROM files AS f
|
||||
JOIN (
|
||||
SELECT size, head, tail, content, MIN(path) AS first
|
||||
FROM files
|
||||
WHERE content <> ''
|
||||
GROUP BY size, head, tail, content
|
||||
HAVING COUNT(*) > 1
|
||||
) AS g USING (size, head, tail, content)
|
||||
ORDER BY f.size DESC, g.first, f.path
|
||||
`
|
||||
|
||||
// loadDupeRows streams the rows of dupeRowsSQL to fn and returns the
|
||||
// number of records in the database. The count and the rows are read
|
||||
// in one transaction, so they agree while a scan is committing. An
|
||||
// error from fn stops the reading and is returned as it is.
|
||||
func loadDupeRows(ctx context.Context, db *sql.DB,
|
||||
fn func(first, path string, size int64) error,
|
||||
) (int, error) {
|
||||
// Everything goes through tx: the report connection is the only
|
||||
// one, so a query on db would wait for tx forever.
|
||||
tx, err := db.BeginTx(ctx, &sql.TxOptions{ReadOnly: true})
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("read records: %w", err)
|
||||
}
|
||||
|
||||
defer func() { _ = tx.Rollback() }()
|
||||
|
||||
var records int
|
||||
|
||||
err = tx.QueryRowContext(ctx, "SELECT COUNT(*) FROM files").Scan(&records)
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("read records: %w", err)
|
||||
}
|
||||
|
||||
rows, err := tx.QueryContext(ctx, dupeRowsSQL)
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("read records: %w", err)
|
||||
}
|
||||
|
||||
defer func() { _ = rows.Close() }()
|
||||
|
||||
for rows.Next() {
|
||||
var (
|
||||
first, path []byte
|
||||
size int64
|
||||
)
|
||||
|
||||
err = rows.Scan(&first, &path, &size)
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("read record: %w", err)
|
||||
}
|
||||
|
||||
err = fn(string(first), string(path), size)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
}
|
||||
|
||||
err = rows.Err()
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("read records: %w", err)
|
||||
}
|
||||
|
||||
return records, nil
|
||||
}
|
||||
|
||||
// loadFileMeta streams every record's path, size, mtime, and whether
|
||||
|
||||
+53
-24
@@ -5,9 +5,9 @@ import (
|
||||
"database/sql"
|
||||
"errors"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
@@ -72,9 +72,8 @@ func TestOpenScanDatabaseCreates(t *testing.T) {
|
||||
|
||||
defer func() { _ = db.Close() }()
|
||||
|
||||
recs, err := loadFileRows(t.Context(), db)
|
||||
if err != nil || len(recs) != 0 {
|
||||
t.Fatalf("loadFileRows = %v, %v; want empty, nil", recs, err)
|
||||
if recs := dbRecords(t, db); len(recs) != 0 {
|
||||
t.Fatalf("records = %v, want none", recs)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -130,6 +129,49 @@ func TestOpenReportDatabaseOK(t *testing.T) {
|
||||
_ = db.Close()
|
||||
}
|
||||
|
||||
func TestCloseScanDatabaseWhileReportOpen(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// A report holding the database open stops scan from taking it out
|
||||
// of WAL mode. The -wal and -shm files must then stay beside it, so
|
||||
// that a later report still needs only read access.
|
||||
path := testDBPath(t)
|
||||
|
||||
scanDB, err := openScanDatabase(t.Context(), path)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
reportDB, err := openReportDatabase(t.Context(), path)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
closeScanDatabase(t.Context(), scanDB, path)
|
||||
|
||||
_ = reportDB.Close()
|
||||
|
||||
_, err = os.Stat(path + "-wal")
|
||||
if err != nil {
|
||||
t.Fatalf("no -wal left: the switch out of WAL mode was not "+
|
||||
"stopped: %v", err)
|
||||
}
|
||||
|
||||
makeReadOnly(t, path)
|
||||
|
||||
reportDB, err = openReportDatabase(t.Context(), path)
|
||||
if err != nil {
|
||||
t.Fatalf("openReportDatabase: %v", err)
|
||||
}
|
||||
|
||||
defer func() { _ = reportDB.Close() }()
|
||||
|
||||
err = loadFileRows(t.Context(), reportDB, func(scanRec) {})
|
||||
if err != nil {
|
||||
t.Fatalf("loadFileRows: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestApplyChangesRoundTrip(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
@@ -152,15 +194,8 @@ func TestApplyChangesRoundTrip(t *testing.T) {
|
||||
t.Fatalf("applyChanges: %v", err)
|
||||
}
|
||||
|
||||
got, err := loadFileRows(t.Context(), db)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
slices.SortFunc(got, func(a, b scanRec) int {
|
||||
return strings.Compare(a.path, b.path)
|
||||
})
|
||||
|
||||
// The records come back in path order, which is the order of recs.
|
||||
got := dbRecords(t, db)
|
||||
if !slices.Equal(got, recs) {
|
||||
t.Fatalf("rows = %+v, want %+v", got, recs)
|
||||
}
|
||||
@@ -177,11 +212,7 @@ func TestApplyChangesRoundTrip(t *testing.T) {
|
||||
t.Fatalf("applyChanges: %v", err)
|
||||
}
|
||||
|
||||
got, err = loadFileRows(t.Context(), db)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
got = dbRecords(t, db)
|
||||
if len(got) != 1 || got[0] != upd {
|
||||
t.Fatalf("rows = %+v, want just %+v", got, upd)
|
||||
}
|
||||
@@ -210,9 +241,8 @@ func TestApplyChangesBatching(t *testing.T) {
|
||||
t.Fatalf("applyChanges: %v", err)
|
||||
}
|
||||
|
||||
got, err := loadFileRows(t.Context(), db)
|
||||
if err != nil || len(got) != n {
|
||||
t.Fatalf("loadFileRows = %d rows, %v; want %d", len(got), err, n)
|
||||
if got := dbRecords(t, db); len(got) != n {
|
||||
t.Fatalf("records = %d, want %d", len(got), n)
|
||||
}
|
||||
|
||||
deletes := make([]string, 0, n)
|
||||
@@ -226,8 +256,7 @@ func TestApplyChangesBatching(t *testing.T) {
|
||||
t.Fatalf("applyChanges deletes: %v", err)
|
||||
}
|
||||
|
||||
got, err = loadFileRows(t.Context(), db)
|
||||
if err != nil || len(got) != 0 {
|
||||
t.Fatalf("loadFileRows = %d rows, %v; want 0", len(got), err)
|
||||
if got := dbRecords(t, db); len(got) != 0 {
|
||||
t.Fatalf("records = %d, want 0", len(got))
|
||||
}
|
||||
}
|
||||
|
||||
@@ -5,6 +5,7 @@ go 1.25.7
|
||||
require (
|
||||
github.com/schollz/progressbar/v3 v3.19.1
|
||||
github.com/spf13/cobra v1.10.2
|
||||
golang.org/x/sys v0.46.0
|
||||
modernc.org/sqlite v1.54.0
|
||||
)
|
||||
|
||||
@@ -18,7 +19,6 @@ require (
|
||||
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
|
||||
github.com/rivo/uniseg v0.4.7 // indirect
|
||||
github.com/spf13/pflag v1.0.9 // indirect
|
||||
golang.org/x/sys v0.46.0 // indirect
|
||||
golang.org/x/term v0.44.0 // indirect
|
||||
modernc.org/libc v1.74.1 // indirect
|
||||
modernc.org/mathutil v1.7.1 // indirect
|
||||
|
||||
@@ -57,22 +57,27 @@ var errNoSubcommand = errors.New("no subcommand")
|
||||
var Version = "dev"
|
||||
|
||||
func main() {
|
||||
os.Exit(run(os.Args[1:], os.Stderr))
|
||||
// Once the reader of a stdout pipe has gone, as in "sfdupes report |
|
||||
// head", the Go runtime ends the process with SIGPIPE on the next
|
||||
// write instead of returning an error (README "Error handling").
|
||||
// Registering for SIGPIPE with os/signal would change that.
|
||||
os.Exit(run(os.Args[1:], os.Stdout, os.Stderr))
|
||||
}
|
||||
|
||||
// run executes args against the command tree and returns the process
|
||||
// exit code. It is the program's single exit point: the subcommands
|
||||
// return their errors instead of exiting, so every deferred cleanup —
|
||||
// above all closing the database, which checkpoints the SQLite WAL —
|
||||
// runs before the process ends.
|
||||
func run(args []string, stderr io.Writer) int {
|
||||
// runs before the process ends. The report and trees subcommands write
|
||||
// their data to stdout.
|
||||
func run(args []string, stdout, stderr io.Writer) int {
|
||||
// A nil slice makes cobra fall back to os.Args, which would let a
|
||||
// test binary's own flags reach the command tree.
|
||||
if args == nil {
|
||||
args = []string{}
|
||||
}
|
||||
|
||||
root := newRootCommand(stderr)
|
||||
root := newRootCommand(stdout, stderr)
|
||||
root.SetArgs(args)
|
||||
|
||||
err := root.Execute()
|
||||
@@ -98,7 +103,7 @@ func run(args []string, stderr io.Writer) int {
|
||||
// newRootCommand builds the command tree. Everything on stdout is
|
||||
// machine-readable data; all human-facing output (help, usage, errors)
|
||||
// goes to stderr.
|
||||
func newRootCommand(stderr io.Writer) *cobra.Command {
|
||||
func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
|
||||
root := &cobra.Command{
|
||||
Use: "sfdupes",
|
||||
Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
|
||||
@@ -140,7 +145,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
|
||||
Short: "Read the scan database and print the file-level duplicates report",
|
||||
Args: cobra.NoArgs,
|
||||
RunE: runE(func(ctx context.Context, _ []string) error {
|
||||
return runReport(ctx)
|
||||
return runReport(ctx, stdout)
|
||||
}),
|
||||
}
|
||||
|
||||
@@ -149,7 +154,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
|
||||
Short: "Read the scan database and print the duplicate-tree report",
|
||||
Args: cobra.NoArgs,
|
||||
RunE: runE(func(ctx context.Context, _ []string) error {
|
||||
return runTrees(ctx)
|
||||
return runTrees(ctx, stdout)
|
||||
}),
|
||||
}
|
||||
|
||||
|
||||
+376
-55
@@ -44,23 +44,61 @@ func assertNoSidecars(t *testing.T, path string) {
|
||||
}
|
||||
}
|
||||
|
||||
// captureStdout redirects os.Stdout to a file for the rest of the test
|
||||
// and returns a function reading back everything written to it. Only
|
||||
// machine-readable data belongs on stdout (README design goal 4), so
|
||||
// the tests assert on it directly.
|
||||
func captureStdout(t *testing.T) func() string {
|
||||
// makeReadOnly takes write permission away from the database at path,
|
||||
// from any WAL sidecar beside it, and from their directory, as for a
|
||||
// user reading a database that a root cron scan keeps. Root ignores
|
||||
// file permissions, so it skips the test when run as root.
|
||||
func makeReadOnly(t *testing.T, path string) {
|
||||
t.Helper()
|
||||
|
||||
f, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
|
||||
if os.Geteuid() == 0 {
|
||||
t.Skip("root ignores file permissions")
|
||||
}
|
||||
|
||||
err := os.Chmod(path, 0o400)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
saved := os.Stdout
|
||||
os.Stdout = f
|
||||
for _, suffix := range walSuffixes {
|
||||
err = os.Chmod(path+suffix, 0o400)
|
||||
if err != nil && !errors.Is(err, fs.ErrNotExist) {
|
||||
t.Fatal(err)
|
||||
}
|
||||
}
|
||||
|
||||
dir := filepath.Dir(path)
|
||||
|
||||
//nolint:gosec // reaching the database needs the search bit
|
||||
err = os.Chmod(dir, 0o500)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
// Runs before t.TempDir's own cleanup, which must delete the files.
|
||||
t.Cleanup(func() {
|
||||
//nolint:gosec // removing the directory needs its search bit back
|
||||
_ = os.Chmod(dir, 0o700)
|
||||
})
|
||||
}
|
||||
|
||||
// captureStderr redirects os.Stderr to a file for the rest of the test
|
||||
// and returns a function reading back everything written to it. scan
|
||||
// writes its warnings and summary straight to os.Stderr, not to the
|
||||
// stderr writer run is given.
|
||||
func captureStderr(t *testing.T) func() string {
|
||||
t.Helper()
|
||||
|
||||
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
saved := os.Stderr
|
||||
os.Stderr = f
|
||||
|
||||
t.Cleanup(func() {
|
||||
os.Stdout = saved
|
||||
os.Stderr = saved
|
||||
|
||||
_ = f.Close()
|
||||
})
|
||||
@@ -91,13 +129,13 @@ func captureStdout(t *testing.T) func() string {
|
||||
// brokenDatabase writes a database that opens cleanly and passes the
|
||||
// schema-version check but has no files table, so the first query
|
||||
// fails with the database already open: a fatal error on a path that
|
||||
// owns an open database.
|
||||
// owns an open database. It closes the database the way scan does.
|
||||
func brokenDatabase(t *testing.T) string {
|
||||
t.Helper()
|
||||
|
||||
path := testDBPath(t)
|
||||
|
||||
db, err := openDB(path)
|
||||
db, err := openDB(path, scanParams)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
@@ -108,10 +146,7 @@ func brokenDatabase(t *testing.T) string {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
err = db.Close()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
closeScanDatabase(t.Context(), db, path)
|
||||
|
||||
return path
|
||||
}
|
||||
@@ -144,7 +179,10 @@ func TestOpenDatabaseKeepsWALWhileOpen(t *testing.T) {
|
||||
|
||||
func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
|
||||
// Every subcommand that owns an open database must close it when
|
||||
// it fails: no os.Exit between the open and the return.
|
||||
// it fails: no os.Exit between the open and the return. The
|
||||
// sidecar check is evidence of the close only for scan: report and
|
||||
// trees only read a database that is out of WAL mode, which leaves
|
||||
// nothing on disk whether they close it or not.
|
||||
cases := map[string][]string{
|
||||
cmdScan: {cmdScan},
|
||||
cmdReport: {cmdReport},
|
||||
@@ -160,17 +198,15 @@ func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
|
||||
args = append(args, t.TempDir())
|
||||
}
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run(args, &stderr)
|
||||
code := run(args, &stdout, &stderr)
|
||||
if code != exitFatal {
|
||||
t.Errorf("run(%v) = %d, want %d", args, code, exitFatal)
|
||||
}
|
||||
|
||||
assertNoSidecars(t, path)
|
||||
assertFatalOutput(t, stderr.String(), stdout())
|
||||
assertFatalOutput(t, stderr.String(), stdout.String())
|
||||
|
||||
// Proof that the failure happened after the open: only a
|
||||
// query against the opened database can report this.
|
||||
@@ -188,18 +224,16 @@ func TestRunMissingOperandIsFatalNotUsage(t *testing.T) {
|
||||
// must not dump the usage text.
|
||||
t.Setenv(databaseEnv, testDBPath(t))
|
||||
|
||||
var stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
missing := filepath.Join(t.TempDir(), "nope")
|
||||
|
||||
code := run([]string{cmdScan, missing}, &stderr)
|
||||
code := run([]string{cmdScan, missing}, &stdout, &stderr)
|
||||
if code != exitFatal {
|
||||
t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal)
|
||||
}
|
||||
|
||||
assertFatalOutput(t, stderr.String(), stdout())
|
||||
assertFatalOutput(t, stderr.String(), stdout.String())
|
||||
}
|
||||
|
||||
// assertFatalOutput checks that a fatal error was reported the way
|
||||
@@ -243,11 +277,9 @@ func TestRunUsageErrors(t *testing.T) {
|
||||
// path that does not exist.
|
||||
t.Setenv(databaseEnv, testDBPath(t))
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run(tc.args, &stderr)
|
||||
code := run(tc.args, &stdout, &stderr)
|
||||
if code != exitUsage {
|
||||
t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage)
|
||||
}
|
||||
@@ -256,7 +288,7 @@ func TestRunUsageErrors(t *testing.T) {
|
||||
t.Errorf("stderr = %q, want %q", stderr.String(), tc.want)
|
||||
}
|
||||
|
||||
if got := stdout(); got != "" {
|
||||
if got := stdout.String(); got != "" {
|
||||
t.Errorf("stdout = %q, want nothing (data only)", got)
|
||||
}
|
||||
})
|
||||
@@ -265,9 +297,9 @@ func TestRunUsageErrors(t *testing.T) {
|
||||
|
||||
// TestRunHelpAndVersionSucceed checks that the two informational flags
|
||||
// exit 0 and keep their human-facing output on stderr.
|
||||
//
|
||||
//nolint:paralleltest // captureStdout replaces the process-wide os.Stdout
|
||||
func TestRunHelpAndVersionSucceed(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
assertHumanOutput(t, "--help")
|
||||
assertHumanOutput(t, "--version")
|
||||
}
|
||||
@@ -278,11 +310,9 @@ func TestRunHelpAndVersionSucceed(t *testing.T) {
|
||||
func assertHumanOutput(t *testing.T, arg string) {
|
||||
t.Helper()
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run([]string{arg}, &stderr)
|
||||
code := run([]string{arg}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
|
||||
}
|
||||
@@ -291,7 +321,7 @@ func assertHumanOutput(t *testing.T, arg string) {
|
||||
t.Errorf("run(%s) wrote nothing to stderr", arg)
|
||||
}
|
||||
|
||||
if got := stdout(); got != "" {
|
||||
if got := stdout.String(); got != "" {
|
||||
t.Errorf("stdout = %q, want nothing (data only)", got)
|
||||
}
|
||||
}
|
||||
@@ -320,21 +350,31 @@ func scanFixture(t *testing.T) []string {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
var stderr bytes.Buffer
|
||||
scanOK(t, dir)
|
||||
|
||||
stdout := captureStdout(t)
|
||||
return dupes
|
||||
}
|
||||
|
||||
code := run([]string{cmdScan, dir}, &stderr)
|
||||
// scanOK runs scan over operands, fails the test unless it exits 0 with
|
||||
// nothing on stdout, and returns everything it printed to stderr.
|
||||
func scanOK(t *testing.T, operands ...string) string {
|
||||
t.Helper()
|
||||
|
||||
var stdout bytes.Buffer
|
||||
|
||||
stderr := captureStderr(t)
|
||||
|
||||
code := run(append([]string{cmdScan}, operands...), &stdout, os.Stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
|
||||
code, exitOK, stderr.String())
|
||||
t.Fatalf("run(scan %q) = %d, want %d; stderr: %s",
|
||||
operands, code, exitOK, stderr())
|
||||
}
|
||||
|
||||
if got := stdout(); got != "" {
|
||||
if got := stdout.String(); got != "" {
|
||||
t.Errorf("scan stdout = %q, want nothing (data only)", got)
|
||||
}
|
||||
|
||||
return dupes
|
||||
return stderr()
|
||||
}
|
||||
|
||||
func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
|
||||
@@ -345,24 +385,119 @@ func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
|
||||
assertNoSidecars(t, path)
|
||||
}
|
||||
|
||||
func TestRunScanSkipsSymlinkOperand(t *testing.T) {
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
dir := t.TempDir()
|
||||
writeFile(t, dir, "target/sub/f", pattern(1, 10))
|
||||
|
||||
link := filepath.Join(dir, "link")
|
||||
|
||||
err := os.Symlink(filepath.Join(dir, "target"), link)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
// Scanning a directory through the symlink stores a record beneath
|
||||
// the symlink's own path for a file beneath its target.
|
||||
scanOK(t, filepath.Join(link, "sub"))
|
||||
|
||||
assertOperandSkipped(t, path, link, "symlink",
|
||||
filepath.Join(link, "sub", "f"))
|
||||
}
|
||||
|
||||
func TestRunScanWalksOperandUnderSymlinkOperand(t *testing.T) {
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
dir := t.TempDir()
|
||||
writeFile(t, dir, "target/sub/f", pattern(1, 10))
|
||||
|
||||
link := filepath.Join(dir, "link")
|
||||
|
||||
err := os.Symlink(filepath.Join(dir, "target"), link)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
// link is dropped as a symlink, but link/sub must still be scanned,
|
||||
// not dropped as lying under link.
|
||||
scanOK(t, link, filepath.Join(link, "sub"))
|
||||
|
||||
db, err := openDB(path, reportParams)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
t.Cleanup(func() { _ = db.Close() })
|
||||
|
||||
recordByPath(t, dbRecords(t, db), filepath.Join(link, "sub", "f"))
|
||||
}
|
||||
|
||||
func TestRunScanSkipsZFSOperand(t *testing.T) {
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
zfs := filepath.Join(t.TempDir(), ".zfs")
|
||||
snapshot := filepath.Join(zfs, "snapshot", "hourly")
|
||||
f := writeFile(t, snapshot, "f", pattern(1, 10))
|
||||
|
||||
// An operand beneath a .zfs directory is walked, because it is not
|
||||
// itself named .zfs.
|
||||
scanOK(t, snapshot)
|
||||
|
||||
assertOperandSkipped(t, path, zfs, ".zfs directory", f)
|
||||
}
|
||||
|
||||
// assertOperandSkipped scans operand alone and checks that it is skipped
|
||||
// as kind: a warning naming it, one skip in the summary, exit 0, and the
|
||||
// record for kept, which an earlier scan stored beneath operand, still
|
||||
// in the database at dbPath.
|
||||
func assertOperandSkipped(t *testing.T, dbPath, operand, kind,
|
||||
kept string,
|
||||
) {
|
||||
t.Helper()
|
||||
|
||||
stderr := scanOK(t, operand)
|
||||
|
||||
warning := "walk " + operand + ": skipping " + kind + " operand\n"
|
||||
if !strings.Contains(stderr, warning) {
|
||||
t.Errorf("stderr = %q, want %q", stderr, warning)
|
||||
}
|
||||
|
||||
summary := "scan: 0 files seen (0 added, 0 updated, 0 removed, " +
|
||||
"0 unchanged), 1 skipped\n"
|
||||
if !strings.Contains(stderr, summary) {
|
||||
t.Errorf("stderr = %q, want %q", stderr, summary)
|
||||
}
|
||||
|
||||
db, err := openDB(dbPath, reportParams)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
t.Cleanup(func() { _ = db.Close() })
|
||||
|
||||
recordByPath(t, dbRecords(t, db), kept)
|
||||
}
|
||||
|
||||
func TestRunReportSucceeds(t *testing.T) {
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
dupes := scanFixture(t)
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run([]string{cmdReport}, &stderr)
|
||||
code := run([]string{cmdReport}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(report) = %d, want %d; stderr: %s",
|
||||
code, exitOK, stderr.String())
|
||||
}
|
||||
|
||||
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
|
||||
if got := stdout(); got != want {
|
||||
if got := stdout.String(); got != want {
|
||||
t.Errorf("stdout = %q, want %q", got, want)
|
||||
}
|
||||
|
||||
@@ -375,11 +510,9 @@ func TestRunTreesSucceeds(t *testing.T) {
|
||||
|
||||
dupes := scanFixture(t)
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run([]string{cmdTrees}, &stderr)
|
||||
code := run([]string{cmdTrees}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
|
||||
code, exitOK, stderr.String())
|
||||
@@ -389,9 +522,197 @@ func TestRunTreesSucceeds(t *testing.T) {
|
||||
// trees of each other.
|
||||
want := "first\tdupe\tfiles\tsize\n" +
|
||||
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
|
||||
if got := stdout(); got != want {
|
||||
if got := stdout.String(); got != want {
|
||||
t.Errorf("stdout = %q, want %q", got, want)
|
||||
}
|
||||
|
||||
assertNoSidecars(t, path)
|
||||
}
|
||||
|
||||
func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
|
||||
// README §Database: report and trees need only read access to the
|
||||
// database file. With its directory read-only as well, SQLite
|
||||
// cannot create any file beside it.
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
dupes := scanFixture(t)
|
||||
assertNoSidecars(t, path)
|
||||
makeReadOnly(t, path)
|
||||
|
||||
cases := map[string]string{
|
||||
cmdReport: "first\tdupe\tsize\n" +
|
||||
dupes[0] + "\t" + dupes[1] + "\t300\n",
|
||||
cmdTrees: "first\tdupe\tfiles\tsize\n" +
|
||||
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) +
|
||||
"\t1\t300\n",
|
||||
}
|
||||
|
||||
for name, want := range cases {
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := run([]string{name}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Errorf("run(%s) = %d, want %d; stderr: %s",
|
||||
name, code, exitOK, stderr.String())
|
||||
|
||||
continue
|
||||
}
|
||||
|
||||
if got := stdout.String(); got != want {
|
||||
t.Errorf("%s stdout = %q, want %q", name, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// holdScanLock takes the lock on the database at path, as a running
|
||||
// scan does, and holds it until the test ends. It fails the test when
|
||||
// the lock is already held.
|
||||
func holdScanLock(t *testing.T, path string) {
|
||||
t.Helper()
|
||||
|
||||
lock, err := lockScanDatabase(path)
|
||||
if err != nil {
|
||||
t.Fatalf("lock %s: %v", path, err)
|
||||
}
|
||||
|
||||
t.Cleanup(func() { _ = lock.Close() })
|
||||
}
|
||||
|
||||
func TestRunSecondScanFails(t *testing.T) {
|
||||
// README §Database: while one scan holds the lock, a second scan
|
||||
// fails at once, naming the lock file, without creating the
|
||||
// database.
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
holdScanLock(t, path)
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := run([]string{cmdScan, t.TempDir()}, &stdout, &stderr)
|
||||
if code != exitFatal {
|
||||
t.Errorf("run(scan) = %d, want %d", code, exitFatal)
|
||||
}
|
||||
|
||||
want := "sfdupes: another scan is running (lock held on " +
|
||||
path + ".lock)\n"
|
||||
if got := stderr.String(); got != want {
|
||||
t.Errorf("stderr = %q, want %q", got, want)
|
||||
}
|
||||
|
||||
if got := stdout.String(); got != "" {
|
||||
t.Errorf("stdout = %q, want nothing (data only)", got)
|
||||
}
|
||||
|
||||
_, err := os.Stat(path)
|
||||
if !errors.Is(err, fs.ErrNotExist) {
|
||||
t.Errorf("stat %s = %v, want the database not created", path, err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunScanReleasesLock(t *testing.T) {
|
||||
// README §Database: a scan releases the lock however it ends.
|
||||
t.Run("success", func(t *testing.T) {
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
scanFixture(t)
|
||||
holdScanLock(t, path)
|
||||
})
|
||||
|
||||
t.Run("fatal error", func(t *testing.T) {
|
||||
path := brokenDatabase(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
code := run([]string{cmdScan, t.TempDir()}, io.Discard, io.Discard)
|
||||
if code != exitFatal {
|
||||
t.Fatalf("run(scan) = %d, want %d", code, exitFatal)
|
||||
}
|
||||
|
||||
holdScanLock(t, path)
|
||||
})
|
||||
}
|
||||
|
||||
func TestRunReportsDuringScan(t *testing.T) {
|
||||
// README §Database: report and trees never take the lock, so they
|
||||
// run while a scan holds it.
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
scanFixture(t)
|
||||
holdScanLock(t, path)
|
||||
|
||||
for _, name := range []string{cmdReport, cmdTrees} {
|
||||
var stderr bytes.Buffer
|
||||
|
||||
code := run([]string{name}, io.Discard, &stderr)
|
||||
if code != exitOK {
|
||||
t.Errorf("run(%s) = %d, want %d; stderr: %s",
|
||||
name, code, exitOK, stderr.String())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunStdoutClosedIsFatal(t *testing.T) {
|
||||
// README §Error handling: a stdout write failure exits 1, reported
|
||||
// in one line on stderr.
|
||||
for _, name := range []string{cmdReport, cmdTrees} {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
t.Setenv(databaseEnv, testDBPath(t))
|
||||
|
||||
scanFixture(t)
|
||||
|
||||
stdout, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
err = stdout.Close()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
var stderr bytes.Buffer
|
||||
|
||||
code := run([]string{name}, stdout, &stderr)
|
||||
if code != exitFatal {
|
||||
t.Errorf("run(%s) = %d, want %d", name, code, exitFatal)
|
||||
}
|
||||
|
||||
got := stderr.String()
|
||||
if !strings.HasPrefix(got, "sfdupes: write stdout: ") ||
|
||||
!strings.Contains(got, os.ErrClosed.Error()) ||
|
||||
strings.Count(got, "\n") != 1 {
|
||||
t.Errorf("stderr = %q, want one line reporting the "+
|
||||
"failed stdout write", got)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// errWriteFailed is the error failingWriter returns.
|
||||
var errWriteFailed = errors.New("write failed")
|
||||
|
||||
// failingWriter is a stdout that fails every write.
|
||||
type failingWriter struct{}
|
||||
|
||||
func (failingWriter) Write([]byte) (int, error) { return 0, errWriteFailed }
|
||||
|
||||
func TestStdoutWriteErrorPropagates(t *testing.T) {
|
||||
t.Setenv(databaseEnv, testDBPath(t))
|
||||
|
||||
scanFixture(t)
|
||||
|
||||
cases := map[string]func(context.Context, io.Writer) error{
|
||||
cmdReport: runReport,
|
||||
cmdTrees: runTrees,
|
||||
}
|
||||
|
||||
for name, fn := range cases {
|
||||
err := fn(t.Context(), failingWriter{})
|
||||
if !errors.Is(err, errWriteFailed) {
|
||||
t.Errorf("%s: error = %v, want %v", name, err, errWriteFailed)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
+3
-1
@@ -106,6 +106,8 @@ func (p *progress) increment() {
|
||||
}
|
||||
|
||||
// warnf prints a one-line warning to stderr without corrupting the bar.
|
||||
// The whole message is escaped like a report's path columns, so a path
|
||||
// holding a newline cannot split the warning.
|
||||
func (p *progress) warnf(format string, args ...any) {
|
||||
if p == nil {
|
||||
return
|
||||
@@ -115,7 +117,7 @@ func (p *progress) warnf(format string, args ...any) {
|
||||
_ = p.bar.Clear()
|
||||
}
|
||||
|
||||
fmt.Fprintf(os.Stderr, format+"\n", args...)
|
||||
fmt.Fprintln(os.Stderr, escapePath(fmt.Sprintf(format, args...)))
|
||||
}
|
||||
|
||||
// finish terminates the pass's display.
|
||||
|
||||
@@ -4,8 +4,8 @@ import (
|
||||
"bufio"
|
||||
"context"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"slices"
|
||||
"strings"
|
||||
)
|
||||
|
||||
@@ -28,73 +28,59 @@ type scanRec struct {
|
||||
path string
|
||||
}
|
||||
|
||||
// loadRecords opens the database and reads every file record for the
|
||||
// report and trees subcommands. Any database problem — including a
|
||||
// missing database — is fatal. The error is returned rather than
|
||||
// exiting, so that the deferred close — which checkpoints the SQLite
|
||||
// WAL — always runs; the database is closed before the caller formats
|
||||
// its output, so it stays closed even if that output fails.
|
||||
func loadRecords(ctx context.Context) ([]scanRec, error) {
|
||||
// runReport implements the report subcommand: it prints the file-level
|
||||
// duplicates report as TSV on stdout. SQLite groups and orders the
|
||||
// records, and each row is written as it is read, so no group is held
|
||||
// in memory. It never touches the scanned filesystem; its only I/O is
|
||||
// the database (with SQLite's temporary sort file), stdout, and stderr.
|
||||
// Any database problem, including a missing database, is fatal.
|
||||
func runReport(ctx context.Context, stdout io.Writer) error {
|
||||
dbPath := databasePath()
|
||||
|
||||
db, err := openReportDatabase(ctx, dbPath)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
return err
|
||||
}
|
||||
|
||||
defer func() { _ = db.Close() }()
|
||||
|
||||
recs, err := loadFileRows(ctx, db)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("database %s: %w", dbPath, err)
|
||||
}
|
||||
|
||||
return recs, nil
|
||||
}
|
||||
|
||||
// dupeGroup is one set of candidate-duplicate files: identical size,
|
||||
// head hash, tail hash, and content hash. paths is sorted
|
||||
// lexicographically; the first entry is the group's "first", the rest
|
||||
// are dupes.
|
||||
type dupeGroup struct {
|
||||
size int64
|
||||
paths []string
|
||||
}
|
||||
|
||||
// runReport implements the report subcommand: it reads every record
|
||||
// from the database and prints the file-level duplicates report as TSV
|
||||
// on stdout. It never touches the scanned filesystem; its only I/O is
|
||||
// the database, stdout, and stderr.
|
||||
func runReport(ctx context.Context) error {
|
||||
recs, err := loadRecords(ctx)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
dupes := collectDupeGroups(recs)
|
||||
|
||||
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
|
||||
out := bufio.NewWriterSize(stdout, ioBufSize)
|
||||
|
||||
_, err = fmt.Fprintln(out, "first\tdupe\tsize")
|
||||
if err != nil {
|
||||
return fmt.Errorf("write stdout: %w", err)
|
||||
}
|
||||
|
||||
dupeFiles := 0
|
||||
var (
|
||||
groups, dupeFiles int
|
||||
reclaimable int64
|
||||
writeErr error
|
||||
)
|
||||
|
||||
var reclaimable int64
|
||||
records, err := loadDupeRows(ctx, db,
|
||||
func(first, path string, size int64) error {
|
||||
// A group's first path is its first row; every other
|
||||
// path is a dupe.
|
||||
if path == first {
|
||||
groups++
|
||||
|
||||
for _, g := range dupes {
|
||||
for _, p := range g.paths[1:] {
|
||||
_, err = fmt.Fprintf(out, "%s\t%s\t%d\n",
|
||||
g.paths[0], p, g.size)
|
||||
if err != nil {
|
||||
return fmt.Errorf("write stdout: %w", err)
|
||||
return nil
|
||||
}
|
||||
|
||||
_, writeErr = fmt.Fprintf(out, "%s\t%s\t%d\n",
|
||||
escapePath(first), escapePath(path), size)
|
||||
dupeFiles++
|
||||
reclaimable += g.size
|
||||
}
|
||||
reclaimable += size
|
||||
|
||||
return writeErr
|
||||
})
|
||||
|
||||
if writeErr != nil {
|
||||
return fmt.Errorf("write stdout: %w", writeErr)
|
||||
}
|
||||
|
||||
if err != nil {
|
||||
return fmt.Errorf("database %s: %w", dbPath, err)
|
||||
}
|
||||
|
||||
err = out.Flush()
|
||||
@@ -105,55 +91,24 @@ func runReport(ctx context.Context) error {
|
||||
fmt.Fprintf(os.Stderr,
|
||||
"report: %d records read, %d duplicate groups, %d dupe files, "+
|
||||
"%s reclaimable\n",
|
||||
len(recs), len(dupes), dupeFiles, humanBytes(reclaimable))
|
||||
records, groups, dupeFiles, humanBytes(reclaimable))
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// collectDupeGroups groups records by signature and returns every group
|
||||
// with two or more paths, each group's paths sorted lexicographically,
|
||||
// groups ordered by size descending then by first path ascending.
|
||||
func collectDupeGroups(recs []scanRec) []dupeGroup {
|
||||
groups := make(map[fileSig][]string)
|
||||
|
||||
for _, r := range recs {
|
||||
// A record without a content hash has unknown content and is
|
||||
// never reported as a duplicate (README "Database").
|
||||
if r.content == "" {
|
||||
continue
|
||||
}
|
||||
|
||||
k := fileSig{
|
||||
size: r.size, head: r.head, tail: r.tail, content: r.content,
|
||||
}
|
||||
groups[k] = append(groups[k], r.path)
|
||||
// escapePath returns a path as it is written in a report column (README
|
||||
// "Report output format"): a backslash, tab, newline or carriage return
|
||||
// becomes \\, \t, \n or \r, and every other byte is kept as it is.
|
||||
// Grouping and sorting use the raw path, never this form.
|
||||
func escapePath(p string) string {
|
||||
// Most paths need no escaping; skip building a replacer for them.
|
||||
if !strings.ContainsAny(p, "\\\t\n\r") {
|
||||
return p
|
||||
}
|
||||
|
||||
var dupes []dupeGroup
|
||||
|
||||
for k, paths := range groups {
|
||||
if len(paths) < minGroupSize {
|
||||
continue
|
||||
}
|
||||
|
||||
slices.Sort(paths)
|
||||
dupes = append(dupes, dupeGroup{size: k.size, paths: paths})
|
||||
}
|
||||
|
||||
// Biggest reclaimable space first; ties broken by first path.
|
||||
slices.SortFunc(dupes, func(a, b dupeGroup) int {
|
||||
if a.size != b.size {
|
||||
if a.size > b.size {
|
||||
return -1
|
||||
}
|
||||
|
||||
return 1
|
||||
}
|
||||
|
||||
return strings.Compare(a.paths[0], b.paths[0])
|
||||
})
|
||||
|
||||
return dupes
|
||||
return strings.NewReplacer(
|
||||
`\`, `\\`, "\t", `\t`, "\n", `\n`, "\r", `\r`,
|
||||
).Replace(p)
|
||||
}
|
||||
|
||||
// humanBytes formats a byte count in human units (binary prefixes).
|
||||
|
||||
+229
-11
@@ -1,11 +1,229 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"database/sql"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"slices"
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func TestCollectDupeGroups(t *testing.T) {
|
||||
// awkwardDir is a directory name holding every byte the reports escape.
|
||||
const awkwardDir = "/d/\tone\ntwo\rthree\\four"
|
||||
|
||||
// awkwardPairRecs is a duplicate pair in sibling directories /d/A and
|
||||
// awkwardDir. A raw tab sorts before "A" but its escaped form `\t`
|
||||
// sorts after it, so awkwardDir coming first shows that sorting uses
|
||||
// the raw path.
|
||||
func awkwardPairRecs() []scanRec {
|
||||
return []scanRec{
|
||||
{size: 5, head: "h", tail: "t", content: "c", path: "/d/A/f"},
|
||||
{size: 5, head: "h", tail: "t", content: "c", path: awkwardDir + "/f"},
|
||||
}
|
||||
}
|
||||
|
||||
// seedDatabase writes recs into a fresh database and returns its path.
|
||||
func seedDatabase(t *testing.T, recs []scanRec) string {
|
||||
t.Helper()
|
||||
|
||||
path := testDBPath(t)
|
||||
|
||||
db, err := openScanDatabase(t.Context(), path)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
err = applyChanges(t.Context(), db, recs, nil, nil)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
err = db.Close()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
return path
|
||||
}
|
||||
|
||||
// dupeGroup is one duplicate group as report reads it: the size, and
|
||||
// the paths in report order, first path first.
|
||||
type dupeGroup struct {
|
||||
size int64
|
||||
paths []string
|
||||
}
|
||||
|
||||
// dupeGroups returns the duplicate groups report reads from db, in
|
||||
// report order.
|
||||
func dupeGroups(t *testing.T, db *sql.DB) []dupeGroup {
|
||||
t.Helper()
|
||||
|
||||
var groups []dupeGroup
|
||||
|
||||
_, err := loadDupeRows(t.Context(), db,
|
||||
func(first, path string, size int64) error {
|
||||
if path == first {
|
||||
groups = append(groups, dupeGroup{size: size})
|
||||
}
|
||||
|
||||
g := &groups[len(groups)-1]
|
||||
g.paths = append(g.paths, path)
|
||||
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
return groups
|
||||
}
|
||||
|
||||
// dupeGroupsOf writes recs into a fresh database and returns the
|
||||
// duplicate groups report reads from it.
|
||||
func dupeGroupsOf(t *testing.T, recs []scanRec) []dupeGroup {
|
||||
t.Helper()
|
||||
|
||||
db := openTestDB(t)
|
||||
|
||||
err := applyChanges(t.Context(), db, recs, nil, nil)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
return dupeGroups(t, db)
|
||||
}
|
||||
|
||||
func TestRunReportEscapesPaths(t *testing.T) {
|
||||
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := run([]string{cmdReport}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(report) = %d, want %d; stderr: %s",
|
||||
code, exitOK, stderr.String())
|
||||
}
|
||||
|
||||
want := "first\tdupe\tsize\n" +
|
||||
`/d/\tone\ntwo\rthree\\four/f` + "\t/d/A/f\t5\n"
|
||||
if got := stdout.String(); got != want {
|
||||
t.Errorf("stdout = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunReportsIgnoreInsertionOrder(t *testing.T) {
|
||||
// README §Constraints: identical database contents give identical
|
||||
// output, whatever order the records were inserted in.
|
||||
recs := append(smokeTreeRecs(), awkwardPairRecs()...)
|
||||
recs = append(recs,
|
||||
scanRec{size: 50, head: "b", tail: "b", content: "b", path: "/y/2"},
|
||||
scanRec{size: 50, head: "b", tail: "b", content: "b", path: "/y/1"},
|
||||
scanRec{size: 50, head: "a", tail: "a", content: "a", path: "/x/2"},
|
||||
scanRec{size: 50, head: "a", tail: "a", content: "a", path: "/x/1"},
|
||||
scanRec{size: 50, path: "/x/unhashed"},
|
||||
)
|
||||
|
||||
reversed := slices.Clone(recs)
|
||||
slices.Reverse(reversed)
|
||||
|
||||
for _, name := range []string{cmdReport, cmdTrees} {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
t.Setenv(databaseEnv, seedDatabase(t, recs))
|
||||
|
||||
forward := runStdout(t, name)
|
||||
|
||||
t.Setenv(databaseEnv, seedDatabase(t, reversed))
|
||||
|
||||
backward := runStdout(t, name)
|
||||
|
||||
if strings.Count(forward, "\n") < 3 {
|
||||
t.Errorf("stdout = %q, want at least two rows", forward)
|
||||
}
|
||||
|
||||
if forward != backward {
|
||||
t.Errorf("stdout depends on insertion order: %q vs %q",
|
||||
forward, backward)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// runStdout runs the subcommand name and returns its stdout, failing
|
||||
// the test unless it succeeds.
|
||||
func runStdout(t *testing.T, name string) string {
|
||||
t.Helper()
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := run([]string{name}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(%s) = %d, want %d; stderr: %s",
|
||||
name, code, exitOK, stderr.String())
|
||||
}
|
||||
|
||||
return stdout.String()
|
||||
}
|
||||
|
||||
func TestEscapePath(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
cases := map[string]string{
|
||||
"/srv/plain": "/srv/plain",
|
||||
"/a\tb": `/a\tb`,
|
||||
"/a\nb": `/a\nb`,
|
||||
"/a\rb": `/a\rb`,
|
||||
`/a\b`: `/a\\b`,
|
||||
`/a\tb`: `/a\\tb`,
|
||||
"/not-utf8\xff": "/not-utf8\xff",
|
||||
}
|
||||
for in, want := range cases {
|
||||
if got := escapePath(in); got != want {
|
||||
t.Errorf("escapePath(%q) = %q, want %q", in, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// TestWarnfEscapes checks that a warning naming a path that holds a
|
||||
// newline is still one line.
|
||||
//
|
||||
//nolint:paralleltest // replaces the process-wide os.Stderr
|
||||
func TestWarnfEscapes(t *testing.T) {
|
||||
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
saved := os.Stderr
|
||||
os.Stderr = f
|
||||
|
||||
t.Cleanup(func() {
|
||||
os.Stderr = saved
|
||||
|
||||
_ = f.Close()
|
||||
})
|
||||
|
||||
(&progress{}).warnf("stat %s: %s", "/d/a\nb", "gone")
|
||||
|
||||
_, err = f.Seek(0, io.SeekStart)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
got, err := io.ReadAll(f)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
want := `stat /d/a\nb: gone` + "\n"
|
||||
if string(got) != want {
|
||||
t.Errorf("warning = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestDupeGroups(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
recs := []scanRec{
|
||||
@@ -20,7 +238,7 @@ func TestCollectDupeGroups(t *testing.T) {
|
||||
{size: 7, head: "u", tail: "u", content: "u", path: "/lonely"},
|
||||
}
|
||||
|
||||
groups := collectDupeGroups(recs)
|
||||
groups := dupeGroupsOf(t, recs)
|
||||
if len(groups) != 2 {
|
||||
t.Fatalf("len(groups) = %d, want 2", len(groups))
|
||||
}
|
||||
@@ -38,7 +256,7 @@ func TestCollectDupeGroups(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestCollectDupeGroupsContentSeparates(t *testing.T) {
|
||||
func TestDupeGroupsContentSeparates(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// Same size, head, and tail, but different content hashes: the final
|
||||
@@ -53,7 +271,7 @@ func TestCollectDupeGroupsContentSeparates(t *testing.T) {
|
||||
{size: 100, head: "h", tail: "t", path: "/e"},
|
||||
}
|
||||
|
||||
groups := collectDupeGroups(recs)
|
||||
groups := dupeGroupsOf(t, recs)
|
||||
if len(groups) != 1 {
|
||||
t.Fatalf("len(groups) = %d, want 1 (only the matching content)",
|
||||
len(groups))
|
||||
@@ -64,7 +282,7 @@ func TestCollectDupeGroupsContentSeparates(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestCollectDupeGroupsMtimeExcluded(t *testing.T) {
|
||||
func TestDupeGroupsMtimeExcluded(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// mtime is informational only; records differing only in mtime
|
||||
@@ -74,13 +292,13 @@ func TestCollectDupeGroupsMtimeExcluded(t *testing.T) {
|
||||
{size: 9, mtime: 200, head: "h", tail: "t", content: "c", path: "/m/2"},
|
||||
}
|
||||
|
||||
groups := collectDupeGroups(recs)
|
||||
groups := dupeGroupsOf(t, recs)
|
||||
if len(groups) != 1 {
|
||||
t.Fatalf("len(groups) = %d, want 1", len(groups))
|
||||
}
|
||||
}
|
||||
|
||||
func TestCollectDupeGroupsTieBreak(t *testing.T) {
|
||||
func TestDupeGroupsTieBreak(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
recs := []scanRec{
|
||||
@@ -90,7 +308,7 @@ func TestCollectDupeGroupsTieBreak(t *testing.T) {
|
||||
{size: 50, head: "a", tail: "a", content: "a", path: "/alpha/1"},
|
||||
}
|
||||
|
||||
groups := collectDupeGroups(recs)
|
||||
groups := dupeGroupsOf(t, recs)
|
||||
if len(groups) != 2 {
|
||||
t.Fatalf("len(groups) = %d, want 2", len(groups))
|
||||
}
|
||||
@@ -102,7 +320,7 @@ func TestCollectDupeGroupsTieBreak(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestCollectDupeGroupsDeterministic(t *testing.T) {
|
||||
func TestDupeGroupsDeterministic(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
recs := []scanRec{
|
||||
@@ -112,12 +330,12 @@ func TestCollectDupeGroupsDeterministic(t *testing.T) {
|
||||
{size: 2, head: "b", tail: "b", content: "b", path: "/q/2"},
|
||||
}
|
||||
|
||||
forward := collectDupeGroups(recs)
|
||||
forward := dupeGroupsOf(t, recs)
|
||||
|
||||
reversed := slices.Clone(recs)
|
||||
slices.Reverse(reversed)
|
||||
|
||||
backward := collectDupeGroups(reversed)
|
||||
backward := dupeGroupsOf(t, reversed)
|
||||
if !slices.EqualFunc(forward, backward, func(a, b dupeGroup) bool {
|
||||
return a.size == b.size && slices.Equal(a.paths, b.paths)
|
||||
}) {
|
||||
|
||||
@@ -85,10 +85,13 @@ type fileMeta struct {
|
||||
// least one other file shares are ever hashed: a size-unique file
|
||||
// cannot be a duplicate. A file of headTailMin or more gets its content
|
||||
// hash only when its size, head, and tail match another file's. Flag
|
||||
// parsing and the at-least-one-operand check are done by cobra. Errors
|
||||
// are returned rather than exiting, so that the deferred close — which
|
||||
// checkpoints the SQLite WAL — always runs. Cancelling ctx unwinds the
|
||||
// worker pools and aborts the scan with the context's error.
|
||||
// parsing and the at-least-one-operand check are done by cobra. The
|
||||
// scan holds the lock on the database for its whole run, so a second
|
||||
// scan fails before it walks the filesystem or opens the database.
|
||||
// Errors are returned rather than exiting, so that the deferred close —
|
||||
// which takes the database out of WAL mode — always runs, and the lock
|
||||
// is released after it. Cancelling ctx unwinds the worker pools and
|
||||
// aborts the scan with the context's error.
|
||||
func runScan(ctx context.Context, roots []string, workers int,
|
||||
oneFS bool,
|
||||
) error {
|
||||
@@ -103,12 +106,19 @@ func runScan(ctx context.Context, roots []string, workers int,
|
||||
|
||||
dbPath := databasePath()
|
||||
|
||||
lock, err := lockScanDatabase(dbPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
defer func() { _ = lock.Close() }()
|
||||
|
||||
db, err := openScanDatabase(ctx, dbPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
defer func() { _ = db.Close() }()
|
||||
defer closeScanDatabase(ctx, db, dbPath)
|
||||
|
||||
st, err := syncScan(ctx, db, roots, workers, oneFS)
|
||||
if err != nil {
|
||||
@@ -210,14 +220,18 @@ type scanState struct {
|
||||
// in the content hash of every record of headTailMin or more whose
|
||||
// size, head, and tail match another record's). Records outside the
|
||||
// roots are never touched, except that the content phase fills in
|
||||
// their content hash.
|
||||
// their content hash. Operands the walk cannot start from are dropped
|
||||
// first, so the records beneath them count as outside the roots unless
|
||||
// they lie under another root.
|
||||
func syncScan(ctx context.Context, db *sql.DB, roots []string,
|
||||
workers int, oneFS bool,
|
||||
) (scanStats, error) {
|
||||
roots = pruneRoots(roots)
|
||||
|
||||
s := &scanState{db: db}
|
||||
|
||||
// Types are checked before pruning so that an operand under a
|
||||
// dropped one is still scanned, not dropped as lying under it.
|
||||
roots = pruneRoots(s.walkableRoots(roots))
|
||||
|
||||
err := s.loadIndex(ctx, roots)
|
||||
if err != nil {
|
||||
return s.st, err
|
||||
@@ -253,6 +267,35 @@ func syncScan(ctx context.Context, db *sql.DB, roots []string,
|
||||
return s.st, s.contentPhase(ctx, workers)
|
||||
}
|
||||
|
||||
// walkableRoots returns the operands the walk can start from: regular
|
||||
// files, and directories not named .zfs. Every other operand is warned
|
||||
// about, counted as skipped, and dropped. A dropped operand is no
|
||||
// longer a root, so the records stored beneath it count as outside the
|
||||
// roots and are not deleted as unverified, unless it lies under another
|
||||
// root. An operand that fails lstat here is kept, and the walk warns
|
||||
// about it.
|
||||
func (s *scanState) walkableRoots(roots []string) []string {
|
||||
kept := make([]string, 0, len(roots))
|
||||
|
||||
for _, root := range roots {
|
||||
fi, err := os.Lstat(root)
|
||||
if err == nil {
|
||||
warn := operandWarning(root, fi)
|
||||
if warn != "" {
|
||||
s.st.skipped++
|
||||
|
||||
fmt.Fprintln(os.Stderr, escapePath(warn))
|
||||
|
||||
continue
|
||||
}
|
||||
}
|
||||
|
||||
kept = append(kept, root)
|
||||
}
|
||||
|
||||
return kept
|
||||
}
|
||||
|
||||
// loadIndex indexes the database records under the scan roots for
|
||||
// change detection and collects the sizes of every record outside
|
||||
// them: out-of-scope records join the size census so a scanned file
|
||||
@@ -789,11 +832,43 @@ func sendEvent(ctx context.Context, events chan<- walkEvent,
|
||||
}
|
||||
}
|
||||
|
||||
// operandWarning returns the one-line warning for an operand the walk
|
||||
// does not start from, naming the path and what it is, or "" for one it
|
||||
// does: a regular file, or a directory not named .zfs. Symlinks are
|
||||
// never followed, including as operands.
|
||||
func operandWarning(root string, fi fs.FileInfo) string {
|
||||
var kind string
|
||||
|
||||
switch mode := fi.Mode(); {
|
||||
case mode.IsRegular():
|
||||
return ""
|
||||
case mode.IsDir():
|
||||
if filepath.Base(root) != ".zfs" {
|
||||
return ""
|
||||
}
|
||||
|
||||
kind = ".zfs directory"
|
||||
case mode&fs.ModeSymlink != 0:
|
||||
kind = "symlink"
|
||||
case mode&fs.ModeSocket != 0:
|
||||
kind = "socket"
|
||||
case mode&fs.ModeNamedPipe != 0:
|
||||
kind = "FIFO"
|
||||
case mode&fs.ModeDevice != 0:
|
||||
kind = "device node"
|
||||
default:
|
||||
kind = "non-regular file"
|
||||
}
|
||||
|
||||
return fmt.Sprintf("walk %s: skipping %s operand", root, kind)
|
||||
}
|
||||
|
||||
// seedRoot turns one PATH operand into the walk's starting state: a
|
||||
// regular-file operand is statted and emitted directly, a directory
|
||||
// operand becomes an initial job, and a symlink or other non-regular
|
||||
// operand yields nothing (symlinks are never followed, including as
|
||||
// operands).
|
||||
// regular-file operand is statted and emitted directly, and a directory
|
||||
// operand becomes an initial job. walkableRoots has already dropped
|
||||
// every other operand. One that has changed into something else since
|
||||
// is warned about and skipped here; it is still a root, so the records
|
||||
// stored beneath it are deleted as unverified.
|
||||
func seedRoot(ctx context.Context, root string,
|
||||
events chan<- walkEvent,
|
||||
) []dirJob {
|
||||
@@ -807,30 +882,30 @@ func seedRoot(ctx context.Context, root string,
|
||||
return nil
|
||||
}
|
||||
|
||||
switch {
|
||||
case fi.IsDir():
|
||||
if filepath.Base(root) == ".zfs" {
|
||||
return nil
|
||||
}
|
||||
warn := operandWarning(root, fi)
|
||||
if warn != "" {
|
||||
sendEvent(ctx, events, walkEvent{warn: warn, fail: true})
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
if fi.IsDir() {
|
||||
dev, ok := deviceOfInfo(fi)
|
||||
|
||||
return []dirJob{{path: root, rootDev: dev, rootDevOK: ok}}
|
||||
case fi.Mode().IsRegular():
|
||||
dev, ino := inodeOfInfo(fi)
|
||||
|
||||
sendEvent(ctx, events, walkEvent{rec: fileRec{
|
||||
path: root,
|
||||
size: fi.Size(),
|
||||
mtime: fi.ModTime().Unix(),
|
||||
dev: dev,
|
||||
ino: ino,
|
||||
}})
|
||||
|
||||
return nil
|
||||
default:
|
||||
return nil
|
||||
}
|
||||
|
||||
dev, ino := inodeOfInfo(fi)
|
||||
|
||||
sendEvent(ctx, events, walkEvent{rec: fileRec{
|
||||
path: root,
|
||||
size: fi.Size(),
|
||||
mtime: fi.ModTime().Unix(),
|
||||
dev: dev,
|
||||
ino: ino,
|
||||
}})
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// startWalkWorkers starts the walk worker pool. Each worker processes
|
||||
|
||||
+31
-27
@@ -7,6 +7,7 @@ import (
|
||||
"database/sql"
|
||||
"encoding/hex"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"runtime"
|
||||
@@ -399,7 +400,7 @@ func TestScanContentGate(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
groups := collectDupeGroups(recs)
|
||||
groups := dupeGroups(t, db)
|
||||
if len(groups) != 1 || !slices.Equal(groups[0].paths, same) {
|
||||
t.Fatalf("groups = %+v, want only the identical pair %q",
|
||||
groups, same)
|
||||
@@ -434,7 +435,7 @@ func TestScanContentAcrossOperands(t *testing.T) {
|
||||
want := []string{a, b}
|
||||
slices.Sort(want)
|
||||
|
||||
groups := collectDupeGroups(recs)
|
||||
groups := dupeGroups(t, db)
|
||||
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
|
||||
t.Fatalf("groups = %+v, want the pair %q", groups, want)
|
||||
}
|
||||
@@ -461,7 +462,7 @@ func TestScanContentWithinOperand(t *testing.T) {
|
||||
|
||||
want := []string{stored, added}
|
||||
|
||||
groups := collectDupeGroups(dbRecords(t, db))
|
||||
groups := dupeGroups(t, db)
|
||||
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
|
||||
t.Fatalf("groups = %+v, want the pair %q", groups, want)
|
||||
}
|
||||
@@ -519,7 +520,7 @@ func TestScanContentStalePartners(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
if groups := collectDupeGroups(recs); len(groups) != 0 {
|
||||
if groups := dupeGroups(t, db); len(groups) != 0 {
|
||||
t.Errorf("groups = %+v, want none", groups)
|
||||
}
|
||||
}
|
||||
@@ -570,7 +571,7 @@ func TestScanContentHashedStalePartners(t *testing.T) {
|
||||
|
||||
// The stored records lie outside the operand and are left as they
|
||||
// are, so they still group with each other, but not with the copy.
|
||||
groups := collectDupeGroups(recs)
|
||||
groups := dupeGroups(t, db)
|
||||
if len(groups) != 1 || !slices.Equal(groups[0].paths, stored) {
|
||||
t.Errorf("groups = %+v, want only the stored pair %q", groups, stored)
|
||||
}
|
||||
@@ -619,7 +620,7 @@ func TestScanContentReadFailure(t *testing.T) {
|
||||
want := []string{a, b}
|
||||
slices.Sort(want)
|
||||
|
||||
groups := collectDupeGroups(dbRecords(t, db))
|
||||
groups := dupeGroups(t, db)
|
||||
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
|
||||
t.Fatalf("groups = %+v, want the pair %q after the retry",
|
||||
groups, want)
|
||||
@@ -856,9 +857,11 @@ func TestWalkFileAndSymlinkOperands(t *testing.T) {
|
||||
t.Fatalf("file operand: recs = %+v, errs = %d", recs, errs)
|
||||
}
|
||||
|
||||
// A symlink operand is not followed and yields nothing.
|
||||
// A symlink operand that reaches the walk (it became one after
|
||||
// walkableRoots checked it) is not followed: it yields a warning and
|
||||
// no records.
|
||||
recs, errs = collectWalk(t, []string{link}, false, 2)
|
||||
if errs != 0 || len(recs) != 0 {
|
||||
if errs != 1 || len(recs) != 0 {
|
||||
t.Fatalf("symlink operand: recs = %+v, errs = %d", recs, errs)
|
||||
}
|
||||
}
|
||||
@@ -953,11 +956,16 @@ func syncTree(t *testing.T, db *sql.DB, roots ...string) scanStats {
|
||||
return st
|
||||
}
|
||||
|
||||
// dbRecords returns every record currently in the database.
|
||||
// dbRecords returns every record currently in the database, in path
|
||||
// order.
|
||||
func dbRecords(t *testing.T, db *sql.DB) []scanRec {
|
||||
t.Helper()
|
||||
|
||||
recs, err := loadFileRows(t.Context(), db)
|
||||
var recs []scanRec
|
||||
|
||||
err := loadFileRows(t.Context(), db, func(r scanRec) {
|
||||
recs = append(recs, r)
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
@@ -994,10 +1002,10 @@ func recordPaths(recs []scanRec) []string {
|
||||
|
||||
// assertSmokeDupeGroups checks the file-level duplicate groups for the
|
||||
// smoke tree rooted at dir.
|
||||
func assertSmokeDupeGroups(t *testing.T, dir string, parsed []scanRec) {
|
||||
func assertSmokeDupeGroups(t *testing.T, dir string, db *sql.DB) {
|
||||
t.Helper()
|
||||
|
||||
groups := collectDupeGroups(parsed)
|
||||
groups := dupeGroups(t, db)
|
||||
if len(groups) != 5 {
|
||||
t.Fatalf("len(groups) = %d, want 5", len(groups))
|
||||
}
|
||||
@@ -1022,11 +1030,10 @@ func assertSmokeDupeGroups(t *testing.T, dir string, parsed []scanRec) {
|
||||
|
||||
// assertSmokeTreeGroups checks the duplicate-tree groups for the smoke
|
||||
// tree rooted at dir.
|
||||
func assertSmokeTreeGroups(t *testing.T, dir string, parsed []scanRec) {
|
||||
func assertSmokeTreeGroups(t *testing.T, dir string, db *sql.DB) {
|
||||
t.Helper()
|
||||
|
||||
super, dirs := buildHierarchy(parsed)
|
||||
super.compute()
|
||||
super, dirs := dbTree(t, db)
|
||||
|
||||
tg := collectTreeGroups(dirs, super)
|
||||
if len(tg) != 1 {
|
||||
@@ -1060,8 +1067,8 @@ func TestScanPipeline(t *testing.T) {
|
||||
t.Fatalf("len(records) = %d, want %d", len(parsed), smokeTreeFiles)
|
||||
}
|
||||
|
||||
assertSmokeDupeGroups(t, dir, parsed)
|
||||
assertSmokeTreeGroups(t, dir, parsed)
|
||||
assertSmokeDupeGroups(t, dir, db)
|
||||
assertSmokeTreeGroups(t, dir, db)
|
||||
}
|
||||
|
||||
func TestSyncScanUnchangedReuse(t *testing.T) {
|
||||
@@ -1295,7 +1302,7 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
if groups := collectDupeGroups(recs); len(groups) != 0 {
|
||||
if groups := dupeGroups(t, db); len(groups) != 0 {
|
||||
t.Fatalf("groups = %+v, want none from unhashed records", groups)
|
||||
}
|
||||
|
||||
@@ -1310,7 +1317,7 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
|
||||
st)
|
||||
}
|
||||
|
||||
groups := collectDupeGroups(dbRecords(t, db))
|
||||
groups := dupeGroups(t, db)
|
||||
if len(groups) != 1 {
|
||||
t.Fatalf("groups = %+v, want the a/c pair", groups)
|
||||
}
|
||||
@@ -1346,8 +1353,7 @@ func TestTreesUnhashedNeverEqual(t *testing.T) {
|
||||
}
|
||||
|
||||
for name, unknown := range cases {
|
||||
super, dirs := buildHierarchy(append(slices.Clone(shared), unknown...))
|
||||
super.compute()
|
||||
super, dirs := treeOf(t, append(slices.Clone(shared), unknown...))
|
||||
|
||||
if tg := collectTreeGroups(dirs, super); len(tg) != 0 {
|
||||
t.Errorf("%s: tree groups = %d, want 0 (the files may differ)",
|
||||
@@ -1384,7 +1390,7 @@ func TestScanHardlinksReadOnce(t *testing.T) {
|
||||
t.Fatalf("hardlink hashes differ: %+v vs %+v", ra, rb)
|
||||
}
|
||||
|
||||
if groups := collectDupeGroups(recs); len(groups) != 1 {
|
||||
if groups := dupeGroups(t, db); len(groups) != 1 {
|
||||
t.Fatalf("groups = %+v, want the hardlink pair", groups)
|
||||
}
|
||||
}
|
||||
@@ -1548,7 +1554,7 @@ func TestScanHashWriteFailureUnwindsPool(t *testing.T) {
|
||||
|
||||
code := run([]string{
|
||||
cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir,
|
||||
}, &stderr)
|
||||
}, io.Discard, &stderr)
|
||||
if code != exitFatal {
|
||||
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
|
||||
code, exitFatal, stderr.String())
|
||||
@@ -1625,10 +1631,8 @@ func TestReportsNeverTouchFilesystem(t *testing.T) {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
recs := dbRecords(t, db)
|
||||
|
||||
assertSmokeDupeGroups(t, dir, recs)
|
||||
assertSmokeTreeGroups(t, dir, recs)
|
||||
assertSmokeDupeGroups(t, dir, db)
|
||||
assertSmokeTreeGroups(t, dir, db)
|
||||
}
|
||||
|
||||
func TestUnderRoot(t *testing.T) {
|
||||
|
||||
+4
-4
@@ -16,10 +16,10 @@
|
||||
# implies the repo is green.
|
||||
#
|
||||
# That implication holds only because of CHECK_EPOCH. A COPY layer is
|
||||
# invalidated by changed content, and a merge commit's tree is
|
||||
# byte-identical to the branch head it merges, so without a fresh value
|
||||
# here Docker serves the gate layers from cache and the build reports a
|
||||
# green it never earned. Passing the current epoch invalidates the gate
|
||||
# invalidated only by changed content, and a rebuild of an unchanged
|
||||
# checkout sends the same content, so without a fresh value here Docker
|
||||
# serves the gate layers from cache and the build reports a green it
|
||||
# never earned. Passing the current epoch invalidates the gate
|
||||
# layers on every run while leaving the pinned base images and
|
||||
# go mod download cached; see the Dockerfile for the placement.
|
||||
set -eu
|
||||
|
||||
@@ -5,49 +5,59 @@ import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"slices"
|
||||
"strconv"
|
||||
"strings"
|
||||
)
|
||||
|
||||
// fileSig is a file's duplicate signature; mtime is excluded.
|
||||
type fileSig struct {
|
||||
size int64
|
||||
head string
|
||||
tail string
|
||||
content string
|
||||
}
|
||||
|
||||
// treeNode is one directory reconstructed from the scan stream.
|
||||
// treeNode is one directory reconstructed from the record paths.
|
||||
type treeNode struct {
|
||||
path string
|
||||
parent *treeNode
|
||||
dirs map[string]*treeNode
|
||||
files map[string]fileSig
|
||||
path string
|
||||
parent *treeNode
|
||||
// entries holds the serialized child entries until the digest is
|
||||
// computed from them, and is then dropped.
|
||||
entries []string
|
||||
digest [sha256.Size]byte
|
||||
fileCount int64
|
||||
totalSize int64
|
||||
}
|
||||
|
||||
// runTrees implements the trees subcommand: it reads every record from
|
||||
// the database, reconstructs the directory hierarchy from the record
|
||||
// paths, computes a Merkle-style digest per directory, and prints
|
||||
// maximal duplicate-tree groups as TSV on stdout. It never touches the
|
||||
// scanned filesystem; its only I/O is the database, stdout, and
|
||||
// stderr.
|
||||
func runTrees(ctx context.Context) error {
|
||||
recs, err := loadRecords(ctx)
|
||||
// the database in path order, reconstructs the directory hierarchy from
|
||||
// the record paths, computes a Merkle-style digest per directory, and
|
||||
// prints maximal duplicate-tree groups as TSV on stdout. It never
|
||||
// touches the scanned filesystem; its only I/O is the database, stdout,
|
||||
// and stderr. Any database problem, including a missing database, is
|
||||
// fatal.
|
||||
func runTrees(ctx context.Context, stdout io.Writer) error {
|
||||
dbPath := databasePath()
|
||||
|
||||
db, err := openReportDatabase(ctx, dbPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
super, allDirs := buildHierarchy(recs)
|
||||
super.compute()
|
||||
defer func() { _ = db.Close() }()
|
||||
|
||||
records := 0
|
||||
tree := newTreeBuilder()
|
||||
|
||||
err = loadFileRows(ctx, db, func(r scanRec) {
|
||||
records++
|
||||
|
||||
tree.add(r)
|
||||
})
|
||||
if err != nil {
|
||||
return fmt.Errorf("database %s: %w", dbPath, err)
|
||||
}
|
||||
|
||||
super, allDirs := tree.finish()
|
||||
|
||||
dupes := collectTreeGroups(allDirs, super)
|
||||
|
||||
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
|
||||
out := bufio.NewWriterSize(stdout, ioBufSize)
|
||||
|
||||
_, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
|
||||
if err != nil {
|
||||
@@ -62,7 +72,8 @@ func runTrees(ctx context.Context) error {
|
||||
first := g[0]
|
||||
for _, n := range g[1:] {
|
||||
_, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n",
|
||||
first.path, n.path, first.fileCount, first.totalSize)
|
||||
escapePath(first.path), escapePath(n.path),
|
||||
first.fileCount, first.totalSize)
|
||||
if err != nil {
|
||||
return fmt.Errorf("write stdout: %w", err)
|
||||
}
|
||||
@@ -80,65 +91,128 @@ func runTrees(ctx context.Context) error {
|
||||
fmt.Fprintf(os.Stderr,
|
||||
"trees: %d records read, %d duplicate tree groups, %d dupe trees, "+
|
||||
"%s reclaimable\n",
|
||||
len(recs), len(dupes), dupeTrees, humanBytes(reclaimable))
|
||||
records, len(dupes), dupeTrees, humanBytes(reclaimable))
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// buildHierarchy reconstructs the directory hierarchy from the record
|
||||
// paths under a synthetic super-root. Paths are split on "/"; for
|
||||
// absolute paths the first component is empty, which simply becomes a
|
||||
// top-level node representing "/". It returns the super-root and every
|
||||
// directory node created.
|
||||
func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
|
||||
// treeBuilder reconstructs the directory hierarchy from records added
|
||||
// in path order, under a synthetic super-root. Paths are split on "/";
|
||||
// for absolute paths the first component is empty, which becomes the
|
||||
// top-level directory with path "/". In path order all the paths under
|
||||
// one directory come together, so a directory is complete once a path
|
||||
// outside it is added: its digest is computed then and its entries are
|
||||
// dropped. Only the directories holding the latest path keep entries.
|
||||
type treeBuilder struct {
|
||||
super *treeNode
|
||||
// open lists the directories holding the latest path, outermost
|
||||
// first, starting with the super-root; names[i] is open[i]'s name.
|
||||
open []*treeNode
|
||||
names []string
|
||||
// dirs lists every completed directory.
|
||||
dirs []*treeNode
|
||||
}
|
||||
|
||||
func newTreeBuilder() *treeBuilder {
|
||||
super := &treeNode{}
|
||||
|
||||
var allDirs []*treeNode
|
||||
return &treeBuilder{
|
||||
super: super,
|
||||
open: []*treeNode{super},
|
||||
names: []string{""},
|
||||
}
|
||||
}
|
||||
|
||||
for _, r := range recs {
|
||||
comps := strings.Split(r.path, "/")
|
||||
// add adds one record. Each record must come after the previous one in
|
||||
// path order (byte order); otherwise a completed directory would be
|
||||
// started again as a second directory with the same path.
|
||||
func (b *treeBuilder) add(r scanRec) {
|
||||
comps := strings.Split(r.path, "/")
|
||||
dirNames, name := comps[:len(comps)-1], comps[len(comps)-1]
|
||||
|
||||
node := super
|
||||
for _, c := range comps[:len(comps)-1] {
|
||||
child := node.dirs[c]
|
||||
if child == nil {
|
||||
childPath := c
|
||||
if node != super {
|
||||
childPath = node.path + "/" + c
|
||||
}
|
||||
|
||||
child = &treeNode{path: childPath, parent: node}
|
||||
if node.dirs == nil {
|
||||
node.dirs = make(map[string]*treeNode)
|
||||
}
|
||||
|
||||
node.dirs[c] = child
|
||||
allDirs = append(allDirs, child)
|
||||
}
|
||||
|
||||
node = child
|
||||
}
|
||||
|
||||
if node.files == nil {
|
||||
node.files = make(map[string]fileSig)
|
||||
}
|
||||
|
||||
sig := fileSig{
|
||||
size: r.size, head: r.head, tail: r.tail, content: r.content,
|
||||
}
|
||||
|
||||
// A record without a content hash has unknown content (README
|
||||
// "Database"): give it a signature no other file can share, so
|
||||
// trees containing it never compare equal. Real hashes are
|
||||
// hex, so the NUL-prefixed form cannot collide.
|
||||
if sig.content == "" {
|
||||
sig.content = "unhashed\x00" + r.path
|
||||
}
|
||||
|
||||
node.files[comps[len(comps)-1]] = sig
|
||||
// Keep the open directories that hold this path; complete the rest.
|
||||
depth := 1
|
||||
for depth < len(b.open) && depth <= len(dirNames) &&
|
||||
b.names[depth] == dirNames[depth-1] {
|
||||
depth++
|
||||
}
|
||||
|
||||
return super, allDirs
|
||||
b.closeTo(depth)
|
||||
|
||||
for _, c := range dirNames[depth-1:] {
|
||||
b.openDir(c)
|
||||
}
|
||||
|
||||
dir := b.open[len(b.open)-1]
|
||||
dir.entries = append(dir.entries, fileEntry(name, r))
|
||||
dir.fileCount++
|
||||
dir.totalSize += r.size
|
||||
}
|
||||
|
||||
// openDir starts the directory called name inside the innermost open
|
||||
// one.
|
||||
func (b *treeBuilder) openDir(name string) {
|
||||
parent := b.open[len(b.open)-1]
|
||||
path := parent.path + "/" + name
|
||||
|
||||
// The root directory's path is "/", not empty, and its children's
|
||||
// paths start with one slash, not two.
|
||||
switch {
|
||||
case parent == b.super && name == "":
|
||||
path = "/"
|
||||
case parent == b.super:
|
||||
path = name
|
||||
case parent.path == "/":
|
||||
path = "/" + name
|
||||
}
|
||||
|
||||
b.open = append(b.open, &treeNode{path: path, parent: parent})
|
||||
b.names = append(b.names, name)
|
||||
}
|
||||
|
||||
// closeTo completes the open directories after the first n, innermost
|
||||
// first: each one's digest is computed and entered in its parent along
|
||||
// with its totals.
|
||||
func (b *treeBuilder) closeTo(n int) {
|
||||
for len(b.open) > n {
|
||||
last := len(b.open) - 1
|
||||
dir, name := b.open[last], b.names[last]
|
||||
b.open, b.names = b.open[:last], b.names[:last]
|
||||
|
||||
dir.computeDigest()
|
||||
|
||||
dir.parent.entries = append(dir.parent.entries,
|
||||
"d\x00"+name+"\x00"+string(dir.digest[:]))
|
||||
dir.parent.fileCount += dir.fileCount
|
||||
dir.parent.totalSize += dir.totalSize
|
||||
b.dirs = append(b.dirs, dir)
|
||||
}
|
||||
}
|
||||
|
||||
// finish completes every open directory and returns the super-root and
|
||||
// every directory.
|
||||
func (b *treeBuilder) finish() (*treeNode, []*treeNode) {
|
||||
b.closeTo(1)
|
||||
|
||||
return b.super, b.dirs
|
||||
}
|
||||
|
||||
// fileEntry serializes a file child for its directory's digest: its
|
||||
// name and its signature (size, head, tail, content); mtime is
|
||||
// excluded.
|
||||
func fileEntry(name string, r scanRec) string {
|
||||
content := r.content
|
||||
|
||||
// A record without a content hash has unknown content (README
|
||||
// "Database"): give it a signature no other file can share, so
|
||||
// trees containing it never compare equal. Real hashes are hex, so
|
||||
// the NUL-prefixed form cannot collide.
|
||||
if content == "" {
|
||||
content = "unhashed\x00" + r.path
|
||||
}
|
||||
|
||||
return "f\x00" + name + "\x00" + strconv.FormatInt(r.size, 10) +
|
||||
"\x00" + r.head + "\x00" + r.tail + "\x00" + content
|
||||
}
|
||||
|
||||
// collectTreeGroups groups directories by digest and returns every
|
||||
@@ -180,38 +254,22 @@ func collectTreeGroups(allDirs []*treeNode, super *treeNode) [][]*treeNode {
|
||||
return dupes
|
||||
}
|
||||
|
||||
// compute fills in digest, fileCount, and totalSize for n and all of
|
||||
// its descendants. A directory's digest is the SHA-256 of its child
|
||||
// entries — files serialized with name and signature, subdirectories
|
||||
// with name and recursive digest — sorted byte-lexicographically.
|
||||
// Filenames cannot contain NUL or "/", so NUL delimiters are
|
||||
// unambiguous.
|
||||
func (n *treeNode) compute() {
|
||||
entries := make([]string, 0, len(n.dirs)+len(n.files))
|
||||
for name, sig := range n.files {
|
||||
entries = append(entries,
|
||||
"f\x00"+name+"\x00"+strconv.FormatInt(sig.size, 10)+
|
||||
"\x00"+sig.head+"\x00"+sig.tail+"\x00"+sig.content)
|
||||
n.fileCount++
|
||||
n.totalSize += sig.size
|
||||
}
|
||||
|
||||
for name, child := range n.dirs {
|
||||
child.compute()
|
||||
entries = append(entries, "d\x00"+name+"\x00"+string(child.digest[:]))
|
||||
n.fileCount += child.fileCount
|
||||
n.totalSize += child.totalSize
|
||||
}
|
||||
|
||||
slices.Sort(entries)
|
||||
// computeDigest sets n's digest and drops its entries. A directory's
|
||||
// digest is the SHA-256 of its child entries — files serialized with
|
||||
// name and signature, subdirectories with name and recursive digest —
|
||||
// sorted byte-lexicographically. Filenames cannot contain NUL or "/",
|
||||
// so NUL delimiters are unambiguous.
|
||||
func (n *treeNode) computeDigest() {
|
||||
slices.Sort(n.entries)
|
||||
|
||||
h := sha256.New()
|
||||
for _, e := range entries {
|
||||
for _, e := range n.entries {
|
||||
h.Write([]byte(e))
|
||||
h.Write([]byte{0})
|
||||
}
|
||||
|
||||
copy(n.digest[:], h.Sum(nil))
|
||||
n.entries = nil
|
||||
}
|
||||
|
||||
// suppressed reports whether a duplicate-tree group is non-maximal: its
|
||||
|
||||
+129
-17
@@ -1,6 +1,8 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"database/sql"
|
||||
"slices"
|
||||
"testing"
|
||||
)
|
||||
@@ -29,6 +31,36 @@ func smokeTreeRecs() []scanRec {
|
||||
}
|
||||
}
|
||||
|
||||
// dbTree builds the directory hierarchy from the records in db the way
|
||||
// trees does, and returns the super-root and every directory.
|
||||
func dbTree(t *testing.T, db *sql.DB) (*treeNode, []*treeNode) {
|
||||
t.Helper()
|
||||
|
||||
tree := newTreeBuilder()
|
||||
|
||||
err := loadFileRows(t.Context(), db, tree.add)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
return tree.finish()
|
||||
}
|
||||
|
||||
// treeOf writes recs into a fresh database and builds the directory
|
||||
// hierarchy from it the way trees does.
|
||||
func treeOf(t *testing.T, recs []scanRec) (*treeNode, []*treeNode) {
|
||||
t.Helper()
|
||||
|
||||
db := openTestDB(t)
|
||||
|
||||
err := applyChanges(t.Context(), db, recs, nil, nil)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
return dbTree(t, db)
|
||||
}
|
||||
|
||||
// nodeByPath finds the directory node with the given path.
|
||||
func nodeByPath(t *testing.T, dirs []*treeNode, path string) *treeNode {
|
||||
t.Helper()
|
||||
@@ -59,11 +91,10 @@ func groupPaths(groups [][]*treeNode) [][]string {
|
||||
return out
|
||||
}
|
||||
|
||||
func TestBuildHierarchyCounts(t *testing.T) {
|
||||
func TestTreeCounts(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
super, dirs := buildHierarchy(smokeTreeRecs())
|
||||
super.compute()
|
||||
_, dirs := treeOf(t, smokeTreeRecs())
|
||||
|
||||
d := nodeByPath(t, dirs, "/d")
|
||||
if d.fileCount != 6 || d.totalSize != 9300 {
|
||||
@@ -84,11 +115,98 @@ func TestBuildHierarchyCounts(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
func TestTreeRootPath(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// The root directory's path is "/", never empty, and its
|
||||
// children's paths start with a single slash.
|
||||
_, dirs := treeOf(t, []scanRec{{path: "/f"}, {path: "/srv/g"}})
|
||||
|
||||
got := make([]string, 0, len(dirs))
|
||||
for _, d := range dirs {
|
||||
got = append(got, d.path)
|
||||
}
|
||||
|
||||
slices.Sort(got)
|
||||
|
||||
want := []string{"/", "/srv"}
|
||||
if !slices.Equal(got, want) {
|
||||
t.Fatalf("directory paths = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTreeNamesSortingBeforeSlash(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// In path order "/a/b-x/f" and "/a/b.txt" come between the file
|
||||
// "/a/b" and "/a/b/f", because "-" and "." sort before "/". Each
|
||||
// directory must still be built once, whole, so /a matches /c.
|
||||
recs := make([]scanRec, 0, 8)
|
||||
|
||||
for _, top := range []string{"/a", "/c"} {
|
||||
for _, p := range []string{"/b", "/b-x/f", "/b.txt", "/b/f"} {
|
||||
content := "c"
|
||||
if p == "/b-x/f" {
|
||||
content = "other"
|
||||
}
|
||||
|
||||
recs = append(recs, scanRec{
|
||||
size: 1, head: "h", tail: "t", content: content, path: top + p,
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
super, dirs := treeOf(t, recs)
|
||||
|
||||
got := make([]string, 0, len(dirs))
|
||||
for _, d := range dirs {
|
||||
got = append(got, d.path)
|
||||
}
|
||||
|
||||
slices.Sort(got)
|
||||
|
||||
want := []string{"/", "/a", "/a/b", "/a/b-x", "/c", "/c/b", "/c/b-x"}
|
||||
if !slices.Equal(got, want) {
|
||||
t.Fatalf("directory paths = %q, want %q", got, want)
|
||||
}
|
||||
|
||||
groups := collectTreeGroups(dirs, super)
|
||||
|
||||
gotGroups := groupPaths(groups)
|
||||
wantGroups := [][]string{{"/a", "/c"}}
|
||||
|
||||
if !slices.EqualFunc(gotGroups, wantGroups, slices.Equal) {
|
||||
t.Fatalf("groups = %v, want %v", gotGroups, wantGroups)
|
||||
}
|
||||
|
||||
if groups[0][0].fileCount != 4 || groups[0][0].totalSize != 4 {
|
||||
t.Errorf("group totals: %d files %d bytes, want 4 4",
|
||||
groups[0][0].fileCount, groups[0][0].totalSize)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunTreesEscapesPaths(t *testing.T) {
|
||||
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := run([]string{cmdTrees}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
|
||||
code, exitOK, stderr.String())
|
||||
}
|
||||
|
||||
want := "first\tdupe\tfiles\tsize\n" +
|
||||
`/d/\tone\ntwo\rthree\\four` + "\t/d/A\t1\t5\n"
|
||||
if got := stdout.String(); got != want {
|
||||
t.Errorf("stdout = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
func TestTreeDigests(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
super, dirs := buildHierarchy(smokeTreeRecs())
|
||||
super.compute()
|
||||
_, dirs := treeOf(t, smokeTreeRecs())
|
||||
|
||||
t1 := nodeByPath(t, dirs, "/d/t1")
|
||||
t2 := nodeByPath(t, dirs, "/d/t2")
|
||||
@@ -121,8 +239,7 @@ func TestTreeDigestContentSensitivity(t *testing.T) {
|
||||
{size: 10, head: "DIFF", tail: sharedTail, content: "c", path: "/r/b/f"},
|
||||
}
|
||||
|
||||
super, dirs := buildHierarchy(recs)
|
||||
super.compute()
|
||||
_, dirs := treeOf(t, recs)
|
||||
|
||||
a := nodeByPath(t, dirs, "/r/a")
|
||||
b := nodeByPath(t, dirs, "/r/b")
|
||||
@@ -135,8 +252,7 @@ func TestTreeDigestContentSensitivity(t *testing.T) {
|
||||
func TestCollectTreeGroupsMaximal(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
super, dirs := buildHierarchy(smokeTreeRecs())
|
||||
super.compute()
|
||||
super, dirs := treeOf(t, smokeTreeRecs())
|
||||
|
||||
groups := collectTreeGroups(dirs, super)
|
||||
|
||||
@@ -160,16 +276,14 @@ func TestCollectTreeGroupsDeterministic(t *testing.T) {
|
||||
|
||||
recs := smokeTreeRecs()
|
||||
|
||||
super, dirs := buildHierarchy(recs)
|
||||
super.compute()
|
||||
super, dirs := treeOf(t, recs)
|
||||
|
||||
forward := groupPaths(collectTreeGroups(dirs, super))
|
||||
|
||||
reversed := slices.Clone(recs)
|
||||
slices.Reverse(reversed)
|
||||
|
||||
superR, dirsR := buildHierarchy(reversed)
|
||||
superR.compute()
|
||||
superR, dirsR := treeOf(t, reversed)
|
||||
|
||||
backward := groupPaths(collectTreeGroups(dirsR, superR))
|
||||
if !slices.EqualFunc(forward, backward, slices.Equal) {
|
||||
@@ -188,8 +302,7 @@ func TestCollectTreeGroupsSiblings(t *testing.T) {
|
||||
{size: 10, head: "h", tail: "t", content: "c", path: "/p/x2/f"},
|
||||
}
|
||||
|
||||
super, dirs := buildHierarchy(recs)
|
||||
super.compute()
|
||||
super, dirs := treeOf(t, recs)
|
||||
|
||||
got := groupPaths(collectTreeGroups(dirs, super))
|
||||
|
||||
@@ -211,8 +324,7 @@ func TestCollectTreeGroupsDifferingParents(t *testing.T) {
|
||||
{size: 10, head: "h", tail: "t", content: "c", path: "/q/b/x/f"},
|
||||
}
|
||||
|
||||
super, dirs := buildHierarchy(recs)
|
||||
super.compute()
|
||||
super, dirs := treeOf(t, recs)
|
||||
|
||||
got := groupPaths(collectTreeGroups(dirs, super))
|
||||
|
||||
|
||||
Reference in New Issue
Block a user