Make scan paths required operands; add -x/--one-file-system
All checks were successful
check / check (push) Successful in 4s

scan now takes one or more PATH operands (directories or regular
files) via cobra flags instead of the -root flag with its /srv
default; invoking scan with no operand is a usage error and a
nonexistent operand is fatal. Filesystem boundaries are crossed by
default; the new -x/--one-file-system flag (GNU du/rsync convention)
stops the walk at each operand's filesystem, implemented by comparing
lstat device IDs with build-tagged helpers for darwin's int32 Dev.
Verified against a real mounted disk image: default crosses, -x does
not, --one-file-system is identical to -x.
This commit is contained in:
2026-07-23 08:55:17 +07:00
parent 3391207f8d
commit 4b1c3cbf70
7 changed files with 262 additions and 84 deletions

View File

@@ -20,16 +20,17 @@ This README is the complete and authoritative specification.
```sh ```sh
make build make build
./sfdupes scan -root /srv > files.dat ./sfdupes scan /srv > files.dat
./sfdupes report files.dat > dupes.tsv ./sfdupes report files.dat > dupes.tsv
./sfdupes trees files.dat > dupetrees.tsv ./sfdupes trees files.dat > dupetrees.tsv
``` ```
`scan` walks a filesystem tree and emits one record per regular file `scan` walks one or more filesystem trees and emits one record per
(path, size, mtime, head hash, tail hash). `report` ingests that stream regular file (path, size, mtime, head hash, tail hash). `report`
and prints the file-level duplicates report. `trees` ingests the same ingests that stream and prints the file-level duplicates report.
stream and prints the duplicate-tree report. A missing/invalid `trees` ingests the same stream and prints the duplicate-tree report. A
subcommand prints a usage message and exits 2. missing/invalid subcommand — or a `scan` invocation with no `PATH`
operand — prints a usage message and exits 2.
## Rationale ## Rationale
@@ -92,18 +93,26 @@ Three subcommands, all implemented:
directory, and report maximal groups of identical trees. directory, and report maximal groups of identical trees.
``` ```
sfdupes scan [-root /srv] [-workers N] > files.dat sfdupes scan [--workers N] [-x] PATH... > files.dat
sfdupes report [files.dat|-] > dupes.tsv sfdupes report [files.dat|-] > dupes.tsv
sfdupes trees [files.dat|-] > dupetrees.tsv sfdupes trees [files.dat|-] > dupetrees.tsv
``` ```
### `scan` mode ### `scan` mode
`scan` requires one or more `PATH` operands naming the trees to scan.
There is no default path; invoking `scan` with no operand is a usage
error (usage message on stderr, exit 2). An operand may be a directory
or a regular file; an operand that does not exist is a fatal error
(exit 1). Operands are walked in the order given; overlapping operands
(one containing another) emit their common files once per operand, so
callers should pass disjoint paths.
`scan` runs **three sequential passes**, in this order, so that every `scan` runs **three sequential passes**, in this order, so that every
expensive pass has an exact total for meaningful progress and ETA: expensive pass has an exact total for meaningful progress and ETA:
1. **walk** — recursively enumerate the tree under `-root` (default 1. **walk** — recursively enumerate the tree under each `PATH` in
`/srv`), collecting the list of regular-file paths. Total unknown turn, collecting the list of regular-file paths. Total unknown
while running: show a live count, not a percentage. while running: show a live count, not a percentage.
2. **stat**`lstat` every collected path, recording size and mtime. 2. **stat**`lstat` every collected path, recording size and mtime.
3. **hash** — for each file, read the first `min(1024, size)` bytes and 3. **hash** — for each file, read the first `min(1024, size)` bytes and
@@ -113,16 +122,21 @@ expensive pass has an exact total for meaningful progress and ETA:
Rules for the walk: Rules for the walk:
- Only regular files. Skip directories, symlinks (do not follow), - Only regular files. Skip directories, symlinks (do not follow,
sockets, FIFOs, and device nodes. including symlink operands), sockets, FIFOs, and device nodes.
- Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs; - Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs;
walking them would list every file once per snapshot). walking them would list every file once per snapshot).
- Filesystem boundaries are crossed by default. With `-x`
(long form `--one-file-system`, following the GNU `du`/`rsync`
convention), never descend into a directory on a different
filesystem than its `PATH` operand; each operand is bounded by its
own filesystem.
- On any per-path error (permission denied, file vanished between - On any per-path error (permission denied, file vanished between
passes, unreadable): print a one-line warning to stderr, skip the passes, unreadable): print a one-line warning to stderr, skip the
path, and continue. Per-file errors never abort the run; the final path, and continue. Per-file errors never abort the run; the final
summary reports how many were skipped. summary reports how many were skipped.
Concurrency: the stat and hash passes use a worker pool (`-workers`, Concurrency: the stat and hash passes use a worker pool (`--workers`,
default `runtime.NumCPU()`). The main goroutine owns stdout writing and default `runtime.NumCPU()`). The main goroutine owns stdout writing and
progress rendering; progress display must never block the workers. progress rendering; progress display must never block the workers.
@@ -278,9 +292,9 @@ Additional requirements:
### Error handling and exit codes ### Error handling and exit codes
- `0`: success, even if individual files were skipped with warnings. - `0`: success, even if individual files were skipped with warnings.
- `1`: fatal error (e.g., `-root` does not exist, cannot read the scan - `1`: fatal error (e.g., a `PATH` operand does not exist, cannot
input, stdout write failure). read the scan input, stdout write failure).
- `2`: usage error. - `2`: usage error (including `scan` with no `PATH` operand).
## Build ## Build
@@ -325,7 +339,7 @@ All of the following, run in this directory, must pass:
cp "$d/t1/sub/f2" "$d/t2/sub/f2" cp "$d/t1/sub/f2" "$d/t2/sub/f2"
cp "$d/t1/f1" "$d/t3/f1" cp "$d/t1/f1" "$d/t3/f1"
cp "$d/t1/sub/f2" "$d/t3/sub/f2renamed" cp "$d/t1/sub/f2" "$d/t3/sub/f2renamed"
./sfdupes scan -root "$d" > files.dat ./sfdupes scan "$d" > files.dat
./sfdupes report files.dat ./sfdupes report files.dat
./sfdupes trees files.dat ./sfdupes trees files.dat
``` ```
@@ -336,7 +350,7 @@ All of the following, run in this directory, must pass:
`t2/sub/f2`/`t3/sub/f2renamed` form one group; `tiny1`/`tiny2` pair; `t2/sub/f2`/`t3/sub/f2renamed` form one group; `tiny1`/`tiny2` pair;
`empty1`/`empty2` pair; `unique.bin` and `tiny3` appear nowhere; `empty1`/`empty2` pair; `unique.bin` and `tiny3` appear nowhere;
groups ordered by size descending; piping scan directly into report groups ordered by size descending; piping scan directly into report
(`./sfdupes scan -root "$d" | ./sfdupes report`) gives the same (`./sfdupes scan "$d" | ./sfdupes report`) gives the same
rows. rows.
Expected from `trees`: exactly one row — `first` `$d/t1`, `dupe` Expected from `trees`: exactly one row — `first` `$d/t1`, `dupe`

View File

@@ -18,6 +18,10 @@
# Completed Steps # Completed Steps
- `scan` CLI rework (2026-07-23, branch `scan-required-paths`): required
`PATH...` operands via cobra flags replacing the `/srv` `-root`
default; new `-x`/`--one-file-system` flag (GNU convention) to stop
at filesystem boundaries, which are crossed by default
- bring the repo into full policy compliance (2026-07-23, branch - bring the repo into full policy compliance (2026-07-23, branch
`repo-policy-compliance`; checklist below) `repo-policy-compliance`; checklist below)
- `git init` with README-only first commit; code baseline committed on - `git init` with README-only first commit; code baseline committed on

22
main.go
View File

@@ -6,7 +6,7 @@
// //
// Usage: // Usage:
// //
// sfdupes scan [-root /srv] [-workers N] > files.dat // sfdupes scan [--workers N] [-x] PATH... > files.dat
// sfdupes report [files.dat|-] > dupes.tsv // sfdupes report [files.dat|-] > dupes.tsv
// sfdupes trees [files.dat|-] > dupetrees.tsv // sfdupes trees [files.dat|-] > dupetrees.tsv
// //
@@ -16,6 +16,7 @@ package main
import ( import (
"fmt" "fmt"
"os" "os"
"runtime"
"github.com/spf13/cobra" "github.com/spf13/cobra"
) )
@@ -52,16 +53,23 @@ func main() {
root.SetErr(os.Stderr) root.SetErr(os.Stderr)
root.CompletionOptions.DisableDefaultCmd = true root.CompletionOptions.DisableDefaultCmd = true
var (
scanWorkers int
scanOneFS bool
)
scanCmd := &cobra.Command{ scanCmd := &cobra.Command{
Use: "scan [-root /srv] [-workers N]", Use: "scan [--workers N] [-x] PATH...",
Short: "Walk a tree and emit one record per regular file on stdout", Short: "Walk trees and emit one record per regular file on stdout",
// The README specifies single-dash flags (-root, -workers); Args: cobra.MinimumNArgs(1),
// parse them with the stdlib flag package inside runScan.
DisableFlagParsing: true,
Run: func(_ *cobra.Command, args []string) { Run: func(_ *cobra.Command, args []string) {
runScan(args) runScan(args, scanWorkers, scanOneFS)
}, },
} }
scanCmd.Flags().IntVar(&scanWorkers, "workers", runtime.NumCPU(),
"concurrent workers for the stat and hash passes")
scanCmd.Flags().BoolVarP(&scanOneFS, "one-file-system", "x", false,
"do not cross filesystem boundaries")
reportCmd := &cobra.Command{ reportCmd := &cobra.Command{
Use: "report [files.dat|-]", Use: "report [files.dat|-]",

155
scan.go
View File

@@ -5,12 +5,11 @@ import (
"crypto/sha256" "crypto/sha256"
"encoding/hex" "encoding/hex"
"errors" "errors"
"flag"
"fmt" "fmt"
"io/fs" "io/fs"
"os" "os"
"path/filepath" "path/filepath"
"runtime" "syscall"
) )
// chunk is the number of bytes hashed from each end of a file. // chunk is the number of bytes hashed from each end of a file.
@@ -31,63 +30,66 @@ type fileRec struct {
mtime int64 mtime int64
} }
// runScan implements the scan subcommand: three sequential passes (walk, // runScan implements the scan subcommand: three sequential passes
// stat, hash) over the tree under -root, emitting one NUL-terminated // (walk, stat, hash) over the trees named by the PATH operands,
// record per regular file on stdout. // emitting one NUL-terminated record per regular file on stdout. Flag
func runScan(args []string) { // parsing and the at-least-one-operand check are done by cobra.
fl := flag.NewFlagSet("scan", flag.ExitOnError) func runScan(roots []string, workers int, oneFS bool) {
root := fl.String("root", "/srv", "directory tree to scan") if workers < 1 {
workers := fl.Int("workers", runtime.NumCPU(), workers = 1
"concurrent workers for the stat and hash passes") }
// The flag set uses ExitOnError, so Parse cannot return a non-nil
// error; the check keeps the error handled explicitly. // A nonexistent operand is a fatal error before any scanning.
err := fl.Parse(args) for _, root := range roots {
_, err := os.Lstat(root)
if err != nil { if err != nil {
os.Exit(exitUsage) fatalf("%v", err)
}
} }
if fl.NArg() > 0 { paths, walkErrs := walkPass(roots, oneFS)
fmt.Fprintf(os.Stderr, "scan: unexpected argument %q\n", fl.Arg(0)) recs, statErrs := statPass(paths, workers)
os.Exit(exitUsage) emitted, hashErrs := hashPass(recs, workers)
}
if *workers < 1 {
*workers = 1
}
st, err := os.Stat(*root)
if err != nil {
fatalf("root %s: %v", *root, err)
}
if !st.IsDir() {
fatalf("root %s: not a directory", *root)
}
paths, walkErrs := walkPass(*root)
recs, statErrs := statPass(paths, *workers)
emitted, hashErrs := hashPass(recs, *workers)
skipped := walkErrs + statErrs + hashErrs skipped := walkErrs + statErrs + hashErrs
fmt.Fprintf(os.Stderr, "scan: %d files emitted, %d skipped\n", fmt.Fprintf(os.Stderr, "scan: %d files emitted, %d skipped\n",
emitted, skipped) emitted, skipped)
} }
// walkPass enumerates every regular file under root. It never follows // treeWalker carries the walk-pass state shared by all PATH operands.
// symlinks, never descends into directories named .zfs, and warns and type treeWalker struct {
// continues on any per-path error. prog *progress
func walkPass(root string) ([]string, int) { oneFS bool
var (
paths []string paths []string
errs int errs int
) }
// walkPass enumerates every regular file under each root operand in
// order. It never follows symlinks, never descends into directories
// named .zfs, and warns and continues on any per-path error. With
// oneFS set it never descends into a directory on a different
// filesystem than its root operand.
func walkPass(roots []string, oneFS bool) ([]string, int) {
w := &treeWalker{prog: newProgress("walk", -1), oneFS: oneFS}
for _, root := range roots {
w.walkRoot(root)
}
w.prog.finish()
return w.paths, w.errs
}
// walkRoot walks a single PATH operand, appending regular-file paths.
func (w *treeWalker) walkRoot(root string) {
rootDev, rootDevOK := deviceOf(root)
prog := newProgress("walk", -1)
walkErr := filepath.WalkDir(root, func(p string, d fs.DirEntry, err error) error { walkErr := filepath.WalkDir(root, func(p string, d fs.DirEntry, err error) error {
if err != nil { if err != nil {
errs++ w.errs++
prog.warnf("walk %s: %v", p, err) w.prog.warnf("walk %s: %v", p, err)
if d != nil && d.IsDir() { if d != nil && d.IsDir() {
return filepath.SkipDir return filepath.SkipDir
@@ -97,13 +99,7 @@ func walkPass(root string) ([]string, int) {
} }
if d.IsDir() { if d.IsDir() {
// ZFS snapshot pseudo-dirs would list every file return w.dirAction(p, d, rootDev, rootDevOK)
// once per snapshot; never descend.
if d.Name() == ".zfs" {
return filepath.SkipDir
}
return nil
} }
// Regular files only: skip symlinks, sockets, FIFOs, and // Regular files only: skip symlinks, sockets, FIFOs, and
// device nodes. // device nodes.
@@ -111,20 +107,65 @@ func walkPass(root string) ([]string, int) {
return nil return nil
} }
paths = append(paths, p) w.paths = append(w.paths, p)
prog.increment() w.prog.increment()
return nil return nil
}) })
prog.finish()
if walkErr != nil { if walkErr != nil {
fatalf("walk %s: %v", root, walkErr) fatalf("walk %s: %v", root, walkErr)
} }
}
return paths, errs // dirAction decides whether the walk descends into directory p.
func (w *treeWalker) dirAction(p string, d fs.DirEntry, rootDev uint64, rootDevOK bool) error {
// ZFS snapshot pseudo-dirs would list every file once per
// snapshot; never descend.
if d.Name() == ".zfs" {
return filepath.SkipDir
}
if !w.oneFS || !rootDevOK {
return nil
}
info, err := d.Info()
if err != nil {
w.errs++
w.prog.warnf("walk %s: %v", p, err)
return filepath.SkipDir
}
if dev, ok := deviceOfInfo(info); ok && dev != rootDev {
return filepath.SkipDir
}
return nil
}
// deviceOf returns the filesystem device ID of path without following
// symlinks.
func deviceOf(path string) (uint64, bool) {
fi, err := os.Lstat(path)
if err != nil {
return 0, false
}
return deviceOfInfo(fi)
}
// deviceOfInfo extracts the filesystem device ID from a FileInfo, when
// the platform exposes one.
func deviceOfInfo(fi fs.FileInfo) (uint64, bool) {
st, ok := fi.Sys().(*syscall.Stat_t)
if !ok {
return 0, false
}
return statDev(st), true
} }
// statPass lstats every collected path in a worker pool, recording size // statPass lstats every collected path in a worker pool, recording size

14
scan_dev_darwin.go Normal file
View File

@@ -0,0 +1,14 @@
//go:build darwin
package main
import "syscall"
// statDev returns the filesystem device ID of st as uint64. Device IDs
// are opaque; the conversion only needs to be consistent so equality
// comparisons work, not numerically meaningful.
//
//nolint:gosec // int32 device IDs convert consistently; only equality matters
func statDev(st *syscall.Stat_t) uint64 {
return uint64(st.Dev)
}

10
scan_dev_other.go Normal file
View File

@@ -0,0 +1,10 @@
//go:build !darwin
package main
import "syscall"
// statDev returns the filesystem device ID of st.
func statDev(st *syscall.Stat_t) uint64 {
return st.Dev
}

View File

@@ -129,7 +129,7 @@ func TestWalkPass(t *testing.T) {
t.Fatal(err) t.Fatal(err)
} }
paths, errs := walkPass(dir) paths, errs := walkPass([]string{dir}, false)
if errs != 0 { if errs != 0 {
t.Fatalf("errs = %d, want 0", errs) t.Fatalf("errs = %d, want 0", errs)
} }
@@ -141,6 +141,93 @@ func TestWalkPass(t *testing.T) {
} }
} }
func TestWalkPassMultipleRoots(t *testing.T) {
t.Parallel()
rootA := t.TempDir()
rootB := t.TempDir()
want := []string{
writeFile(t, rootA, "a1", []byte("1")),
writeFile(t, rootA, "sub/a2", []byte("2")),
writeFile(t, rootB, "b1", []byte("3")),
}
paths, errs := walkPass([]string{rootA, rootB}, false)
if errs != 0 {
t.Fatalf("errs = %d, want 0", errs)
}
// Operands are walked in the order given.
if !slices.Equal(paths, want) {
t.Fatalf("paths = %q, want %q", paths, want)
}
}
func TestWalkPassFileAndSymlinkOperands(t *testing.T) {
t.Parallel()
dir := t.TempDir()
f := writeFile(t, dir, "plain", []byte("data"))
link := filepath.Join(dir, "link")
err := os.Symlink(f, link)
if err != nil {
t.Fatal(err)
}
// A regular-file operand is emitted as itself.
paths, errs := walkPass([]string{f}, false)
if errs != 0 || !slices.Equal(paths, []string{f}) {
t.Fatalf("file operand: paths = %q, errs = %d", paths, errs)
}
// A symlink operand is not followed and yields nothing.
paths, errs = walkPass([]string{link}, false)
if errs != 0 || len(paths) != 0 {
t.Fatalf("symlink operand: paths = %q, errs = %d", paths, errs)
}
}
func TestWalkPassOneFilesystemSameFS(t *testing.T) {
t.Parallel()
// Everything in one filesystem: -x must not skip anything.
dir := t.TempDir()
want := []string{
writeFile(t, dir, "a", []byte("a")),
writeFile(t, dir, "sub/deep/b", []byte("b")),
}
paths, errs := walkPass([]string{dir}, true)
if errs != 0 {
t.Fatalf("errs = %d, want 0", errs)
}
slices.Sort(paths)
if !slices.Equal(paths, want) {
t.Fatalf("paths = %q, want %q", paths, want)
}
}
func TestDeviceOf(t *testing.T) {
t.Parallel()
dir := t.TempDir()
dev1, ok1 := deviceOf(dir)
dev2, ok2 := deviceOf(dir)
if !ok1 || !ok2 || dev1 != dev2 {
t.Fatalf("deviceOf unstable: %d/%v vs %d/%v", dev1, ok1, dev2, ok2)
}
if _, ok := deviceOf(filepath.Join(dir, "missing")); ok {
t.Fatal("deviceOf reported ok for a missing path")
}
}
func TestStatPass(t *testing.T) { func TestStatPass(t *testing.T) {
t.Parallel() t.Parallel()
@@ -234,7 +321,7 @@ func buildSmokeTree(t *testing.T) string {
func scanToRecords(t *testing.T, dir string) []scanRec { func scanToRecords(t *testing.T, dir string) []scanRec {
t.Helper() t.Helper()
paths, walkErrs := walkPass(dir) paths, walkErrs := walkPass([]string{dir}, false)
if walkErrs != 0 { if walkErrs != 0 {
t.Fatalf("walk errors: %d", walkErrs) t.Fatalf("walk errors: %d", walkErrs)
} }