Compare commits
2
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
efc2b62a77 | ||
|
|
01ff3bb5f0 |
@@ -113,8 +113,9 @@ Goals, in order:
|
||||
`sfdupes`.
|
||||
- Dependencies: standard library, `github.com/spf13/cobra` for the
|
||||
CLI, **one progress-bar library**
|
||||
(`github.com/schollz/progressbar/v3`), and **one SQLite driver**
|
||||
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled).
|
||||
(`github.com/schollz/progressbar/v3`), **one SQLite driver**
|
||||
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled), and
|
||||
`golang.org/x/sys` for `flock(2)` (the scan lock, see "Database").
|
||||
`github.com/spf13/viper` is permitted if configuration-file support
|
||||
is ever needed, but is not currently used. No other third-party
|
||||
deps.
|
||||
@@ -152,6 +153,20 @@ All three subcommands operate on a single SQLite database file:
|
||||
use. `report` and `trees` require an existing database; a missing
|
||||
database file is a fatal error (exit 1) telling the user to run
|
||||
`scan` first.
|
||||
- Only one `scan` runs against a database at a time. For its whole
|
||||
run, `scan` holds an exclusive `flock(2)` lock on a lock file
|
||||
beside the database, named by appending `.lock` to the database
|
||||
path (`/var/lib/sfdupes/db.sqlite.lock` by default), taken before
|
||||
it walks the filesystem or opens the database. A second `scan`
|
||||
against the same database does not wait: it fails at once with a
|
||||
one-line error naming the lock file and exits 1, without walking
|
||||
anything or opening the database, and the running scan carries on.
|
||||
The lock file is created on first use, open to its owner only, and
|
||||
left in place: a leftover file blocks nothing, because the lock
|
||||
ends with the process holding it however it ends, a fatal error or
|
||||
an interrupt included, and deleting the file while a scan runs
|
||||
would let a second scan start. `report` and `trees` never take the
|
||||
lock, so they run during a scan.
|
||||
- While `scan` runs, the database is in WAL journal mode with a busy
|
||||
timeout, so running a report while a cron `scan` is in progress is
|
||||
safe. The filesystem is authoritative; the database is an
|
||||
@@ -567,12 +582,25 @@ Additional requirements:
|
||||
### Error handling and exit codes
|
||||
|
||||
- `0`: success, even if individual files were skipped with warnings.
|
||||
- `1`: fatal error (e.g., a `PATH` operand does not exist, the
|
||||
database cannot be created/opened/read/written, a missing database
|
||||
for `report`/`trees`, stdout write failure).
|
||||
- `1`: fatal error (e.g., a `PATH` operand does not exist, another
|
||||
`scan` is already running against the same database, the database
|
||||
cannot be created/opened/read/written, a missing database for
|
||||
`report`/`trees`, stdout write failure).
|
||||
- `2`: usage error (including `scan` with no `PATH` operand and
|
||||
`report`/`trees` with any positional argument).
|
||||
|
||||
A stdout write failure, such as a full disk, is reported in one line on
|
||||
stderr and exits 1. Two cases never reach sfdupes as a failed write:
|
||||
|
||||
- When the reader of a stdout pipe exits early, as in
|
||||
`sfdupes report | head`, the next write ends sfdupes with `SIGPIPE`,
|
||||
quietly and without a summary, the way `cat` or `sort` end. The
|
||||
shell reports the signal (status 141 in most shells), not exit 1.
|
||||
- When stdout is closed outright (`sfdupes report >&-`), the Go
|
||||
runtime opens `/dev/null` in its place before sfdupes starts, so
|
||||
the output is discarded and the run succeeds, as with
|
||||
`> /dev/null`.
|
||||
|
||||
## Entrypoints
|
||||
|
||||
This repository adheres to the
|
||||
@@ -710,7 +738,7 @@ All of the following, run in this directory, must pass:
|
||||
|
||||
```sh
|
||||
d=$(mktemp -d)
|
||||
export SFDUPES_DATABASE="$d/db.sqlite"
|
||||
export SFDUPES_DATABASE="$(mktemp -d)/db.sqlite"
|
||||
mkdir -p "$d/a" "$d/b"
|
||||
head -c 2000 /dev/urandom > "$d/a/one.bin"
|
||||
cp "$d/a/one.bin" "$d/b/copy.bin"
|
||||
@@ -738,9 +766,9 @@ All of the following, run in this directory, must pass:
|
||||
./sfdupes report
|
||||
```
|
||||
|
||||
(The scan database lives inside `$d` here purely for test hygiene;
|
||||
scanning `$d` therefore also records the SQLite file itself, which
|
||||
is harmless.)
|
||||
(The database lives in a temp directory of its own: inside `$d`,
|
||||
the scan would record it, and its empty lock file would join the
|
||||
`empty1`/`empty2` group.)
|
||||
|
||||
Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin`
|
||||
form one group (two dupe rows, `first` is the lexicographically
|
||||
|
||||
@@ -29,6 +29,14 @@
|
||||
|
||||
# Completed Steps
|
||||
|
||||
- `scan` holds a lock on a lock file beside the database for its whole run,
|
||||
so a second `scan` fails at once with exit 1 (2026-10-03,
|
||||
https://git.eeqj.de/sneak/sfdupes/issues/53)
|
||||
|
||||
- test stdout write failures in `report` and `trees`; README states that
|
||||
`| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null`
|
||||
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/30)
|
||||
|
||||
- `report` and `trees` open the database read-only, and `scan` leaves it
|
||||
out of WAL mode, so reading needs only read access (2026-10-03, closes
|
||||
https://git.eeqj.de/sneak/sfdupes/issues/8)
|
||||
|
||||
@@ -11,6 +11,7 @@ import (
|
||||
"slices"
|
||||
"strconv"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
// The pure-Go SQLite driver, registered as "sqlite"; keeps cgo
|
||||
// disabled.
|
||||
_ "modernc.org/sqlite"
|
||||
@@ -32,6 +33,11 @@ const schemaVersion = 1
|
||||
// scan.
|
||||
const dbDirPerm = 0o755
|
||||
|
||||
// lockFilePerm is the mode for the scan lock file. Anyone who can open
|
||||
// the file can hold the lock and keep every scan from running, so it
|
||||
// is open to its owner only.
|
||||
const lockFilePerm = 0o600
|
||||
|
||||
// createTableSQL is the schema applied to a fresh database. Paths are
|
||||
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
|
||||
const createTableSQL = `
|
||||
@@ -64,6 +70,10 @@ var errNoDatabase = errors.New(
|
||||
// does not understand.
|
||||
var errSchemaVersion = errors.New("unsupported database schema version")
|
||||
|
||||
// errScanRunning reports that another scan holds the lock on the
|
||||
// database.
|
||||
var errScanRunning = errors.New("another scan is running")
|
||||
|
||||
// databasePath resolves the database location: SFDUPES_DATABASE when
|
||||
// set and non-empty, the compiled-in default otherwise.
|
||||
func databasePath() string {
|
||||
@@ -104,6 +114,43 @@ func openDB(path, params string) (*sql.DB, error) {
|
||||
return db, nil
|
||||
}
|
||||
|
||||
// lockScanDatabase takes the lock that keeps a second scan off the
|
||||
// database at path: an exclusive flock(2) on the file beside it named
|
||||
// path with ".lock" appended, created along with the database's parent
|
||||
// directory if missing. A lock held by another scan fails at once
|
||||
// instead of waiting. The lock lasts until the returned file is closed
|
||||
// or the process ends. The file is never deleted: a scan that deleted
|
||||
// it would let the next scan lock a new file while another still holds
|
||||
// the old one.
|
||||
func lockScanDatabase(path string) (*os.File, error) {
|
||||
err := os.MkdirAll(filepath.Dir(path), dbDirPerm)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("create database directory: %w", err)
|
||||
}
|
||||
|
||||
lockPath := path + ".lock"
|
||||
|
||||
//nolint:gosec // the operator chooses the database path
|
||||
f, err := os.OpenFile(lockPath, os.O_RDWR|os.O_CREATE, lockFilePerm)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
err = unix.Flock(int(f.Fd()), unix.LOCK_EX|unix.LOCK_NB)
|
||||
if err != nil {
|
||||
_ = f.Close()
|
||||
|
||||
if errors.Is(err, unix.EWOULDBLOCK) {
|
||||
return nil, fmt.Errorf("%w (lock held on %s)",
|
||||
errScanRunning, lockPath)
|
||||
}
|
||||
|
||||
return nil, fmt.Errorf("lock %s: %w", lockPath, err)
|
||||
}
|
||||
|
||||
return f, nil
|
||||
}
|
||||
|
||||
// openScanDatabase opens the database for the scan subcommand, creating
|
||||
// the file, its parent directory, and the schema as needed.
|
||||
func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
|
||||
|
||||
@@ -5,6 +5,7 @@ go 1.25.7
|
||||
require (
|
||||
github.com/schollz/progressbar/v3 v3.19.1
|
||||
github.com/spf13/cobra v1.10.2
|
||||
golang.org/x/sys v0.46.0
|
||||
modernc.org/sqlite v1.54.0
|
||||
)
|
||||
|
||||
@@ -18,7 +19,6 @@ require (
|
||||
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
|
||||
github.com/rivo/uniseg v0.4.7 // indirect
|
||||
github.com/spf13/pflag v1.0.9 // indirect
|
||||
golang.org/x/sys v0.46.0 // indirect
|
||||
golang.org/x/term v0.44.0 // indirect
|
||||
modernc.org/libc v1.74.1 // indirect
|
||||
modernc.org/mathutil v1.7.1 // indirect
|
||||
|
||||
@@ -57,22 +57,27 @@ var errNoSubcommand = errors.New("no subcommand")
|
||||
var Version = "dev"
|
||||
|
||||
func main() {
|
||||
os.Exit(run(os.Args[1:], os.Stderr))
|
||||
// Once the reader of a stdout pipe has gone, as in "sfdupes report |
|
||||
// head", the Go runtime ends the process with SIGPIPE on the next
|
||||
// write instead of returning an error (README "Error handling").
|
||||
// Registering for SIGPIPE with os/signal would change that.
|
||||
os.Exit(run(os.Args[1:], os.Stdout, os.Stderr))
|
||||
}
|
||||
|
||||
// run executes args against the command tree and returns the process
|
||||
// exit code. It is the program's single exit point: the subcommands
|
||||
// return their errors instead of exiting, so every deferred cleanup —
|
||||
// above all closing the database, which checkpoints the SQLite WAL —
|
||||
// runs before the process ends.
|
||||
func run(args []string, stderr io.Writer) int {
|
||||
// runs before the process ends. The report and trees subcommands write
|
||||
// their data to stdout.
|
||||
func run(args []string, stdout, stderr io.Writer) int {
|
||||
// A nil slice makes cobra fall back to os.Args, which would let a
|
||||
// test binary's own flags reach the command tree.
|
||||
if args == nil {
|
||||
args = []string{}
|
||||
}
|
||||
|
||||
root := newRootCommand(stderr)
|
||||
root := newRootCommand(stdout, stderr)
|
||||
root.SetArgs(args)
|
||||
|
||||
err := root.Execute()
|
||||
@@ -98,7 +103,7 @@ func run(args []string, stderr io.Writer) int {
|
||||
// newRootCommand builds the command tree. Everything on stdout is
|
||||
// machine-readable data; all human-facing output (help, usage, errors)
|
||||
// goes to stderr.
|
||||
func newRootCommand(stderr io.Writer) *cobra.Command {
|
||||
func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
|
||||
root := &cobra.Command{
|
||||
Use: "sfdupes",
|
||||
Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
|
||||
@@ -140,7 +145,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
|
||||
Short: "Read the scan database and print the file-level duplicates report",
|
||||
Args: cobra.NoArgs,
|
||||
RunE: runE(func(ctx context.Context, _ []string) error {
|
||||
return runReport(ctx)
|
||||
return runReport(ctx, stdout)
|
||||
}),
|
||||
}
|
||||
|
||||
@@ -149,7 +154,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
|
||||
Short: "Read the scan database and print the duplicate-tree report",
|
||||
Args: cobra.NoArgs,
|
||||
RunE: runE(func(ctx context.Context, _ []string) error {
|
||||
return runTrees(ctx)
|
||||
return runTrees(ctx, stdout)
|
||||
}),
|
||||
}
|
||||
|
||||
|
||||
+178
-86
@@ -82,50 +82,6 @@ func makeReadOnly(t *testing.T, path string) {
|
||||
})
|
||||
}
|
||||
|
||||
// captureStdout redirects os.Stdout to a file for the rest of the test
|
||||
// and returns a function reading back everything written to it. Only
|
||||
// machine-readable data belongs on stdout (README design goal 4), so
|
||||
// the tests assert on it directly.
|
||||
func captureStdout(t *testing.T) func() string {
|
||||
t.Helper()
|
||||
|
||||
f, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
saved := os.Stdout
|
||||
os.Stdout = f
|
||||
|
||||
t.Cleanup(func() {
|
||||
os.Stdout = saved
|
||||
|
||||
_ = f.Close()
|
||||
})
|
||||
|
||||
return func() string {
|
||||
// Read what has been written without disturbing the write
|
||||
// offset, so the capture can be inspected more than once.
|
||||
size, err := f.Seek(0, io.SeekCurrent)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
if size == 0 {
|
||||
return ""
|
||||
}
|
||||
|
||||
b := make([]byte, size)
|
||||
|
||||
_, err = f.ReadAt(b, 0)
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
return string(b)
|
||||
}
|
||||
}
|
||||
|
||||
// brokenDatabase writes a database that opens cleanly and passes the
|
||||
// schema-version check but has no files table, so the first query
|
||||
// fails with the database already open: a fatal error on a path that
|
||||
@@ -198,17 +154,15 @@ func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
|
||||
args = append(args, t.TempDir())
|
||||
}
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run(args, &stderr)
|
||||
code := run(args, &stdout, &stderr)
|
||||
if code != exitFatal {
|
||||
t.Errorf("run(%v) = %d, want %d", args, code, exitFatal)
|
||||
}
|
||||
|
||||
assertNoSidecars(t, path)
|
||||
assertFatalOutput(t, stderr.String(), stdout())
|
||||
assertFatalOutput(t, stderr.String(), stdout.String())
|
||||
|
||||
// Proof that the failure happened after the open: only a
|
||||
// query against the opened database can report this.
|
||||
@@ -226,18 +180,16 @@ func TestRunMissingOperandIsFatalNotUsage(t *testing.T) {
|
||||
// must not dump the usage text.
|
||||
t.Setenv(databaseEnv, testDBPath(t))
|
||||
|
||||
var stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
missing := filepath.Join(t.TempDir(), "nope")
|
||||
|
||||
code := run([]string{cmdScan, missing}, &stderr)
|
||||
code := run([]string{cmdScan, missing}, &stdout, &stderr)
|
||||
if code != exitFatal {
|
||||
t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal)
|
||||
}
|
||||
|
||||
assertFatalOutput(t, stderr.String(), stdout())
|
||||
assertFatalOutput(t, stderr.String(), stdout.String())
|
||||
}
|
||||
|
||||
// assertFatalOutput checks that a fatal error was reported the way
|
||||
@@ -281,11 +233,9 @@ func TestRunUsageErrors(t *testing.T) {
|
||||
// path that does not exist.
|
||||
t.Setenv(databaseEnv, testDBPath(t))
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run(tc.args, &stderr)
|
||||
code := run(tc.args, &stdout, &stderr)
|
||||
if code != exitUsage {
|
||||
t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage)
|
||||
}
|
||||
@@ -294,7 +244,7 @@ func TestRunUsageErrors(t *testing.T) {
|
||||
t.Errorf("stderr = %q, want %q", stderr.String(), tc.want)
|
||||
}
|
||||
|
||||
if got := stdout(); got != "" {
|
||||
if got := stdout.String(); got != "" {
|
||||
t.Errorf("stdout = %q, want nothing (data only)", got)
|
||||
}
|
||||
})
|
||||
@@ -303,9 +253,9 @@ func TestRunUsageErrors(t *testing.T) {
|
||||
|
||||
// TestRunHelpAndVersionSucceed checks that the two informational flags
|
||||
// exit 0 and keep their human-facing output on stderr.
|
||||
//
|
||||
//nolint:paralleltest // captureStdout replaces the process-wide os.Stdout
|
||||
func TestRunHelpAndVersionSucceed(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
assertHumanOutput(t, "--help")
|
||||
assertHumanOutput(t, "--version")
|
||||
}
|
||||
@@ -316,11 +266,9 @@ func TestRunHelpAndVersionSucceed(t *testing.T) {
|
||||
func assertHumanOutput(t *testing.T, arg string) {
|
||||
t.Helper()
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run([]string{arg}, &stderr)
|
||||
code := run([]string{arg}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
|
||||
}
|
||||
@@ -329,7 +277,7 @@ func assertHumanOutput(t *testing.T, arg string) {
|
||||
t.Errorf("run(%s) wrote nothing to stderr", arg)
|
||||
}
|
||||
|
||||
if got := stdout(); got != "" {
|
||||
if got := stdout.String(); got != "" {
|
||||
t.Errorf("stdout = %q, want nothing (data only)", got)
|
||||
}
|
||||
}
|
||||
@@ -358,17 +306,15 @@ func scanFixture(t *testing.T) []string {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run([]string{cmdScan, dir}, &stderr)
|
||||
code := run([]string{cmdScan, dir}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
|
||||
code, exitOK, stderr.String())
|
||||
}
|
||||
|
||||
if got := stdout(); got != "" {
|
||||
if got := stdout.String(); got != "" {
|
||||
t.Errorf("scan stdout = %q, want nothing (data only)", got)
|
||||
}
|
||||
|
||||
@@ -389,18 +335,16 @@ func TestRunReportSucceeds(t *testing.T) {
|
||||
|
||||
dupes := scanFixture(t)
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run([]string{cmdReport}, &stderr)
|
||||
code := run([]string{cmdReport}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(report) = %d, want %d; stderr: %s",
|
||||
code, exitOK, stderr.String())
|
||||
}
|
||||
|
||||
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
|
||||
if got := stdout(); got != want {
|
||||
if got := stdout.String(); got != want {
|
||||
t.Errorf("stdout = %q, want %q", got, want)
|
||||
}
|
||||
|
||||
@@ -413,11 +357,9 @@ func TestRunTreesSucceeds(t *testing.T) {
|
||||
|
||||
dupes := scanFixture(t)
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run([]string{cmdTrees}, &stderr)
|
||||
code := run([]string{cmdTrees}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
|
||||
code, exitOK, stderr.String())
|
||||
@@ -427,7 +369,7 @@ func TestRunTreesSucceeds(t *testing.T) {
|
||||
// trees of each other.
|
||||
want := "first\tdupe\tfiles\tsize\n" +
|
||||
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
|
||||
if got := stdout(); got != want {
|
||||
if got := stdout.String(); got != want {
|
||||
t.Errorf("stdout = %q, want %q", got, want)
|
||||
}
|
||||
|
||||
@@ -454,11 +396,9 @@ func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
|
||||
}
|
||||
|
||||
for name, want := range cases {
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run([]string{name}, &stderr)
|
||||
code := run([]string{name}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Errorf("run(%s) = %d, want %d; stderr: %s",
|
||||
name, code, exitOK, stderr.String())
|
||||
@@ -466,8 +406,160 @@ func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
|
||||
continue
|
||||
}
|
||||
|
||||
if got := stdout(); got != want {
|
||||
if got := stdout.String(); got != want {
|
||||
t.Errorf("%s stdout = %q, want %q", name, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// holdScanLock takes the lock on the database at path, as a running
|
||||
// scan does, and holds it until the test ends. It fails the test when
|
||||
// the lock is already held.
|
||||
func holdScanLock(t *testing.T, path string) {
|
||||
t.Helper()
|
||||
|
||||
lock, err := lockScanDatabase(path)
|
||||
if err != nil {
|
||||
t.Fatalf("lock %s: %v", path, err)
|
||||
}
|
||||
|
||||
t.Cleanup(func() { _ = lock.Close() })
|
||||
}
|
||||
|
||||
func TestRunSecondScanFails(t *testing.T) {
|
||||
// README §Database: while one scan holds the lock, a second scan
|
||||
// fails at once, naming the lock file, without creating the
|
||||
// database.
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
holdScanLock(t, path)
|
||||
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
code := run([]string{cmdScan, t.TempDir()}, &stdout, &stderr)
|
||||
if code != exitFatal {
|
||||
t.Errorf("run(scan) = %d, want %d", code, exitFatal)
|
||||
}
|
||||
|
||||
want := "sfdupes: another scan is running (lock held on " +
|
||||
path + ".lock)\n"
|
||||
if got := stderr.String(); got != want {
|
||||
t.Errorf("stderr = %q, want %q", got, want)
|
||||
}
|
||||
|
||||
if got := stdout.String(); got != "" {
|
||||
t.Errorf("stdout = %q, want nothing (data only)", got)
|
||||
}
|
||||
|
||||
_, err := os.Stat(path)
|
||||
if !errors.Is(err, fs.ErrNotExist) {
|
||||
t.Errorf("stat %s = %v, want the database not created", path, err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunScanReleasesLock(t *testing.T) {
|
||||
// README §Database: a scan releases the lock however it ends.
|
||||
t.Run("success", func(t *testing.T) {
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
scanFixture(t)
|
||||
holdScanLock(t, path)
|
||||
})
|
||||
|
||||
t.Run("fatal error", func(t *testing.T) {
|
||||
path := brokenDatabase(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
code := run([]string{cmdScan, t.TempDir()}, io.Discard, io.Discard)
|
||||
if code != exitFatal {
|
||||
t.Fatalf("run(scan) = %d, want %d", code, exitFatal)
|
||||
}
|
||||
|
||||
holdScanLock(t, path)
|
||||
})
|
||||
}
|
||||
|
||||
func TestRunReportsDuringScan(t *testing.T) {
|
||||
// README §Database: report and trees never take the lock, so they
|
||||
// run while a scan holds it.
|
||||
path := testDBPath(t)
|
||||
t.Setenv(databaseEnv, path)
|
||||
|
||||
scanFixture(t)
|
||||
holdScanLock(t, path)
|
||||
|
||||
for _, name := range []string{cmdReport, cmdTrees} {
|
||||
var stderr bytes.Buffer
|
||||
|
||||
code := run([]string{name}, io.Discard, &stderr)
|
||||
if code != exitOK {
|
||||
t.Errorf("run(%s) = %d, want %d; stderr: %s",
|
||||
name, code, exitOK, stderr.String())
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestRunStdoutClosedIsFatal(t *testing.T) {
|
||||
// README §Error handling: a stdout write failure exits 1, reported
|
||||
// in one line on stderr.
|
||||
for _, name := range []string{cmdReport, cmdTrees} {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
t.Setenv(databaseEnv, testDBPath(t))
|
||||
|
||||
scanFixture(t)
|
||||
|
||||
stdout, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
err = stdout.Close()
|
||||
if err != nil {
|
||||
t.Fatal(err)
|
||||
}
|
||||
|
||||
var stderr bytes.Buffer
|
||||
|
||||
code := run([]string{name}, stdout, &stderr)
|
||||
if code != exitFatal {
|
||||
t.Errorf("run(%s) = %d, want %d", name, code, exitFatal)
|
||||
}
|
||||
|
||||
got := stderr.String()
|
||||
if !strings.HasPrefix(got, "sfdupes: write stdout: ") ||
|
||||
!strings.Contains(got, os.ErrClosed.Error()) ||
|
||||
strings.Count(got, "\n") != 1 {
|
||||
t.Errorf("stderr = %q, want one line reporting the "+
|
||||
"failed stdout write", got)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// errWriteFailed is the error failingWriter returns.
|
||||
var errWriteFailed = errors.New("write failed")
|
||||
|
||||
// failingWriter is a stdout that fails every write.
|
||||
type failingWriter struct{}
|
||||
|
||||
func (failingWriter) Write([]byte) (int, error) { return 0, errWriteFailed }
|
||||
|
||||
func TestStdoutWriteErrorPropagates(t *testing.T) {
|
||||
t.Setenv(databaseEnv, testDBPath(t))
|
||||
|
||||
scanFixture(t)
|
||||
|
||||
cases := map[string]func(context.Context, io.Writer) error{
|
||||
cmdReport: runReport,
|
||||
cmdTrees: runTrees,
|
||||
}
|
||||
|
||||
for name, fn := range cases {
|
||||
err := fn(t.Context(), failingWriter{})
|
||||
if !errors.Is(err, errWriteFailed) {
|
||||
t.Errorf("%s: error = %v, want %v", name, err, errWriteFailed)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
@@ -4,6 +4,7 @@ import (
|
||||
"bufio"
|
||||
"context"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"slices"
|
||||
"strings"
|
||||
@@ -65,7 +66,7 @@ type dupeGroup struct {
|
||||
// from the database and prints the file-level duplicates report as TSV
|
||||
// on stdout. It never touches the scanned filesystem; its only I/O is
|
||||
// the database, stdout, and stderr.
|
||||
func runReport(ctx context.Context) error {
|
||||
func runReport(ctx context.Context, stdout io.Writer) error {
|
||||
recs, err := loadRecords(ctx)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -73,7 +74,7 @@ func runReport(ctx context.Context) error {
|
||||
|
||||
dupes := collectDupeGroups(recs)
|
||||
|
||||
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
|
||||
out := bufio.NewWriterSize(stdout, ioBufSize)
|
||||
|
||||
_, err = fmt.Fprintln(out, "first\tdupe\tsize")
|
||||
if err != nil {
|
||||
|
||||
+3
-5
@@ -50,11 +50,9 @@ func seedDatabase(t *testing.T, recs []scanRec) string {
|
||||
func TestRunReportEscapesPaths(t *testing.T) {
|
||||
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run([]string{cmdReport}, &stderr)
|
||||
code := run([]string{cmdReport}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(report) = %d, want %d; stderr: %s",
|
||||
code, exitOK, stderr.String())
|
||||
@@ -62,7 +60,7 @@ func TestRunReportEscapesPaths(t *testing.T) {
|
||||
|
||||
want := "first\tdupe\tsize\n" +
|
||||
`/d/\tone\ntwo\rthree\\four/f` + "\t/d/A/f\t5\n"
|
||||
if got := stdout(); got != want {
|
||||
if got := stdout.String(); got != want {
|
||||
t.Errorf("stdout = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -85,11 +85,13 @@ type fileMeta struct {
|
||||
// least one other file shares are ever hashed: a size-unique file
|
||||
// cannot be a duplicate. A file of headTailMin or more gets its content
|
||||
// hash only when its size, head, and tail match another file's. Flag
|
||||
// parsing and the at-least-one-operand check are done by cobra. Errors
|
||||
// are returned rather than exiting, so that the deferred close — which
|
||||
// takes the database out of WAL mode — always runs. Cancelling ctx
|
||||
// unwinds the worker pools and aborts the scan with the context's
|
||||
// error.
|
||||
// parsing and the at-least-one-operand check are done by cobra. The
|
||||
// scan holds the lock on the database for its whole run, so a second
|
||||
// scan fails before it walks the filesystem or opens the database.
|
||||
// Errors are returned rather than exiting, so that the deferred close —
|
||||
// which takes the database out of WAL mode — always runs, and the lock
|
||||
// is released after it. Cancelling ctx unwinds the worker pools and
|
||||
// aborts the scan with the context's error.
|
||||
func runScan(ctx context.Context, roots []string, workers int,
|
||||
oneFS bool,
|
||||
) error {
|
||||
@@ -104,6 +106,13 @@ func runScan(ctx context.Context, roots []string, workers int,
|
||||
|
||||
dbPath := databasePath()
|
||||
|
||||
lock, err := lockScanDatabase(dbPath)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
defer func() { _ = lock.Close() }()
|
||||
|
||||
db, err := openScanDatabase(ctx, dbPath)
|
||||
if err != nil {
|
||||
return err
|
||||
|
||||
+2
-1
@@ -7,6 +7,7 @@ import (
|
||||
"database/sql"
|
||||
"encoding/hex"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"runtime"
|
||||
@@ -1548,7 +1549,7 @@ func TestScanHashWriteFailureUnwindsPool(t *testing.T) {
|
||||
|
||||
code := run([]string{
|
||||
cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir,
|
||||
}, &stderr)
|
||||
}, io.Discard, &stderr)
|
||||
if code != exitFatal {
|
||||
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
|
||||
code, exitFatal, stderr.String())
|
||||
|
||||
@@ -5,6 +5,7 @@ import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"slices"
|
||||
"strconv"
|
||||
@@ -36,7 +37,7 @@ type treeNode struct {
|
||||
// maximal duplicate-tree groups as TSV on stdout. It never touches the
|
||||
// scanned filesystem; its only I/O is the database, stdout, and
|
||||
// stderr.
|
||||
func runTrees(ctx context.Context) error {
|
||||
func runTrees(ctx context.Context, stdout io.Writer) error {
|
||||
recs, err := loadRecords(ctx)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -47,7 +48,7 @@ func runTrees(ctx context.Context) error {
|
||||
|
||||
dupes := collectTreeGroups(allDirs, super)
|
||||
|
||||
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
|
||||
out := bufio.NewWriterSize(stdout, ioBufSize)
|
||||
|
||||
_, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
|
||||
if err != nil {
|
||||
|
||||
+3
-5
@@ -108,11 +108,9 @@ func TestBuildHierarchyRootPath(t *testing.T) {
|
||||
func TestRunTreesEscapesPaths(t *testing.T) {
|
||||
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
|
||||
|
||||
var stderr bytes.Buffer
|
||||
var stdout, stderr bytes.Buffer
|
||||
|
||||
stdout := captureStdout(t)
|
||||
|
||||
code := run([]string{cmdTrees}, &stderr)
|
||||
code := run([]string{cmdTrees}, &stdout, &stderr)
|
||||
if code != exitOK {
|
||||
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
|
||||
code, exitOK, stderr.String())
|
||||
@@ -120,7 +118,7 @@ func TestRunTreesEscapesPaths(t *testing.T) {
|
||||
|
||||
want := "first\tdupe\tfiles\tsize\n" +
|
||||
`/d/\tone\ntwo\rthree\\four` + "\t/d/A\t1\t5\n"
|
||||
if got := stdout(); got != want {
|
||||
if got := stdout.String(); got != want {
|
||||
t.Errorf("stdout = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user