Author SHA1 Message Date
sneak 933a7ced69 Refuse an unversioned database that already has a files table (closes #11)
check / check (push) Successful in 1m46s
An unversioned database (user_version 0) that already contains a files
table was not created by this build; it is a foreign or partially
initialized file. Adopting it silently could corrupt unrelated data, so
createSchema now checks for a files table first and, when one exists,
returns the schema-version error telling the operator to remove the file
and rescan. A genuinely empty database is still created and stamped as
before.

Model: opus-4-8 (implementation); opus-5-5 (rebase)
2026-10-03 13:32:17 +00:00
13 changed files with 196 additions and 490 deletions
+19 -59
View File
@@ -113,9 +113,8 @@ Goals, in order:
`sfdupes`.
- Dependencies: standard library, `github.com/spf13/cobra` for the
CLI, **one progress-bar library**
(`github.com/schollz/progressbar/v3`), **one SQLite driver**
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled), and
`golang.org/x/sys` for `flock(2)` (the scan lock, see "Database").
(`github.com/schollz/progressbar/v3`), and **one SQLite driver**
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled).
`github.com/spf13/viper` is permitted if configuration-file support
is ever needed, but is not currently used. No other third-party
deps.
@@ -153,42 +152,16 @@ All three subcommands operate on a single SQLite database file:
use. `report` and `trees` require an existing database; a missing
database file is a fatal error (exit 1) telling the user to run
`scan` first.
- Only one `scan` runs against a database at a time. For its whole
run, `scan` holds an exclusive `flock(2)` lock on a lock file
beside the database, named by appending `.lock` to the database
path (`/var/lib/sfdupes/db.sqlite.lock` by default), taken before
it walks the filesystem or opens the database. A second `scan`
against the same database does not wait: it fails at once with a
one-line error naming the lock file and exits 1, without walking
anything or opening the database, and the running scan carries on.
The lock file is created on first use, open to its owner only, and
left in place: a leftover file blocks nothing, because the lock
ends with the process holding it however it ends, a fatal error or
an interrupt included, and deleting the file while a scan runs
would let a second scan start. `report` and `trees` never take the
lock, so they run during a scan.
- While `scan` runs, the database is in WAL journal mode with a busy
timeout, so running a report while a cron `scan` is in progress is
safe. The filesystem is authoritative; the database is an
eventually-consistent reflection of it. Hashed records are
committed in batched transactions while the scan is still running
(keeping the WAL small and letting concurrent reports observe
progress), so a report may see a scan's changes partially applied,
and a scan that dies partway leaves a valid database holding
everything hashed so far; the next scan skips those records and
converges toward the filesystem.
- `scan` switches the database back to rollback-journal mode when it
closes it, so between scans the database file alone holds the whole
database. Each switch needs the database to itself: a `scan` that
starts while a report is still reading waits for it up to the
10-second busy timeout, then fails; a `scan` that ends while a
report has the database open warns and leaves the database in WAL
mode until the next scan.
- `report` and `trees` open the database read-only and need only read
access to the database file, and no write access to its directory.
While the database is in WAL mode they also read the `-wal` and
`-shm` files beside it, which SQLite creates with the database
file's permissions.
- The database uses WAL journal mode and a busy timeout, so running a
report while a cron `scan` is in progress is safe. The filesystem
is authoritative; the database is an eventually-consistent
reflection of it. Hashed records are committed in batched
transactions while the scan is still running (keeping the WAL
small and letting concurrent reports observe progress), so a
report may see a scan's changes partially applied, and a scan
that dies partway leaves a valid database holding everything
hashed so far; the next scan skips those records and converges
toward the filesystem.
- Schema (`PRAGMA user_version` is the schema version, currently 1; a
database with any other version is a fatal error):
@@ -582,25 +555,12 @@ Additional requirements:
### Error handling and exit codes
- `0`: success, even if individual files were skipped with warnings.
- `1`: fatal error (e.g., a `PATH` operand does not exist, another
`scan` is already running against the same database, the database
cannot be created/opened/read/written, a missing database for
`report`/`trees`, stdout write failure).
- `1`: fatal error (e.g., a `PATH` operand does not exist, the
database cannot be created/opened/read/written, a missing database
for `report`/`trees`, stdout write failure).
- `2`: usage error (including `scan` with no `PATH` operand and
`report`/`trees` with any positional argument).
A stdout write failure, such as a full disk, is reported in one line on
stderr and exits 1. Two cases never reach sfdupes as a failed write:
- When the reader of a stdout pipe exits early, as in
`sfdupes report | head`, the next write ends sfdupes with `SIGPIPE`,
quietly and without a summary, the way `cat` or `sort` end. The
shell reports the signal (status 141 in most shells), not exit 1.
- When stdout is closed outright (`sfdupes report >&-`), the Go
runtime opens `/dev/null` in its place before sfdupes starts, so
the output is discarded and the run succeeds, as with
`> /dev/null`.
## Entrypoints
This repository adheres to the
@@ -738,7 +698,7 @@ All of the following, run in this directory, must pass:
```sh
d=$(mktemp -d)
export SFDUPES_DATABASE="$(mktemp -d)/db.sqlite"
export SFDUPES_DATABASE="$d/db.sqlite"
mkdir -p "$d/a" "$d/b"
head -c 2000 /dev/urandom > "$d/a/one.bin"
cp "$d/a/one.bin" "$d/b/copy.bin"
@@ -766,9 +726,9 @@ All of the following, run in this directory, must pass:
./sfdupes report
```
(The database lives in a temp directory of its own: inside `$d`,
the scan would record it, and its empty lock file would join the
`empty1`/`empty2` group.)
(The scan database lives inside `$d` here purely for test hygiene;
scanning `$d` therefore also records the SQLite file itself, which
is harmless.)
Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin`
form one group (two dupe rows, `first` is the lexicographically
+3 -11
View File
@@ -29,17 +29,9 @@
# Completed Steps
- `scan` holds a lock on a lock file beside the database for its whole run,
so a second `scan` fails at once with exit 1 (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/53)
- test stdout write failures in `report` and `trees`; README states that
`| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null`
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/30)
- `report` and `trees` open the database read-only, and `scan` leaves it
out of WAL mode, so reading needs only read access (2026-10-03, closes
https://git.eeqj.de/sneak/sfdupes/issues/8)
- refuse an unversioned database that already has a `files` table with
a clear schema-version error (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/11)
- escape tabs, newlines, carriage returns and backslashes in report,
trees and warning paths; the root directory's path is `/`
+34 -86
View File
@@ -11,7 +11,6 @@ import (
"slices"
"strconv"
"golang.org/x/sys/unix"
// The pure-Go SQLite driver, registered as "sqlite"; keeps cgo
// disabled.
_ "modernc.org/sqlite"
@@ -33,11 +32,6 @@ const schemaVersion = 1
// scan.
const dbDirPerm = 0o755
// lockFilePerm is the mode for the scan lock file. Anyone who can open
// the file can hold the lock and keep every scan from running, so it
// is open to its owner only.
const lockFilePerm = 0o600
// createTableSQL is the schema applied to a fresh database. Paths are
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
const createTableSQL = `
@@ -70,10 +64,6 @@ var errNoDatabase = errors.New(
// does not understand.
var errSchemaVersion = errors.New("unsupported database schema version")
// errScanRunning reports that another scan holds the lock on the
// database.
var errScanRunning = errors.New("another scan is running")
// databasePath resolves the database location: SFDUPES_DATABASE when
// set and non-empty, the compiled-in default otherwise.
func databasePath() string {
@@ -84,24 +74,16 @@ func databasePath() string {
return defaultDatabasePath
}
// scanParams are the connection parameters for scan: read-write, with
// WAL journaling and a busy timeout, so a report can run while a cron
// scan is in progress. closeScanDatabase leaves WAL mode again.
const scanParams = "_pragma=busy_timeout(10000)" +
"&_pragma=journal_mode(WAL)" +
"&_pragma=synchronous(NORMAL)"
// openDB opens the SQLite database at path with WAL journaling and a
// busy timeout, so a report can run while a cron scan is in progress.
// It does not create or verify the schema.
func openDB(path string) (*sql.DB, error) {
dsn := "file:" + path +
"?_pragma=busy_timeout(10000)" +
"&_pragma=journal_mode(WAL)" +
"&_pragma=synchronous(NORMAL)"
// reportParams are the connection parameters for report and trees:
// read-only, with the same busy timeout. They set no journal mode,
// because setting one is a write.
const reportParams = "mode=ro" +
"&_pragma=busy_timeout(10000)" +
"&_pragma=query_only(1)"
// openDB opens the SQLite database at path with the connection
// parameters params. It does not create or verify the schema.
func openDB(path, params string) (*sql.DB, error) {
db, err := sql.Open("sqlite", "file:"+path+"?"+params)
db, err := sql.Open("sqlite", dsn)
if err != nil {
return nil, fmt.Errorf("open database %s: %w", path, err)
}
@@ -114,43 +96,6 @@ func openDB(path, params string) (*sql.DB, error) {
return db, nil
}
// lockScanDatabase takes the lock that keeps a second scan off the
// database at path: an exclusive flock(2) on the file beside it named
// path with ".lock" appended, created along with the database's parent
// directory if missing. A lock held by another scan fails at once
// instead of waiting. The lock lasts until the returned file is closed
// or the process ends. The file is never deleted: a scan that deleted
// it would let the next scan lock a new file while another still holds
// the old one.
func lockScanDatabase(path string) (*os.File, error) {
err := os.MkdirAll(filepath.Dir(path), dbDirPerm)
if err != nil {
return nil, fmt.Errorf("create database directory: %w", err)
}
lockPath := path + ".lock"
//nolint:gosec // the operator chooses the database path
f, err := os.OpenFile(lockPath, os.O_RDWR|os.O_CREATE, lockFilePerm)
if err != nil {
return nil, err
}
err = unix.Flock(int(f.Fd()), unix.LOCK_EX|unix.LOCK_NB)
if err != nil {
_ = f.Close()
if errors.Is(err, unix.EWOULDBLOCK) {
return nil, fmt.Errorf("%w (lock held on %s)",
errScanRunning, lockPath)
}
return nil, fmt.Errorf("lock %s: %w", lockPath, err)
}
return f, nil
}
// openScanDatabase opens the database for the scan subcommand, creating
// the file, its parent directory, and the schema as needed.
func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
@@ -159,7 +104,7 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return nil, fmt.Errorf("create database directory: %w", err)
}
db, err := openDB(path, scanParams)
db, err := openDB(path)
if err != nil {
return nil, err
}
@@ -174,24 +119,6 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return db, nil
}
// closeScanDatabase switches the database at path from WAL back to
// rollback-journal mode and closes it. Out of WAL mode the database
// file alone holds the whole database, so a reader needs no -wal or
// -shm file beside it, nor write access to create them. The switch
// fails while a report has the database open; the database then stays
// in WAL mode, still readable, until a later scan closes it.
func closeScanDatabase(ctx context.Context, db *sql.DB, path string) {
// Runs on the way out of a cancelled scan too.
_, err := db.ExecContext(context.WithoutCancel(ctx),
"PRAGMA journal_mode = DELETE")
if err != nil {
fmt.Fprintf(os.Stderr, "scan: database %s left in WAL mode: %v\n",
path, err)
}
_ = db.Close()
}
// openReportDatabase opens an existing database for the report and
// trees subcommands. A missing database file is an error directing the
// user to run scan first; the schema version must match exactly.
@@ -207,7 +134,7 @@ func openReportDatabase(ctx context.Context,
return nil, fmt.Errorf("database: %w", err)
}
db, err := openDB(path, reportParams)
db, err := openDB(path)
if err != nil {
return nil, err
}
@@ -249,9 +176,30 @@ func initSchema(ctx context.Context, db *sql.DB) error {
}
// createSchema applies the schema to a fresh database and stamps the
// schema version.
// schema version. A database with user_version 0 that already has a
// files table was not created by this build — a foreign or partially
// initialized file. Adopting it silently could corrupt unrelated data,
// so that is a fatal schema-version error telling the operator to
// remove the file and rescan.
func createSchema(ctx context.Context, db *sql.DB) error {
_, err := db.ExecContext(ctx, createTableSQL)
var name string
err := db.QueryRowContext(ctx,
"SELECT name FROM sqlite_master "+
"WHERE type = 'table' AND name = 'files'").Scan(&name)
switch {
case err == nil:
return fmt.Errorf(
"has a files table but no schema version; "+
"remove the file and rescan: %w", errSchemaVersion)
case errors.Is(err, sql.ErrNoRows):
// Genuinely empty: create the schema below.
default:
return fmt.Errorf("check for files table: %w", err)
}
_, err = db.ExecContext(ctx, createTableSQL)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
+32 -44
View File
@@ -5,7 +5,6 @@ import (
"database/sql"
"errors"
"fmt"
"os"
"path/filepath"
"slices"
"strings"
@@ -79,6 +78,38 @@ func TestOpenScanDatabaseCreates(t *testing.T) {
}
}
func TestOpenScanDatabaseUnversionedForeign(t *testing.T) {
t.Parallel()
path := testDBPath(t)
// A database that has a files table but user_version 0 — a foreign
// or partially initialized file. scan must refuse it with a clear
// schema-version error, not adopt it and not emit a raw SQLite
// "table files already exists".
db, err := openDB(path)
if err != nil {
t.Fatal(err)
}
_, err = db.ExecContext(t.Context(), "CREATE TABLE files (x INTEGER)")
if err != nil {
t.Fatal(err)
}
_ = db.Close()
_, err = openScanDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) {
t.Fatalf("err = %v, want errSchemaVersion", err)
}
if !strings.Contains(err.Error(), "remove the file and rescan") {
t.Fatalf("err = %v, want it to tell the operator to remove and rescan",
err)
}
}
func TestOpenReportDatabaseMissing(t *testing.T) {
t.Parallel()
@@ -131,49 +162,6 @@ func TestOpenReportDatabaseOK(t *testing.T) {
_ = db.Close()
}
func TestCloseScanDatabaseWhileReportOpen(t *testing.T) {
t.Parallel()
// A report holding the database open stops scan from taking it out
// of WAL mode. The -wal and -shm files must then stay beside it, so
// that a later report still needs only read access.
path := testDBPath(t)
scanDB, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
reportDB, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
closeScanDatabase(t.Context(), scanDB, path)
_ = reportDB.Close()
_, err = os.Stat(path + "-wal")
if err != nil {
t.Fatalf("no -wal left: the switch out of WAL mode was not "+
"stopped: %v", err)
}
makeReadOnly(t, path)
reportDB, err = openReportDatabase(t.Context(), path)
if err != nil {
t.Fatalf("openReportDatabase: %v", err)
}
defer func() { _ = reportDB.Close() }()
_, err = loadFileRows(t.Context(), reportDB)
if err != nil {
t.Fatalf("loadFileRows: %v", err)
}
}
func TestApplyChangesRoundTrip(t *testing.T) {
t.Parallel()
+1 -1
View File
@@ -5,7 +5,6 @@ go 1.25.7
require (
github.com/schollz/progressbar/v3 v3.19.1
github.com/spf13/cobra v1.10.2
golang.org/x/sys v0.46.0
modernc.org/sqlite v1.54.0
)
@@ -19,6 +18,7 @@ require (
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/rivo/uniseg v0.4.7 // indirect
github.com/spf13/pflag v1.0.9 // indirect
golang.org/x/sys v0.46.0 // indirect
golang.org/x/term v0.44.0 // indirect
modernc.org/libc v1.74.1 // indirect
modernc.org/mathutil v1.7.1 // indirect
+7 -12
View File
@@ -57,27 +57,22 @@ var errNoSubcommand = errors.New("no subcommand")
var Version = "dev"
func main() {
// Once the reader of a stdout pipe has gone, as in "sfdupes report |
// head", the Go runtime ends the process with SIGPIPE on the next
// write instead of returning an error (README "Error handling").
// Registering for SIGPIPE with os/signal would change that.
os.Exit(run(os.Args[1:], os.Stdout, os.Stderr))
os.Exit(run(os.Args[1:], os.Stderr))
}
// run executes args against the command tree and returns the process
// exit code. It is the program's single exit point: the subcommands
// return their errors instead of exiting, so every deferred cleanup —
// above all closing the database, which checkpoints the SQLite WAL —
// runs before the process ends. The report and trees subcommands write
// their data to stdout.
func run(args []string, stdout, stderr io.Writer) int {
// runs before the process ends.
func run(args []string, stderr io.Writer) int {
// A nil slice makes cobra fall back to os.Args, which would let a
// test binary's own flags reach the command tree.
if args == nil {
args = []string{}
}
root := newRootCommand(stdout, stderr)
root := newRootCommand(stderr)
root.SetArgs(args)
err := root.Execute()
@@ -103,7 +98,7 @@ func run(args []string, stdout, stderr io.Writer) int {
// newRootCommand builds the command tree. Everything on stdout is
// machine-readable data; all human-facing output (help, usage, errors)
// goes to stderr.
func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
func newRootCommand(stderr io.Writer) *cobra.Command {
root := &cobra.Command{
Use: "sfdupes",
Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
@@ -145,7 +140,7 @@ func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the file-level duplicates report",
Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error {
return runReport(ctx, stdout)
return runReport(ctx)
}),
}
@@ -154,7 +149,7 @@ func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the duplicate-tree report",
Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error {
return runTrees(ctx, stdout)
return runTrees(ctx)
}),
}
+77 -245
View File
@@ -44,54 +44,60 @@ func assertNoSidecars(t *testing.T, path string) {
}
}
// makeReadOnly takes write permission away from the database at path,
// from any WAL sidecar beside it, and from their directory, as for a
// user reading a database that a root cron scan keeps. Root ignores
// file permissions, so it skips the test when run as root.
func makeReadOnly(t *testing.T, path string) {
// captureStdout redirects os.Stdout to a file for the rest of the test
// and returns a function reading back everything written to it. Only
// machine-readable data belongs on stdout (README design goal 4), so
// the tests assert on it directly.
func captureStdout(t *testing.T) func() string {
t.Helper()
if os.Geteuid() == 0 {
t.Skip("root ignores file permissions")
}
err := os.Chmod(path, 0o400)
f, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
if err != nil {
t.Fatal(err)
}
for _, suffix := range walSuffixes {
err = os.Chmod(path+suffix, 0o400)
if err != nil && !errors.Is(err, fs.ErrNotExist) {
saved := os.Stdout
os.Stdout = f
t.Cleanup(func() {
os.Stdout = saved
_ = f.Close()
})
return func() string {
// Read what has been written without disturbing the write
// offset, so the capture can be inspected more than once.
size, err := f.Seek(0, io.SeekCurrent)
if err != nil {
t.Fatal(err)
}
if size == 0 {
return ""
}
b := make([]byte, size)
_, err = f.ReadAt(b, 0)
if err != nil {
t.Fatal(err)
}
return string(b)
}
dir := filepath.Dir(path)
//nolint:gosec // reaching the database needs the search bit
err = os.Chmod(dir, 0o500)
if err != nil {
t.Fatal(err)
}
// Runs before t.TempDir's own cleanup, which must delete the files.
t.Cleanup(func() {
//nolint:gosec // removing the directory needs its search bit back
_ = os.Chmod(dir, 0o700)
})
}
// brokenDatabase writes a database that opens cleanly and passes the
// schema-version check but has no files table, so the first query
// fails with the database already open: a fatal error on a path that
// owns an open database. It closes the database the way scan does.
// owns an open database.
func brokenDatabase(t *testing.T) string {
t.Helper()
path := testDBPath(t)
db, err := openDB(path, scanParams)
db, err := openDB(path)
if err != nil {
t.Fatal(err)
}
@@ -102,7 +108,10 @@ func brokenDatabase(t *testing.T) string {
t.Fatal(err)
}
closeScanDatabase(t.Context(), db, path)
err = db.Close()
if err != nil {
t.Fatal(err)
}
return path
}
@@ -135,10 +144,7 @@ func TestOpenDatabaseKeepsWALWhileOpen(t *testing.T) {
func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
// Every subcommand that owns an open database must close it when
// it fails: no os.Exit between the open and the return. The
// sidecar check is evidence of the close only for scan: report and
// trees only read a database that is out of WAL mode, which leaves
// nothing on disk whether they close it or not.
// it fails: no os.Exit between the open and the return.
cases := map[string][]string{
cmdScan: {cmdScan},
cmdReport: {cmdReport},
@@ -154,15 +160,17 @@ func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
args = append(args, t.TempDir())
}
var stdout, stderr bytes.Buffer
var stderr bytes.Buffer
code := run(args, &stdout, &stderr)
stdout := captureStdout(t)
code := run(args, &stderr)
if code != exitFatal {
t.Errorf("run(%v) = %d, want %d", args, code, exitFatal)
}
assertNoSidecars(t, path)
assertFatalOutput(t, stderr.String(), stdout.String())
assertFatalOutput(t, stderr.String(), stdout())
// Proof that the failure happened after the open: only a
// query against the opened database can report this.
@@ -180,16 +188,18 @@ func TestRunMissingOperandIsFatalNotUsage(t *testing.T) {
// must not dump the usage text.
t.Setenv(databaseEnv, testDBPath(t))
var stdout, stderr bytes.Buffer
var stderr bytes.Buffer
stdout := captureStdout(t)
missing := filepath.Join(t.TempDir(), "nope")
code := run([]string{cmdScan, missing}, &stdout, &stderr)
code := run([]string{cmdScan, missing}, &stderr)
if code != exitFatal {
t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal)
}
assertFatalOutput(t, stderr.String(), stdout.String())
assertFatalOutput(t, stderr.String(), stdout())
}
// assertFatalOutput checks that a fatal error was reported the way
@@ -233,9 +243,11 @@ func TestRunUsageErrors(t *testing.T) {
// path that does not exist.
t.Setenv(databaseEnv, testDBPath(t))
var stdout, stderr bytes.Buffer
var stderr bytes.Buffer
code := run(tc.args, &stdout, &stderr)
stdout := captureStdout(t)
code := run(tc.args, &stderr)
if code != exitUsage {
t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage)
}
@@ -244,7 +256,7 @@ func TestRunUsageErrors(t *testing.T) {
t.Errorf("stderr = %q, want %q", stderr.String(), tc.want)
}
if got := stdout.String(); got != "" {
if got := stdout(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
})
@@ -253,9 +265,9 @@ func TestRunUsageErrors(t *testing.T) {
// TestRunHelpAndVersionSucceed checks that the two informational flags
// exit 0 and keep their human-facing output on stderr.
//
//nolint:paralleltest // captureStdout replaces the process-wide os.Stdout
func TestRunHelpAndVersionSucceed(t *testing.T) {
t.Parallel()
assertHumanOutput(t, "--help")
assertHumanOutput(t, "--version")
}
@@ -266,9 +278,11 @@ func TestRunHelpAndVersionSucceed(t *testing.T) {
func assertHumanOutput(t *testing.T, arg string) {
t.Helper()
var stdout, stderr bytes.Buffer
var stderr bytes.Buffer
code := run([]string{arg}, &stdout, &stderr)
stdout := captureStdout(t)
code := run([]string{arg}, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
}
@@ -277,7 +291,7 @@ func assertHumanOutput(t *testing.T, arg string) {
t.Errorf("run(%s) wrote nothing to stderr", arg)
}
if got := stdout.String(); got != "" {
if got := stdout(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
}
@@ -306,15 +320,17 @@ func scanFixture(t *testing.T) []string {
t.Fatal(err)
}
var stdout, stderr bytes.Buffer
var stderr bytes.Buffer
code := run([]string{cmdScan, dir}, &stdout, &stderr)
stdout := captureStdout(t)
code := run([]string{cmdScan, dir}, &stderr)
if code != exitOK {
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
if got := stdout.String(); got != "" {
if got := stdout(); got != "" {
t.Errorf("scan stdout = %q, want nothing (data only)", got)
}
@@ -335,16 +351,18 @@ func TestRunReportSucceeds(t *testing.T) {
dupes := scanFixture(t)
var stdout, stderr bytes.Buffer
var stderr bytes.Buffer
code := run([]string{cmdReport}, &stdout, &stderr)
stdout := captureStdout(t)
code := run([]string{cmdReport}, &stderr)
if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
if got := stdout.String(); got != want {
if got := stdout(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
@@ -357,9 +375,11 @@ func TestRunTreesSucceeds(t *testing.T) {
dupes := scanFixture(t)
var stdout, stderr bytes.Buffer
var stderr bytes.Buffer
code := run([]string{cmdTrees}, &stdout, &stderr)
stdout := captureStdout(t)
code := run([]string{cmdTrees}, &stderr)
if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
@@ -369,197 +389,9 @@ func TestRunTreesSucceeds(t *testing.T) {
// trees of each other.
want := "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
if got := stdout.String(); got != want {
if got := stdout(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
assertNoSidecars(t, path)
}
func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
// README §Database: report and trees need only read access to the
// database file. With its directory read-only as well, SQLite
// cannot create any file beside it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dupes := scanFixture(t)
assertNoSidecars(t, path)
makeReadOnly(t, path)
cases := map[string]string{
cmdReport: "first\tdupe\tsize\n" +
dupes[0] + "\t" + dupes[1] + "\t300\n",
cmdTrees: "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) +
"\t1\t300\n",
}
for name, want := range cases {
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
continue
}
if got := stdout.String(); got != want {
t.Errorf("%s stdout = %q, want %q", name, got, want)
}
}
}
// holdScanLock takes the lock on the database at path, as a running
// scan does, and holds it until the test ends. It fails the test when
// the lock is already held.
func holdScanLock(t *testing.T, path string) {
t.Helper()
lock, err := lockScanDatabase(path)
if err != nil {
t.Fatalf("lock %s: %v", path, err)
}
t.Cleanup(func() { _ = lock.Close() })
}
func TestRunSecondScanFails(t *testing.T) {
// README §Database: while one scan holds the lock, a second scan
// fails at once, naming the lock file, without creating the
// database.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
holdScanLock(t, path)
var stdout, stderr bytes.Buffer
code := run([]string{cmdScan, t.TempDir()}, &stdout, &stderr)
if code != exitFatal {
t.Errorf("run(scan) = %d, want %d", code, exitFatal)
}
want := "sfdupes: another scan is running (lock held on " +
path + ".lock)\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
_, err := os.Stat(path)
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("stat %s = %v, want the database not created", path, err)
}
}
func TestRunScanReleasesLock(t *testing.T) {
// README §Database: a scan releases the lock however it ends.
t.Run("success", func(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
})
t.Run("fatal error", func(t *testing.T) {
path := brokenDatabase(t)
t.Setenv(databaseEnv, path)
code := run([]string{cmdScan, t.TempDir()}, io.Discard, io.Discard)
if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d", code, exitFatal)
}
holdScanLock(t, path)
})
}
func TestRunReportsDuringScan(t *testing.T) {
// README §Database: report and trees never take the lock, so they
// run while a scan holds it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
for _, name := range []string{cmdReport, cmdTrees} {
var stderr bytes.Buffer
code := run([]string{name}, io.Discard, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
}
}
}
func TestRunStdoutClosedIsFatal(t *testing.T) {
// README §Error handling: a stdout write failure exits 1, reported
// in one line on stderr.
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
stdout, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
if err != nil {
t.Fatal(err)
}
err = stdout.Close()
if err != nil {
t.Fatal(err)
}
var stderr bytes.Buffer
code := run([]string{name}, stdout, &stderr)
if code != exitFatal {
t.Errorf("run(%s) = %d, want %d", name, code, exitFatal)
}
got := stderr.String()
if !strings.HasPrefix(got, "sfdupes: write stdout: ") ||
!strings.Contains(got, os.ErrClosed.Error()) ||
strings.Count(got, "\n") != 1 {
t.Errorf("stderr = %q, want one line reporting the "+
"failed stdout write", got)
}
})
}
}
// errWriteFailed is the error failingWriter returns.
var errWriteFailed = errors.New("write failed")
// failingWriter is a stdout that fails every write.
type failingWriter struct{}
func (failingWriter) Write([]byte) (int, error) { return 0, errWriteFailed }
func TestStdoutWriteErrorPropagates(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
cases := map[string]func(context.Context, io.Writer) error{
cmdReport: runReport,
cmdTrees: runTrees,
}
for name, fn := range cases {
err := fn(t.Context(), failingWriter{})
if !errors.Is(err, errWriteFailed) {
t.Errorf("%s: error = %v, want %v", name, err, errWriteFailed)
}
}
}
+5 -6
View File
@@ -4,7 +4,6 @@ import (
"bufio"
"context"
"fmt"
"io"
"os"
"slices"
"strings"
@@ -32,9 +31,9 @@ type scanRec struct {
// loadRecords opens the database and reads every file record for the
// report and trees subcommands. Any database problem — including a
// missing database — is fatal. The error is returned rather than
// exiting, so that the deferred close always runs; the database is
// closed before the caller formats its output, so it stays closed even
// if that output fails.
// exiting, so that the deferred close — which checkpoints the SQLite
// WAL — always runs; the database is closed before the caller formats
// its output, so it stays closed even if that output fails.
func loadRecords(ctx context.Context) ([]scanRec, error) {
dbPath := databasePath()
@@ -66,7 +65,7 @@ type dupeGroup struct {
// from the database and prints the file-level duplicates report as TSV
// on stdout. It never touches the scanned filesystem; its only I/O is
// the database, stdout, and stderr.
func runReport(ctx context.Context, stdout io.Writer) error {
func runReport(ctx context.Context) error {
recs, err := loadRecords(ctx)
if err != nil {
return err
@@ -74,7 +73,7 @@ func runReport(ctx context.Context, stdout io.Writer) error {
dupes := collectDupeGroups(recs)
out := bufio.NewWriterSize(stdout, ioBufSize)
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tsize")
if err != nil {
+5 -3
View File
@@ -50,9 +50,11 @@ func seedDatabase(t *testing.T, recs []scanRec) string {
func TestRunReportEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
var stderr bytes.Buffer
code := run([]string{cmdReport}, &stdout, &stderr)
stdout := captureStdout(t)
code := run([]string{cmdReport}, &stderr)
if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
@@ -60,7 +62,7 @@ func TestRunReportEscapesPaths(t *testing.T) {
want := "first\tdupe\tsize\n" +
`/d/\tone\ntwo\rthree\\four/f` + "\t/d/A/f\t5\n"
if got := stdout.String(); got != want {
if got := stdout(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
+5 -15
View File
@@ -85,13 +85,10 @@ type fileMeta struct {
// least one other file shares are ever hashed: a size-unique file
// cannot be a duplicate. A file of headTailMin or more gets its content
// hash only when its size, head, and tail match another file's. Flag
// parsing and the at-least-one-operand check are done by cobra. The
// scan holds the lock on the database for its whole run, so a second
// scan fails before it walks the filesystem or opens the database.
// Errors are returned rather than exiting, so that the deferred close —
// which takes the database out of WAL mode — always runs, and the lock
// is released after it. Cancelling ctx unwinds the worker pools and
// aborts the scan with the context's error.
// parsing and the at-least-one-operand check are done by cobra. Errors
// are returned rather than exiting, so that the deferred close — which
// checkpoints the SQLite WAL — always runs. Cancelling ctx unwinds the
// worker pools and aborts the scan with the context's error.
func runScan(ctx context.Context, roots []string, workers int,
oneFS bool,
) error {
@@ -106,19 +103,12 @@ func runScan(ctx context.Context, roots []string, workers int,
dbPath := databasePath()
lock, err := lockScanDatabase(dbPath)
if err != nil {
return err
}
defer func() { _ = lock.Close() }()
db, err := openScanDatabase(ctx, dbPath)
if err != nil {
return err
}
defer closeScanDatabase(ctx, db, dbPath)
defer func() { _ = db.Close() }()
st, err := syncScan(ctx, db, roots, workers, oneFS)
if err != nil {
+1 -2
View File
@@ -7,7 +7,6 @@ import (
"database/sql"
"encoding/hex"
"fmt"
"io"
"os"
"path/filepath"
"runtime"
@@ -1549,7 +1548,7 @@ func TestScanHashWriteFailureUnwindsPool(t *testing.T) {
code := run([]string{
cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir,
}, io.Discard, &stderr)
}, &stderr)
if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
code, exitFatal, stderr.String())
+2 -3
View File
@@ -5,7 +5,6 @@ import (
"context"
"crypto/sha256"
"fmt"
"io"
"os"
"slices"
"strconv"
@@ -37,7 +36,7 @@ type treeNode struct {
// maximal duplicate-tree groups as TSV on stdout. It never touches the
// scanned filesystem; its only I/O is the database, stdout, and
// stderr.
func runTrees(ctx context.Context, stdout io.Writer) error {
func runTrees(ctx context.Context) error {
recs, err := loadRecords(ctx)
if err != nil {
return err
@@ -48,7 +47,7 @@ func runTrees(ctx context.Context, stdout io.Writer) error {
dupes := collectTreeGroups(allDirs, super)
out := bufio.NewWriterSize(stdout, ioBufSize)
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
if err != nil {
+5 -3
View File
@@ -108,9 +108,11 @@ func TestBuildHierarchyRootPath(t *testing.T) {
func TestRunTreesEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
var stderr bytes.Buffer
code := run([]string{cmdTrees}, &stdout, &stderr)
stdout := captureStdout(t)
code := run([]string{cmdTrees}, &stderr)
if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
@@ -118,7 +120,7 @@ func TestRunTreesEscapesPaths(t *testing.T) {
want := "first\tdupe\tfiles\tsize\n" +
`/d/\tone\ntwo\rthree\\four` + "\t/d/A\t1\t5\n"
if got := stdout.String(); got != want {
if got := stdout(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}