Compare commits
4
Commits
1f64f4d6de
..
next
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
01ff3bb5f0 | ||
|
|
d63d3cc7fc | ||
|
|
c887f80f57 | ||
|
|
c9bf22d483 |
@@ -152,16 +152,28 @@ All three subcommands operate on a single SQLite database file:
|
|||||||
use. `report` and `trees` require an existing database; a missing
|
use. `report` and `trees` require an existing database; a missing
|
||||||
database file is a fatal error (exit 1) telling the user to run
|
database file is a fatal error (exit 1) telling the user to run
|
||||||
`scan` first.
|
`scan` first.
|
||||||
- The database uses WAL journal mode and a busy timeout, so running a
|
- While `scan` runs, the database is in WAL journal mode with a busy
|
||||||
report while a cron `scan` is in progress is safe. The filesystem
|
timeout, so running a report while a cron `scan` is in progress is
|
||||||
is authoritative; the database is an eventually-consistent
|
safe. The filesystem is authoritative; the database is an
|
||||||
reflection of it. Hashed records are committed in batched
|
eventually-consistent reflection of it. Hashed records are
|
||||||
transactions while the scan is still running (keeping the WAL
|
committed in batched transactions while the scan is still running
|
||||||
small and letting concurrent reports observe progress), so a
|
(keeping the WAL small and letting concurrent reports observe
|
||||||
report may see a scan's changes partially applied, and a scan
|
progress), so a report may see a scan's changes partially applied,
|
||||||
that dies partway leaves a valid database holding everything
|
and a scan that dies partway leaves a valid database holding
|
||||||
hashed so far; the next scan skips those records and converges
|
everything hashed so far; the next scan skips those records and
|
||||||
toward the filesystem.
|
converges toward the filesystem.
|
||||||
|
- `scan` switches the database back to rollback-journal mode when it
|
||||||
|
closes it, so between scans the database file alone holds the whole
|
||||||
|
database. Each switch needs the database to itself: a `scan` that
|
||||||
|
starts while a report is still reading waits for it up to the
|
||||||
|
10-second busy timeout, then fails; a `scan` that ends while a
|
||||||
|
report has the database open warns and leaves the database in WAL
|
||||||
|
mode until the next scan.
|
||||||
|
- `report` and `trees` open the database read-only and need only read
|
||||||
|
access to the database file, and no write access to its directory.
|
||||||
|
While the database is in WAL mode they also read the `-wal` and
|
||||||
|
`-shm` files beside it, which SQLite creates with the database
|
||||||
|
file's permissions.
|
||||||
- Schema (`PRAGMA user_version` is the schema version, currently 1; a
|
- Schema (`PRAGMA user_version` is the schema version, currently 1; a
|
||||||
database with any other version is a fatal error):
|
database with any other version is a fatal error):
|
||||||
|
|
||||||
@@ -428,6 +440,15 @@ first dupe size
|
|||||||
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296
|
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296
|
||||||
```
|
```
|
||||||
|
|
||||||
|
Paths are raw bytes and may hold any byte except NUL, so the path
|
||||||
|
columns (`first` and `dupe`) are escaped to keep every row one line of
|
||||||
|
tab-separated fields: a backslash is written as `\\`, a tab as `\t`, a
|
||||||
|
newline as `\n`, and a carriage return as `\r`. Every other byte is
|
||||||
|
written unchanged, including bytes that are not valid UTF-8. Undoing
|
||||||
|
those four escapes gives back the stored path. Grouping and ordering
|
||||||
|
use the stored path, not the escaped one. The warnings `scan` prints on
|
||||||
|
stderr are escaped the same way, so each warning is one line.
|
||||||
|
|
||||||
Summary to stderr: records read, number of duplicate groups, number of
|
Summary to stderr: records read, number of duplicate groups, number of
|
||||||
dupe files, and total reclaimable bytes (sum of `size` over all dupe
|
dupe files, and total reclaimable bytes (sum of `size` over all dupe
|
||||||
rows) in human units.
|
rows) in human units.
|
||||||
@@ -499,6 +520,9 @@ first dupe files size
|
|||||||
/srv/a/project /srv/backup/project 3417 104857600
|
/srv/a/project /srv/backup/project 3417 104857600
|
||||||
```
|
```
|
||||||
|
|
||||||
|
The `first` and `dupe` paths are escaped as described under "Report
|
||||||
|
output format". The root directory's path is `/`.
|
||||||
|
|
||||||
Summary to stderr: records read, number of duplicate-tree groups,
|
Summary to stderr: records read, number of duplicate-tree groups,
|
||||||
number of dupe trees, and total reclaimable bytes (sum of `size` over
|
number of dupe trees, and total reclaimable bytes (sum of `size` over
|
||||||
all dupe rows) in human units.
|
all dupe rows) in human units.
|
||||||
@@ -549,6 +573,18 @@ Additional requirements:
|
|||||||
- `2`: usage error (including `scan` with no `PATH` operand and
|
- `2`: usage error (including `scan` with no `PATH` operand and
|
||||||
`report`/`trees` with any positional argument).
|
`report`/`trees` with any positional argument).
|
||||||
|
|
||||||
|
A stdout write failure, such as a full disk, is reported in one line on
|
||||||
|
stderr and exits 1. Two cases never reach sfdupes as a failed write:
|
||||||
|
|
||||||
|
- When the reader of a stdout pipe exits early, as in
|
||||||
|
`sfdupes report | head`, the next write ends sfdupes with `SIGPIPE`,
|
||||||
|
quietly and without a summary, the way `cat` or `sort` end. The
|
||||||
|
shell reports the signal (status 141 in most shells), not exit 1.
|
||||||
|
- When stdout is closed outright (`sfdupes report >&-`), the Go
|
||||||
|
runtime opens `/dev/null` in its place before sfdupes starts, so
|
||||||
|
the output is discarded and the run succeeds, as with
|
||||||
|
`> /dev/null`.
|
||||||
|
|
||||||
## Entrypoints
|
## Entrypoints
|
||||||
|
|
||||||
This repository adheres to the
|
This repository adheres to the
|
||||||
|
|||||||
@@ -29,6 +29,18 @@
|
|||||||
|
|
||||||
# Completed Steps
|
# Completed Steps
|
||||||
|
|
||||||
|
- test stdout write failures in `report` and `trees`; README states that
|
||||||
|
`| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null`
|
||||||
|
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/30)
|
||||||
|
|
||||||
|
- `report` and `trees` open the database read-only, and `scan` leaves it
|
||||||
|
out of WAL mode, so reading needs only read access (2026-10-03, closes
|
||||||
|
https://git.eeqj.de/sneak/sfdupes/issues/8)
|
||||||
|
|
||||||
|
- escape tabs, newlines, carriage returns and backslashes in report,
|
||||||
|
trees and warning paths; the root directory's path is `/`
|
||||||
|
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/7)
|
||||||
|
|
||||||
- stamp the git tag or short commit in a plain `docker build .`
|
- stamp the git tag or short commit in a plain `docker build .`
|
||||||
instead of `dev` (2026-10-02, branch `next`, closes
|
instead of `dev` (2026-10-02, branch `next`, closes
|
||||||
https://git.eeqj.de/sneak/sfdupes/issues/67): `.dockerignore` now
|
https://git.eeqj.de/sneak/sfdupes/issues/67): `.dockerignore` now
|
||||||
|
|||||||
@@ -74,16 +74,24 @@ func databasePath() string {
|
|||||||
return defaultDatabasePath
|
return defaultDatabasePath
|
||||||
}
|
}
|
||||||
|
|
||||||
// openDB opens the SQLite database at path with WAL journaling and a
|
// scanParams are the connection parameters for scan: read-write, with
|
||||||
// busy timeout, so a report can run while a cron scan is in progress.
|
// WAL journaling and a busy timeout, so a report can run while a cron
|
||||||
// It does not create or verify the schema.
|
// scan is in progress. closeScanDatabase leaves WAL mode again.
|
||||||
func openDB(path string) (*sql.DB, error) {
|
const scanParams = "_pragma=busy_timeout(10000)" +
|
||||||
dsn := "file:" + path +
|
"&_pragma=journal_mode(WAL)" +
|
||||||
"?_pragma=busy_timeout(10000)" +
|
"&_pragma=synchronous(NORMAL)"
|
||||||
"&_pragma=journal_mode(WAL)" +
|
|
||||||
"&_pragma=synchronous(NORMAL)"
|
|
||||||
|
|
||||||
db, err := sql.Open("sqlite", dsn)
|
// reportParams are the connection parameters for report and trees:
|
||||||
|
// read-only, with the same busy timeout. They set no journal mode,
|
||||||
|
// because setting one is a write.
|
||||||
|
const reportParams = "mode=ro" +
|
||||||
|
"&_pragma=busy_timeout(10000)" +
|
||||||
|
"&_pragma=query_only(1)"
|
||||||
|
|
||||||
|
// openDB opens the SQLite database at path with the connection
|
||||||
|
// parameters params. It does not create or verify the schema.
|
||||||
|
func openDB(path, params string) (*sql.DB, error) {
|
||||||
|
db, err := sql.Open("sqlite", "file:"+path+"?"+params)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, fmt.Errorf("open database %s: %w", path, err)
|
return nil, fmt.Errorf("open database %s: %w", path, err)
|
||||||
}
|
}
|
||||||
@@ -104,7 +112,7 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
|
|||||||
return nil, fmt.Errorf("create database directory: %w", err)
|
return nil, fmt.Errorf("create database directory: %w", err)
|
||||||
}
|
}
|
||||||
|
|
||||||
db, err := openDB(path)
|
db, err := openDB(path, scanParams)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
@@ -119,6 +127,24 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
|
|||||||
return db, nil
|
return db, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// closeScanDatabase switches the database at path from WAL back to
|
||||||
|
// rollback-journal mode and closes it. Out of WAL mode the database
|
||||||
|
// file alone holds the whole database, so a reader needs no -wal or
|
||||||
|
// -shm file beside it, nor write access to create them. The switch
|
||||||
|
// fails while a report has the database open; the database then stays
|
||||||
|
// in WAL mode, still readable, until a later scan closes it.
|
||||||
|
func closeScanDatabase(ctx context.Context, db *sql.DB, path string) {
|
||||||
|
// Runs on the way out of a cancelled scan too.
|
||||||
|
_, err := db.ExecContext(context.WithoutCancel(ctx),
|
||||||
|
"PRAGMA journal_mode = DELETE")
|
||||||
|
if err != nil {
|
||||||
|
fmt.Fprintf(os.Stderr, "scan: database %s left in WAL mode: %v\n",
|
||||||
|
path, err)
|
||||||
|
}
|
||||||
|
|
||||||
|
_ = db.Close()
|
||||||
|
}
|
||||||
|
|
||||||
// openReportDatabase opens an existing database for the report and
|
// openReportDatabase opens an existing database for the report and
|
||||||
// trees subcommands. A missing database file is an error directing the
|
// trees subcommands. A missing database file is an error directing the
|
||||||
// user to run scan first; the schema version must match exactly.
|
// user to run scan first; the schema version must match exactly.
|
||||||
@@ -134,7 +160,7 @@ func openReportDatabase(ctx context.Context,
|
|||||||
return nil, fmt.Errorf("database: %w", err)
|
return nil, fmt.Errorf("database: %w", err)
|
||||||
}
|
}
|
||||||
|
|
||||||
db, err := openDB(path)
|
db, err := openDB(path, reportParams)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return nil, err
|
return nil, err
|
||||||
}
|
}
|
||||||
|
|||||||
+44
@@ -5,6 +5,7 @@ import (
|
|||||||
"database/sql"
|
"database/sql"
|
||||||
"errors"
|
"errors"
|
||||||
"fmt"
|
"fmt"
|
||||||
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"slices"
|
"slices"
|
||||||
"strings"
|
"strings"
|
||||||
@@ -130,6 +131,49 @@ func TestOpenReportDatabaseOK(t *testing.T) {
|
|||||||
_ = db.Close()
|
_ = db.Close()
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestCloseScanDatabaseWhileReportOpen(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
// A report holding the database open stops scan from taking it out
|
||||||
|
// of WAL mode. The -wal and -shm files must then stay beside it, so
|
||||||
|
// that a later report still needs only read access.
|
||||||
|
path := testDBPath(t)
|
||||||
|
|
||||||
|
scanDB, err := openScanDatabase(t.Context(), path)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
reportDB, err := openReportDatabase(t.Context(), path)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
closeScanDatabase(t.Context(), scanDB, path)
|
||||||
|
|
||||||
|
_ = reportDB.Close()
|
||||||
|
|
||||||
|
_, err = os.Stat(path + "-wal")
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("no -wal left: the switch out of WAL mode was not "+
|
||||||
|
"stopped: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
makeReadOnly(t, path)
|
||||||
|
|
||||||
|
reportDB, err = openReportDatabase(t.Context(), path)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("openReportDatabase: %v", err)
|
||||||
|
}
|
||||||
|
|
||||||
|
defer func() { _ = reportDB.Close() }()
|
||||||
|
|
||||||
|
_, err = loadFileRows(t.Context(), reportDB)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatalf("loadFileRows: %v", err)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestApplyChangesRoundTrip(t *testing.T) {
|
func TestApplyChangesRoundTrip(t *testing.T) {
|
||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
|
|||||||
@@ -57,22 +57,27 @@ var errNoSubcommand = errors.New("no subcommand")
|
|||||||
var Version = "dev"
|
var Version = "dev"
|
||||||
|
|
||||||
func main() {
|
func main() {
|
||||||
os.Exit(run(os.Args[1:], os.Stderr))
|
// Once the reader of a stdout pipe has gone, as in "sfdupes report |
|
||||||
|
// head", the Go runtime ends the process with SIGPIPE on the next
|
||||||
|
// write instead of returning an error (README "Error handling").
|
||||||
|
// Registering for SIGPIPE with os/signal would change that.
|
||||||
|
os.Exit(run(os.Args[1:], os.Stdout, os.Stderr))
|
||||||
}
|
}
|
||||||
|
|
||||||
// run executes args against the command tree and returns the process
|
// run executes args against the command tree and returns the process
|
||||||
// exit code. It is the program's single exit point: the subcommands
|
// exit code. It is the program's single exit point: the subcommands
|
||||||
// return their errors instead of exiting, so every deferred cleanup —
|
// return their errors instead of exiting, so every deferred cleanup —
|
||||||
// above all closing the database, which checkpoints the SQLite WAL —
|
// above all closing the database, which checkpoints the SQLite WAL —
|
||||||
// runs before the process ends.
|
// runs before the process ends. The report and trees subcommands write
|
||||||
func run(args []string, stderr io.Writer) int {
|
// their data to stdout.
|
||||||
|
func run(args []string, stdout, stderr io.Writer) int {
|
||||||
// A nil slice makes cobra fall back to os.Args, which would let a
|
// A nil slice makes cobra fall back to os.Args, which would let a
|
||||||
// test binary's own flags reach the command tree.
|
// test binary's own flags reach the command tree.
|
||||||
if args == nil {
|
if args == nil {
|
||||||
args = []string{}
|
args = []string{}
|
||||||
}
|
}
|
||||||
|
|
||||||
root := newRootCommand(stderr)
|
root := newRootCommand(stdout, stderr)
|
||||||
root.SetArgs(args)
|
root.SetArgs(args)
|
||||||
|
|
||||||
err := root.Execute()
|
err := root.Execute()
|
||||||
@@ -98,7 +103,7 @@ func run(args []string, stderr io.Writer) int {
|
|||||||
// newRootCommand builds the command tree. Everything on stdout is
|
// newRootCommand builds the command tree. Everything on stdout is
|
||||||
// machine-readable data; all human-facing output (help, usage, errors)
|
// machine-readable data; all human-facing output (help, usage, errors)
|
||||||
// goes to stderr.
|
// goes to stderr.
|
||||||
func newRootCommand(stderr io.Writer) *cobra.Command {
|
func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
|
||||||
root := &cobra.Command{
|
root := &cobra.Command{
|
||||||
Use: "sfdupes",
|
Use: "sfdupes",
|
||||||
Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
|
Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
|
||||||
@@ -140,7 +145,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
|
|||||||
Short: "Read the scan database and print the file-level duplicates report",
|
Short: "Read the scan database and print the file-level duplicates report",
|
||||||
Args: cobra.NoArgs,
|
Args: cobra.NoArgs,
|
||||||
RunE: runE(func(ctx context.Context, _ []string) error {
|
RunE: runE(func(ctx context.Context, _ []string) error {
|
||||||
return runReport(ctx)
|
return runReport(ctx, stdout)
|
||||||
}),
|
}),
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -149,7 +154,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
|
|||||||
Short: "Read the scan database and print the duplicate-tree report",
|
Short: "Read the scan database and print the duplicate-tree report",
|
||||||
Args: cobra.NoArgs,
|
Args: cobra.NoArgs,
|
||||||
RunE: runE(func(ctx context.Context, _ []string) error {
|
RunE: runE(func(ctx context.Context, _ []string) error {
|
||||||
return runTrees(ctx)
|
return runTrees(ctx, stdout)
|
||||||
}),
|
}),
|
||||||
}
|
}
|
||||||
|
|
||||||
|
|||||||
+156
-77
@@ -44,60 +44,54 @@ func assertNoSidecars(t *testing.T, path string) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// captureStdout redirects os.Stdout to a file for the rest of the test
|
// makeReadOnly takes write permission away from the database at path,
|
||||||
// and returns a function reading back everything written to it. Only
|
// from any WAL sidecar beside it, and from their directory, as for a
|
||||||
// machine-readable data belongs on stdout (README design goal 4), so
|
// user reading a database that a root cron scan keeps. Root ignores
|
||||||
// the tests assert on it directly.
|
// file permissions, so it skips the test when run as root.
|
||||||
func captureStdout(t *testing.T) func() string {
|
func makeReadOnly(t *testing.T, path string) {
|
||||||
t.Helper()
|
t.Helper()
|
||||||
|
|
||||||
f, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
|
if os.Geteuid() == 0 {
|
||||||
|
t.Skip("root ignores file permissions")
|
||||||
|
}
|
||||||
|
|
||||||
|
err := os.Chmod(path, 0o400)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
|
|
||||||
saved := os.Stdout
|
for _, suffix := range walSuffixes {
|
||||||
os.Stdout = f
|
err = os.Chmod(path+suffix, 0o400)
|
||||||
|
if err != nil && !errors.Is(err, fs.ErrNotExist) {
|
||||||
t.Cleanup(func() {
|
|
||||||
os.Stdout = saved
|
|
||||||
|
|
||||||
_ = f.Close()
|
|
||||||
})
|
|
||||||
|
|
||||||
return func() string {
|
|
||||||
// Read what has been written without disturbing the write
|
|
||||||
// offset, so the capture can be inspected more than once.
|
|
||||||
size, err := f.Seek(0, io.SeekCurrent)
|
|
||||||
if err != nil {
|
|
||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
|
|
||||||
if size == 0 {
|
|
||||||
return ""
|
|
||||||
}
|
|
||||||
|
|
||||||
b := make([]byte, size)
|
|
||||||
|
|
||||||
_, err = f.ReadAt(b, 0)
|
|
||||||
if err != nil {
|
|
||||||
t.Fatal(err)
|
|
||||||
}
|
|
||||||
|
|
||||||
return string(b)
|
|
||||||
}
|
}
|
||||||
|
|
||||||
|
dir := filepath.Dir(path)
|
||||||
|
|
||||||
|
//nolint:gosec // reaching the database needs the search bit
|
||||||
|
err = os.Chmod(dir, 0o500)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
// Runs before t.TempDir's own cleanup, which must delete the files.
|
||||||
|
t.Cleanup(func() {
|
||||||
|
//nolint:gosec // removing the directory needs its search bit back
|
||||||
|
_ = os.Chmod(dir, 0o700)
|
||||||
|
})
|
||||||
}
|
}
|
||||||
|
|
||||||
// brokenDatabase writes a database that opens cleanly and passes the
|
// brokenDatabase writes a database that opens cleanly and passes the
|
||||||
// schema-version check but has no files table, so the first query
|
// schema-version check but has no files table, so the first query
|
||||||
// fails with the database already open: a fatal error on a path that
|
// fails with the database already open: a fatal error on a path that
|
||||||
// owns an open database.
|
// owns an open database. It closes the database the way scan does.
|
||||||
func brokenDatabase(t *testing.T) string {
|
func brokenDatabase(t *testing.T) string {
|
||||||
t.Helper()
|
t.Helper()
|
||||||
|
|
||||||
path := testDBPath(t)
|
path := testDBPath(t)
|
||||||
|
|
||||||
db, err := openDB(path)
|
db, err := openDB(path, scanParams)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
@@ -108,10 +102,7 @@ func brokenDatabase(t *testing.T) string {
|
|||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
|
|
||||||
err = db.Close()
|
closeScanDatabase(t.Context(), db, path)
|
||||||
if err != nil {
|
|
||||||
t.Fatal(err)
|
|
||||||
}
|
|
||||||
|
|
||||||
return path
|
return path
|
||||||
}
|
}
|
||||||
@@ -144,7 +135,10 @@ func TestOpenDatabaseKeepsWALWhileOpen(t *testing.T) {
|
|||||||
|
|
||||||
func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
|
func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
|
||||||
// Every subcommand that owns an open database must close it when
|
// Every subcommand that owns an open database must close it when
|
||||||
// it fails: no os.Exit between the open and the return.
|
// it fails: no os.Exit between the open and the return. The
|
||||||
|
// sidecar check is evidence of the close only for scan: report and
|
||||||
|
// trees only read a database that is out of WAL mode, which leaves
|
||||||
|
// nothing on disk whether they close it or not.
|
||||||
cases := map[string][]string{
|
cases := map[string][]string{
|
||||||
cmdScan: {cmdScan},
|
cmdScan: {cmdScan},
|
||||||
cmdReport: {cmdReport},
|
cmdReport: {cmdReport},
|
||||||
@@ -160,17 +154,15 @@ func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
|
|||||||
args = append(args, t.TempDir())
|
args = append(args, t.TempDir())
|
||||||
}
|
}
|
||||||
|
|
||||||
var stderr bytes.Buffer
|
var stdout, stderr bytes.Buffer
|
||||||
|
|
||||||
stdout := captureStdout(t)
|
code := run(args, &stdout, &stderr)
|
||||||
|
|
||||||
code := run(args, &stderr)
|
|
||||||
if code != exitFatal {
|
if code != exitFatal {
|
||||||
t.Errorf("run(%v) = %d, want %d", args, code, exitFatal)
|
t.Errorf("run(%v) = %d, want %d", args, code, exitFatal)
|
||||||
}
|
}
|
||||||
|
|
||||||
assertNoSidecars(t, path)
|
assertNoSidecars(t, path)
|
||||||
assertFatalOutput(t, stderr.String(), stdout())
|
assertFatalOutput(t, stderr.String(), stdout.String())
|
||||||
|
|
||||||
// Proof that the failure happened after the open: only a
|
// Proof that the failure happened after the open: only a
|
||||||
// query against the opened database can report this.
|
// query against the opened database can report this.
|
||||||
@@ -188,18 +180,16 @@ func TestRunMissingOperandIsFatalNotUsage(t *testing.T) {
|
|||||||
// must not dump the usage text.
|
// must not dump the usage text.
|
||||||
t.Setenv(databaseEnv, testDBPath(t))
|
t.Setenv(databaseEnv, testDBPath(t))
|
||||||
|
|
||||||
var stderr bytes.Buffer
|
var stdout, stderr bytes.Buffer
|
||||||
|
|
||||||
stdout := captureStdout(t)
|
|
||||||
|
|
||||||
missing := filepath.Join(t.TempDir(), "nope")
|
missing := filepath.Join(t.TempDir(), "nope")
|
||||||
|
|
||||||
code := run([]string{cmdScan, missing}, &stderr)
|
code := run([]string{cmdScan, missing}, &stdout, &stderr)
|
||||||
if code != exitFatal {
|
if code != exitFatal {
|
||||||
t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal)
|
t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal)
|
||||||
}
|
}
|
||||||
|
|
||||||
assertFatalOutput(t, stderr.String(), stdout())
|
assertFatalOutput(t, stderr.String(), stdout.String())
|
||||||
}
|
}
|
||||||
|
|
||||||
// assertFatalOutput checks that a fatal error was reported the way
|
// assertFatalOutput checks that a fatal error was reported the way
|
||||||
@@ -243,11 +233,9 @@ func TestRunUsageErrors(t *testing.T) {
|
|||||||
// path that does not exist.
|
// path that does not exist.
|
||||||
t.Setenv(databaseEnv, testDBPath(t))
|
t.Setenv(databaseEnv, testDBPath(t))
|
||||||
|
|
||||||
var stderr bytes.Buffer
|
var stdout, stderr bytes.Buffer
|
||||||
|
|
||||||
stdout := captureStdout(t)
|
code := run(tc.args, &stdout, &stderr)
|
||||||
|
|
||||||
code := run(tc.args, &stderr)
|
|
||||||
if code != exitUsage {
|
if code != exitUsage {
|
||||||
t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage)
|
t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage)
|
||||||
}
|
}
|
||||||
@@ -256,7 +244,7 @@ func TestRunUsageErrors(t *testing.T) {
|
|||||||
t.Errorf("stderr = %q, want %q", stderr.String(), tc.want)
|
t.Errorf("stderr = %q, want %q", stderr.String(), tc.want)
|
||||||
}
|
}
|
||||||
|
|
||||||
if got := stdout(); got != "" {
|
if got := stdout.String(); got != "" {
|
||||||
t.Errorf("stdout = %q, want nothing (data only)", got)
|
t.Errorf("stdout = %q, want nothing (data only)", got)
|
||||||
}
|
}
|
||||||
})
|
})
|
||||||
@@ -265,9 +253,9 @@ func TestRunUsageErrors(t *testing.T) {
|
|||||||
|
|
||||||
// TestRunHelpAndVersionSucceed checks that the two informational flags
|
// TestRunHelpAndVersionSucceed checks that the two informational flags
|
||||||
// exit 0 and keep their human-facing output on stderr.
|
// exit 0 and keep their human-facing output on stderr.
|
||||||
//
|
|
||||||
//nolint:paralleltest // captureStdout replaces the process-wide os.Stdout
|
|
||||||
func TestRunHelpAndVersionSucceed(t *testing.T) {
|
func TestRunHelpAndVersionSucceed(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
assertHumanOutput(t, "--help")
|
assertHumanOutput(t, "--help")
|
||||||
assertHumanOutput(t, "--version")
|
assertHumanOutput(t, "--version")
|
||||||
}
|
}
|
||||||
@@ -278,11 +266,9 @@ func TestRunHelpAndVersionSucceed(t *testing.T) {
|
|||||||
func assertHumanOutput(t *testing.T, arg string) {
|
func assertHumanOutput(t *testing.T, arg string) {
|
||||||
t.Helper()
|
t.Helper()
|
||||||
|
|
||||||
var stderr bytes.Buffer
|
var stdout, stderr bytes.Buffer
|
||||||
|
|
||||||
stdout := captureStdout(t)
|
code := run([]string{arg}, &stdout, &stderr)
|
||||||
|
|
||||||
code := run([]string{arg}, &stderr)
|
|
||||||
if code != exitOK {
|
if code != exitOK {
|
||||||
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
|
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
|
||||||
}
|
}
|
||||||
@@ -291,7 +277,7 @@ func assertHumanOutput(t *testing.T, arg string) {
|
|||||||
t.Errorf("run(%s) wrote nothing to stderr", arg)
|
t.Errorf("run(%s) wrote nothing to stderr", arg)
|
||||||
}
|
}
|
||||||
|
|
||||||
if got := stdout(); got != "" {
|
if got := stdout.String(); got != "" {
|
||||||
t.Errorf("stdout = %q, want nothing (data only)", got)
|
t.Errorf("stdout = %q, want nothing (data only)", got)
|
||||||
}
|
}
|
||||||
}
|
}
|
||||||
@@ -320,17 +306,15 @@ func scanFixture(t *testing.T) []string {
|
|||||||
t.Fatal(err)
|
t.Fatal(err)
|
||||||
}
|
}
|
||||||
|
|
||||||
var stderr bytes.Buffer
|
var stdout, stderr bytes.Buffer
|
||||||
|
|
||||||
stdout := captureStdout(t)
|
code := run([]string{cmdScan, dir}, &stdout, &stderr)
|
||||||
|
|
||||||
code := run([]string{cmdScan, dir}, &stderr)
|
|
||||||
if code != exitOK {
|
if code != exitOK {
|
||||||
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
|
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
|
||||||
code, exitOK, stderr.String())
|
code, exitOK, stderr.String())
|
||||||
}
|
}
|
||||||
|
|
||||||
if got := stdout(); got != "" {
|
if got := stdout.String(); got != "" {
|
||||||
t.Errorf("scan stdout = %q, want nothing (data only)", got)
|
t.Errorf("scan stdout = %q, want nothing (data only)", got)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -351,18 +335,16 @@ func TestRunReportSucceeds(t *testing.T) {
|
|||||||
|
|
||||||
dupes := scanFixture(t)
|
dupes := scanFixture(t)
|
||||||
|
|
||||||
var stderr bytes.Buffer
|
var stdout, stderr bytes.Buffer
|
||||||
|
|
||||||
stdout := captureStdout(t)
|
code := run([]string{cmdReport}, &stdout, &stderr)
|
||||||
|
|
||||||
code := run([]string{cmdReport}, &stderr)
|
|
||||||
if code != exitOK {
|
if code != exitOK {
|
||||||
t.Fatalf("run(report) = %d, want %d; stderr: %s",
|
t.Fatalf("run(report) = %d, want %d; stderr: %s",
|
||||||
code, exitOK, stderr.String())
|
code, exitOK, stderr.String())
|
||||||
}
|
}
|
||||||
|
|
||||||
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
|
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
|
||||||
if got := stdout(); got != want {
|
if got := stdout.String(); got != want {
|
||||||
t.Errorf("stdout = %q, want %q", got, want)
|
t.Errorf("stdout = %q, want %q", got, want)
|
||||||
}
|
}
|
||||||
|
|
||||||
@@ -375,11 +357,9 @@ func TestRunTreesSucceeds(t *testing.T) {
|
|||||||
|
|
||||||
dupes := scanFixture(t)
|
dupes := scanFixture(t)
|
||||||
|
|
||||||
var stderr bytes.Buffer
|
var stdout, stderr bytes.Buffer
|
||||||
|
|
||||||
stdout := captureStdout(t)
|
code := run([]string{cmdTrees}, &stdout, &stderr)
|
||||||
|
|
||||||
code := run([]string{cmdTrees}, &stderr)
|
|
||||||
if code != exitOK {
|
if code != exitOK {
|
||||||
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
|
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
|
||||||
code, exitOK, stderr.String())
|
code, exitOK, stderr.String())
|
||||||
@@ -389,9 +369,108 @@ func TestRunTreesSucceeds(t *testing.T) {
|
|||||||
// trees of each other.
|
// trees of each other.
|
||||||
want := "first\tdupe\tfiles\tsize\n" +
|
want := "first\tdupe\tfiles\tsize\n" +
|
||||||
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
|
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
|
||||||
if got := stdout(); got != want {
|
if got := stdout.String(); got != want {
|
||||||
t.Errorf("stdout = %q, want %q", got, want)
|
t.Errorf("stdout = %q, want %q", got, want)
|
||||||
}
|
}
|
||||||
|
|
||||||
assertNoSidecars(t, path)
|
assertNoSidecars(t, path)
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
|
||||||
|
// README §Database: report and trees need only read access to the
|
||||||
|
// database file. With its directory read-only as well, SQLite
|
||||||
|
// cannot create any file beside it.
|
||||||
|
path := testDBPath(t)
|
||||||
|
t.Setenv(databaseEnv, path)
|
||||||
|
|
||||||
|
dupes := scanFixture(t)
|
||||||
|
assertNoSidecars(t, path)
|
||||||
|
makeReadOnly(t, path)
|
||||||
|
|
||||||
|
cases := map[string]string{
|
||||||
|
cmdReport: "first\tdupe\tsize\n" +
|
||||||
|
dupes[0] + "\t" + dupes[1] + "\t300\n",
|
||||||
|
cmdTrees: "first\tdupe\tfiles\tsize\n" +
|
||||||
|
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) +
|
||||||
|
"\t1\t300\n",
|
||||||
|
}
|
||||||
|
|
||||||
|
for name, want := range cases {
|
||||||
|
var stdout, stderr bytes.Buffer
|
||||||
|
|
||||||
|
code := run([]string{name}, &stdout, &stderr)
|
||||||
|
if code != exitOK {
|
||||||
|
t.Errorf("run(%s) = %d, want %d; stderr: %s",
|
||||||
|
name, code, exitOK, stderr.String())
|
||||||
|
|
||||||
|
continue
|
||||||
|
}
|
||||||
|
|
||||||
|
if got := stdout.String(); got != want {
|
||||||
|
t.Errorf("%s stdout = %q, want %q", name, got, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRunStdoutClosedIsFatal(t *testing.T) {
|
||||||
|
// README §Error handling: a stdout write failure exits 1, reported
|
||||||
|
// in one line on stderr.
|
||||||
|
for _, name := range []string{cmdReport, cmdTrees} {
|
||||||
|
t.Run(name, func(t *testing.T) {
|
||||||
|
t.Setenv(databaseEnv, testDBPath(t))
|
||||||
|
|
||||||
|
scanFixture(t)
|
||||||
|
|
||||||
|
stdout, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
err = stdout.Close()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
var stderr bytes.Buffer
|
||||||
|
|
||||||
|
code := run([]string{name}, stdout, &stderr)
|
||||||
|
if code != exitFatal {
|
||||||
|
t.Errorf("run(%s) = %d, want %d", name, code, exitFatal)
|
||||||
|
}
|
||||||
|
|
||||||
|
got := stderr.String()
|
||||||
|
if !strings.HasPrefix(got, "sfdupes: write stdout: ") ||
|
||||||
|
!strings.Contains(got, os.ErrClosed.Error()) ||
|
||||||
|
strings.Count(got, "\n") != 1 {
|
||||||
|
t.Errorf("stderr = %q, want one line reporting the "+
|
||||||
|
"failed stdout write", got)
|
||||||
|
}
|
||||||
|
})
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// errWriteFailed is the error failingWriter returns.
|
||||||
|
var errWriteFailed = errors.New("write failed")
|
||||||
|
|
||||||
|
// failingWriter is a stdout that fails every write.
|
||||||
|
type failingWriter struct{}
|
||||||
|
|
||||||
|
func (failingWriter) Write([]byte) (int, error) { return 0, errWriteFailed }
|
||||||
|
|
||||||
|
func TestStdoutWriteErrorPropagates(t *testing.T) {
|
||||||
|
t.Setenv(databaseEnv, testDBPath(t))
|
||||||
|
|
||||||
|
scanFixture(t)
|
||||||
|
|
||||||
|
cases := map[string]func(context.Context, io.Writer) error{
|
||||||
|
cmdReport: runReport,
|
||||||
|
cmdTrees: runTrees,
|
||||||
|
}
|
||||||
|
|
||||||
|
for name, fn := range cases {
|
||||||
|
err := fn(t.Context(), failingWriter{})
|
||||||
|
if !errors.Is(err, errWriteFailed) {
|
||||||
|
t.Errorf("%s: error = %v, want %v", name, err, errWriteFailed)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|||||||
+3
-1
@@ -106,6 +106,8 @@ func (p *progress) increment() {
|
|||||||
}
|
}
|
||||||
|
|
||||||
// warnf prints a one-line warning to stderr without corrupting the bar.
|
// warnf prints a one-line warning to stderr without corrupting the bar.
|
||||||
|
// The whole message is escaped like a report's path columns, so a path
|
||||||
|
// holding a newline cannot split the warning.
|
||||||
func (p *progress) warnf(format string, args ...any) {
|
func (p *progress) warnf(format string, args ...any) {
|
||||||
if p == nil {
|
if p == nil {
|
||||||
return
|
return
|
||||||
@@ -115,7 +117,7 @@ func (p *progress) warnf(format string, args ...any) {
|
|||||||
_ = p.bar.Clear()
|
_ = p.bar.Clear()
|
||||||
}
|
}
|
||||||
|
|
||||||
fmt.Fprintf(os.Stderr, format+"\n", args...)
|
fmt.Fprintln(os.Stderr, escapePath(fmt.Sprintf(format, args...)))
|
||||||
}
|
}
|
||||||
|
|
||||||
// finish terminates the pass's display.
|
// finish terminates the pass's display.
|
||||||
|
|||||||
@@ -4,6 +4,7 @@ import (
|
|||||||
"bufio"
|
"bufio"
|
||||||
"context"
|
"context"
|
||||||
"fmt"
|
"fmt"
|
||||||
|
"io"
|
||||||
"os"
|
"os"
|
||||||
"slices"
|
"slices"
|
||||||
"strings"
|
"strings"
|
||||||
@@ -31,9 +32,9 @@ type scanRec struct {
|
|||||||
// loadRecords opens the database and reads every file record for the
|
// loadRecords opens the database and reads every file record for the
|
||||||
// report and trees subcommands. Any database problem — including a
|
// report and trees subcommands. Any database problem — including a
|
||||||
// missing database — is fatal. The error is returned rather than
|
// missing database — is fatal. The error is returned rather than
|
||||||
// exiting, so that the deferred close — which checkpoints the SQLite
|
// exiting, so that the deferred close always runs; the database is
|
||||||
// WAL — always runs; the database is closed before the caller formats
|
// closed before the caller formats its output, so it stays closed even
|
||||||
// its output, so it stays closed even if that output fails.
|
// if that output fails.
|
||||||
func loadRecords(ctx context.Context) ([]scanRec, error) {
|
func loadRecords(ctx context.Context) ([]scanRec, error) {
|
||||||
dbPath := databasePath()
|
dbPath := databasePath()
|
||||||
|
|
||||||
@@ -65,7 +66,7 @@ type dupeGroup struct {
|
|||||||
// from the database and prints the file-level duplicates report as TSV
|
// from the database and prints the file-level duplicates report as TSV
|
||||||
// on stdout. It never touches the scanned filesystem; its only I/O is
|
// on stdout. It never touches the scanned filesystem; its only I/O is
|
||||||
// the database, stdout, and stderr.
|
// the database, stdout, and stderr.
|
||||||
func runReport(ctx context.Context) error {
|
func runReport(ctx context.Context, stdout io.Writer) error {
|
||||||
recs, err := loadRecords(ctx)
|
recs, err := loadRecords(ctx)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return err
|
return err
|
||||||
@@ -73,7 +74,7 @@ func runReport(ctx context.Context) error {
|
|||||||
|
|
||||||
dupes := collectDupeGroups(recs)
|
dupes := collectDupeGroups(recs)
|
||||||
|
|
||||||
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
|
out := bufio.NewWriterSize(stdout, ioBufSize)
|
||||||
|
|
||||||
_, err = fmt.Fprintln(out, "first\tdupe\tsize")
|
_, err = fmt.Fprintln(out, "first\tdupe\tsize")
|
||||||
if err != nil {
|
if err != nil {
|
||||||
@@ -87,7 +88,7 @@ func runReport(ctx context.Context) error {
|
|||||||
for _, g := range dupes {
|
for _, g := range dupes {
|
||||||
for _, p := range g.paths[1:] {
|
for _, p := range g.paths[1:] {
|
||||||
_, err = fmt.Fprintf(out, "%s\t%s\t%d\n",
|
_, err = fmt.Fprintf(out, "%s\t%s\t%d\n",
|
||||||
g.paths[0], p, g.size)
|
escapePath(g.paths[0]), escapePath(p), g.size)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return fmt.Errorf("write stdout: %w", err)
|
return fmt.Errorf("write stdout: %w", err)
|
||||||
}
|
}
|
||||||
@@ -156,6 +157,21 @@ func collectDupeGroups(recs []scanRec) []dupeGroup {
|
|||||||
return dupes
|
return dupes
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// escapePath returns a path as it is written in a report column (README
|
||||||
|
// "Report output format"): a backslash, tab, newline or carriage return
|
||||||
|
// becomes \\, \t, \n or \r, and every other byte is kept as it is.
|
||||||
|
// Grouping and sorting use the raw path, never this form.
|
||||||
|
func escapePath(p string) string {
|
||||||
|
// Most paths need no escaping; skip building a replacer for them.
|
||||||
|
if !strings.ContainsAny(p, "\\\t\n\r") {
|
||||||
|
return p
|
||||||
|
}
|
||||||
|
|
||||||
|
return strings.NewReplacer(
|
||||||
|
`\`, `\\`, "\t", `\t`, "\n", `\n`, "\r", `\r`,
|
||||||
|
).Replace(p)
|
||||||
|
}
|
||||||
|
|
||||||
// humanBytes formats a byte count in human units (binary prefixes).
|
// humanBytes formats a byte count in human units (binary prefixes).
|
||||||
func humanBytes(n int64) string {
|
func humanBytes(n int64) string {
|
||||||
const unit = 1024
|
const unit = 1024
|
||||||
|
|||||||
+116
@@ -1,10 +1,126 @@
|
|||||||
package main
|
package main
|
||||||
|
|
||||||
import (
|
import (
|
||||||
|
"bytes"
|
||||||
|
"io"
|
||||||
|
"os"
|
||||||
|
"path/filepath"
|
||||||
"slices"
|
"slices"
|
||||||
"testing"
|
"testing"
|
||||||
)
|
)
|
||||||
|
|
||||||
|
// awkwardDir is a directory name holding every byte the reports escape.
|
||||||
|
const awkwardDir = "/d/\tone\ntwo\rthree\\four"
|
||||||
|
|
||||||
|
// awkwardPairRecs is a duplicate pair in sibling directories /d/A and
|
||||||
|
// awkwardDir. A raw tab sorts before "A" but its escaped form `\t`
|
||||||
|
// sorts after it, so awkwardDir coming first shows that sorting uses
|
||||||
|
// the raw path.
|
||||||
|
func awkwardPairRecs() []scanRec {
|
||||||
|
return []scanRec{
|
||||||
|
{size: 5, head: "h", tail: "t", content: "c", path: "/d/A/f"},
|
||||||
|
{size: 5, head: "h", tail: "t", content: "c", path: awkwardDir + "/f"},
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// seedDatabase writes recs into a fresh database and returns its path.
|
||||||
|
func seedDatabase(t *testing.T, recs []scanRec) string {
|
||||||
|
t.Helper()
|
||||||
|
|
||||||
|
path := testDBPath(t)
|
||||||
|
|
||||||
|
db, err := openScanDatabase(t.Context(), path)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
err = applyChanges(t.Context(), db, recs, nil, nil)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
err = db.Close()
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
return path
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRunReportEscapesPaths(t *testing.T) {
|
||||||
|
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
|
||||||
|
|
||||||
|
var stdout, stderr bytes.Buffer
|
||||||
|
|
||||||
|
code := run([]string{cmdReport}, &stdout, &stderr)
|
||||||
|
if code != exitOK {
|
||||||
|
t.Fatalf("run(report) = %d, want %d; stderr: %s",
|
||||||
|
code, exitOK, stderr.String())
|
||||||
|
}
|
||||||
|
|
||||||
|
want := "first\tdupe\tsize\n" +
|
||||||
|
`/d/\tone\ntwo\rthree\\four/f` + "\t/d/A/f\t5\n"
|
||||||
|
if got := stdout.String(); got != want {
|
||||||
|
t.Errorf("stdout = %q, want %q", got, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestEscapePath(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
cases := map[string]string{
|
||||||
|
"/srv/plain": "/srv/plain",
|
||||||
|
"/a\tb": `/a\tb`,
|
||||||
|
"/a\nb": `/a\nb`,
|
||||||
|
"/a\rb": `/a\rb`,
|
||||||
|
`/a\b`: `/a\\b`,
|
||||||
|
`/a\tb`: `/a\\tb`,
|
||||||
|
"/not-utf8\xff": "/not-utf8\xff",
|
||||||
|
}
|
||||||
|
for in, want := range cases {
|
||||||
|
if got := escapePath(in); got != want {
|
||||||
|
t.Errorf("escapePath(%q) = %q, want %q", in, got, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestWarnfEscapes checks that a warning naming a path that holds a
|
||||||
|
// newline is still one line.
|
||||||
|
//
|
||||||
|
//nolint:paralleltest // replaces the process-wide os.Stderr
|
||||||
|
func TestWarnfEscapes(t *testing.T) {
|
||||||
|
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
saved := os.Stderr
|
||||||
|
os.Stderr = f
|
||||||
|
|
||||||
|
t.Cleanup(func() {
|
||||||
|
os.Stderr = saved
|
||||||
|
|
||||||
|
_ = f.Close()
|
||||||
|
})
|
||||||
|
|
||||||
|
(&progress{}).warnf("stat %s: %s", "/d/a\nb", "gone")
|
||||||
|
|
||||||
|
_, err = f.Seek(0, io.SeekStart)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
got, err := io.ReadAll(f)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
want := `stat /d/a\nb: gone` + "\n"
|
||||||
|
if string(got) != want {
|
||||||
|
t.Errorf("warning = %q, want %q", got, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestCollectDupeGroups(t *testing.T) {
|
func TestCollectDupeGroups(t *testing.T) {
|
||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
|
|||||||
@@ -87,8 +87,9 @@ type fileMeta struct {
|
|||||||
// hash only when its size, head, and tail match another file's. Flag
|
// hash only when its size, head, and tail match another file's. Flag
|
||||||
// parsing and the at-least-one-operand check are done by cobra. Errors
|
// parsing and the at-least-one-operand check are done by cobra. Errors
|
||||||
// are returned rather than exiting, so that the deferred close — which
|
// are returned rather than exiting, so that the deferred close — which
|
||||||
// checkpoints the SQLite WAL — always runs. Cancelling ctx unwinds the
|
// takes the database out of WAL mode — always runs. Cancelling ctx
|
||||||
// worker pools and aborts the scan with the context's error.
|
// unwinds the worker pools and aborts the scan with the context's
|
||||||
|
// error.
|
||||||
func runScan(ctx context.Context, roots []string, workers int,
|
func runScan(ctx context.Context, roots []string, workers int,
|
||||||
oneFS bool,
|
oneFS bool,
|
||||||
) error {
|
) error {
|
||||||
@@ -108,7 +109,7 @@ func runScan(ctx context.Context, roots []string, workers int,
|
|||||||
return err
|
return err
|
||||||
}
|
}
|
||||||
|
|
||||||
defer func() { _ = db.Close() }()
|
defer closeScanDatabase(ctx, db, dbPath)
|
||||||
|
|
||||||
st, err := syncScan(ctx, db, roots, workers, oneFS)
|
st, err := syncScan(ctx, db, roots, workers, oneFS)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
|
|||||||
+2
-1
@@ -7,6 +7,7 @@ import (
|
|||||||
"database/sql"
|
"database/sql"
|
||||||
"encoding/hex"
|
"encoding/hex"
|
||||||
"fmt"
|
"fmt"
|
||||||
|
"io"
|
||||||
"os"
|
"os"
|
||||||
"path/filepath"
|
"path/filepath"
|
||||||
"runtime"
|
"runtime"
|
||||||
@@ -1548,7 +1549,7 @@ func TestScanHashWriteFailureUnwindsPool(t *testing.T) {
|
|||||||
|
|
||||||
code := run([]string{
|
code := run([]string{
|
||||||
cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir,
|
cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir,
|
||||||
}, &stderr)
|
}, io.Discard, &stderr)
|
||||||
if code != exitFatal {
|
if code != exitFatal {
|
||||||
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
|
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
|
||||||
code, exitFatal, stderr.String())
|
code, exitFatal, stderr.String())
|
||||||
|
|||||||
@@ -5,6 +5,7 @@ import (
|
|||||||
"context"
|
"context"
|
||||||
"crypto/sha256"
|
"crypto/sha256"
|
||||||
"fmt"
|
"fmt"
|
||||||
|
"io"
|
||||||
"os"
|
"os"
|
||||||
"slices"
|
"slices"
|
||||||
"strconv"
|
"strconv"
|
||||||
@@ -36,7 +37,7 @@ type treeNode struct {
|
|||||||
// maximal duplicate-tree groups as TSV on stdout. It never touches the
|
// maximal duplicate-tree groups as TSV on stdout. It never touches the
|
||||||
// scanned filesystem; its only I/O is the database, stdout, and
|
// scanned filesystem; its only I/O is the database, stdout, and
|
||||||
// stderr.
|
// stderr.
|
||||||
func runTrees(ctx context.Context) error {
|
func runTrees(ctx context.Context, stdout io.Writer) error {
|
||||||
recs, err := loadRecords(ctx)
|
recs, err := loadRecords(ctx)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return err
|
return err
|
||||||
@@ -47,7 +48,7 @@ func runTrees(ctx context.Context) error {
|
|||||||
|
|
||||||
dupes := collectTreeGroups(allDirs, super)
|
dupes := collectTreeGroups(allDirs, super)
|
||||||
|
|
||||||
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
|
out := bufio.NewWriterSize(stdout, ioBufSize)
|
||||||
|
|
||||||
_, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
|
_, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
|
||||||
if err != nil {
|
if err != nil {
|
||||||
@@ -62,7 +63,8 @@ func runTrees(ctx context.Context) error {
|
|||||||
first := g[0]
|
first := g[0]
|
||||||
for _, n := range g[1:] {
|
for _, n := range g[1:] {
|
||||||
_, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n",
|
_, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n",
|
||||||
first.path, n.path, first.fileCount, first.totalSize)
|
escapePath(first.path), escapePath(n.path),
|
||||||
|
first.fileCount, first.totalSize)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return fmt.Errorf("write stdout: %w", err)
|
return fmt.Errorf("write stdout: %w", err)
|
||||||
}
|
}
|
||||||
@@ -87,8 +89,8 @@ func runTrees(ctx context.Context) error {
|
|||||||
|
|
||||||
// buildHierarchy reconstructs the directory hierarchy from the record
|
// buildHierarchy reconstructs the directory hierarchy from the record
|
||||||
// paths under a synthetic super-root. Paths are split on "/"; for
|
// paths under a synthetic super-root. Paths are split on "/"; for
|
||||||
// absolute paths the first component is empty, which simply becomes a
|
// absolute paths the first component is empty, which becomes the
|
||||||
// top-level node representing "/". It returns the super-root and every
|
// top-level node with path "/". It returns the super-root and every
|
||||||
// directory node created.
|
// directory node created.
|
||||||
func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
|
func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
|
||||||
super := &treeNode{}
|
super := &treeNode{}
|
||||||
@@ -102,9 +104,17 @@ func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
|
|||||||
for _, c := range comps[:len(comps)-1] {
|
for _, c := range comps[:len(comps)-1] {
|
||||||
child := node.dirs[c]
|
child := node.dirs[c]
|
||||||
if child == nil {
|
if child == nil {
|
||||||
childPath := c
|
childPath := node.path + "/" + c
|
||||||
if node != super {
|
|
||||||
childPath = node.path + "/" + c
|
// The root directory's path is "/", not empty, and its
|
||||||
|
// children's paths start with one slash, not two.
|
||||||
|
switch {
|
||||||
|
case node == super && c == "":
|
||||||
|
childPath = "/"
|
||||||
|
case node == super:
|
||||||
|
childPath = c
|
||||||
|
case node.path == "/":
|
||||||
|
childPath = "/" + c
|
||||||
}
|
}
|
||||||
|
|
||||||
child = &treeNode{path: childPath, parent: node}
|
child = &treeNode{path: childPath, parent: node}
|
||||||
|
|||||||
@@ -1,6 +1,7 @@
|
|||||||
package main
|
package main
|
||||||
|
|
||||||
import (
|
import (
|
||||||
|
"bytes"
|
||||||
"slices"
|
"slices"
|
||||||
"testing"
|
"testing"
|
||||||
)
|
)
|
||||||
@@ -84,6 +85,44 @@ func TestBuildHierarchyCounts(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
func TestBuildHierarchyRootPath(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
// The root directory's path is "/", never empty, and its
|
||||||
|
// children's paths start with a single slash.
|
||||||
|
_, dirs := buildHierarchy([]scanRec{{path: "/f"}, {path: "/srv/g"}})
|
||||||
|
|
||||||
|
got := make([]string, 0, len(dirs))
|
||||||
|
for _, d := range dirs {
|
||||||
|
got = append(got, d.path)
|
||||||
|
}
|
||||||
|
|
||||||
|
slices.Sort(got)
|
||||||
|
|
||||||
|
want := []string{"/", "/srv"}
|
||||||
|
if !slices.Equal(got, want) {
|
||||||
|
t.Fatalf("directory paths = %q, want %q", got, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
func TestRunTreesEscapesPaths(t *testing.T) {
|
||||||
|
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
|
||||||
|
|
||||||
|
var stdout, stderr bytes.Buffer
|
||||||
|
|
||||||
|
code := run([]string{cmdTrees}, &stdout, &stderr)
|
||||||
|
if code != exitOK {
|
||||||
|
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
|
||||||
|
code, exitOK, stderr.String())
|
||||||
|
}
|
||||||
|
|
||||||
|
want := "first\tdupe\tfiles\tsize\n" +
|
||||||
|
`/d/\tone\ntwo\rthree\\four` + "\t/d/A\t1\t5\n"
|
||||||
|
if got := stdout.String(); got != want {
|
||||||
|
t.Errorf("stdout = %q, want %q", got, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
func TestTreeDigests(t *testing.T) {
|
func TestTreeDigests(t *testing.T) {
|
||||||
t.Parallel()
|
t.Parallel()
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user