5 Commits
Author SHA1 Message Date
clawbot f08f4c5b3b Print progress at once off a terminal, keep warnings out of redraws (closes #13)
check / check (push) Successful in 1m20s
When stderr is not a terminal, each phase prints its zero-state line
as it starts instead of after its first item. stderrIsTTY uses
term.IsTerminal from golang.org/x/term, now a direct dependency, so
/dev/null is no longer taken for a terminal.

A spinner keeps the library's background redraw, so its count and
elapsed time stay current while a phase waits for its next item. A
warning printed during a spinner phase goes through the bar
(progressbar.Bprintln), which prints it before its next redraw instead
of racing it. Bars with a total have no background redraw and still
print warnings directly.

Model: opus-5-5
2026-10-04 00:28:10 +00:00
clawbot 01ff3bb5f0 Test report and trees stdout write failures (closes #30)
check / check (push) Successful in 1m34s
report and trees already checked every stdout write and the final
flush. run now takes the stdout it hands to them, so tests pass a
closed file or a failing writer instead of swapping os.Stdout: a
closed stdout exits 1 with a one-line diagnostic, and the writer's
error reaches the caller.

README "Error handling" now states the two cases that never reach
sfdupes as a failed write: a pipe reader that exits early ends the
process with SIGPIPE, as with cat; and stdout closed with >&- is
replaced by /dev/null by the Go runtime before main runs, so the run
succeeds.

Model: opus-5-5
2026-10-03 18:01:27 +02:00
clawbot d63d3cc7fc Open the database read-only for report and trees (closes #8)
check / check (push) Successful in 1m17s
report and trees now connect read-only (mode=ro, query_only, the same
busy timeout) and no longer set the journal mode, which is a write. A
read-only connection to a WAL database still needs its -wal and -shm
files, or write access to the directory to create them, so scan now
switches the database back to rollback-journal mode whenever it closes
it: between scans the file alone holds the database. If a report has
the database open at that moment the switch is refused; scan warns and
the database stays in WAL mode, with its -wal and -shm files, until the
next scan. README §Database states what readers need.

Model: opus-5-5
2026-10-03 16:30:19 +02:00
clawbot c887f80f57 Escape tab, newline, CR and backslash in report paths (closes #7)
check / check (push) Successful in 1m29s
A path holding a tab or newline split a row of the report or trees
output. The path columns of both now write a backslash, tab, newline
and carriage return as \\, \t, \n and \r; every other byte is written
unchanged. Grouping and sorting still use the stored path. Warnings
on stderr are escaped the same way in warnf, so each stays one line.

In trees, the root directory's node now has the path "/" instead of
an empty string, and its children's paths start with a single slash.

README states the rule under "Report output format".

Model: opus-5-5
2026-10-03 15:30:37 +02:00
clawbot c9bf22d483 Stamp the git tag or short commit in a plain docker build (closes #67)
check / check (push) Successful in 1m1s
.dockerignore now sends .git, without .git/config, which can hold a
credential. The build stage takes the VERSION build argument when one
is given, otherwise git describe --tags --always of that .git, and
fails if the context carries .git and still yields no version. A plain
docker build . used to stamp dev. The CI checkout fetches full history
so CI sees the tag and stamps the same value as make build.

Model: opus-5-5
2026-10-02 08:49:03 +02:00
19 changed files with 791 additions and 150 deletions
+6 -1
View File
@@ -1,4 +1,9 @@
.git # .git is sent without its config. Without a VERSION build argument the
# stage that compiles runs `git describe --tags --always` on .git, which
# does not need .git/config; that file can hold a credential, such as a
# password in a remote URL or the token the CI checkout step stores there.
.git/config
.claude .claude
.DS_Store .DS_Store
sfdupes sfdupes
+2
View File
@@ -6,4 +6,6 @@ jobs:
steps: steps:
# actions/checkout v4.2.2, 2026-02-22 # actions/checkout v4.2.2, 2026-02-22
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683 - uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
with:
fetch-depth: 0
- run: script/cibuild - run: script/cibuild
+13 -1
View File
@@ -108,7 +108,19 @@ ARG CHECK_EPOCH
RUN echo "gate test, epoch ${CHECK_EPOCH}" && make test RUN echo "gate test, epoch ${CHECK_EPOCH}" && make test
RUN echo "gate fmt-check, epoch ${CHECK_EPOCH}" && make fmt-check RUN echo "gate fmt-check, epoch ${CHECK_EPOCH}" && make fmt-check
RUN make build # The version stamped into the binary: the VERSION build argument when
# one is given, otherwise `git describe --tags --always` of the .git in
# the build context (git is installed by script/bootstrap above). A
# context that carries .git and still yields no version fails the build;
# with neither, as from a source tarball, it is "dev".
ARG VERSION
RUN version="${VERSION:-$(git describe --tags --always || echo dev)}"; \
if [ -e .git ] && { [ -z "$version" ] || [ "$version" = dev ] || \
[ "$version" = unknown ]; }; then \
echo "no version could be derived although the build context carries .git" >&2; \
exit 1; \
fi; \
make build VERSION="$version"
# Runtime stage # Runtime stage
# alpine:3.22, 2026-07-23 # alpine:3.22, 2026-07-23
+57 -14
View File
@@ -113,7 +113,8 @@ Goals, in order:
`sfdupes`. `sfdupes`.
- Dependencies: standard library, `github.com/spf13/cobra` for the - Dependencies: standard library, `github.com/spf13/cobra` for the
CLI, **one progress-bar library** CLI, **one progress-bar library**
(`github.com/schollz/progressbar/v3`), and **one SQLite driver** (`github.com/schollz/progressbar/v3`), `golang.org/x/term` to tell
whether stderr is a terminal, and **one SQLite driver**
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled). (`modernc.org/sqlite`, pure Go, so builds keep cgo disabled).
`github.com/spf13/viper` is permitted if configuration-file support `github.com/spf13/viper` is permitted if configuration-file support
is ever needed, but is not currently used. No other third-party is ever needed, but is not currently used. No other third-party
@@ -152,16 +153,28 @@ All three subcommands operate on a single SQLite database file:
use. `report` and `trees` require an existing database; a missing use. `report` and `trees` require an existing database; a missing
database file is a fatal error (exit 1) telling the user to run database file is a fatal error (exit 1) telling the user to run
`scan` first. `scan` first.
- The database uses WAL journal mode and a busy timeout, so running a - While `scan` runs, the database is in WAL journal mode with a busy
report while a cron `scan` is in progress is safe. The filesystem timeout, so running a report while a cron `scan` is in progress is
is authoritative; the database is an eventually-consistent safe. The filesystem is authoritative; the database is an
reflection of it. Hashed records are committed in batched eventually-consistent reflection of it. Hashed records are
transactions while the scan is still running (keeping the WAL committed in batched transactions while the scan is still running
small and letting concurrent reports observe progress), so a (keeping the WAL small and letting concurrent reports observe
report may see a scan's changes partially applied, and a scan progress), so a report may see a scan's changes partially applied,
that dies partway leaves a valid database holding everything and a scan that dies partway leaves a valid database holding
hashed so far; the next scan skips those records and converges everything hashed so far; the next scan skips those records and
toward the filesystem. converges toward the filesystem.
- `scan` switches the database back to rollback-journal mode when it
closes it, so between scans the database file alone holds the whole
database. Each switch needs the database to itself: a `scan` that
starts while a report is still reading waits for it up to the
10-second busy timeout, then fails; a `scan` that ends while a
report has the database open warns and leaves the database in WAL
mode until the next scan.
- `report` and `trees` open the database read-only and need only read
access to the database file, and no write access to its directory.
While the database is in WAL mode they also read the `-wal` and
`-shm` files beside it, which SQLite creates with the database
file's permissions.
- Schema (`PRAGMA user_version` is the schema version, currently 1; a - Schema (`PRAGMA user_version` is the schema version, currently 1; a
database with any other version is a fatal error): database with any other version is a fatal error):
@@ -428,6 +441,15 @@ first dupe size
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296 /srv/a/big.iso /srv/c/big-copy2.iso 4294967296
``` ```
Paths are raw bytes and may hold any byte except NUL, so the path
columns (`first` and `dupe`) are escaped to keep every row one line of
tab-separated fields: a backslash is written as `\\`, a tab as `\t`, a
newline as `\n`, and a carriage return as `\r`. Every other byte is
written unchanged, including bytes that are not valid UTF-8. Undoing
those four escapes gives back the stored path. Grouping and ordering
use the stored path, not the escaped one. The warnings `scan` prints on
stderr are escaped the same way, so each warning is one line.
Summary to stderr: records read, number of duplicate groups, number of Summary to stderr: records read, number of duplicate groups, number of
dupe files, and total reclaimable bytes (sum of `size` over all dupe dupe files, and total reclaimable bytes (sum of `size` over all dupe
rows) in human units. rows) in human units.
@@ -499,6 +521,9 @@ first dupe files size
/srv/a/project /srv/backup/project 3417 104857600 /srv/a/project /srv/backup/project 3417 104857600
``` ```
The `first` and `dupe` paths are escaped as described under "Report
output format". The root directory's path is `/`.
Summary to stderr: records read, number of duplicate-tree groups, Summary to stderr: records read, number of duplicate-tree groups,
number of dupe trees, and total reclaimable bytes (sum of `size` over number of dupe trees, and total reclaimable bytes (sum of `size` over
all dupe rows) in human units. all dupe rows) in human units.
@@ -533,10 +558,16 @@ hash: [12345/98765] 12% |████ | 92 files/s elapsed 2:32 eta 17:54
Additional requirements: Additional requirements:
- When stderr is not a TTY, do not emit ANSI redraws: print a plain - When stderr is not a terminal (a pipe, a file, `/dev/null`), do not
one-line progress update no more often than every 5 seconds instead. emit ANSI redraws: print a plain one-line progress update the moment
each phase starts, then no more often than every 5 seconds.
- Progress updates are driven from the main goroutine and must be - Progress updates are driven from the main goroutine and must be
non-blocking with respect to the worker pool. non-blocking with respect to the worker pool. On a terminal the
spinner-style displays also redraw on their own several times a
second, so their count and elapsed time stay current while a phase
waits for its next item.
- A warning printed during a phase always lands on a line of its own,
never inside the progress display.
- `report` and `trees` modes need no progress display, only their - `report` and `trees` modes need no progress display, only their
stderr summaries. stderr summaries.
@@ -549,6 +580,18 @@ Additional requirements:
- `2`: usage error (including `scan` with no `PATH` operand and - `2`: usage error (including `scan` with no `PATH` operand and
`report`/`trees` with any positional argument). `report`/`trees` with any positional argument).
A stdout write failure, such as a full disk, is reported in one line on
stderr and exits 1. Two cases never reach sfdupes as a failed write:
- When the reader of a stdout pipe exits early, as in
`sfdupes report | head`, the next write ends sfdupes with `SIGPIPE`,
quietly and without a summary, the way `cat` or `sort` end. The
shell reports the signal (status 141 in most shells), not exit 1.
- When stdout is closed outright (`sfdupes report >&-`), the Go
runtime opens `/dev/null` in its place before sfdupes starts, so
the output is discarded and the run succeeds, as with
`> /dev/null`.
## Entrypoints ## Entrypoints
This repository adheres to the This repository adheres to the
+27
View File
@@ -29,6 +29,33 @@
# Completed Steps # Completed Steps
- progress prints at once on a non-terminal, uses a real terminal test,
and prints warnings through a spinner instead of racing its redraw
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/13)
- test stdout write failures in `report` and `trees`; README states that
`| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null`
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/30)
- `report` and `trees` open the database read-only, and `scan` leaves it
out of WAL mode, so reading needs only read access (2026-10-03, closes
https://git.eeqj.de/sneak/sfdupes/issues/8)
- escape tabs, newlines, carriage returns and backslashes in report,
trees and warning paths; the root directory's path is `/`
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/7)
- stamp the git tag or short commit in a plain `docker build .`
instead of `dev` (2026-10-02, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/67): `.dockerignore` now
sends `.git`, without `.git/config`, and the `Dockerfile` build
stage takes the `VERSION` build argument when one is given,
otherwise `git describe --tags --always` of that `.git`. The build
fails if the context carries `.git` and the version still comes out
empty, `dev` or `unknown`. The CI checkout step fetches the full
history (`fetch-depth: 0`) so CI sees the tag and stamps the same
value as `make build`.
- replace the 1 KiB end-window sampling with the head/tail plus - replace the 1 KiB end-window sampling with the head/tail plus
content-hash ladder (2026-09-22, branch `next`, closes content-hash ladder (2026-09-22, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/61): a file under 10 MiB is https://git.eeqj.de/sneak/sfdupes/issues/61): a file under 10 MiB is
+37 -11
View File
@@ -74,16 +74,24 @@ func databasePath() string {
return defaultDatabasePath return defaultDatabasePath
} }
// openDB opens the SQLite database at path with WAL journaling and a // scanParams are the connection parameters for scan: read-write, with
// busy timeout, so a report can run while a cron scan is in progress. // WAL journaling and a busy timeout, so a report can run while a cron
// It does not create or verify the schema. // scan is in progress. closeScanDatabase leaves WAL mode again.
func openDB(path string) (*sql.DB, error) { const scanParams = "_pragma=busy_timeout(10000)" +
dsn := "file:" + path + "&_pragma=journal_mode(WAL)" +
"?_pragma=busy_timeout(10000)" + "&_pragma=synchronous(NORMAL)"
"&_pragma=journal_mode(WAL)" +
"&_pragma=synchronous(NORMAL)"
db, err := sql.Open("sqlite", dsn) // reportParams are the connection parameters for report and trees:
// read-only, with the same busy timeout. They set no journal mode,
// because setting one is a write.
const reportParams = "mode=ro" +
"&_pragma=busy_timeout(10000)" +
"&_pragma=query_only(1)"
// openDB opens the SQLite database at path with the connection
// parameters params. It does not create or verify the schema.
func openDB(path, params string) (*sql.DB, error) {
db, err := sql.Open("sqlite", "file:"+path+"?"+params)
if err != nil { if err != nil {
return nil, fmt.Errorf("open database %s: %w", path, err) return nil, fmt.Errorf("open database %s: %w", path, err)
} }
@@ -104,7 +112,7 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return nil, fmt.Errorf("create database directory: %w", err) return nil, fmt.Errorf("create database directory: %w", err)
} }
db, err := openDB(path) db, err := openDB(path, scanParams)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -119,6 +127,24 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return db, nil return db, nil
} }
// closeScanDatabase switches the database at path from WAL back to
// rollback-journal mode and closes it. Out of WAL mode the database
// file alone holds the whole database, so a reader needs no -wal or
// -shm file beside it, nor write access to create them. The switch
// fails while a report has the database open; the database then stays
// in WAL mode, still readable, until a later scan closes it.
func closeScanDatabase(ctx context.Context, db *sql.DB, path string) {
// Runs on the way out of a cancelled scan too.
_, err := db.ExecContext(context.WithoutCancel(ctx),
"PRAGMA journal_mode = DELETE")
if err != nil {
fmt.Fprintf(os.Stderr, "scan: database %s left in WAL mode: %v\n",
path, err)
}
_ = db.Close()
}
// openReportDatabase opens an existing database for the report and // openReportDatabase opens an existing database for the report and
// trees subcommands. A missing database file is an error directing the // trees subcommands. A missing database file is an error directing the
// user to run scan first; the schema version must match exactly. // user to run scan first; the schema version must match exactly.
@@ -134,7 +160,7 @@ func openReportDatabase(ctx context.Context,
return nil, fmt.Errorf("database: %w", err) return nil, fmt.Errorf("database: %w", err)
} }
db, err := openDB(path) db, err := openDB(path, reportParams)
if err != nil { if err != nil {
return nil, err return nil, err
} }
+44
View File
@@ -5,6 +5,7 @@ import (
"database/sql" "database/sql"
"errors" "errors"
"fmt" "fmt"
"os"
"path/filepath" "path/filepath"
"slices" "slices"
"strings" "strings"
@@ -130,6 +131,49 @@ func TestOpenReportDatabaseOK(t *testing.T) {
_ = db.Close() _ = db.Close()
} }
func TestCloseScanDatabaseWhileReportOpen(t *testing.T) {
t.Parallel()
// A report holding the database open stops scan from taking it out
// of WAL mode. The -wal and -shm files must then stay beside it, so
// that a later report still needs only read access.
path := testDBPath(t)
scanDB, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
reportDB, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
closeScanDatabase(t.Context(), scanDB, path)
_ = reportDB.Close()
_, err = os.Stat(path + "-wal")
if err != nil {
t.Fatalf("no -wal left: the switch out of WAL mode was not "+
"stopped: %v", err)
}
makeReadOnly(t, path)
reportDB, err = openReportDatabase(t.Context(), path)
if err != nil {
t.Fatalf("openReportDatabase: %v", err)
}
defer func() { _ = reportDB.Close() }()
_, err = loadFileRows(t.Context(), reportDB)
if err != nil {
t.Fatalf("loadFileRows: %v", err)
}
}
func TestApplyChangesRoundTrip(t *testing.T) { func TestApplyChangesRoundTrip(t *testing.T) {
t.Parallel() t.Parallel()
+1 -1
View File
@@ -5,6 +5,7 @@ go 1.25.7
require ( require (
github.com/schollz/progressbar/v3 v3.19.1 github.com/schollz/progressbar/v3 v3.19.1
github.com/spf13/cobra v1.10.2 github.com/spf13/cobra v1.10.2
golang.org/x/term v0.44.0
modernc.org/sqlite v1.54.0 modernc.org/sqlite v1.54.0
) )
@@ -19,7 +20,6 @@ require (
github.com/rivo/uniseg v0.4.7 // indirect github.com/rivo/uniseg v0.4.7 // indirect
github.com/spf13/pflag v1.0.9 // indirect github.com/spf13/pflag v1.0.9 // indirect
golang.org/x/sys v0.46.0 // indirect golang.org/x/sys v0.46.0 // indirect
golang.org/x/term v0.44.0 // indirect
modernc.org/libc v1.74.1 // indirect modernc.org/libc v1.74.1 // indirect
modernc.org/mathutil v1.7.1 // indirect modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect modernc.org/memory v1.11.0 // indirect
+12 -7
View File
@@ -57,22 +57,27 @@ var errNoSubcommand = errors.New("no subcommand")
var Version = "dev" var Version = "dev"
func main() { func main() {
os.Exit(run(os.Args[1:], os.Stderr)) // Once the reader of a stdout pipe has gone, as in "sfdupes report |
// head", the Go runtime ends the process with SIGPIPE on the next
// write instead of returning an error (README "Error handling").
// Registering for SIGPIPE with os/signal would change that.
os.Exit(run(os.Args[1:], os.Stdout, os.Stderr))
} }
// run executes args against the command tree and returns the process // run executes args against the command tree and returns the process
// exit code. It is the program's single exit point: the subcommands // exit code. It is the program's single exit point: the subcommands
// return their errors instead of exiting, so every deferred cleanup — // return their errors instead of exiting, so every deferred cleanup —
// above all closing the database, which checkpoints the SQLite WAL — // above all closing the database, which checkpoints the SQLite WAL —
// runs before the process ends. // runs before the process ends. The report and trees subcommands write
func run(args []string, stderr io.Writer) int { // their data to stdout.
func run(args []string, stdout, stderr io.Writer) int {
// A nil slice makes cobra fall back to os.Args, which would let a // A nil slice makes cobra fall back to os.Args, which would let a
// test binary's own flags reach the command tree. // test binary's own flags reach the command tree.
if args == nil { if args == nil {
args = []string{} args = []string{}
} }
root := newRootCommand(stderr) root := newRootCommand(stdout, stderr)
root.SetArgs(args) root.SetArgs(args)
err := root.Execute() err := root.Execute()
@@ -98,7 +103,7 @@ func run(args []string, stderr io.Writer) int {
// newRootCommand builds the command tree. Everything on stdout is // newRootCommand builds the command tree. Everything on stdout is
// machine-readable data; all human-facing output (help, usage, errors) // machine-readable data; all human-facing output (help, usage, errors)
// goes to stderr. // goes to stderr.
func newRootCommand(stderr io.Writer) *cobra.Command { func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
root := &cobra.Command{ root := &cobra.Command{
Use: "sfdupes", Use: "sfdupes",
Short: "Find candidate duplicate files by size and head/tail/content SHA-256", Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
@@ -140,7 +145,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the file-level duplicates report", Short: "Read the scan database and print the file-level duplicates report",
Args: cobra.NoArgs, Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error { RunE: runE(func(ctx context.Context, _ []string) error {
return runReport(ctx) return runReport(ctx, stdout)
}), }),
} }
@@ -149,7 +154,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the duplicate-tree report", Short: "Read the scan database and print the duplicate-tree report",
Args: cobra.NoArgs, Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error { RunE: runE(func(ctx context.Context, _ []string) error {
return runTrees(ctx) return runTrees(ctx, stdout)
}), }),
} }
+156 -77
View File
@@ -44,60 +44,54 @@ func assertNoSidecars(t *testing.T, path string) {
} }
} }
// captureStdout redirects os.Stdout to a file for the rest of the test // makeReadOnly takes write permission away from the database at path,
// and returns a function reading back everything written to it. Only // from any WAL sidecar beside it, and from their directory, as for a
// machine-readable data belongs on stdout (README design goal 4), so // user reading a database that a root cron scan keeps. Root ignores
// the tests assert on it directly. // file permissions, so it skips the test when run as root.
func captureStdout(t *testing.T) func() string { func makeReadOnly(t *testing.T, path string) {
t.Helper() t.Helper()
f, err := os.Create(filepath.Join(t.TempDir(), "stdout")) if os.Geteuid() == 0 {
t.Skip("root ignores file permissions")
}
err := os.Chmod(path, 0o400)
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
saved := os.Stdout for _, suffix := range walSuffixes {
os.Stdout = f err = os.Chmod(path+suffix, 0o400)
if err != nil && !errors.Is(err, fs.ErrNotExist) {
t.Cleanup(func() {
os.Stdout = saved
_ = f.Close()
})
return func() string {
// Read what has been written without disturbing the write
// offset, so the capture can be inspected more than once.
size, err := f.Seek(0, io.SeekCurrent)
if err != nil {
t.Fatal(err) t.Fatal(err)
} }
if size == 0 {
return ""
}
b := make([]byte, size)
_, err = f.ReadAt(b, 0)
if err != nil {
t.Fatal(err)
}
return string(b)
} }
dir := filepath.Dir(path)
//nolint:gosec // reaching the database needs the search bit
err = os.Chmod(dir, 0o500)
if err != nil {
t.Fatal(err)
}
// Runs before t.TempDir's own cleanup, which must delete the files.
t.Cleanup(func() {
//nolint:gosec // removing the directory needs its search bit back
_ = os.Chmod(dir, 0o700)
})
} }
// brokenDatabase writes a database that opens cleanly and passes the // brokenDatabase writes a database that opens cleanly and passes the
// schema-version check but has no files table, so the first query // schema-version check but has no files table, so the first query
// fails with the database already open: a fatal error on a path that // fails with the database already open: a fatal error on a path that
// owns an open database. // owns an open database. It closes the database the way scan does.
func brokenDatabase(t *testing.T) string { func brokenDatabase(t *testing.T) string {
t.Helper() t.Helper()
path := testDBPath(t) path := testDBPath(t)
db, err := openDB(path) db, err := openDB(path, scanParams)
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -108,10 +102,7 @@ func brokenDatabase(t *testing.T) string {
t.Fatal(err) t.Fatal(err)
} }
err = db.Close() closeScanDatabase(t.Context(), db, path)
if err != nil {
t.Fatal(err)
}
return path return path
} }
@@ -144,7 +135,10 @@ func TestOpenDatabaseKeepsWALWhileOpen(t *testing.T) {
func TestRunFatalAfterOpenClosesDatabase(t *testing.T) { func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
// Every subcommand that owns an open database must close it when // Every subcommand that owns an open database must close it when
// it fails: no os.Exit between the open and the return. // it fails: no os.Exit between the open and the return. The
// sidecar check is evidence of the close only for scan: report and
// trees only read a database that is out of WAL mode, which leaves
// nothing on disk whether they close it or not.
cases := map[string][]string{ cases := map[string][]string{
cmdScan: {cmdScan}, cmdScan: {cmdScan},
cmdReport: {cmdReport}, cmdReport: {cmdReport},
@@ -160,17 +154,15 @@ func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
args = append(args, t.TempDir()) args = append(args, t.TempDir())
} }
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run(args, &stdout, &stderr)
code := run(args, &stderr)
if code != exitFatal { if code != exitFatal {
t.Errorf("run(%v) = %d, want %d", args, code, exitFatal) t.Errorf("run(%v) = %d, want %d", args, code, exitFatal)
} }
assertNoSidecars(t, path) assertNoSidecars(t, path)
assertFatalOutput(t, stderr.String(), stdout()) assertFatalOutput(t, stderr.String(), stdout.String())
// Proof that the failure happened after the open: only a // Proof that the failure happened after the open: only a
// query against the opened database can report this. // query against the opened database can report this.
@@ -188,18 +180,16 @@ func TestRunMissingOperandIsFatalNotUsage(t *testing.T) {
// must not dump the usage text. // must not dump the usage text.
t.Setenv(databaseEnv, testDBPath(t)) t.Setenv(databaseEnv, testDBPath(t))
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t)
missing := filepath.Join(t.TempDir(), "nope") missing := filepath.Join(t.TempDir(), "nope")
code := run([]string{cmdScan, missing}, &stderr) code := run([]string{cmdScan, missing}, &stdout, &stderr)
if code != exitFatal { if code != exitFatal {
t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal) t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal)
} }
assertFatalOutput(t, stderr.String(), stdout()) assertFatalOutput(t, stderr.String(), stdout.String())
} }
// assertFatalOutput checks that a fatal error was reported the way // assertFatalOutput checks that a fatal error was reported the way
@@ -243,11 +233,9 @@ func TestRunUsageErrors(t *testing.T) {
// path that does not exist. // path that does not exist.
t.Setenv(databaseEnv, testDBPath(t)) t.Setenv(databaseEnv, testDBPath(t))
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run(tc.args, &stdout, &stderr)
code := run(tc.args, &stderr)
if code != exitUsage { if code != exitUsage {
t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage) t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage)
} }
@@ -256,7 +244,7 @@ func TestRunUsageErrors(t *testing.T) {
t.Errorf("stderr = %q, want %q", stderr.String(), tc.want) t.Errorf("stderr = %q, want %q", stderr.String(), tc.want)
} }
if got := stdout(); got != "" { if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got) t.Errorf("stdout = %q, want nothing (data only)", got)
} }
}) })
@@ -265,9 +253,9 @@ func TestRunUsageErrors(t *testing.T) {
// TestRunHelpAndVersionSucceed checks that the two informational flags // TestRunHelpAndVersionSucceed checks that the two informational flags
// exit 0 and keep their human-facing output on stderr. // exit 0 and keep their human-facing output on stderr.
//
//nolint:paralleltest // captureStdout replaces the process-wide os.Stdout
func TestRunHelpAndVersionSucceed(t *testing.T) { func TestRunHelpAndVersionSucceed(t *testing.T) {
t.Parallel()
assertHumanOutput(t, "--help") assertHumanOutput(t, "--help")
assertHumanOutput(t, "--version") assertHumanOutput(t, "--version")
} }
@@ -278,11 +266,9 @@ func TestRunHelpAndVersionSucceed(t *testing.T) {
func assertHumanOutput(t *testing.T, arg string) { func assertHumanOutput(t *testing.T, arg string) {
t.Helper() t.Helper()
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run([]string{arg}, &stdout, &stderr)
code := run([]string{arg}, &stderr)
if code != exitOK { if code != exitOK {
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK) t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
} }
@@ -291,7 +277,7 @@ func assertHumanOutput(t *testing.T, arg string) {
t.Errorf("run(%s) wrote nothing to stderr", arg) t.Errorf("run(%s) wrote nothing to stderr", arg)
} }
if got := stdout(); got != "" { if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got) t.Errorf("stdout = %q, want nothing (data only)", got)
} }
} }
@@ -320,17 +306,15 @@ func scanFixture(t *testing.T) []string {
t.Fatal(err) t.Fatal(err)
} }
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run([]string{cmdScan, dir}, &stdout, &stderr)
code := run([]string{cmdScan, dir}, &stderr)
if code != exitOK { if code != exitOK {
t.Fatalf("run(scan) = %d, want %d; stderr: %s", t.Fatalf("run(scan) = %d, want %d; stderr: %s",
code, exitOK, stderr.String()) code, exitOK, stderr.String())
} }
if got := stdout(); got != "" { if got := stdout.String(); got != "" {
t.Errorf("scan stdout = %q, want nothing (data only)", got) t.Errorf("scan stdout = %q, want nothing (data only)", got)
} }
@@ -351,18 +335,16 @@ func TestRunReportSucceeds(t *testing.T) {
dupes := scanFixture(t) dupes := scanFixture(t)
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run([]string{cmdReport}, &stdout, &stderr)
code := run([]string{cmdReport}, &stderr)
if code != exitOK { if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s", t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String()) code, exitOK, stderr.String())
} }
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n" want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
if got := stdout(); got != want { if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want) t.Errorf("stdout = %q, want %q", got, want)
} }
@@ -375,11 +357,9 @@ func TestRunTreesSucceeds(t *testing.T) {
dupes := scanFixture(t) dupes := scanFixture(t)
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run([]string{cmdTrees}, &stdout, &stderr)
code := run([]string{cmdTrees}, &stderr)
if code != exitOK { if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s", t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String()) code, exitOK, stderr.String())
@@ -389,9 +369,108 @@ func TestRunTreesSucceeds(t *testing.T) {
// trees of each other. // trees of each other.
want := "first\tdupe\tfiles\tsize\n" + want := "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n" filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
if got := stdout(); got != want { if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want) t.Errorf("stdout = %q, want %q", got, want)
} }
assertNoSidecars(t, path) assertNoSidecars(t, path)
} }
func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
// README §Database: report and trees need only read access to the
// database file. With its directory read-only as well, SQLite
// cannot create any file beside it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dupes := scanFixture(t)
assertNoSidecars(t, path)
makeReadOnly(t, path)
cases := map[string]string{
cmdReport: "first\tdupe\tsize\n" +
dupes[0] + "\t" + dupes[1] + "\t300\n",
cmdTrees: "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) +
"\t1\t300\n",
}
for name, want := range cases {
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
continue
}
if got := stdout.String(); got != want {
t.Errorf("%s stdout = %q, want %q", name, got, want)
}
}
}
func TestRunStdoutClosedIsFatal(t *testing.T) {
// README §Error handling: a stdout write failure exits 1, reported
// in one line on stderr.
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
stdout, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
if err != nil {
t.Fatal(err)
}
err = stdout.Close()
if err != nil {
t.Fatal(err)
}
var stderr bytes.Buffer
code := run([]string{name}, stdout, &stderr)
if code != exitFatal {
t.Errorf("run(%s) = %d, want %d", name, code, exitFatal)
}
got := stderr.String()
if !strings.HasPrefix(got, "sfdupes: write stdout: ") ||
!strings.Contains(got, os.ErrClosed.Error()) ||
strings.Count(got, "\n") != 1 {
t.Errorf("stderr = %q, want one line reporting the "+
"failed stdout write", got)
}
})
}
}
// errWriteFailed is the error failingWriter returns.
var errWriteFailed = errors.New("write failed")
// failingWriter is a stdout that fails every write.
type failingWriter struct{}
func (failingWriter) Write([]byte) (int, error) { return 0, errWriteFailed }
func TestStdoutWriteErrorPropagates(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
cases := map[string]func(context.Context, io.Writer) error{
cmdReport: runReport,
cmdTrees: runTrees,
}
for name, fn := range cases {
err := fn(t.Context(), failingWriter{})
if !errors.Is(err, errWriteFailed) {
t.Errorf("%s: error = %v, want %v", name, err, errWriteFailed)
}
}
}
+39 -16
View File
@@ -6,6 +6,7 @@ import (
"time" "time"
"github.com/schollz/progressbar/v3" "github.com/schollz/progressbar/v3"
"golang.org/x/term"
) )
// plainInterval is the minimum time between progress lines when stderr // plainInterval is the minimum time between progress lines when stderr
@@ -24,24 +25,23 @@ const percentScale = 100
// stderrIsTTY reports whether stderr is attached to a terminal. // stderrIsTTY reports whether stderr is attached to a terminal.
func stderrIsTTY() bool { func stderrIsTTY() bool {
fi, err := os.Stderr.Stat() return term.IsTerminal(int(os.Stderr.Fd()))
if err != nil {
return false
}
return fi.Mode()&os.ModeCharDevice != 0
} }
// progress renders one scan pass's progress on stderr. On a TTY it // progress renders one scan pass's progress on stderr. On a TTY it
// delegates to the progressbar library (spinner style when the total is // delegates to the progressbar library (spinner style when the total is
// unknown, full bar with count/percent/rate/elapsed/ETA otherwise). When // unknown, full bar with count/percent/rate/elapsed/ETA otherwise). When
// stderr is not a TTY it emits no ANSI redraws: it prints a plain // stderr is not a TTY it emits no ANSI redraws: it prints a plain
// one-line update no more often than every plainInterval. // one-line update as the pass starts, then no more often than every
// plainInterval.
// //
// All methods must be called from the main goroutine only. A nil // All methods must be called from the main goroutine only. On a TTY
// *progress is a valid no-display receiver: every method is a no-op, // the library also redraws a spinner from its own goroutine, several
// so batched database flushes during the streaming pass can reuse the // times a second, so its count and elapsed time stay current while a
// update-pass helpers without rendering anything. // pass waits for its next item. A nil *progress is a valid
// no-display receiver: every method is a no-op, so batched database
// flushes during the streaming pass can reuse the update-pass helpers
// without rendering anything.
type progress struct { type progress struct {
label string label string
total int64 // -1 when unknown (walk pass) total int64 // -1 when unknown (walk pass)
@@ -53,10 +53,22 @@ type progress struct {
func newProgress(label string, total int64) *progress { func newProgress(label string, total int64) *progress {
p := &progress{label: label, total: total, start: time.Now()} p := &progress{label: label, total: total, start: time.Now()}
if !stderrIsTTY() { if stderrIsTTY() {
p.bar = newBar(label, total)
return p return p
} }
// Print the zero state at once: the first item may take minutes,
// and a pass must never look hung.
p.last = p.start
fmt.Fprintln(os.Stderr, p.plainLine())
return p
}
// newBar builds the TTY display for newProgress.
func newBar(label string, total int64) *progressbar.ProgressBar {
opts := []progressbar.Option{ opts := []progressbar.Option{
progressbar.OptionSetWriter(os.Stderr), progressbar.OptionSetWriter(os.Stderr),
progressbar.OptionSetDescription(label), progressbar.OptionSetDescription(label),
@@ -81,9 +93,7 @@ func newProgress(label string, total int64) *progress {
) )
} }
p.bar = progressbar.NewOptions64(total, opts...) return progressbar.NewOptions64(total, opts...)
return p
} }
// increment records one completed item and refreshes the display. // increment records one completed item and refreshes the display.
@@ -106,16 +116,29 @@ func (p *progress) increment() {
} }
// warnf prints a one-line warning to stderr without corrupting the bar. // warnf prints a one-line warning to stderr without corrupting the bar.
// The whole message is escaped like a report's path columns, so a path
// holding a newline cannot split the warning.
func (p *progress) warnf(format string, args ...any) { func (p *progress) warnf(format string, args ...any) {
if p == nil { if p == nil {
return return
} }
msg := escapePath(fmt.Sprintf(format, args...))
if p.bar != nil && p.total < 0 {
// The library also redraws a spinner from its own goroutine, so
// a direct write could land inside a redraw. The bar prints the
// warning itself, just before its next redraw.
_, _ = progressbar.Bprintln(p.bar, msg)
return
}
if p.bar != nil { if p.bar != nil {
_ = p.bar.Clear() _ = p.bar.Clear()
} }
fmt.Fprintf(os.Stderr, format+"\n", args...) fmt.Fprintln(os.Stderr, msg)
} }
// finish terminates the pass's display. // finish terminates the pass's display.
+192
View File
@@ -0,0 +1,192 @@
package main
import (
"os"
"path/filepath"
"strings"
"testing"
"time"
)
// spinnerIdle comfortably outlasts the 100ms interval at which the
// progressbar library redraws a spinner from its own goroutine.
const spinnerIdle = 500 * time.Millisecond
// captureStderr points os.Stderr at a file for the rest of the test and
// returns a function reading back everything written to it.
func captureStderr(t *testing.T) func() string {
t.Helper()
path := filepath.Join(t.TempDir(), "stderr")
f, err := os.Create(path) //nolint:gosec // test-controlled path
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
os.Stderr = f
t.Cleanup(func() {
os.Stderr = saved
_ = f.Close()
})
return func() string {
b, err := os.ReadFile(path) //nolint:gosec // test-controlled path
if err != nil {
t.Fatal(err)
}
return string(b)
}
}
//nolint:paralleltest // replaces the process-wide os.Stderr
func TestStderrIsTTYFalseForNonTerminals(t *testing.T) {
r, pipe, err := os.Pipe()
if err != nil {
t.Fatal(err)
}
regular, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
devNull, err := os.OpenFile(os.DevNull, os.O_WRONLY, 0)
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
t.Cleanup(func() {
os.Stderr = saved
for _, f := range []*os.File{r, pipe, regular, devNull} {
_ = f.Close()
}
})
cases := map[string]*os.File{
"a pipe": pipe,
"a regular file": regular,
os.DevNull: devNull,
}
for name, f := range cases {
os.Stderr = f
if stderrIsTTY() {
t.Errorf("stderrIsTTY() = true with stderr on %s", name)
}
}
}
// TestNewProgressPrintsBeforeFirstItem checks that each pass shows its
// zero state the moment it starts when stderr is not a terminal, and
// that the next line still waits for plainInterval.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestNewProgressPrintsBeforeFirstItem(t *testing.T) {
stderr := captureStderr(t)
newProgress("walk", -1).increment()
newProgress("hash", 10).increment()
want := "walk: 0 files, elapsed 0s\n" +
"hash: [0/10] 0% 0 files/s elapsed 0s eta ?\n"
if got := stderr(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
}
// newWalkSpinner returns the walk pass's terminal display, writing to
// os.Stderr whether or not it is a terminal, and stops the library's
// redraws when the test ends.
func newWalkSpinner(t *testing.T) *progress {
t.Helper()
p := &progress{
label: "walk", total: -1, start: time.Now(),
bar: newBar("walk", -1),
}
t.Cleanup(p.finish)
return p
}
// TestProgressWarningsOnOwnLines drives the terminal display of the walk
// pass through a run of warnings with no items between them, as when the
// walk meets many unreadable paths, for several of the spinner's
// redraws: every warning must land on a line of its own, never inside a
// redraw.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestProgressWarningsOnOwnLines(t *testing.T) {
stderr := captureStderr(t)
p := newWalkSpinner(t)
// No pause between warnings: one written straight to stderr is
// garbled only if a redraw lands while it is being written.
issued := 0
for start := time.Now(); time.Since(start) < spinnerIdle; issued++ {
p.warnf("warning")
}
// The spinner prints the warnings at its next redraw.
time.Sleep(spinnerIdle)
// A terminal shows each line as the text after its last carriage
// return.
shown := 0
for line := range strings.SplitSeq(stderr(), "\n") {
if !strings.Contains(line, "warning") {
continue
}
shown++
if text := line[strings.LastIndex(line, "\r")+1:]; text != "warning" {
t.Errorf("terminal shows %q, want %q", text, "warning")
}
}
if shown != issued {
t.Errorf("%d warning lines, want %d", shown, issued)
}
}
// TestSpinnerShowsCountAfterBurst checks that once a burst of items
// faster than the redraw limit is over, the walk display shows every
// item completed while it waits for the next one.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestSpinnerShowsCountAfterBurst(t *testing.T) {
stderr := captureStderr(t)
p := newWalkSpinner(t)
for range 50 {
p.increment()
}
time.Sleep(spinnerIdle)
// A terminal shows the last frame drawn. The library starts each
// frame with a carriage return and erases the previous one with
// spaces first.
var shown string
for frame := range strings.SplitSeq(stderr(), "\r") {
if strings.TrimSpace(frame) != "" {
shown = frame
}
}
if !strings.Contains(shown, "(50/-,") {
t.Errorf("terminal shows %q, want a count of 50", shown)
}
}
+22 -6
View File
@@ -4,6 +4,7 @@ import (
"bufio" "bufio"
"context" "context"
"fmt" "fmt"
"io"
"os" "os"
"slices" "slices"
"strings" "strings"
@@ -31,9 +32,9 @@ type scanRec struct {
// loadRecords opens the database and reads every file record for the // loadRecords opens the database and reads every file record for the
// report and trees subcommands. Any database problem — including a // report and trees subcommands. Any database problem — including a
// missing database — is fatal. The error is returned rather than // missing database — is fatal. The error is returned rather than
// exiting, so that the deferred close — which checkpoints the SQLite // exiting, so that the deferred close always runs; the database is
// WAL — always runs; the database is closed before the caller formats // closed before the caller formats its output, so it stays closed even
// its output, so it stays closed even if that output fails. // if that output fails.
func loadRecords(ctx context.Context) ([]scanRec, error) { func loadRecords(ctx context.Context) ([]scanRec, error) {
dbPath := databasePath() dbPath := databasePath()
@@ -65,7 +66,7 @@ type dupeGroup struct {
// from the database and prints the file-level duplicates report as TSV // from the database and prints the file-level duplicates report as TSV
// on stdout. It never touches the scanned filesystem; its only I/O is // on stdout. It never touches the scanned filesystem; its only I/O is
// the database, stdout, and stderr. // the database, stdout, and stderr.
func runReport(ctx context.Context) error { func runReport(ctx context.Context, stdout io.Writer) error {
recs, err := loadRecords(ctx) recs, err := loadRecords(ctx)
if err != nil { if err != nil {
return err return err
@@ -73,7 +74,7 @@ func runReport(ctx context.Context) error {
dupes := collectDupeGroups(recs) dupes := collectDupeGroups(recs)
out := bufio.NewWriterSize(os.Stdout, ioBufSize) out := bufio.NewWriterSize(stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tsize") _, err = fmt.Fprintln(out, "first\tdupe\tsize")
if err != nil { if err != nil {
@@ -87,7 +88,7 @@ func runReport(ctx context.Context) error {
for _, g := range dupes { for _, g := range dupes {
for _, p := range g.paths[1:] { for _, p := range g.paths[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\n", _, err = fmt.Fprintf(out, "%s\t%s\t%d\n",
g.paths[0], p, g.size) escapePath(g.paths[0]), escapePath(p), g.size)
if err != nil { if err != nil {
return fmt.Errorf("write stdout: %w", err) return fmt.Errorf("write stdout: %w", err)
} }
@@ -156,6 +157,21 @@ func collectDupeGroups(recs []scanRec) []dupeGroup {
return dupes return dupes
} }
// escapePath returns a path as it is written in a report column (README
// "Report output format"): a backslash, tab, newline or carriage return
// becomes \\, \t, \n or \r, and every other byte is kept as it is.
// Grouping and sorting use the raw path, never this form.
func escapePath(p string) string {
// Most paths need no escaping; skip building a replacer for them.
if !strings.ContainsAny(p, "\\\t\n\r") {
return p
}
return strings.NewReplacer(
`\`, `\\`, "\t", `\t`, "\n", `\n`, "\r", `\r`,
).Replace(p)
}
// humanBytes formats a byte count in human units (binary prefixes). // humanBytes formats a byte count in human units (binary prefixes).
func humanBytes(n int64) string { func humanBytes(n int64) string {
const unit = 1024 const unit = 1024
+116
View File
@@ -1,10 +1,126 @@
package main package main
import ( import (
"bytes"
"io"
"os"
"path/filepath"
"slices" "slices"
"testing" "testing"
) )
// awkwardDir is a directory name holding every byte the reports escape.
const awkwardDir = "/d/\tone\ntwo\rthree\\four"
// awkwardPairRecs is a duplicate pair in sibling directories /d/A and
// awkwardDir. A raw tab sorts before "A" but its escaped form `\t`
// sorts after it, so awkwardDir coming first shows that sorting uses
// the raw path.
func awkwardPairRecs() []scanRec {
return []scanRec{
{size: 5, head: "h", tail: "t", content: "c", path: "/d/A/f"},
{size: 5, head: "h", tail: "t", content: "c", path: awkwardDir + "/f"},
}
}
// seedDatabase writes recs into a fresh database and returns its path.
func seedDatabase(t *testing.T, recs []scanRec) string {
t.Helper()
path := testDBPath(t)
db, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
err = applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
err = db.Close()
if err != nil {
t.Fatal(err)
}
return path
}
func TestRunReportEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
code := run([]string{cmdReport}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tsize\n" +
`/d/\tone\ntwo\rthree\\four/f` + "\t/d/A/f\t5\n"
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
func TestEscapePath(t *testing.T) {
t.Parallel()
cases := map[string]string{
"/srv/plain": "/srv/plain",
"/a\tb": `/a\tb`,
"/a\nb": `/a\nb`,
"/a\rb": `/a\rb`,
`/a\b`: `/a\\b`,
`/a\tb`: `/a\\tb`,
"/not-utf8\xff": "/not-utf8\xff",
}
for in, want := range cases {
if got := escapePath(in); got != want {
t.Errorf("escapePath(%q) = %q, want %q", in, got, want)
}
}
}
// TestWarnfEscapes checks that a warning naming a path that holds a
// newline is still one line.
//
//nolint:paralleltest // replaces the process-wide os.Stderr
func TestWarnfEscapes(t *testing.T) {
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
os.Stderr = f
t.Cleanup(func() {
os.Stderr = saved
_ = f.Close()
})
(&progress{}).warnf("stat %s: %s", "/d/a\nb", "gone")
_, err = f.Seek(0, io.SeekStart)
if err != nil {
t.Fatal(err)
}
got, err := io.ReadAll(f)
if err != nil {
t.Fatal(err)
}
want := `stat /d/a\nb: gone` + "\n"
if string(got) != want {
t.Errorf("warning = %q, want %q", got, want)
}
}
func TestCollectDupeGroups(t *testing.T) { func TestCollectDupeGroups(t *testing.T) {
t.Parallel() t.Parallel()
+4 -3
View File
@@ -87,8 +87,9 @@ type fileMeta struct {
// hash only when its size, head, and tail match another file's. Flag // hash only when its size, head, and tail match another file's. Flag
// parsing and the at-least-one-operand check are done by cobra. Errors // parsing and the at-least-one-operand check are done by cobra. Errors
// are returned rather than exiting, so that the deferred close — which // are returned rather than exiting, so that the deferred close — which
// checkpoints the SQLite WAL — always runs. Cancelling ctx unwinds the // takes the database out of WAL mode — always runs. Cancelling ctx
// worker pools and aborts the scan with the context's error. // unwinds the worker pools and aborts the scan with the context's
// error.
func runScan(ctx context.Context, roots []string, workers int, func runScan(ctx context.Context, roots []string, workers int,
oneFS bool, oneFS bool,
) error { ) error {
@@ -108,7 +109,7 @@ func runScan(ctx context.Context, roots []string, workers int,
return err return err
} }
defer func() { _ = db.Close() }() defer closeScanDatabase(ctx, db, dbPath)
st, err := syncScan(ctx, db, roots, workers, oneFS) st, err := syncScan(ctx, db, roots, workers, oneFS)
if err != nil { if err != nil {
+2 -1
View File
@@ -7,6 +7,7 @@ import (
"database/sql" "database/sql"
"encoding/hex" "encoding/hex"
"fmt" "fmt"
"io"
"os" "os"
"path/filepath" "path/filepath"
"runtime" "runtime"
@@ -1548,7 +1549,7 @@ func TestScanHashWriteFailureUnwindsPool(t *testing.T) {
code := run([]string{ code := run([]string{
cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir, cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir,
}, &stderr) }, io.Discard, &stderr)
if code != exitFatal { if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d; stderr: %s", t.Fatalf("run(scan) = %d, want %d; stderr: %s",
code, exitFatal, stderr.String()) code, exitFatal, stderr.String())
+4 -4
View File
@@ -16,10 +16,10 @@
# implies the repo is green. # implies the repo is green.
# #
# That implication holds only because of CHECK_EPOCH. A COPY layer is # That implication holds only because of CHECK_EPOCH. A COPY layer is
# invalidated by changed content, and a merge commit's tree is # invalidated only by changed content, and a rebuild of an unchanged
# byte-identical to the branch head it merges, so without a fresh value # checkout sends the same content, so without a fresh value here Docker
# here Docker serves the gate layers from cache and the build reports a # serves the gate layers from cache and the build reports a green it
# green it never earned. Passing the current epoch invalidates the gate # never earned. Passing the current epoch invalidates the gate
# layers on every run while leaving the pinned base images and # layers on every run while leaving the pinned base images and
# go mod download cached; see the Dockerfile for the placement. # go mod download cached; see the Dockerfile for the placement.
set -eu set -eu
+18 -8
View File
@@ -5,6 +5,7 @@ import (
"context" "context"
"crypto/sha256" "crypto/sha256"
"fmt" "fmt"
"io"
"os" "os"
"slices" "slices"
"strconv" "strconv"
@@ -36,7 +37,7 @@ type treeNode struct {
// maximal duplicate-tree groups as TSV on stdout. It never touches the // maximal duplicate-tree groups as TSV on stdout. It never touches the
// scanned filesystem; its only I/O is the database, stdout, and // scanned filesystem; its only I/O is the database, stdout, and
// stderr. // stderr.
func runTrees(ctx context.Context) error { func runTrees(ctx context.Context, stdout io.Writer) error {
recs, err := loadRecords(ctx) recs, err := loadRecords(ctx)
if err != nil { if err != nil {
return err return err
@@ -47,7 +48,7 @@ func runTrees(ctx context.Context) error {
dupes := collectTreeGroups(allDirs, super) dupes := collectTreeGroups(allDirs, super)
out := bufio.NewWriterSize(os.Stdout, ioBufSize) out := bufio.NewWriterSize(stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize") _, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
if err != nil { if err != nil {
@@ -62,7 +63,8 @@ func runTrees(ctx context.Context) error {
first := g[0] first := g[0]
for _, n := range g[1:] { for _, n := range g[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n", _, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n",
first.path, n.path, first.fileCount, first.totalSize) escapePath(first.path), escapePath(n.path),
first.fileCount, first.totalSize)
if err != nil { if err != nil {
return fmt.Errorf("write stdout: %w", err) return fmt.Errorf("write stdout: %w", err)
} }
@@ -87,8 +89,8 @@ func runTrees(ctx context.Context) error {
// buildHierarchy reconstructs the directory hierarchy from the record // buildHierarchy reconstructs the directory hierarchy from the record
// paths under a synthetic super-root. Paths are split on "/"; for // paths under a synthetic super-root. Paths are split on "/"; for
// absolute paths the first component is empty, which simply becomes a // absolute paths the first component is empty, which becomes the
// top-level node representing "/". It returns the super-root and every // top-level node with path "/". It returns the super-root and every
// directory node created. // directory node created.
func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) { func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
super := &treeNode{} super := &treeNode{}
@@ -102,9 +104,17 @@ func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
for _, c := range comps[:len(comps)-1] { for _, c := range comps[:len(comps)-1] {
child := node.dirs[c] child := node.dirs[c]
if child == nil { if child == nil {
childPath := c childPath := node.path + "/" + c
if node != super {
childPath = node.path + "/" + c // The root directory's path is "/", not empty, and its
// children's paths start with one slash, not two.
switch {
case node == super && c == "":
childPath = "/"
case node == super:
childPath = c
case node.path == "/":
childPath = "/" + c
} }
child = &treeNode{path: childPath, parent: node} child = &treeNode{path: childPath, parent: node}
+39
View File
@@ -1,6 +1,7 @@
package main package main
import ( import (
"bytes"
"slices" "slices"
"testing" "testing"
) )
@@ -84,6 +85,44 @@ func TestBuildHierarchyCounts(t *testing.T) {
} }
} }
func TestBuildHierarchyRootPath(t *testing.T) {
t.Parallel()
// The root directory's path is "/", never empty, and its
// children's paths start with a single slash.
_, dirs := buildHierarchy([]scanRec{{path: "/f"}, {path: "/srv/g"}})
got := make([]string, 0, len(dirs))
for _, d := range dirs {
got = append(got, d.path)
}
slices.Sort(got)
want := []string{"/", "/srv"}
if !slices.Equal(got, want) {
t.Fatalf("directory paths = %q, want %q", got, want)
}
}
func TestRunTreesEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
code := run([]string{cmdTrees}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tfiles\tsize\n" +
`/d/\tone\ntwo\rthree\\four` + "\t/d/A\t1\t5\n"
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
func TestTreeDigests(t *testing.T) { func TestTreeDigests(t *testing.T) {
t.Parallel() t.Parallel()