1 Commits
Author SHA1 Message Date
sneak bd8d41b174 Stream report and trees instead of loading every record (closes #14)
check / check (push) Successful in 1m49s
report now has SQLite group the records and put the rows in report
order, helped by a new files_signature index on (size, head, tail,
content), and writes each row as it reads it. trees reads the records
in path order, where all the paths under a directory come together, so
it computes each directory's digest as soon as the stream leaves it and
keeps only its path, parent, digest and totals. Output is unchanged.

The tests that called the removed in-memory grouping functions now group
records stored in a database. A new test checks that both commands give
the same output whatever order the records were inserted in.

Model: opus-5-5
2026-10-04 02:04:37 +00:00
6 changed files with 33 additions and 257 deletions
+12 -23
View File
@@ -95,14 +95,13 @@ Goals, in order:
(ZFS pool under resilver). Holding one small record (path, size,
mtime) per file in memory during a scan is acceptable; holding
every file's hashes is not (they stay in the database). The
reporting commands do not hold every file's hashes either: `report`
lets SQLite group and order the records and writes each row as it
reads it, so its memory does not grow with the database, and
`trees` reads the records in path order and keeps each directory's
path, digest and totals, plus the hashes of only the files in the
directories holding the record being read, so its memory grows with
the number of directories and with the size of the largest
directory.
reporting commands hold no file's hashes either: `report` lets
SQLite group and order the records and writes each row as it reads
it, so its memory does not grow with the database, and `trees`
reads the records in path order and keeps each directory's path,
digest and totals, plus the files of the directories holding the
record being read, so its memory grows with the number of
directories, not files.
3. **Scan incrementally, analyze offline.** The expensive filesystem
scan maintains a persistent database; an unchanged file is never
read again on a rescan, except to compute its content hash once a
@@ -121,8 +120,7 @@ Goals, in order:
`sfdupes`.
- Dependencies: standard library, `github.com/spf13/cobra` for the
CLI, **one progress-bar library**
(`github.com/schollz/progressbar/v3`), `golang.org/x/term` to tell
whether stderr is a terminal, **one SQLite driver**
(`github.com/schollz/progressbar/v3`), **one SQLite driver**
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled), and
`golang.org/x/sys` for `flock(2)` (the scan lock, see "Database").
`github.com/spf13/viper` is permitted if configuration-file support
@@ -192,10 +190,7 @@ All three subcommands operate on a single SQLite database file:
starts while a report is still reading waits for it up to the
10-second busy timeout, then fails; a `scan` that ends while a
report has the database open warns and leaves the database in WAL
mode until the next scan. `report` writes each row as it reads it,
so it is still reading while its output is paused (a pager, a
stalled pipe), and a `scan` started then fails after the busy
timeout.
mode until the next scan.
- `report` and `trees` open the database read-only and need only read
access to the database file, and no write access to its directory.
While the database is in WAL mode they also read the `-wal` and
@@ -606,16 +601,10 @@ hash: [12345/98765] 12% |████ | 92 files/s elapsed 2:32 eta 17:54
Additional requirements:
- When stderr is not a terminal (a pipe, a file, `/dev/null`), do not
emit ANSI redraws: print a plain one-line progress update the moment
each phase starts, then no more often than every 5 seconds.
- When stderr is not a TTY, do not emit ANSI redraws: print a plain
one-line progress update no more often than every 5 seconds instead.
- Progress updates are driven from the main goroutine and must be
non-blocking with respect to the worker pool. On a terminal the
spinner-style displays also redraw on their own several times a
second, so their count and elapsed time stay current while a phase
waits for its next item.
- A warning printed during a phase always lands on a line of its own,
never inside the progress display.
non-blocking with respect to the worker pool.
- `report` and `trees` modes need no progress display, only their
stderr summaries.
-4
View File
@@ -33,10 +33,6 @@
memory; the schema gains the `files_signature` index (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/14)
- progress prints at once on a non-terminal, uses a real terminal test,
and prints warnings through a spinner instead of racing its redraw
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/13)
- warn about and skip symlink, socket, FIFO, device and `.zfs`
operands, keeping the records beneath them (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/9)
+1 -1
View File
@@ -6,7 +6,6 @@ require (
github.com/schollz/progressbar/v3 v3.19.1
github.com/spf13/cobra v1.10.2
golang.org/x/sys v0.46.0
golang.org/x/term v0.44.0
modernc.org/sqlite v1.54.0
)
@@ -20,6 +19,7 @@ require (
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/rivo/uniseg v0.4.7 // indirect
github.com/spf13/pflag v1.0.9 // indirect
golang.org/x/term v0.44.0 // indirect
modernc.org/libc v1.74.1 // indirect
modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect
+16 -37
View File
@@ -6,7 +6,6 @@ import (
"time"
"github.com/schollz/progressbar/v3"
"golang.org/x/term"
)
// plainInterval is the minimum time between progress lines when stderr
@@ -25,23 +24,24 @@ const percentScale = 100
// stderrIsTTY reports whether stderr is attached to a terminal.
func stderrIsTTY() bool {
return term.IsTerminal(int(os.Stderr.Fd()))
fi, err := os.Stderr.Stat()
if err != nil {
return false
}
return fi.Mode()&os.ModeCharDevice != 0
}
// progress renders one scan pass's progress on stderr. On a TTY it
// delegates to the progressbar library (spinner style when the total is
// unknown, full bar with count/percent/rate/elapsed/ETA otherwise). When
// stderr is not a TTY it emits no ANSI redraws: it prints a plain
// one-line update as the pass starts, then no more often than every
// plainInterval.
// one-line update no more often than every plainInterval.
//
// All methods must be called from the main goroutine only. On a TTY
// the library also redraws a spinner from its own goroutine, several
// times a second, so its count and elapsed time stay current while a
// pass waits for its next item. A nil *progress is a valid
// no-display receiver: every method is a no-op, so batched database
// flushes during the streaming pass can reuse the update-pass helpers
// without rendering anything.
// All methods must be called from the main goroutine only. A nil
// *progress is a valid no-display receiver: every method is a no-op,
// so batched database flushes during the streaming pass can reuse the
// update-pass helpers without rendering anything.
type progress struct {
label string
total int64 // -1 when unknown (walk pass)
@@ -53,22 +53,10 @@ type progress struct {
func newProgress(label string, total int64) *progress {
p := &progress{label: label, total: total, start: time.Now()}
if stderrIsTTY() {
p.bar = newBar(label, total)
if !stderrIsTTY() {
return p
}
// Print the zero state at once: the first item may take minutes,
// and a pass must never look hung.
p.last = p.start
fmt.Fprintln(os.Stderr, p.plainLine())
return p
}
// newBar builds the TTY display for newProgress.
func newBar(label string, total int64) *progressbar.ProgressBar {
opts := []progressbar.Option{
progressbar.OptionSetWriter(os.Stderr),
progressbar.OptionSetDescription(label),
@@ -93,7 +81,9 @@ func newBar(label string, total int64) *progressbar.ProgressBar {
)
}
return progressbar.NewOptions64(total, opts...)
p.bar = progressbar.NewOptions64(total, opts...)
return p
}
// increment records one completed item and refreshes the display.
@@ -123,22 +113,11 @@ func (p *progress) warnf(format string, args ...any) {
return
}
msg := escapePath(fmt.Sprintf(format, args...))
if p.bar != nil && p.total < 0 {
// The library also redraws a spinner from its own goroutine, so
// a direct write could land inside a redraw. The bar prints the
// warning itself, just before its next redraw.
_, _ = progressbar.Bprintln(p.bar, msg)
return
}
if p.bar != nil {
_ = p.bar.Clear()
}
fmt.Fprintln(os.Stderr, msg)
fmt.Fprintln(os.Stderr, escapePath(fmt.Sprintf(format, args...)))
}
// finish terminates the pass's display.
-161
View File
@@ -1,161 +0,0 @@
package main
import (
"os"
"path/filepath"
"strings"
"testing"
"time"
)
// spinnerIdle comfortably outlasts the 100ms interval at which the
// progressbar library redraws a spinner from its own goroutine.
const spinnerIdle = 500 * time.Millisecond
//nolint:paralleltest // replaces the process-wide os.Stderr
func TestStderrIsTTYFalseForNonTerminals(t *testing.T) {
r, pipe, err := os.Pipe()
if err != nil {
t.Fatal(err)
}
regular, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
devNull, err := os.OpenFile(os.DevNull, os.O_WRONLY, 0)
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
t.Cleanup(func() {
os.Stderr = saved
for _, f := range []*os.File{r, pipe, regular, devNull} {
_ = f.Close()
}
})
cases := map[string]*os.File{
"a pipe": pipe,
"a regular file": regular,
os.DevNull: devNull,
}
for name, f := range cases {
os.Stderr = f
if stderrIsTTY() {
t.Errorf("stderrIsTTY() = true with stderr on %s", name)
}
}
}
// TestNewProgressPrintsBeforeFirstItem checks that each pass shows its
// zero state the moment it starts when stderr is not a terminal, and
// that the next line still waits for plainInterval.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestNewProgressPrintsBeforeFirstItem(t *testing.T) {
stderr := captureStderr(t)
newProgress("walk", -1).increment()
newProgress("hash", 10).increment()
want := "walk: 0 files, elapsed 0s\n" +
"hash: [0/10] 0% 0 files/s elapsed 0s eta ?\n"
if got := stderr(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
}
// newWalkSpinner returns the walk pass's terminal display, writing to
// os.Stderr whether or not it is a terminal, and stops the library's
// redraws when the test ends.
func newWalkSpinner(t *testing.T) *progress {
t.Helper()
p := &progress{
label: "walk", total: -1, start: time.Now(),
bar: newBar("walk", -1),
}
t.Cleanup(p.finish)
return p
}
// TestProgressWarningsOnOwnLines drives the terminal display of the walk
// pass through a run of warnings with no items between them, as when the
// walk meets many unreadable paths, for several of the spinner's
// redraws: every warning must land on a line of its own, never inside a
// redraw.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestProgressWarningsOnOwnLines(t *testing.T) {
stderr := captureStderr(t)
p := newWalkSpinner(t)
// No pause between warnings: one written straight to stderr is
// garbled only if a redraw lands while it is being written.
issued := 0
for start := time.Now(); time.Since(start) < spinnerIdle; issued++ {
p.warnf("warning")
}
// The spinner prints the warnings at its next redraw.
time.Sleep(spinnerIdle)
// A terminal shows each line as the text after its last carriage
// return.
shown := 0
for line := range strings.SplitSeq(stderr(), "\n") {
if !strings.Contains(line, "warning") {
continue
}
shown++
if text := line[strings.LastIndex(line, "\r")+1:]; text != "warning" {
t.Errorf("terminal shows %q, want %q", text, "warning")
}
}
if shown != issued {
t.Errorf("%d warning lines, want %d", shown, issued)
}
}
// TestSpinnerShowsCountAfterBurst checks that once a burst of items
// faster than the redraw limit is over, the walk display shows every
// item completed while it waits for the next one.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestSpinnerShowsCountAfterBurst(t *testing.T) {
stderr := captureStderr(t)
p := newWalkSpinner(t)
for range 50 {
p.increment()
}
time.Sleep(spinnerIdle)
// A terminal shows the last frame drawn. The library starts each
// frame with a carriage return and erases the previous one with
// spaces first.
var shown string
for frame := range strings.SplitSeq(stderr(), "\r") {
if strings.TrimSpace(frame) != "" {
shown = frame
}
}
if !strings.Contains(shown, "(50/-,") {
t.Errorf("terminal shows %q, want a count of 50", shown)
}
}
+4 -31
View File
@@ -3,8 +3,6 @@ package main
import (
"bytes"
"database/sql"
"errors"
"fmt"
"io"
"os"
"path/filepath"
@@ -116,29 +114,6 @@ func TestRunReportEscapesPaths(t *testing.T) {
}
}
func TestReportStdoutFailsWhileReading(t *testing.T) {
// Each row holds two paths longer than dir, so the report is more
// than twice the stdout buffer and stdout fails while rows are
// still being read, not at the final flush.
dir := "/" + strings.Repeat("d", 4096)
var recs []scanRec
for i := range ioBufSize / len(dir) {
recs = append(recs, scanRec{
size: 1, head: "h", tail: "t", content: "c",
path: fmt.Sprintf("%s/%d", dir, i),
})
}
t.Setenv(databaseEnv, seedDatabase(t, recs))
err := runReport(t.Context(), failingWriter{})
if !errors.Is(err, errWriteFailed) ||
!strings.HasPrefix(err.Error(), "write stdout: ") {
t.Errorf("error = %v, want write stdout: %v", err, errWriteFailed)
}
}
func TestRunReportsIgnoreInsertionOrder(t *testing.T) {
// README §Constraints: identical database contents give identical
// output, whatever order the records were inserted in.
@@ -326,13 +301,11 @@ func TestDupeGroupsMtimeExcluded(t *testing.T) {
func TestDupeGroupsTieBreak(t *testing.T) {
t.Parallel()
// The hashes sort opposite to the first paths, so ordering the
// groups by hash instead of by first path fails this test.
recs := []scanRec{
{size: 50, head: "a", tail: "a", content: "a", path: "/beta/2"},
{size: 50, head: "a", tail: "a", content: "a", path: "/beta/1"},
{size: 50, head: "b", tail: "b", content: "b", path: "/alpha/2"},
{size: 50, head: "b", tail: "b", content: "b", path: "/alpha/1"},
{size: 50, head: "b", tail: "b", content: "b", path: "/beta/2"},
{size: 50, head: "b", tail: "b", content: "b", path: "/beta/1"},
{size: 50, head: "a", tail: "a", content: "a", path: "/alpha/2"},
{size: 50, head: "a", tail: "a", content: "a", path: "/alpha/1"},
}
groups := dupeGroupsOf(t, recs)