Author SHA1 Message Date
clawbot 705c8729ca Hold a lock so a second scan fails at once (closes #53)
check / check (push) Waiting to run
scan takes an exclusive flock(2) on a lock file beside the database
(its path with .lock appended) before it walks anything or opens the
database, and holds it until it returns. A second scan against the
same database fails at once with a one-line error naming the lock
file and exits 1. report and trees never take the lock. The lock ends
with the process, so a fatal error or an interrupt releases it; the
file is never deleted. golang.org/x/sys becomes a direct dependency.

The README smoke test now keeps the database outside the scanned
tree, where its empty lock file would have joined the empty-file
group.

Model: opus-5-5
2026-10-04 02:30:21 +02:00
7 changed files with 217 additions and 293 deletions
+30 -31
View File
@@ -113,8 +113,9 @@ Goals, in order:
`sfdupes`. `sfdupes`.
- Dependencies: standard library, `github.com/spf13/cobra` for the - Dependencies: standard library, `github.com/spf13/cobra` for the
CLI, **one progress-bar library** CLI, **one progress-bar library**
(`github.com/schollz/progressbar/v3`), and **one SQLite driver** (`github.com/schollz/progressbar/v3`), **one SQLite driver**
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled). (`modernc.org/sqlite`, pure Go, so builds keep cgo disabled), and
`golang.org/x/sys` for `flock(2)` (the scan lock, see "Database").
`github.com/spf13/viper` is permitted if configuration-file support `github.com/spf13/viper` is permitted if configuration-file support
is ever needed, but is not currently used. No other third-party is ever needed, but is not currently used. No other third-party
deps. deps.
@@ -152,6 +153,20 @@ All three subcommands operate on a single SQLite database file:
use. `report` and `trees` require an existing database; a missing use. `report` and `trees` require an existing database; a missing
database file is a fatal error (exit 1) telling the user to run database file is a fatal error (exit 1) telling the user to run
`scan` first. `scan` first.
- Only one `scan` runs against a database at a time. For its whole
run, `scan` holds an exclusive `flock(2)` lock on a lock file
beside the database, named by appending `.lock` to the database
path (`/var/lib/sfdupes/db.sqlite.lock` by default), taken before
it walks the filesystem or opens the database. A second `scan`
against the same database does not wait: it fails at once with a
one-line error naming the lock file and exits 1, without walking
anything or opening the database, and the running scan carries on.
The lock file is created on first use, open to its owner only, and
left in place: a leftover file blocks nothing, because the lock
ends with the process holding it however it ends, a fatal error or
an interrupt included, and deleting the file while a scan runs
would let a second scan start. `report` and `trees` never take the
lock, so they run during a scan.
- While `scan` runs, the database is in WAL journal mode with a busy - While `scan` runs, the database is in WAL journal mode with a busy
timeout, so running a report while a cron `scan` is in progress is timeout, so running a report while a cron `scan` is in progress is
safe. The filesystem is authoritative; the database is an safe. The filesystem is authoritative; the database is an
@@ -263,18 +278,6 @@ duplicates another or lies under another is dropped before walking,
so every file is reached exactly once and produces one database so every file is reached exactly once and produces one database
record. record.
An operand that is a symlink (never followed, not even as an operand),
socket, FIFO, or device node, or a directory named `.zfs`, is not
scanned. `scan` prints a one-line warning naming the path and what it
is, counts it as skipped, and drops it from the scanned operands before
reading the database. Another operand beneath it is still scanned. The
records stored beneath it are not deleted: they are treated like any
other record outside the scanned operands, including the content-phase
exception below. If it lies under another operand, they are under that
operand instead, and are deleted like any other record there that this
scan did not verify. This is not an error: a scan whose every operand
is dropped walks nothing and exits 0.
`scan` synchronizes the database with the filesystem state under the `scan` synchronizes the database with the filesystem state under the
scanned operands: scanned operands:
@@ -380,12 +383,9 @@ during the hash phase:
Rules for the walk: Rules for the walk:
- Only regular files. Skip directories, symlinks (do not follow, - Only regular files. Skip directories, symlinks (do not follow,
including symlink operands), sockets, FIFOs, and device nodes. An including symlink operands), sockets, FIFOs, and device nodes.
operand that is a symlink, socket, FIFO, or device node is dropped
as described in "`scan` mode" above.
- Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs; - Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs;
walking them would list every file once per snapshot), not even walking them would list every file once per snapshot).
when it is an operand; such an operand is dropped the same way.
- Filesystem boundaries are crossed by default. With `-x` - Filesystem boundaries are crossed by default. With `-x`
(long form `--one-file-system`, following the GNU `du`/`rsync` (long form `--one-file-system`, following the GNU `du`/`rsync`
convention), never descend into a directory on a different convention), never descend into a directory on a different
@@ -396,11 +396,9 @@ Rules for the walk:
path, and continue. Per-file errors never abort the run; the final path, and continue. Per-file errors never abort the run; the final
summary reports how many were skipped. As specified above, a summary reports how many were skipped. As specified above, a
skipped path that has a database record from an earlier scan loses skipped path that has a database record from an earlier scan loses
that record, unless it failed only in the content phase, or is an that record, unless it failed only in the content phase; an
operand dropped before the database was read that lies under no unreadable directory subtree likewise loses its records (accepted:
other operand; an unreadable directory subtree likewise loses its the database mirrors what the latest scan could actually verify).
records (accepted: the database mirrors what the latest scan could
actually verify).
Concurrency: the walk phase (which also stats files), the hash phase, Concurrency: the walk phase (which also stats files), the hash phase,
and the content phase each use a worker pool of `--workers` workers and the content phase each use a worker pool of `--workers` workers
@@ -584,9 +582,10 @@ Additional requirements:
### Error handling and exit codes ### Error handling and exit codes
- `0`: success, even if individual files were skipped with warnings. - `0`: success, even if individual files were skipped with warnings.
- `1`: fatal error (e.g., a `PATH` operand does not exist, the - `1`: fatal error (e.g., a `PATH` operand does not exist, another
database cannot be created/opened/read/written, a missing database `scan` is already running against the same database, the database
for `report`/`trees`, stdout write failure). cannot be created/opened/read/written, a missing database for
`report`/`trees`, stdout write failure).
- `2`: usage error (including `scan` with no `PATH` operand and - `2`: usage error (including `scan` with no `PATH` operand and
`report`/`trees` with any positional argument). `report`/`trees` with any positional argument).
@@ -739,7 +738,7 @@ All of the following, run in this directory, must pass:
```sh ```sh
d=$(mktemp -d) d=$(mktemp -d)
export SFDUPES_DATABASE="$d/db.sqlite" export SFDUPES_DATABASE="$(mktemp -d)/db.sqlite"
mkdir -p "$d/a" "$d/b" mkdir -p "$d/a" "$d/b"
head -c 2000 /dev/urandom > "$d/a/one.bin" head -c 2000 /dev/urandom > "$d/a/one.bin"
cp "$d/a/one.bin" "$d/b/copy.bin" cp "$d/a/one.bin" "$d/b/copy.bin"
@@ -767,9 +766,9 @@ All of the following, run in this directory, must pass:
./sfdupes report ./sfdupes report
``` ```
(The scan database lives inside `$d` here purely for test hygiene; (The database lives in a temp directory of its own: inside `$d`,
scanning `$d` therefore also records the SQLite file itself, which the scan would record it, and its empty lock file would join the
is harmless.) `empty1`/`empty2` group.)
Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin` Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin`
form one group (two dupe rows, `first` is the lexicographically form one group (two dupe rows, `first` is the lexicographically
+3 -3
View File
@@ -29,9 +29,9 @@
# Completed Steps # Completed Steps
- warn about and skip symlink, socket, FIFO, device and `.zfs` - `scan` holds a lock on a lock file beside the database for its whole run,
operands, keeping the records beneath them (2026-10-03, so a second `scan` fails at once with exit 1 (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/9) https://git.eeqj.de/sneak/sfdupes/issues/53)
- test stdout write failures in `report` and `trees`; README states that - test stdout write failures in `report` and `trees`; README states that
`| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null` `| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null`
+47
View File
@@ -11,6 +11,7 @@ import (
"slices" "slices"
"strconv" "strconv"
"golang.org/x/sys/unix"
// The pure-Go SQLite driver, registered as "sqlite"; keeps cgo // The pure-Go SQLite driver, registered as "sqlite"; keeps cgo
// disabled. // disabled.
_ "modernc.org/sqlite" _ "modernc.org/sqlite"
@@ -32,6 +33,11 @@ const schemaVersion = 1
// scan. // scan.
const dbDirPerm = 0o755 const dbDirPerm = 0o755
// lockFilePerm is the mode for the scan lock file. Anyone who can open
// the file can hold the lock and keep every scan from running, so it
// is open to its owner only.
const lockFilePerm = 0o600
// createTableSQL is the schema applied to a fresh database. Paths are // createTableSQL is the schema applied to a fresh database. Paths are
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8. // BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
const createTableSQL = ` const createTableSQL = `
@@ -64,6 +70,10 @@ var errNoDatabase = errors.New(
// does not understand. // does not understand.
var errSchemaVersion = errors.New("unsupported database schema version") var errSchemaVersion = errors.New("unsupported database schema version")
// errScanRunning reports that another scan holds the lock on the
// database.
var errScanRunning = errors.New("another scan is running")
// databasePath resolves the database location: SFDUPES_DATABASE when // databasePath resolves the database location: SFDUPES_DATABASE when
// set and non-empty, the compiled-in default otherwise. // set and non-empty, the compiled-in default otherwise.
func databasePath() string { func databasePath() string {
@@ -104,6 +114,43 @@ func openDB(path, params string) (*sql.DB, error) {
return db, nil return db, nil
} }
// lockScanDatabase takes the lock that keeps a second scan off the
// database at path: an exclusive flock(2) on the file beside it named
// path with ".lock" appended, created along with the database's parent
// directory if missing. A lock held by another scan fails at once
// instead of waiting. The lock lasts until the returned file is closed
// or the process ends. The file is never deleted: a scan that deleted
// it would let the next scan lock a new file while another still holds
// the old one.
func lockScanDatabase(path string) (*os.File, error) {
err := os.MkdirAll(filepath.Dir(path), dbDirPerm)
if err != nil {
return nil, fmt.Errorf("create database directory: %w", err)
}
lockPath := path + ".lock"
//nolint:gosec // the operator chooses the database path
f, err := os.OpenFile(lockPath, os.O_RDWR|os.O_CREATE, lockFilePerm)
if err != nil {
return nil, err
}
err = unix.Flock(int(f.Fd()), unix.LOCK_EX|unix.LOCK_NB)
if err != nil {
_ = f.Close()
if errors.Is(err, unix.EWOULDBLOCK) {
return nil, fmt.Errorf("%w (lock held on %s)",
errScanRunning, lockPath)
}
return nil, fmt.Errorf("lock %s: %w", lockPath, err)
}
return f, nil
}
// openScanDatabase opens the database for the scan subcommand, creating // openScanDatabase opens the database for the scan subcommand, creating
// the file, its parent directory, and the schema as needed. // the file, its parent directory, and the schema as needed.
func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) { func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
+1 -1
View File
@@ -5,6 +5,7 @@ go 1.25.7
require ( require (
github.com/schollz/progressbar/v3 v3.19.1 github.com/schollz/progressbar/v3 v3.19.1
github.com/spf13/cobra v1.10.2 github.com/spf13/cobra v1.10.2
golang.org/x/sys v0.46.0
modernc.org/sqlite v1.54.0 modernc.org/sqlite v1.54.0
) )
@@ -18,7 +19,6 @@ require (
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/rivo/uniseg v0.4.7 // indirect github.com/rivo/uniseg v0.4.7 // indirect
github.com/spf13/pflag v1.0.9 // indirect github.com/spf13/pflag v1.0.9 // indirect
golang.org/x/sys v0.46.0 // indirect
golang.org/x/term v0.44.0 // indirect golang.org/x/term v0.44.0 // indirect
modernc.org/libc v1.74.1 // indirect modernc.org/libc v1.74.1 // indirect
modernc.org/mathutil v1.7.1 // indirect modernc.org/mathutil v1.7.1 // indirect
+94 -158
View File
@@ -82,50 +82,6 @@ func makeReadOnly(t *testing.T, path string) {
}) })
} }
// captureStderr redirects os.Stderr to a file for the rest of the test
// and returns a function reading back everything written to it. scan
// writes its warnings and summary straight to os.Stderr, not to the
// stderr writer run is given.
func captureStderr(t *testing.T) func() string {
t.Helper()
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
os.Stderr = f
t.Cleanup(func() {
os.Stderr = saved
_ = f.Close()
})
return func() string {
// Read what has been written without disturbing the write
// offset, so the capture can be inspected more than once.
size, err := f.Seek(0, io.SeekCurrent)
if err != nil {
t.Fatal(err)
}
if size == 0 {
return ""
}
b := make([]byte, size)
_, err = f.ReadAt(b, 0)
if err != nil {
t.Fatal(err)
}
return string(b)
}
}
// brokenDatabase writes a database that opens cleanly and passes the // brokenDatabase writes a database that opens cleanly and passes the
// schema-version check but has no files table, so the first query // schema-version check but has no files table, so the first query
// fails with the database already open: a fatal error on a path that // fails with the database already open: a fatal error on a path that
@@ -350,31 +306,19 @@ func scanFixture(t *testing.T) []string {
t.Fatal(err) t.Fatal(err)
} }
scanOK(t, dir) var stdout, stderr bytes.Buffer
return dupes code := run([]string{cmdScan, dir}, &stdout, &stderr)
}
// scanOK runs scan over operands, fails the test unless it exits 0 with
// nothing on stdout, and returns everything it printed to stderr.
func scanOK(t *testing.T, operands ...string) string {
t.Helper()
var stdout bytes.Buffer
stderr := captureStderr(t)
code := run(append([]string{cmdScan}, operands...), &stdout, os.Stderr)
if code != exitOK { if code != exitOK {
t.Fatalf("run(scan %q) = %d, want %d; stderr: %s", t.Fatalf("run(scan) = %d, want %d; stderr: %s",
operands, code, exitOK, stderr()) code, exitOK, stderr.String())
} }
if got := stdout.String(); got != "" { if got := stdout.String(); got != "" {
t.Errorf("scan stdout = %q, want nothing (data only)", got) t.Errorf("scan stdout = %q, want nothing (data only)", got)
} }
return stderr() return dupes
} }
func TestRunScanSucceedsDespiteWarnings(t *testing.T) { func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
@@ -385,103 +329,6 @@ func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
assertNoSidecars(t, path) assertNoSidecars(t, path)
} }
func TestRunScanSkipsSymlinkOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dir := t.TempDir()
writeFile(t, dir, "target/sub/f", pattern(1, 10))
link := filepath.Join(dir, "link")
err := os.Symlink(filepath.Join(dir, "target"), link)
if err != nil {
t.Fatal(err)
}
// Scanning a directory through the symlink stores a record beneath
// the symlink's own path for a file beneath its target.
scanOK(t, filepath.Join(link, "sub"))
assertOperandSkipped(t, path, link, "symlink",
filepath.Join(link, "sub", "f"))
}
func TestRunScanWalksOperandUnderSymlinkOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dir := t.TempDir()
writeFile(t, dir, "target/sub/f", pattern(1, 10))
link := filepath.Join(dir, "link")
err := os.Symlink(filepath.Join(dir, "target"), link)
if err != nil {
t.Fatal(err)
}
// link is dropped as a symlink, but link/sub must still be scanned,
// not dropped as lying under link.
scanOK(t, link, filepath.Join(link, "sub"))
db, err := openDB(path, reportParams)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
recordByPath(t, dbRecords(t, db), filepath.Join(link, "sub", "f"))
}
func TestRunScanSkipsZFSOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
zfs := filepath.Join(t.TempDir(), ".zfs")
snapshot := filepath.Join(zfs, "snapshot", "hourly")
f := writeFile(t, snapshot, "f", pattern(1, 10))
// An operand beneath a .zfs directory is walked, because it is not
// itself named .zfs.
scanOK(t, snapshot)
assertOperandSkipped(t, path, zfs, ".zfs directory", f)
}
// assertOperandSkipped scans operand alone and checks that it is skipped
// as kind: a warning naming it, one skip in the summary, exit 0, and the
// record for kept, which an earlier scan stored beneath operand, still
// in the database at dbPath.
func assertOperandSkipped(t *testing.T, dbPath, operand, kind,
kept string,
) {
t.Helper()
stderr := scanOK(t, operand)
warning := "walk " + operand + ": skipping " + kind + " operand\n"
if !strings.Contains(stderr, warning) {
t.Errorf("stderr = %q, want %q", stderr, warning)
}
summary := "scan: 0 files seen (0 added, 0 updated, 0 removed, " +
"0 unchanged), 1 skipped\n"
if !strings.Contains(stderr, summary) {
t.Errorf("stderr = %q, want %q", stderr, summary)
}
db, err := openDB(dbPath, reportParams)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
recordByPath(t, dbRecords(t, db), kept)
}
func TestRunReportSucceeds(t *testing.T) { func TestRunReportSucceeds(t *testing.T) {
path := testDBPath(t) path := testDBPath(t)
t.Setenv(databaseEnv, path) t.Setenv(databaseEnv, path)
@@ -565,6 +412,95 @@ func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
} }
} }
// holdScanLock takes the lock on the database at path, as a running
// scan does, and holds it until the test ends. It fails the test when
// the lock is already held.
func holdScanLock(t *testing.T, path string) {
t.Helper()
lock, err := lockScanDatabase(path)
if err != nil {
t.Fatalf("lock %s: %v", path, err)
}
t.Cleanup(func() { _ = lock.Close() })
}
func TestRunSecondScanFails(t *testing.T) {
// README §Database: while one scan holds the lock, a second scan
// fails at once, naming the lock file, without creating the
// database.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
holdScanLock(t, path)
var stdout, stderr bytes.Buffer
code := run([]string{cmdScan, t.TempDir()}, &stdout, &stderr)
if code != exitFatal {
t.Errorf("run(scan) = %d, want %d", code, exitFatal)
}
want := "sfdupes: another scan is running (lock held on " +
path + ".lock)\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
_, err := os.Stat(path)
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("stat %s = %v, want the database not created", path, err)
}
}
func TestRunScanReleasesLock(t *testing.T) {
// README §Database: a scan releases the lock however it ends.
t.Run("success", func(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
})
t.Run("fatal error", func(t *testing.T) {
path := brokenDatabase(t)
t.Setenv(databaseEnv, path)
code := run([]string{cmdScan, t.TempDir()}, io.Discard, io.Discard)
if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d", code, exitFatal)
}
holdScanLock(t, path)
})
}
func TestRunReportsDuringScan(t *testing.T) {
// README §Database: report and trees never take the lock, so they
// run while a scan holds it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
for _, name := range []string{cmdReport, cmdTrees} {
var stderr bytes.Buffer
code := run([]string{name}, io.Discard, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
}
}
}
func TestRunStdoutClosedIsFatal(t *testing.T) { func TestRunStdoutClosedIsFatal(t *testing.T) {
// README §Error handling: a stdout write failure exits 1, reported // README §Error handling: a stdout write failure exits 1, reported
// in one line on stderr. // in one line on stderr.
+28 -84
View File
@@ -85,11 +85,13 @@ type fileMeta struct {
// least one other file shares are ever hashed: a size-unique file // least one other file shares are ever hashed: a size-unique file
// cannot be a duplicate. A file of headTailMin or more gets its content // cannot be a duplicate. A file of headTailMin or more gets its content
// hash only when its size, head, and tail match another file's. Flag // hash only when its size, head, and tail match another file's. Flag
// parsing and the at-least-one-operand check are done by cobra. Errors // parsing and the at-least-one-operand check are done by cobra. The
// are returned rather than exiting, so that the deferred close — which // scan holds the lock on the database for its whole run, so a second
// takes the database out of WAL mode — always runs. Cancelling ctx // scan fails before it walks the filesystem or opens the database.
// unwinds the worker pools and aborts the scan with the context's // Errors are returned rather than exiting, so that the deferred close —
// error. // which takes the database out of WAL mode — always runs, and the lock
// is released after it. Cancelling ctx unwinds the worker pools and
// aborts the scan with the context's error.
func runScan(ctx context.Context, roots []string, workers int, func runScan(ctx context.Context, roots []string, workers int,
oneFS bool, oneFS bool,
) error { ) error {
@@ -104,6 +106,13 @@ func runScan(ctx context.Context, roots []string, workers int,
dbPath := databasePath() dbPath := databasePath()
lock, err := lockScanDatabase(dbPath)
if err != nil {
return err
}
defer func() { _ = lock.Close() }()
db, err := openScanDatabase(ctx, dbPath) db, err := openScanDatabase(ctx, dbPath)
if err != nil { if err != nil {
return err return err
@@ -211,17 +220,13 @@ type scanState struct {
// in the content hash of every record of headTailMin or more whose // in the content hash of every record of headTailMin or more whose
// size, head, and tail match another record's). Records outside the // size, head, and tail match another record's). Records outside the
// roots are never touched, except that the content phase fills in // roots are never touched, except that the content phase fills in
// their content hash. Operands the walk cannot start from are dropped // their content hash.
// first, so the records beneath them count as outside the roots unless
// they lie under another root.
func syncScan(ctx context.Context, db *sql.DB, roots []string, func syncScan(ctx context.Context, db *sql.DB, roots []string,
workers int, oneFS bool, workers int, oneFS bool,
) (scanStats, error) { ) (scanStats, error) {
s := &scanState{db: db} roots = pruneRoots(roots)
// Types are checked before pruning so that an operand under a s := &scanState{db: db}
// dropped one is still scanned, not dropped as lying under it.
roots = pruneRoots(s.walkableRoots(roots))
err := s.loadIndex(ctx, roots) err := s.loadIndex(ctx, roots)
if err != nil { if err != nil {
@@ -258,35 +263,6 @@ func syncScan(ctx context.Context, db *sql.DB, roots []string,
return s.st, s.contentPhase(ctx, workers) return s.st, s.contentPhase(ctx, workers)
} }
// walkableRoots returns the operands the walk can start from: regular
// files, and directories not named .zfs. Every other operand is warned
// about, counted as skipped, and dropped. A dropped operand is no
// longer a root, so the records stored beneath it count as outside the
// roots and are not deleted as unverified, unless it lies under another
// root. An operand that fails lstat here is kept, and the walk warns
// about it.
func (s *scanState) walkableRoots(roots []string) []string {
kept := make([]string, 0, len(roots))
for _, root := range roots {
fi, err := os.Lstat(root)
if err == nil {
warn := operandWarning(root, fi)
if warn != "" {
s.st.skipped++
fmt.Fprintln(os.Stderr, escapePath(warn))
continue
}
}
kept = append(kept, root)
}
return kept
}
// loadIndex indexes the database records under the scan roots for // loadIndex indexes the database records under the scan roots for
// change detection and collects the sizes of every record outside // change detection and collects the sizes of every record outside
// them: out-of-scope records join the size census so a scanned file // them: out-of-scope records join the size census so a scanned file
@@ -823,43 +799,11 @@ func sendEvent(ctx context.Context, events chan<- walkEvent,
} }
} }
// operandWarning returns the one-line warning for an operand the walk
// does not start from, naming the path and what it is, or "" for one it
// does: a regular file, or a directory not named .zfs. Symlinks are
// never followed, including as operands.
func operandWarning(root string, fi fs.FileInfo) string {
var kind string
switch mode := fi.Mode(); {
case mode.IsRegular():
return ""
case mode.IsDir():
if filepath.Base(root) != ".zfs" {
return ""
}
kind = ".zfs directory"
case mode&fs.ModeSymlink != 0:
kind = "symlink"
case mode&fs.ModeSocket != 0:
kind = "socket"
case mode&fs.ModeNamedPipe != 0:
kind = "FIFO"
case mode&fs.ModeDevice != 0:
kind = "device node"
default:
kind = "non-regular file"
}
return fmt.Sprintf("walk %s: skipping %s operand", root, kind)
}
// seedRoot turns one PATH operand into the walk's starting state: a // seedRoot turns one PATH operand into the walk's starting state: a
// regular-file operand is statted and emitted directly, and a directory // regular-file operand is statted and emitted directly, a directory
// operand becomes an initial job. walkableRoots has already dropped // operand becomes an initial job, and a symlink or other non-regular
// every other operand. One that has changed into something else since // operand yields nothing (symlinks are never followed, including as
// is warned about and skipped here; it is still a root, so the records // operands).
// stored beneath it are deleted as unverified.
func seedRoot(ctx context.Context, root string, func seedRoot(ctx context.Context, root string,
events chan<- walkEvent, events chan<- walkEvent,
) []dirJob { ) []dirJob {
@@ -873,19 +817,16 @@ func seedRoot(ctx context.Context, root string,
return nil return nil
} }
warn := operandWarning(root, fi) switch {
if warn != "" { case fi.IsDir():
sendEvent(ctx, events, walkEvent{warn: warn, fail: true}) if filepath.Base(root) == ".zfs" {
return nil return nil
} }
if fi.IsDir() {
dev, ok := deviceOfInfo(fi) dev, ok := deviceOfInfo(fi)
return []dirJob{{path: root, rootDev: dev, rootDevOK: ok}} return []dirJob{{path: root, rootDev: dev, rootDevOK: ok}}
} case fi.Mode().IsRegular():
dev, ino := inodeOfInfo(fi) dev, ino := inodeOfInfo(fi)
sendEvent(ctx, events, walkEvent{rec: fileRec{ sendEvent(ctx, events, walkEvent{rec: fileRec{
@@ -897,6 +838,9 @@ func seedRoot(ctx context.Context, root string,
}}) }})
return nil return nil
default:
return nil
}
} }
// startWalkWorkers starts the walk worker pool. Each worker processes // startWalkWorkers starts the walk worker pool. Each worker processes
+2 -4
View File
@@ -857,11 +857,9 @@ func TestWalkFileAndSymlinkOperands(t *testing.T) {
t.Fatalf("file operand: recs = %+v, errs = %d", recs, errs) t.Fatalf("file operand: recs = %+v, errs = %d", recs, errs)
} }
// A symlink operand that reaches the walk (it became one after // A symlink operand is not followed and yields nothing.
// walkableRoots checked it) is not followed: it yields a warning and
// no records.
recs, errs = collectWalk(t, []string{link}, false, 2) recs, errs = collectWalk(t, []string{link}, false, 2)
if errs != 1 || len(recs) != 0 { if errs != 0 || len(recs) != 0 {
t.Fatalf("symlink operand: recs = %+v, errs = %d", recs, errs) t.Fatalf("symlink operand: recs = %+v, errs = %d", recs, errs)
} }
} }