19 Commits

Author SHA1 Message Date
a0f0050ada Announce each operand on stderr before its passes
All checks were successful
check / check (push) Successful in 6s
With per-operand walk/hash/update cycles, a multi-operand invocation
(e.g. scan /srv/*) showed pass totals that looked like the whole
run's: an operator watching operand 3 of 14 hash 300k files concluded
the other 20M files were being skipped. Print the operand path and
its position before each cycle.
2026-07-24 07:52:54 +07:00
6a15b879de Merge branch 'parallel-walk': parallel per-directory walk, per-operand commits
All checks were successful
check / check (push) Successful in 1m10s
2026-07-24 07:44:28 +07:00
dced5cf0d2 Parallelize the walk and commit per operand
All checks were successful
check / check (push) Successful in 5s
Replace the single-goroutine WalkDir traversal with a per-directory
worker pool: workers read directories concurrently and lstat entries
while each directory is fresh in cache, recording size and mtime
during the walk. This folds the separate stat pass away (halving
metadata I/O per run) and overlaps metadata latency, which dominated
on busy pools — a sequential walk of a ~22M-file tree was observed
taking over 4 hours.

Each PATH operand now loads its scope, walks, hashes, and commits in
its own transaction, so an interrupted scan keeps every operand
completed so far; a later overlapping operand sees the records
committed by earlier ones and reuses them unchanged.
2026-07-24 07:44:18 +07:00
3ebf98940a Specify parallel per-directory walk and per-operand commits
The walk pass is a single goroutine; on a busy ZFS pool it manages
only a few thousand directory entries per second and takes hours at
~20M files. Respecify it as a worker-pool traversal that reads
directories concurrently and records size/mtime during the walk,
folding away the separate stat pass and halving metadata I/O. Each
PATH operand now commits in its own transaction so an interrupted
scan keeps the operands completed so far.
2026-07-24 06:28:12 +07:00
abea945730 Merge branch 'persistent-database': persistent SQLite scan database
All checks were successful
check / check (push) Successful in 1m2s
2026-07-24 03:08:58 +07:00
d8fbcb32c2 Implement persistent SQLite scan database
All checks were successful
check / check (push) Successful in 3s
scan now synchronizes a database that survives between runs
(SFDUPES_DATABASE, default /var/lib/sfdupes/db.sqlite) instead of
emitting a stream: operands are resolved to absolute paths, unchanged
files (same size, mtime not newer than recorded) are never re-read, new
and changed files are hashed, and records under the scanned operands
that were not verified this run are deleted; records outside the
operands are untouched. All changes commit in a single transaction, and
WAL journaling with a busy timeout keeps a report run during a cron
scan safe.

report and trees read the database (no positional arguments); the
NUL-terminated stream format, its parser, and the malformed-record
handling are gone. The driver is modernc.org/sqlite (pure Go), so
builds keep cgo disabled.
2026-07-24 03:08:49 +07:00
e0d578a707 Specify persistent SQLite scan database in README; plan in TODO.md
scan will maintain a persistent database of file signatures
(default /var/lib/sfdupes/db.sqlite, overridable via
SFDUPES_DATABASE) that survives between runs; rescans hash only new
or changed files (mtime/size) and remove records for vanished files,
so scan can be cronned daily. report and trees will read the
database instead of a scan stream.
2026-07-24 02:54:08 +07:00
90c9ef3546 Record remote setup and v0.0.1 tag in TODO.md
All checks were successful
check / check (push) Successful in 40s
2026-07-23 09:08:28 +07:00
13b9839e73 Disable background-session worktree isolation for this repo 2026-07-23 09:06:24 +07:00
4b1c3cbf70 Make scan paths required operands; add -x/--one-file-system
All checks were successful
check / check (push) Successful in 4s
scan now takes one or more PATH operands (directories or regular
files) via cobra flags instead of the -root flag with its /srv
default; invoking scan with no operand is a usage error and a
nonexistent operand is fatal. Filesystem boundaries are crossed by
default; the new -x/--one-file-system flag (GNU du/rsync convention)
stops the walk at each operand's filesystem, implemented by comparing
lstat device IDs with build-tagged helpers for darwin's int32 Dev.
Verified against a real mounted disk image: default crosses, -x does
not, --one-file-system is identical to -x.
2026-07-23 08:55:17 +07:00
3391207f8d Check off completed repo policy compliance items in TODO.md 2026-07-23 07:59:12 +07:00
5b20171dbd Install pre-commit hook via make hooks; resolve shared hooks dir
The hooks target uses git rev-parse --git-common-dir so it works from
both the main checkout and linked worktrees.
2026-07-23 06:45:02 +07:00
375233fff4 Restructure README with required policy sections
Adds Description (name/purpose/category/license/author), Getting
Started, Rationale, TODO, License, and Author sections; the full
normative specification is preserved under Design. The stale non-goal
about having no git repository or CI is removed, and the Build section
now documents the Makefile targets.
2026-07-23 06:44:32 +07:00
b4d013cb34 Add hash-pinned Dockerfile, .dockerignore, and Gitea CI workflow
Multistage build per policy: a fail-fast lint stage on the pinned
golangci-lint image runs fmt-check and lint, the builder stage reuses
its linter binary (which also forces stage ordering), runs make check,
and builds; the runtime stage is pinned alpine with just the binary.
CI runs docker build . on push with the checkout action pinned by
commit SHA. All image references pinned by sha256 digest with
version/date comments.
2026-07-23 06:42:11 +07:00
7a6adae0d2 Add test suite covering parsing, grouping, digests, and scan passes
Covers splitNUL framing, parseScanStream (malformed records, tabs and
newlines in paths, unterminated trailing record), duplicate grouping
with ordering/tie-break/determinism, humanBytes units, Merkle digest
equality and name/content sensitivity, maximal-tree suppression
(including sibling and differing-parent cases), hashHeadTail edge
sizes and error paths, walkPass .zfs/symlink skipping, statPass
vanished files, and an end-to-end walk/stat/hash pipeline test of the
README smoke-test scenario. 64% statement coverage.
2026-07-23 05:42:38 +07:00
316ea9072d Rewrite Makefile with required policy targets
test/lint/fmt/fmt-check/check/docker/hooks plus build, all, and clean.
make check no longer builds the binary (it must not modify files); the
test target uses the 30s-timeout conditional -v rerun pattern. Version
is injected via -ldflags and exposed as sfdupes --version.
2026-07-23 05:37:29 +07:00
065a533224 Make code lint-clean under the canonical golangci-lint config
150 findings fixed: linter autofixes for whitespace style (wsl_v5,
nlreturn, noinlineerr), named returns removed, magic numbers replaced
with named constants, static errNotRegular error, explicit Close/Parse
error handling, unused cobra params renamed to _, and runReport/
runTrees/hashPass split into helpers to satisfy cyclop, gocognit, and
funlen. Behavior verified unchanged against the README smoke test.
2026-07-23 05:36:05 +07:00
733b33d9d1 Add canonical .golangci.yml 2026-07-22 22:24:04 +07:00
2512ae797b Add REPO_POLICIES.md from canonical source (2026-07-06) 2026-07-22 22:24:04 +07:00
24 changed files with 3251 additions and 490 deletions

5
.claude/settings.json Normal file
View File

@@ -0,0 +1,5 @@
{
"worktree": {
"bgIsolation": "none"
}
}

8
.dockerignore Normal file
View File

@@ -0,0 +1,8 @@
.git
.claude
.DS_Store
sfdupes
files.dat
*.log
*.out
*.test

View File

@@ -0,0 +1,9 @@
name: check
on: [push]
jobs:
check:
runs-on: ubuntu-latest
steps:
# actions/checkout v4.2.2, 2026-02-22
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
- run: docker build .

3
.gitignore vendored
View File

@@ -28,6 +28,9 @@ node_modules/
# Local scan data
files.dat
*.sqlite
*.sqlite-shm
*.sqlite-wal
# Agent worktrees
.claude/worktrees/

32
.golangci.yml Normal file
View File

@@ -0,0 +1,32 @@
version: "2"
run:
timeout: 5m
modules-download-mode: readonly
linters:
default: all
disable:
# Genuinely incompatible with project patterns
- exhaustruct # Requires all struct fields
- depguard # Dependency allow/block lists
- godot # Requires comments to end with periods
- wsl # Deprecated, replaced by wsl_v5
- wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go
linters-settings:
lll:
line-length: 88
funlen:
lines: 80
statements: 50
cyclop:
max-complexity: 15
dupl:
threshold: 100
issues:
exclude-use-default: false
max-issues-per-linter: 0
max-same-issues: 0

38
Dockerfile Normal file
View File

@@ -0,0 +1,38 @@
# Lint stage — fast feedback on formatting and lint issues
# golangci/golangci-lint:v2.12.1, 2026-07-23
FROM golangci/golangci-lint@sha256:c9843d374ca80ecbac86081ec4dd7fe2bb6187b03224f59a0cc2f80759e1845b AS lint
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make fmt-check
RUN make lint
# Build stage
# golang:1.25-alpine, 2026-07-23
FROM golang@sha256:56961d79ea8129efddcc0b8643fd8a5416b4e6228cfd477e3fd61deb2672c587 AS builder
RUN apk add --no-cache make
WORKDIR /src
# Reuse the linter binary from the lint stage; the copy also forces
# BuildKit to complete linting before this stage proceeds.
COPY --from=lint /usr/bin/golangci-lint /usr/local/bin/golangci-lint
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Fail the build unless the branch is green.
RUN make check
RUN make build
# Runtime stage
# alpine:3.22, 2026-07-23
FROM alpine@sha256:14358309a308569c32bdc37e2e0e9694be33a9d99e68afb0f5ff33cc1f695dce
COPY --from=builder /src/sfdupes /usr/local/bin/sfdupes
ENTRYPOINT ["sfdupes"]

View File

@@ -2,16 +2,47 @@
# toolchain.
export CGO_ENABLED = 0
.PHONY: all build check clean
BINARY := sfdupes
VERSION := $(shell git describe --tags --always --dirty 2>/dev/null || echo dev)
LDFLAGS := -X main.Version=$(VERSION)
all: build
.PHONY: all build test lint fmt fmt-check check docker hooks clean
all: check build
build:
go build
go build -ldflags "$(LDFLAGS)" -o $(BINARY)
check: build
test -z "$$(gofmt -l .)"
go vet ./...
test:
@go test -timeout 30s -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 30s -v ./...; exit 1; }
lint:
golangci-lint run --config .golangci.yml ./...
fmt:
gofmt -s -w .
fmt-check:
@files="$$(gofmt -l -s .)"; if [ -n "$$files" ]; then \
echo "gofmt: files not formatted:"; echo "$$files"; exit 1; fi
check: test lint fmt-check
docker:
docker build -t $(BINARY) .
# Hooks are shared between the main checkout and all worktrees, so
# resolve the common git dir instead of assuming .git is a directory.
HOOKS_DIR := $(shell git rev-parse --git-common-dir)/hooks
hooks:
@printf '#!/bin/sh\nset -e\n' > $(HOOKS_DIR)/pre-commit
@printf 'go mod tidy\ngo fmt ./...\n' >> $(HOOKS_DIR)/pre-commit
@printf 'git diff --exit-code -- go.mod go.sum || { echo "go mod tidy changed files; please stage and retry"; exit 1; }\n' >> $(HOOKS_DIR)/pre-commit
@printf 'make check\n' >> $(HOOKS_DIR)/pre-commit
@chmod +x $(HOOKS_DIR)/pre-commit
clean:
rm -f sfdupes files.dat
rm -f $(BINARY) files.dat

432
README.md
View File

@@ -1,18 +1,67 @@
# sfdupes
`sfdupes` quickly identifies *candidate* duplicate files — and, ultimately,
entire duplicate directory trees — across very large filesystems without
reading full file contents. Files are considered duplicates when they have
identical size, identical SHA-256 of their first 1024 bytes, and identical
SHA-256 of their last 1024 bytes. This is a strong candidate signal, not
proof of identical content (the middle of the file is never read); the
intended use is finding duplicate downloads and duplicated directory trees
on multi-terabyte ZFS servers where reading every byte is prohibitively
expensive.
## Description
`sfdupes` is an MIT-licensed Go CLI tool by
[@sneak](https://sneak.berlin) that quickly identifies *candidate*
duplicate files — and, ultimately, entire duplicate directory trees —
across very large filesystems without reading full file contents. Files
are considered duplicates when they have identical size, identical
SHA-256 of their first 1024 bytes, and identical SHA-256 of their last
1024 bytes. This is a strong candidate signal, not proof of identical
content (the middle of the file is never read); the intended use is
finding duplicate downloads and duplicated directory trees on
multi-terabyte ZFS servers where reading every byte is prohibitively
expensive. `scan` maintains a persistent SQLite database of file
signatures that survives between runs, so it can be run from cron and
the reports can be generated at any time from the most recent scan.
This tool was created by [@sneak](https://sneak.berlin) to scratch an itch,
using Claude Code/Fable.
This README is the complete and authoritative specification.
## Design goals
## Getting Started
```sh
make build
export SFDUPES_DATABASE="$HOME/.local/share/sfdupes/db.sqlite"
./sfdupes scan /srv
./sfdupes report > dupes.tsv
./sfdupes trees > dupetrees.tsv
```
`scan` walks one or more filesystem trees and maintains one database
record per regular file (path, size, mtime, head hash, tail hash). The
database persists between runs; a rescan only hashes files that are new
or changed, and removes records for files that no longer exist.
`report` reads the database and prints the file-level duplicates
report. `trees` reads the same database and prints the duplicate-tree
report. A missing/invalid subcommand — or a `scan` invocation with no
`PATH` operand — prints a usage message and exits 2.
The database defaults to `/var/lib/sfdupes/db.sqlite` and can be placed
anywhere by setting `SFDUPES_DATABASE`. The intended deployment is a
daily `sfdupes scan` cron job, with the reporting commands run
interactively whenever needed; their results are as fresh as the last
completed scan.
## Rationale
Duplicate finders that hash entire files do not scale to the target
environment: ~10 million files and ~150 TB on possibly slow or busy
disks (a ZFS pool under resilver). Reading at most 2 KiB per file makes
a full-filesystem sweep tractable, and the signatures are kept in a
persistent database, so the expensive filesystem pass is incremental: a
rescan re-hashes only files whose recorded mtime or size changed, and
all analysis happens offline from the database alone. The end goal is
not individual files but whole duplicated trees — duplicate
extractions, duplicate downloads, copied project trees — which an
operator can consider removing as a unit.
## Design
Goals, in order:
1. **Find whole duplicate trees, not just files.** The end goal is to
identify places where the exact same set of files and directories
@@ -25,129 +74,208 @@ This README is the complete and authoritative specification.
filesystem, possibly slow or busy disks (ZFS pool under resilver).
Holding the full file list in memory is acceptable; reading file
contents beyond 2 KiB per file is not.
3. **Scan once, analyze offline.** The expensive filesystem scan
produces a self-contained stream; all analysis (`report`, `trees`)
works from that stream alone and must never touch the scanned
filesystem again.
3. **Scan incrementally, analyze offline.** The expensive filesystem
scan maintains a persistent database; an unchanged file is never
read again on a rescan. All analysis (`report`, `trees`) works from
the database alone and must never touch the scanned filesystem
again. `scan` is designed to be cronned; the reports run at any
time against the last completed scan.
4. **Clean stream separation.** Everything on stdout is machine-readable
data. All progress, warnings, and summaries go to stderr. Never mix
them.
## Constraints
### Constraints
- Language: Go (module `sneak.berlin/go/sfdupes`, already initialized
in `go.mod`). Binary name: `sfdupes`.
- Dependencies: standard library, `github.com/spf13/cobra` for the CLI,
and **one progress-bar library**
(`github.com/schollz/progressbar/v3`). `github.com/spf13/viper` is
permitted if configuration-file support is ever needed, but is not
currently used. No other third-party deps.
- Language: Go (module `sneak.berlin/go/sfdupes`). Binary name:
`sfdupes`.
- Dependencies: standard library, `github.com/spf13/cobra` for the
CLI, **one progress-bar library**
(`github.com/schollz/progressbar/v3`), and **one SQLite driver**
(`modernc.org/sqlite`, pure Go, so builds keep cgo disabled).
`github.com/spf13/viper` is permitted if configuration-file support
is ever needed, but is not currently used. No other third-party
deps.
- Cross-compilation is not a concern. Builds run with cgo disabled (the
`Makefile` exports `CGO_ENABLED=0`); the code must remain pure Go.
- Analysis modes (`report`, `trees`) must be deterministic: identical
input stream, identical output, regardless of record order.
database contents, identical output, regardless of the order in
which records were inserted.
## Plan
### Subcommands
Three subcommands, built in this order:
Three subcommands, all implemented:
1. `scan` — walk the filesystem and emit one signature record per
regular file (implemented).
2. `report` — file-level duplicate report from the scan stream
(implemented).
1. `scan` — walk the filesystem and synchronize the database: one
signature record per regular file.
2. `report` — file-level duplicate report from the database.
3. `trees` — tree-level duplicate report: reconstruct the directory
hierarchy from the scan stream, compute a Merkle-style digest per
directory, and report maximal groups of identical trees
(implemented).
Possible later work, explicitly out of scope for now: full-content
verification of candidates, and helpers that emit removal scripts.
## Usage
hierarchy from the database records, compute a Merkle-style digest
per directory, and report maximal groups of identical trees.
```
sfdupes scan [-root /srv] [-workers N] > files.dat
sfdupes report [files.dat|-] > dupes.tsv
sfdupes trees [files.dat|-] > dupetrees.tsv
sfdupes scan [--workers N] [-x] PATH...
sfdupes report > dupes.tsv
sfdupes trees > dupetrees.tsv
```
`scan` walks a filesystem tree and emits one record per regular file
(path, size, mtime, head hash, tail hash). `report` ingests that stream
and prints the file-level duplicates report. `trees` ingests the same
stream and prints the duplicate-tree report. A missing/invalid subcommand
prints a usage message and exits 2.
### Database
## `scan` mode
All three subcommands operate on a single SQLite database file:
`scan` runs **three sequential passes**, in this order, so that every
expensive pass has an exact total for meaningful progress and ETA:
- Location: the value of the `SFDUPES_DATABASE` environment variable
when set and non-empty, otherwise `/var/lib/sfdupes/db.sqlite`.
There is no command-line flag.
- `scan` creates the database (and its parent directory) on first
use. `report` and `trees` require an existing database; a missing
database file is a fatal error (exit 1) telling the user to run
`scan` first.
- The database uses WAL journal mode and a busy timeout, so running a
report while a cron `scan` is in progress is safe; the reports see
the last committed state. Each `PATH` operand's changes are
committed in a single transaction when that operand completes, so
a report never observes a half-finished operand, and a scan that
dies partway keeps every operand completed so far (the interrupted
operand's changes are lost, not corrupted).
- Schema (`PRAGMA user_version` is the schema version, currently 1; a
database with any other version is a fatal error):
1. **walk** — recursively enumerate the tree under `-root` (default
`/srv`), collecting the list of regular-file paths. Total unknown
while running: show a live count, not a percentage.
2. **stat**`lstat` every collected path, recording size and mtime.
3. **hash** — for each file, read the first `min(1024, size)` bytes and
the last `min(1024, size)` bytes (the two reads overlap when
`size < 2048`; for `size == 0` hash the empty input) and compute the
SHA-256 of each. Emit the output record.
```sql
CREATE TABLE files (
path BLOB PRIMARY KEY, -- absolute path, raw bytes
size INTEGER NOT NULL, -- bytes, from lstat
mtime INTEGER NOT NULL, -- Unix seconds, from lstat
head TEXT NOT NULL, -- lowercase-hex SHA-256, first 1 KiB
tail TEXT NOT NULL -- lowercase-hex SHA-256, last 1 KiB
) WITHOUT ROWID;
```
Paths are stored as BLOBs because Unix paths are raw bytes, not
guaranteed UTF-8. `mtime` is used only for change detection; it is
not part of the duplicate key.
### `scan` mode
`scan` requires one or more `PATH` operands naming the trees to scan.
There is no default path; invoking `scan` with no operand is a usage
error (usage message on stderr, exit 2). An operand may be a directory
or a regular file; an operand that does not exist is a fatal error
(exit 1). Because database records persist between runs and are keyed
by absolute path, each operand is resolved to an absolute, lexically
cleaned path (symlinks are not resolved) before walking, so results do
not depend on the working directory. Operands are processed in the
order given, each one walked and committed independently; overlapping
operands (one containing another) are harmless — a file reached via
multiple operands produces one database record (later operands see the
records committed by earlier ones and reuse them unchanged).
`scan` synchronizes the database with the filesystem state under the
scanned operands:
- A file not yet in the database is hashed and inserted.
- A file already in the database is **skipped without reading its
contents** when its lstat size equals the recorded size and its
lstat mtime is not newer than the recorded mtime. This is what
makes a daily rescan cheap.
- A file whose mtime is newer than recorded, or whose size differs,
is re-hashed and its record updated.
- A database record whose path lies under one of the scanned operands
but was not successfully processed this run is deleted. This
removes records for deleted files. It also removes records for
paths that failed to stat or hash this run: the database only ever
contains signatures verified by the most recent scan that covered
them (a subsequent successful scan re-adds such files).
- Database records outside the scanned operands are untouched, so
disjoint trees can be scanned on different schedules into the same
database.
`scan` runs **three sequential passes per operand**, committing each
operand before starting the next:
1. **walk** — enumerate the tree with the worker pool: each worker
reads one directory at a time, records size and mtime for every
regular-file entry (`lstat` while the directory is fresh in
cache), and hands discovered subdirectories back to the shared
queue. Sequential directory enumeration is metadata-latency-bound
and takes hours at tens of millions of files; per-directory
parallelism is what makes the walk tractable on large or busy
pools. Total unknown while running: show a live count, not a
percentage.
2. **hash** — for each new or changed file (per the rules above), read
the first `min(1024, size)` bytes and the last `min(1024, size)`
bytes (the two reads overlap when `size < 2048`; for `size == 0`
hash the empty input) and compute the SHA-256 of each. Unchanged
files are not read and do not appear in this pass's total, which
is therefore exact for meaningful progress and ETA.
3. **update** — apply the operand's insertions, updates, and
deletions to the database in a single transaction.
Because every operand runs its own three passes with its own totals,
`scan` announces each operand on stderr before starting it, so a
multi-operand invocation (e.g. a shell glob) is not mistaken for a
nearly-finished run:
```
scan: /srv/media (operand 3 of 14)
```
Rules for the walk:
- Only regular files. Skip directories, symlinks (do not follow),
sockets, FIFOs, and device nodes.
- Only regular files. Skip directories, symlinks (do not follow,
including symlink operands), sockets, FIFOs, and device nodes.
- Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs;
walking them would list every file once per snapshot).
- Filesystem boundaries are crossed by default. With `-x`
(long form `--one-file-system`, following the GNU `du`/`rsync`
convention), never descend into a directory on a different
filesystem than its `PATH` operand; each operand is bounded by its
own filesystem.
- On any per-path error (permission denied, file vanished between
passes, unreadable): print a one-line warning to stderr, skip the
path, and continue. Per-file errors never abort the run; the final
summary reports how many were skipped.
summary reports how many were skipped. As specified above, a
skipped path that has a database record from an earlier scan loses
that record; an unreadable directory subtree likewise loses its
records (accepted: the database mirrors what the latest scan could
actually verify).
Concurrency: the stat and hash passes use a worker pool (`-workers`,
default `runtime.NumCPU()`). The main goroutine owns stdout writing and
progress rendering; progress display must never block the workers.
Concurrency: the walk and hash passes use a worker pool (`--workers`,
default `runtime.NumCPU()`); the walk parallelizes across directories,
so raising `--workers` can speed up metadata-bound walks on busy
pools. The main goroutine owns database writes and progress rendering;
progress display must never block the workers.
### Output record format
One record per file on stdout, NUL-terminated (`\x00`), with
tab-separated fields, **path last** so tabs or newlines embedded in
paths cannot corrupt the record structure:
`scan` writes nothing to stdout. The summary line on stderr reports the
files seen this run broken down by disposition, plus skips:
```
<size>\t<mtime_unix>\t<sha256_first1k_hex>\t<sha256_last1k_hex>\t<path>\x00
scan: 123456 files seen (1200 added, 34 updated, 56 removed, 122166 unchanged), 3 skipped
```
- `size`: decimal bytes, from the stat pass.
- `mtime_unix`: decimal Unix seconds. Informational only; not part of
the duplicate key.
- Hashes: lowercase hex, 64 chars each.
- Record order is unspecified (workers complete out of order); the
analysis modes must not depend on ordering.
(`removed` counts deleted database records, which are not part of the
files-seen total.)
## `report` mode
### `report` mode
`report` reads the scan stream from the file named in its first
positional argument, or from stdin if the argument is absent or `-`.
`report` reads every record from the database and takes no positional
arguments.
**`report` must never touch the filesystem being analyzed.** It does not
stat, open, or otherwise access any path that appears in the records; its
only I/O is reading the scan file/stdin and writing stdout/stderr. It must
only I/O is reading the database and writing stdout/stderr. It must
produce identical output whether or not the scanned filesystem is still
mounted.
Processing:
- Parse records; a record that does not have exactly 5 fields or whose
size is non-numeric is counted as malformed and skipped (warn once
with the total malformed count in the summary, not per record).
- Group records by the key `(size, head_hash, tail_hash)`.
- Every group with two or more paths is a duplicate group.
- Within each group, sort paths lexicographically (byte order). The
first path is the group's `first`; every other path is a `dupe`.
- Order groups by size descending (biggest reclaimable space first),
tie-broken by `first` path ascending. Output must be fully
deterministic for a given input.
deterministic for a given database state.
### Report output format
#### Report output format
TSV on stdout: a header line, then one row per duplicate file (N-1 rows
for a group of N):
@@ -158,16 +286,16 @@ first dupe size
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296
```
Summary to stderr: records read, malformed count (if any), number of
duplicate groups, number of dupe files, and total reclaimable bytes
(sum of `size` over all dupe rows) in human units.
Summary to stderr: records read, number of duplicate groups, number of
dupe files, and total reclaimable bytes (sum of `size` over all dupe
rows) in human units.
## `trees` mode
### `trees` mode
`trees` reads the same scan stream as `report` (same argument handling,
same parsing and malformed-record rules) and reports **entire duplicate
directory trees**: directories under which the exact same set of relative
paths exists with the exact same file signatures.
`trees` reads the same database as `report` (no positional arguments)
and reports **entire duplicate directory trees**: directories under
which the exact same set of relative paths exists with the exact same
file signatures.
**`trees` must never touch the filesystem being analyzed** — the same
rule as `report`. The directory hierarchy is reconstructed purely from
@@ -188,10 +316,10 @@ Definitions:
equal. Equal digests imply equal recursive file count and equal
total byte size.
Known limitation (accepted): only regular files that appear in the scan
stream define a tree. Empty directories are invisible, and a file skipped
during the scan (e.g. permission error) in one copy but not the other
will make otherwise-identical trees compare as different.
Known limitation (accepted): only regular files that appear in the
database define a tree. Empty directories are invisible, and a file
skipped during the scan (e.g. permission error) in one copy but not the
other will make otherwise-identical trees compare as different.
Processing:
@@ -209,7 +337,7 @@ Processing:
path ascending. Output must be fully deterministic for a given
input.
### Trees output format
#### Trees output format
TSV on stdout: a header line, then one row per duplicate tree (N-1 rows
for a group of N). `files` is the recursive regular-file count of one
@@ -220,22 +348,22 @@ first dupe files size
/srv/a/project /srv/backup/project 3417 104857600
```
Summary to stderr: records read, malformed count (if any), number of
duplicate-tree groups, number of dupe trees, and total reclaimable bytes
(sum of `size` over all dupe rows) in human units.
Summary to stderr: records read, number of duplicate-tree groups,
number of dupe trees, and total reclaimable bytes (sum of `size` over
all dupe rows) in human units.
## Progress
### Progress
Use the progress-bar library for all scan-pass progress; rendering in the
style of `pv` is the model. All progress goes to stderr.
Each scan pass gets its own bar. Required elements for the stat and hash
passes (known totals):
Each scan pass gets its own bar, repeated per operand. Required
elements for the hash and update passes (known totals):
- elapsed time
- estimated time remaining
- a `[m/n] x%` display (files processed / total files, percent)
- current rate (files/s)
- a `[m/n] x%` display (items processed / total items, percent)
- current rate (items/s)
Example shape (exact layout is flexible, content is not):
@@ -255,31 +383,43 @@ Additional requirements:
- `report` and `trees` modes need no progress display, only their
stderr summaries.
## Error handling and exit codes
### Error handling and exit codes
- `0`: success, even if individual files were skipped with warnings.
- `1`: fatal error (e.g., `-root` does not exist, cannot read the scan
input, stdout write failure).
- `2`: usage error.
- `1`: fatal error (e.g., a `PATH` operand does not exist, the
database cannot be created/opened/read/written, a missing database
for `report`/`trees`, stdout write failure).
- `2`: usage error (including `scan` with no `PATH` operand and
`report`/`trees` with any positional argument).
## Build
`make` (or `make build`) builds the binary with cgo disabled. `make
check` builds and runs `gofmt` and `go vet`. `make clean` removes the
binary and any local `files.dat`.
The `Makefile` is the single source of truth for all operations:
## Definition of done
- `make build` — build the `sfdupes` binary (cgo disabled).
- `make test` — run the test suite (30-second timeout; reruns with
`-v` on failure).
- `make lint` — run `golangci-lint` with the repo config.
- `make fmt` / `make fmt-check` — format Go sources / verify
formatting without writing.
- `make check` — `test`, `lint`, and `fmt-check`; modifies nothing.
- `make docker` — build the Docker image, which runs `make check` as
a build stage.
- `make hooks` — install the pre-commit hook.
- `make clean` — remove the binary and any legacy local `files.dat`.
### Definition of done
All of the following, run in this directory, must pass:
1. `gofmt -l .` prints nothing.
2. `go vet ./...` is clean.
3. `go build` succeeds (equivalently, `make check` passes).
4. Smoke test — create a throwaway tree in a temp dir (never test
1. `make check` passes (tests, lint, `gofmt`).
2. `make docker` succeeds.
3. Smoke test — create a throwaway tree in a temp dir (never test
against real data):
```sh
d=$(mktemp -d)
export SFDUPES_DATABASE="$d/db.sqlite"
mkdir -p "$d/a" "$d/b"
head -c 2000 /dev/urandom > "$d/a/one.bin"
cp "$d/a/one.bin" "$d/b/copy.bin"
@@ -296,35 +436,59 @@ All of the following, run in this directory, must pass:
cp "$d/t1/sub/f2" "$d/t2/sub/f2"
cp "$d/t1/f1" "$d/t3/f1"
cp "$d/t1/sub/f2" "$d/t3/sub/f2renamed"
./sfdupes scan -root "$d" > files.dat
./sfdupes report files.dat
./sfdupes trees files.dat
./sfdupes scan "$d"
./sfdupes report
./sfdupes trees
# incremental behavior (scan a subtree; records elsewhere persist):
./sfdupes scan "$d/a" # everything unchanged, nothing hashed
printf 'z' >> "$d/a/one.bin" # modify: next scan re-hashes it
rm "$d/a/unique.bin" # delete: next scan removes its record
./sfdupes scan "$d/a" # 1 updated, 1 removed
./sfdupes report
```
Expected from `report`: `one.bin`/`copy.bin`/`copy2.bin` form one
group (two dupe rows, `first` is the lexicographically smallest
path); `t1/f1`/`t2/f1`/`t3/f1` form one group; `t1/sub/f2`/
`t2/sub/f2`/`t3/sub/f2renamed` form one group; `tiny1`/`tiny2` pair;
`empty1`/`empty2` pair; `unique.bin` and `tiny3` appear nowhere;
groups ordered by size descending; piping scan directly into report
(`./sfdupes scan -root "$d" | ./sfdupes report`) gives the same
rows.
(The scan database lives inside `$d` here purely for test hygiene;
scanning `$d` therefore also records the SQLite file itself, which
is harmless.)
Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin`
form one group (two dupe rows, `first` is the lexicographically
smallest path); `t1/f1`/`t2/f1`/`t3/f1` form one group;
`t1/sub/f2`/ `t2/sub/f2`/`t3/sub/f2renamed` form one group;
`tiny1`/`tiny2` pair; `empty1`/`empty2` pair; `unique.bin` and
`tiny3` appear nowhere; groups ordered by size descending.
Expected from `trees`: exactly one row — `first` `$d/t1`, `dupe`
`$d/t2`, 2 files, 3100 bytes. `$d/t1/sub` vs `$d/t2/sub` is
suppressed as non-maximal (implied by the `t1`/`t2` group), and `t3`
appears nowhere (its file set differs by name).
5. A negative check: run `report` and `trees` after deleting the temp
tree — output must be unchanged (proves the analysis modes never
touch the filesystem).
Expected from the second `report` (after the modify/delete rescan):
`one.bin` has left its group (its content changed), so
`copy.bin`/`copy2.bin` remain as one pair, and `unique.bin` is
gone from the database.
The test suite automates this scenario (see `scan_test.go`), plus a
negative check: `report` and `trees` operate on the database alone
and never touch the scanned filesystem.
## TODO
Tracked in [TODO.md](TODO.md).
## Non-goals
- No full-content verification, no byte-for-byte compare, no deletion
or linking of duplicates. The reports are advisory; acting on them is
the user's job.
- No persistence formats beyond the scan stream described above.
- No git repository setup and no CI — code, `go.mod`/`go.sum`, the
`Makefile`, and this README only. Do not write anything outside this
directory (module cache aside).
- No persistence beyond the SQLite database described above; no
export/import formats.
- No daemon or filesystem watcher; scheduling rescans is cron's job.
## License
MIT. See [LICENSE](LICENSE).
## Author
[@sneak](https://sneak.berlin)

368
REPO_POLICIES.md Normal file
View File

@@ -0,0 +1,368 @@
---
title: Repository Policies
last_modified: 2026-07-06
---
This document covers repository structure, tooling, and workflow standards. Code
style conventions are in separate documents:
- [Code Styleguide](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE.md)
(general, bash, Docker)
- [Go](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE_GO.md)
- [JavaScript](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE_JS.md)
- [Python](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE_PYTHON.md)
- [Go HTTP Server Conventions](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/GO_HTTP_SERVER_CONVENTIONS.md)
---
- Cross-project documentation (such as this file) must include
`last_modified: YYYY-MM-DD` in the YAML front matter so it can be kept in sync
with the authoritative source as policies evolve.
- **ALL external references must be pinned by cryptographic hash.** This
includes Docker base images, Go modules, npm packages, GitHub Actions, and
anything else fetched from a remote source. Version tags (`@v4`, `@latest`,
`:3.21`, etc.) are server-mutable and therefore remote code execution
vulnerabilities. The ONLY acceptable way to reference an external dependency
is by its content hash (Docker `@sha256:...`, Go module hash in `go.sum`, npm
integrity hash in lockfile, GitHub Actions `@<commit-sha>`). No exceptions.
This also means never `curl | bash` to install tools like pyenv, nvm, rustup,
etc. Instead, download a specific release archive from GitHub, verify its hash
(hardcoded in the Dockerfile or script), and only then install. Unverified
install scripts are arbitrary remote code execution. This is the single most
important rule in this document. Double-check every external reference in
every file before committing. There are zero exceptions to this rule.
- Every repo with software must have a root `Makefile` with these targets:
`make test`, `make lint`, `make fmt` (writes), `make fmt-check` (read-only),
`make check` (prereqs: `test`, `lint`, `fmt-check`), `make docker`, and
`make hooks` (installs pre-commit hook). A model Makefile is at
`https://git.eeqj.de/sneak/prompts/raw/branch/main/Makefile`.
- Always use Makefile targets (`make fmt`, `make test`, `make lint`, etc.)
instead of invoking the underlying tools directly. The Makefile is the single
source of truth for how these operations are run.
- The Makefile is authoritative documentation for how the repo is used. Beyond
the required targets above, it should have targets for every common operation:
running a local development server (`make run`, `make dev`), re-initializing
or migrating the database (`make db-reset`, `make migrate`), building
artifacts (`make build`), generating code, seeding data, or anything else a
developer would do regularly. If someone checks out the repo and types
`make<tab>`, they should see every meaningful operation available. A new
contributor should be able to understand the entire development workflow by
reading the Makefile.
- Every repo should have a `Dockerfile`. All Dockerfiles must run `make check`
as a build step so the build fails if the branch is not green. For non-server
repos, the Dockerfile should bring up a development environment and run
`make check`. For server repos, `make check` should run as an early build
stage before the final image is assembled.
- **Dockerfiles must use a separate lint stage for fail-fast feedback.** Go
repos use a multistage build where linting runs in an independent stage based
on the `golangci/golangci-lint` image (pinned by hash). This stage runs
`make fmt-check` and `make lint` before the full build begins. The build stage
then declares an explicit dependency on the lint stage via
`COPY --from=lint /src/go.sum /dev/null`, which forces BuildKit to complete
linting before proceeding to compilation and tests. This ensures lint failures
surface in seconds rather than minutes, without blocking on dependency
download or compilation in the build stage.
The standard pattern for a Go repo Dockerfile is:
```dockerfile
# Lint stage — fast feedback on formatting and lint issues
# golangci/golangci-lint:v2.x.x, YYYY-MM-DD
FROM golangci/golangci-lint@sha256:... AS lint
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make fmt-check
RUN make lint
# Build stage
# golang:1.x-alpine, YYYY-MM-DD
FROM golang@sha256:... AS builder
WORKDIR /src
# Force BuildKit to run the lint stage before proceeding
COPY --from=lint /src/go.sum /dev/null
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make test
ARG VERSION=dev
RUN CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X main.Version=${VERSION}" \
-o /app ./cmd/app/
# Runtime stage
FROM alpine@sha256:...
COPY --from=builder /app /usr/local/bin/app
ENTRYPOINT ["app"]
```
Key points:
- The lint stage uses the `golangci/golangci-lint` image directly (it
includes both Go and the linter), so there is no need to install the
linter separately.
- `COPY --from=lint /src/go.sum /dev/null` is a no-op file copy that creates
a stage dependency. BuildKit runs stages in parallel by default; without
this line, the build stage would not wait for lint to finish and a lint
failure might not fail the overall build.
- If the project uses `//go:embed` directives that reference build artifacts
(e.g. a web frontend compiled in a separate stage), the lint stage must
create placeholder files so the embed directives resolve. Example:
`RUN mkdir -p web/dist && touch web/dist/index.html web/dist/style.css`.
The lint stage should not depend on the actual build output — it exists to
fail fast.
- If the project requires CGO or system libraries for linting (e.g.
`vips-dev`), install them in the lint stage with `apk add`.
- The build stage runs `make test` after compilation setup. Tests run in the
build stage, not the lint stage, because they may require compiled
artifacts or heavier dependencies.
- Every repo should have a Gitea Actions workflow (`.gitea/workflows/`) that
runs `docker build .` on push. Since the Dockerfile already runs `make check`,
a successful build implies all checks pass.
- Use platform-standard formatters: `black` for Python, `prettier` for
JS/CSS/Markdown/HTML, `go fmt` for Go. Always use default configuration with
two exceptions: four-space indents (except Go), and `proseWrap: always` for
Markdown (hard-wrap at 80 columns). Documentation and writing repos (Markdown,
HTML, CSS) should also have `.prettierrc` and `.prettierignore`.
- Pre-commit hook: `make check` if local testing is possible, otherwise
`make lint && make fmt-check`. The Makefile should provide a `make hooks`
target to install the pre-commit hook.
- All repos with software must have tests that run via the platform-standard
test framework (`go test`, `pytest`, `jest`/`vitest`, etc.). If no meaningful
tests exist yet, add the most minimal test possible — e.g. importing the
module under test to verify it compiles/parses. There is no excuse for
`make test` to be a no-op.
- `make test` must complete in under 20 seconds. Add a 30-second timeout in the
Makefile.
- **`make test` should use the conditional verbose rerun pattern.** Run tests
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to
show full output. This keeps CI logs and `docker build` output clean on
success (just package/suite summaries) while providing full diagnostic detail
on failure (every test case, every assertion). The general shell pattern:
```makefile
test:
@<test-command> || \
{ echo "--- Rerunning with -v for details ---"; \
<test-command-with-v>; exit 1; }
```
Go example:
```makefile
test:
@go test -timeout 30s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 30s -race -v ./...; exit 1; }
```
Python example:
```makefile
test:
@python -m pytest || \
{ echo "--- Rerunning with -v for details ---"; \
python -m pytest -v; exit 1; }
```
The `exit 1` ensures the target always fails after a rerun — the first run
already proved the tests are broken, so the build must not pass even if a
flaky test happens to succeed on the second attempt. The rerun exists solely
for diagnostic output.
- Docker builds must complete in under 5 minutes.
- `make check` must not modify any files in the repo. Tests may use temporary
directories.
- `main` must always pass `make check`, no exceptions.
- Never commit secrets. `.env` files, credentials, API keys, and private keys
must be in `.gitignore`. No exceptions.
- `.gitignore` should be comprehensive from the start: OS files (`.DS_Store`),
editor files (`.swp`, `*~`), language build artifacts, and `node_modules/`.
Fetch the standard `.gitignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.gitignore` when setting up
a new repo.
- **No build artifacts in version control.** Code-derived data (compiled
bundles, minified output, generated assets) must never be committed to the
repository if it can be avoided. The build process (e.g. Dockerfile, Makefile)
should generate these at build time. Notable exception: Go protobuf generated
files (`.pb.go`) ARE committed because repos need to work with `go get`, which
downloads code but does not execute code generation.
- Never use `git add -A` or `git add .`. Always stage files explicitly by name.
- Never force-push to `main`.
- Make all changes on a feature branch. You can do whatever you want on a
feature branch.
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only
manually by the user. Fetch from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`.
- When pinning images or packages by hash, add a comment above the reference
with the version and date (YYYY-MM-DD).
- Use `yarn`, not `npm`.
- Write all dates as YYYY-MM-DD (ISO 8601).
- Simple projects should be configured with environment variables.
- Dockerized web services listen on port 8080 by default, overridable with
`PORT`.
- **HTTP/web services must be hardened for production internet exposure before
tagging 1.0.** This means full compliance with security best practices
including, without limitation, all of the following:
- **Security headers** on every response:
- `Strict-Transport-Security` (HSTS) with `max-age` of at least one year
and `includeSubDomains`.
- `Content-Security-Policy` (CSP) with a restrictive default policy
(`default-src 'self'` as a baseline, tightened per-resource as
needed). Never use `unsafe-inline` or `unsafe-eval` unless
unavoidable, and document the reason.
- `X-Frame-Options: DENY` (or `SAMEORIGIN` if framing is required).
Prefer the `frame-ancestors` CSP directive as the primary control.
- `X-Content-Type-Options: nosniff`.
- `Referrer-Policy: strict-origin-when-cross-origin` (or stricter).
- `Permissions-Policy` restricting access to browser features the
application does not use (camera, microphone, geolocation, etc.).
- **Request and response limits:**
- Maximum request body size enforced on all endpoints (e.g. Go
`http.MaxBytesReader`). Choose a sane default per-route; never accept
unbounded input.
- Maximum response body size where applicable (e.g. paginated APIs).
- `ReadTimeout` and `ReadHeaderTimeout` on the `http.Server` to defend
against slowloris attacks.
- `WriteTimeout` on the `http.Server`.
- `IdleTimeout` on the `http.Server`.
- Per-handler execution time limits via `context.WithTimeout` or
chi/stdlib `middleware.Timeout`.
- **Authentication and session security:**
- Rate limiting on password-based authentication endpoints. API keys are
high-entropy and not susceptible to brute force, so they are exempt.
- CSRF tokens on all state-mutating HTML forms. API endpoints
authenticated via `Authorization` header (Bearer token, API key) are
exempt because the browser does not attach these automatically.
- Passwords stored using bcrypt, scrypt, or argon2 — never plain-text,
MD5, or SHA.
- Session cookies set with `HttpOnly`, `Secure`, and `SameSite=Lax` (or
`Strict`) attributes.
- **Reverse proxy awareness:**
- True client IP detection when behind a reverse proxy
(`X-Forwarded-For`, `X-Real-IP`). The application must accept
forwarded headers only from a configured set of trusted proxy
addresses — never trust `X-Forwarded-For` unconditionally.
- **CORS:**
- Authenticated endpoints must restrict `Access-Control-Allow-Origin` to
an explicit allowlist of known origins. Wildcard (`*`) is acceptable
only for public, unauthenticated read-only APIs.
- **Error handling:**
- Internal errors must never leak stack traces, SQL queries, file paths,
or other implementation details to the client. Return generic error
messages in production; detailed errors only when `DEBUG` is enabled.
- **TLS:**
- Services never terminate TLS directly. They are always deployed behind
a TLS-terminating reverse proxy. The service itself listens on plain
HTTP. However, HSTS headers and `Secure` cookie flags must still be
set by the application so that the browser enforces HTTPS end-to-end.
This list is non-exhaustive. Apply defense-in-depth: if a standard security
hardening measure exists for HTTP services and is not listed here, it is
still expected. When in doubt, harden.
- `README.md` is the primary documentation. Required sections:
- **Description**: First line must include the project name, purpose,
category (web server, SPA, CLI tool, etc.), license, and author. Example:
"µPaaS is an MIT-licensed Go web application by @sneak that receives
git-frontend webhooks and deploys applications via Docker in realtime."
- **Getting Started**: Copy-pasteable install/usage code block.
- **Rationale**: Why does this exist?
- **Design**: How is the program structured?
- **TODO**: Update meticulously, even between commits. When planning, put
the todo list in the README so a new agent can pick up where the last one
left off.
- **License**: MIT, GPL, or WTFPL. Ask the user for new projects. Include a
`LICENSE` file in the repo root and a License section in the README.
- **Author**: [@sneak](https://sneak.berlin).
- First commit of a new repo should contain only `README.md`.
- Go module root: `sneak.berlin/go/<name>`. Always run `go mod tidy` before
committing.
- Use SemVer.
- Database migrations live in `internal/db/migrations/` and must be embedded in
the binary.
- `000_migration.sql` — contains ONLY the creation of the migrations
tracking table itself. Nothing else.
- `001_schema.sql` — the full application schema.
- **Pre-1.0.0:** never add additional migration files (002, 003, etc.).
There is no installed base to migrate. Edit `001_schema.sql` directly.
- **Post-1.0.0:** add new numbered migration files for each schema change.
Never edit existing migrations after release.
- All repos should have an `.editorconfig` enforcing the project's indentation
settings.
- **Claude Code repo memory is versioned in the repo**, not left only in
`~/.claude` on one machine. Each memory is one file at
`.claude/memory/<memory>.md`, and every memory file must be `@`-imported
from `.claude/CLAUDE.md` (one `- @memory/<memory>.md` list line per file;
relative import paths resolve against `.claude/`, and Claude Code expands
the imports into context at session launch). When adding a memory, add both
the file and its import line. A root `MEMORY.md` is a violation — Claude
Code never auto-loads it; split it into `.claude/memory/` files. Repos with
no memories yet need no `.claude/` scaffolding.
- Avoid putting files in the repo root unless necessary. Root should contain
only project-level config files (`README.md`, `Makefile`, `Dockerfile`,
`LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`, and
language-specific config). Everything else goes in a subdirectory. Canonical
subdirectory names:
- `bin/` — executable scripts and tools
- `cmd/` — Go command entrypoints
- `configs/` — configuration templates and examples
- `deploy/` — deployment manifests (k8s, compose, terraform)
- `docs/` — documentation and markdown (README.md stays in root)
- `internal/` — Go internal packages
- `internal/db/migrations/` — database migrations
- `pkg/` — Go library packages
- `share/` — systemd units, data files
- `static/` — static assets (images, fonts, etc.)
- `web/` — web frontend source
- When setting up a new repo, files from the `prompts` repo may be used as
templates. Fetch them from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/<path>`.
- New repos must contain at minimum:
- `README.md`, `.git`, `.gitignore`, `.editorconfig`
- `LICENSE`, `REPO_POLICIES.md` (copy from the `prompts` repo)
- `Makefile`
- `Dockerfile`, `.dockerignore`
- `.gitea/workflows/check.yml`
- Go: `go.mod`, `go.sum`, `.golangci.yml`
- JS: `package.json`, `yarn.lock`, `.prettierrc`, `.prettierignore`
- Python: `pyproject.toml`

71
TODO.md
View File

@@ -14,21 +14,49 @@
# Next Step
- bring the repo into full policy compliance (work the Repo Policy
Compliance checklist below)
- convert Makefile targets to scripts-to-rule-them-all `script/`
entrypoints like the other managed repos
# Completed Steps
- announce each operand on stderr before its passes (2026-07-24,
branch `scan-operand-progress`): with per-operand walk/hash/update
cycles, a multi-operand run (e.g. `scan /srv/*`) showed pass totals
that looked like the whole run's — an operator watching operand 3 of
14 hash 300k files concluded 20M files were being skipped
- parallel walk (2026-07-24, branch `parallel-walk`): the walk pass
was a single goroutine and took hours at ~20M files on a busy pool
(observed: 22M files in 4h on a ZFS server); it is now a
per-directory worker-pool traversal that records size/mtime during
the walk (folding away the separate stat pass, halving metadata
I/O), and each `PATH` operand commits in its own transaction so an
interrupted scan keeps completed operands
- persistent scan database (2026-07-24, branch `persistent-database`):
`scan` now maintains a SQLite database (`modernc.org/sqlite`, pure
Go, cgo stays disabled) keyed by absolute path that survives between
runs — a rescan hashes only new or changed files (by mtime/size),
deletes records for files vanished from under the scanned operands,
and leaves records outside them untouched, so `scan` can be cronned
daily; `report` and `trees` read the database (no positional
arguments) instead of a scan stream. Database at
`/var/lib/sfdupes/db.sqlite`, overridable via `SFDUPES_DATABASE`;
WAL journaling plus a single-transaction update keep a report run
during a scan safe
- add the `origin` remote (`git@git.eeqj.de:sneak/sfdupes.git`), tag
`v0.0.1`, and push `main` plus tags (2026-07-23)
- `scan` CLI rework (2026-07-23, branch `scan-required-paths`): required
`PATH...` operands via cobra flags replacing the `/srv` `-root`
default; new `-x`/`--one-file-system` flag (GNU convention) to stop
at filesystem boundaries, which are crossed by default
- bring the repo into full policy compliance (2026-07-23, branch
`repo-policy-compliance`; checklist below)
- `git init` with README-only first commit; code baseline committed on
`main` (2026-07-22)
- implement `scan`, `report`, and `trees` subcommands (pre-git history)
# Future Steps
- add a remote on git.eeqj.de and push (`main` plus tags)
- convert Makefile targets to scripts-to-rule-them-all `script/`
entrypoints like the other managed repos
- tag `v0.0.1` once the compliance branch is merged
- possible later features (explicitly out of scope per README):
full-content verification of candidates, removal-script helpers
@@ -38,33 +66,34 @@ Audited 2026-07-22 against `REPO_POLICIES.md` (2026-07-06), the existing
repo checklist, and the Go styleguide. Code is already gofmt-clean, so no
standalone formatting commit is needed.
- [ ] `.gitignore` missing — the compiled `sfdupes` binary and
- [x] `.gitignore` missing — the compiled `sfdupes` binary and
`files.dat` sit untracked in the tree; needs OS/editor/Go
artifacts plus secrets patterns
- [ ] `.editorconfig` missing
- [ ] `LICENSE` missing and README has no License section (MIT assumed
- [x] `.editorconfig` missing
- [x] `LICENSE` missing and README has no License section (MIT assumed
from house convention — user to confirm)
- [ ] `REPO_POLICIES.md` missing from repo root
- [ ] `.golangci.yml` missing (install canonical copy); code must then
pass `make lint`
- [ ] `Makefile` lacks required targets `test`, `lint`, `fmt`,
- [x] `REPO_POLICIES.md` missing from repo root
- [x] `.golangci.yml` missing (install canonical copy); code must then
pass `make lint` (150 findings fixed; `make lint` is clean)
- [x] `Makefile` lacks required targets `test`, `lint`, `fmt`,
`fmt-check`, `docker`, `hooks`; `check` currently depends on
`build`, which writes the binary (`make check` must not modify
files)
- [ ] no tests — `go test ./...` has nothing to run; policy requires
- [x] no tests — `go test ./...` has nothing to run; policy requires
real tests with a 30-second timeout and the conditional `-v`
rerun pattern
- [ ] `Dockerfile` missing — Go multistage with hash-pinned images:
rerun pattern (suite covers parsing, grouping, digests,
suppression, hashing, and the scan pipeline; 64% coverage)
- [x] `Dockerfile` missing — Go multistage with hash-pinned images:
fail-fast lint stage, build stage running `make check`
- [ ] `.dockerignore` missing
- [ ] `.gitea/workflows/check.yml` missing (`docker build .` on push,
- [x] `.dockerignore` missing
- [x] `.gitea/workflows/check.yml` missing (`docker build .` on push,
checkout action pinned by commit SHA)
- [ ] README lacks required sections: Description first line
- [x] README lacks required sections: Description first line
(name/purpose/category/license/author), Getting Started,
Rationale, TODO, License, Author
- [ ] README non-goal "no git repository setup and no CI" is stale now
- [x] README non-goal "no git repository setup and no CI" is stale now
that the repo is under git with CI
- [ ] pre-commit hook not installed (`make hooks` once the target
- [x] pre-commit hook not installed (`make hooks` once the target
exists)
Accepted divergences (no action):

319
db.go Normal file
View File

@@ -0,0 +1,319 @@
package main
import (
"context"
"database/sql"
"errors"
"fmt"
"io/fs"
"os"
"path/filepath"
"strconv"
// The pure-Go SQLite driver, registered as "sqlite"; keeps cgo
// disabled.
_ "modernc.org/sqlite"
)
// defaultDatabasePath is where the persistent scan database lives when
// SFDUPES_DATABASE is not set.
const defaultDatabasePath = "/var/lib/sfdupes/db.sqlite"
// databaseEnv is the environment variable that overrides the database
// path.
const databaseEnv = "SFDUPES_DATABASE"
// schemaVersion is the database schema version this build reads and
// writes, stored in PRAGMA user_version.
const schemaVersion = 1
// dbDirPerm is the mode for a database parent directory created by
// scan.
const dbDirPerm = 0o755
// createTableSQL is the schema applied to a fresh database. Paths are
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
const createTableSQL = `
CREATE TABLE files (
path BLOB PRIMARY KEY,
size INTEGER NOT NULL,
mtime INTEGER NOT NULL,
head TEXT NOT NULL,
tail TEXT NOT NULL
) WITHOUT ROWID
`
// upsertSQL inserts one file record, replacing any existing record for
// the same path.
const upsertSQL = `
INSERT INTO files (path, size, mtime, head, tail)
VALUES (?, ?, ?, ?, ?)
ON CONFLICT (path) DO UPDATE SET
size = excluded.size, mtime = excluded.mtime,
head = excluded.head, tail = excluded.tail
`
// errNoDatabase reports a missing database file for report/trees.
var errNoDatabase = errors.New(
"no database (run \"sfdupes scan\" first, or set " + databaseEnv + ")")
// errSchemaVersion reports a database whose schema version this build
// does not understand.
var errSchemaVersion = errors.New("unsupported database schema version")
// databasePath resolves the database location: SFDUPES_DATABASE when
// set and non-empty, the compiled-in default otherwise.
func databasePath() string {
if p := os.Getenv(databaseEnv); p != "" {
return p
}
return defaultDatabasePath
}
// openDB opens the SQLite database at path with WAL journaling and a
// busy timeout, so a report can run while a cron scan is in progress.
// It does not create or verify the schema.
func openDB(path string) (*sql.DB, error) {
dsn := "file:" + path +
"?_pragma=busy_timeout(10000)" +
"&_pragma=journal_mode(WAL)" +
"&_pragma=synchronous(NORMAL)"
db, err := sql.Open("sqlite", dsn)
if err != nil {
return nil, fmt.Errorf("open database %s: %w", path, err)
}
// A single connection avoids SQLITE_BUSY between this process's
// own connections; concurrency lives in the worker pools, not in
// parallel database access.
db.SetMaxOpenConns(1)
return db, nil
}
// openScanDatabase opens the database for the scan subcommand, creating
// the file, its parent directory, and the schema as needed.
func openScanDatabase(path string) (*sql.DB, error) {
err := os.MkdirAll(filepath.Dir(path), dbDirPerm)
if err != nil {
return nil, fmt.Errorf("create database directory: %w", err)
}
db, err := openDB(path)
if err != nil {
return nil, err
}
err = initSchema(db)
if err != nil {
_ = db.Close()
return nil, fmt.Errorf("database %s: %w", path, err)
}
return db, nil
}
// openReportDatabase opens an existing database for the report and
// trees subcommands. A missing database file is an error directing the
// user to run scan first; the schema version must match exactly.
func openReportDatabase(path string) (*sql.DB, error) {
_, err := os.Stat(path)
if errors.Is(err, fs.ErrNotExist) {
return nil, fmt.Errorf("%s: %w", path, errNoDatabase)
}
if err != nil {
return nil, fmt.Errorf("database: %w", err)
}
db, err := openDB(path)
if err != nil {
return nil, err
}
v, err := userVersion(db)
if err != nil {
_ = db.Close()
return nil, fmt.Errorf("database %s: %w", path, err)
}
if v != schemaVersion {
_ = db.Close()
return nil, fmt.Errorf("database %s: version %d, want %d: %w",
path, v, schemaVersion, errSchemaVersion)
}
return db, nil
}
// initSchema creates the schema on a fresh database and verifies the
// schema version on an existing one.
func initSchema(db *sql.DB) error {
v, err := userVersion(db)
if err != nil {
return err
}
switch v {
case 0:
return createSchema(db)
case schemaVersion:
return nil
default:
return fmt.Errorf("version %d, want %d: %w",
v, schemaVersion, errSchemaVersion)
}
}
// createSchema applies the schema to a fresh database and stamps the
// schema version.
func createSchema(db *sql.DB) error {
ctx := context.Background()
_, err := db.ExecContext(ctx, createTableSQL)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
_, err = db.ExecContext(ctx,
"PRAGMA user_version = "+strconv.Itoa(schemaVersion))
if err != nil {
return fmt.Errorf("set schema version: %w", err)
}
return nil
}
// userVersion reads the database's PRAGMA user_version.
func userVersion(db *sql.DB) (int, error) {
var v int
err := db.QueryRowContext(context.Background(),
"PRAGMA user_version").Scan(&v)
if err != nil {
return 0, fmt.Errorf("read schema version: %w", err)
}
return v, nil
}
// loadFileRows reads every record from the files table.
func loadFileRows(db *sql.DB) ([]scanRec, error) {
rows, err := db.QueryContext(context.Background(),
"SELECT path, size, mtime, head, tail FROM files")
if err != nil {
return nil, fmt.Errorf("read records: %w", err)
}
defer func() { _ = rows.Close() }()
var recs []scanRec
for rows.Next() {
var (
path []byte
r scanRec
)
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail)
if err != nil {
return nil, fmt.Errorf("read record: %w", err)
}
r.path = string(path)
recs = append(recs, r)
}
err = rows.Err()
if err != nil {
return nil, fmt.Errorf("read records: %w", err)
}
return recs, nil
}
// applyChanges writes one scan's database changes — upserts for new and
// changed files, deletes for vanished ones — in a single transaction,
// so a concurrent report never observes a half-finished scan. Progress
// is rendered on prog (one increment per change).
func applyChanges(db *sql.DB, upserts []scanRec, deletes []string,
prog *progress,
) error {
ctx := context.Background()
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return fmt.Errorf("begin transaction: %w", err)
}
defer func() { _ = tx.Rollback() }()
err = execUpserts(ctx, tx, upserts, prog)
if err != nil {
return err
}
err = execDeletes(ctx, tx, deletes, prog)
if err != nil {
return err
}
err = tx.Commit()
if err != nil {
return fmt.Errorf("commit: %w", err)
}
return nil
}
// execUpserts inserts or updates one record per new or changed file.
func execUpserts(ctx context.Context, tx *sql.Tx, upserts []scanRec,
prog *progress,
) error {
st, err := tx.PrepareContext(ctx, upsertSQL)
if err != nil {
return fmt.Errorf("prepare upsert: %w", err)
}
defer func() { _ = st.Close() }()
for _, r := range upserts {
_, err = st.ExecContext(ctx,
[]byte(r.path), r.size, r.mtime, r.head, r.tail)
if err != nil {
return fmt.Errorf("upsert %s: %w", r.path, err)
}
prog.increment()
}
return nil
}
// execDeletes removes the records for paths no longer present.
func execDeletes(ctx context.Context, tx *sql.Tx, deletes []string,
prog *progress,
) error {
st, err := tx.PrepareContext(ctx, "DELETE FROM files WHERE path = ?")
if err != nil {
return fmt.Errorf("prepare delete: %w", err)
}
defer func() { _ = st.Close() }()
for _, p := range deletes {
_, err = st.ExecContext(ctx, []byte(p))
if err != nil {
return fmt.Errorf("delete %s: %w", p, err)
}
prog.increment()
}
return nil
}

180
db_test.go Normal file
View File

@@ -0,0 +1,180 @@
package main
import (
"context"
"database/sql"
"errors"
"path/filepath"
"slices"
"strings"
"testing"
)
// testDBPath returns a database path inside a fresh temp dir.
func testDBPath(t *testing.T) string {
t.Helper()
return filepath.Join(t.TempDir(), "db.sqlite")
}
// openTestDB creates a fresh scan database in a temp dir.
func openTestDB(t *testing.T) *sql.DB {
t.Helper()
db, err := openScanDatabase(testDBPath(t))
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
return db
}
func TestDatabasePath(t *testing.T) {
t.Setenv(databaseEnv, "")
if got := databasePath(); got != defaultDatabasePath {
t.Errorf("databasePath() = %q, want %q", got, defaultDatabasePath)
}
t.Setenv(databaseEnv, "/custom/place.sqlite")
if got := databasePath(); got != "/custom/place.sqlite" {
t.Errorf("databasePath() = %q, want the env override", got)
}
}
func TestOpenScanDatabaseCreates(t *testing.T) {
t.Parallel()
// The parent directory does not exist yet; scan must create it.
path := filepath.Join(t.TempDir(), "nested", "dir", "db.sqlite")
db, err := openScanDatabase(path)
if err != nil {
t.Fatalf("openScanDatabase: %v", err)
}
v, err := userVersion(db)
if err != nil || v != schemaVersion {
t.Fatalf("userVersion = %d, %v; want %d, nil", v, err, schemaVersion)
}
_ = db.Close()
// Reopening an existing database must succeed and find the schema.
db, err = openScanDatabase(path)
if err != nil {
t.Fatalf("reopen: %v", err)
}
defer func() { _ = db.Close() }()
recs, err := loadFileRows(db)
if err != nil || len(recs) != 0 {
t.Fatalf("loadFileRows = %v, %v; want empty, nil", recs, err)
}
}
func TestOpenReportDatabaseMissing(t *testing.T) {
t.Parallel()
_, err := openReportDatabase(testDBPath(t))
if !errors.Is(err, errNoDatabase) {
t.Fatalf("err = %v, want errNoDatabase", err)
}
}
func TestOpenReportDatabaseVersionMismatch(t *testing.T) {
t.Parallel()
path := testDBPath(t)
db, err := openScanDatabase(path)
if err != nil {
t.Fatal(err)
}
_, err = db.ExecContext(context.Background(), "PRAGMA user_version = 99")
if err != nil {
t.Fatal(err)
}
_ = db.Close()
_, err = openReportDatabase(path)
if !errors.Is(err, errSchemaVersion) {
t.Fatalf("err = %v, want errSchemaVersion", err)
}
}
func TestOpenReportDatabaseOK(t *testing.T) {
t.Parallel()
path := testDBPath(t)
db, err := openScanDatabase(path)
if err != nil {
t.Fatal(err)
}
_ = db.Close()
db, err = openReportDatabase(path)
if err != nil {
t.Fatalf("openReportDatabase: %v", err)
}
_ = db.Close()
}
func TestApplyChangesRoundTrip(t *testing.T) {
t.Parallel()
db := openTestDB(t)
// Paths may contain tabs and newlines; the database must store
// them byte-exactly.
recs := []scanRec{
{size: 2, mtime: 20, head: "h2", tail: "t2", path: "/a/tab\tnew\nline"},
{size: 1, mtime: 10, head: "h1", tail: "t1", path: "/a/x"},
}
err := applyChanges(db, recs, nil, newProgress("update", 2))
if err != nil {
t.Fatalf("applyChanges: %v", err)
}
got, err := loadFileRows(db)
if err != nil {
t.Fatal(err)
}
slices.SortFunc(got, func(a, b scanRec) int {
return strings.Compare(a.path, b.path)
})
if !slices.Equal(got, recs) {
t.Fatalf("rows = %+v, want %+v", got, recs)
}
// An upsert for an existing path updates in place; a delete
// removes exactly its path.
upd := scanRec{size: 3, mtime: 30, head: "h3", tail: "t3", path: "/a/x"}
err = applyChanges(db, []scanRec{upd},
[]string{"/a/tab\tnew\nline"}, newProgress("update", 2))
if err != nil {
t.Fatalf("applyChanges: %v", err)
}
got, err = loadFileRows(db)
if err != nil {
t.Fatal(err)
}
if len(got) != 1 || got[0] != upd {
t.Fatalf("rows = %+v, want just %+v", got, upd)
}
}

9
go.mod
View File

@@ -5,13 +5,22 @@ go 1.25.7
require (
github.com/schollz/progressbar/v3 v3.19.1
github.com/spf13/cobra v1.10.2
modernc.org/sqlite v1.54.0
)
require (
github.com/dustin/go-humanize v1.0.1 // indirect
github.com/google/uuid v1.6.0 // indirect
github.com/inconshreveable/mousetrap v1.1.0 // indirect
github.com/mattn/go-isatty v0.0.22 // indirect
github.com/mitchellh/colorstring v0.0.0-20190213212951-d06e56a500db // indirect
github.com/ncruces/go-strftime v1.0.0 // indirect
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/rivo/uniseg v0.4.7 // indirect
github.com/spf13/pflag v1.0.9 // indirect
golang.org/x/sys v0.46.0 // indirect
golang.org/x/term v0.44.0 // indirect
modernc.org/libc v1.74.1 // indirect
modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect
)

48
go.sum
View File

@@ -3,14 +3,28 @@ github.com/chengxilo/virtualterm v1.0.4/go.mod h1:DyxxBZz/x1iqJjFxTFcr6/x+jSpqN0
github.com/cpuguy83/go-md2man/v2 v2.0.6/go.mod h1:oOW0eioCTA6cOiMLiUPZOpcVxMig6NIQQ7OS05n1F4g=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/dustin/go-humanize v1.0.1 h1:GzkhY7T5VNhEkwH0PVJgjz+fX1rhBrR7pRT3mDkpeCY=
github.com/dustin/go-humanize v1.0.1/go.mod h1:Mu1zIs6XwVuF/gI1OepvI0qD18qycQx+mFykh5fBlto=
github.com/google/pprof v0.0.0-20250317173921-a4b03ec1a45e h1:ijClszYn+mADRFY17kjQEVQ1XRhq2/JR1M3sGqeJoxs=
github.com/google/pprof v0.0.0-20250317173921-a4b03ec1a45e/go.mod h1:boTsfXsheKC2y+lKOCMpSfarhxDeIzfZG1jqGcPl3cA=
github.com/google/uuid v1.6.0 h1:NIvaJDMOsjHA8n1jAhLSgzrAzy1Hgr+hNrb57e+94F0=
github.com/google/uuid v1.6.0/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
github.com/hashicorp/golang-lru/v2 v2.0.7 h1:a+bsQ5rvGLjzHuww6tVxozPZFVghXaHOwFs4luLUK2k=
github.com/hashicorp/golang-lru/v2 v2.0.7/go.mod h1:QeFd9opnmA6QUJc5vARoKUSoFhyfM2/ZepoAG6RGpeM=
github.com/inconshreveable/mousetrap v1.1.0 h1:wN+x4NVGpMsO7ErUn/mUI3vEoE6Jt13X2s0bqwp9tc8=
github.com/inconshreveable/mousetrap v1.1.0/go.mod h1:vpF70FUmC8bwa3OWnCshd2FqLfsEA9PFc4w1p2J65bw=
github.com/mattn/go-isatty v0.0.22 h1:j8l17JJ9i6VGPUFUYoTUKPSgKe/83EYU2zBC7YNKMw4=
github.com/mattn/go-isatty v0.0.22/go.mod h1:ZXfXG4SQHsB/w3ZeOYbR0PrPwLy+n6xiMrJlRFqopa4=
github.com/mattn/go-runewidth v0.0.16 h1:E5ScNMtiwvlvB5paMFdw9p4kSQzbXFikJ5SQO6TULQc=
github.com/mattn/go-runewidth v0.0.16/go.mod h1:Jdepj2loyihRzMpdS35Xk/zdY8IAYHsh153qUoGf23w=
github.com/mitchellh/colorstring v0.0.0-20190213212951-d06e56a500db h1:62I3jR2EmQ4l5rM/4FEfDWcRD+abF5XlKShorW5LRoQ=
github.com/mitchellh/colorstring v0.0.0-20190213212951-d06e56a500db/go.mod h1:l0dey0ia/Uv7NcFFVbCLtqEBQbrT4OCwCSKTEv6enCw=
github.com/ncruces/go-strftime v1.0.0 h1:HMFp8mLCTPp341M/ZnA4qaf7ZlsbTc+miZjCLOFAw7w=
github.com/ncruces/go-strftime v1.0.0/go.mod h1:Fwc5htZGVVkseilnfgOVb9mKy6w1naJmn9CehxcKcls=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec h1:W09IVJc94icq4NjY3clb7Lk8O1qJ8BdBEF8z0ibU0rE=
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec/go.mod h1:qqbHyh8v60DhA7CoWK5oRCqLrMHRGoxYCSS9EjAz6Eo=
github.com/rivo/uniseg v0.4.7 h1:WUdvkW8uEhrYfLC4ZzdpI2ztxP1I582+49Oc5Mq64VQ=
github.com/rivo/uniseg v0.4.7/go.mod h1:FN3SvrM+Zdj16jyLfmOkMNblXMcoc8DfTHruCPUcx88=
github.com/russross/blackfriday/v2 v2.1.0/go.mod h1:+Rmxgy9KzJVeS9/2gXHxylqXiyQDYRxCVz55jmeOWTM=
@@ -23,10 +37,44 @@ github.com/spf13/pflag v1.0.9/go.mod h1:McXfInJRrz4CZXVZOBLb0bTZqETkiAhM9Iw0y3An
github.com/stretchr/testify v1.9.0 h1:HtqpIVDClZ4nwg75+f6Lvsy/wHu+3BoSGCbBAcpTsTg=
github.com/stretchr/testify v1.9.0/go.mod h1:r2ic/lqez/lEtzL7wO/rwa5dbSLXVDPFyf8C91i36aY=
go.yaml.in/yaml/v3 v3.0.4/go.mod h1:DhzuOOF2ATzADvBadXxruRBLzYTpT36CKvDb3+aBEFg=
golang.org/x/mod v0.37.0 h1:vF1DjpVEshcIqoEaauuHebaLk1O1forxjxBaVn884JQ=
golang.org/x/mod v0.37.0/go.mod h1:m8S8VeM9r4dzDwjrKO0a1sZP3YjeMamRRlD+fmR2Q/0=
golang.org/x/sync v0.21.0 h1:HLII4xRRTtCRkxYp4HNFF0Js/Og6q2i++KXbg0gHCwM=
golang.org/x/sync v0.21.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.46.0 h1:noSf2Fq6F8DBgS+LysIkx7rIExoNHJsxOAtPp4rthXw=
golang.org/x/sys v0.46.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/term v0.44.0 h1:0rLvDRCtNj0gZkyIXhCyOb2OAzEhLVqc4B+hrsBhrmc=
golang.org/x/term v0.44.0/go.mod h1:7ze4MdzUzLXpSAoFP1H0bOI9aXDqveSvatT5vKcFh2Y=
golang.org/x/tools v0.47.0 h1:7Kn5x/d1svx/PzryTsqeoZN4TZwqeH5pGWjefhLi/1Q=
golang.org/x/tools v0.47.0/go.mod h1:dFHnyTvFWY212G+h7ZY4Vsp/K3U4/7W9TyVaAul8uCA=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
modernc.org/cc/v4 v4.29.0 h1:CXgwL8cvxmyzBQZzbSl/6xFtMCryb6u8IOqDci39cgc=
modernc.org/cc/v4 v4.29.0/go.mod h1:OnovgIhbbMXMu1aISnJ0wvVD1KnW+cAUJkIrAWh+kVI=
modernc.org/ccgo/v4 v4.34.6 h1:sBgfIwyN0TQ9C5hwIeuqyeAKyMWnbvj2fvpF4L11uzU=
modernc.org/ccgo/v4 v4.34.6/go.mod h1:SZ8YcN9NG7XVsQYdm6jYBvi8PQP1qi+kqB6OhjqI3Fk=
modernc.org/fileutil v1.4.0 h1:j6ZzNTftVS054gi281TyLjHPp6CPHr2KCxEXjEbD6SM=
modernc.org/fileutil v1.4.0/go.mod h1:EqdKFDxiByqxLk8ozOxObDSfcVOv/54xDs/DUHdvCUU=
modernc.org/gc/v2 v2.6.5 h1:nyqdV8q46KvTpZlsw66kWqwXRHdjIlJOhG6kxiV/9xI=
modernc.org/gc/v2 v2.6.5/go.mod h1:YgIahr1ypgfe7chRuJi2gD7DBQiKSLMPgBQe9oIiito=
modernc.org/gc/v3 v3.1.4 h1:2g65LGVSmFQrXeITAw97x7hCRvZFcyE1uDP+7Vng7JI=
modernc.org/gc/v3 v3.1.4/go.mod h1:HFK/6AGESC7Ex+EZJhJ2Gni6cTaYpSMmU/cT9RmlfYY=
modernc.org/goabi0 v0.2.0 h1:HvEowk7LxcPd0eq6mVOAEMai46V+i7Jrj13t4AzuNks=
modernc.org/goabi0 v0.2.0/go.mod h1:CEFRnnJhKvWT1c1JTI3Avm+tgOWbkOu5oPA8eH8LnMI=
modernc.org/libc v1.74.1 h1:bdR4VTKFMC4966QSNZ05XLGI/VwzVa2kTUX51Dm0riQ=
modernc.org/libc v1.74.1/go.mod h1:uH4t5bOx3G3g9Xcmj10YKlTcVISlRDwv8VoQJG9n8Os=
modernc.org/mathutil v1.7.1 h1:GCZVGXdaN8gTqB1Mf/usp1Y/hSqgI2vAGGP4jZMCxOU=
modernc.org/mathutil v1.7.1/go.mod h1:4p5IwJITfppl0G4sUEDtCr4DthTaT47/N3aT6MhfgJg=
modernc.org/memory v1.11.0 h1:o4QC8aMQzmcwCK3t3Ux/ZHmwFPzE6hf2Y5LbkRs+hbI=
modernc.org/memory v1.11.0/go.mod h1:/JP4VbVC+K5sU2wZi9bHoq2MAkCnrt2r98UGeSK7Mjw=
modernc.org/opt v0.2.0 h1:tGyef5ApycA7FSEOMraay9SaTk5zmbx7Tu+cJs4QKZg=
modernc.org/opt v0.2.0/go.mod h1:03fq9lsNfvkYSfxrfUhZCWPk1lm4cq4N+Bh//bEtgns=
modernc.org/sortutil v1.2.1 h1:+xyoGf15mM3NMlPDnFqrteY07klSFxLElE2PVuWIJ7w=
modernc.org/sortutil v1.2.1/go.mod h1:7ZI3a3REbai7gzCLcotuw9AC4VZVpYMjDzETGsSMqJE=
modernc.org/sqlite v1.54.0 h1:JCxR4qwkJvOaqAoYcgDoO25Nc+ROg6EJ2LfBVzdrgog=
modernc.org/sqlite v1.54.0/go.mod h1:4ntCLuNmnH8+GNqjka1wNg7KJd5/Hi5FYp8K+XQ7GZw=
modernc.org/strutil v1.2.1 h1:UneZBkQA+DX2Rp35KcM69cSsNES9ly8mQWD71HKlOA0=
modernc.org/strutil v1.2.1/go.mod h1:EHkiggD70koQxjVdSBM3JKM7k6L0FbGE5eymy9i3B9A=
modernc.org/token v1.1.0 h1:Xl7Ap9dKaEs5kLoOQeQmPWevfnk/DM5qcLcYlA8ys6Y=
modernc.org/token v1.1.0/go.mod h1:UGzOrNV1mAFSEB63lOFHIpNRUVMvYTc6yu1SMY/XTDM=

85
main.go
View File

@@ -2,13 +2,15 @@
// very large filesystems without reading full file contents. Files are
// considered duplicates when they have identical size, identical SHA-256
// of their first 1024 bytes, and identical SHA-256 of their last 1024
// bytes.
// bytes. scan maintains a persistent SQLite database of file signatures
// (SFDUPES_DATABASE, default /var/lib/sfdupes/db.sqlite) that the
// reporting subcommands read.
//
// Usage:
//
// sfdupes scan [-root /srv] [-workers N] > files.dat
// sfdupes report [files.dat|-] > dupes.tsv
// sfdupes trees [files.dat|-] > dupetrees.tsv
// sfdupes scan [--workers N] [-x] PATH...
// sfdupes report > dupes.tsv
// sfdupes trees > dupetrees.tsv
//
// See README.md for the complete specification.
package main
@@ -16,19 +18,35 @@ package main
import (
"fmt"
"os"
"runtime"
"github.com/spf13/cobra"
)
// Exit codes: 0 is success (even with per-file warnings), exitFatal is
// a fatal error, exitUsage is a usage error.
const (
exitFatal = 1
exitUsage = 2
)
// Version is the build version, injected at link time via -ldflags
// (see the Makefile); "dev" for a plain go build.
//
//nolint:gochecknoglobals // written only by the linker
var Version = "dev"
func main() {
root := &cobra.Command{
Use: "sfdupes",
Short: "Find candidate duplicate files by size and head/tail SHA-256",
Args: cobra.NoArgs,
Run: func(cmd *cobra.Command, args []string) {
Use: "sfdupes",
Short: "Find candidate duplicate files by size and head/tail SHA-256",
Version: Version,
Args: cobra.NoArgs,
Run: func(cmd *cobra.Command, _ []string) {
// A missing subcommand prints usage and exits 2.
_ = cmd.Usage()
os.Exit(2)
os.Exit(exitUsage)
},
}
// Everything on stdout is machine-readable data; all human-facing
@@ -37,45 +55,54 @@ func main() {
root.SetErr(os.Stderr)
root.CompletionOptions.DisableDefaultCmd = true
var (
scanWorkers int
scanOneFS bool
)
scanCmd := &cobra.Command{
Use: "scan [-root /srv] [-workers N]",
Short: "Walk a tree and emit one record per regular file on stdout",
// The README specifies single-dash flags (-root, -workers);
// parse them with the stdlib flag package inside runScan.
DisableFlagParsing: true,
Run: func(cmd *cobra.Command, args []string) {
runScan(args)
Use: "scan [--workers N] [-x] PATH...",
Short: "Walk trees and synchronize the scan database",
Args: cobra.MinimumNArgs(1),
Run: func(_ *cobra.Command, args []string) {
runScan(args, scanWorkers, scanOneFS)
},
}
scanCmd.Flags().IntVar(&scanWorkers, "workers", runtime.NumCPU(),
"concurrent workers for the stat and hash passes")
scanCmd.Flags().BoolVarP(&scanOneFS, "one-file-system", "x", false,
"do not cross filesystem boundaries")
reportCmd := &cobra.Command{
Use: "report [files.dat|-]",
Short: "Read a scan stream and print the file-level duplicates report",
Args: cobra.MaximumNArgs(1),
Run: func(cmd *cobra.Command, args []string) {
runReport(args)
Use: "report",
Short: "Read the scan database and print the file-level duplicates report",
Args: cobra.NoArgs,
Run: func(_ *cobra.Command, _ []string) {
runReport()
},
}
treesCmd := &cobra.Command{
Use: "trees [files.dat|-]",
Short: "Read a scan stream and print the duplicate-tree report",
Args: cobra.MaximumNArgs(1),
Run: func(cmd *cobra.Command, args []string) {
runTrees(args)
Use: "trees",
Short: "Read the scan database and print the duplicate-tree report",
Args: cobra.NoArgs,
Run: func(_ *cobra.Command, _ []string) {
runTrees()
},
}
root.AddCommand(scanCmd, reportCmd, treesCmd)
if err := root.Execute(); err != nil {
err := root.Execute()
if err != nil {
// Cobra has already printed the error and usage to stderr;
// an invalid subcommand or bad arguments is a usage error.
os.Exit(2)
os.Exit(exitUsage)
}
}
// fatalf reports a fatal error and exits 1.
func fatalf(format string, args ...any) {
fmt.Fprintf(os.Stderr, "sfdupes: "+format+"\n", args...)
os.Exit(1)
os.Exit(exitFatal)
}

View File

@@ -12,12 +12,23 @@ import (
// is not a TTY.
const plainInterval = 5 * time.Second
// barThrottle caps how often the TTY progress bar redraws.
const barThrottle = 100 * time.Millisecond
// walkSpinnerType selects the progressbar library's spinner style used
// for the walk pass, which has no known total.
const walkSpinnerType = 14
// percentScale converts a fraction to a percentage.
const percentScale = 100
// stderrIsTTY reports whether stderr is attached to a terminal.
func stderrIsTTY() bool {
fi, err := os.Stderr.Stat()
if err != nil {
return false
}
return fi.Mode()&os.ModeCharDevice != 0
}
@@ -42,6 +53,7 @@ func newProgress(label string, total int64) *progress {
if !stderrIsTTY() {
return p
}
opts := []progressbar.Option{
progressbar.OptionSetWriter(os.Stderr),
progressbar.OptionSetDescription(label),
@@ -49,7 +61,7 @@ func newProgress(label string, total int64) *progress {
progressbar.OptionShowIts(),
progressbar.OptionSetItsString("files"),
progressbar.OptionSetElapsedTime(true),
progressbar.OptionThrottle(100 * time.Millisecond),
progressbar.OptionThrottle(barThrottle),
}
if total >= 0 {
opts = append(opts,
@@ -59,10 +71,12 @@ func newProgress(label string, total int64) *progress {
} else {
opts = append(opts,
progressbar.OptionSetPredictTime(false),
progressbar.OptionSpinnerType(14),
progressbar.OptionSpinnerType(walkSpinnerType),
)
}
p.bar = progressbar.NewOptions64(total, opts...)
return p
}
@@ -71,8 +85,10 @@ func (p *progress) increment() {
p.count++
if p.bar != nil {
_ = p.bar.Add(1)
return
}
if time.Since(p.last) >= plainInterval {
p.last = time.Now()
fmt.Fprintln(os.Stderr, p.plainLine())
@@ -84,6 +100,7 @@ func (p *progress) warnf(format string, args ...any) {
if p.bar != nil {
_ = p.bar.Clear()
}
fmt.Fprintf(os.Stderr, format+"\n", args...)
}
@@ -91,9 +108,12 @@ func (p *progress) warnf(format string, args ...any) {
func (p *progress) finish() {
if p.bar != nil {
_ = p.bar.Finish()
fmt.Fprintln(os.Stderr)
return
}
fmt.Fprintln(os.Stderr, p.plainLine())
}
@@ -104,19 +124,24 @@ func (p *progress) plainLine() string {
return fmt.Sprintf("%s: %d files, elapsed %s",
p.label, p.count, elapsed.Round(time.Second))
}
pct := 100.0
pct := float64(percentScale)
if p.total > 0 {
pct = 100 * float64(p.count) / float64(p.total)
pct = percentScale * float64(p.count) / float64(p.total)
}
rate := 0.0
if s := elapsed.Seconds(); s > 0 {
rate = float64(p.count) / s
}
eta := "?"
if rate > 0 && p.count <= p.total {
remaining := time.Duration(float64(p.total-p.count) / rate * float64(time.Second))
eta = remaining.Round(time.Second).String()
}
return fmt.Sprintf("%s: [%d/%d] %.0f%% %.0f files/s elapsed %s eta %s",
p.label, p.count, p.total, pct, rate,
elapsed.Round(time.Second), eta)

204
report.go
View File

@@ -2,71 +2,49 @@ package main
import (
"bufio"
"bytes"
"fmt"
"io"
"os"
"slices"
"strconv"
"strings"
)
// scanRec is one well-formed record parsed from a scan stream. The
// signature (size, head, tail) is the duplicate key; mtime is
// informational and not retained.
// ioBufSize is the buffer size for the buffered stdout writers.
const ioBufSize = 1 << 20
// minGroupSize is the smallest number of members that makes a
// duplicate group.
const minGroupSize = 2
// scanRec is one file record from the database. The signature (size,
// head, tail) is the duplicate key; mtime is informational only and
// used by scan for change detection.
type scanRec struct {
size int64
head string
tail string
path string
size int64
mtime int64
head string
tail string
path string
}
// openScanInput resolves the analysis-mode input: the file named by the
// single optional positional argument, or stdin when it is absent or
// "-". The returned closer must be called when reading is done.
func openScanInput(args []string) (in io.Reader, name string, closer func()) {
if len(args) == 1 && args[0] != "-" {
f, err := os.Open(args[0])
if err != nil {
fatalf("open %s: %v", args[0], err)
}
return f, args[0], func() { f.Close() }
}
return os.Stdin, "stdin", func() {}
}
// loadRecords opens the database and reads every file record for the
// report and trees subcommands. Any database problem — including a
// missing database — is fatal.
func loadRecords() []scanRec {
dbPath := databasePath()
// parseScanStream reads every NUL-terminated record from in. A record
// that does not have exactly 5 fields or whose size is non-numeric is
// counted as malformed and skipped. Total records read is
// len(recs) + malformed.
func parseScanStream(in io.Reader, name string) (recs []scanRec, malformed int) {
sc := bufio.NewScanner(bufio.NewReaderSize(in, 1<<20))
sc.Buffer(make([]byte, 0, 64<<10), 4<<20)
sc.Split(splitNUL)
for sc.Scan() {
// The path is the last field and may itself contain tabs,
// so split into at most 5 fields.
fields := strings.SplitN(sc.Text(), "\t", 5)
if len(fields) != 5 {
malformed++
continue
}
size, err := strconv.ParseInt(fields[0], 10, 64)
if err != nil {
malformed++
continue
}
recs = append(recs, scanRec{
size: size,
head: fields[2],
tail: fields[3],
path: fields[4],
})
db, err := openReportDatabase(dbPath)
if err != nil {
fatalf("%v", err)
}
if err := sc.Err(); err != nil {
fatalf("read %s: %v", name, err)
defer func() { _ = db.Close() }()
recs, err := loadFileRows(db)
if err != nil {
fatalf("database %s: %v", dbPath, err)
}
return recs, malformed
return recs
}
// dupeGroup is one set of candidate-duplicate files: identical size,
@@ -77,93 +55,85 @@ type dupeGroup struct {
paths []string
}
// runReport implements the report subcommand: it reads a scan stream
// from the named file (or stdin when absent or "-") and prints the
// file-level duplicates report as TSV on stdout. It never touches the
// scanned filesystem; its only I/O is the scan input, stdout, and
// stderr. args holds the positional arguments already validated by
// cobra (at most one).
func runReport(args []string) {
in, name, closer := openScanInput(args)
defer closer()
recs, malformed := parseScanStream(in, name)
records := len(recs) + malformed
// runReport implements the report subcommand: it reads every record
// from the database and prints the file-level duplicates report as TSV
// on stdout. It never touches the scanned filesystem; its only I/O is
// the database, stdout, and stderr.
func runReport() {
recs := loadRecords()
dupes := collectDupeGroups(recs)
type dupeKey struct {
size int64
head string
tail string
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
_, err := fmt.Fprintln(out, "first\tdupe\tsize")
if err != nil {
fatalf("write stdout: %v", err)
}
groups := make(map[dupeKey][]string)
dupeFiles := 0
var reclaimable int64
for _, g := range dupes {
for _, p := range g.paths[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\n",
g.paths[0], p, g.size)
if err != nil {
fatalf("write stdout: %v", err)
}
dupeFiles++
reclaimable += g.size
}
}
err = out.Flush()
if err != nil {
fatalf("write stdout: %v", err)
}
fmt.Fprintf(os.Stderr,
"report: %d records read, %d duplicate groups, %d dupe files, "+
"%s reclaimable\n",
len(recs), len(dupes), dupeFiles, humanBytes(reclaimable))
}
// collectDupeGroups groups records by signature and returns every group
// with two or more paths, each group's paths sorted lexicographically,
// groups ordered by size descending then by first path ascending.
func collectDupeGroups(recs []scanRec) []dupeGroup {
groups := make(map[fileSig][]string)
for _, r := range recs {
k := dupeKey{size: r.size, head: r.head, tail: r.tail}
k := fileSig{size: r.size, head: r.head, tail: r.tail}
groups[k] = append(groups[k], r.path)
}
var dupes []dupeGroup
for k, paths := range groups {
if len(paths) < 2 {
if len(paths) < minGroupSize {
continue
}
slices.Sort(paths)
dupes = append(dupes, dupeGroup{size: k.size, paths: paths})
}
// Biggest reclaimable space first; ties broken by first path.
slices.SortFunc(dupes, func(a, b dupeGroup) int {
if a.size != b.size {
if a.size > b.size {
return -1
}
return 1
}
return strings.Compare(a.paths[0], b.paths[0])
})
out := bufio.NewWriterSize(os.Stdout, 1<<20)
if _, err := fmt.Fprintln(out, "first\tdupe\tsize"); err != nil {
fatalf("write stdout: %v", err)
}
dupeFiles := 0
var reclaimable int64
for _, g := range dupes {
for _, p := range g.paths[1:] {
if _, err := fmt.Fprintf(out, "%s\t%s\t%d\n",
g.paths[0], p, g.size); err != nil {
fatalf("write stdout: %v", err)
}
dupeFiles++
reclaimable += g.size
}
}
if err := out.Flush(); err != nil {
fatalf("write stdout: %v", err)
}
fmt.Fprintf(os.Stderr,
"report: %d records read%s, %d duplicate groups, %d dupe files, %s reclaimable\n",
records, malformedNote(malformed), len(dupes), dupeFiles,
humanBytes(reclaimable))
}
// malformedNote formats the optional malformed-record note for the
// stderr summaries.
func malformedNote(malformed int) string {
if malformed == 0 {
return ""
}
return fmt.Sprintf(" (%d malformed, skipped)", malformed)
}
// splitNUL is a bufio.SplitFunc for NUL-terminated records. Trailing
// data without a terminator at EOF is returned as a final record.
func splitNUL(data []byte, atEOF bool) (advance int, token []byte, err error) {
if i := bytes.IndexByte(data, 0); i >= 0 {
return i + 1, data[:i], nil
}
if atEOF && len(data) > 0 {
return len(data), data, nil
}
return 0, nil, nil
return dupes
}
// humanBytes formats a byte count in human units (binary prefixes).
@@ -172,10 +142,12 @@ func humanBytes(n int64) string {
if n < unit {
return fmt.Sprintf("%d B", n)
}
div, exp := int64(unit), 0
for m := n / unit; m >= unit; m /= unit {
div *= unit
exp++
}
return fmt.Sprintf("%.1f %ciB", float64(n)/float64(div), "KMGTPE"[exp])
}

126
report_test.go Normal file
View File

@@ -0,0 +1,126 @@
package main
import (
"slices"
"testing"
)
func TestCollectDupeGroups(t *testing.T) {
t.Parallel()
recs := []scanRec{
{size: 100, head: "h", tail: "t", path: "/z/b"},
{size: 100, head: "h", tail: "t", path: "/z/a"},
{size: 100, head: "h", tail: "t", path: "/z/c"},
{size: 4000, head: "H", tail: "T", path: "/big/2"},
{size: 4000, head: "H", tail: "T", path: "/big/1"},
// Same size as the /z group but a different head hash.
{size: 100, head: "other", tail: "t", path: "/z/d"},
// A singleton signature must not form a group.
{size: 7, head: "u", tail: "u", path: "/lonely"},
}
groups := collectDupeGroups(recs)
if len(groups) != 2 {
t.Fatalf("len(groups) = %d, want 2", len(groups))
}
if groups[0].size != 4000 ||
!slices.Equal(groups[0].paths, []string{"/big/1", "/big/2"}) {
t.Errorf("groups[0] = %+v, want size 4000, paths /big/1 /big/2",
groups[0])
}
if groups[1].size != 100 ||
!slices.Equal(groups[1].paths, []string{"/z/a", "/z/b", "/z/c"}) {
t.Errorf("groups[1] = %+v, want size 100, paths /z/a /z/b /z/c",
groups[1])
}
}
func TestCollectDupeGroupsMtimeExcluded(t *testing.T) {
t.Parallel()
// mtime is informational only; records differing only in mtime
// still group together.
recs := []scanRec{
{size: 9, mtime: 100, head: "h", tail: "t", path: "/m/1"},
{size: 9, mtime: 200, head: "h", tail: "t", path: "/m/2"},
}
groups := collectDupeGroups(recs)
if len(groups) != 1 {
t.Fatalf("len(groups) = %d, want 1", len(groups))
}
}
func TestCollectDupeGroupsTieBreak(t *testing.T) {
t.Parallel()
recs := []scanRec{
{size: 50, head: "b", tail: "b", path: "/beta/2"},
{size: 50, head: "b", tail: "b", path: "/beta/1"},
{size: 50, head: "a", tail: "a", path: "/alpha/2"},
{size: 50, head: "a", tail: "a", path: "/alpha/1"},
}
groups := collectDupeGroups(recs)
if len(groups) != 2 {
t.Fatalf("len(groups) = %d, want 2", len(groups))
}
// Equal sizes: ordered by first path ascending.
if groups[0].paths[0] != "/alpha/1" || groups[1].paths[0] != "/beta/1" {
t.Fatalf("tie-break order wrong: %q then %q",
groups[0].paths[0], groups[1].paths[0])
}
}
func TestCollectDupeGroupsDeterministic(t *testing.T) {
t.Parallel()
recs := []scanRec{
{size: 1, head: "a", tail: "a", path: "/p/1"},
{size: 1, head: "a", tail: "a", path: "/p/2"},
{size: 2, head: "b", tail: "b", path: "/q/1"},
{size: 2, head: "b", tail: "b", path: "/q/2"},
}
forward := collectDupeGroups(recs)
reversed := slices.Clone(recs)
slices.Reverse(reversed)
backward := collectDupeGroups(reversed)
if !slices.EqualFunc(forward, backward, func(a, b dupeGroup) bool {
return a.size == b.size && slices.Equal(a.paths, b.paths)
}) {
t.Fatalf("output depends on record order: %+v vs %+v",
forward, backward)
}
}
func TestHumanBytes(t *testing.T) {
t.Parallel()
cases := []struct {
n int64
want string
}{
{0, "0 B"},
{1, "1 B"},
{1023, "1023 B"},
{1024, "1.0 KiB"},
{1536, "1.5 KiB"},
{1 << 20, "1.0 MiB"},
{5 << 30, "5.0 GiB"},
{1 << 40, "1.0 TiB"},
{1 << 50, "1.0 PiB"},
{1 << 60, "1.0 EiB"},
}
for _, c := range cases {
if got := humanBytes(c.n); got != c.want {
t.Errorf("humanBytes(%d) = %q, want %q", c.n, got, c.want)
}
}
}

666
scan.go
View File

@@ -1,20 +1,31 @@
package main
import (
"bufio"
"crypto/sha256"
"database/sql"
"encoding/hex"
"flag"
"errors"
"fmt"
"io/fs"
"os"
"path/filepath"
"runtime"
"slices"
"strings"
"sync"
"syscall"
)
// chunk is the number of bytes hashed from each end of a file.
const chunk = 1024
// workQueueDepth bounds the job and result channels feeding the stat
// and hash worker pools.
const workQueueDepth = 1024
// errNotRegular reports a path whose type changed between the
// directory read and its lstat.
var errNotRegular = errors.New("no longer a regular file")
// fileRec carries one file between the stat and hash passes.
type fileRec struct {
path string
@@ -22,183 +33,559 @@ type fileRec struct {
mtime int64
}
// runScan implements the scan subcommand: three sequential passes (walk,
// stat, hash) over the tree under -root, emitting one NUL-terminated
// record per regular file on stdout.
func runScan(args []string) {
fl := flag.NewFlagSet("scan", flag.ExitOnError)
root := fl.String("root", "/srv", "directory tree to scan")
workers := fl.Int("workers", runtime.NumCPU(),
"concurrent workers for the stat and hash passes")
fl.Parse(args)
if fl.NArg() > 0 {
fmt.Fprintf(os.Stderr, "scan: unexpected argument %q\n", fl.Arg(0))
os.Exit(2)
}
if *workers < 1 {
*workers = 1
// runScan implements the scan subcommand: four sequential passes
// (walk, stat, hash, update) that synchronize the persistent database
// with the filesystem state under the PATH operands. Flag parsing and
// the at-least-one-operand check are done by cobra.
func runScan(roots []string, workers int, oneFS bool) {
if workers < 1 {
workers = 1
}
st, err := os.Stat(*root)
roots = resolveRoots(roots)
dbPath := databasePath()
db, err := openScanDatabase(dbPath)
if err != nil {
fatalf("root %s: %v", *root, err)
}
if !st.IsDir() {
fatalf("root %s: not a directory", *root)
fatalf("%v", err)
}
paths, walkErrs := walkPass(*root)
recs, statErrs := statPass(paths, *workers)
emitted, hashErrs := hashPass(recs, *workers)
defer func() { _ = db.Close() }()
skipped := walkErrs + statErrs + hashErrs
fmt.Fprintf(os.Stderr, "scan: %d files emitted, %d skipped\n",
emitted, skipped)
st, err := syncScan(db, roots, workers, oneFS)
if err != nil {
fatalf("update database %s: %v", dbPath, err)
}
fmt.Fprintf(os.Stderr,
"scan: %d files seen (%d added, %d updated, %d removed, "+
"%d unchanged), %d skipped\n",
st.added+st.updated+st.unchanged, st.added, st.updated,
st.removed, st.unchanged, st.skipped)
}
// walkPass enumerates every regular file under root. It never follows
// symlinks, never descends into directories named .zfs, and warns and
// continues on any per-path error.
func walkPass(root string) (paths []string, errs int) {
prog := newProgress("walk", -1)
walkErr := filepath.WalkDir(root, func(p string, d fs.DirEntry, err error) error {
// resolveRoots converts each PATH operand to an absolute, lexically
// cleaned path (symlinks are not resolved) and verifies that it
// exists. Database records are keyed by absolute path, so scan results
// must not depend on the working directory.
func resolveRoots(roots []string) []string {
abs := make([]string, 0, len(roots))
for _, root := range roots {
a, err := filepath.Abs(root)
if err != nil {
errs++
prog.warnf("walk %s: %v", p, err)
if d != nil && d.IsDir() {
return filepath.SkipDir
}
return nil
fatalf("resolve %s: %v", root, err)
}
if d.IsDir() {
// ZFS snapshot pseudo-dirs would list every file
// once per snapshot; never descend.
if d.Name() == ".zfs" {
return filepath.SkipDir
}
return nil
// A nonexistent operand is a fatal error before any scanning.
_, err = os.Lstat(a)
if err != nil {
fatalf("%v", err)
}
// Regular files only: skip symlinks, sockets, FIFOs, and
// device nodes.
if !d.Type().IsRegular() {
return nil
}
paths = append(paths, p)
prog.increment()
return nil
})
prog.finish()
if walkErr != nil {
fatalf("walk %s: %v", root, walkErr)
abs = append(abs, a)
}
return paths, errs
return abs
}
// statPass lstats every collected path in a worker pool, recording size
// and mtime. Paths that fail to stat (or are no longer regular files)
// are warned about and dropped.
func statPass(paths []string, workers int) (recs []fileRec, errs int) {
type result struct {
rec fileRec
err error
}
jobs := make(chan string, 1024)
results := make(chan result, 1024)
for range workers {
go func() {
for p := range jobs {
fi, err := os.Lstat(p)
switch {
case err != nil:
results <- result{rec: fileRec{path: p}, err: err}
case !fi.Mode().IsRegular():
results <- result{
rec: fileRec{path: p},
err: fmt.Errorf("no longer a regular file"),
}
default:
results <- result{rec: fileRec{
path: p,
size: fi.Size(),
mtime: fi.ModTime().Unix(),
}}
}
}
}()
}
go func() {
for _, p := range paths {
jobs <- p
}
close(jobs)
}()
// scanStats summarizes one scan's database synchronization for the
// final stderr summary.
type scanStats struct {
added int
updated int
removed int
unchanged int
skipped int
}
prog := newProgress("stat", int64(len(paths)))
recs = make([]fileRec, 0, len(paths))
for range paths {
r := <-results
if r.err != nil {
errs++
prog.warnf("stat %s: %v", r.rec.path, r.err)
} else {
recs = append(recs, r.rec)
// syncScan synchronizes the database with the filesystem under roots.
// Operands are processed in order, each one walked, hashed, and
// committed independently, so an interrupted scan keeps every operand
// completed so far. Records outside the roots are never touched.
func syncScan(db *sql.DB, roots []string, workers int,
oneFS bool,
) (scanStats, error) {
var st scanStats
for i, root := range roots {
// Each operand runs its own walk/hash/update sequence, so
// announce it: without this, the per-operand pass totals look
// like the whole run's.
fmt.Fprintf(os.Stderr, "scan: %s (operand %d of %d)\n",
root, i+1, len(roots))
err := syncRoot(db, root, workers, oneFS, &st)
if err != nil {
return st, err
}
}
return st, nil
}
// syncRoot walks one PATH operand, hashes its new and changed files,
// and commits the operand's record insertions, updates, and deletions
// in a single transaction.
func syncRoot(db *sql.DB, root string, workers int, oneFS bool,
st *scanStats,
) error {
existing, err := loadScopedRows(db, root)
if err != nil {
return err
}
walked, walkErrs := walkPass(root, oneFS, workers)
toHash, unchanged := partitionChanged(walked, existing)
hashed, hashErrs := hashPass(toHash, workers)
st.unchanged += len(unchanged)
st.skipped += walkErrs + hashErrs
for _, r := range hashed {
if _, ok := existing[r.path]; ok {
st.updated++
} else {
st.added++
}
}
deletes := collectDeletes(existing, unchanged, hashed)
st.removed += len(deletes)
return applyPass(db, hashed, deletes)
}
// loadScopedRows loads the database records whose paths lie under
// root, keyed by path. Records outside the scanned operands belong to
// other trees and are left untouched.
func loadScopedRows(db *sql.DB, root string) (map[string]scanRec, error) {
all, err := loadFileRows(db)
if err != nil {
return nil, err
}
scoped := make(map[string]scanRec)
for _, r := range all {
if underRoot(r.path, root) {
scoped[r.path] = r
}
}
return scoped, nil
}
// underRoot reports whether path is root itself or lies under it. Both
// must be absolute and lexically clean.
func underRoot(path, root string) bool {
if path == root {
return true
}
prefix := root
if !strings.HasSuffix(prefix, "/") {
prefix += "/"
}
return strings.HasPrefix(path, prefix)
}
// partitionChanged splits the stat results into files that must be
// hashed (new, or changed) and files whose existing records are reused
// without reading them: a file is unchanged when its statted size
// equals the recorded size and its statted mtime is not newer than the
// recorded mtime.
func partitionChanged(recs []fileRec,
existing map[string]scanRec,
) ([]fileRec, []scanRec) {
var (
toHash []fileRec
unchanged []scanRec
)
for _, rec := range recs {
old, ok := existing[rec.path]
if ok && old.size == rec.size && old.mtime >= rec.mtime {
unchanged = append(unchanged, old)
continue
}
toHash = append(toHash, rec)
}
return toHash, unchanged
}
// collectDeletes returns the existing record paths that were not
// successfully processed this run: vanished files, plus paths that
// failed to stat or hash. The database keeps only records verified by
// the latest scan covering them. The result is sorted so the update
// pass is deterministic.
func collectDeletes(existing map[string]scanRec,
unchanged, hashed []scanRec,
) []string {
kept := make(map[string]bool, len(unchanged)+len(hashed))
for _, r := range unchanged {
kept[r.path] = true
}
for _, r := range hashed {
kept[r.path] = true
}
var deletes []string
for p := range existing {
if !kept[p] {
deletes = append(deletes, p)
}
}
slices.Sort(deletes)
return deletes
}
// applyPass writes the scan's changes to the database under an update
// progress display (one item per insertion, update, or deletion).
func applyPass(db *sql.DB, upserts []scanRec, deletes []string) error {
prog := newProgress("update", int64(len(upserts)+len(deletes)))
err := applyChanges(db, upserts, deletes, prog)
prog.finish()
return err
}
// dirJob is one directory awaiting traversal by the walk workers. It
// carries its operand's filesystem device so -x can stop at
// filesystem boundaries.
type dirJob struct {
path string
rootDev uint64
rootDevOK bool
}
// walkEvent is one walk result delivered to the main goroutine: a
// regular file's stat record, or a warning when fail is set.
type walkEvent struct {
rec fileRec
warn string
fail bool
}
// walkPass enumerates every regular file under root with a
// per-directory worker pool, recording size and mtime from lstat
// while each directory is fresh in cache. It never follows symlinks,
// never descends into directories named .zfs, and warns and continues
// on any per-path error. With oneFS set it never descends into a
// directory on a different filesystem than root.
func walkPass(root string, oneFS bool, workers int) ([]fileRec, int) {
prog := newProgress("walk", -1)
recs, initial, errs := seedRoot(root, prog)
jobs, subdirs, events := startWalkWorkers(workers, oneFS)
dispatchDirs(initial, jobs, subdirs)
for ev := range events {
if ev.fail {
errs++
prog.warnf("%s", ev.warn)
continue
}
recs = append(recs, ev.rec)
prog.increment()
}
prog.finish()
return recs, errs
}
// hashPass hashes the first and last chunk bytes of every file in a
// worker pool and emits the output records on stdout from the main
// goroutine. Files that fail to open or read are warned about and
// dropped.
func hashPass(recs []fileRec, workers int) (emitted, errs int) {
type result struct {
rec fileRec
head string
tail string
err error
// seedRoot turns the PATH operand into the walk's starting state: a
// regular-file operand becomes a record directly, a directory operand
// becomes the initial job, and a symlink or other non-regular operand
// yields nothing (symlinks are never followed, including as
// operands).
func seedRoot(root string, prog *progress) ([]fileRec, []dirJob, int) {
fi, err := os.Lstat(root)
if err != nil {
prog.warnf("walk %s: %v", root, err)
return nil, nil, 1
}
jobs := make(chan fileRec, 1024)
results := make(chan result, 1024)
switch {
case fi.IsDir():
if filepath.Base(root) == ".zfs" {
return nil, nil, 0
}
dev, ok := deviceOfInfo(fi)
return nil, []dirJob{{path: root, rootDev: dev, rootDevOK: ok}}, 0
case fi.Mode().IsRegular():
rec := fileRec{path: root, size: fi.Size(), mtime: fi.ModTime().Unix()}
prog.increment()
return []fileRec{rec}, nil, 0
default:
return nil, nil, 0
}
}
// startWalkWorkers starts the walk worker pool. Each worker processes
// one directory at a time, emitting an event per regular file and
// handing discovered subdirectories back to the dispatcher; events is
// closed once every worker has finished.
func startWalkWorkers(workers int,
oneFS bool,
) (chan dirJob, chan []dirJob, chan walkEvent) {
jobs := make(chan dirJob, workQueueDepth)
subdirs := make(chan []dirJob, workers)
events := make(chan walkEvent, workQueueDepth)
var wg sync.WaitGroup
for range workers {
wg.Go(func() {
for job := range jobs {
subdirs <- walkOneDir(job, oneFS, events)
}
})
}
go func() {
wg.Wait()
close(events)
}()
return jobs, subdirs, events
}
// dispatchDirs feeds directory jobs to the walk workers, queueing
// newly discovered subdirectories (newest first, which keeps the
// frontier small) until every directory has been processed, then
// closes jobs.
func dispatchDirs(initial []dirJob, jobs chan<- dirJob,
subdirs <-chan []dirJob,
) {
go func() {
queue := slices.Clone(initial)
pending := len(queue)
for pending > 0 {
var (
out chan<- dirJob
next dirJob
)
if len(queue) > 0 {
out = jobs
next = queue[len(queue)-1]
}
select {
case out <- next:
queue = queue[:len(queue)-1]
case subs := <-subdirs:
pending += len(subs) - 1
queue = append(queue, subs...)
}
}
close(jobs)
}()
}
// walkOneDir reads one directory, emitting an event per regular-file
// entry and a warning event per unreadable one, and returns the
// subdirectories to descend into.
func walkOneDir(job dirJob, oneFS bool, events chan<- walkEvent) []dirJob {
entries, err := os.ReadDir(job.path)
if err != nil {
events <- walkEvent{
warn: fmt.Sprintf("walk %s: %v", job.path, err),
fail: true,
}
return nil
}
var subs []dirJob
for _, e := range entries {
p := filepath.Join(job.path, e.Name())
if e.IsDir() {
if sub, ok := subdirJob(p, e, job, oneFS, events); ok {
subs = append(subs, sub)
}
continue
}
// Regular files only: skip symlinks, sockets, FIFOs, and
// device nodes.
if !e.Type().IsRegular() {
continue
}
events <- fileEvent(p, e)
}
return subs
}
// fileEvent lstats one regular-file directory entry into its walk
// event.
func fileEvent(p string, e fs.DirEntry) walkEvent {
fi, err := e.Info()
switch {
case err != nil:
return walkEvent{
warn: fmt.Sprintf("walk %s: %v", p, err),
fail: true,
}
case !fi.Mode().IsRegular():
return walkEvent{
warn: fmt.Sprintf("walk %s: %v", p, errNotRegular),
fail: true,
}
default:
return walkEvent{rec: fileRec{
path: p,
size: fi.Size(),
mtime: fi.ModTime().Unix(),
}}
}
}
// subdirJob applies the descent rules to directory p: never enter
// .zfs (ZFS snapshot pseudo-dirs would list every file once per
// snapshot), and with -x never enter a directory on a different
// filesystem than its operand.
func subdirJob(p string, e fs.DirEntry, parent dirJob, oneFS bool,
events chan<- walkEvent,
) (dirJob, bool) {
if e.Name() == ".zfs" {
return dirJob{}, false
}
job := dirJob{path: p, rootDev: parent.rootDev, rootDevOK: parent.rootDevOK}
if !oneFS || !parent.rootDevOK {
return job, true
}
info, err := e.Info()
if err != nil {
events <- walkEvent{
warn: fmt.Sprintf("walk %s: %v", p, err),
fail: true,
}
return dirJob{}, false
}
if dev, ok := deviceOfInfo(info); ok && dev != parent.rootDev {
return dirJob{}, false
}
return job, true
}
// deviceOfInfo extracts the filesystem device ID from a FileInfo, when
// the platform exposes one.
func deviceOfInfo(fi fs.FileInfo) (uint64, bool) {
st, ok := fi.Sys().(*syscall.Stat_t)
if !ok {
return 0, false
}
return statDev(st), true
}
// hashResult carries one file's head/tail hashes (or the error that
// prevented hashing it) from the hash workers to the main goroutine.
type hashResult struct {
rec fileRec
head string
tail string
err error
}
// startHashWorkers starts the hash worker pool over recs and returns
// the channel its results arrive on (one per record, in completion
// order).
func startHashWorkers(recs []fileRec, workers int) <-chan hashResult {
jobs := make(chan fileRec, workQueueDepth)
results := make(chan hashResult, workQueueDepth)
for range workers {
go func() {
for rec := range jobs {
head, tail, err := hashHeadTail(rec.path, rec.size)
results <- result{rec: rec, head: head, tail: tail, err: err}
results <- hashResult{rec: rec, head: head, tail: tail, err: err}
}
}()
}
go func() {
for _, rec := range recs {
jobs <- rec
}
close(jobs)
}()
return results
}
// hashPass hashes the first and last chunk bytes of every file in a
// worker pool and collects the resulting records on the main
// goroutine. Files that fail to open or read are warned about and
// dropped. Only new or changed files reach this pass.
func hashPass(recs []fileRec, workers int) ([]scanRec, int) {
results := startHashWorkers(recs, workers)
prog := newProgress("hash", int64(len(recs)))
out := bufio.NewWriterSize(os.Stdout, 1<<20)
out := make([]scanRec, 0, len(recs))
var errs int
for range recs {
r := <-results
if r.err != nil {
errs++
prog.warnf("hash %s: %v", r.rec.path, r.err)
prog.increment()
continue
}
if _, err := fmt.Fprintf(out, "%d\t%d\t%s\t%s\t%s\x00",
r.rec.size, r.rec.mtime, r.head, r.tail, r.rec.path); err != nil {
fatalf("write stdout: %v", err)
}
emitted++
out = append(out, scanRec{
size: r.rec.size,
mtime: r.rec.mtime,
head: r.head,
tail: r.tail,
path: r.rec.path,
})
prog.increment()
}
prog.finish()
if err := out.Flush(); err != nil {
fatalf("write stdout: %v", err)
}
return emitted, errs
return out, errs
}
// hashHeadTail returns the lowercase-hex SHA-256 of the first
@@ -206,26 +593,35 @@ func hashPass(recs []fileRec, workers int) (emitted, errs int) {
// file at path. The two reads overlap when size < 2*chunk; for
// size == 0 both hashes are of the empty input. size is the value
// recorded by the stat pass.
func hashHeadTail(path string, size int64) (head, tail string, err error) {
func hashHeadTail(path string, size int64) (string, string, error) {
//nolint:gosec // hashing operator-supplied paths is the tool's purpose
f, err := os.Open(path)
if err != nil {
return "", "", err
}
defer f.Close()
defer func() { _ = f.Close() }()
n := min(int64(chunk), size)
buf := make([]byte, n)
if n > 0 {
if _, err := f.ReadAt(buf, 0); err != nil {
_, err = f.ReadAt(buf, 0)
if err != nil {
return "", "", err
}
}
h := sha256.Sum256(buf)
if n > 0 {
if _, err := f.ReadAt(buf, size-n); err != nil {
_, err = f.ReadAt(buf, size-n)
if err != nil {
return "", "", err
}
}
t := sha256.Sum256(buf)
return hex.EncodeToString(h[:]), hex.EncodeToString(t[:]), nil
}

14
scan_dev_darwin.go Normal file
View File

@@ -0,0 +1,14 @@
//go:build darwin
package main
import "syscall"
// statDev returns the filesystem device ID of st as uint64. Device IDs
// are opaque; the conversion only needs to be consistent so equality
// comparisons work, not numerically meaningful.
//
//nolint:gosec // int32 device IDs convert consistently; only equality matters
func statDev(st *syscall.Stat_t) uint64 {
return uint64(st.Dev)
}

10
scan_dev_other.go Normal file
View File

@@ -0,0 +1,10 @@
//go:build !darwin
package main
import "syscall"
// statDev returns the filesystem device ID of st.
func statDev(st *syscall.Stat_t) uint64 {
return st.Dev
}

684
scan_test.go Normal file
View File

@@ -0,0 +1,684 @@
package main
import (
"context"
"crypto/sha256"
"database/sql"
"encoding/hex"
"fmt"
"os"
"path/filepath"
"slices"
"strings"
"testing"
"time"
)
// writeFile creates a file with the given content and returns its path.
func writeFile(t *testing.T, dir, name string, data []byte) string {
t.Helper()
p := filepath.Join(dir, name)
err := os.MkdirAll(filepath.Dir(p), 0o750)
if err != nil {
t.Fatal(err)
}
err = os.WriteFile(p, data, 0o600)
if err != nil {
t.Fatal(err)
}
return p
}
// hexSum returns the lowercase-hex SHA-256 of data.
func hexSum(data []byte) string {
s := sha256.Sum256(data)
return hex.EncodeToString(s[:])
}
// pattern returns n bytes of deterministic content seeded by tag.
func pattern(tag byte, n int) []byte {
data := make([]byte, n)
for i := range data {
data[i] = tag ^ byte(i)
}
return data
}
func TestHashHeadTail(t *testing.T) {
t.Parallel()
dir := t.TempDir()
cases := []struct {
name string
data []byte
}{
{"empty", nil},
{"one-byte", []byte("x")},
{"under-one-chunk", pattern(1, chunk-1)},
{"exactly-one-chunk", pattern(2, chunk)},
{"overlapping-reads", pattern(3, chunk+chunk/2)},
{"exactly-two-chunks", pattern(4, 2*chunk)},
{"beyond-two-chunks", pattern(5, 3*chunk)},
}
for _, c := range cases {
t.Run(c.name, func(t *testing.T) {
t.Parallel()
p := writeFile(t, dir, c.name, c.data)
head, tail, err := hashHeadTail(p, int64(len(c.data)))
if err != nil {
t.Fatalf("hashHeadTail: %v", err)
}
n := min(chunk, len(c.data))
if want := hexSum(c.data[:n]); head != want {
t.Errorf("head = %s, want %s", head, want)
}
if want := hexSum(c.data[len(c.data)-n:]); tail != want {
t.Errorf("tail = %s, want %s", tail, want)
}
})
}
}
func TestHashHeadTailErrors(t *testing.T) {
t.Parallel()
dir := t.TempDir()
_, _, err := hashHeadTail(filepath.Join(dir, "missing"), 1)
if err == nil {
t.Error("no error for a missing file")
}
// A file that shrank between the stat and hash passes: reading at
// the stat-reported size must fail rather than emit wrong hashes.
p := writeFile(t, dir, "shrunk", []byte("tiny"))
_, _, err = hashHeadTail(p, int64(2*chunk))
if err == nil {
t.Error("no error when the stat size exceeds the file size")
}
}
// walkedPaths returns the sorted paths of walked records.
func walkedPaths(recs []fileRec) []string {
paths := make([]string, 0, len(recs))
for _, r := range recs {
paths = append(paths, r.path)
}
slices.Sort(paths)
return paths
}
// walkedByPath finds the walked record with the given path.
func walkedByPath(t *testing.T, recs []fileRec, path string) fileRec {
t.Helper()
for _, r := range recs {
if r.path == path {
return r
}
}
t.Fatalf("no walked record for %q", path)
return fileRec{}
}
func TestWalkPass(t *testing.T) {
t.Parallel()
dir := t.TempDir()
want := []string{
writeFile(t, dir, "a.txt", []byte("a")),
writeFile(t, dir, "sub/b.txt", []byte("b")),
writeFile(t, dir, "sub/deeper/c.txt", []byte("c")),
}
slices.Sort(want)
// Files under a .zfs directory must never be walked.
writeFile(t, dir, ".zfs/snapshot/hourly/a.txt", []byte("a"))
// Symlinks are skipped, not followed.
err := os.Symlink(want[0], filepath.Join(dir, "link"))
if err != nil {
t.Fatal(err)
}
recs, errs := walkPass(dir, false, 4)
if errs != 0 {
t.Fatalf("errs = %d, want 0", errs)
}
if got := walkedPaths(recs); !slices.Equal(got, want) {
t.Fatalf("paths = %q, want %q", got, want)
}
// The walk records lstat sizes and mtimes.
if r := walkedByPath(t, recs, want[0]); r.size != 1 || r.mtime <= 0 {
t.Fatalf("rec = %+v, want size 1 and a positive mtime", r)
}
}
func TestWalkPassDeepAndWide(t *testing.T) {
t.Parallel()
// Exercise the dispatcher with more directories than workers and
// with nesting deeper than the worker count.
dir := t.TempDir()
deep := "deep" + strings.Repeat("/d", 30)
want := make([]string, 0, 41)
want = append(want, writeFile(t, dir, deep+"/f", []byte("x")))
for i := range 40 {
want = append(want, writeFile(t, dir,
fmt.Sprintf("wide/%02d/f", i), []byte("y")))
}
slices.Sort(want)
recs, errs := walkPass(dir, false, 8)
if errs != 0 {
t.Fatalf("errs = %d, want 0", errs)
}
if got := walkedPaths(recs); !slices.Equal(got, want) {
t.Fatalf("walked %d paths, want %d", len(got), len(want))
}
}
func TestWalkPassFileAndSymlinkOperands(t *testing.T) {
t.Parallel()
dir := t.TempDir()
f := writeFile(t, dir, "plain", []byte("data"))
link := filepath.Join(dir, "link")
err := os.Symlink(f, link)
if err != nil {
t.Fatal(err)
}
// A regular-file operand is emitted as itself.
recs, errs := walkPass(f, false, 2)
if errs != 0 || len(recs) != 1 || recs[0].path != f || recs[0].size != 4 {
t.Fatalf("file operand: recs = %+v, errs = %d", recs, errs)
}
// A symlink operand is not followed and yields nothing.
recs, errs = walkPass(link, false, 2)
if errs != 0 || len(recs) != 0 {
t.Fatalf("symlink operand: recs = %+v, errs = %d", recs, errs)
}
}
func TestWalkPassOneFilesystemSameFS(t *testing.T) {
t.Parallel()
// Everything in one filesystem: -x must not skip anything.
dir := t.TempDir()
want := []string{
writeFile(t, dir, "a", []byte("a")),
writeFile(t, dir, "sub/deep/b", []byte("b")),
}
recs, errs := walkPass(dir, true, 4)
if errs != 0 {
t.Fatalf("errs = %d, want 0", errs)
}
if got := walkedPaths(recs); !slices.Equal(got, want) {
t.Fatalf("paths = %q, want %q", got, want)
}
}
func TestDeviceOfInfo(t *testing.T) {
t.Parallel()
dir := t.TempDir()
fi1, err := os.Lstat(dir)
if err != nil {
t.Fatal(err)
}
fi2, err := os.Lstat(dir)
if err != nil {
t.Fatal(err)
}
dev1, ok1 := deviceOfInfo(fi1)
dev2, ok2 := deviceOfInfo(fi2)
if !ok1 || !ok2 || dev1 != dev2 {
t.Fatalf("deviceOfInfo unstable: %d/%v vs %d/%v",
dev1, ok1, dev2, ok2)
}
}
// buildSmokeTree recreates the README smoke-test filesystem layout
// with deterministic content and returns the tree root.
func buildSmokeTree(t *testing.T) string {
t.Helper()
dir := t.TempDir()
one := pattern(10, 2000)
f1 := pattern(30, 3000)
f2 := pattern(40, 100)
writeFile(t, dir, "a/one.bin", one)
writeFile(t, dir, "b/copy.bin", one)
writeFile(t, dir, "b/copy2.bin", one)
// Same size as one.bin, different content.
writeFile(t, dir, "a/unique.bin", pattern(20, 2000))
writeFile(t, dir, "tiny1", []byte("x"))
writeFile(t, dir, "tiny2", []byte("x"))
writeFile(t, dir, "tiny3", []byte("y"))
writeFile(t, dir, "empty1", nil)
writeFile(t, dir, "empty2", nil)
writeFile(t, dir, "t1/f1", f1)
writeFile(t, dir, "t1/sub/f2", f2)
writeFile(t, dir, "t2/f1", f1)
writeFile(t, dir, "t2/sub/f2", f2)
writeFile(t, dir, "t3/f1", f1)
writeFile(t, dir, "t3/sub/f2renamed", f2)
return dir
}
// smokeTreeFiles is the number of regular files buildSmokeTree creates.
const smokeTreeFiles = 15
// syncTree synchronizes the database with the given roots and returns
// the scan stats.
func syncTree(t *testing.T, db *sql.DB, roots ...string) scanStats {
t.Helper()
st, err := syncScan(db, roots, 4, false)
if err != nil {
t.Fatalf("syncScan: %v", err)
}
return st
}
// dbRecords returns every record currently in the database.
func dbRecords(t *testing.T, db *sql.DB) []scanRec {
t.Helper()
recs, err := loadFileRows(db)
if err != nil {
t.Fatal(err)
}
return recs
}
// recordByPath finds the record with the given path.
func recordByPath(t *testing.T, recs []scanRec, path string) scanRec {
t.Helper()
for _, r := range recs {
if r.path == path {
return r
}
}
t.Fatalf("no record for %q", path)
return scanRec{}
}
// recordPaths returns the sorted paths of recs.
func recordPaths(recs []scanRec) []string {
paths := make([]string, 0, len(recs))
for _, r := range recs {
paths = append(paths, r.path)
}
slices.Sort(paths)
return paths
}
// assertSmokeDupeGroups checks the file-level duplicate groups for the
// smoke tree rooted at dir.
func assertSmokeDupeGroups(t *testing.T, dir string, parsed []scanRec) {
t.Helper()
groups := collectDupeGroups(parsed)
if len(groups) != 5 {
t.Fatalf("len(groups) = %d, want 5", len(groups))
}
wantSizes := []int64{3000, 2000, 100, 1, 0}
for i, g := range groups {
if g.size != wantSizes[i] {
t.Errorf("groups[%d].size = %d, want %d",
i, g.size, wantSizes[i])
}
}
wantF1 := []string{
filepath.Join(dir, "t1/f1"),
filepath.Join(dir, "t2/f1"),
filepath.Join(dir, "t3/f1"),
}
if !slices.Equal(groups[0].paths, wantF1) {
t.Errorf("groups[0].paths = %q, want %q", groups[0].paths, wantF1)
}
}
// assertSmokeTreeGroups checks the duplicate-tree groups for the smoke
// tree rooted at dir.
func assertSmokeTreeGroups(t *testing.T, dir string, parsed []scanRec) {
t.Helper()
super, dirs := buildHierarchy(parsed)
super.compute()
tg := collectTreeGroups(dirs, super)
if len(tg) != 1 {
t.Fatalf("len(tree groups) = %d, want 1", len(tg))
}
wantTrees := []string{filepath.Join(dir, "t1"), filepath.Join(dir, "t2")}
if got := groupPaths(tg)[0]; !slices.Equal(got, wantTrees) {
t.Fatalf("tree group = %q, want %q", got, wantTrees)
}
if tg[0][0].fileCount != 2 || tg[0][0].totalSize != 3100 {
t.Fatalf("tree totals: %d files %d bytes, want 2 3100",
tg[0][0].fileCount, tg[0][0].totalSize)
}
}
func TestScanPipeline(t *testing.T) {
t.Parallel()
dir := buildSmokeTree(t)
db := openTestDB(t)
st := syncTree(t, db, dir)
if st != (scanStats{added: smokeTreeFiles}) {
t.Fatalf("stats = %+v, want %d added only", st, smokeTreeFiles)
}
parsed := dbRecords(t, db)
if len(parsed) != smokeTreeFiles {
t.Fatalf("len(records) = %d, want %d", len(parsed), smokeTreeFiles)
}
assertSmokeDupeGroups(t, dir, parsed)
assertSmokeTreeGroups(t, dir, parsed)
}
func TestSyncScanUnchangedReuse(t *testing.T) {
t.Parallel()
dir := t.TempDir()
db := openTestDB(t)
a := writeFile(t, dir, "a.bin", pattern(1, 500))
writeFile(t, dir, "b.bin", pattern(2, 600))
st := syncTree(t, db, dir)
if st != (scanStats{added: 2}) {
t.Fatalf("first scan stats = %+v, want 2 added", st)
}
// An immediate rescan reuses every record without reading file
// contents. Prove the files are not re-read by corrupting a stored
// hash and observing that it survives the rescan.
_, err := db.ExecContext(context.Background(),
"UPDATE files SET head = 'sentinel' WHERE path = ?", []byte(a))
if err != nil {
t.Fatal(err)
}
st = syncTree(t, db, dir)
if st != (scanStats{unchanged: 2}) {
t.Fatalf("rescan stats = %+v, want 2 unchanged", st)
}
if r := recordByPath(t, dbRecords(t, db), a); r.head != "sentinel" {
t.Fatalf("head = %q, want sentinel (file must not be re-read)",
r.head)
}
}
func TestSyncScanMtimeBump(t *testing.T) {
t.Parallel()
dir := t.TempDir()
db := openTestDB(t)
a := writeFile(t, dir, "a.bin", pattern(1, 500))
syncTree(t, db, dir)
// Bump the mtime forward: the file must be re-hashed even though
// its size is unchanged.
future := time.Now().Add(time.Hour)
err := os.Chtimes(a, future, future)
if err != nil {
t.Fatal(err)
}
st := syncTree(t, db, dir)
if st != (scanStats{updated: 1}) {
t.Fatalf("mtime-bump stats = %+v, want 1 updated", st)
}
if r := recordByPath(t, dbRecords(t, db), a); r.mtime != future.Unix() {
t.Fatalf("mtime = %d, want %d", r.mtime, future.Unix())
}
}
func TestSyncScanAddRemove(t *testing.T) {
t.Parallel()
dir := t.TempDir()
db := openTestDB(t)
a := writeFile(t, dir, "a.bin", pattern(1, 500))
writeFile(t, dir, "b.bin", pattern(2, 600))
syncTree(t, db, dir)
// Add one file, remove another.
c := writeFile(t, dir, "c.bin", pattern(3, 700))
err := os.Remove(a)
if err != nil {
t.Fatal(err)
}
st := syncTree(t, db, dir)
if st != (scanStats{added: 1, removed: 1, unchanged: 1}) {
t.Fatalf("add/remove stats = %+v, want 1 added 1 removed 1 unchanged",
st)
}
want := []string{filepath.Join(dir, "b.bin"), c}
if got := recordPaths(dbRecords(t, db)); !slices.Equal(got, want) {
t.Fatalf("paths = %q, want %q", got, want)
}
}
func TestSyncScanSizeChange(t *testing.T) {
t.Parallel()
dir := t.TempDir()
db := openTestDB(t)
p := writeFile(t, dir, "f", pattern(1, 100))
syncTree(t, db, dir)
// Rewrite with a different size but force the mtime back to the
// recorded value: the size mismatch alone must trigger a re-hash.
old := recordByPath(t, dbRecords(t, db), p)
writeFile(t, dir, "f", pattern(1, 200))
mt := time.Unix(old.mtime, 0)
err := os.Chtimes(p, mt, mt)
if err != nil {
t.Fatal(err)
}
st := syncTree(t, db, dir)
if st.updated != 1 {
t.Fatalf("stats = %+v, want 1 updated", st)
}
if got := recordByPath(t, dbRecords(t, db), p); got.size != 200 {
t.Fatalf("size = %d, want 200", got.size)
}
}
func TestSyncScanScope(t *testing.T) {
t.Parallel()
dir := t.TempDir()
db := openTestDB(t)
writeFile(t, dir, "a/keep", pattern(1, 10))
gone := writeFile(t, dir, "b/gone", pattern(2, 10))
syncTree(t, db, dir)
// Deleting a file outside the rescanned root must not remove its
// record: records outside the scanned operands are untouched.
err := os.Remove(gone)
if err != nil {
t.Fatal(err)
}
st := syncTree(t, db, filepath.Join(dir, "a"))
if st.removed != 0 || st.unchanged != 1 {
t.Fatalf("subtree stats = %+v, want 0 removed 1 unchanged", st)
}
if got := recordPaths(dbRecords(t, db)); len(got) != 2 {
t.Fatalf("records = %q, want both retained", got)
}
// Rescanning the parent now removes the vanished file's record.
st = syncTree(t, db, dir)
if st.removed != 1 {
t.Fatalf("parent stats = %+v, want 1 removed", st)
}
}
func TestSyncScanRemovesNonRegular(t *testing.T) {
t.Parallel()
dir := t.TempDir()
db := openTestDB(t)
p := writeFile(t, dir, "f", pattern(1, 10))
keep := writeFile(t, dir, "g", pattern(2, 10))
syncTree(t, db, dir)
// Replace the file with a symlink: it is no longer walked, so its
// record must be deleted.
err := os.Remove(p)
if err != nil {
t.Fatal(err)
}
err = os.Symlink(keep, p)
if err != nil {
t.Fatal(err)
}
st := syncTree(t, db, dir)
if st.removed != 1 || st.unchanged != 1 {
t.Fatalf("stats = %+v, want 1 removed 1 unchanged", st)
}
}
func TestSyncScanOverlappingRoots(t *testing.T) {
t.Parallel()
dir := t.TempDir()
db := openTestDB(t)
writeFile(t, dir, "sub/f", pattern(1, 10))
// A file reachable via two overlapping operands yields one
// record: the second operand sees the rows the first operand
// committed and reuses them unchanged, which also proves each
// operand commits before the next starts.
st := syncTree(t, db, dir, filepath.Join(dir, "sub"))
if st != (scanStats{added: 1, unchanged: 1}) {
t.Fatalf("stats = %+v, want 1 added 1 unchanged", st)
}
if got := recordPaths(dbRecords(t, db)); len(got) != 1 {
t.Fatalf("records = %q, want exactly one", got)
}
}
func TestReportsNeverTouchFilesystem(t *testing.T) {
t.Parallel()
dir := buildSmokeTree(t)
db := openTestDB(t)
syncTree(t, db, dir)
// Remove the scanned tree entirely; the analysis must be
// unaffected because it reads the database alone.
err := os.RemoveAll(dir)
if err != nil {
t.Fatal(err)
}
recs := dbRecords(t, db)
assertSmokeDupeGroups(t, dir, recs)
assertSmokeTreeGroups(t, dir, recs)
}
func TestUnderRoot(t *testing.T) {
t.Parallel()
const abRoot = "/a/b"
cases := []struct {
path string
root string
want bool
}{
{"/a/b/c", abRoot, true},
{abRoot, abRoot, true},
{"/a/bc", abRoot, false},
{"/a", abRoot, false},
{"/x/y", "/", true},
{"/", "/", true},
}
for _, c := range cases {
if got := underRoot(c.path, c.root); got != c.want {
t.Errorf("underRoot(%q, %q) = %v, want %v",
c.path, c.root, got, c.want)
}
}
}

132
trees.go
View File

@@ -28,26 +28,69 @@ type treeNode struct {
totalSize int64
}
// runTrees implements the trees subcommand: it reads a scan stream from
// the named file (or stdin when absent or "-"), reconstructs the
// directory hierarchy from the record paths, computes a Merkle-style
// digest per directory, and prints maximal duplicate-tree groups as TSV
// on stdout. It never touches the scanned filesystem; its only I/O is
// the scan input, stdout, and stderr. args holds the positional
// arguments already validated by cobra (at most one).
func runTrees(args []string) {
in, name, closer := openScanInput(args)
defer closer()
recs, malformed := parseScanStream(in, name)
records := len(recs) + malformed
// runTrees implements the trees subcommand: it reads every record from
// the database, reconstructs the directory hierarchy from the record
// paths, computes a Merkle-style digest per directory, and prints
// maximal duplicate-tree groups as TSV on stdout. It never touches the
// scanned filesystem; its only I/O is the database, stdout, and
// stderr.
func runTrees() {
recs := loadRecords()
// Build the hierarchy under a synthetic super-root. Paths are
// split on "/"; for absolute paths the first component is empty,
// which simply becomes a top-level node representing "/".
super, allDirs := buildHierarchy(recs)
super.compute()
dupes := collectTreeGroups(allDirs, super)
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
_, err := fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
if err != nil {
fatalf("write stdout: %v", err)
}
dupeTrees := 0
var reclaimable int64
for _, g := range dupes {
first := g[0]
for _, n := range g[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n",
first.path, n.path, first.fileCount, first.totalSize)
if err != nil {
fatalf("write stdout: %v", err)
}
dupeTrees++
reclaimable += first.totalSize
}
}
err = out.Flush()
if err != nil {
fatalf("write stdout: %v", err)
}
fmt.Fprintf(os.Stderr,
"trees: %d records read, %d duplicate tree groups, %d dupe trees, "+
"%s reclaimable\n",
len(recs), len(dupes), dupeTrees, humanBytes(reclaimable))
}
// buildHierarchy reconstructs the directory hierarchy from the record
// paths under a synthetic super-root. Paths are split on "/"; for
// absolute paths the first component is empty, which simply becomes a
// top-level node representing "/". It returns the super-root and every
// directory node created.
func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
super := &treeNode{}
var allDirs []*treeNode
for _, r := range recs {
comps := strings.Split(r.path, "/")
node := super
for _, c := range comps[:len(comps)-1] {
child := node.dirs[c]
@@ -56,76 +99,68 @@ func runTrees(args []string) {
if node != super {
childPath = node.path + "/" + c
}
child = &treeNode{path: childPath, parent: node}
if node.dirs == nil {
node.dirs = make(map[string]*treeNode)
}
node.dirs[c] = child
allDirs = append(allDirs, child)
}
node = child
}
if node.files == nil {
node.files = make(map[string]fileSig)
}
node.files[comps[len(comps)-1]] = fileSig{
size: r.size, head: r.head, tail: r.tail,
}
}
super.compute()
// Group directories by digest; two or more dirs sharing a digest
// are candidate duplicate trees.
return super, allDirs
}
// collectTreeGroups groups directories by digest and returns every
// maximal group with two or more members, each group's members sorted
// by path, groups ordered by tree size descending then by first path
// ascending.
func collectTreeGroups(allDirs []*treeNode, super *treeNode) [][]*treeNode {
groups := make(map[[sha256.Size]byte][]*treeNode)
for _, d := range allDirs {
groups[d.digest] = append(groups[d.digest], d)
}
var dupes [][]*treeNode
for _, g := range groups {
if len(g) < 2 || suppressed(g, super) {
if len(g) < minGroupSize || suppressed(g, super) {
continue
}
slices.SortFunc(g, func(a, b *treeNode) int {
return strings.Compare(a.path, b.path)
})
dupes = append(dupes, g)
}
// Biggest reclaimable space first; ties broken by first path.
slices.SortFunc(dupes, func(a, b []*treeNode) int {
if a[0].totalSize != b[0].totalSize {
if a[0].totalSize > b[0].totalSize {
return -1
}
return 1
}
return strings.Compare(a[0].path, b[0].path)
})
out := bufio.NewWriterSize(os.Stdout, 1<<20)
if _, err := fmt.Fprintln(out, "first\tdupe\tfiles\tsize"); err != nil {
fatalf("write stdout: %v", err)
}
dupeTrees := 0
var reclaimable int64
for _, g := range dupes {
first := g[0]
for _, n := range g[1:] {
if _, err := fmt.Fprintf(out, "%s\t%s\t%d\t%d\n",
first.path, n.path, first.fileCount, first.totalSize); err != nil {
fatalf("write stdout: %v", err)
}
dupeTrees++
reclaimable += first.totalSize
}
}
if err := out.Flush(); err != nil {
fatalf("write stdout: %v", err)
}
fmt.Fprintf(os.Stderr,
"trees: %d records read%s, %d duplicate tree groups, %d dupe trees, %s reclaimable\n",
records, malformedNote(malformed), len(dupes), dupeTrees,
humanBytes(reclaimable))
return dupes
}
// compute fills in digest, fileCount, and totalSize for n and all of
@@ -143,18 +178,22 @@ func (n *treeNode) compute() {
n.fileCount++
n.totalSize += sig.size
}
for name, child := range n.dirs {
child.compute()
entries = append(entries, "d\x00"+name+"\x00"+string(child.digest[:]))
n.fileCount += child.fileCount
n.totalSize += child.totalSize
}
slices.Sort(entries)
h := sha256.New()
for _, e := range entries {
h.Write([]byte(e))
h.Write([]byte{0})
}
copy(n.digest[:], h.Sum(nil))
}
@@ -165,15 +204,19 @@ func (n *treeNode) compute() {
// parent) or members whose parents differ are always reported.
func suppressed(g []*treeNode, super *treeNode) bool {
seen := make(map[*treeNode]bool, len(g))
var parentDigest [sha256.Size]byte
for i, n := range g {
p := n.parent
if p == nil || p == super {
return false
}
if seen[p] {
return false // siblings: not implied by any parent group
}
seen[p] = true
if i == 0 {
parentDigest = p.digest
@@ -181,5 +224,6 @@ func suppressed(g []*treeNode, super *treeNode) bool {
return false
}
}
return true
}

220
trees_test.go Normal file
View File

@@ -0,0 +1,220 @@
package main
import (
"slices"
"testing"
)
// Signature hashes shared by the smoke-test records.
const (
f1Head = "f1h"
f1Tail = "f1t"
f2Head = "f2h"
f2Tail = "f2t"
)
// smokeTreeRecs mirrors the README smoke-test tree layout: /d/t1 and
// /d/t2 are identical, /d/t3 differs from them only by one filename.
func smokeTreeRecs() []scanRec {
return []scanRec{
{size: 3000, head: f1Head, tail: f1Tail, path: "/d/t1/f1"},
{size: 100, head: f2Head, tail: f2Tail, path: "/d/t1/sub/f2"},
{size: 3000, head: f1Head, tail: f1Tail, path: "/d/t2/f1"},
{size: 100, head: f2Head, tail: f2Tail, path: "/d/t2/sub/f2"},
{size: 3000, head: f1Head, tail: f1Tail, path: "/d/t3/f1"},
{size: 100, head: f2Head, tail: f2Tail, path: "/d/t3/sub/f2renamed"},
}
}
// nodeByPath finds the directory node with the given path.
func nodeByPath(t *testing.T, dirs []*treeNode, path string) *treeNode {
t.Helper()
for _, d := range dirs {
if d.path == path {
return d
}
}
t.Fatalf("no directory node with path %q", path)
return nil
}
// groupPaths flattens tree groups into their member path lists.
func groupPaths(groups [][]*treeNode) [][]string {
out := make([][]string, 0, len(groups))
for _, g := range groups {
paths := make([]string, 0, len(g))
for _, n := range g {
paths = append(paths, n.path)
}
out = append(out, paths)
}
return out
}
func TestBuildHierarchyCounts(t *testing.T) {
t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs())
super.compute()
d := nodeByPath(t, dirs, "/d")
if d.fileCount != 6 || d.totalSize != 9300 {
t.Errorf("/d: fileCount %d size %d, want 6 9300",
d.fileCount, d.totalSize)
}
t1 := nodeByPath(t, dirs, "/d/t1")
if t1.fileCount != 2 || t1.totalSize != 3100 {
t.Errorf("/d/t1: fileCount %d size %d, want 2 3100",
t1.fileCount, t1.totalSize)
}
sub := nodeByPath(t, dirs, "/d/t1/sub")
if sub.fileCount != 1 || sub.totalSize != 100 {
t.Errorf("/d/t1/sub: fileCount %d size %d, want 1 100",
sub.fileCount, sub.totalSize)
}
}
func TestTreeDigests(t *testing.T) {
t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs())
super.compute()
t1 := nodeByPath(t, dirs, "/d/t1")
t2 := nodeByPath(t, dirs, "/d/t2")
t3 := nodeByPath(t, dirs, "/d/t3")
if t1.digest != t2.digest {
t.Error("identical trees /d/t1 and /d/t2 have different digests")
}
// t3 differs only in a filename; names are part of the digest.
if t1.digest == t3.digest {
t.Error("/d/t3 digest equals /d/t1 despite a renamed file")
}
sub1 := nodeByPath(t, dirs, "/d/t1/sub")
sub3 := nodeByPath(t, dirs, "/d/t3/sub")
if sub1.digest == sub3.digest {
t.Error("subdirs with differently-named files share a digest")
}
}
func TestTreeDigestContentSensitivity(t *testing.T) {
t.Parallel()
const sharedTail = "same"
recs := []scanRec{
{size: 10, head: sharedTail, tail: sharedTail, path: "/r/a/f"},
{size: 10, head: "DIFF", tail: sharedTail, path: "/r/b/f"},
}
super, dirs := buildHierarchy(recs)
super.compute()
a := nodeByPath(t, dirs, "/r/a")
b := nodeByPath(t, dirs, "/r/b")
if a.digest == b.digest {
t.Error("trees with different file content share a digest")
}
}
func TestCollectTreeGroupsMaximal(t *testing.T) {
t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs())
super.compute()
groups := collectTreeGroups(dirs, super)
want := [][]string{{"/d/t1", "/d/t2"}}
if got := groupPaths(groups); !slices.EqualFunc(got, want,
slices.Equal) {
t.Fatalf("groups = %v, want %v", got, want)
}
// The /d/t1/sub vs /d/t2/sub group must be suppressed as implied
// by its parents' group; the winning group reports one copy's
// recursive totals.
if groups[0][0].fileCount != 2 || groups[0][0].totalSize != 3100 {
t.Errorf("group totals: %d files %d bytes, want 2 3100",
groups[0][0].fileCount, groups[0][0].totalSize)
}
}
func TestCollectTreeGroupsDeterministic(t *testing.T) {
t.Parallel()
recs := smokeTreeRecs()
super, dirs := buildHierarchy(recs)
super.compute()
forward := groupPaths(collectTreeGroups(dirs, super))
reversed := slices.Clone(recs)
slices.Reverse(reversed)
superR, dirsR := buildHierarchy(reversed)
superR.compute()
backward := groupPaths(collectTreeGroups(dirsR, superR))
if !slices.EqualFunc(forward, backward, slices.Equal) {
t.Fatalf("output depends on record order: %v vs %v",
forward, backward)
}
}
func TestCollectTreeGroupsSiblings(t *testing.T) {
t.Parallel()
// Identical sibling dirs share a parent, so their group cannot be
// implied by a parent group and must be reported.
recs := []scanRec{
{size: 10, head: "h", tail: "t", path: "/p/x1/f"},
{size: 10, head: "h", tail: "t", path: "/p/x2/f"},
}
super, dirs := buildHierarchy(recs)
super.compute()
got := groupPaths(collectTreeGroups(dirs, super))
want := [][]string{{"/p/x1", "/p/x2"}}
if !slices.EqualFunc(got, want, slices.Equal) {
t.Fatalf("groups = %v, want %v", got, want)
}
}
func TestCollectTreeGroupsDifferingParents(t *testing.T) {
t.Parallel()
// /p/a and /q/b contain identical x subtrees, but /p/a has an
// extra file, so the parents' digests differ and the x group must
// be reported.
recs := []scanRec{
{size: 10, head: "h", tail: "t", path: "/p/a/x/f"},
{size: 99, head: "e", tail: "e", path: "/p/a/extra"},
{size: 10, head: "h", tail: "t", path: "/q/b/x/f"},
}
super, dirs := buildHierarchy(recs)
super.compute()
got := groupPaths(collectTreeGroups(dirs, super))
want := [][]string{{"/p/a/x", "/q/b/x"}}
if !slices.EqualFunc(got, want, slices.Equal) {
t.Fatalf("groups = %v, want %v", got, want)
}
}