Author SHA1 Message Date
clawbot 2acb657a44 Re-vendor the canonical files from sneak/prompts at c55a0cb (closes #95)
check / check (push) Canceled after 0s
Every vendored file, REPO_POLICIES.md and every model script is the
copy at sneak/prompts c55a0cb, with this repository's own entries kept
after the canonical content. Lint and test are phases of the Dockerfile
that write no image, built uncached. make test runs the suite under the
race detector as nobody, because root reads the files the tests make
unreadable. Dockerfile.lint, script/verify-lint-image-pin and
make test-race are gone. Prettier runs on the host, from the node and
yarn that script/bootstrap installs. golangci-lint v2.14.0 raises no
findings. .claude/settings.json is deleted.

Deviation: the set comes from c55a0cb on next rather than dd4027b, as
the instructions on sneak/prompts#78 allow.

Model: opus-5-5
2026-10-08 01:23:35 +00:00
clawbot 5900feb515 Store mtime to the nanosecond so a same-second rewrite is re-hashed (closes #12)
check / check (push) Waiting to run
scan recorded mtime in whole seconds, so a file rewritten in place at
the same size within the same second as its recorded mtime was classed
unchanged and kept its old hashes. The files table keeps mtime as whole
Unix seconds and gains mtime_nsec, the nanoseconds within that second.
scan holds the mtime as a time.Time and decides "newer" by comparing
Unix() and then Nanosecond(), so any time a filesystem can record
compares in the right order; After would misorder one too late for a
time.Time to hold without wrapping. The walk, a file given as an
operand, and the content phase's recheck all move over. PRAGMA
user_version stays 1, per the owner's ruling. README states what both
columns hold.

Model: opus-5-5
2026-10-08 02:51:53 +02:00
clawbot 0064eba542 Cut the narration from TODO.md and the script and Dockerfile comments (closes #49)
check / check (push) Failing after 1s
Completed Steps entries keep what landed, the traps, every disclosure
and every record that a check ran; the argument and history go, with
bare issue numbers turned into full links. Comment blocks in script/,
Dockerfile and Dockerfile.lint keep the trap and drop the defence of
past decisions. TODO.md Workflow now branches from next, targets next,
and leaves merging next to main to the owner. Only comments and
Markdown change.

Model: opus-5-5
2026-10-04 20:47:21 +02:00
clawbot 546203afe5 Run the tests under the race detector with make test-race (closes #18)
check / check (push) Failing after 3s
script/test-race runs go test -race in a digest-pinned Debian golang
image that has gcc, since the detector needs cgo and the build keeps it
off. The checkout is mounted read-only and the container is removed
afterwards. The tests run as the calling user, or as nobody when that is
root, so the tests that make a file unreadable still see the read fail.
It is not part of make check. The detector found no races.

Model: opus-5-5
2026-10-04 20:01:26 +02:00
clawbot bebfac1dcb Fail a bare docker build instead of serving cached gates (closes #39)
check / check (push) Failing after 3s
Each Dockerfile stage that runs gates now checks, right after its
ARG CHECK_EPOCH, that the value is not empty, and stops with a message
naming script/cibuild and script/docker. A plain `docker build .` can
no longer report a green from cached gate layers.

script/cibuild and script/docker now append the process id to the
epoch, the form script/lint already uses, so two runs started in the
same second still get different values.

README says both. TODO.md corrects the steady-state CACHED count
recorded for issue 32 from twelve to thirteen.

Model: opus-5-5
2026-10-04 19:30:20 +02:00
clawbot 313aa0fc12 Format Markdown with prettier in make fmt and make fmt-check (closes #19)
check / check (push) Failing after 3s
script/fmt and script/fmt-check run prettier over every Markdown file
again, next to gofmt. prettier is pinned by hash through package.json
and yarn.lock, copied from the prompts repo with .prettierrc and
.prettierignore, and is never installed on a host: a new prettier stage
of the Dockerfile installs it into a digest-pinned node image, and both
scripts build that stage and run it with the repository mounted. CI
checks the Markdown in a markdown stage that the build stage waits on.
Because make fmt-check now runs docker, the Dockerfile runs gofmt
directly in its lint stage instead. All Markdown is reformatted.

Model: opus-5-5
2026-10-04 18:30:25 +02:00
clawbot 1317d66589 Stop script/lint writing an image it never uses (closes #48)
check / check (push) Failing after 3s
script/lint builds Dockerfile.lint only for the exit status, but every
run exported the result as an image: seconds spent exporting, and one
untagged image left behind each time. It now builds with
--output=type=cacheonly, so nothing is exported. CHECK_EPOCH still
changes on every run, so the gate steps still run each time; the build
cache is kept as before.

Model: opus-5-5
2026-10-04 18:01:29 +02:00
clawbot 0b7078301d Test the remaining CLI cases (closes #16)
check / check (push) Failing after 3s
Most of the command-line contract was already tested through run. This
adds what was missing: report and trees with no database exit 1 with
the message telling the user to run scan; a scan that skips an
unreadable file prints its warning and counts the skip in its summary;
the report and trees summary lines are checked exactly. The scan tests
now capture the process's own stdout, which scan would write to
directly, so a stray stdout write in scan fails them. The fatal-path
test takes its subcommands from the command tree, so a new subcommand
wired without runE fails it. The nonexistent-operand test gets an
accurate name.

Model: opus-5-5
2026-10-04 17:47:26 +02:00
clawbot e2227ac07e Copy the current canonical .golangci.yml (closes #26)
check / check (push) Failing after 10s
The shared lint config in the prompts repo moved from the deprecated
gomodguard linter to gomodguard_v2, with a module block list, and now
enables depguard to keep test-support packages out of non-test files.
This replaces the repo's copy with that file unchanged, so lint no
longer prints the gomodguard deprecation warning.

Model: opus-5-5
2026-10-04 16:13:20 +02:00
clawbot 4a16a41bd7 Test both hashWorker cancellation checks on their own (closes #83)
check / check (push) Failing after 2s
TestHashWorkerDropsQueuedRuns now passes hashWorker a hash function
that records being called, so a worker that hashes a run after the
scan is cancelled fails the test every time instead of only when it
then chose to send its result.

TestHashWorkerAbandonsBlockedSend cancels the scan from inside the
hash function and leaves the result channel unread, so the worker can
only return through the cancellation case beside its send. The scan
tests could not show this, because stop drains results and frees a
parked worker anyway.

Model: opus-5-5
2026-10-04 16:01:36 +02:00
clawbot 722675f153 Escape the database path in the SQLite connection string (closes #55)
check / check (push) Failing after 2s
openDB put the path into the connection string unescaped, so a ? or #
in it ended the file name and a % started an escape: scan could
silently fill a database under a shortened name. The path now goes
through net/url as a file: URI. An absolute path gets an empty host and
a relative path none, because SQLite reads what follows file:// up to
the next slash as a host name. The path is not cleaned, so it stays
exactly what the operator gave.

A test runs scan, report and trees against such a file name given as an
absolute path, as one starting with //, and as a relative path, and
checks that only that file and its lock file exist afterwards.

Model: opus-5-5
2026-10-04 15:13:19 +02:00
clawbot c9c8b1d06c Reject scan --workers below 1 as a usage error (closes #10)
check / check (push) Failing after 1s
A --workers value of 0 or less used to be quietly raised to 1, so a
typo ran the whole scan on one worker with nothing on stderr to say
why. scan now refuses it before anything is scanned: one line on
stderr and exit 2, like the other usage errors. The clamp in runScan
is gone, and README states the rule and the default.

Model: opus-5-5
2026-10-04 14:30:21 +02:00
clawbot 7278c354f0 Test both walk cancellation checks on their own (closes #81)
check / check (push) Failing after 1s
walkOneDir's check was hiding the worker's: a worker that walked a
queued directory on a cancelled scan still emitted nothing, because
walkOneDir stopped at its first entry. The worker test now queues a
missing directory, whose read fails and sends a warning before
walkOneDir's check is reached. A new test calls walkOneDir directly on
a cancelled scan and checks it returns no subdirectory to descend into.

Model: opus-5-5
2026-10-04 14:01:38 +02:00
clawbot 8032ea682b Test that scan refuses another schema version (closes #64)
check / check (push) Failing after 2s
The version-mismatch test only opened its database through
openReportDatabase, so nothing exercised the branch of initSchema that
stops scan on a database stamped with an unknown schema version. The
test now opens the same database through openScanDatabase too and
requires errSchemaVersion, and is renamed to match the unversioned-file
test beside it, which also covers both paths.

Model: opus-5-5
2026-10-04 13:30:23 +02:00
clawbot 948fb03630 Correct four inaccurate comments in cancel_test.go and rename a constant (closes #33)
check / check (push) Failing after 2s
Fix the comments the re-review of
#6 found misdescribing their
tests; no test's behaviour changes.

State the property the walkClock tests rely on, that the index load's
Done cost does not grow with the record count, instead of a wrong fixed
figure. Describe poolUnwind so it is true of every use, and mark the
tests that have no bound. Record that hashWorker's results-send exit is
reachable through stop but has no test that fails without it. Rename
walkCancelInFlightDirs to walkCancelInFlightFiles, since it counts
files. Note at the top why the file departs from the
one-test-file-per-source-file convention.

Model: opus-4-8 (implementation); opus-5-5 (rework)
2026-10-04 13:01:31 +02:00
clawbot 7ff0dbb6e1 Test the -x filesystem-boundary rejection branch (closes #17)
check / check (push) Failing after 2s
No test ever ran the part of subdirJob that refuses to descend across
a filesystem boundary under -x. The new tests call subdirJob directly
with a made-up device for the operand, so no second real filesystem is
needed. They cover refusal across a boundary, skipping the check when
the operand's device is unknown, the warning when the subdirectory
cannot be statted, and crossing by default without -x. When descent is
accepted, the whole returned job is compared, so losing the operand's
device on the way down fails a test. scan.go is unchanged.

Model: opus-4-8 (implementation); opus-5-5 (rebase)
2026-10-04 12:13:23 +02:00
clawbot bccc14ffef Refuse an unversioned database that already has a files table (closes #11)
check / check (push) Failing after 2s
A database at user_version 0 that already has a files table was made
by something else: scan used to run its CREATE TABLE on it and fail
with a raw SQLite error, and report and trees gave only a bare version
mismatch. All three now refuse such a database with the schema-version
error telling the operator to remove the file and rescan. scan creates
the table and index and sets the version in one transaction, so a first
scan stopped partway leaves an empty database the next scan sets up,
never a files table at version 0. A genuinely empty database is
unchanged.

Model: opus-4-8 (implementation); opus-5-5 (rebase)
2026-10-04 12:01:36 +02:00
clawbot 33f8607e17 Document install, a daily cron scan and reading the reports (closes #54)
check / check (push) Failing after 1s
Getting Started gains three parts: installing with go install, from a
clone or as the Docker image; a crontab line for a daily root scan,
with where its stderr and failures go; and how to read the two reports,
why a row is a candidate rather than proof, and how to check a pair
with cmp before removing anything. The usage block names --help, and
the text below it --workers and -x.

The install line uses @main, not @latest: @latest resolves to the
v0.0.1 tag, which predates the database.

Model: opus-5-5
2026-10-04 11:47:20 +02:00
clawbot 2eeba3df6f Keep the module cache out of the build stage's chown (closes #43)
check / check (push) Failing after 2s
The build stage handed /src and the whole Go module cache to the
unprivileged user with chown -R, in a layer that re-ran on every source
change. On this host that step took from about 80 s to over ten minutes,
depending on load.

The module cache now sits at /go/pkg/mod and stays root's: bootstrap
fills it as root. The sources are copied with --chown. A small chown,
cached with bootstrap, hands builder the /src directory itself, the
module cache's cache/download directory (where make build saves its
lookup of this module's own version), and the telemetry files root's go
commands left in its home. Tests and the build still run as builder.

Model: opus-5-5
2026-10-04 10:01:28 +02:00
clawbot cb5dda4f45 Print --version to stdout (closes #15)
check / check (push) Failing after 3s
Cobra's built-in version flag prints through the writer that carries
help and usage, which is stderr here. The root command now defines
its own -v/--version flag and prints one line, "sfdupes VERSION", to
stdout; a failed write is a fatal error (exit 1). Help and usage stay
on stderr. README documents --version and --help, their streams and
exit codes. Tests cover both flags and the failed write.

Model: opus-5-5
2026-10-04 09:47:28 +02:00
clawbot 4847882e46 Stop scan cleanly on SIGINT or SIGTERM (closes #5)
check / check (push) Failing after 3s
A first SIGINT or SIGTERM cancels the scan. It commits the hashed
records still in its batch, with a context that is not cancelled for
that one write, and starts no other write or deletion; deletions need
a complete walk, so records under paths an interrupted walk never
reached are kept. A batch whose commit failed is now kept for that
final commit instead of dropped. The progress display is finished (a
bar stopped short is no longer filled up), `scan: interrupted after N
files` goes to stderr, and the exit code is 1. A second signal ends the
process at once. A SIGINT inherited as ignored stays ignored.

Model: opus-5-5
2026-10-04 09:30:38 +02:00
clawbot 33cf3dd29a Stream report and trees instead of loading every record (closes #14)
check / check (push) Successful in 1m46s
report now has SQLite group the records and put the rows in report
order, helped by a new files_signature index on (size, head, tail,
content), and writes each row as it reads it. trees reads the records
in path order, where all the paths under a directory come together, so
it computes each directory's digest as soon as the stream leaves it and
keeps only its path, parent, digest and totals. Output is unchanged.

The tests that called the removed in-memory grouping functions now group
records stored in a database. New tests check that both commands give
the same output whatever order the records were inserted in, and that a
stdout failure partway through a long report is reported as one.

Model: opus-5-5
2026-10-04 06:47:25 +02:00
clawbot 9abf81535a Print progress at once off a terminal, keep warnings out of redraws (closes #13)
check / check (push) Successful in 1m56s
When stderr is not a terminal, each phase prints its zero-state line
as it starts instead of after its first item. stderrIsTTY uses
term.IsTerminal from golang.org/x/term, now a direct dependency, so
/dev/null is no longer taken for a terminal.

A spinner keeps the library's background redraw, so its count and
elapsed time stay current while a phase waits for its next item. A
warning printed during a spinner phase goes through the bar
(progressbar.Bprintln), which prints it before its next redraw instead
of racing it. Bars with a total have no background redraw and still
print warnings directly.

Model: opus-5-5
2026-10-04 04:47:35 +02:00
clawbot 2dd1194f33 Warn about and skip non-regular and .zfs operands, keeping their records (closes #9)
check / check (push) Successful in 2m48s
A symlink, socket, FIFO or device-node operand, or a directory operand
named .zfs, was silently ignored yet stayed in the scanned operands, so
the update phase deleted every record stored beneath it. Such an
operand now gets a one-line warning, counts as skipped, and is dropped
before overlapping operands are pruned and the database index is
loaded: another operand beneath it is still scanned, and the records
beneath it count as outside the scanned operands and are not deleted,
unless it lies under another operand. The exit status stays 0. An
operand that turns into one of these after that check is warned about
and skipped by the walk instead. README "scan mode" and "Rules for the
walk" say so.

Model: opus-5-5
2026-10-04 04:01:43 +02:00
clawbot 705c8729ca Hold a lock so a second scan fails at once (closes #53)
check / check (push) Successful in 1m43s
scan takes an exclusive flock(2) on a lock file beside the database
(its path with .lock appended) before it walks anything or opens the
database, and holds it until it returns. A second scan against the
same database fails at once with a one-line error naming the lock
file and exits 1. report and trees never take the lock. The lock ends
with the process, so a fatal error or an interrupt releases it; the
file is never deleted. golang.org/x/sys becomes a direct dependency.

The README smoke test now keeps the database outside the scanned
tree, where its empty lock file would have joined the empty-file
group.

Model: opus-5-5
2026-10-04 02:30:21 +02:00
clawbot 01ff3bb5f0 Test report and trees stdout write failures (closes #30)
check / check (push) Successful in 1m34s
report and trees already checked every stdout write and the final
flush. run now takes the stdout it hands to them, so tests pass a
closed file or a failing writer instead of swapping os.Stdout: a
closed stdout exits 1 with a one-line diagnostic, and the writer's
error reaches the caller.

README "Error handling" now states the two cases that never reach
sfdupes as a failed write: a pipe reader that exits early ends the
process with SIGPIPE, as with cat; and stdout closed with >&- is
replaced by /dev/null by the Go runtime before main runs, so the run
succeeds.

Model: opus-5-5
2026-10-03 18:01:27 +02:00
clawbot d63d3cc7fc Open the database read-only for report and trees (closes #8)
check / check (push) Successful in 1m17s
report and trees now connect read-only (mode=ro, query_only, the same
busy timeout) and no longer set the journal mode, which is a write. A
read-only connection to a WAL database still needs its -wal and -shm
files, or write access to the directory to create them, so scan now
switches the database back to rollback-journal mode whenever it closes
it: between scans the file alone holds the database. If a report has
the database open at that moment the switch is refused; scan warns and
the database stays in WAL mode, with its -wal and -shm files, until the
next scan. README §Database states what readers need.

Model: opus-5-5
2026-10-03 16:30:19 +02:00
clawbot c887f80f57 Escape tab, newline, CR and backslash in report paths (closes #7)
check / check (push) Successful in 1m29s
A path holding a tab or newline split a row of the report or trees
output. The path columns of both now write a backslash, tab, newline
and carriage return as \\, \t, \n and \r; every other byte is written
unchanged. Grouping and sorting still use the stored path. Warnings
on stderr are escaped the same way in warnf, so each stays one line.

In trees, the root directory's node now has the path "/" instead of
an empty string, and its children's paths start with a single slash.

README states the rule under "Report output format".

Model: opus-5-5
2026-10-03 15:30:37 +02:00
clawbot c9bf22d483 Stamp the git tag or short commit in a plain docker build (closes #67)
check / check (push) Successful in 1m1s
.dockerignore now sends .git, without .git/config, which can hold a
credential. The build stage takes the VERSION build argument when one
is given, otherwise git describe --tags --always of that .git, and
fails if the context carries .git and still yields no version. A plain
docker build . used to stamp dev. The CI checkout fetches full history
so CI sees the tag and stamps the same value as make build.

Model: opus-5-5
2026-10-02 08:49:03 +02:00
clawbot c737490a53 Compute the content hash only when head and tail match (closes #61)
check / check (push) Successful in 49s
A file of 10 MiB or more now gets only its 64 KiB head and tail in the
hash phase, so its content is read only when it can be a duplicate. A
new content phase after the update phase finds every group of records,
anywhere in the database, that share size, head and tail and include
one without a content hash. It checks every member with lstat and, when
at least two pass, reads those without a content hash through the
existing worker pool; a stale file does not count as a match. report
and trees leave out records without a content hash. The README, help
text and TODO entry describe the gate; the schema stays at version 1.

Lint suppressed: gosec on the file open in hashContentOnly, as in
hashSignature, and on one chmod in a test.

Model: opus-5-5
2026-09-23 16:06:09 +02:00
clawbot 09a39ddf37 Keep the database schema at version 1 (closes #61)
check / check (push) Successful in 42s
sfdupes is pre-1.0, with no installed base and no databases anywhere,
so the schema is changed in place and its version stays 1.
schemaVersion goes back to 1; the six-column files table, content
included, is the version 1 schema. The check that stops on a database
with any other version stays. README.md and TODO.md no longer describe
a version 2 or rejecting and rescanning version 1 databases. The
main.go package comment still described 1024-byte end windows and said
full file contents are never read; it now describes the hashes the
code computes.

Model: opus-5-5
2026-09-23 13:38:06 +02:00
clawbot 29a65016d0 Add 64 KiB head/tail and content-hash duplicate ladder (closes #61) (#62)
check / check (push) Successful in 57s
2026-09-22 16:40:43 +02:00
clawbot 7ac4f6b723 Remove dead files.dat references from build config (closes #22)
check / check (push) Failing after 0s
files.dat was the scan format before the SQLite database; nothing has
produced it since. Drop the stale references from the Makefile clean
target, .gitignore and .dockerignore. make clean still removes the
binary and .gitignore still covers the database files. The only
remaining mention is the historical entry in TODO.md.

Model: opus-4-8 (implementation); fable-5-1 (merge)
2026-09-21 15:01:57 +02:00
38 changed files with 5786 additions and 1675 deletions
-5
View File
@@ -1,5 +0,0 @@
{
"worktree": {
"bgIsolation": "none"
}
}
+78 -8
View File
@@ -1,9 +1,79 @@
.git
# .dockerignore does NOT use .gitignore semantics. Docker matches with
# moby/patternmatcher: filepath.Match plus `**`, so `*` does not cross
# `/` and an unprefixed pattern is anchored at the context root. Every
# depth-independent pattern therefore needs `**/`, or `config/.env` and
# `certs/server.key` still ship while this file reads as solved. Only
# genuinely root-anchored entries go unprefixed. Never transplant these
# into .gitignore, where `**/` is wrong.
#
# Matching is case-sensitive, so secrets use character ranges rather
# than an ALL-CAPS twin, which would still miss `Server.Key`.
#
# Extend with this repo's own host-built artifacts, written anchored:
# `/myapp`, never `**/myapp`, which also matches `cmd/myapp/` and
# deletes the package directory from the context.
# .git is sent without its config. Without a VERSION build argument the
# stage that compiles runs `git describe --tags --always` on .git, which
# does not need .git/config; that file can hold a credential, such as a
# password in a remote URL or the token the CI checkout step stores there.
# Each submodule keeps a config with the same exposure in its git directory
# under .git/modules/, nested again for a submodule's own submodules, or in
# its own .git directory when it keeps one.
# KNOWN GAP: a submodule whose name has a `config` segment (`config`,
# `deploy/config`, `config/lib`) loses its whole git directory, because
# `**/.git/modules/**/config` also matches that segment's directory
# under .git/modules/. Go's version stamping then fails the build;
# nothing leaks. Name such a submodule without that segment:
# `git submodule add --name`.
**/.git/config
**/.git/modules/**/config
# Agent scratch: one full checkout of the repo per in-flight agent.
# Anchored because it occurs once where agents run at the repo root.
# KNOWN GAP: a repo running agents in subdirectories still ships
# `services/api/.claude/` and must add its own anchored entry.
.claude
.DS_Store
sfdupes
files.dat
node_modules
*.log
*.out
*.test
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Re-include a committed template with a negation if the
# build needs one: `!docs/example.env`.
**/*.[eE][nN][vV]
**/.[eE][nN][vV].*
**/.[eE][nN][vV][rR][cC]
# Private keys and the bundles carrying them. Public certificates
# (*.crt, *.cer) are deliberately absent: they are legitimate inputs.
**/*.[pP][eE][mM]
**/*.[kK][eE][yY]
**/*.[pP]12
**/*.[pP][fF][xX]
**/[iI][dD]_[rR][sS][aA]
**/[iI][dD]_[dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]_[sS][kK]
**/[iI][dD]_[eE][dD]25519
**/[iI][dD]_[eE][dD]25519_[sS][kK]
# Dependencies: restored inside the image, never copied in.
**/node_modules
# OS metadata.
**/.DS_Store
**/Thumbs.db
# Editor state: never a build input, and it churns COPY.
**/*.swp
**/*.swo
**/*~
**/*.bak
**/.idea
**/.vscode
**/*.sublime-*
# This repository's host-built artifacts: the binary `make build` writes,
# and test binaries, coverage output and logs.
/sfdupes
/*.test
/*.out
/*.log
+3
View File
@@ -13,3 +13,6 @@ indent_style = tab
[*.go]
indent_style = tab
# This repository's own sections, such as one for another language it
# uses, go below this comment, and a re-vendor keeps them.
+11
View File
@@ -1,9 +1,20 @@
name: check
on: [push]
# Free the shared runner: a new push cancels only the same branch's older run.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true
jobs:
check:
runs-on: ubuntu-latest
# Free the shared runner from a hung build.
timeout-minutes: 20
steps:
# actions/checkout v4.2.2, 2026-02-22
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
# script/cibuild needs no token, so none is left in .git/config.
with:
persist-credentials: false
# All history and tags, so git describe finds the version tag.
fetch-depth: 0
- run: script/cibuild
+37 -13
View File
@@ -11,26 +11,50 @@ Thumbs.db
.vscode/
*.sublime-*
# Agent scratch (worktrees of this repo, created and destroyed by
# in-flight tooling). Unanchored: .gitignore patterns already match at
# every depth, so no prefix is wanted here. This is not a .dockerignore
# entry and must not be given a `**/` prefix on the way into one.
.claude/
# Node
node_modules/
# Environment / secrets
.env
.env.*
*.pem
*.key
# Secrets. Unanchored like every entry above, so each matches at every
# depth. Matching is case-sensitive on Linux, so names use character
# ranges rather than a lowercase form that misses `Server.Key`.
# Go build artifacts
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Only the templates `example.env` and `sample.env` are
# re-included below. A repository that commits any other template adds
# its own negation at the end of this file, for example `!.env.example`.
*.[eE][nN][vV]
.[eE][nN][vV].*
.[eE][nN][vV][rR][cC]
!example.env
!sample.env
# Private keys and the bundles carrying them.
*.[pP][eE][mM]
*.[kK][eE][yY]
*.[pP]12
*.[pP][fF][xX]
[iI][dD]_[rR][sS][aA]
[iI][dD]_[dD][sS][aA]
[iI][dD]_[eE][cC][dD][sS][aA]
[iI][dD]_[eE][cC][dD][sS][aA]_[sS][kK]
[iI][dD]_[eE][dD]25519
[iI][dD]_[eE][dD]25519_[sS][kK]
# This repository's own entries, such as its build outputs, go below
# this comment, and a re-vendor keeps them. Anchor a binary built at the
# root: `/myapp`, never `myapp`, which also ignores `cmd/myapp/`.
/sfdupes
*.test
*.out
*.log
*.out
*.test
# Local scan data
files.dat
# A scan database lists every path it scanned.
*.sqlite
*.sqlite-shm
*.sqlite-wal
# Agent worktrees
.claude/worktrees/
+69 -2
View File
@@ -10,14 +10,23 @@ run:
linters:
default: all
enable:
# Successor to the deprecated gomodguard. Named explicitly, rather than
# left to `default: all`, because it carries the module policy below.
- gomodguard_v2
disable:
# Genuinely incompatible with project patterns
- exhaustruct # Requires all struct fields
- depguard # Dependency allow/block lists
- exhaustruct_v5 # Requires all struct fields (successor to exhaustruct)
- godot # Requires comments to end with periods
- wsl # Deprecated, replaced by wsl_v5
- wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go
# Deprecated: the warning is attached to the old name, so it is
# silenced by disabling that name, not by enabling the successor.
- wsl # Deprecated, replaced by wsl_v5
- gomodguard # Deprecated, replaced by gomodguard_v2
# Misses findings at random in v2.14.0; back once a pinned release fixes it
- canonicalheader
settings:
lll:
line-length: 88
@@ -28,6 +37,64 @@ linters:
max-complexity: 15
dupl:
threshold: 100
depguard:
# Test-support code must not be compiled into the shipped binary. A
# test-support package exists to hand a test privileges the program
# itself must never have, so a file that is not a test must not import
# one. Test files, and the files inside a package whose directory name
# ends in `test`, are where that code belongs, and are exempt.
#
# The deny list below is the one part of this file a repository is
# expected to extend, and the only part it may. depguard matches an
# import path against a list of prefixes, so it cannot be told "any path
# whose last segment ends in test"; a repository's own test-support
# packages have to be named here one at a time, by full import path,
# under a module path that differs from repository to repository. Add
# them; change nothing else.
rules:
test-support:
list-mode: lax
files:
- "$all"
- "!$test"
- "!**/*test/**"
deny:
- pkg: net/http/httptest
desc: >-
Test-support code belongs in test files and in packages whose
directory name ends in test, not in the shipped binary.
# Only decisions already recorded in the Go package defaults are
# listed here. Every entry matches the module path exactly.
gomodguard_v2:
blocked:
- module: github.com/rs/zerolog
recommendations:
- log/slog
reason: "Structured logging is stdlib log/slog."
# One entry per pre-fork module path, because the later releases
# are separate paths. A prefix match would be shorter but would
# also reach github.com/go-redis/redismock, the test double for
# the successor these entries recommend.
- module: github.com/go-redis/redis
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/go-redis/redis/v7
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/go-redis/redis/v8
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/sergi/go-diff
recommendations:
- github.com/aymanbagabas/go-udiff
reason: "No unified diff output; use go-udiff."
- module: github.com/hexops/gotextdiff
recommendations:
- github.com/aymanbagabas/go-udiff
reason: "Unmaintained fork; use go-udiff."
issues:
max-issues-per-linter: 0
+59 -105
View File
@@ -1,121 +1,75 @@
# Lint stage — fast feedback on formatting and lint issues
# golangci/golangci-lint:v2.12.2, 2026-08-07
FROM golangci/golangci-lint@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 AS lint
# Lint phase, built alone by script/lint. The tools are invoked directly
# rather than through `make lint`, which runs docker itself and so cannot
# run inside a build step.
# golangci/golangci-lint:v2.14.0, 2026-10-07
FROM golangci/golangci-lint@sha256:ad862ba6b3798cbe0fd9fd7408d498fd74fbd2623a92406b2fd3898faf0bf98f AS lint
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Cache-buster for the gate layers, and only for them. Docker
# invalidates COPY only when the copied content changes, so on an
# unchanged tree the gates below would be served from cache and the
# build would exit 0 having run nothing. script/cibuild and
# script/docker pass a fresh CHECK_EPOCH on every invocation.
# The gofmt half of `make fmt-check`. gofmt's output is assigned to a
# variable first so that its own exit status, as when it cannot parse a
# file, still fails the step.
RUN files="$(gofmt -s -l .)" && \
if [ -n "$files" ]; then \
echo "gofmt: files not formatted:" >&2; echo "$files" >&2; exit 1; \
fi
# Validates .golangci.yml against the schema the pinned binary embeds.
RUN golangci-lint config verify --config .golangci.yml
RUN golangci-lint run --config .golangci.yml ./...
# Test phase, built alone by script/test. -race needs cgo and so a C
# compiler, which the Debian Go image ships and the alpine one does not.
#
# Two properties this depends on. ARG is per-stage, so the build stage
# below declares it again; one declaration here would leave that
# stage's gate cacheable. And each gate RUN must reference the value,
# because BuildKit hashes the expanded command: a declared but
# unreferenced ARG invalidates nothing.
#
# It sits below the dependency layers deliberately. Everything above it
# (the pinned base image, go mod download) keeps its cache; only the
# gates go cold.
ARG CHECK_EPOCH
# The tests run as nobody: several of them make a file unreadable and
# expect reading it to fail, and root reads it anyway. nobody has no home
# directory, so HOME is /tmp, where Go puts its build cache.
# golang:1.25-trixie, 2026-10-04
FROM golang@sha256:2c4c60ef415fbfa5e90300722293bef36c5e63fae17570ce18f580af933dbd73 AS test
USER nobody
ENV HOME=/tmp
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN go test -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 90s -race -v ./...; exit 1; }
# The linter is invoked directly here, not through `make lint`. That
# target now runs `docker build -f Dockerfile.lint`, and a docker build
# cannot run a docker build: routing the gate through make would mean
# nesting docker inside this image. Same reason `make check` is gone
# from the build stage below.
#
# `make fmt-check` is not run in this stage: it now also runs prettier
# over Markdown, and this golangci-lint image has no node. The gate runs
# in the build stage below, where script/bootstrap installs node and
# prettier.
# The FROM above and the one in Dockerfile.lint pin the same linter
# twice, and nothing else keeps them in sync; this fails the build when
# they disagree. See the script for why it restates neither pin.
RUN echo "gate lint-image-pin, epoch ${CHECK_EPOCH}" && \
script/verify-lint-image-pin
# Same config-schema check Dockerfile.lint runs, kept here so this build
# gates on exactly what script/lint gates on. It validates against a
# schema the pinned binary embeds, so it needs no network.
RUN echo "gate config verify, epoch ${CHECK_EPOCH}" && \
golangci-lint config verify --config .golangci.yml
RUN echo "gate lint, epoch ${CHECK_EPOCH}" && \
golangci-lint run --config .golangci.yml ./...
# Build stage
# Build stage. Nothing is wanted from either phase above; the copies are
# what make BuildKit build them first, so this stage cannot run unless
# lint and test passed.
# golang:1.25-alpine, 2026-07-23
FROM golang@sha256:56961d79ea8129efddcc0b8643fd8a5416b4e6228cfd477e3fd61deb2672c587 AS builder
# We never build or run as root. Create an unprivileged user and point
# HOME and the Go caches at its home so go build and go test can write
# their caches when we drop to it below. $GOPATH/bin is deliberately not
# on PATH: script/bootstrap no longer `go install`s anything (the linter
# runs from a pinned image, never from a host install), so nothing lands
# there and adding it would only widen what this image resolves.
RUN adduser -D -u 1000 builder
ENV HOME=/home/builder
ENV GOPATH=/home/builder/go
ENV GOCACHE=/home/builder/.cache/go-build
WORKDIR /src
# No-op file copy whose only purpose is the build-graph edge: it is what
# makes this stage depend on the lint stage, and so what forces BuildKit
# to finish fmt-check, the pin guard and lint before compilation and
# tests start. Remove it and the fail-fast design dies silently — the
# build stops gating on lint and still exits 0. It replaces a copy of
# the linter binary itself, which is no longer wanted here: nothing in
# this stage runs the linter, because `make lint` is now a docker build
# and a docker build cannot run inside one.
COPY --from=lint /src/go.sum /dev/null
# Install development prerequisites the same way a developer does,
# rather than duplicating the installs inline. Only script/ and the
# dependency manifests are copied first, nothing else, so this layer
# stays cached until the scripts or the dependencies change — bootstrap
# runs `go mod download` and `yarn install`, which is why there is no
# separate invocation of either here. The JS manifests (package.json,
# yarn.lock) are copied too so the yarn install layer caches alongside
# the Go one.
COPY script/ script/
COPY go.mod go.sum package.json yarn.lock ./
RUN script/bootstrap
COPY --from=test /src/go.sum /dev/null
RUN apk add --no-cache git make
# A tar-stream context keeps the sender's file owners, which git refuses.
RUN git config --system --add safe.directory /src
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Hand the sources and caches to the unprivileged user, then drop root
# before running any checks or builds.
RUN chown -R builder:builder /src /home/builder
USER builder
# The version stamped into the binary: the VERSION build argument when
# one is given, otherwise `git describe --tags --always` of the .git in
# the build context. A context that carries .git and still yields no
# version fails the build; with neither, as from a source tarball, it is
# "dev". `make build` rather than `go build`, so the image and a host
# build share one compile recipe, cgo disabled included.
ARG VERSION
RUN version="${VERSION:-$(git describe --tags --always || echo dev)}"; \
if [ -e .git ] && { [ -z "$version" ] || [ "$version" = dev ] || \
[ "$version" = unknown ]; }; then \
echo "no version could be derived although the build context carries .git" >&2; \
exit 1; \
fi; \
make build VERSION="$version"
# Fail the build unless the branch is green. Runs as non-root so the
# permission-denied test paths are exercised legitimately (root would
# bypass the chmod(0) the tests rely on).
#
# The gates are the individual targets, not `make check`: that aggregate
# runs `script/lint`, which is now a docker build, and nothing inside an
# image build may shell out to docker. Lint is not skipped by this — it
# ran in the lint stage above, which this stage's COPY --from makes a
# prerequisite. `make`, not the scripts directly, because the Makefile's
# `export CGO_ENABLED = 0` applies only to what it invokes.
#
# Second per-stage declaration of the gate cache-buster; see the lint
# stage above for why one is not enough. It is placed after USER so the
# drop to the unprivileged user still happens before the checks run.
ARG CHECK_EPOCH
RUN echo "gate test, epoch ${CHECK_EPOCH}" && make test
RUN echo "gate fmt-check, epoch ${CHECK_EPOCH}" && make fmt-check
RUN make build
# Runtime stage
# Runtime stage, and the last one: a plain `docker build .` builds this
# stage's chain and nothing else.
# alpine:3.22, 2026-07-23
FROM alpine@sha256:14358309a308569c32bdc37e2e0e9694be33a9d99e68afb0f5ff33cc1f695dce
-59
View File
@@ -1,59 +0,0 @@
# Lint-only image: this is how the linter runs, everywhere. The repo is
# COPYed into the pinned golangci-lint image and the linter runs as a
# build step, so a successful build IS a clean lint. golangci-lint is
# never installed on a host — one toolchain, pinned by digest, identical
# on a laptop and in CI — and this works even when the docker daemon is
# remote and bind mounts are impossible.
#
# script/lint builds this file. It is a separate image from the lint
# stage of the main Dockerfile because script/lint must not depend on
# the rest of that build; the two FROM lines are kept identical by
# script/verify-lint-image-pin, run as a gate below.
# golangci/golangci-lint:v2.12.2, 2026-08-07
FROM golangci/golangci-lint@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240
WORKDIR /src
# Dependency layers first, so they stay cached across lint runs.
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Cache-buster for the gate layers, and only for them. Caching of the
# lint run is waived by ruling: COPY is invalidated only by changed
# content, so on an unchanged tree the gates below would be served from
# cache and this build would exit 0 in under a second having run no
# linter at all. That exact false green has bitten this repo twice
# already (#32, #39). script/lint passes a fresh value on every
# invocation.
#
# Each gate RUN must reference the value, because BuildKit hashes the
# expanded command and not the ARG declaration: a declared but
# unreferenced ARG invalidates nothing. The ARG sits below the
# dependency layers deliberately — everything above it keeps its cache,
# only the gates go cold.
ARG CHECK_EPOCH
# The linter version is pinned in two places, here and in the main
# Dockerfile's lint stage. Nothing else keeps them in sync, so a
# half-applied bump is a build failure; see the script.
RUN echo "gate lint-image-pin, epoch ${CHECK_EPOCH}" && \
script/verify-lint-image-pin
# Validates .golangci.yml against golangci-lint's JSON schema. The
# concern about this step was that it fetches that schema over a live,
# unpinned HTTPS call; measured on the pinned image, it does not. The
# binary carries the schema for its own version, so under
# `--network none` this both passes on a valid config and still rejects
# an invalid one with the jsonschema error. That holds for the gate
# steps generally — none of them makes a network call — but not for
# this build as a whole: `go mod download` above needs the network on a
# cold cache, and under `--network none` a first build fails there
# before reaching any gate. That layer stays cached, so only a warm
# cache lints offline, until go.mod or go.sum changes.
RUN echo "gate config verify, epoch ${CHECK_EPOCH}" && \
golangci-lint config verify --config .golangci.yml
RUN echo "gate lint, epoch ${CHECK_EPOCH}" && \
golangci-lint run --config .golangci.yml ./...
+1 -1
View File
@@ -46,4 +46,4 @@ hooks:
@script/install-precommit
clean:
rm -f $(BINARY) files.dat
rm -f $(BINARY)
+491 -177
View File
@@ -4,16 +4,21 @@
`sfdupes` is an MIT-licensed Go CLI tool by [@sneak](https://sneak.berlin) that
quickly identifies _candidate_ duplicate files — and, ultimately, entire
duplicate directory trees — across very large filesystems without reading full
file contents. Files are considered duplicates when they have identical size,
identical SHA-256 of their first 1024 bytes, and identical SHA-256 of their last
1024 bytes. This is a strong candidate signal, not proof of identical content
(the middle of the file is never read); the intended use is finding duplicate
downloads and duplicated directory trees on multi-terabyte ZFS servers where
reading every byte is prohibitively expensive. `scan` maintains a persistent
SQLite database of file signatures that survives between runs, so it can be run
from cron and the reports can be generated at any time from the most recent
scan.
duplicate directory trees — across very large filesystems without reading every
byte of every file. Files are considered duplicates when their sizes are equal
and they agree on a short ladder of hashes. A file under 10 MiB is hashed in
full and compared directly. A larger file is gated first on the SHA-256 of its
first 64 KiB and of its last 64 KiB, and only when its size and both of those
match another file's is it read for a content hash to compare — the SHA-256 of
the whole file when it is under 50 MiB, or of gigabyte-spaced 1 MiB samples when
it is 50 MiB or larger. Below 50 MiB the content hash is proof of identical
content; at or above 50 MiB it is a strong candidate signal rather than proof,
because the gaps between samples are never read. The intended use is finding
duplicate downloads and duplicated directory trees on multi-terabyte ZFS servers
where reading every byte of every file is prohibitively expensive. `scan`
maintains a persistent SQLite database of file signatures that survives between
runs, so it can be run from cron and the reports can be generated at any time
from the most recent scan.
This README is the complete and authoritative specification.
@@ -28,8 +33,9 @@ export SFDUPES_DATABASE="$HOME/.local/share/sfdupes/db.sqlite"
```
`scan` walks one or more filesystem trees and maintains one database record per
regular file (path, size, mtime, head hash, tail hash). The database persists
between runs; a rescan only hashes files that are new or changed, and removes
regular file (path, size, mtime, head hash, tail hash, content hash). The
database persists between runs; a rescan only hashes files that are new or
changed, or that may have gained a duplicate since the last scan, and removes
records for files that no longer exist. `report` reads the database and prints
the file-level duplicates report. `trees` reads the same database and prints the
duplicate-tree report. A missing/invalid subcommand — or a `scan` invocation
@@ -40,18 +46,134 @@ by setting `SFDUPES_DATABASE`. The intended deployment is a daily `sfdupes scan`
cron job, with the reporting commands run interactively whenever needed; their
results are as fresh as the last completed scan.
### Install
With Go installed, this builds and installs the current `main` branch:
```sh
go install sneak.berlin/go/sfdupes@main
```
The binary goes to `$(go env GOPATH)/bin`, or to `$GOBIN` when that is set. A
binary installed this way reports its version as `dev`; one built from a clone
or into the Docker image carries the git tag or commit it was built from.
From a clone, `make build` writes the binary to `./sfdupes`:
```sh
git clone https://git.eeqj.de/sneak/sfdupes.git
cd sfdupes
make build
```
Copy the binary to `/usr/local/bin` for the cron job below.
`make docker` builds the Docker image, tagged `sfdupes`, after running the tests
and the linter (see "Build"). The image runs `sfdupes` as root with the database
at its default path, so a bind mount of `/var/lib/sfdupes` keeps the database
between runs. Mount the scanned tree at the same path inside the container as on
the host; read-only is enough. The database records paths as the container sees
them, so the reports then name the host's paths.
```sh
make docker
docker run --rm -v /srv:/srv:ro -v /var/lib/sfdupes:/var/lib/sfdupes \
sfdupes scan /srv
docker run --rm -v /var/lib/sfdupes:/var/lib/sfdupes sfdupes report > dupes.tsv
```
### Daily scan from cron
Run `scan` as root, so that it can read every file: a path it cannot read is
skipped with a warning and loses its database record (see "Rules for the walk").
As a file `/etc/cron.d/sfdupes`:
```
30 3 * * * root /usr/local/bin/sfdupes scan /srv 2>>/var/log/sfdupes.log || tail -n 3 /var/log/sfdupes.log
```
- The database is `/var/lib/sfdupes/db.sqlite`, created with its directory by
the first scan. To keep it elsewhere, set
`SFDUPES_DATABASE=/path/to/db.sqlite` before the command on the same line.
- `scan` writes nothing to stdout. Its stderr, appended here to
`/var/log/sfdupes.log`, holds a plain progress line as each phase starts and
then at most every 5 seconds, a warning for each path it skips, and the
summary line (see "Progress" and "`scan` mode"). The log grows with every
scan; rotate it like any other.
- Skipped paths do not fail a scan: it still exits 0, and cron sends nothing. A
scan that fails, or is stopped by `SIGINT` or `SIGTERM`, exits 1 with the
reason among the last lines of the log; `tail` prints them, and cron mails
them to root if the host can send mail.
- A scan still running when the next one starts carries on. The new one fails at
once, and the lines cron mails include
`sfdupes: another scan is running (lock held on /var/lib/sfdupes/db.sqlite.lock)`.
- `report` and `trees` need only read access to the database (see "Database").
Under the usual umask of `022` the first scan creates it readable by every
user, so an unprivileged user can run them against root's database.
### Reading the reports
Each row of `report` names two copies of one file, and each row of `trees` two
copies of one directory tree (see "Report output format" and "Trees output
format"). In a group of copies, the path that sorts first byte by byte is
`first` and every other path is a `dupe` of it. `first` says nothing about which
copy is the original or the oldest; which copy to keep is your choice.
A row is a candidate, not proof:
- The reports read only the database, so they show the files as of the last
scan; a file may have changed or gone since.
- A file of 50 MiB or more is compared only on samples of its content (see
"Duplicate detection").
- Paths that are hard links to one file are listed as duplicates, but they share
their data, so removing one frees nothing.
Compare a pair byte for byte before removing either copy. For the row
`/srv/a/big.iso`, `/srv/b/big-copy.iso`, `4294967296`:
```sh
cmp /srv/a/big.iso /srv/b/big-copy.iso && echo identical
[ /srv/a/big.iso -ef /srv/b/big-copy.iso ] && echo "hard links"
```
`cmp` prints nothing and exits 0 only when every byte matches, and otherwise
reports where the files differ. The second line prints `hard links` when the two
paths are the same file, so removing either frees nothing.
A path holding a backslash, tab, newline or carriage return is escaped in the
reports (see "Report output format"). Undo the escapes before using it.
`printf '%b'` does exactly that, because every backslash in an escaped path
starts one of the four escapes. Command substitution drops trailing newlines, so
print an `x` after the path and remove it afterwards, or a path that ends in a
newline names a different file:
```sh
p="$(printf '%bx' '/srv/a/tab\tname.txt')"; p="${p%x}"
cmp "$p" /srv/b/tab-copy.txt
```
Check a `trees` row with `diff -r`, which compares the two trees file by file
and also names anything present in only one of them, such as an empty directory
or a symlink, which `trees` does not see.
## Rationale
Duplicate finders that hash entire files do not scale to the target environment:
~10 million files and ~150 TB on possibly slow or busy disks (a ZFS pool under
resilver). Reading at most 2 KiB per file — and only from files whose size at
least one other file shares, since a size-unique file cannot be a duplicate —
makes a full-filesystem sweep tractable, and the signatures are kept in a
persistent database, so the expensive filesystem pass is incremental: a rescan
re-hashes only files whose recorded mtime or size changed, and all analysis
happens offline from the database alone. The end goal is not individual files
but whole duplicated trees — duplicate extractions, duplicate downloads, copied
project trees — which an operator can consider removing as a unit.
resilver). sfdupes spends disk I/O only on files whose size at least one other
file shares, since a size-unique file cannot be a duplicate. Of those, a file
under 10 MiB is read in full; a larger one has its cheap end windows read first,
and is read for a content hash only when its size and both end windows match
another file's — the whole file below 50 MiB, but only gigabyte-spaced samples
at or above 50 MiB, so the largest files are never read in full. This keeps a
full-filesystem sweep tractable, and the signatures are kept in a persistent
database, so the expensive filesystem pass is incremental: a rescan re-hashes
only files whose recorded mtime or size changed, plus — for its content hash — a
file of 10 MiB or more whose size and end windows have come to match another
file's. All analysis happens offline from the database alone. The end goal is
not individual files but whole duplicated trees — duplicate extractions,
duplicate downloads, copied project trees — which an operator can consider
removing as a unit.
## Design
@@ -63,27 +185,43 @@ Goals, in order:
trees), so the operator can consider removing an entire subtree at once.
File-level duplicate detection is the foundation; tree-level detection is
built on top of it.
2. **Never read full file contents.** At most 2 KiB is read per file (first and
last 1024 bytes), and only files whose size at least one other file shares
are read at all — a size-unique file cannot be a duplicate. Scale target:
tens of millions of files, ~150 TB filesystem, possibly slow or busy disks
(ZFS pool under resilver). Holding one small record (path, size, mtime) per
file in memory during a scan is acceptable; holding every file's hashes is
not (they stay in the database).
2. **Spend I/O in proportion to duplicate likelihood.** Only files whose size
at least one other file shares are read at all — a size-unique file cannot
be a duplicate. Those are compared by the ladder in "Duplicate detection"
below: a file under 10 MiB is hashed in full, while a larger file is gated
on cheap 64 KiB end windows first, and gets a content hash only when its
size and both end windows match another file's. That hash reads the whole
file below 50 MiB but only gigabyte-spaced 1 MiB samples at or above it, so
the very largest files are still never read in full. Scale target: tens of
millions of files, ~150 TB filesystem, possibly slow or busy disks (ZFS pool
under resilver). Holding one small record (path, size, mtime) per file in
memory during a scan is acceptable; holding every file's hashes is not (they
stay in the database). The reporting commands do not hold every file's
hashes either: `report` lets SQLite group and order the records and writes
each row as it reads it, so its memory does not grow with the database, and
`trees` reads the records in path order and keeps each directory's path,
digest and totals, plus the hashes of only the files in the directories
holding the record being read, so its memory grows with the number of
directories and with the size of the largest directory.
3. **Scan incrementally, analyze offline.** The expensive filesystem scan
maintains a persistent database; an unchanged file is never read again on a
rescan. All analysis (`report`, `trees`) works from the database alone and
must never touch the scanned filesystem again. `scan` is designed to be
cronned; the reports run at any time against the last completed scan.
rescan, except to compute its content hash once a file of 10 MiB or more
comes to match another on size and both end windows. All analysis (`report`,
`trees`) works from the database alone and must never touch the scanned
filesystem again. `scan` is designed to be cronned; the reports run at any
time against the last completed scan.
4. **Clean stream separation.** Everything on stdout is machine-readable data.
All progress, warnings, and summaries go to stderr. Never mix them.
All progress, warnings, summaries, and help and usage text go to stderr.
Never mix them.
### Constraints
- Language: Go (module `sneak.berlin/go/sfdupes`). Binary name: `sfdupes`.
- Dependencies: standard library, `github.com/spf13/cobra` for the CLI, **one
progress-bar library** (`github.com/schollz/progressbar/v3`), and **one SQLite
driver** (`modernc.org/sqlite`, pure Go, so builds keep cgo disabled).
progress-bar library** (`github.com/schollz/progressbar/v3`),
`golang.org/x/term` to tell whether stderr is a terminal, **one SQLite
driver** (`modernc.org/sqlite`, pure Go, so builds keep cgo disabled), and
`golang.org/x/sys` for `flock(2)` (the scan lock, see "Database").
`github.com/spf13/viper` is permitted if configuration-file support is ever
needed, but is not currently used. No other third-party deps.
- Cross-compilation is not a concern. Builds run with cgo disabled (the
@@ -107,45 +245,141 @@ Three subcommands, all implemented:
sfdupes scan [--workers N] [-x] PATH...
sfdupes report > dupes.tsv
sfdupes trees > dupetrees.tsv
sfdupes --version
sfdupes [command] --help
```
`--workers N` sets the size of each `scan` worker pool (default: the number of
CPUs), and `-x` (`--one-file-system`) keeps the walk of each operand on that
operand's filesystem; both are described under "`scan` mode".
`sfdupes --version` (or `-v`) prints one line, `sfdupes VERSION`, to stdout and
exits 0, writing nothing to stderr. `-h` or `--help`, alone or after a
subcommand, prints the help text to stderr and exits 0, writing nothing to
stdout.
### Database
All three subcommands operate on a single SQLite database file:
- Location: the value of the `SFDUPES_DATABASE` environment variable when set
and non-empty, otherwise `/var/lib/sfdupes/db.sqlite`. There is no
command-line flag.
command-line flag. The path names the file exactly, whatever characters it
holds (`?`, `#` and `%` included); a relative path is relative to the working
directory.
- `scan` creates the database (and its parent directory) on first use. `report`
and `trees` require an existing database; a missing database file is a fatal
error (exit 1) telling the user to run `scan` first.
- The database uses WAL journal mode and a busy timeout, so running a report
while a cron `scan` is in progress is safe. The filesystem is authoritative;
the database is an eventually-consistent reflection of it. Hashed records are
committed in batched transactions while the scan is still running (keeping the
WAL small and letting concurrent reports observe progress), so a report may
see a scan's changes partially applied, and a scan that dies partway leaves a
valid database holding everything hashed so far; the next scan skips those
records and converges toward the filesystem.
- Only one `scan` runs against a database at a time. For its whole run, `scan`
holds an exclusive `flock(2)` lock on a lock file beside the database, named
by appending `.lock` to the database path (`/var/lib/sfdupes/db.sqlite.lock`
by default), taken before it walks the filesystem or opens the database. A
second `scan` against the same database does not wait: it fails at once with a
one-line error naming the lock file and exits 1, without walking anything or
opening the database, and the running scan carries on. The lock file is
created on first use, open to its owner only, and left in place: a leftover
file blocks nothing, because the lock ends with the process holding it however
it ends, a fatal error or an interrupt included, and deleting the file while a
scan runs would let a second scan start. `report` and `trees` never take the
lock, so they run during a scan.
- While `scan` runs, the database is in WAL journal mode with a busy timeout, so
running a report while a cron `scan` is in progress is safe. The filesystem is
authoritative; the database is an eventually-consistent reflection of it.
Hashed records are committed in batched transactions while the scan is still
running (keeping the WAL small and letting concurrent reports observe
progress), so a report may see a scan's changes partially applied, and a scan
that dies partway leaves a valid database holding every batch committed so far
(an interrupted scan also commits the batch in progress, see "Error handling
and exit codes"); the next scan skips those records and converges toward the
filesystem.
- `scan` switches the database back to rollback-journal mode when it closes it,
so between scans the database file alone holds the whole database. Each switch
needs the database to itself: a `scan` that starts while a report is still
reading waits for it up to the 10-second busy timeout, then fails; a `scan`
that ends while a report has the database open warns and leaves the database
in WAL mode until the next scan. `report` writes each row as it reads it, so
it is still reading while its output is paused (a pager, a stalled pipe), and
a `scan` started then fails after the busy timeout.
- `report` and `trees` open the database read-only and need only read access to
the database file, and no write access to its directory. While the database is
in WAL mode they also read the `-wal` and `-shm` files beside it, which SQLite
creates with the database file's permissions.
- Schema (`PRAGMA user_version` is the schema version, currently 1; a database
with any other version is a fatal error):
with any other version is a fatal error. `scan` creates the schema and sets
the version in one transaction, so a first scan stopped while doing so leaves
an empty database the next scan sets up. A database at version 0 that already
has a `files` table was therefore not made by sfdupes; every subcommand
refuses it with an error telling the user to remove the file and rescan):
```sql
CREATE TABLE files (
path BLOB PRIMARY KEY, -- absolute path, raw bytes
size INTEGER NOT NULL, -- bytes, from lstat
mtime INTEGER NOT NULL, -- Unix seconds, from lstat
head TEXT NOT NULL, -- lowercase-hex SHA-256, first 1 KiB
tail TEXT NOT NULL -- lowercase-hex SHA-256, last 1 KiB
mtime INTEGER NOT NULL, -- whole Unix seconds of the mtime, from lstat
mtime_nsec INTEGER NOT NULL, -- nanoseconds within that second, 0 to 999999999
head TEXT NOT NULL, -- lowercase-hex SHA-256; first 64 KiB, or whole file under 10 MiB
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
) WITHOUT ROWID;
CREATE INDEX files_signature ON files (size, head, tail, content);
```
Paths are stored as BLOBs because Unix paths are raw bytes, not guaranteed
UTF-8. `mtime` is used only for change detection; it is not part of the
duplicate key. `head` and `tail` are empty strings when the file has never
been hashed because its size was unique as of the last scan that covered it;
such records still define the file for tree reconstruction but never
participate in duplicate groups.
UTF-8. `mtime` and `mtime_nsec` hold the file's mtime to the nanosecond:
`mtime` the whole Unix seconds, rounded down, and `mtime_nsec` the
nanoseconds past that second. Split this way they hold any mtime a
filesystem can record, one before 1678 or after 2262 included, which a
single 64-bit count of nanoseconds cannot. They are used only for change
detection and are not part of the duplicate key. For a file under 10 MiB
`head`, `tail`, and `content` all hold the whole-file hash (that range is
hashed in full, with no end windows); for a larger file `head` and `tail`
hold the first- and last-64 KiB hashes and `content` the whole-file or
sampled hash. All three are empty strings when the file has never been
hashed because its size was unique as of the last scan that covered it. For
a file of 10 MiB or more, `content` stays empty until the content phase of a
scan (see "`scan` mode" below) has read the file. A record with an empty
`content` is never part of a duplicate group, though it still defines the
file for tree reconstruction. The `files_signature` index lets SQLite group
the records by signature for `report` without sorting the whole table.
### Duplicate detection
Two files are duplicates only when they agree on every rung of this ladder; a
mismatch at any rung means they are not duplicates. `scan` stores each file's
hashes, and `report` and `trees` group files by the whole signature — size,
`head`, `tail`, and `content` — so the grouping is exactly this ladder applied
across everything scanned into the database, even across separate scans.
1. **Size.** Files of different sizes are never compared. Only files whose size
at least one other file shares are hashed at all.
2. **Under 10 MiB: whole file.** A file smaller than 10 MiB is hashed in full
and compared directly, with no separate end-window step — small files are
cheap to read to the last byte, and doing so makes the comparison exact.
`head`, `tail`, and `content` all hold this whole-file SHA-256, so such a
file's signature is decided entirely by its size and its content.
3. **10 MiB and above: head and tail.** For a larger file, the SHA-256 of the
first 64 KiB (`head`) and of the last 64 KiB (`tail`) are a cheap gate that
eliminates most same-size pairs before any bulk reading: the content hash of
the next two rungs is computed only for a file whose size, `head`, and
`tail` match another file's, whether that file is scanned in the same run or
stored by an earlier scan. A stored file that first gains such a match in a
later scan gets its content hash then; until it has one, its `content` is
empty and it is not a duplicate. At 10 MiB and above the two windows never
overlap.
4. **10 MiB and above, content below 50 MiB.** The SHA-256 of the entire file.
Agreement here is proof of identical content (barring a SHA-256 collision).
5. **10 MiB and above, content 50 MiB and above.** A sampled SHA-256: the 1 MiB
window at each gigabyte-aligned offset (0, 1 GiB, 2 GiB, … while inside the
file, the final window truncated at end of file) is fed, in order, into one
hash. This is **deliberately probabilistic** — the gaps between samples are
never read, so two large files that agree on every sample are reported as
duplicates without being read in full. It is the price of never reading a
150 GB file end to end. Because size is already part of the signature, only
equal-size files reach this rung, so their sample boundaries always align.
`head`, `tail`, and `content` are one column each. A file below 10 MiB and one
at or above it never share a size, and neither do a file below 50 MiB and one at
or above it, so a stored value is never ambiguous between the whole-file,
end-window, and sampled forms.
### `scan` mode
@@ -161,14 +395,25 @@ pool. Overlapping operands are harmless — an operand that duplicates another o
lies under another is dropped before walking, so every file is reached exactly
once and produces one database record.
An operand that is a symlink (never followed, not even as an operand), socket,
FIFO, or device node, or a directory named `.zfs`, is not scanned. `scan` prints
a one-line warning naming the path and what it is, counts it as skipped, and
drops it from the scanned operands before reading the database. Another operand
beneath it is still scanned. The records stored beneath it are not deleted: they
are treated like any other record outside the scanned operands, including the
content-phase exception below. If it lies under another operand, they are under
that operand instead, and are deleted like any other record there that this scan
did not verify. This is not an error: a scan whose every operand is dropped
walks nothing and exits 0.
`scan` synchronizes the database with the filesystem state under the scanned
operands:
- Only a file whose size at least one other file shares is ever read: a
size-unique file cannot be a duplicate, so it is recorded without hashes
(`head` and `tail` empty). The size census covers every file walked this scan
plus every database record outside the scanned operands, so a possible
duplicate of a separately scanned tree is still recognized.
(`head`, `tail`, and `content` empty). The size census covers every file
walked this scan plus every database record outside the scanned operands, so a
possible duplicate of a separately scanned tree is still recognized.
- A file not yet in the database is inserted: hashed when its size is shared,
without hashes otherwise.
- A file already in the database is **skipped without reading its contents**
@@ -176,19 +421,29 @@ operands:
than the recorded mtime. This is what makes a daily rescan cheap. Exception:
an unchanged file whose record lacks hashes is hashed — and its record updated
— once its size becomes shared, so hashing deferred by size-uniqueness happens
as soon as it could matter.
as soon as it could matter. Likewise, an unchanged file of 10 MiB or more
whose record has no `content` hash is read for one by the content phase below
once its size, `head`, and `tail` match another record's.
- A file whose mtime is newer than recorded, or whose size differs, is processed
as if new: re-hashed, or recorded without hashes, per the shared-size rule.
Change detection compares the mtime to the nanosecond, as finely as the
filesystem records it, so a same-size rewrite counts as a change whenever the
filesystem gives it a later mtime than recorded, even within the same second.
- A database record whose path lies under one of the scanned operands but was
not successfully processed this run is deleted. This removes records for
deleted files. It also removes records for paths that failed to stat or hash
this run: the database only ever contains signatures verified by the most
recent scan that covered them (a subsequent successful scan re-adds such
files).
files). A failure in the content phase below removes nothing: the record is
left as it is.
- Database records outside the scanned operands are untouched, so disjoint trees
can be scanned on different schedules into the same database.
can be scanned on different schedules into the same database. The one
exception is the content phase below: a stored file of 10 MiB or more without
a `content` hash is read for one, wherever it lies, once its size, `head`, and
`tail` match another record's. If that file is gone or has changed since its
record was written, the record is left as it is.
`scan` runs **three sequential phases over the whole scan**. Parallelism lives
`scan` runs **four sequential phases over the whole scan**. Parallelism lives
inside each phase; batched database writes begin during the hash phase:
1. **walk + stat** — enumerate the trees under all `PATH` operands concurrently
@@ -204,10 +459,11 @@ inside each phase; batched database writes begin during the hash phase:
2. **hash** — with the census complete, each carried file's size decides its
fate. Size-unique files are never read: new or changed ones are recorded
without hashes in the update phase, unchanged unhashed ones simply keep
their records. Every file with a shared size is hashed by the worker pool:
read the first `min(1024, size)` bytes and the last `min(1024, size)` bytes
(one read when `size <= 1024`, since the two windows coincide) and compute
the SHA-256 of each. Zero-length files have constant hashes and are never
their records. Every file with a shared size is hashed by the worker pool as
described in "Duplicate detection" above: a file under 10 MiB in full, which
gives its `head`, `tail`, and `content` alike, and a larger file only in its
end windows, which give its `head` and `tail`; its content hash is left to
the content phase. Zero-length files have constant hashes and are never
opened. Files are hashed in **inode order** (minimizing seeks on spinning
disks), and paths that are hard links to the same inode are **read once**,
all sharing the one result — a hard-link backup farm costs one read per
@@ -219,13 +475,36 @@ inside each phase; batched database writes begin during the hash phase:
3. **update** — commit the final partial batch, the hash-less records for
size-unique new and changed files, and the deletions for records the scan
did not verify (vanished files, plus paths that failed to stat or hash).
4. **content** — find every record of 10 MiB or more without a `content` hash
whose size, `head`, and `tail` equal another record's, anywhere in the
database: records from this scan and records stored by earlier scans, inside
or outside the scanned operands. SQLite finds them, so only the records to
be read are kept in memory, never every file's hashes. Every record sharing
their size, `head`, and `tail`, including one that already has a `content`
hash, has its file checked with `lstat` first. A file that is gone, is no
longer a regular file, or has changed (a different size, or an mtime newer
than recorded) keeps its record as it is and does not count as a match for
the others. Any other `lstat` error is warned about and counted as skipped,
with the same result. If such a record has no `content` hash, it stays out
of duplicate groups; if it has one, it is still reported until a scan
covering its own tree updates or removes it. The files that pass and have no
`content` hash are read only if at least two of those records pass, so a
file whose only matches are stale costs no read; a file that already has a
`content` hash is never read again. They are read by a worker pool as in the
hash phase, in inode order and once per inode, and their content hashes are
committed in batches. A failed read is warned about and counted as skipped;
its record keeps an empty `content`, so it is not a duplicate, and a later
scan tries again.
Rules for the walk:
- Only regular files. Skip directories, symlinks (do not follow, including
symlink operands), sockets, FIFOs, and device nodes.
symlink operands), sockets, FIFOs, and device nodes. An operand that is a
symlink, socket, FIFO, or device node is dropped as described in "`scan` mode"
above.
- Never descend into a directory named `.zfs` (ZFS snapshot pseudo-dirs; walking
them would list every file once per snapshot).
them would list every file once per snapshot), not even when it is an operand;
such an operand is dropped the same way.
- Filesystem boundaries are crossed by default. With `-x` (long form
`--one-file-system`, following the GNU `du`/`rsync` convention), never descend
into a directory on a different filesystem than its `PATH` operand; each
@@ -234,17 +513,20 @@ Rules for the walk:
unreadable): print a one-line warning to stderr, skip the path, and continue.
Per-file errors never abort the run; the final summary reports how many were
skipped. As specified above, a skipped path that has a database record from an
earlier scan loses that record; an unreadable directory subtree likewise loses
its records (accepted: the database mirrors what the latest scan could
actually verify).
earlier scan loses that record, unless it failed only in the content phase, or
is an operand dropped before the database was read that lies under no other
operand; an unreadable directory subtree likewise loses its records (accepted:
the database mirrors what the latest scan could actually verify).
Concurrency: the walk phase (which also stats files) and the hash phase each use
a worker pool of `--workers` workers (default `runtime.NumCPU()`); the walk
parallelizes across directories, hashing across files. Both phases are
seek-bound on spinning disks, so raising `--workers` well past the core count
can help on pools with many spindles. The main goroutine owns partitioning,
database writes, and progress rendering; progress display must never block the
workers.
Concurrency: the walk phase (which also stats files), the hash phase, and the
content phase each use a worker pool of `--workers` workers (default
`runtime.NumCPU()`); the walk parallelizes across directories, hashing across
files. `--workers` must be at least 1: a smaller value is a usage error,
reported in one line on stderr with exit 2 before anything is scanned. All three
phases are seek-bound on spinning disks, so raising `--workers` well past the
core count can help on pools with many spindles. The main goroutine owns
partitioning, database writes, and progress rendering; progress display must
never block the workers.
`scan` writes nothing to stdout. The summary line on stderr reports the files
seen this run broken down by disposition, plus skips:
@@ -262,14 +544,17 @@ total.)
**`report` must never touch the filesystem being analyzed.** It does not stat,
open, or otherwise access any path that appears in the records; its only I/O is
reading the database and writing stdout/stderr. It must produce identical output
reading the database, writing stdout/stderr, and the temporary file SQLite sorts
in when the duplicate rows do not fit in memory. SQLite puts that file in
`$SQLITE_TMPDIR` or `$TMPDIR` when set, otherwise in `/var/tmp` (or `/tmp`), and
deletes it as soon as it has opened it. `report` must produce identical output
whether or not the scanned filesystem is still mounted.
Processing:
- Records without hashes (size-unique when last scanned) are excluded: their
- Records without a `content` hash (see "Database" above) are excluded: their
content is unknown, so they are never reported as duplicates.
- Group the remaining records by the key `(size, head_hash, tail_hash)`.
- Group the remaining records by the key `(size, head, tail, content)`.
- Every group with two or more paths is a duplicate group.
- Within each group, sort paths lexicographically (byte order). The first path
is the group's `first`; every other path is a `dupe`.
@@ -288,6 +573,14 @@ first dupe size
/srv/a/big.iso /srv/c/big-copy2.iso 4294967296
```
Paths are raw bytes and may hold any byte except NUL, so the path columns
(`first` and `dupe`) are escaped to keep every row one line of tab-separated
fields: a backslash is written as `\\`, a tab as `\t`, a newline as `\n`, and a
carriage return as `\r`. Every other byte is written unchanged, including bytes
that are not valid UTF-8. Undoing those four escapes gives back the stored path.
Grouping and ordering use the stored path, not the escaped one. The warnings
`scan` prints on stderr are escaped the same way, so each warning is one line.
Summary to stderr: records read, number of duplicate groups, number of dupe
files, and total reclaimable bytes (sum of `size` over all dupe rows) in human
units.
@@ -304,10 +597,10 @@ records, split on `/`.
Definitions:
- A file's **signature** is `(size, head_hash, tail_hash)` — mtime is
informational and excluded. An unhashed record (empty hashes) has unknown
- A file's **signature** is `(size, head, tail, content)` — mtime is
informational and excluded. A record without a `content` hash has unknown
content: its signature is treated as unique to that file, so a tree containing
an unhashed file never compares equal to any other tree.
such a file never compares equal to any other tree.
- A directory's **digest** is a SHA-256 Merkle digest computed bottom-up:
serialize the directory's child entries — for a file child, its name and
signature; for a subdirectory child, its name and that subdirectory's digest —
@@ -353,6 +646,9 @@ first dupe files size
/srv/a/project /srv/backup/project 3417 104857600
```
The `first` and `dupe` paths are escaped as described under "Report output
format". The root directory's path is `/`.
Summary to stderr: records read, number of duplicate-tree groups, number of dupe
trees, and total reclaimable bytes (sum of `size` over all dupe rows) in human
units.
@@ -365,9 +661,12 @@ Use the progress-bar library for all scan progress; rendering in the style of
Each phase gets its own display, rendered the moment the phase starts — a scan
must never look hung. Loading the existing-record index (`load`) and the walk
have no known totals while running: show a live count, rate, and elapsed time
(spinner-style, no percentage or ETA). The hash and update phases have exact
totals — only files that actually need hashing appear in the hash total, so its
ETA is meaningful. Required elements for the bars with known totals:
(spinner-style, no percentage or ETA). The content phase's display (`content`)
starts the same way, counting the records checked while SQLite finds the files
to read and `lstat` checks them, then shows a bar once reading starts. The hash
and update phases, and the content phase's reads, have exact totals — only files
that actually need hashing appear in the hash and content totals, so their ETAs
are meaningful. Required elements for the bars with known totals:
- elapsed time
- estimated time remaining
@@ -382,21 +681,57 @@ hash: [12345/98765] 12% |████ | 92 files/s elapsed 2:32 eta 17:54
Additional requirements:
- When stderr is not a TTY, do not emit ANSI redraws: print a plain one-line
progress update no more often than every 5 seconds instead.
- When stderr is not a terminal (a pipe, a file, `/dev/null`), do not emit ANSI
redraws: print a plain one-line progress update the moment each phase starts,
then no more often than every 5 seconds.
- Progress updates are driven from the main goroutine and must be non-blocking
with respect to the worker pool.
with respect to the worker pool. On a terminal the spinner-style displays also
redraw on their own several times a second, so their count and elapsed time
stay current while a phase waits for its next item.
- A warning printed during a phase always lands on a line of its own, never
inside the progress display.
- A bar whose phase stops short of its total, as an interrupted one does, is
left as last drawn rather than filled up.
- `report` and `trees` modes need no progress display, only their stderr
summaries.
### Error handling and exit codes
- `0`: success, even if individual files were skipped with warnings.
- `1`: fatal error (e.g., a `PATH` operand does not exist, the database cannot
be created/opened/read/written, a missing database for `report`/`trees`,
stdout write failure).
- `2`: usage error (including `scan` with no `PATH` operand and `report`/`trees`
with any positional argument).
- `1`: fatal error (e.g., a `PATH` operand does not exist, another `scan` is
already running against the same database, the database cannot be
created/opened/read/written, a missing database for `report`/`trees`, stdout
write failure), or a `scan` stopped by `SIGINT` or `SIGTERM` (see below).
- `2`: usage error (including `scan` with no `PATH` operand, `scan` with
`--workers` below 1, and `report`/`trees` with any positional argument).
A stdout write failure, such as a full disk, is reported in one line on stderr
and exits 1. Two cases never reach sfdupes as a failed write:
- When the reader of a stdout pipe exits early, as in `sfdupes report | head`,
the next write ends sfdupes with `SIGPIPE`, quietly and without a summary, the
way `cat` or `sort` end. The shell reports the signal (status 141 in most
shells), not exit 1.
- When stdout is closed outright (`sfdupes report >&-`), the Go runtime opens
`/dev/null` in its place before sfdupes starts, so the output is discarded and
the run succeeds, as with `> /dev/null`.
`scan` stops cleanly on `SIGINT` (Ctrl-C) or `SIGTERM`. Its workers stop taking
work, each finishing at most the directory listing or file it is reading; the
progress display is finished; and the records it has hashed but not yet
committed are committed, so the next scan does not hash them again. Apart from
that commit it starts no further writes or deletions: records are deleted only
after a complete walk, so those under paths an interrupted walk never reached
are kept. The database is closed and the lock released as on any other exit, the
line `scan: interrupted after N files` goes to stderr, N being the number of
files the walk reached, and the exit code is 1. The next scan skips the records
already written and converges as usual.
After the first signal `scan` stops catching them, so a second one ends it at
once, as an uncaught signal does: the records not yet committed are lost, and
the database is left valid, as when any scan dies (see "Database"). A `SIGINT`
that `scan` inherits as ignored, as a script's background job does, stays
ignored.
## Entrypoints
@@ -409,89 +744,64 @@ from any working directory, and may be invoked directly. The provided
entrypoints are:
- `script/bootstrap` — install everything needed to build and develop this
repository, idempotently, assuming nothing is present. `git`, `make`, `go`,
and `node` come from the first of nix, apt, brew, or apk found on the host,
and are presence-checked only; `node` is an unpinned host runtime like the
rest, because nvm's prebuilt node is glibc-linked and does not run on this
repo's musl/Alpine build image. The Markdown formatter itself — `prettier` —
is pinned by `yarn.lock`'s integrity hash and installed with
`yarn install --frozen-lockfile`. `golangci-lint` is deliberately **not**
installed: it runs from a digest-pinned image via `script/lint` and never from
a host install, so there is no host copy to drift from the pin. A missing
`docker` is warned about rather than installed or treated as fatal —
everything except linting works without it. Ends with `go mod download` and
the `yarn` install.
repository, idempotently, assuming nothing is present. `git`, `make`, and `go`
come from the first of nix, apt, brew, or apk found on the host, and are
presence-checked only. An installed node is used as it is; otherwise node
22.17.0 is installed through nvm, which comes from a release archive whose
sha256 the script checks. A yarn already on `PATH` is used as it is; otherwise
yarn 1.22.22 is activated through corepack, or installed with `npm` when there
is no corepack. `yarn install --frozen-lockfile` then installs the prettier
that `package.json` and `yarn.lock` pin. `golangci-lint` is never installed:
it runs in Docker (see `script/lint`). `docker` is not installed either;
testing, linting and the image build need it. Ends with `go mod download`.
- `script/setup` — make a fresh clone ready for development: runs
`script/bootstrap`, then `script/install-precommit`.
- `script/projectname` — print this project's name (`sfdupes`). Scripts that
need the name call it, so they stay identical across repositories.
- `script/test` — run the test suite with a 30-second timeout and coverage
enabled, rerunning verbosely on failure so the logs show which test failed.
- `script/lint` — run the linter. It builds `Dockerfile.lint`, which copies the
repository into the digest-pinned `golangci/golangci-lint` image and runs
`golangci-lint config verify` and `golangci-lint run` as build steps, so a
successful build is a clean lint. The linter is never run on the host, which
makes a working `docker` the one prerequisite for linting — and therefore for
`make check` and the pre-commit hook. Offline machines: the gate steps
themselves make no network calls. `golangci-lint run` does not, and neither
does `golangci-lint config verify` — it validates against a schema the pinned
binary embeds, measured under `--network none` to both pass a valid config and
reject an invalid one. The build around them does. `Dockerfile.lint` runs
`go mod download` before the gates and this module has external dependencies,
so a first lint on a machine with a cold BuildKit cache reaches the network
there (as well as pulling the pinned image); under `--network none` it fails
at that step, before any gate. That layer sits above the gates and stays
cached, so once it is warm `script/lint` — and with it `make check` — runs
entirely offline, until `go.mod` or `go.sum` changes and the download layer
goes cold again. Because the daemon only ever sees a build context, this works
when the docker daemon is remote and bind mounts are impossible.
- `script/fmt` — format in place: `gofmt -s -w` for Go sources and `prettier`
for Markdown (`--tab-width 4 --prose-wrap always`, the house settings, also
carried in `.prettierrc`). prettier is the pinned devDependency in
`package.json`/`yarn.lock`, installed by `script/bootstrap`.
- `script/fmt-check` — the read-only counterpart of `script/fmt`: runs both
checks, reports each independently so it is clear which failed, and exits
non-zero if either found unformatted files instead of writing.
- `script/test` — run the test suite under the race detector, with a 90-second
timeout and coverage enabled, rerunning verbosely on failure so the logs show
which test failed. It builds the `Dockerfile`'s `test` phase alone and writes
no image. The race detector needs cgo and a C compiler, which the build never
uses, so the phase starts from a digest-pinned Debian `golang` image, which
has `gcc`. The tests run as `nobody`, because several of them make a file
unreadable and root reads it anyway.
- `script/lint` — run the linter. It builds the `Dockerfile`'s `lint` phase
alone and writes no image: in the digest-pinned `golangci/golangci-lint`
image, the gofmt check, `golangci-lint config verify` and `golangci-lint run`
run as build steps, so a successful build is a clean lint. The linter is never
run on the host, which makes a working `docker` the one prerequisite for
linting.
- `script/fmt` — format in place: the Go sources with `gofmt -s -w`, and every
Markdown file with prettier, at the settings in `.prettierrc` (4-space
indents, prose wrapped at 80 columns). Both run on the host, prettier through
yarn at the version `yarn.lock` pins. When yarn is not on `PATH`, this loads
the node 22.17.0 that `script/bootstrap` installed through nvm.
- `script/fmt-check` — the read-only counterpart of `script/fmt`: prints any
unformatted file and exits non-zero instead of writing. gofmt and prettier
both run every time, and each names itself when it fails. The `Dockerfile`'s
`lint` phase runs the same gofmt check.
- `script/check` — run `script/test`, `script/lint`, and `script/fmt-check`, in
that order. Modifies nothing. Needs `docker`, because `script/lint` does.
that order. Modifies nothing. The first two need `docker`, the third what
`script/bootstrap` installs.
- `script/docker` — build the Docker image, tagged with the name from
`script/projectname`. The `Dockerfile` runs the gates as build steps, so this
is also the check a developer or reviewer runs by hand.
- `script/cibuild` — build the Docker image untagged. This is what the Gitea
workflow runs on push; because the gates run as build steps, a successful
build implies the repository is green.
`script/projectname`, passing the output of
`git describe --tags --always --dirty` (or `unknown` when that is empty) as
the `VERSION` build argument. The `Dockerfile`'s build stage depends on its
`lint` and `test` phases, so this also runs the tests and the linter.
- `script/cibuild` — run `script/bootstrap`, then `script/check`, then build the
image as `script/docker` does. This is what the Gitea workflow runs on push; a
successful run means every check passed. The tests and the linter run twice:
once in `script/check`, and again as stages of the image build.
- `script/precommit` — run by the git pre-commit hook: `go mod tidy` must be a
no-op (a resulting change to `go.mod` or `go.sum` fails the commit), then
`script/check`.
- `script/install-precommit` — install the git pre-commit hook that runs
`script/precommit`. The hook is written to the common git directory, so the
main checkout and every worktree share it.
- `script/verify-lint-image-pin` — fail unless the `golangci/golangci-lint`
reference in `Dockerfile.lint` and the one in the `Dockerfile` lint stage are
the same image at the same digest, naming both if not. The linter is pinned in
those two files and nothing else keeps them in sync, so a bump applied to one
alone would leave `make lint` and the `Dockerfile`'s fail-fast lint stage
checking the same tree against different rulesets, both green. The guard
restates neither pin — a third copy would be the same drift one file further
out — and runs as a gate in both files, so `make lint`, `make check` and
`make docker` all catch it.
- `script/install-precommit` — install the git pre-commit hook, as
`.git/hooks/pre-commit`, that runs `script/precommit`.
`script/verify-linter-pin` used to live here. It compared a linter binary
against a version pin in `script/bootstrap`, and both of its subjects are gone:
no linter binary is copied between build stages any more, and bootstrap pins no
version because it installs no linter. The drift it existed to catch has moved
from binary-versus-pin to pin-versus-pin, which is what
`script/verify-lint-image-pin` above checks.
`script/lint`, `script/docker` and `script/cibuild` all pass a freshly computed
`CHECK_EPOCH` build argument, and the gate steps in `Dockerfile.lint` and
`Dockerfile` reference it. Without that, an unchanged tree lets Docker serve the
gate layers from cache and the build exits 0 having executed no tests and no
lint — a green it never earned, and one this repository has produced twice.
`CHECK_EPOCH` invalidates the gate layers on every run while leaving the pinned
base images and the dependency layers cached. `script/lint`'s value carries the
process id as well as the epoch, because two lint runs land inside the same
second easily and a bare epoch would cache the second one.
Every `docker build` in `script/` passes `--no-cache`. On an unchanged tree
Docker would serve the gate steps from its cache, and the build would pass
having run no test and no linter. Every run therefore downloads the Go modules
again and needs the network.
## Build
@@ -503,14 +813,15 @@ compile recipe:
the default target.
- `make bootstrap` — install the build and development dependencies.
- `make setup` — prepare a fresh clone: `bootstrap` plus the pre-commit hook.
- `make test` — run the test suite (30-second timeout; reruns with `-v` on
failure).
- `make lint` — run `golangci-lint` with the repo config, in Docker (see
`script/lint`); requires `docker`.
- `make fmt` / `make fmt-check` — format Go and Markdown sources / verify both
without writing.
- `make test` — run the test suite under the race detector, in Docker (90-second
timeout; reruns with `-v` on failure; see `script/test`); requires `docker`.
- `make lint` — run `golangci-lint` with the repo config and the gofmt check, in
Docker (see `script/lint`); requires `docker`.
- `make fmt` / `make fmt-check` — format the Go sources and the Markdown /
verify formatting without writing, on the host; requires `go` and what
`make bootstrap` installs (see `script/fmt`).
- `make check` — `test`, `lint`, and `fmt-check`; modifies nothing. Requires
`docker`, via `lint`.
`docker` and what `make bootstrap` installs.
- `make docker` — build the Docker image, which runs the gates as build stages.
- `make hooks` — install the pre-commit hook.
- `make clean` — remove the binary.
@@ -519,14 +830,14 @@ compile recipe:
All of the following, run in this directory, must pass:
1. `make check` passes (tests, lint, `gofmt`).
1. `make check` passes (tests, lint, `gofmt`, prettier).
2. `make docker` succeeds.
3. Smoke test — create a throwaway tree in a temp dir (never test against real
data):
```sh
d=$(mktemp -d)
export SFDUPES_DATABASE="$d/db.sqlite"
export SFDUPES_DATABASE="$(mktemp -d)/db.sqlite"
mkdir -p "$d/a" "$d/b"
head -c 2000 /dev/urandom > "$d/a/one.bin"
cp "$d/a/one.bin" "$d/b/copy.bin"
@@ -554,8 +865,9 @@ All of the following, run in this directory, must pass:
./sfdupes report
```
(The scan database lives inside `$d` here purely for test hygiene; scanning
`$d` therefore also records the SQLite file itself, which is harmless.)
(The database lives in a temp directory of its own: inside `$d`, the scan
would record it, and its empty lock file would join the `empty1`/`empty2`
group.)
Expected from the first `report`: `one.bin`/`copy.bin`/`copy2.bin` form one
group (two dupe rows, `first` is the lexicographically smallest path);
@@ -584,8 +896,10 @@ Tracked in [TODO.md](TODO.md).
## Non-goals
- No full-content verification, no byte-for-byte compare, no deletion or linking
of duplicates. The reports are advisory; acting on them is the user's job.
- No byte-for-byte compare, and no deletion or linking of duplicates. Files that
match are compared by a SHA-256 of the whole file below 50 MiB, and only by
samples at 50 MiB and over. The reports are advisory; acting on them is the
user's job.
- No persistence beyond the SQLite database described above; no export/import
formats.
- No daemon or filesystem watcher; scheduling rescans is cron's job.
+388 -84
View File
@@ -1,6 +1,6 @@
---
title: Repository Policies
last_modified: 2026-07-06
last_modified: 2026-10-07
---
This document covers repository structure, tooling, and workflow standards. Code
@@ -60,17 +60,28 @@ style conventions are in separate documents:
prerequisite since nvm requires bash. yarn is then pinned via
`corepack prepare yarn@<version> --activate`. Never install "latest" or "lts";
always exact versions. `script/cibuild` runs the CI build: it changes to the
repo root and runs `docker build .`; the Gitea workflow calls it. Four further
scripts are our own extensions to the standard: `script/check` runs
`script/test`, `script/lint`, and `script/fmt-check`; `script/precommit` is
what the git pre-commit hook runs, and it calls `script/check`;
`script/install-precommit` installs the git pre-commit hook (the `make hooks`
target shims to it); and `script/projectname` (literally that filename) simply
outputs the project's name. Scripts that need the name call
`script/projectname` — e.g. `script/docker` assembles its image tag from it —
so those scripts stay byte-identical across all repos. Repo-type-specific
pre-commit extras (e.g. `go mod tidy` verification in Go repos) belong in
`script/precommit`, not in the hook itself. Model scripts are at
repo root, runs `script/bootstrap`, runs `script/check`, and builds the image
with the version; the Gitea workflow calls it. **`script/cibuild` runs
`script/bootstrap` first**, because the workflow checks out the repo and runs
nothing else, while `script/fmt-check` runs the formatter on the host: on a
pristine checkout with nothing installed the run dies there, after the
containerised gates have passed. **The bootstrap alone is not enough**:
`script/bootstrap` installs node and yarn under nvm and leaves neither on the
`PATH` of the shell that called it, so a bare `yarn` still exits 127. The host
entrypoints that need yarn — `script/fmt` and `script/fmt-check` — therefore
source nvm for the pinned node version before invoking it, exactly as
`script/bootstrap`'s own install step does. A runner carrying nothing but
docker and git then gets through `script/check`. Four further scripts are our
own extensions to the standard: `script/check` runs `script/test`,
`script/lint` and `script/fmt-check`; `script/precommit` is what the git
pre-commit hook runs, and it calls `script/check`; `script/install-precommit`
installs the git pre-commit hook (the `make hooks` target shims to it); and
`script/projectname` (literally that filename) simply outputs the project's
name. Scripts that need the name call `script/projectname` — e.g.
`script/docker` assembles its image tag from it — so those scripts stay
byte-identical across all repos. Repo-type-specific pre-commit extras (e.g.
`go mod tidy` verification in Go repos) belong in `script/precommit`, not in
the hook itself. Model scripts are at
`https://git.eeqj.de/sneak/prompts/raw/branch/main/script/<name>`. The README
must document the provided scripts in an **Entrypoints** section (see the
README requirements below).
@@ -89,87 +100,222 @@ style conventions are in separate documents:
contributor should be able to understand the entire development workflow by
reading the Makefile.
- Every repo should have a `Dockerfile`. All Dockerfiles must run `make check`
as a build step so the build fails if the branch is not green. For non-server
repos, the Dockerfile should bring up a development environment and run
`make check`. For server repos, `make check` should run as an early build
stage before the final image is assembled. Dockerfiles install development
prerequisites by running `script/bootstrap` rather than duplicating installs
inline; COPY `script/` and the dependency manifests (`package.json` +
`yarn.lock`, `go.mod` + `go.sum`, etc.) before running it so the bootstrap
layer stays cached until dependencies change.
- Every repo should have a `Dockerfile`, and it carries the repo's gates: a
`lint` phase and a `test` phase, with the final stage depending on both so the
image cannot be built unless they pass. For non-server repos the final stage
brings up a development environment; for server repos it is the runtime image.
The gate phases and the build stage start from their pinned base images and
install what those images lack either inline, as the canonical Go `Dockerfile`
below does for `git`, or by running `script/bootstrap`, as the `prompts`
repo's own `Dockerfile` does for its yarn packages. The development
environment stage installs development prerequisites by running
`script/bootstrap` rather than duplicating its installs inline. A stage that
runs `script/bootstrap` COPYs `script/` and the dependency manifests
(`package.json` + `yarn.lock`, `go.mod` + `go.sum`, etc.) before running it.
- **Dockerfiles must use a separate lint stage for fail-fast feedback.** Go
repos use a multistage build where linting runs in an independent stage based
on the `golangci/golangci-lint` image (pinned by hash). This stage runs
`make fmt-check` and `make lint` before the full build begins. The build stage
then declares an explicit dependency on the lint stage via
`COPY --from=lint /src/go.sum /dev/null`, which forces BuildKit to complete
linting before proceeding to compilation and tests. This ensures lint failures
surface in seconds rather than minutes, without blocking on dependency
download or compilation in the build stage.
- **Linting and testing run in Docker, as phases of the `Dockerfile`.** There is
no separate lint file. `script/lint` and `script/test` each build one phase
and nothing else:
The standard pattern for a Go repo Dockerfile is:
```sh
docker build --no-cache --target lint --output type=cacheonly .
docker build --no-cache --target test --output type=cacheonly .
```
**A stage that is not the last one in the file is built only when the final
stage's chain depends on it, or when `--target` names it.** That is why the
two gates are always invoked by name here, and why the final stage carries a
`COPY --from=` of a harmless file from each of them: without that edge a
plain `docker build .` builds the last stage alone and exits 0 having linted
and tested nothing.
**The gate builds write no image.** With `--output type=cacheonly` the phase
runs and a failing step fails the build, but the result is not exported.
Nothing uses those images, and writing one out is slow: a Go test phase's
image holds the toolchain and every compiled package. A build given neither
`--output` nor `-t` writes an untagged image and leaves it dangling, on
every developer host and every CI runner. `script/cibuild` and
`script/docker` build the image that ships and tag it, so each build
replaces the previous image; each assigns the tag on its own line before the
build, so `set -e` stops it where `script/projectname` fails.
Inside a phase the tool is invoked directly — `golangci-lint`, `go test`,
`eslint`, `prettier` — never through `make lint` or `script/test`, which are
themselves a `docker build` and would recurse into a daemon that does not
exist in a build step. Formatting is the exception and stays on the host:
`script/fmt` writes the working tree, and `script/fmt-check` is its
read-only twin.
**No lint verdict may come from a host invocation of the linter.** On a
shared host golangci-lint reads a result cache keyed on file content rather
than location, so a second checkout of the same content is served the first
one's findings, and a host-global lock in `$TMPDIR` makes concurrent runs
exit non-zero with `parallel golangci-lint is running` — a status a caller
cannot tell from real findings. Both have produced wrong verdicts in this
org, in both directions. A container has its own cache, its own `TMPDIR` and
a digest-pinned binary, so neither is reachable.
- **Any build that runs checks is built with `--no-cache`.** Docker invalidates
a `COPY` layer only when the copied content changes, so on an unchanged tree
the check `RUN` is served from cache, nothing executes, and the build still
exits 0. Every `docker build` in `script/` therefore passes `--no-cache`:
`script/lint`, `script/test`, `script/cibuild` and `script/docker` are the
four, and there is no fifth — `script/check` runs the two gate phases and
`script/fmt-check`, and builds no image of its own. A bare `docker build .` is
not evidence that anything ran: a sub-second build reporting success is a
cache hit, not a result. Never invalidate by pruning — `docker builder prune`
and friends destroy a build cache shared with every other build on the host.
When a check is added or changed, prove it works by planting a defect it must
catch and watching the run fail on it, then revert the defect. A green run
alone shows neither that the check ran nor that it covers what it should.
- **The gate phases are separate stages, and the build stage depends on both.**
The lint phase is based on the `golangci/golangci-lint` image (pinned by
hash), so lint failures surface in seconds rather than after a full compile,
and the test phase is based on the Debian Go image. The canonical Go repo
`Dockerfile`:
```dockerfile
# Lint stage — fast feedback on formatting and lint issues
# Lint phase
# golangci/golangci-lint:v2.x.x, YYYY-MM-DD
FROM golangci/golangci-lint@sha256:... AS lint
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make fmt-check
RUN make lint
RUN golangci-lint run --config .golangci.yml ./...
# Build stage
# golang:1.x-alpine, YYYY-MM-DD
FROM golang@sha256:... AS builder
# Test phase. -race needs cgo and so a C compiler, which the Debian Go
# image ships and the alpine one does not.
# golang:1.x, YYYY-MM-DD
FROM golang@sha256:... AS test
WORKDIR /src
# Force BuildKit to run the lint stage before proceeding
COPY --from=lint /src/go.sum /dev/null
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make test
RUN go test -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 90s -race -v ./...; exit 1; }
ARG VERSION=dev
RUN CGO_ENABLED=0 go build -trimpath \
# Build stage. Nothing is wanted from either phase above; the copies
# are what make BuildKit build them first, so this stage cannot run
# unless lint and test passed.
# golang:1.x-alpine, YYYY-MM-DD
FROM golang@sha256:... AS builder
COPY --from=lint /src/go.sum /dev/null
COPY --from=test /src/go.sum /dev/null
RUN apk add --no-cache git
# A tar-stream context keeps the sender's file owners, which git refuses.
RUN git config --system --add safe.directory /src
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# The VERSION build arg when one is given, otherwise
# `git describe --tags --always` on the .git in the build context. With
# .git present, a version that is still empty, dev or unknown fails the
# build: git is missing or could not read the checkout.
ARG VERSION
RUN VERSION="${VERSION:-$(git describe --tags --always)}"; \
if [ -e .git ]; then \
case "$VERSION" in ""|dev|unknown) \
echo "version is '$VERSION' although .git is present" >&2; \
exit 1 ;; \
esac; \
fi; \
CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X main.Version=${VERSION}" \
-o /app ./cmd/app/
# Runtime stage
# Runtime stage, and the last one
FROM alpine@sha256:...
COPY --from=builder /app /usr/local/bin/app
ENTRYPOINT ["app"]
```
Key points:
- The lint stage uses the `golangci/golangci-lint` image directly (it
includes both Go and the linter), so there is no need to install the
linter separately.
- `COPY --from=lint /src/go.sum /dev/null` is a no-op file copy that creates
a stage dependency. BuildKit runs stages in parallel by default; without
this line, the build stage would not wait for lint to finish and a lint
failure might not fail the overall build.
- The lint phase uses the `golangci/golangci-lint` image directly (it has
both Go and the linter), so nothing needs installing.
- `COPY --from=<phase> /src/go.sum /dev/null` is a no-op copy whose only
purpose is the ordering edge. BuildKit runs stages in parallel by default,
and a stage nothing depends on is not built at all, so without these two
lines a red gate would not fail the build.
- Keep the runtime stage last, and if you add a stage after it, give it the
same two copies. A plain `docker build .` builds the last stage's chain
and nothing else.
- If the project uses `//go:embed` directives that reference build artifacts
(e.g. a web frontend compiled in a separate stage), the lint stage must
(e.g. a web frontend compiled in a separate stage), the lint phase must
create placeholder files so the embed directives resolve. Example:
`RUN mkdir -p web/dist && touch web/dist/index.html web/dist/style.css`.
The lint stage should not depend on the actual build output — it exists to
fail fast.
- If the project requires CGO or system libraries for linting (e.g.
`vips-dev`), install them in the lint stage with `apk add`.
- The build stage runs `make test` after compilation setup. Tests run in the
build stage, not the lint stage, because they may require compiled
artifacts or heavier dependencies.
- If the project requires CGO or system libraries for linting, install them
in the lint phase. The `golangci/golangci-lint` image is Debian-based and
has no `apk`, so install with `apt-get` under the Debian package name
(`libvips-dev`, where alpine says `vips-dev`), and delete the package
lists in the same `RUN`, so the layer does not keep them:
```dockerfile
RUN apt-get update \
&& apt-get install -y --no-install-recommends libvips-dev \
&& rm -rf /var/lib/apt/lists/*
```
- `.dockerignore` lets `.git` into the build context. It keeps out every git
`config` at any depth (`**/.git/config`, `**/.git/modules/**/config`): the
repository's own, each submodule's under `.git/modules/`, and that of a
submodule keeping its own `.git` directory. `git describe` does not need
them, and each can hold a credential: a password in a remote URL, or the
token the CI checkout step stores there. A submodule whose name has a
`config` segment (`config`, `deploy/config`, `config/lib`) loses its whole
git directory to `**/.git/modules/**/config`, and Go's version stamping
then fails the build: give it a name without that segment
(`git submodule add --name`). The stage that compiles has `git` (the
Debian Go image has it; an alpine one needs `apk add --no-cache git`) and
takes the version from the `VERSION` build argument when one is given,
otherwise from `git describe --tags --always`. That gives the tag on a
tagged commit; on a later commit, the tag, the number of commits since it
and the short commit (`v1.2.3-4-gabc1234`); and the short commit when no
tag is reachable. The stage that compiles also marks its working directory
safe for git (`git config --system --add safe.directory /src`): a context
sent as a tar stream keeps the sender's file owners, and git refuses a
checkout owned by another user, so the version would come out empty.
`ARG VERSION` has no default, and the build fails if the context carries
`.git` and the version still comes out empty, `dev` or `unknown`. A plain
`docker build .` with no build arguments must succeed; a Dockerfile that
refuses an empty build argument drops that refusal and keeps the argument.
A checkout whose `.git` is a file (a linked worktree, or a repository
checked out as a submodule) is the exception: that file points to a git
directory outside the build context, so the build cannot read the version
and a plain `docker build .` fails; pass the version with
`--build-arg VERSION=...`, as `script/docker` and `script/cibuild` already
do.
- Every repo should have a Gitea Actions workflow (`.gitea/workflows/`) that
runs `script/cibuild` (which runs `docker build .`) on push. Since the
Dockerfile already runs `make check`, a successful build implies all checks
pass.
runs `script/cibuild` on push, and checks out the repo as its only other step,
with `persist-credentials: false`: `script/cibuild` needs no token, and
without it the checkout leaves the job's token in `.git/config` for every
later step. The checkout step also sets `fetch-depth: 0`, which fetches the
tags `git describe` needs: by default it clones shallow with no tags, and a
tagged repository's CI build would stamp a bare short commit id. The
workflow's `concurrency` block groups runs by workflow and branch
(`${{ github.workflow }}-${{ github.ref }}`) with `cancel-in-progress: true`,
so a new push cancels the older run on the same branch, queued or running, and
no other: runs for replaced commits do not hold up the shared runner.
`script/cibuild` bootstraps, runs the gate phases, and then builds the image,
so a successful run means every check passed; a bare `docker build .` does not
carry the same guarantee, because its gate phases may come from the cache. The
image build is uncached and so runs the gate phases a second time. That is the
price of the rule above, and it is worth paying: the image that ships is built
from a run of its own gates rather than from a cache entry. The `check` job
sets `timeout-minutes: 20`, so a hung build frees the shared runner after 20
minutes. That allows for the three Docker builds described above (the test
phase, the lint phase, then the image), each held to the 5-minute Docker build
limit below, plus the bootstrap. A separate workflow limited to `main` by a
`branches` list under `on: push` cannot be checked by review: to try a change
to it, add the feature branch to that list and push, then remove the branch
from the list again before merging. Keep any job in it that publishes behind
`if: github.ref_name == 'main'`, so the run from the feature branch publishes
nothing.
- Use platform-standard formatters: `black` for Python, `prettier` for
JS/CSS/Markdown/HTML, `go fmt` for Go. Always use default configuration with
@@ -189,14 +335,21 @@ style conventions are in separate documents:
module under test to verify it compiles/parses. There is no excuse for
`make test` to be a no-op.
- `make test` must complete in under 20 seconds. Add a 30-second timeout in the
Makefile.
- `make test` must complete in under 60 seconds. That is the hard cap, and a
suite that exceeds it fails. Under 20 seconds is the target. A suite between
20 and 60 seconds is still green, but the overage must be filed as an
improvement bug against that repo. Add a 90-second timeout to the test
invocation (`go test -timeout 90s`). The backstop deliberately sits above the
hard cap so that it catches a genuinely hung test rather than a merely slow
one.
- **`make test` should use the conditional verbose rerun pattern.** Run tests
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to
show full output. This keeps CI logs and `docker build` output clean on
success (just package/suite summaries) while providing full diagnostic detail
on failure (every test case, every assertion). The general shell pattern:
- **The test command should use the conditional verbose rerun pattern.** Run
tests without `-v` (verbose) first. If tests fail, automatically rerun with
`-v` to show full output. This keeps CI logs and `docker build` output clean
on success (just package/suite summaries) while providing full diagnostic
detail on failure (every test case, every assertion). The command lives in the
`test` phase of the `Dockerfile`, since `script/test` builds that phase; the
Makefile form below is the same pattern for any repo-local invocation:
```makefile
test:
@@ -209,11 +362,26 @@ style conventions are in separate documents:
```makefile
test:
@go test -timeout 30s -race -cover ./... || \
@go test -count=1 -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 30s -race -v ./...; exit 1; }
go test -count=1 -timeout 90s -race -v ./...; exit 1; }
```
`-count=1` is required on both invocations: it defeats Go's test _result_
cache, so neither run can report a stored pass in place of running the
tests. It leaves the build cache alone, so it costs the runtime of the suite
and no recompilation.
That cache is Go's own, separate from Docker's layer cache. Go stores a
passing result in its cache directory (`GOCACHE`), and when the same tests
run again on unchanged code it prints that result, marked `(cached)`,
without running them. That matters on a developer's machine, where this
target runs and the directory lasts from one run to the next. The `test`
phase of the `Dockerfile` needs no `-count=1`: its base image holds no
result for this repo's tests and nothing before its `go test` step runs a
test, so there is nothing to replay. `--no-cache` (above) is what makes that
step run on an unchanged tree.
Python example:
```makefile
@@ -239,10 +407,89 @@ style conventions are in separate documents:
must be in `.gitignore`. No exceptions.
- `.gitignore` should be comprehensive from the start: OS files (`.DS_Store`),
editor files (`.swp`, `*~`), language build artifacts, and `node_modules/`.
Fetch the standard `.gitignore` from
editor files (`.swp`, `*~`), in-repo agent scratch directories (`.claude/`),
`node_modules/`, and the repo's own build outputs. Fetch the standard
`.gitignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.gitignore` when setting up
a new repo.
a new repo. A repo's `.gitignore` is the standard file followed by the repo's
own entries, such as its binaries; a re-vendor replaces the standard part and
keeps those entries. These patterns are written to `.gitignore`'s own
semantics, in which an unanchored pattern already matches at every depth; they
are not a `.dockerignore` and must not be transplanted into one unmodified.
- **`.dockerignore` does not use `.gitignore` semantics, and copying patterns
across unmodified leaves secrets in the build context.** Docker matches with
`moby/patternmatcher`: `filepath.Match` semantics plus a `**` extension, so
`*` does not cross `/` and a pattern without a leading `**/` is anchored at
the build-context root. A `.dockerignore` listing `.env`, `*.pem` and `*.key`
therefore excludes only the copies at the repository root, while `config/.env`
and `certs/server.key` still reach the context and can land in an image layer
— which is more dangerous than a short file with no secret patterns at all,
because it reads as solved and stops anyone looking. Give every
depth-independent pattern the `**/` prefix and leave only genuinely
root-anchored entries unprefixed: `.claude`, and the repo's own host-built
binary, written `/myapp` and never `**/myapp`, which would also match
`cmd/myapp/` and delete the package directory from the context. Matching is
case-sensitive, and an ALL-CAPS twin per pattern still misses `Server.Key`, so
secret names use character ranges — `**/*.[kK][eE][yY]`, `**/*.[pP][eE][mM]`,
and likewise for `.envrc` and the extensionless SSH keys. Where such a pattern
also catches something the build needs, re-include it with a negation
(`!docs/example.env`); deleting the pattern reopens the exposure for every
other file it covers. Fetch the standard `.dockerignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.dockerignore` and extend
it with the repo's own artifacts.
- **In-repo agent scratch belongs in both files, written to each file's own
semantics.** `.claude/` holds one worktree per in-flight agent — an entire
additional checkout of the repo — so under `COPY . .` the build context
inflates by a multiple of the repo and another session's unreviewed work can
be copied into an image layer. In `.gitignore` the entry is `.claude/`,
unanchored. In `.dockerignore` it is `.claude`, anchored and with **no** `**/`
prefix, because the prefixed form would also delete any nested directory of
that name from the build. Anchoring carries a known gap that the canonical
`.dockerignore` states in its own comment, since consuming repos receive the
file and not the tracker: the directory is created in the agent's working
directory, so a repo running agents in subdirectories still ships
`services/api/.claude/` and must add its own anchored entry there.
- **A plain `docker build .` of a clone stamps the version that
`git describe --tags --always` gives**, derived from the `.git` in the build
context as the canonical `Dockerfile` above shows. Without its failure check,
a missing `git` or an unreadable checkout would leave `-X main.Version=` empty
and the build would still exit 0. `script/docker` and `script/cibuild` pass
the version they compute on the host; it takes precedence. They do this
byte-identically across repos:
```sh
# The version and the tag each get their own line: a failing command
# substitution inside an argument does not trip `set -e`, so the inline
# form degrades to an empty constant.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
tag="$(script/projectname)"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$tag" .
```
`--always` makes an untagged repo yield an abbreviated commit hash rather
than failing, and the `[ -n "$version" ]` line is the single place the
fallback is applied — a live check that fires on a build from an export with
no `.git` and on a repository with no commits yet. Do not fold it into the
substitution as `|| echo unknown`, which makes the guard unreachable. The
Dockerfile's side is `ARG VERSION` in the stage that compiles, declared
there because `ARG` is stage-scoped; passing `VERSION` to a repo whose
Dockerfile declares no such `ARG` is ignored and costs nothing, which is why
the scripts stay byte-identical. One consequence for CI: the standard
checkout action clones shallow and fetches no tags, so the canonical
`.gitea/workflows/check.yml` sets `fetch-depth: 0` on its checkout step.
- **Verify `.dockerignore` by enumerating the image, not by reading the
patterns.** Plant files at the root _and_ at least two directories deep, build
a probe image that does `COPY . .`, and list what actually landed
(`docker run --rm --entrypoint find IMAGE /app`). The `transferring context`
size is not a substitute: a nested secret is a few bytes, and BuildKit
transfers only the delta from the previous build.
- **No build artifacts in version control.** Code-derived data (compiled
bundles, minified output, generated assets) must never be committed to the
@@ -258,9 +505,56 @@ style conventions are in separate documents:
- Make all changes on a feature branch. You can do whatever you want on a
feature branch.
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only
manually by the user. Fetch from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`.
- `.golangci.yml` is standardized. The vendored copy in a consuming repo must
_NEVER_ be modified by an agent: fetch it from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml` and keep it
byte-identical, so that no repo can quietly loosen its own linting. Linter
configuration changes are made to the canonical copy in the `prompts` repo and
reach consuming repos by re-vendoring; an agent may open a PR against
canonical, which only the user merges. One list is exempt from byte-identity,
because it cannot be written once for every repo: the `deny` list of the
`test-support` depguard rule, where a repo names its own test-support packages
by full import path. A repo adds entries there and changes nothing else, and a
re-vendor carries its entries forward. The canonical golangci-lint version is
v2.14.0 (released 2026-09-24), pinned as the digest of the lint phase's base
image
(`golangci/golangci-lint@sha256:ad862ba6b3798cbe0fd9fd7408d498fd74fbd2623a92406b2fd3898faf0bf98f`,
which reports `2.14.0 built with go1.27.0 from 114493f9`). A module's `go`
directive must not name a newer Go minor version than the one golangci-lint
was built with, or golangci-lint refuses to lint it: this release lints
`go 1.27.1` but not `go 1.28`. That digest is the only pin, since no repo
installs golangci-lint on the host. A repo sets the lint phase digest to the
one named here and re-vendors `.golangci.yml` in the same commit, whichever of
the two prompted the change: the canonical copy can name linters that an older
golangci-lint rejects, and a newer golangci-lint can add linters that
`default: all` switches on until the canonical copy disables them.
- **`script/bootstrap` installs a pinned tool by comparing versions, never by
testing presence.** An `if ! command -v <tool>; then install; fi` guard tests
`PATH` only, so on an already-provisioned machine the pin is inert and a
version bump is a silent no-op — while the Dockerfile, installing into a clean
image, gets the pinned version, so a local `make check` and `make docker` can
disagree about what the tool even is. The canonical form:
- compares the installed version against the pin over the **whole** version
token; a parser that stops at the first `-` reports `2.12.2` for a host
running `2.12.2-rc1` and skips the install;
- treats absent, non-zero, empty or unrecognised `--version` output as a
mismatch, so the failure direction is a redundant install and never a
skipped one;
- after installing, re-resolves the binary the way callers do — `hash -r`,
then through `PATH`, not through the directory the installer wrote to —
and fails naming the resolved path, since an install that a shadowing
binary hides succeeds while changing nothing any caller sees;
- is actually called, and prints the version on both success paths: a
function defined and never invoked has the same exit status and the same
empty output as one that worked.
Keep it POSIX sh: no arrays, no `[[`, no `grep -P`.
A Go tool a repo needs on the host is installed with `go install` pinned to
a commit hash (`go install <package>@<commit hash>`). It is never tracked as
a `go.mod` tool dependency or through a `tools.go` file, either of which
pulls the tool's own dependencies into the repo's `go.mod` and `go.sum`.
- When pinning images or packages by hash, add a comment above the reference
with the version and date (YYYY-MM-DD).
@@ -371,15 +665,21 @@ style conventions are in separate documents:
Never edit existing migrations after release.
- All repos should have an `.editorconfig` enforcing the project's indentation
settings.
settings: the standard file from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.editorconfig`, which sets
tabs for `Makefile` and Go files, followed by the repo's own sections, such as
one for another language it uses. A re-vendor replaces the standard part and
keeps those sections.
- Avoid putting files in the repo root unless necessary. Root should contain
only project-level config files (`README.md`, `Makefile`, `Dockerfile`,
`LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`, and
language-specific config). Everything else goes in a subdirectory. Canonical
subdirectory names:
only project-level config files (`README.md`, `AGENTS.md`, `Makefile`,
`Dockerfile`, `LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`,
and language-specific config). Everything else goes in a subdirectory.
Canonical subdirectory names:
- `bin/` — executable scripts and tools
- `cmd/` — Go command entrypoints
- `cmd/` — Go command entrypoints; thin only: one `main.go` per binary whose
body is a single call into `internal/` or `pkg/`, no project logic in
`cmd/`
- `configs/` — configuration templates and examples
- `deploy/` — deployment manifests (k8s, compose, terraform)
- `docs/` — documentation and markdown (README.md stays in root)
@@ -406,3 +706,7 @@ style conventions are in separate documents:
- Go: `go.mod`, `go.sum`, `.golangci.yml`
- JS: `package.json`, `yarn.lock`, `.prettierrc`, `.prettierignore`
- Python: `pyproject.toml`
- Guidance for coding agents lives in one `AGENTS.md` at the repository root. It
is never committed under a file or directory named after one agent tool, such
as `CLAUDE.md` or `.claude/`, and never split into separate memory files.
+278 -255
View File
@@ -2,14 +2,15 @@
- take an issue from the `1.0.0` milestone on the tracker; work not yet on the
tracker gets filed as an issue first
- branch (from `main`)
- branch from `next`
- do the work, with tests, in small focused commits
- record it at the top of Completed Steps (`TODO.md` changes in the same commit
as the work)
- push the branch and open a PR whose title ends with ` (closes #N)`
- an independent review gates the merge; every finding is addressed or
explicitly rebutted on the PR
- merge to `main` once the review passes
- push the branch and open a PR against `next` whose title ends with
` (closes #N)`
- an independent review gates each merge to `next`; every finding is addressed
or explicitly rebutted on the PR
- only the owner merges `next` to `main`
# Status
@@ -28,269 +29,294 @@
# Completed Steps
- restore Markdown formatting in `script/fmt`/`fmt-check` and reformat all
Markdown to the house prettier settings (2026-09-21, closes
- re-vendor the canonical files and model scripts from `sneak/prompts` `next` at
`c55a0cb`: golangci-lint v2.14.0; lint and test are phases of the `Dockerfile`
that write no image, and `make test` runs the suite under the race detector,
so `Dockerfile.lint`, `script/verify-lint-image-pin` and `make test-race` are
gone; every `docker build` in `script/` passes `--no-cache`; prettier runs on
the host, from the node and yarn `script/bootstrap` installs, so the
`prettier` and `markdown` stages are gone and the build stage installs `git`
and `make` itself; a new push cancels the workflow's older run on the same
branch, and a run stops after 20 minutes; `.claude/settings.json` is deleted
(2026-10-08, https://git.eeqj.de/sneak/sfdupes/issues/95)
- `scan` records mtime to the nanosecond, as whole seconds in `mtime` plus
`mtime_nsec`, and compares it at that resolution, so a same-size rewrite
within the same second is re-hashed (2026-10-07,
https://git.eeqj.de/sneak/sfdupes/issues/12)
- cut the narration from `TODO.md` Completed Steps and from the comments in
`script/` and both Dockerfiles; §Workflow now branches from and merges to
`next` (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/49)
- `make test-race` ran the test suite under the race detector in a cgo-enabled
container, outside `make check` (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/18). Since
https://git.eeqj.de/sneak/sfdupes/issues/95 `make test` itself runs the suite
under the race detector, in the `Dockerfile`'s Debian-based `test` phase, and
`make test-race` is gone
- a bare `docker build .` failed, naming `script/cibuild` and `script/docker`,
rather than serve the gates from cache (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/39). Since
https://git.eeqj.de/sneak/sfdupes/issues/95 a bare build succeeds, as
`REPO_POLICIES.md` requires, and may serve the gates from cache; the builds in
`script/` pass `--no-cache`, so theirs always run
- `make fmt` and `make fmt-check` run prettier over all Markdown, and CI checks
it; all Markdown reformatted (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/19)
- `script/lint` writes no image, so a run no longer leaves an untagged one
behind (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/48)
- tests cover a missing database, `scan` keeping stdout empty, its skip warning,
the `report` and `trees` summary lines, and every subcommand going through
`runE` (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/16)
- `.golangci.yml` replaced with the current canonical copy, which uses
`gomodguard_v2`, so lint no longer prints a deprecation warning (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/26)
- a test fails when either `hashWorker` cancellation check in `scan.go` is
removed (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/83)
- a database path holding `?`, `#` or `%` opens exactly the file it names
(2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/55)
- `scan` rejects `--workers` below 1 as a usage error instead of running
single-threaded (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/10)
- a test fails when either walk cancellation check in `scan.go` is removed
(2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/81)
- test that `scan` refuses a database with another schema version (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/64)
- correct four inaccurate comments in `cancel_test.go` and rename
`walkCancelInFlightDirs` to `walkCancelInFlightFiles` (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/33)
- test the `-x` filesystem-boundary rules in `subdirJob` (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/17)
- `scan` creates the schema in one transaction; a version-0 database with a
`files` table is refused with a clear schema-version error (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/11)
- README documents install, Docker, a daily cron scan and how to read and check
the reports (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/54)
- the `Dockerfile` build stage kept the Go module cache out of `builder`'s home
and copied the sources with `--chown`, so no `chown -R` walked them
(2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/43). Since
https://git.eeqj.de/sneak/sfdupes/issues/95 there is no `builder` user and
nothing changes owner: the build stage only compiles, as root, and the tests
run as `nobody` in the `test` phase
- `--version` prints `sfdupes VERSION` to stdout; README documents it and
`--help` (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/15)
- `scan` stops cleanly on `SIGINT` or `SIGTERM`: commits what it has hashed,
deletes nothing more, exits 1 (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/5)
- `report` and `trees` stream the records instead of holding them all in memory;
the schema gains the `files_signature` index (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/14)
- progress prints at once on a non-terminal, uses a real terminal test, and
prints warnings through a spinner instead of racing its redraw (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/13)
- warn about and skip symlink, socket, FIFO, device and `.zfs` operands, keeping
the records beneath them (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/9)
- `scan` holds a lock on a lock file beside the database for its whole run, so a
second `scan` fails at once with exit 1 (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/53)
- test stdout write failures in `report` and `trees`; README states that
`| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null`
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/30)
- `report` and `trees` open the database read-only, and `scan` leaves it out of
WAL mode, so reading needs only read access (2026-10-03, closes
https://git.eeqj.de/sneak/sfdupes/issues/8)
- escape tabs, newlines, carriage returns and backslashes in report, trees and
warning paths; the root directory's path is `/` (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/7)
- stamp the git tag or short commit in a plain `docker build .` instead of `dev`
(2026-10-02, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/67): `.dockerignore` sends `.git`
without `.git/config`; the build stage stamps the `VERSION` build argument,
else `git describe --tags --always`, and fails if the context carries `.git`
and the version is still empty, `dev` or `unknown`. CI checks out the full
history (`fetch-depth: 0`) so it stamps the same value as `make build`.
- replace the 1 KiB end-window sampling with the head/tail plus content-hash
ladder (2026-09-22, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/61); README "Duplicate detection"
documents every rung. A file under 10 MiB is hashed in full, and its `head`,
`tail` and `content` all hold that hash. A larger file gets only its 64 KiB
`head` and `tail` in the hash phase; the content phase, after the update
phase, reads it for `content` (the whole file below 50 MiB, gigabyte-spaced 1
MiB samples at or above) only when its size, `head` and `tail` match another
record's from this scan or an earlier one, and never reads a file gone or
changed since its record was written. `report` and `trees` leave out any
record without a `content` hash. The `content` column is part of the version 1
schema.
- remove the dead `files.dat` references from `Makefile`, `.gitignore` and
`.dockerignore` (2026-09-21, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/22)
- fix the lint-image pin comments and `FROM` form in `Dockerfile` and
`Dockerfile.lint` (2026-08-10, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/25): dropped the false
`(Debian-based)` parenthetical (v2.12.1 was Debian too) and the redundant tag,
so both pins are the policy `# image:vX.Y.Z, YYYY-MM-DD` comment over a bare
`FROM image@sha256:...`. Digest unchanged. `script/verify-lint-image-pin`
parses those `FROM` lines and still matches the tagless form; its advice line
lost the now meaningless "tag and digest". With no tag in either reference, a
tag-only disagreement no longer exists — a one-sided tag is caught as a plain
mismatch.
https://git.eeqj.de/sneak/sfdupes/issues/25): both pins are now the policy
`# image:vX.Y.Z, YYYY-MM-DD` comment over a bare `FROM image@sha256:...`,
without the false `(Debian-based)` note or the tag; digest unchanged.
`script/verify-lint-image-pin` still matches the tagless form, and a tag on
one side only is caught as a plain mismatch.
- run all linting in Docker via `Dockerfile.lint` and `script/lint` (2026-08-10,
branch `next`, closes https://git.eeqj.de/sneak/sfdupes/issues/46): per the
owner ruling, the linter runs inside a container invoked through the `script/`
entrypoint and is never installed on a host. New root `Dockerfile.lint` COPYs
owner ruling the linter is never installed on a host. `Dockerfile.lint` copies
the repo into the digest-pinned `golangci/golangci-lint:v2.12.2` image and
runs `golangci-lint config verify` and `golangci-lint run` as build steps, so
a successful build IS a clean lint; `script/lint` is reduced to building it.
`script/bootstrap` loses the `go install`, the pin constants, the version
parser and `verify_golangci_lint` outright rather than hardening them — with
nothing linting on the host, the `$GOPATH/bin` versus `PATH` problem that
motivated them has no subject — and now warns rather than fails when `docker`
is absent. Two traps handled. A lint build on an unchanged tree returns
success in well under a second having run no linter, which is
https://git.eeqj.de/sneak/sfdupes/issues/32 and
https://git.eeqj.de/sneak/sfdupes/issues/39 again, so `Dockerfile.lint`
carries `ARG CHECK_EPOCH` referenced inside every gate `RUN` (BuildKit hashes
the expanded command, not the declaration) and `script/lint` passes
`"$(date +%s)-$$"` — the PID matters because two lint runs land inside the
same second easily. And nothing inside an image build may shell out to docker,
so the main `Dockerfile`'s lint stage now invokes `golangci-lint` directly
instead of `make lint`, and its build stage runs `make test` and
`make fmt-check` instead of the `make check` aggregate (`make`, not the
scripts bare, because the Makefile's `export CGO_ENABLED = 0` only reaches
what it invokes). `COPY --from=lint` `/usr/bin/golangci-lint` is replaced by
`COPY --from=lint /src/go.sum /dev/null`: the copied binary was the only edge
forcing BuildKit to finish linting before the build stage starts, and dropping
it without replacing the edge would have ended fail-fast linting silently
under a still-green build. That is canonical `REPO_POLICIES.md:107`'s ordering
edge, restored. `ENV PATH=/home/builder/go/bin:$PATH` is gone with the
`go install` that justified it. `script/verify-linter-pin` is retired, deleted
along with its README entry, because both of its subjects ceased to exist in
the same change: it compared a linter binary against `GOLANGCI_LINT_VERSION`
in `script/bootstrap`, and there is now neither a binary crossing between
stages nor a version pin in bootstrap. The drift it guarded has not gone away,
it has moved — the linter is still pinned twice, now as the `FROM` line of
`Dockerfile.lint` and the `FROM` line of the `Dockerfile` lint stage, with
nothing syncing them, which is exactly what
https://git.eeqj.de/sneak/sfdupes/issues/42 made a build failure. Its
replacement is one new `script/verify-lint-image-pin`, run as a gate in both
files, which compares the two references to each other and deliberately
restates neither: a hardcoded expected digest would be a third copy and the
same drift one file further out. `golangci-lint config verify` is included per
the ruling, and the concern about its unpinned live HTTPS schema fetch was
measured rather than assumed — under `--network none` the pinned binary both
passes a valid config and rejects an invalid one with the jsonschema error, so
it validates from an embedded schema and makes no network call of its own. The
README scopes that to the gate steps rather than to linting as a whole:
`Dockerfile.lint` runs `go mod download` above them, so a cold cache still
needs the network and only a warm one lints offline. Verified: `make lint`
green with every `PATH` directory containing a `golangci-lint` removed
(`/home/user/go/bin`, `/home/user/.local/bin`, `/usr/local/bin`;
`command -v golangci-lint` empty); two consecutive `script/lint` runs on an
untouched tree both executed the linter, 27.7s and 28.7s in the lint step
under distinct epochs with the `COPY . .` layer `CACHED` above them, at 42.2s
and 41.8s wall clock — the no-cache rule was not weakened to shorten that.
Negative control: a planted `var unusedIssue46Sentinel = 1` failed
`script/lint` with
`report.go:173:5: var unusedIssue46Sentinel is unused (unused)`, and failed
`make docker` at `[lint 9/9]` with the build stage stopped at `[builder 3/12]`
— `COPY --from=lint`, `script/bootstrap`, the test gate and `make build` all
zero occurrences — then reverted clean. The drift guard fails on a tag-only
disagreement, on a digest-only disagreement, and on an unreadable reference,
naming both sides. `make docker` green in 5m35s with all six gates executing
under one epoch (lint 37.6s, test 25.2s reporting
`ok sneak.berlin/go/sfdupes 1.938s coverage: 88.5%`, not `(cached)`). The
non-root quirk still holds: in the builder image with the Go test cache off,
`--user 0:0` fails `TestScanHardlinkRunFailsTogether` (exit 1) where the
unprivileged user passes (exit 0). Noted for follow-up, not fixed here:
`golangci-lint` warns that the `gomodguard` linter is deprecated since v2.12.0
in favour of `gomodguard_v2`.
runs `golangci-lint config verify` and `golangci-lint run` as build steps;
`script/lint` builds it. `script/bootstrap` no longer installs or pins the
linter, and warns rather than fails when `docker` is absent;
`ENV PATH=/home/builder/go/bin:$PATH` went with its `go install`.
`script/verify-linter-pin` is retired; `script/verify-lint-image-pin`, a gate
in both files, compares their two `FROM` lines and restates neither pin.
Traps: nothing inside an image build may shell out to docker, so the
`Dockerfile` lint stage calls `golangci-lint` directly and the build stage
runs `make test` and `make fmt-check` instead of `make check`, through `make`
because the Makefile's `export CGO_ENABLED = 0` only reaches what it invokes.
`COPY --from=lint /src/go.sum /dev/null` replaces the copied linter binary as
the only edge making the build stage wait for lint; dropping it would end
fail-fast linting under a still-green build. `golangci-lint config verify`,
included per the ruling, validates from an embedded schema with no network
call, but `go mod download` above the gates still needs the network on a cold
cache. Verified: `make lint` green with no `golangci-lint` on `PATH`; two
back-to-back `script/lint` runs on an untouched tree both ran the linter
(27.7s and 28.7s in the lint step, `COPY . .` `CACHED` above); a planted
unused variable failed `script/lint`, and failed `make docker` at `[lint 9/9]`
with the build stage stopped at `[builder 3/12]`; the drift guard fails on a
tag-only, a digest-only and an unreadable reference, naming both sides; under
`--network none` config verify passes a valid config and rejects an invalid
one; `make docker` green in 5m35s with all six gates run (lint 37.6s, test
25.2s reporting `ok sneak.berlin/go/sfdupes 1.938s coverage: 88.5%`, not
`(cached)`); in the builder image with the Go test cache off, `--user 0:0`
still fails `TestScanHardlinkRunFailsTogether` where the unprivileged user
passes. Noted for follow-up, not fixed here: `golangci-lint` warns that
`gomodguard` is deprecated since v2.12.0 in favour of `gomodguard_v2`.
- install the Docker build stage's prerequisites by running `script/bootstrap`
instead of `apk add --no-cache make` inline (2026-08-09, branch
`dockerfile-bootstrap`, closes #42): canonical `REPO_POLICIES.md:97` requires
it, and the inline install left the build stage maintaining its own notion of
the toolchain — exactly the divergence #24 exists to close, one layer down.
The stage now copies `script/` plus `go.mod`/`go.sum` and runs
`script/bootstrap`, which ends in `go mod download`, so the separate
invocation of that is gone. `COPY --from=lint /usr/bin/golangci-lint` stays,
and moves above the bootstrap layer. It is the only edge making this stage
depend on the lint stage, so deleting it as redundant would end fail-fast
linting silently. Letting bootstrap install its own linter here would have
reintroduced the second toolchain and paid for a from-source build of it. What
makes the two stages provably one toolchain rather than two that happen to
agree is a new `script/verify-linter-pin`, run in the build stage on the
binary that arrives from the lint stage, before bootstrap: it fails the build
naming both versions unless that binary is the version `script/bootstrap`
pins. Bootstrap's own check could not serve that purpose — it reinstalls its
pin from source and then verifies whatever `PATH` resolves, so drift
self-heals silently and a lint stage image bumped on its own would lint at the
new version while `make check` ran at the old one, green. The linter version
is pinned in two independent places (the lint stage image digest and
`GOLANGCI_LINT_VERSION`) and nothing else keeps them in sync, so a
half-applied bump is now a build failure. The pin is read out of
`script/bootstrap`, which stays the single source of truth; a pin that cannot
be read is a hard failure, not a skip. The check needs no `CHECK_EPOCH`: its
only inputs are the copied binary and `script/`, so Docker invalidates the
layer exactly when a cached result would stop being true, and it is documented
with the other entrypoints in the README. `$GOPATH/bin` joins `PATH` because
that is where bootstrap's `go install` lands and bootstrap verifies its
installs against what `PATH` resolves — nothing in the image is shadowed by
it, the directory does not exist until bootstrap runs. Everything added sits
above `ARG CHECK_EPOCH`, and the `chown` and `USER builder` still precede
`make check`. Verified: the guard fails the build with both versions named
when the lint stage's linter is faked to a different version, and an
unmodified build still passes it; bootstrap runs clean under Alpine's `sh` and
its `apk` branch, installing `git` and `make` and finding the copied linter
already at the pin; a second build served the bootstrap and dependency layers
`CACHED` while both gates ran with a fresh epoch; a planted `unused` finding
failed the build at the lint gate in 48.9s with the build stage's `make check`
never starting; and the suite run in the image as `--user 0:0` fails
`TestScanHardlinkRunFailsTogether`, so the drop to the unprivileged user is
still load-bearing. That last check needs the Go test cache disabled — the
first attempt reported `ok ... (cached)` as root, reusing the result the
build-time run had left in the shared cache, which would have read as a pass.
Build wall time, on a shared host running many concurrent builds and so noisy:
2m13s on an unchanged tree, 2m17s and 4m29s for two builds after a source
change, 5m14s cold. Only the cold one breaches the policy ceiling, and not
because of this change — `chown -R builder:builder /src /home/builder` walks
the module cache and re-runs on every source change, and it alone varied
between 77s and 210s across those four builds, which is also the whole spread
in the totals. The same cold measurement against `main` is 5m03s with a 209s
`chown`. Filed as #43
`dockerfile-bootstrap`, closes https://git.eeqj.de/sneak/sfdupes/issues/42):
the stage copies `script/` plus `go.mod`/`go.sum` and runs `script/bootstrap`,
which ends in `go mod download`, so the separate call to it is gone.
`COPY --from=lint /usr/bin/golangci-lint` stays and moves above the bootstrap
layer: it is the only edge making this stage depend on the lint stage, so
deleting it would end fail-fast linting silently. A new
`script/verify-linter-pin`, run in the build stage before bootstrap, fails the
build naming both versions unless that copied binary is the version
`script/bootstrap` pins; a pin it cannot read is a hard failure, not a skip.
`$GOPATH/bin` joins `PATH`, where bootstrap's `go install` lands. The `chown`
and `USER builder` still precede `make check`. Verified: the guard fails the
build with both versions named when the lint stage's linter is faked to
another version, and passes an unmodified build; bootstrap runs clean under
Alpine's `sh` and `apk`, finding the copied linter already at the pin; a
second build served the bootstrap and dependency layers `CACHED` while both
gates ran; a planted `unused` finding failed the build at the lint gate in
48.9s with the build stage's `make check` never starting; and the suite run in
the image as `--user 0:0` fails `TestScanHardlinkRunFailsTogether`, so the
drop to the unprivileged user is still needed. That last check needs the Go
test cache off: as root it first reported `ok ... (cached)`, reusing the
build-time result. Build times on a noisy shared host: 2m13s on an unchanged
tree, 2m17s and 4m29s after a source change, 5m14s cold, which breaches the
policy ceiling; `chown -R builder:builder /src /home/builder` walks the module
cache and alone varied from 77s to 210s across those builds, and `main`
measured 5m03s cold with a 209s `chown`. Filed as
https://git.eeqj.de/sneak/sfdupes/issues/43
- bust the Docker layer cache for the gate steps, so `script/cibuild` and
`script/docker` cannot report a green they did not earn (2026-08-09, branch
`cibuild-cache-bust`, closes #32): both scripts were bare `docker build`
invocations with no cache control, and the `Dockerfile` copies the tree before
running its gates, so on an unchanged tree Docker served those layers from
cache and the build exited 0 having executed nothing. That is not hypothetical
here — every merge this repo has done is a non-fast-forward merge of an
undiverged branch, so each merge commit's tree is byte-identical to the branch
head's and each merge CI run was almost certainly a full cache hit; and PR
#31's reviewer found `make docker` returning success as a 17-layer cache hit,
catching it only by being suspicious. The fix is `ARG CHECK_EPOCH` with the
scripts passing `--build-arg CHECK_EPOCH="$(date +%s)"`. Two details make or
break it. `ARG` is scoped per stage and this `Dockerfile` has three gates
across two — `make fmt-check` and `make lint` in the lint stage, `make check`
in the build stage — so a single declaration would have left one stage
silently cacheable; it is declared in both. And BuildKit hashes the expanded
command, not the declaration, so a declared-but-unreferenced `ARG` invalidates
nothing: each gate `RUN` echoes the epoch, which also puts the value in the
build log as evidence the layer really ran. Placement is below the dependency
layers on purpose — a build that goes cold every time would be a different
bug, not a fix. Verified by running each script twice back to back on an
unchanged tree under `BUILDKIT_PROGRESS=plain`: all three gates executed on
all four runs, each with a fresh epoch in the log (`script/cibuild` 78.8s then
61.1s; `script/docker` 61.1s then 53.4s), and twelve steps were still served
`CACHED` in the steady state — both `go mod download`s, `apk add`, `adduser`,
the `chown`, every `go.mod`/`go.sum` and source copy, the linter copy out of
the lint stage, and the binary copy into the runtime stage. The lint stage
still gates the build stage: with a deliberate `unused` finding planted in the
tree, the build failed at `make lint` in 36.1s and the build-stage
`make check` never started. The build stage also still drops to the
unprivileged `builder` user before `make check`, which the suite depends on
rather than merely prefers: forcing the same image to run the tests as root
fails `TestScanHardlinkRunFailsTogether`, because root reads straight through
the `chmod(0)` the test uses to prove hard links are read once. This is the
local fix only; propagating it to the canonical templates is `prompts` #26
`cibuild-cache-bust`, closes https://git.eeqj.de/sneak/sfdupes/issues/32): the
`Dockerfile` copies the tree before its gates, so on an unchanged tree Docker
served them from cache and the build exited 0 having run nothing. Every
`docker build` in `script/` now passes `--no-cache` instead
(https://git.eeqj.de/sneak/sfdupes/issues/95). Run as root, the tests fail
`TestScanHardlinkRunFailsTogether`, because root reads through the `chmod(0)`
the test relies on, so they run as an unprivileged user
- check the installed golangci-lint version in `script/bootstrap` instead of
only its presence (2026-08-09, branch `bootstrap-version-check`, closes #24):
`missing golangci-lint` meant any linter already on `PATH` satisfied the
check, so the pin was never consulted and the v2.12.2 bump from #3 was inert
on every host that already had one — this host ran v2.10.1 against a v2.12.2
pin, `make check` went green, and `make docker` then rejected the same commit
with findings the local gate never saw. The version now lives in one place,
`GOLANGCI_LINT_VERSION`, with the `go install` module ref derived from it so a
bump cannot half-apply; a `golangci_lint_version` helper parses
`golangci-lint --version` (taking the field after the word `version` and
tolerating an optional leading `v`, which the module ref carries and the
binary's output does not), and any version that is not the pin — older, newer,
absent or unparseable — is reinstalled. The install is then verified against
the binary `PATH` actually resolves: `go install` writes into `GOBIN` (or
`GOPATH/bin`) while `make lint` runs whichever `golangci-lint` comes first on
`PATH`, so a wrong-version one sitting ahead of it — nix, apt, brew, apk, or
the `/usr/local/bin` copy the `Dockerfile` builder stage makes — would swallow
the install and leave the local gate disagreeing with CI under an affirmative
`bootstrap complete`. Bootstrap now re-reads the effective version after
installing and, on a mismatch, prints both paths and both versions to stderr
and exits non-zero instead of claiming success; it does not reorder anyone's
`PATH` or delete their binary. The `--version` call keeps its stderr
connected, so a present-but-broken binary says why rather than reinstalling
forever in silence, and is bounded by `timeout(1)` where that exists, so a
wedged binary cannot hang bootstrap. `git`, `make` and `go` keep their
presence-only checks and now say why in a comment: they are host
package-manager tools the repo deliberately does not pin, with `go.mod`
governing the language version and the digest-pinned images covering
reproducible builds. Verified on this host by bootstrapping from v2.10.1 to
v2.12.2 and running it again to a no-op, plus stub runs of the real script
under `dash` covering a thirteen-input parse matrix (absent, older, newer,
host-style, image-style, leading-`v`, stderr-only, empty, non-zero exit,
impostor binary, `(devel)`, trailing `version`), a shadowed install that must
exit non-zero, an install destination not on `PATH` at all, `GOBIN` set, and a
wedged binary that must hit the timeout; `make check` and `make lint` are
clean at v2.12.2, so v2.10.1 was not hiding any findings on `main`
only its presence (2026-08-09, branch `bootstrap-version-check`, closes
https://git.eeqj.de/sneak/sfdupes/issues/24): the version lives only in
`GOLANGCI_LINT_VERSION`, with the `go install` module ref derived from it, and
any installed version that is not the pin — older, newer, absent or
unparseable — is reinstalled. `go install` writes into `GOBIN` (or
`GOPATH/bin`) while `make lint` runs the first `golangci-lint` on `PATH`, so
bootstrap re-reads the effective version after installing and, on a mismatch,
prints both paths and both versions and exits non-zero; it does not reorder
`PATH` or delete anyone's binary. The `--version` call keeps its stderr and is
bounded by `timeout(1)` where that exists. `git`, `make` and `go` keep
presence-only checks. Verified by bootstrapping this host from v2.10.1 to
v2.12.2 and again to a no-op, and by stub runs of the script under `dash`
covering a thirteen-input version-parse matrix, a shadowed install that must
exit non-zero, an install destination not on `PATH`, `GOBIN` set, and a wedged
binary that must hit the timeout; `make check` and `make lint` are clean at
v2.12.2, so v2.10.1 was not hiding any findings on `main`
- unwind the hash worker pool on the error path (2026-08-09, branch
`hash-pool-cleanup`, closes #6): `hashPhase` used to return the moment
`recordRun` failed and abandon the pool — the feeder parked forever on a full
`jobs` channel and every worker on a full `results` channel. That only stopped
being invisible when #4 landed and `runScan` began unwinding instead of
calling `os.Exit`. The pool is now an owned, context-aware `hashPool`: every
blocking send in the feeder and the workers selects on `ctx.Done()`, `jobs` is
closed on every path out, and `hashPhase` defers `pool.stop()`, which cancels
and then drains `results` until the last goroutine has exited — draining is
what frees a worker already parked on a send. `ctx` is threaded from
`cmd.Context()` through `runScan`, `syncScan`, both worker pools and the whole
database layer (it is the first parameter everywhere), so #5 can hand this
path a signal and needs to add nothing else. The walk pool never leaked,
because `walkPhase` always drains its events to close, but it has the same
unbounded-send shape and #5 will give it an early return, so it gets the same
treatment plus a `ctx.Err()` guard after the walk: a cancelled walk yields a
partial size census, and every file it never reached looks vanished to the
update phase. That phase's own `BeginTx` fails on the same cancelled context
before deleting anything, so the guard is defence in depth rather than the
only barrier — but it is the one that survives #5 deciding an interrupted scan
may commit what it has. Tests drive `run(scan)` against a database whose
insert trigger aborts, and assert both that the scan fails instead of hanging
and that `runtime.NumGoroutine()` polls back to its pre-scan baseline; a
second set cancels a scan part-way through the walk — deterministically, by
counting the scan's own consultations of `ctx.Done()` rather than racing a
timer — and asserts that it stops at the guard holding a partial census and a
still-populated record index, with every record intact. The remaining
cancellation branches of both pools are covered by direct tests of
`sendEvent`, the walk workers, `dispatchDirs`, `feedHashJobs`, `hashWorker`
and `hashPhase`
`hash-pool-cleanup`, closes https://git.eeqj.de/sneak/sfdupes/issues/6): the
pool is now an owned, context-aware `hashPool`: every blocking send in the
feeder and the workers selects on `ctx.Done()`, `jobs` is closed on every path
out, and `hashPhase` defers `pool.stop()`, which cancels and then drains
`results` until the last goroutine has exited — draining is what frees a
worker already parked on a send. `ctx` is threaded from `cmd.Context()`
through `runScan`, `syncScan`, both worker pools and the whole database layer,
as the first parameter everywhere. The walk pool gets the same treatment plus
a `ctx.Err()` guard after the walk: a cancelled walk yields a partial size
census, and every file it never reached looks vanished to the update phase.
That phase's own `BeginTx` also fails on the cancelled context before deleting
anything, but the guard is the barrier that still holds once an interrupted
scan may commit what it has. Tests drive `run(scan)` against a database whose
insert trigger aborts and assert that the scan fails instead of hanging and
that `runtime.NumGoroutine()` polls back to its pre-scan baseline; others
cancel a scan part-way through the walk, deterministically, by counting its
own consultations of `ctx.Done()`, and assert that it stops at the guard
holding a partial census and a still-populated record index, with every record
intact. Direct tests of `sendEvent`, the walk workers, `dispatchDirs`,
`feedHashJobs`, `hashWorker` and `hashPhase` cover the remaining cancellation
branches of both pools
- guarantee the database is closed on every fatal exit path (2026-08-09, branch
`db-close-on-fatal`, closes #4): `fatalf` and its `os.Exit(1)` are gone, so
the deferred `db.Close()` — and with it the SQLite WAL checkpoint — now
actually runs when a subcommand fails; `runScan`, `runReport`, `runTrees`,
`loadRecords` and `resolveRoots` return errors instead. The single exit point
is `run` in `main.go`: it maps a `fatalError` (anything a subcommand returned)
to exit 1 and cobra's own argument and flag errors to exit 2, which keeps a
runtime failure from being reported as a usage error or printing the usage
text. New `main_test.go` drives the CLI in-process and asserts the exit codes
from README §Error handling plus the stdout/stderr split, including that a
fatal error raised after the database is open leaves no `-wal`/`-shm` sidecar
behind for `scan`, `report` or `trees`
`db-close-on-fatal`, closes https://git.eeqj.de/sneak/sfdupes/issues/4):
`fatalf` and its `os.Exit(1)` are gone, so the deferred `db.Close()` — and
with it the SQLite WAL checkpoint — now actually runs when a subcommand fails;
`runScan`, `runReport`, `runTrees`, `loadRecords` and `resolveRoots` return
errors instead. The single exit point is `run` in `main.go`: it maps a
`fatalError` (anything a subcommand returned) to exit 1 and cobra's own
argument and flag errors to exit 2, which keeps a runtime failure from being
reported as a usage error or printing the usage text. New `main_test.go`
drives the CLI in-process and asserts the exit codes from README §Error
handling plus the stdout/stderr split, including that a fatal error raised
after the database is open leaves no `-wal`/`-shm` sidecar behind for `scan`,
`report` or `trees`
- update golangci-lint to v2.12.2 with the canonical config (2026-08-09, branch
`golangci-v2.12.2`, merged as `38a01bd`, closes #3): bumped the pinned linter
in the `Dockerfile` lint stage and `script/bootstrap` from v2.12.1 to v2.12.2,
and replaced `.golangci.yml` with the canonical file — the linter settings
(`lll`, `funlen`, `cyclop`, `dupl` thresholds) now live under
`linters.settings` per the v2 schema, so they are actually applied; no new
lint findings surfaced
`golangci-v2.12.2`, merged as `38a01bd`, closes
https://git.eeqj.de/sneak/sfdupes/issues/3): bumped the pinned linter in the
`Dockerfile` lint stage and `script/bootstrap` from v2.12.1 to v2.12.2, and
replaced `.golangci.yml` with the canonical file — the linter settings (`lll`,
`funlen`, `cyclop`, `dupl` thresholds) now live under `linters.settings` per
the v2 schema, so they are actually applied; no new lint findings surfaced
- convert Makefile targets to scripts-to-rule-them-all `script/` entrypoints
like the other managed repos (2026-07-26, commit `3abeacf`, closes #1): all 12
`script/` entrypoints exist (`bootstrap`, `setup`, `projectname`, `test`,
`lint`, `fmt`, `fmt-check`, `check`, `docker`, `cibuild`, `precommit`,
`install-precommit`) and every Makefile target is now a thin shim over them,
matching the other managed repos
like the other managed repos (2026-07-26, commit `3abeacf`, closes
https://git.eeqj.de/sneak/sfdupes/issues/1): all 12 `script/` entrypoints
exist (`bootstrap`, `setup`, `projectname`, `test`, `lint`, `fmt`,
`fmt-check`, `check`, `docker`, `cibuild`, `precommit`, `install-precommit`)
and every Makefile target is now a thin shim over them
- make the binary the default Make target (2026-07-24, branch
`make-default-target`): plain `make` now builds `sfdupes` (previously it ran
`check` plus `build`); `make build` remains as an alias
@@ -309,8 +335,7 @@
- announce each operand on stderr before its passes (2026-07-24, branch
`scan-operand-progress`): with per-operand walk/hash/update cycles, a
multi-operand run (e.g. `scan /srv/*`) showed pass totals that looked like the
whole run's — an operator watching operand 3 of 14 hash 300k files concluded
20M files were being skipped
whole run's
- parallel walk (2026-07-24, branch `parallel-walk`): the walk pass was a single
goroutine and took hours at ~20M files on a busy pool (observed: 22M files in
4h on a ZFS server); it is now a per-directory worker-pool traversal that
@@ -382,5 +407,3 @@ Accepted divergences (no action):
- flat single-package layout with `.go` files in the repo root — fine for a
small single-binary tool per the Go styleguide; the tracker audit agrees
- `go test` runs without `-race` — the repo mandates `CGO_ENABLED=0` (pure-Go
builds) and the race detector requires cgo
+392 -48
View File
@@ -4,20 +4,35 @@ import (
"context"
"database/sql"
"errors"
"fmt"
"os"
"os/signal"
"path/filepath"
"slices"
"strconv"
"strings"
"sync"
"sync/atomic"
"syscall"
"testing"
"time"
)
// poolUnwind bounds how long a goroutine is given to leave a pool
// after its context is cancelled. Only a failing run ever waits this
// long: a pool that ignored its cancellation parks forever, and this
// is what turns that into a failed assertion instead of a suite that
// hangs until the test binary's own timeout.
// This file gathers the tests for scan cancellation and worker-pool
// unwinding. Everything it exercises lives in scan.go, so by the repo's
// convention of one test file per source file it would belong in
// scan_test.go. It is kept separate on purpose: cancellation behaviour
// cuts across both the walk pool and the hash pool as a single concern,
// and scan_test.go is already over 1,600 lines. That is the deliberate
// exception the convention otherwise expects to be stated.
// poolUnwind bounds how long a test waits for a cancellation to take
// effect: for a goroutine to return or a channel to close once its
// context is cancelled, or for a signal to cancel the scan's context.
// Only a failing run waits this long, and the bound is what makes that
// failure an assertion instead of a hang. A call made without it, as
// most of this file's scans are, has no bound: a regression that parks
// it is caught only as the test binary's own timeout.
const poolUnwind = 2 * time.Second
// walkClock is a context whose cancellation is driven by the scan's
@@ -28,9 +43,14 @@ const poolUnwind = 2 * time.Second
//
// The accounting behind the n chosen by each test: every blocking
// channel operation in the walk selects on Done, so the walk spends
// one consultation per file event plus a couple per directory, while
// the index load that runs ahead of it spends a small fixed number
// (three) whatever the record count.
// one consultation per file event plus a couple per directory. The
// index load that runs ahead of it also consults Done, but a bounded
// number of times that does not grow with the record count. The tests
// depend on that property, not on the bound's exact value: each test
// sets n from the consultations of the walk, plus those of the hash
// phase when it cancels mid-hash, far from both ends of the phase it
// interrupts, so the cancellation lands inside that phase whatever the
// record count.
type walkClock struct {
n int64
seen atomic.Int64
@@ -90,7 +110,10 @@ const (
walkCancelFilesPerDir = 20
walkCancelFiles = walkCancelDirs * walkCancelFilesPerDir
walkCancelWorkers = 4
walkCancelInFlightDirs = walkCancelWorkers * walkCancelFilesPerDir
// The most files the walkCancelWorkers directories already in
// flight when the scan is cancelled can still emit, at
// walkCancelFilesPerDir each. A file count, not a directory count.
walkCancelInFlightFiles = walkCancelWorkers * walkCancelFilesPerDir
)
// walkCancelAtDone is the consultation on which the fixture's context
@@ -150,9 +173,13 @@ func assertRecordsIntact(t *testing.T, db *sql.DB, before []string) {
// Every one of those records would look vanished to the update phase.
// The guard is what stops the scan there, and this test is what
// notices if it stops doing so: deleting the guard, or making it
// unreachable, makes the scan carry its truncated view into a later
// phase and fail there instead, with a wrapped error rather than the
// bare cancellation.
// unreachable, makes the scan carry its truncated view into the update
// phase, which counts every record the walk never reached for removal.
//
// The syncScan call here is not bounded by poolUnwind: a regression
// that left a worker pool parked would hang it, and that regression is
// caught only by the test binary's own timeout, not by a quick
// assertion.
//
//nolint:paralleltest // counts goroutines: must not run beside others
func TestSyncScanCancelledMidWalkKeepsRecords(t *testing.T) {
@@ -182,10 +209,10 @@ func TestSyncScanCancelledMidWalkKeepsRecords(t *testing.T) {
// assertWalkGuardAborted checks that the scan stopped at the post-walk
// guard: with a census that is neither empty (the walk really ran)
// nor complete (it really was cut short), and with the guard's own
// bare cancellation as the error. A wrapped error means the partial
// census was carried past the guard into the hash or update phase,
// which is the failure this test exists to catch.
// nor complete (it really was cut short), and with no record counted
// for removal. A removal count means the partial census was carried
// past the guard into the update phase, which is the failure this test
// exists to catch.
func assertWalkGuardAborted(t *testing.T, st scanStats, err error) {
t.Helper()
@@ -194,12 +221,6 @@ func assertWalkGuardAborted(t *testing.T, st scanStats, err error) {
err, context.Canceled)
}
if errors.Unwrap(err) != nil {
t.Errorf("syncScan reported %q, want the guard's bare "+
"cancellation: a wrapped error means the truncated census "+
"reached a later phase", err)
}
if st.unchanged == 0 {
t.Fatalf("stats = %+v: the census is empty, so the walk never "+
"ran and the guard was reached for the wrong reason", st)
@@ -211,10 +232,11 @@ func assertWalkGuardAborted(t *testing.T, st scanStats, err error) {
}
// The workers drop every directory still queued once the scan is
// cancelled, so only the directories already in flight can add to
// the census after the fact. A census beyond that bound would mean
// the cancellation was not observed where it should have been.
limit := walkCancelAtDone + walkCancelInFlightDirs
// cancelled, so only the files in the directories already in flight
// can add to the census after the fact. A census beyond that bound
// would mean the cancellation was not observed where it should have
// been.
limit := walkCancelAtDone + walkCancelInFlightFiles
if st.unchanged > limit {
t.Errorf("census covers %d files, want at most %d: the walk kept "+
"taking directories off the queue after cancellation",
@@ -260,6 +282,252 @@ func TestSyncScanCancelledBeforeLoadIndex(t *testing.T) {
assertRecordsIntact(t, db, before)
}
// hashCancelAtDone is the consultation on which the mid-hash test's
// context cancels itself. The walk of buildWalkCancelTree spends about
// one per file and three per directory, and the hash phase then one per
// file hashed, so this lands about half way through the hash phase.
const hashCancelAtDone = walkCancelFiles + 3*walkCancelDirs +
walkCancelFiles/2
// TestSyncScanCancelledMidHashKeepsHashedRecords cancels a first scan
// part-way through its hash phase. The fixture holds fewer files than a
// batch, so every file hashed is still waiting to be committed: the scan
// must commit them all before it returns, and the next scan must hash
// only the rest.
func TestSyncScanCancelledMidHashKeepsHashedRecords(t *testing.T) {
t.Parallel()
dir := buildWalkCancelTree(t)
db := openTestDB(t)
st, err := syncScan(newWalkClock(hashCancelAtDone), db,
[]string{dir}, walkCancelWorkers, false)
if !errors.Is(err, context.Canceled) {
t.Fatalf("syncScan cancelled mid-hash = %v, want %v",
err, context.Canceled)
}
if st.walked != walkCancelFiles || st.added == 0 ||
st.added >= walkCancelFiles {
t.Fatalf("stats = %+v: want the walk complete and the hash phase "+
"cut short", st)
}
if got := len(dbRecords(t, db)); got != st.added {
t.Errorf("%d records after the cancelled scan, want the %d it hashed",
got, st.added)
}
hashed := st.added
st = syncTree(t, db, dir)
if st.added != walkCancelFiles-hashed || st.unchanged != hashed {
t.Errorf("next scan stats = %+v, want %d added %d unchanged",
st, walkCancelFiles-hashed, hashed)
}
}
// storedPaths opens the database at path as report does, which fails
// unless it is a valid database, and returns its records' paths.
func storedPaths(t *testing.T, path string) []string {
t.Helper()
db, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
defer func() { _ = db.Close() }()
return recordPaths(dbRecords(t, db))
}
// TestRunScanInterrupted calls the scan entrypoint with a context that
// is already cancelled, as when a signal arrives at once. It must return
// errInterrupted promptly with its one line on stderr and nothing on
// stdout, leave the database valid and as it was, and leave nothing in
// the way of the next scan, which must bring the database up to date.
func TestRunScanInterrupted(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
stdout := captureStdout(t)
stderr := captureStderr(t)
dir := buildSmokeTree(t)
err := runScan(t.Context(), []string{dir}, walkCancelWorkers, false)
if err != nil {
t.Fatal(err)
}
before := storedPaths(t, path)
// A vanished file and a new one: the interrupted scan records
// neither.
gone := filepath.Join(dir, "a", "unique.bin")
err = os.Remove(gone)
if err != nil {
t.Fatal(err)
}
added := writeFile(t, dir, "a/new.bin", pattern(50, 10))
shown := len(stderr())
done := make(chan struct{})
go func() {
defer close(done)
err = runScan(cancelledContext(t), []string{dir}, walkCancelWorkers,
false)
}()
awaitReturn(t, done, "runScan")
if !errors.Is(err, errInterrupted) {
t.Fatalf("runScan on a cancelled context = %v, want %v",
err, errInterrupted)
}
want := "scan: interrupted after 0 files\n"
if got := stderr()[shown:]; got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
if got := stdout(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
assertNoSidecars(t, path)
if got := storedPaths(t, path); !slices.Equal(got, before) {
t.Errorf("records = %q after the interrupted scan, want %q",
got, before)
}
err = runScan(t.Context(), []string{dir}, walkCancelWorkers, false)
if err != nil {
t.Fatal(err)
}
got := storedPaths(t, path)
if slices.Contains(got, gone) || !slices.Contains(got, added) {
t.Errorf("records = %q after the next scan, want %q gone and %q "+
"added", got, gone, added)
}
}
// TestRunScanInterruptedMidHash interrupts the scan entrypoint part-way
// through its hash phase, after the database is open. It must return
// errInterrupted, release the lock, end stderr with its line counting
// every file the walk reached, write nothing to stdout, close the
// database out of WAL mode, and keep the records it hashed.
func TestRunScanInterruptedMidHash(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
stdout := captureStdout(t)
stderr := captureStderr(t)
dir := buildWalkCancelTree(t)
err := runScan(newWalkClock(hashCancelAtDone), []string{dir},
walkCancelWorkers, false)
if !errors.Is(err, errInterrupted) {
t.Fatalf("runScan interrupted mid-hash = %v, want %v",
err, errInterrupted)
}
holdScanLock(t, path)
want := fmt.Sprintf("scan: interrupted after %d files\n", walkCancelFiles)
if got := stderr(); !strings.HasSuffix(got, want) {
t.Errorf("stderr = %q, want it to end with %q", got, want)
}
if got := stdout(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
assertNoSidecars(t, path)
db, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
defer func() { _ = db.Close() }()
// A plain close also removes the sidecars, but leaves WAL mode on.
var mode string
err = db.QueryRowContext(t.Context(), "PRAGMA journal_mode").Scan(&mode)
if err != nil {
t.Fatal(err)
}
if mode != "delete" {
t.Errorf("journal mode = %q after the interrupted scan, want %q",
mode, "delete")
}
kept := len(dbRecords(t, db))
if kept == 0 || kept >= walkCancelFiles {
t.Errorf("%d records after the interrupted scan, want those it "+
"hashed: some but not all of the %d files", kept, walkCancelFiles)
}
}
// TestInterruptContextCatchesSIGTERM sends SIGTERM to the test process
// while the scan's handler is installed, and checks that it cancels the
// scan's context.
//
//nolint:paralleltest // signals the whole process: must not run beside a scan
func TestInterruptContextCatchesSIGTERM(t *testing.T) {
// Caught here as well, so that a handler that misses SIGTERM fails
// this test instead of ending the test process.
caught := make(chan os.Signal, 1)
signal.Notify(caught, syscall.SIGTERM)
defer signal.Stop(caught)
ctx, stop := interruptContext(t.Context())
defer stop()
err := syscall.Kill(os.Getpid(), syscall.SIGTERM)
if err != nil {
t.Fatal(err)
}
select {
case <-ctx.Done():
case <-time.After(poolUnwind):
t.Fatal("SIGTERM did not cancel the scan's context")
}
}
// TestCommitFullBatchKeepsFailedBatch checks that a full batch whose
// commit fails, as it does once the scan is interrupted, stays in the
// batch, so that syncScan's final commit saves it.
func TestCommitFullBatchKeepsFailedBatch(t *testing.T) {
t.Parallel()
s := &scanState{db: openTestDB(t)}
for i := range updateBatchSize {
s.batch = append(s.batch, scanRec{path: "/f" + strconv.Itoa(i)})
}
err := s.commitFullBatch(cancelledContext(t))
if !errors.Is(err, context.Canceled) {
t.Fatalf("commitFullBatch on a cancelled context = %v, want %v",
err, context.Canceled)
}
if len(s.batch) != updateBatchSize {
t.Errorf("batch holds %d records after the failed commit, want %d",
len(s.batch), updateBatchSize)
}
}
// drainClosed counts the values received from ch until it closes,
// failing the test if it does not close within poolUnwind. A pool that
// ignored its cancellation leaves its channel open with its goroutines
@@ -328,20 +596,49 @@ func TestSendEventAbandonsBlockedSend(t *testing.T) {
awaitReturn(t, done, "sendEvent")
}
// TestWalkWorkersDropQueuedDirs checks that cancelled walk workers keep
// reading jobs and drop the directories rather than stopping their
// read: the range over jobs has to run out for the pool to tear down
// and close its event stream.
func TestWalkWorkersDropQueuedDirs(t *testing.T) {
// TestWalkOneDirStopsWhenCancelled checks that a cancelled scan stops
// reading a directory instead of going through the rest of its
// entries. A walk that kept going would return the subdirectory below
// to descend into. Unlike a file event, that return is not a send the
// cancellation can abandon, so the test catches the regression every
// time.
func TestWalkOneDirStopsWhenCancelled(t *testing.T) {
t.Parallel()
dir := t.TempDir()
writeEmptyFiles(t, dir, walkCancelFilesPerDir)
err := os.Mkdir(filepath.Join(dir, "sub"), 0o750)
if err != nil {
t.Fatal(err)
}
// Unbuffered and unread: on a cancelled scan every send gives up.
events := make(chan walkEvent)
subs := walkOneDir(cancelledContext(t), dirJob{path: dir}, false, events)
if len(subs) != 0 {
t.Errorf("cancelled walkOneDir returned %+v to descend into, "+
"want none", subs)
}
}
// TestWalkWorkersDropQueuedDirs checks that cancelled walk workers keep
// reading jobs and drop the directories rather than stopping their
// read: the range over jobs has to run out for the pool to tear down
// and close its event stream. The queued directory does not exist, so
// a worker that walked it anyway would send a warning before
// walkOneDir's own cancellation check could stop it. On a cancelled
// scan that send delivers or gives up at random, so with 64 jobs
// queued the regression has a one in 2^64 chance of passing.
func TestWalkWorkersDropQueuedDirs(t *testing.T) {
t.Parallel()
missing := filepath.Join(t.TempDir(), "missing")
jobs, _, events := startWalkWorkers(cancelledContext(t), 2, false)
for range 4 {
jobs <- dirJob{path: dir}
for range 64 {
jobs <- dirJob{path: missing}
}
close(jobs)
@@ -412,7 +709,11 @@ func TestDispatchDirsClosesJobsWhenCancelled(t *testing.T) {
// TestFeedHashJobsClosesJobsWhenCancelled checks that the hash feeder
// abandons the runs it has not queued yet and still closes the job
// channel, which is what lets the workers' range terminate.
// channel, which is what lets the workers' range terminate. The
// receive on jobs below is not bounded: a feeder that returned without
// closing jobs would leave that receive with no sender and no close, so
// this regression is caught by the test binary's timeout rather than by
// a bounded assertion.
func TestFeedHashJobsClosesJobsWhenCancelled(t *testing.T) {
t.Parallel()
@@ -438,41 +739,84 @@ func TestFeedHashJobsClosesJobsWhenCancelled(t *testing.T) {
// TestHashWorkerDropsQueuedRuns checks that a cancelled hash worker
// keeps reading jobs and drops the runs rather than reading files
// nobody wants the hashes of — while still letting the range run out
// so the pool tears down. The queued run names a file that does not
// exist, so a worker that hashed it anyway would produce a result.
// so the pool tears down. The hash function records that it was
// called, so a worker that hashed the queued run anyway is caught
// every time.
func TestHashWorkerDropsQueuedRuns(t *testing.T) {
t.Parallel()
done := make(chan struct{})
jobs := make(chan []fileRec, 1)
results := make(chan hashResult, 1)
results := make(chan hashResult)
run := []fileRec{{path: filepath.Join(t.TempDir(), "missing"), size: 1}}
jobs <- run
jobs <- []fileRec{{path: filepath.Join(t.TempDir(), "missing"), size: 1}}
close(jobs)
var hashed atomic.Bool
hash := func(path string, size int64) (string, string, string, error) {
hashed.Store(true)
return hashSignature(path, size)
}
go func() {
defer close(done)
hashWorker(cancelledContext(t), jobs, results)
hashWorker(cancelledContext(t), jobs, results, hash)
}()
awaitReturn(t, done, "hashWorker")
select {
case r := <-results:
t.Errorf("cancelled hash worker produced %+v, want the run dropped",
r)
default:
if hashed.Load() {
t.Error("cancelled hash worker hashed the queued run, want it dropped")
}
}
// TestHashWorkerAbandonsBlockedSend checks that a hash worker with a
// result to deliver and nobody to deliver it to leaves once the scan
// is cancelled, instead of holding the pool open. The scan tests do
// not catch this: stop drains results, which frees a parked worker
// anyway.
func TestHashWorkerAbandonsBlockedSend(t *testing.T) {
t.Parallel()
ctx, cancel := context.WithCancel(t.Context())
defer cancel()
done := make(chan struct{})
jobs := make(chan []fileRec, 1)
// Unbuffered and unread, with jobs left open: the worker's only way
// out is the cancellation case beside its send.
results := make(chan hashResult)
jobs <- []fileRec{{path: filepath.Join(t.TempDir(), "missing"), size: 1}}
// The scan is cancelled while the worker hashes, so the worker has
// already passed the check that drops queued runs.
hash := func(path string, size int64) (string, string, string, error) {
cancel()
return hashSignature(path, size)
}
go func() {
defer close(done)
hashWorker(ctx, jobs, results, hash)
}()
awaitReturn(t, done, "hashWorker")
}
// TestHashPhaseCancelledReturnsContextError checks the result loop's
// own exit: with the pool cancelled, no result will ever arrive, and
// the loop must leave through the cancellation rather than wait for a
// receive that cannot happen.
// receive that cannot happen. This call is not bounded by poolUnwind: a
// loop that dropped its cancellation case would block on that receive,
// so the regression surfaces as the test binary's timeout rather than
// as a bounded assertion.
func TestHashPhaseCancelledReturnsContextError(t *testing.T) {
t.Parallel()
+356 -66
View File
@@ -6,11 +6,14 @@ import (
"errors"
"fmt"
"io/fs"
"net/url"
"os"
"path/filepath"
"slices"
"strconv"
"time"
"golang.org/x/sys/unix"
// The pure-Go SQLite driver, registered as "sqlite"; keeps cgo
// disabled.
_ "modernc.org/sqlite"
@@ -32,26 +35,43 @@ const schemaVersion = 1
// scan.
const dbDirPerm = 0o755
// lockFilePerm is the mode for the scan lock file. Anyone who can open
// the file can hold the lock and keep every scan from running, so it
// is open to its owner only.
const lockFilePerm = 0o600
// createTableSQL is the schema applied to a fresh database. Paths are
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8. mtime
// holds whole Unix seconds and mtime_nsec the nanoseconds within that
// second.
const createTableSQL = `
CREATE TABLE files (
path BLOB PRIMARY KEY,
size INTEGER NOT NULL,
mtime INTEGER NOT NULL,
mtime_nsec INTEGER NOT NULL,
head TEXT NOT NULL,
tail TEXT NOT NULL
tail TEXT NOT NULL,
content TEXT NOT NULL
) WITHOUT ROWID
`
// createIndexSQL indexes the records by signature, so report can have
// SQLite group them without sorting the whole table.
const createIndexSQL = `
CREATE INDEX files_signature ON files (size, head, tail, content)
`
// upsertSQL inserts one file record, replacing any existing record for
// the same path.
const upsertSQL = `
INSERT INTO files (path, size, mtime, head, tail)
VALUES (?, ?, ?, ?, ?)
INSERT INTO files (path, size, mtime, mtime_nsec, head, tail, content)
VALUES (?, ?, ?, ?, ?, ?, ?)
ON CONFLICT (path) DO UPDATE SET
size = excluded.size, mtime = excluded.mtime,
head = excluded.head, tail = excluded.tail
mtime_nsec = excluded.mtime_nsec,
head = excluded.head, tail = excluded.tail,
content = excluded.content
`
// errNoDatabase reports a missing database file for report/trees.
@@ -62,6 +82,10 @@ var errNoDatabase = errors.New(
// does not understand.
var errSchemaVersion = errors.New("unsupported database schema version")
// errScanRunning reports that another scan holds the lock on the
// database.
var errScanRunning = errors.New("another scan is running")
// databasePath resolves the database location: SFDUPES_DATABASE when
// set and non-empty, the compiled-in default otherwise.
func databasePath() string {
@@ -72,16 +96,36 @@ func databasePath() string {
return defaultDatabasePath
}
// openDB opens the SQLite database at path with WAL journaling and a
// busy timeout, so a report can run while a cron scan is in progress.
// It does not create or verify the schema.
func openDB(path string) (*sql.DB, error) {
dsn := "file:" + path +
"?_pragma=busy_timeout(10000)" +
// scanParams are the connection parameters for scan: read-write, with
// WAL journaling and a busy timeout, so a report can run while a cron
// scan is in progress. closeScanDatabase leaves WAL mode again.
const scanParams = "_pragma=busy_timeout(10000)" +
"&_pragma=journal_mode(WAL)" +
"&_pragma=synchronous(NORMAL)"
db, err := sql.Open("sqlite", dsn)
// reportParams are the connection parameters for report and trees:
// read-only, with the same busy timeout. They set no journal mode,
// because setting one is a write.
const reportParams = "mode=ro" +
"&_pragma=busy_timeout(10000)" +
"&_pragma=query_only(1)"
// openDB opens the SQLite database at path with the connection
// parameters params. It does not create or verify the schema.
func openDB(path, params string) (*sql.DB, error) {
// The path is escaped into a file: URI, so ?, # and % in it stay
// part of the file name. SQLite reads what follows file:// up to
// the next / as a host name, so an absolute path goes after an
// empty host (file:///abs) and a relative path goes without one
// (file:rel).
uri := url.URL{
Scheme: "file",
OmitHost: !filepath.IsAbs(path),
Path: path,
RawQuery: params,
}
db, err := sql.Open("sqlite", uri.String())
if err != nil {
return nil, fmt.Errorf("open database %s: %w", path, err)
}
@@ -94,6 +138,43 @@ func openDB(path string) (*sql.DB, error) {
return db, nil
}
// lockScanDatabase takes the lock that keeps a second scan off the
// database at path: an exclusive flock(2) on the file beside it named
// path with ".lock" appended, created along with the database's parent
// directory if missing. A lock held by another scan fails at once
// instead of waiting. The lock lasts until the returned file is closed
// or the process ends. The file is never deleted: a scan that deleted
// it would let the next scan lock a new file while another still holds
// the old one.
func lockScanDatabase(path string) (*os.File, error) {
err := os.MkdirAll(filepath.Dir(path), dbDirPerm)
if err != nil {
return nil, fmt.Errorf("create database directory: %w", err)
}
lockPath := path + ".lock"
//nolint:gosec // the operator chooses the database path
f, err := os.OpenFile(lockPath, os.O_RDWR|os.O_CREATE, lockFilePerm)
if err != nil {
return nil, err
}
err = unix.Flock(int(f.Fd()), unix.LOCK_EX|unix.LOCK_NB)
if err != nil {
_ = f.Close()
if errors.Is(err, unix.EWOULDBLOCK) {
return nil, fmt.Errorf("%w (lock held on %s)",
errScanRunning, lockPath)
}
return nil, fmt.Errorf("lock %s: %w", lockPath, err)
}
return f, nil
}
// openScanDatabase opens the database for the scan subcommand, creating
// the file, its parent directory, and the schema as needed.
func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
@@ -102,7 +183,7 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return nil, fmt.Errorf("create database directory: %w", err)
}
db, err := openDB(path)
db, err := openDB(path, scanParams)
if err != nil {
return nil, err
}
@@ -117,6 +198,24 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return db, nil
}
// closeScanDatabase switches the database at path from WAL back to
// rollback-journal mode and closes it. Out of WAL mode the database
// file alone holds the whole database, so a reader needs no -wal or
// -shm file beside it, nor write access to create them. The switch
// fails while a report has the database open; the database then stays
// in WAL mode, still readable, until a later scan closes it.
func closeScanDatabase(ctx context.Context, db *sql.DB, path string) {
// Runs on the way out of a cancelled scan too.
_, err := db.ExecContext(context.WithoutCancel(ctx),
"PRAGMA journal_mode = DELETE")
if err != nil {
fmt.Fprintf(os.Stderr, "scan: database %s left in WAL mode: %v\n",
path, err)
}
_ = db.Close()
}
// openReportDatabase opens an existing database for the report and
// trees subcommands. A missing database file is an error directing the
// user to run scan first; the schema version must match exactly.
@@ -132,12 +231,18 @@ func openReportDatabase(ctx context.Context,
return nil, fmt.Errorf("database: %w", err)
}
db, err := openDB(path)
db, err := openDB(path, reportParams)
if err != nil {
return nil, err
}
v, err := userVersion(ctx, db)
if err == nil && v == 0 {
// An empty database passes this check and fails the version
// check below.
err = checkUnversioned(ctx, db)
}
if err != nil {
_ = db.Close()
@@ -164,6 +269,11 @@ func initSchema(ctx context.Context, db *sql.DB) error {
switch v {
case 0:
err = checkUnversioned(ctx, db)
if err != nil {
return err
}
return createSchema(ctx, db)
case schemaVersion:
return nil
@@ -173,20 +283,65 @@ func initSchema(ctx context.Context, db *sql.DB) error {
}
}
// checkUnversioned checks a database at user_version 0 before it is
// taken for an empty one. createSchema creates the files table and
// sets the version together, so a files table at version 0 was made by
// something else. Adopting it could corrupt unrelated data, so that is
// a schema-version error telling the operator to remove the file and
// rescan.
func checkUnversioned(ctx context.Context, db *sql.DB) error {
var name string
err := db.QueryRowContext(ctx,
"SELECT name FROM sqlite_master "+
"WHERE type = 'table' AND name = 'files'").Scan(&name)
switch {
case err == nil:
return fmt.Errorf(
"has a files table but no schema version; "+
"remove the file and rescan: %w", errSchemaVersion)
case errors.Is(err, sql.ErrNoRows):
return nil
default:
return fmt.Errorf("check for files table: %w", err)
}
}
// createSchema applies the schema to a fresh database and stamps the
// schema version.
// schema version in one transaction, so a creation stopped partway, by
// an interrupt or an error, leaves an empty database the next scan
// sets up, never a files table at version 0, which checkUnversioned
// refuses.
func createSchema(ctx context.Context, db *sql.DB) error {
_, err := db.ExecContext(ctx, createTableSQL)
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
_, err = db.ExecContext(ctx,
defer func() { _ = tx.Rollback() }()
_, err = tx.ExecContext(ctx, createTableSQL)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
_, err = tx.ExecContext(ctx, createIndexSQL)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
_, err = tx.ExecContext(ctx,
"PRAGMA user_version = "+strconv.Itoa(schemaVersion))
if err != nil {
return fmt.Errorf("set schema version: %w", err)
}
err = tx.Commit()
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
return nil
}
@@ -202,50 +357,12 @@ func userVersion(ctx context.Context, db *sql.DB) (int, error) {
return v, nil
}
// loadFileRows reads every record from the files table.
func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) {
// loadFileRows streams every record to fn in path order: byte order,
// which is the order of the primary key, so SQLite does not sort.
func loadFileRows(ctx context.Context, db *sql.DB, fn func(r scanRec)) error {
rows, err := db.QueryContext(ctx,
"SELECT path, size, mtime, head, tail FROM files")
if err != nil {
return nil, fmt.Errorf("read records: %w", err)
}
defer func() { _ = rows.Close() }()
var recs []scanRec
for rows.Next() {
var (
path []byte
r scanRec
)
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail)
if err != nil {
return nil, fmt.Errorf("read record: %w", err)
}
r.path = string(path)
recs = append(recs, r)
}
err = rows.Err()
if err != nil {
return nil, fmt.Errorf("read records: %w", err)
}
return recs, nil
}
// loadFileMeta streams every record's path, size, mtime, and whether
// it carries hashes to fn. Scan change detection needs no hash
// values, and skipping the hash columns keeps the scan's in-memory
// index small on multi-million-file databases.
func loadFileMeta(ctx context.Context, db *sql.DB,
fn func(path string, size, mtime int64, hashed bool),
) error {
rows, err := db.QueryContext(ctx,
"SELECT path, size, mtime, head <> '' FROM files")
"SELECT path, size, mtime, mtime_nsec, head, tail, content "+
"FROM files ORDER BY path")
if err != nil {
return fmt.Errorf("read records: %w", err)
}
@@ -255,16 +372,189 @@ func loadFileMeta(ctx context.Context, db *sql.DB,
for rows.Next() {
var (
path []byte
size, mtime int64
hashed int64
sec, nsec int64
r scanRec
)
err = rows.Scan(&path, &size, &mtime, &hashed)
err = rows.Scan(&path, &r.size, &sec, &nsec, &r.head, &r.tail,
&r.content)
if err != nil {
return fmt.Errorf("read record: %w", err)
}
fn(string(path), size, mtime, hashed != 0)
r.path = string(path)
r.mtime = time.Unix(sec, nsec)
fn(r)
}
err = rows.Err()
if err != nil {
return fmt.Errorf("read records: %w", err)
}
return nil
}
// dupeRowsSQL selects every record in a duplicate group, with the
// group's first path. A group is the records with a content hash that
// share a size, head, tail, and content, when there are two or more of
// them. The rows come in report order: groups by size descending, then
// by first path, and each group's paths ascending.
const dupeRowsSQL = `
SELECT g.first, f.path, f.size
FROM files AS f
JOIN (
SELECT size, head, tail, content, MIN(path) AS first
FROM files
WHERE content <> ''
GROUP BY size, head, tail, content
HAVING COUNT(*) > 1
) AS g USING (size, head, tail, content)
ORDER BY f.size DESC, g.first, f.path
`
// loadDupeRows streams the rows of dupeRowsSQL to fn and returns the
// number of records in the database. The count and the rows are read
// in one transaction, so they agree while a scan is committing. An
// error from fn stops the reading and is returned as it is.
func loadDupeRows(ctx context.Context, db *sql.DB,
fn func(first, path string, size int64) error,
) (int, error) {
// Everything goes through tx: the report connection is the only
// one, so a query on db would wait for tx forever.
tx, err := db.BeginTx(ctx, &sql.TxOptions{ReadOnly: true})
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
defer func() { _ = tx.Rollback() }()
var records int
err = tx.QueryRowContext(ctx, "SELECT COUNT(*) FROM files").Scan(&records)
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
rows, err := tx.QueryContext(ctx, dupeRowsSQL)
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
defer func() { _ = rows.Close() }()
for rows.Next() {
var (
first, path []byte
size int64
)
err = rows.Scan(&first, &path, &size)
if err != nil {
return 0, fmt.Errorf("read record: %w", err)
}
err = fn(string(first), string(path), size)
if err != nil {
return 0, err
}
}
err = rows.Err()
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
return records, nil
}
// loadFileMeta streams every record's path, size, mtime, and whether
// it carries hashes to fn. Scan change detection needs no hash
// values, and skipping the hash columns keeps the scan's in-memory
// index small on multi-million-file databases.
func loadFileMeta(ctx context.Context, db *sql.DB,
fn func(path string, size int64, mtime time.Time, hashed bool),
) error {
rows, err := db.QueryContext(ctx,
"SELECT path, size, mtime, mtime_nsec, head <> '' FROM files")
if err != nil {
return fmt.Errorf("read records: %w", err)
}
defer func() { _ = rows.Close() }()
for rows.Next() {
var (
path []byte
size, sec, nsec int64
hashed int64
)
err = rows.Scan(&path, &size, &sec, &nsec, &hashed)
if err != nil {
return fmt.Errorf("read record: %w", err)
}
fn(string(path), size, time.Unix(sec, nsec), hashed != 0)
}
err = rows.Err()
if err != nil {
return fmt.Errorf("read records: %w", err)
}
return nil
}
// contentCandidatesSQL selects every record of at least headTailMin
// bytes whose size, head, and tail equal another record's, in each
// group (the records sharing a size, head, and tail) where at least one
// record has no content hash, with whether each record has one. SQLite
// does the grouping, so no other record's hashes are loaded into
// memory; the rows come ordered by size, head, and tail, so each
// group's rows arrive together.
const contentCandidatesSQL = `
SELECT f.path, f.size, f.mtime, f.mtime_nsec, f.head, f.tail, f.content <> ''
FROM files AS f
JOIN (
SELECT size, head, tail
FROM files
WHERE size >= ? AND head <> ''
GROUP BY size, head, tail
HAVING COUNT(*) > 1 AND SUM(content = '') > 0
) AS g USING (size, head, tail)
ORDER BY size, head, tail
`
// loadContentCandidates streams the rows of contentCandidatesSQL to fn:
// each record, without its content hash, and whether it has one.
func loadContentCandidates(ctx context.Context, db *sql.DB,
fn func(r scanRec, hashed bool),
) error {
rows, err := db.QueryContext(ctx, contentCandidatesSQL, headTailMin)
if err != nil {
return fmt.Errorf("read records: %w", err)
}
defer func() { _ = rows.Close() }()
for rows.Next() {
var (
path []byte
sec, nsec int64
r scanRec
hashed int64
)
err = rows.Scan(&path, &r.size, &sec, &nsec, &r.head, &r.tail,
&hashed)
if err != nil {
return fmt.Errorf("read record: %w", err)
}
r.path = string(path)
r.mtime = time.Unix(sec, nsec)
fn(r, hashed != 0)
}
err = rows.Err()
@@ -347,8 +637,8 @@ func execUpserts(ctx context.Context, tx *sql.Tx, upserts []scanRec,
defer func() { _ = st.Close() }()
for _, r := range upserts {
_, err = st.ExecContext(ctx,
[]byte(r.path), r.size, r.mtime, r.head, r.tail)
_, err = st.ExecContext(ctx, []byte(r.path), r.size,
r.mtime.Unix(), r.mtime.Nanosecond(), r.head, r.tail, r.content)
if err != nil {
return fmt.Errorf("upsert %s: %w", r.path, err)
}
+148 -30
View File
@@ -5,10 +5,12 @@ import (
"database/sql"
"errors"
"fmt"
"os"
"path/filepath"
"slices"
"strings"
"testing"
"time"
)
// testDBPath returns a database path inside a fresh temp dir.
@@ -72,9 +74,78 @@ func TestOpenScanDatabaseCreates(t *testing.T) {
defer func() { _ = db.Close() }()
recs, err := loadFileRows(t.Context(), db)
if err != nil || len(recs) != 0 {
t.Fatalf("loadFileRows = %v, %v; want empty, nil", recs, err)
if recs := dbRecords(t, db); len(recs) != 0 {
t.Fatalf("records = %v, want none", recs)
}
}
func TestOpenDatabaseUnversionedForeign(t *testing.T) {
t.Parallel()
// A database that has a files table but user_version 0, written by
// some other tool. report, trees and scan must refuse it with the
// schema-version error, not adopt it and not emit a raw SQLite
// "table files already exists".
path := testDBPath(t)
db, err := sql.Open("sqlite", path)
if err != nil {
t.Fatal(err)
}
_, err = db.ExecContext(t.Context(), "CREATE TABLE files (x INTEGER)")
if err != nil {
t.Fatal(err)
}
_ = db.Close()
_, err = openReportDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) ||
!strings.Contains(err.Error(), "remove the file and rescan") {
t.Fatalf("report: err = %v, want errSchemaVersion telling the "+
"operator to remove the file and rescan", err)
}
_, err = openScanDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) ||
!strings.Contains(err.Error(), "remove the file and rescan") {
t.Fatalf("scan: err = %v, want errSchemaVersion telling the "+
"operator to remove the file and rescan", err)
}
}
func TestSchemaCreationStoppedPartway(t *testing.T) {
t.Parallel()
// A first scan stopped while creating the schema must leave a
// database the next scan accepts. max_page_count(2) leaves room for
// the files table but not its index, so schema creation fails right
// after CREATE TABLE, a point an interrupt could also stop it at.
path := testDBPath(t)
db, err := openDB(path, scanParams+"&_pragma=max_page_count(2)")
if err != nil {
t.Fatal(err)
}
err = initSchema(t.Context(), db)
_ = db.Close()
if err == nil {
t.Fatal("initSchema with no room for the index succeeded")
}
db, err = openScanDatabase(t.Context(), path)
if err != nil {
t.Fatalf("next scan: %v", err)
}
defer func() { _ = db.Close() }()
v, err := userVersion(t.Context(), db)
if err != nil || v != schemaVersion {
t.Fatalf("userVersion = %d, %v; want %d, nil", v, err, schemaVersion)
}
}
@@ -87,9 +158,11 @@ func TestOpenReportDatabaseMissing(t *testing.T) {
}
}
func TestOpenReportDatabaseVersionMismatch(t *testing.T) {
func TestOpenDatabaseVersionMismatch(t *testing.T) {
t.Parallel()
// A database stamped with a schema version other than 0 and
// schemaVersion. report, trees and scan must all refuse it.
path := testDBPath(t)
db, err := openScanDatabase(t.Context(), path)
@@ -106,7 +179,12 @@ func TestOpenReportDatabaseVersionMismatch(t *testing.T) {
_, err = openReportDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) {
t.Fatalf("err = %v, want errSchemaVersion", err)
t.Fatalf("report: err = %v, want errSchemaVersion", err)
}
_, err = openScanDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) {
t.Fatalf("scan: err = %v, want errSchemaVersion", err)
}
}
@@ -130,16 +208,66 @@ func TestOpenReportDatabaseOK(t *testing.T) {
_ = db.Close()
}
func TestCloseScanDatabaseWhileReportOpen(t *testing.T) {
t.Parallel()
// A report holding the database open stops scan from taking it out
// of WAL mode. The -wal and -shm files must then stay beside it, so
// that a later report still needs only read access.
path := testDBPath(t)
scanDB, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
reportDB, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
closeScanDatabase(t.Context(), scanDB, path)
_ = reportDB.Close()
_, err = os.Stat(path + "-wal")
if err != nil {
t.Fatalf("no -wal left: the switch out of WAL mode was not "+
"stopped: %v", err)
}
makeReadOnly(t, path)
reportDB, err = openReportDatabase(t.Context(), path)
if err != nil {
t.Fatalf("openReportDatabase: %v", err)
}
defer func() { _ = reportDB.Close() }()
err = loadFileRows(t.Context(), reportDB, func(scanRec) {})
if err != nil {
t.Fatalf("loadFileRows: %v", err)
}
}
func TestApplyChangesRoundTrip(t *testing.T) {
t.Parallel()
db := openTestDB(t)
// Paths may contain tabs and newlines; the database must store
// them byte-exactly.
// them byte-exactly. Every hash, content included, comes back as
// written.
recs := []scanRec{
{size: 2, mtime: 20, head: "h2", tail: "t2", path: "/a/tab\tnew\nline"},
{size: 1, mtime: 10, head: "h1", tail: "t1", path: "/a/x"},
{
size: 2, mtime: time.Unix(20, 999_999_999), head: "h2", tail: "t2",
content: "c2", path: "/a/tab\tnew\nline",
},
{
size: 1, mtime: time.Unix(10, 0), head: "h1", tail: "t1",
content: "c1", path: "/a/x",
},
}
err := applyChanges(t.Context(), db, recs, nil,
@@ -148,22 +276,18 @@ func TestApplyChangesRoundTrip(t *testing.T) {
t.Fatalf("applyChanges: %v", err)
}
got, err := loadFileRows(t.Context(), db)
if err != nil {
t.Fatal(err)
}
slices.SortFunc(got, func(a, b scanRec) int {
return strings.Compare(a.path, b.path)
})
// The records come back in path order, which is the order of recs.
got := dbRecords(t, db)
if !slices.Equal(got, recs) {
t.Fatalf("rows = %+v, want %+v", got, recs)
}
// An upsert for an existing path updates in place; a delete
// removes exactly its path.
upd := scanRec{size: 3, mtime: 30, head: "h3", tail: "t3", path: "/a/x"}
upd := scanRec{
size: 3, mtime: time.Unix(30, 0), head: "h3", tail: "t3", content: "c3",
path: "/a/x",
}
err = applyChanges(t.Context(), db, []scanRec{upd},
[]string{"/a/tab\tnew\nline"}, newProgress("update", 2))
@@ -171,11 +295,7 @@ func TestApplyChangesRoundTrip(t *testing.T) {
t.Fatalf("applyChanges: %v", err)
}
got, err = loadFileRows(t.Context(), db)
if err != nil {
t.Fatal(err)
}
got = dbRecords(t, db)
if len(got) != 1 || got[0] != upd {
t.Fatalf("rows = %+v, want just %+v", got, upd)
}
@@ -193,7 +313,7 @@ func TestApplyChangesBatching(t *testing.T) {
recs := make([]scanRec, 0, n)
for i := range n {
recs = append(recs, scanRec{
size: int64(i), mtime: 1, head: "h", tail: "t",
size: int64(i), mtime: time.Unix(1, 0), head: "h", tail: "t",
path: fmt.Sprintf("/batch/%07d", i),
})
}
@@ -204,9 +324,8 @@ func TestApplyChangesBatching(t *testing.T) {
t.Fatalf("applyChanges: %v", err)
}
got, err := loadFileRows(t.Context(), db)
if err != nil || len(got) != n {
t.Fatalf("loadFileRows = %d rows, %v; want %d", len(got), err, n)
if got := dbRecords(t, db); len(got) != n {
t.Fatalf("records = %d, want %d", len(got), n)
}
deletes := make([]string, 0, n)
@@ -220,8 +339,7 @@ func TestApplyChangesBatching(t *testing.T) {
t.Fatalf("applyChanges deletes: %v", err)
}
got, err = loadFileRows(t.Context(), db)
if err != nil || len(got) != 0 {
t.Fatalf("loadFileRows = %d rows, %v; want 0", len(got), err)
if got := dbRecords(t, db); len(got) != 0 {
t.Fatalf("records = %d, want 0", len(got))
}
}
+2 -2
View File
@@ -5,6 +5,8 @@ go 1.25.7
require (
github.com/schollz/progressbar/v3 v3.19.1
github.com/spf13/cobra v1.10.2
golang.org/x/sys v0.46.0
golang.org/x/term v0.44.0
modernc.org/sqlite v1.54.0
)
@@ -18,8 +20,6 @@ require (
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/rivo/uniseg v0.4.7 // indirect
github.com/spf13/pflag v1.0.9 // indirect
golang.org/x/sys v0.46.0 // indirect
golang.org/x/term v0.44.0 // indirect
modernc.org/libc v1.74.1 // indirect
modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect
+82 -28
View File
@@ -1,16 +1,21 @@
// Command sfdupes quickly identifies candidate duplicate files across
// very large filesystems without reading full file contents. Files are
// considered duplicates when they have identical size, identical SHA-256
// of their first 1024 bytes, and identical SHA-256 of their last 1024
// bytes. scan maintains a persistent SQLite database of file signatures
// (SFDUPES_DATABASE, default /var/lib/sfdupes/db.sqlite) that the
// reporting subcommands read.
// very large filesystems without reading every byte of every file.
// Files are considered duplicates when their sizes are equal and they
// agree on a short ladder of SHA-256 hashes. A file under 10 MiB is
// hashed in full. A larger file is compared on the hashes of its first
// and last 64 KiB, and only when those match another file's is its
// content hash computed and compared: of the whole file when it is
// under 50 MiB, or of gigabyte-spaced 1 MiB samples when it is 50 MiB
// or larger. scan maintains a persistent SQLite database of file
// signatures (SFDUPES_DATABASE, default /var/lib/sfdupes/db.sqlite)
// that the reporting subcommands read.
//
// Usage:
//
// sfdupes scan [--workers N] [-x] PATH...
// sfdupes report > dupes.tsv
// sfdupes trees > dupetrees.tsv
// sfdupes --version
//
// See README.md for the complete specification.
package main
@@ -46,6 +51,10 @@ const (
// cobra prints for it is the whole message.
var errNoSubcommand = errors.New("no subcommand")
// errWorkersBelowOne is the usage error for a scan --workers value
// below 1.
var errWorkersBelowOne = errors.New("--workers must be at least 1")
// Version is the build version, injected at link time via -ldflags
// (see the Makefile); "dev" for a plain go build.
//
@@ -53,22 +62,27 @@ var errNoSubcommand = errors.New("no subcommand")
var Version = "dev"
func main() {
os.Exit(run(os.Args[1:], os.Stderr))
// Once the reader of a stdout pipe has gone, as in "sfdupes report |
// head", the Go runtime ends the process with SIGPIPE on the next
// write instead of returning an error (README "Error handling").
// Registering for SIGPIPE with os/signal would change that.
os.Exit(run(os.Args[1:], os.Stdout, os.Stderr))
}
// run executes args against the command tree and returns the process
// exit code. It is the program's single exit point: the subcommands
// return their errors instead of exiting, so every deferred cleanup —
// above all closing the database, which checkpoints the SQLite WAL —
// runs before the process ends.
func run(args []string, stderr io.Writer) int {
// runs before the process ends. The report and trees subcommands write
// their data to stdout.
func run(args []string, stdout, stderr io.Writer) int {
// A nil slice makes cobra fall back to os.Args, which would let a
// test binary's own flags reach the command tree.
if args == nil {
args = []string{}
}
root := newRootCommand(stderr)
root := newRootCommand(stdout, stderr)
root.SetArgs(args)
err := root.Execute()
@@ -78,6 +92,9 @@ func run(args []string, stderr io.Writer) int {
switch {
case err == nil:
return exitOK
case errors.Is(err, errInterrupted):
// The interrupted scan has printed its own line.
return exitFatal
case errors.As(err, &fatal):
// The command ran and failed: a runtime error, reported
// without the usage text that a usage error gets.
@@ -85,22 +102,35 @@ func run(args []string, stderr io.Writer) int {
return exitFatal
default:
// A usage error: cobra has already printed the message and
// the usage text.
// A usage error, which cobra has already reported on stderr.
return exitUsage
}
}
// newRootCommand builds the command tree. Everything on stdout is
// machine-readable data; all human-facing output (help, usage, errors)
// goes to stderr.
func newRootCommand(stderr io.Writer) *cobra.Command {
// machine-readable data, the version line included; all human-facing
// output (help, usage, errors) goes to stderr.
func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
var showVersion bool
printVersion := runE(func(context.Context, []string) error {
_, err := fmt.Fprintf(stdout, "sfdupes %s\n", Version)
if err != nil {
return fmt.Errorf("write stdout: %w", err)
}
return nil
})
root := &cobra.Command{
Use: "sfdupes",
Short: "Find candidate duplicate files by size and head/tail SHA-256",
Version: Version,
Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, _ []string) error {
RunE: func(cmd *cobra.Command, args []string) error {
if showVersion {
return printVersion(cmd, args)
}
// A missing subcommand prints usage and exits 2: cobra
// prints the usage text for the returned error, and run
// maps everything that is not a fatal error to exit 2.
@@ -113,6 +143,11 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
root.SetErr(stderr)
root.CompletionOptions.DisableDefaultCmd = true
// Cobra's built-in version flag prints through the help writer,
// stderr; this one prints to stdout.
root.Flags().BoolVarP(&showVersion, "version", "v", false,
"print the version to stdout")
var (
scanWorkers int
scanOneFS bool
@@ -122,12 +157,18 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Use: cmdScan + " [--workers N] [-x] PATH...",
Short: "Walk trees and synchronize the scan database",
Args: cobra.MinimumNArgs(1),
PreRunE: func(cmd *cobra.Command, _ []string) error {
return checkScanWorkers(cmd, scanWorkers)
},
RunE: runE(func(ctx context.Context, args []string) error {
ctx, stop := interruptContext(ctx)
defer stop()
return runScan(ctx, args, scanWorkers, scanOneFS)
}),
}
scanCmd.Flags().IntVar(&scanWorkers, "workers", runtime.NumCPU(),
"concurrent workers for the walk and hash phases")
"concurrent workers for the walk, hash, and content phases")
scanCmd.Flags().BoolVarP(&scanOneFS, "one-file-system", "x", false,
"do not cross filesystem boundaries")
@@ -136,7 +177,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the file-level duplicates report",
Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error {
return runReport(ctx)
return runReport(ctx, stdout)
}),
}
@@ -145,7 +186,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the duplicate-tree report",
Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error {
return runTrees(ctx)
return runTrees(ctx, stdout)
}),
}
@@ -154,13 +195,26 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
return root
}
// runE adapts a subcommand implementation to cobra's RunE. Cobra
// prints the error and the command's usage text for every error RunE
// returns, but a subcommand that ran and failed has no usage problem
// to report: both are silenced here, and the error is marked fatal so
// that run reports it on stderr and exits 1 rather than 2. The command's
// context is handed to the implementation: cancelling it unwinds the
// scan's worker pools.
// checkScanWorkers rejects a scan --workers value below 1. That is a
// usage error reported in one line: cobra prints the returned message
// without the usage text, and run exits 2.
func checkScanWorkers(cmd *cobra.Command, workers int) error {
if workers >= 1 {
return nil
}
cmd.SilenceUsage = true
return fmt.Errorf("%w, got %d", errWorkersBelowOne, workers)
}
// runE adapts a subcommand implementation, or the version print, to
// cobra's RunE. Cobra prints the error and the command's usage text for
// every error RunE returns, but a subcommand that ran and failed has no
// usage problem to report: both are silenced here, and the error is
// marked fatal so that run reports it on stderr and exits 1 rather than
// 2. The command's context is handed to the implementation: cancelling
// it unwinds the scan's worker pools.
func runE(
fn func(ctx context.Context, args []string) error,
) func(*cobra.Command, []string) error {
+633 -69
View File
@@ -8,6 +8,7 @@ import (
"io/fs"
"os"
"path/filepath"
"slices"
"strconv"
"strings"
"testing"
@@ -44,23 +45,81 @@ func assertNoSidecars(t *testing.T, path string) {
}
}
// captureStdout redirects os.Stdout to a file for the rest of the test
// and returns a function reading back everything written to it. Only
// machine-readable data belongs on stdout (README design goal 4), so
// the tests assert on it directly.
func captureStdout(t *testing.T) func() string {
// makeReadOnly takes write permission away from the database at path,
// from any WAL sidecar beside it, and from their directory, as for a
// user reading a database that a root cron scan keeps. Root ignores
// file permissions, so it skips the test when run as root.
func makeReadOnly(t *testing.T, path string) {
t.Helper()
f, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
if os.Geteuid() == 0 {
t.Skip("root ignores file permissions")
}
err := os.Chmod(path, 0o400)
if err != nil {
t.Fatal(err)
}
saved := os.Stdout
os.Stdout = f
for _, suffix := range walSuffixes {
err = os.Chmod(path+suffix, 0o400)
if err != nil && !errors.Is(err, fs.ErrNotExist) {
t.Fatal(err)
}
}
dir := filepath.Dir(path)
//nolint:gosec // reaching the database needs the search bit
err = os.Chmod(dir, 0o500)
if err != nil {
t.Fatal(err)
}
// Runs before t.TempDir's own cleanup, which must delete the files.
t.Cleanup(func() {
//nolint:gosec // removing the directory needs its search bit back
_ = os.Chmod(dir, 0o700)
})
}
// captureStderr redirects os.Stderr to a file for the rest of the test
// and returns a function reading back everything written to it. scan
// writes its warnings and summary, and report and trees their
// summaries, straight to os.Stderr, not to the stderr writer run is
// given.
func captureStderr(t *testing.T) func() string {
t.Helper()
return capture(t, &os.Stderr)
}
// captureStdout does for os.Stdout what captureStderr does for
// os.Stderr. scan is never given run's stdout writer, so anything it
// printed would go straight to os.Stdout. The scan tests pass os.Stdout
// as run's stdout too, so the one capture sees both.
func captureStdout(t *testing.T) func() string {
t.Helper()
return capture(t, &os.Stdout)
}
// capture redirects *std, which is os.Stdout or os.Stderr, to a file
// for the rest of the test and returns a function reading back
// everything written to it.
func capture(t *testing.T, std **os.File) func() string {
t.Helper()
f, err := os.Create(filepath.Join(t.TempDir(), "output"))
if err != nil {
t.Fatal(err)
}
saved := *std
*std = f
t.Cleanup(func() {
os.Stdout = saved
*std = saved
_ = f.Close()
})
@@ -91,13 +150,13 @@ func captureStdout(t *testing.T) func() string {
// brokenDatabase writes a database that opens cleanly and passes the
// schema-version check but has no files table, so the first query
// fails with the database already open: a fatal error on a path that
// owns an open database.
// owns an open database. It closes the database the way scan does.
func brokenDatabase(t *testing.T) string {
t.Helper()
path := testDBPath(t)
db, err := openDB(path)
db, err := openDB(path, scanParams)
if err != nil {
t.Fatal(err)
}
@@ -108,10 +167,7 @@ func brokenDatabase(t *testing.T) string {
t.Fatal(err)
}
err = db.Close()
if err != nil {
t.Fatal(err)
}
closeScanDatabase(t.Context(), db, path)
return path
}
@@ -144,27 +200,31 @@ func TestOpenDatabaseKeepsWALWhileOpen(t *testing.T) {
func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
// Every subcommand that owns an open database must close it when
// it fails: no os.Exit between the open and the return.
cases := map[string][]string{
cmdScan: {cmdScan},
cmdReport: {cmdReport},
cmdTrees: {cmdTrees},
}
// it fails: no os.Exit between the open and the return. The
// sidecar check is evidence of the close only for scan: report and
// trees only read a database that is out of WAL mode, which leaves
// nothing on disk whether they close it or not.
//
// The subcommands come from the command tree, so a new one is
// checked too: one wired with a bare RunE instead of runE reports
// its failure as a usage error, exit 2 with the usage text.
for _, cmd := range newRootCommand(io.Discard, io.Discard).Commands() {
name := cmd.Name()
for name, args := range cases {
t.Run(name, func(t *testing.T) {
path := brokenDatabase(t)
t.Setenv(databaseEnv, path)
args := []string{name}
if name == cmdScan {
args = append(args, t.TempDir())
}
var stderr bytes.Buffer
stdout := captureStdout(t)
code := run(args, &stderr)
var stderr bytes.Buffer
code := run(args, os.Stdout, &stderr)
if code != exitFatal {
t.Errorf("run(%v) = %d, want %d", args, code, exitFatal)
}
@@ -182,19 +242,19 @@ func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
}
}
func TestRunMissingOperandIsFatalNotUsage(t *testing.T) {
func TestRunNonexistentPathIsFatalNotUsage(t *testing.T) {
// README §Error handling: a PATH operand that does not exist is a
// fatal error (1), not a usage error (2) — and a runtime failure
// must not dump the usage text.
t.Setenv(databaseEnv, testDBPath(t))
var stderr bytes.Buffer
stdout := captureStdout(t)
var stderr bytes.Buffer
missing := filepath.Join(t.TempDir(), "nope")
code := run([]string{cmdScan, missing}, &stderr)
code := run([]string{cmdScan, missing}, os.Stdout, &stderr)
if code != exitFatal {
t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal)
}
@@ -223,6 +283,34 @@ func assertFatalOutput(t *testing.T, stderr, stdout string) {
}
}
func TestRunMissingDatabaseIsFatal(t *testing.T) {
// README §Database: report and trees need an existing database; a
// missing one exits 1 with a message telling the user to run scan.
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitFatal {
t.Errorf("run(%s) = %d, want %d", name, code, exitFatal)
}
want := "sfdupes: " + path + ": no database (run \"sfdupes " +
"scan\" first, or set " + databaseEnv + ")\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
})
}
}
func TestRunUsageErrors(t *testing.T) {
// Usage errors keep exiting 2 with cobra's own report on stderr.
cases := map[string]struct {
@@ -243,11 +331,9 @@ func TestRunUsageErrors(t *testing.T) {
// path that does not exist.
t.Setenv(databaseEnv, testDBPath(t))
var stderr bytes.Buffer
var stdout, stderr bytes.Buffer
stdout := captureStdout(t)
code := run(tc.args, &stderr)
code := run(tc.args, &stdout, &stderr)
if code != exitUsage {
t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage)
}
@@ -256,43 +342,115 @@ func TestRunUsageErrors(t *testing.T) {
t.Errorf("stderr = %q, want %q", stderr.String(), tc.want)
}
if got := stdout(); got != "" {
if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
})
}
}
// TestRunHelpAndVersionSucceed checks that the two informational flags
// exit 0 and keep their human-facing output on stderr.
//
//nolint:paralleltest // captureStdout replaces the process-wide os.Stdout
func TestRunHelpAndVersionSucceed(t *testing.T) {
assertHumanOutput(t, "--help")
assertHumanOutput(t, "--version")
func TestRunScanRejectsWorkersBelowOne(t *testing.T) {
// README §scan mode: --workers below 1 is a usage error reported in
// one line on stderr, before the scan opens the database.
for _, workers := range []string{"0", "-1"} {
t.Run(workers, func(t *testing.T) {
dbPath := testDBPath(t)
t.Setenv(databaseEnv, dbPath)
var stdout, stderr bytes.Buffer
args := []string{cmdScan, "--workers", workers, t.TempDir()}
code := run(args, &stdout, &stderr)
if code != exitUsage {
t.Errorf("run(%v) = %d, want %d", args, code, exitUsage)
}
want := "Error: --workers must be at least 1, got " + workers +
"\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
_, err := os.Stat(dbPath)
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("stat %s: %v, want the database never created",
dbPath, err)
}
})
}
}
// assertHumanOutput runs sfdupes with one informational flag and checks
// that it succeeds with its output on stderr and stdout untouched
// (README design goal 4).
func assertHumanOutput(t *testing.T, arg string) {
t.Helper()
func TestRunHelp(t *testing.T) {
t.Parallel()
var stderr bytes.Buffer
// README §Subcommands: help goes to stderr, exits 0, and leaves
// stdout empty.
cases := [][]string{{"--help"}, {"-h"}, {cmdScan, "--help"}}
stdout := captureStdout(t)
for _, args := range cases {
var stdout, stderr bytes.Buffer
code := run([]string{arg}, &stderr)
code := run(args, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%v) = %d, want %d", args, code, exitOK)
}
if !strings.Contains(stderr.String(), usageMarker) {
t.Errorf("run(%v) stderr = %q, want the help text",
args, stderr.String())
}
if got := stdout.String(); got != "" {
t.Errorf("run(%v) stdout = %q, want nothing (data only)",
args, got)
}
}
}
func TestRunVersion(t *testing.T) {
t.Parallel()
// README §Subcommands: the version is one line on stdout, with
// nothing on stderr, and exits 0.
for _, arg := range []string{"--version", "-v"} {
var stdout, stderr bytes.Buffer
code := run([]string{arg}, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
}
if stderr.Len() == 0 {
t.Errorf("run(%s) wrote nothing to stderr", arg)
want := "sfdupes " + Version + "\n"
if got := stdout.String(); got != want {
t.Errorf("run(%s) stdout = %q, want %q", arg, got, want)
}
if got := stdout(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
if got := stderr.String(); got != "" {
t.Errorf("run(%s) stderr = %q, want nothing", arg, got)
}
}
}
func TestRunVersionWriteFailureIsFatal(t *testing.T) {
t.Parallel()
// README §Error handling: a stdout write failure exits 1, reported
// in one line on stderr.
var stderr bytes.Buffer
code := run([]string{"--version"}, failingWriter{}, &stderr)
if code != exitFatal {
t.Errorf("run(--version) = %d, want %d", code, exitFatal)
}
want := "sfdupes: write stdout: " + errWriteFailed.Error() + "\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
}
@@ -320,52 +478,195 @@ func scanFixture(t *testing.T) []string {
t.Fatal(err)
}
var stderr bytes.Buffer
scanOK(t, dir)
return dupes
}
// scanOK runs scan over operands, fails the test unless it exits 0 with
// nothing on stdout, and returns everything it printed to stderr.
func scanOK(t *testing.T, operands ...string) string {
t.Helper()
stdout := captureStdout(t)
stderr := captureStderr(t)
code := run([]string{cmdScan, dir}, &stderr)
code := run(append([]string{cmdScan}, operands...), os.Stdout, os.Stderr)
if code != exitOK {
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
t.Fatalf("run(scan %q) = %d, want %d; stderr: %s",
operands, code, exitOK, stderr())
}
if got := stdout(); got != "" {
t.Errorf("scan stdout = %q, want nothing (data only)", got)
}
return dupes
return stderr()
}
func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
// README §Error handling: a scan that skips a file it cannot read
// warns, counts the skip in its summary, and still exits 0, which
// scanOK checks along with the empty stdout.
if os.Geteuid() == 0 {
t.Skip("root ignores file permissions")
}
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
dir := t.TempDir()
writeFile(t, dir, "a.bin", pattern(1, 300))
// Same size as a.bin, so the scan reads it, and the read fails.
unreadable := writeFile(t, dir, "unreadable.bin", pattern(2, 300))
err := os.Chmod(unreadable, 0)
if err != nil {
t.Fatal(err)
}
stderr := scanOK(t, dir)
warning := "hash " + unreadable + ": open " + unreadable +
": permission denied\n"
if !strings.Contains(stderr, warning) {
t.Errorf("stderr = %q, want %q", stderr, warning)
}
summary := "scan: 1 files seen (1 added, 0 updated, 0 removed, " +
"0 unchanged), 1 skipped\n"
if !strings.Contains(stderr, summary) {
t.Errorf("stderr = %q, want %q", stderr, summary)
}
assertNoSidecars(t, path)
}
func TestRunScanSkipsSymlinkOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dir := t.TempDir()
writeFile(t, dir, "target/sub/f", pattern(1, 10))
link := filepath.Join(dir, "link")
err := os.Symlink(filepath.Join(dir, "target"), link)
if err != nil {
t.Fatal(err)
}
// Scanning a directory through the symlink stores a record beneath
// the symlink's own path for a file beneath its target.
scanOK(t, filepath.Join(link, "sub"))
assertOperandSkipped(t, path, link, "symlink",
filepath.Join(link, "sub", "f"))
}
func TestRunScanWalksOperandUnderSymlinkOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dir := t.TempDir()
writeFile(t, dir, "target/sub/f", pattern(1, 10))
link := filepath.Join(dir, "link")
err := os.Symlink(filepath.Join(dir, "target"), link)
if err != nil {
t.Fatal(err)
}
// link is dropped as a symlink, but link/sub must still be scanned,
// not dropped as lying under link.
scanOK(t, link, filepath.Join(link, "sub"))
db, err := openDB(path, reportParams)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
recordByPath(t, dbRecords(t, db), filepath.Join(link, "sub", "f"))
}
func TestRunScanSkipsZFSOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
zfs := filepath.Join(t.TempDir(), ".zfs")
snapshot := filepath.Join(zfs, "snapshot", "hourly")
f := writeFile(t, snapshot, "f", pattern(1, 10))
// An operand beneath a .zfs directory is walked, because it is not
// itself named .zfs.
scanOK(t, snapshot)
assertOperandSkipped(t, path, zfs, ".zfs directory", f)
}
// assertOperandSkipped scans operand alone and checks that it is skipped
// as kind: a warning naming it, one skip in the summary, exit 0, and the
// record for kept, which an earlier scan stored beneath operand, still
// in the database at dbPath.
func assertOperandSkipped(t *testing.T, dbPath, operand, kind,
kept string,
) {
t.Helper()
stderr := scanOK(t, operand)
warning := "walk " + operand + ": skipping " + kind + " operand\n"
if !strings.Contains(stderr, warning) {
t.Errorf("stderr = %q, want %q", stderr, warning)
}
summary := "scan: 0 files seen (0 added, 0 updated, 0 removed, " +
"0 unchanged), 1 skipped\n"
if !strings.Contains(stderr, summary) {
t.Errorf("stderr = %q, want %q", stderr, summary)
}
db, err := openDB(dbPath, reportParams)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
recordByPath(t, dbRecords(t, db), kept)
}
func TestRunReportSucceeds(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dupes := scanFixture(t)
var stderr bytes.Buffer
var stdout bytes.Buffer
stdout := captureStdout(t)
stderr := captureStderr(t)
code := run([]string{cmdReport}, &stderr)
code := run([]string{cmdReport}, &stdout, os.Stderr)
if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
code, exitOK, stderr())
}
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
if got := stdout(); got != want {
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
want = "report: 2 records read, 1 duplicate groups, 1 dupe files, " +
"300 B reclaimable\n"
if got := stderr(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
assertNoSidecars(t, path)
}
@@ -375,23 +676,286 @@ func TestRunTreesSucceeds(t *testing.T) {
dupes := scanFixture(t)
var stderr bytes.Buffer
var stdout bytes.Buffer
stdout := captureStdout(t)
stderr := captureStderr(t)
code := run([]string{cmdTrees}, &stderr)
code := run([]string{cmdTrees}, &stdout, os.Stderr)
if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
code, exitOK, stderr())
}
// The two directories holding the duplicate pair are duplicate
// trees of each other.
want := "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
if got := stdout(); got != want {
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
want = "trees: 2 records read, 1 duplicate tree groups, 1 dupe trees, " +
"300 B reclaimable\n"
if got := stderr(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
assertNoSidecars(t, path)
}
func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
// README §Database: report and trees need only read access to the
// database file. With its directory read-only as well, SQLite
// cannot create any file beside it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dupes := scanFixture(t)
assertNoSidecars(t, path)
makeReadOnly(t, path)
cases := map[string]string{
cmdReport: "first\tdupe\tsize\n" +
dupes[0] + "\t" + dupes[1] + "\t300\n",
cmdTrees: "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) +
"\t1\t300\n",
}
for name, want := range cases {
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
continue
}
if got := stdout.String(); got != want {
t.Errorf("%s stdout = %q, want %q", name, got, want)
}
}
}
// assertRunsUseDatabase runs scan, then report and trees, against the
// database that SFDUPES_DATABASE names, the file name in dir. It fails
// unless the reports find the duplicate pair the scan recorded and dir
// then holds only that file and its lock file: nothing was created
// under a shortened name.
func assertRunsUseDatabase(t *testing.T, dir, name string) {
t.Helper()
dupes := scanFixture(t)
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
if got := runStdout(t, cmdReport); got != want {
t.Errorf("report stdout = %q, want %q", got, want)
}
want = "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
if got := runStdout(t, cmdTrees); got != want {
t.Errorf("trees stdout = %q, want %q", got, want)
}
entries, err := os.ReadDir(dir)
if err != nil {
t.Fatal(err)
}
got := make([]string, 0, len(entries))
for _, e := range entries {
got = append(got, e.Name())
}
if wantFiles := []string{name, name + ".lock"}; !slices.Equal(got, wantFiles) {
t.Errorf("%s holds %q, want %q", dir, got, wantFiles)
}
}
func TestRunDatabasePathUsedAsGiven(t *testing.T) {
// README §Database: the path names the database file exactly. In
// SQLite's connection string an unescaped ? or # would end the file
// name and % would start an escape, and a path starting with //
// could be read as a host name. %25 is a valid escape, so unescaped
// this name opens a file named a without any error.
const name = "a?b#c%25d e.sqlite"
t.Run("absolute", func(t *testing.T) {
dir := t.TempDir()
t.Setenv(databaseEnv, filepath.Join(dir, name))
assertRunsUseDatabase(t, dir, name)
})
t.Run("leading double slash", func(t *testing.T) {
dir := t.TempDir()
t.Setenv(databaseEnv, "/"+filepath.Join(dir, name))
assertRunsUseDatabase(t, dir, name)
})
t.Run("relative", func(t *testing.T) {
dir := t.TempDir()
t.Chdir(dir)
t.Setenv(databaseEnv, name)
assertRunsUseDatabase(t, dir, name)
})
}
// holdScanLock takes the lock on the database at path, as a running
// scan does, and holds it until the test ends. It fails the test when
// the lock is already held.
func holdScanLock(t *testing.T, path string) {
t.Helper()
lock, err := lockScanDatabase(path)
if err != nil {
t.Fatalf("lock %s: %v", path, err)
}
t.Cleanup(func() { _ = lock.Close() })
}
func TestRunSecondScanFails(t *testing.T) {
// README §Database: while one scan holds the lock, a second scan
// fails at once, naming the lock file, without creating the
// database.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
holdScanLock(t, path)
stdout := captureStdout(t)
var stderr bytes.Buffer
code := run([]string{cmdScan, t.TempDir()}, os.Stdout, &stderr)
if code != exitFatal {
t.Errorf("run(scan) = %d, want %d", code, exitFatal)
}
want := "sfdupes: another scan is running (lock held on " +
path + ".lock)\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
if got := stdout(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
_, err := os.Stat(path)
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("stat %s = %v, want the database not created", path, err)
}
}
func TestRunScanReleasesLock(t *testing.T) {
// README §Database: a scan releases the lock however it ends.
t.Run("success", func(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
})
t.Run("fatal error", func(t *testing.T) {
path := brokenDatabase(t)
t.Setenv(databaseEnv, path)
code := run([]string{cmdScan, t.TempDir()}, io.Discard, io.Discard)
if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d", code, exitFatal)
}
holdScanLock(t, path)
})
}
func TestRunReportsDuringScan(t *testing.T) {
// README §Database: report and trees never take the lock, so they
// run while a scan holds it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
for _, name := range []string{cmdReport, cmdTrees} {
var stderr bytes.Buffer
code := run([]string{name}, io.Discard, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
}
}
}
func TestRunStdoutClosedIsFatal(t *testing.T) {
// README §Error handling: a stdout write failure exits 1, reported
// in one line on stderr.
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
stdout, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
if err != nil {
t.Fatal(err)
}
err = stdout.Close()
if err != nil {
t.Fatal(err)
}
var stderr bytes.Buffer
code := run([]string{name}, stdout, &stderr)
if code != exitFatal {
t.Errorf("run(%s) = %d, want %d", name, code, exitFatal)
}
got := stderr.String()
if !strings.HasPrefix(got, "sfdupes: write stdout: ") ||
!strings.Contains(got, os.ErrClosed.Error()) ||
strings.Count(got, "\n") != 1 {
t.Errorf("stderr = %q, want one line reporting the "+
"failed stdout write", got)
}
})
}
}
// errWriteFailed is the error failingWriter returns.
var errWriteFailed = errors.New("write failed")
// failingWriter is a stdout that fails every write.
type failingWriter struct{}
func (failingWriter) Write([]byte) (int, error) { return 0, errWriteFailed }
func TestStdoutWriteErrorPropagates(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
cases := map[string]func(context.Context, io.Writer) error{
cmdReport: runReport,
cmdTrees: runTrees,
}
for name, fn := range cases {
err := fn(t.Context(), failingWriter{})
if !errors.Is(err, errWriteFailed) {
t.Errorf("%s: error = %v, want %v", name, err, errWriteFailed)
}
}
}
+46 -17
View File
@@ -6,6 +6,7 @@ import (
"time"
"github.com/schollz/progressbar/v3"
"golang.org/x/term"
)
// plainInterval is the minimum time between progress lines when stderr
@@ -24,24 +25,23 @@ const percentScale = 100
// stderrIsTTY reports whether stderr is attached to a terminal.
func stderrIsTTY() bool {
fi, err := os.Stderr.Stat()
if err != nil {
return false
}
return fi.Mode()&os.ModeCharDevice != 0
return term.IsTerminal(int(os.Stderr.Fd()))
}
// progress renders one scan pass's progress on stderr. On a TTY it
// delegates to the progressbar library (spinner style when the total is
// unknown, full bar with count/percent/rate/elapsed/ETA otherwise). When
// stderr is not a TTY it emits no ANSI redraws: it prints a plain
// one-line update no more often than every plainInterval.
// one-line update as the pass starts, then no more often than every
// plainInterval.
//
// All methods must be called from the main goroutine only. A nil
// *progress is a valid no-display receiver: every method is a no-op,
// so batched database flushes during the streaming pass can reuse the
// update-pass helpers without rendering anything.
// All methods must be called from the main goroutine only. On a TTY
// the library also redraws a spinner from its own goroutine, several
// times a second, so its count and elapsed time stay current while a
// pass waits for its next item. A nil *progress is a valid
// no-display receiver: every method is a no-op, so batched database
// flushes during the streaming pass can reuse the update-pass helpers
// without rendering anything.
type progress struct {
label string
total int64 // -1 when unknown (walk pass)
@@ -53,10 +53,22 @@ type progress struct {
func newProgress(label string, total int64) *progress {
p := &progress{label: label, total: total, start: time.Now()}
if !stderrIsTTY() {
if stderrIsTTY() {
p.bar = newBar(label, total)
return p
}
// Print the zero state at once: the first item may take minutes,
// and a pass must never look hung.
p.last = p.start
fmt.Fprintln(os.Stderr, p.plainLine())
return p
}
// newBar builds the TTY display for newProgress.
func newBar(label string, total int64) *progressbar.ProgressBar {
opts := []progressbar.Option{
progressbar.OptionSetWriter(os.Stderr),
progressbar.OptionSetDescription(label),
@@ -81,9 +93,7 @@ func newProgress(label string, total int64) *progress {
)
}
p.bar = progressbar.NewOptions64(total, opts...)
return p
return progressbar.NewOptions64(total, opts...)
}
// increment records one completed item and refreshes the display.
@@ -106,26 +116,45 @@ func (p *progress) increment() {
}
// warnf prints a one-line warning to stderr without corrupting the bar.
// The whole message is escaped like a report's path columns, so a path
// holding a newline cannot split the warning.
func (p *progress) warnf(format string, args ...any) {
if p == nil {
return
}
msg := escapePath(fmt.Sprintf(format, args...))
if p.bar != nil && p.total < 0 {
// The library also redraws a spinner from its own goroutine, so
// a direct write could land inside a redraw. The bar prints the
// warning itself, just before its next redraw.
_, _ = progressbar.Bprintln(p.bar, msg)
return
}
if p.bar != nil {
_ = p.bar.Clear()
}
fmt.Fprintf(os.Stderr, format+"\n", args...)
fmt.Fprintln(os.Stderr, msg)
}
// finish terminates the pass's display.
// finish terminates the pass's display. A bar whose pass stopped short
// of its total, as an interrupted one does, is left as last drawn; the
// library's Finish would fill it up.
func (p *progress) finish() {
if p == nil {
return
}
if p.bar != nil {
if p.total >= 0 && p.count < p.total {
_ = p.bar.Exit()
} else {
_ = p.bar.Finish()
}
fmt.Fprintln(os.Stderr)
+189
View File
@@ -0,0 +1,189 @@
package main
import (
"os"
"path/filepath"
"strings"
"testing"
"time"
)
// spinnerIdle comfortably outlasts the 100ms interval at which the
// progressbar library redraws a spinner from its own goroutine.
const spinnerIdle = 500 * time.Millisecond
//nolint:paralleltest // replaces the process-wide os.Stderr
func TestStderrIsTTYFalseForNonTerminals(t *testing.T) {
r, pipe, err := os.Pipe()
if err != nil {
t.Fatal(err)
}
regular, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
devNull, err := os.OpenFile(os.DevNull, os.O_WRONLY, 0)
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
t.Cleanup(func() {
os.Stderr = saved
for _, f := range []*os.File{r, pipe, regular, devNull} {
_ = f.Close()
}
})
cases := map[string]*os.File{
"a pipe": pipe,
"a regular file": regular,
os.DevNull: devNull,
}
for name, f := range cases {
os.Stderr = f
if stderrIsTTY() {
t.Errorf("stderrIsTTY() = true with stderr on %s", name)
}
}
}
// TestNewProgressPrintsBeforeFirstItem checks that each pass shows its
// zero state the moment it starts when stderr is not a terminal, and
// that the next line still waits for plainInterval.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestNewProgressPrintsBeforeFirstItem(t *testing.T) {
stderr := captureStderr(t)
newProgress("walk", -1).increment()
newProgress("hash", 10).increment()
want := "walk: 0 files, elapsed 0s\n" +
"hash: [0/10] 0% 0 files/s elapsed 0s eta ?\n"
if got := stderr(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
}
// newWalkSpinner returns the walk pass's terminal display, writing to
// os.Stderr whether or not it is a terminal, and stops the library's
// redraws when the test ends.
func newWalkSpinner(t *testing.T) *progress {
t.Helper()
p := &progress{
label: "walk", total: -1, start: time.Now(),
bar: newBar("walk", -1),
}
t.Cleanup(p.finish)
return p
}
// TestProgressWarningsOnOwnLines drives the terminal display of the walk
// pass through a run of warnings with no items between them, as when the
// walk meets many unreadable paths, for several of the spinner's
// redraws: every warning must land on a line of its own, never inside a
// redraw.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestProgressWarningsOnOwnLines(t *testing.T) {
stderr := captureStderr(t)
p := newWalkSpinner(t)
// No pause between warnings: one written straight to stderr is
// garbled only if a redraw lands while it is being written.
issued := 0
for start := time.Now(); time.Since(start) < spinnerIdle; issued++ {
p.warnf("warning")
}
// The spinner prints the warnings at its next redraw.
time.Sleep(spinnerIdle)
// A terminal shows each line as the text after its last carriage
// return.
shown := 0
for line := range strings.SplitSeq(stderr(), "\n") {
if !strings.Contains(line, "warning") {
continue
}
shown++
if text := line[strings.LastIndex(line, "\r")+1:]; text != "warning" {
t.Errorf("terminal shows %q, want %q", text, "warning")
}
}
if shown != issued {
t.Errorf("%d warning lines, want %d", shown, issued)
}
}
// TestSpinnerShowsCountAfterBurst checks that once a burst of items
// faster than the redraw limit is over, the walk display shows every
// item completed while it waits for the next one.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestSpinnerShowsCountAfterBurst(t *testing.T) {
stderr := captureStderr(t)
p := newWalkSpinner(t)
for range 50 {
p.increment()
}
time.Sleep(spinnerIdle)
if shown := lastFrame(stderr()); !strings.Contains(shown, "(50/-,") {
t.Errorf("terminal shows %q, want a count of 50", shown)
}
}
// lastFrame returns what a terminal shows of the frames a bar drew: the
// last one. The library starts each frame with a carriage return and
// erases the previous one with spaces first.
func lastFrame(out string) string {
var shown string
for frame := range strings.SplitSeq(out, "\r") {
if strings.TrimSpace(frame) != "" {
shown = frame
}
}
return shown
}
// TestBarStoppedShortKeepsCount checks that the terminal display of a
// pass that stops before its total, as an interrupted one does, is left
// as last drawn instead of being filled up.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestBarStoppedShortKeepsCount(t *testing.T) {
stderr := captureStderr(t)
p := &progress{
label: "hash", total: 10, start: time.Now(),
bar: newBar("hash", 10),
}
p.increment()
// Past the redraw limit, so the bar draws the next count.
time.Sleep(2 * barThrottle)
p.increment()
p.finish()
if shown := lastFrame(stderr()); !strings.Contains(shown, "(2/10,") {
t.Errorf("terminal shows %q, want a count of 2 of 10", shown)
}
}
+51 -92
View File
@@ -4,9 +4,10 @@ import (
"bufio"
"context"
"fmt"
"io"
"os"
"slices"
"strings"
"time"
)
// ioBufSize is the buffer size for the buffered stdout writers.
@@ -17,82 +18,70 @@ const ioBufSize = 1 << 20
const minGroupSize = 2
// scanRec is one file record from the database. The signature (size,
// head, tail) is the duplicate key; mtime is informational only and
// used by scan for change detection.
// head, tail, content) is the duplicate key; mtime is informational
// only and used by scan for change detection.
type scanRec struct {
size int64
mtime int64
mtime time.Time
head string
tail string
content string
path string
}
// loadRecords opens the database and reads every file record for the
// report and trees subcommands. Any database problem — including a
// missing database — is fatal. The error is returned rather than
// exiting, so that the deferred close — which checkpoints the SQLite
// WAL — always runs; the database is closed before the caller formats
// its output, so it stays closed even if that output fails.
func loadRecords(ctx context.Context) ([]scanRec, error) {
// runReport implements the report subcommand: it prints the file-level
// duplicates report as TSV on stdout. SQLite groups and orders the
// records, and each row is written as it is read, so no group is held
// in memory. It never touches the scanned filesystem; its only I/O is
// the database (with SQLite's temporary sort file), stdout, and stderr.
// Any database problem, including a missing database, is fatal.
func runReport(ctx context.Context, stdout io.Writer) error {
dbPath := databasePath()
db, err := openReportDatabase(ctx, dbPath)
if err != nil {
return nil, err
return err
}
defer func() { _ = db.Close() }()
recs, err := loadFileRows(ctx, db)
if err != nil {
return nil, fmt.Errorf("database %s: %w", dbPath, err)
}
return recs, nil
}
// dupeGroup is one set of candidate-duplicate files: identical size,
// head hash, and tail hash. paths is sorted lexicographically; the
// first entry is the group's "first", the rest are dupes.
type dupeGroup struct {
size int64
paths []string
}
// runReport implements the report subcommand: it reads every record
// from the database and prints the file-level duplicates report as TSV
// on stdout. It never touches the scanned filesystem; its only I/O is
// the database, stdout, and stderr.
func runReport(ctx context.Context) error {
recs, err := loadRecords(ctx)
if err != nil {
return err
}
dupes := collectDupeGroups(recs)
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
out := bufio.NewWriterSize(stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tsize")
if err != nil {
return fmt.Errorf("write stdout: %w", err)
}
dupeFiles := 0
var (
groups, dupeFiles int
reclaimable int64
writeErr error
)
var reclaimable int64
records, err := loadDupeRows(ctx, db,
func(first, path string, size int64) error {
// A group's first path is its first row; every other
// path is a dupe.
if path == first {
groups++
for _, g := range dupes {
for _, p := range g.paths[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\n",
g.paths[0], p, g.size)
if err != nil {
return fmt.Errorf("write stdout: %w", err)
return nil
}
_, writeErr = fmt.Fprintf(out, "%s\t%s\t%d\n",
escapePath(first), escapePath(path), size)
dupeFiles++
reclaimable += g.size
reclaimable += size
return writeErr
})
if writeErr != nil {
return fmt.Errorf("write stdout: %w", writeErr)
}
if err != nil {
return fmt.Errorf("database %s: %w", dbPath, err)
}
err = out.Flush()
@@ -103,54 +92,24 @@ func runReport(ctx context.Context) error {
fmt.Fprintf(os.Stderr,
"report: %d records read, %d duplicate groups, %d dupe files, "+
"%s reclaimable\n",
len(recs), len(dupes), dupeFiles, humanBytes(reclaimable))
records, groups, dupeFiles, humanBytes(reclaimable))
return nil
}
// collectDupeGroups groups records by signature and returns every group
// with two or more paths, each group's paths sorted lexicographically,
// groups ordered by size descending then by first path ascending.
func collectDupeGroups(recs []scanRec) []dupeGroup {
groups := make(map[fileSig][]string)
for _, r := range recs {
// A record without hashes (its size was unique when last
// scanned) has unknown content and is never reported as a
// duplicate.
if r.head == "" {
continue
// escapePath returns a path as it is written in a report column (README
// "Report output format"): a backslash, tab, newline or carriage return
// becomes \\, \t, \n or \r, and every other byte is kept as it is.
// Grouping and sorting use the raw path, never this form.
func escapePath(p string) string {
// Most paths need no escaping; skip building a replacer for them.
if !strings.ContainsAny(p, "\\\t\n\r") {
return p
}
k := fileSig{size: r.size, head: r.head, tail: r.tail}
groups[k] = append(groups[k], r.path)
}
var dupes []dupeGroup
for k, paths := range groups {
if len(paths) < minGroupSize {
continue
}
slices.Sort(paths)
dupes = append(dupes, dupeGroup{size: k.size, paths: paths})
}
// Biggest reclaimable space first; ties broken by first path.
slices.SortFunc(dupes, func(a, b dupeGroup) int {
if a.size != b.size {
if a.size > b.size {
return -1
}
return 1
}
return strings.Compare(a.paths[0], b.paths[0])
})
return dupes
return strings.NewReplacer(
`\`, `\\`, "\t", `\t`, "\n", `\n`, "\r", `\r`,
).Replace(p)
}
// humanBytes formats a byte count in human units (binary prefixes).
+304 -26
View File
@@ -1,26 +1,270 @@
package main
import (
"bytes"
"database/sql"
"errors"
"fmt"
"io"
"os"
"path/filepath"
"slices"
"strings"
"testing"
"time"
)
func TestCollectDupeGroups(t *testing.T) {
// awkwardDir is a directory name holding every byte the reports escape.
const awkwardDir = "/d/\tone\ntwo\rthree\\four"
// awkwardPairRecs is a duplicate pair in sibling directories /d/A and
// awkwardDir. A raw tab sorts before "A" but its escaped form `\t`
// sorts after it, so awkwardDir coming first shows that sorting uses
// the raw path.
func awkwardPairRecs() []scanRec {
return []scanRec{
{size: 5, head: "h", tail: "t", content: "c", path: "/d/A/f"},
{size: 5, head: "h", tail: "t", content: "c", path: awkwardDir + "/f"},
}
}
// seedDatabase writes recs into a fresh database and returns its path.
func seedDatabase(t *testing.T, recs []scanRec) string {
t.Helper()
path := testDBPath(t)
db, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
err = applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
err = db.Close()
if err != nil {
t.Fatal(err)
}
return path
}
// dupeGroup is one duplicate group as report reads it: the size, and
// the paths in report order, first path first.
type dupeGroup struct {
size int64
paths []string
}
// dupeGroups returns the duplicate groups report reads from db, in
// report order.
func dupeGroups(t *testing.T, db *sql.DB) []dupeGroup {
t.Helper()
var groups []dupeGroup
_, err := loadDupeRows(t.Context(), db,
func(first, path string, size int64) error {
if path == first {
groups = append(groups, dupeGroup{size: size})
}
g := &groups[len(groups)-1]
g.paths = append(g.paths, path)
return nil
})
if err != nil {
t.Fatal(err)
}
return groups
}
// dupeGroupsOf writes recs into a fresh database and returns the
// duplicate groups report reads from it.
func dupeGroupsOf(t *testing.T, recs []scanRec) []dupeGroup {
t.Helper()
db := openTestDB(t)
err := applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
return dupeGroups(t, db)
}
func TestRunReportEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
code := run([]string{cmdReport}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tsize\n" +
`/d/\tone\ntwo\rthree\\four/f` + "\t/d/A/f\t5\n"
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
func TestReportStdoutFailsWhileReading(t *testing.T) {
// Each row holds two paths longer than dir, so the report is more
// than twice the stdout buffer and stdout fails while rows are
// still being read, not at the final flush.
dir := "/" + strings.Repeat("d", 4096)
recs := make([]scanRec, ioBufSize/len(dir))
for i := range recs {
recs[i] = scanRec{
size: 1, head: "h", tail: "t", content: "c",
path: fmt.Sprintf("%s/%d", dir, i),
}
}
t.Setenv(databaseEnv, seedDatabase(t, recs))
err := runReport(t.Context(), failingWriter{})
if !errors.Is(err, errWriteFailed) ||
!strings.HasPrefix(err.Error(), "write stdout: ") {
t.Errorf("error = %v, want write stdout: %v", err, errWriteFailed)
}
}
func TestRunReportsIgnoreInsertionOrder(t *testing.T) {
// README §Constraints: identical database contents give identical
// output, whatever order the records were inserted in.
recs := append(smokeTreeRecs(), awkwardPairRecs()...)
recs = append(recs,
scanRec{size: 50, head: "b", tail: "b", content: "b", path: "/y/2"},
scanRec{size: 50, head: "b", tail: "b", content: "b", path: "/y/1"},
scanRec{size: 50, head: "a", tail: "a", content: "a", path: "/x/2"},
scanRec{size: 50, head: "a", tail: "a", content: "a", path: "/x/1"},
scanRec{size: 50, path: "/x/unhashed"},
)
reversed := slices.Clone(recs)
slices.Reverse(reversed)
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, recs))
forward := runStdout(t, name)
t.Setenv(databaseEnv, seedDatabase(t, reversed))
backward := runStdout(t, name)
if strings.Count(forward, "\n") < 3 {
t.Errorf("stdout = %q, want at least two rows", forward)
}
if forward != backward {
t.Errorf("stdout depends on insertion order: %q vs %q",
forward, backward)
}
})
}
}
// runStdout runs the subcommand name and returns its stdout, failing
// the test unless it succeeds.
func runStdout(t *testing.T, name string) string {
t.Helper()
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
}
return stdout.String()
}
func TestEscapePath(t *testing.T) {
t.Parallel()
cases := map[string]string{
"/srv/plain": "/srv/plain",
"/a\tb": `/a\tb`,
"/a\nb": `/a\nb`,
"/a\rb": `/a\rb`,
`/a\b`: `/a\\b`,
`/a\tb`: `/a\\tb`,
"/not-utf8\xff": "/not-utf8\xff",
}
for in, want := range cases {
if got := escapePath(in); got != want {
t.Errorf("escapePath(%q) = %q, want %q", in, got, want)
}
}
}
// TestWarnfEscapes checks that a warning naming a path that holds a
// newline is still one line.
//
//nolint:paralleltest // replaces the process-wide os.Stderr
func TestWarnfEscapes(t *testing.T) {
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
os.Stderr = f
t.Cleanup(func() {
os.Stderr = saved
_ = f.Close()
})
(&progress{}).warnf("stat %s: %s", "/d/a\nb", "gone")
_, err = f.Seek(0, io.SeekStart)
if err != nil {
t.Fatal(err)
}
got, err := io.ReadAll(f)
if err != nil {
t.Fatal(err)
}
want := `stat /d/a\nb: gone` + "\n"
if string(got) != want {
t.Errorf("warning = %q, want %q", got, want)
}
}
func TestDupeGroups(t *testing.T) {
t.Parallel()
recs := []scanRec{
{size: 100, head: "h", tail: "t", path: "/z/b"},
{size: 100, head: "h", tail: "t", path: "/z/a"},
{size: 100, head: "h", tail: "t", path: "/z/c"},
{size: 4000, head: "H", tail: "T", path: "/big/2"},
{size: 4000, head: "H", tail: "T", path: "/big/1"},
{size: 100, head: "h", tail: "t", content: "c", path: "/z/b"},
{size: 100, head: "h", tail: "t", content: "c", path: "/z/a"},
{size: 100, head: "h", tail: "t", content: "c", path: "/z/c"},
{size: 4000, head: "H", tail: "T", content: "C", path: "/big/2"},
{size: 4000, head: "H", tail: "T", content: "C", path: "/big/1"},
// Same size as the /z group but a different head hash.
{size: 100, head: "other", tail: "t", path: "/z/d"},
{size: 100, head: "other", tail: "t", content: "c", path: "/z/d"},
// A singleton signature must not form a group.
{size: 7, head: "u", tail: "u", path: "/lonely"},
{size: 7, head: "u", tail: "u", content: "u", path: "/lonely"},
}
groups := collectDupeGroups(recs)
groups := dupeGroupsOf(t, recs)
if len(groups) != 2 {
t.Fatalf("len(groups) = %d, want 2", len(groups))
}
@@ -38,33 +282,67 @@ func TestCollectDupeGroups(t *testing.T) {
}
}
func TestCollectDupeGroupsMtimeExcluded(t *testing.T) {
func TestDupeGroupsContentSeparates(t *testing.T) {
t.Parallel()
// Same size, head, and tail, but different content hashes: the final
// rung keeps them apart, so no group forms. Matching content groups.
// Records without a content hash never group, not even with each
// other.
recs := []scanRec{
{size: 100, head: "h", tail: "t", content: "c1", path: "/a"},
{size: 100, head: "h", tail: "t", content: "c2", path: "/b"},
{size: 100, head: "h", tail: "t", content: "c1", path: "/c"},
{size: 100, head: "h", tail: "t", path: "/d"},
{size: 100, head: "h", tail: "t", path: "/e"},
}
groups := dupeGroupsOf(t, recs)
if len(groups) != 1 {
t.Fatalf("len(groups) = %d, want 1 (only the matching content)",
len(groups))
}
if !slices.Equal(groups[0].paths, []string{"/a", "/c"}) {
t.Errorf("group paths = %q, want /a /c", groups[0].paths)
}
}
func TestDupeGroupsMtimeExcluded(t *testing.T) {
t.Parallel()
// mtime is informational only; records differing only in mtime
// still group together.
recs := []scanRec{
{size: 9, mtime: 100, head: "h", tail: "t", path: "/m/1"},
{size: 9, mtime: 200, head: "h", tail: "t", path: "/m/2"},
{
size: 9, mtime: time.Unix(100, 0), head: "h", tail: "t",
content: "c", path: "/m/1",
},
{
size: 9, mtime: time.Unix(200, 0), head: "h", tail: "t",
content: "c", path: "/m/2",
},
}
groups := collectDupeGroups(recs)
groups := dupeGroupsOf(t, recs)
if len(groups) != 1 {
t.Fatalf("len(groups) = %d, want 1", len(groups))
}
}
func TestCollectDupeGroupsTieBreak(t *testing.T) {
func TestDupeGroupsTieBreak(t *testing.T) {
t.Parallel()
// The hashes sort opposite to the first paths, so ordering the
// groups by hash instead of by first path fails this test.
recs := []scanRec{
{size: 50, head: "b", tail: "b", path: "/beta/2"},
{size: 50, head: "b", tail: "b", path: "/beta/1"},
{size: 50, head: "a", tail: "a", path: "/alpha/2"},
{size: 50, head: "a", tail: "a", path: "/alpha/1"},
{size: 50, head: "a", tail: "a", content: "a", path: "/beta/2"},
{size: 50, head: "a", tail: "a", content: "a", path: "/beta/1"},
{size: 50, head: "b", tail: "b", content: "b", path: "/alpha/2"},
{size: 50, head: "b", tail: "b", content: "b", path: "/alpha/1"},
}
groups := collectDupeGroups(recs)
groups := dupeGroupsOf(t, recs)
if len(groups) != 2 {
t.Fatalf("len(groups) = %d, want 2", len(groups))
}
@@ -76,22 +354,22 @@ func TestCollectDupeGroupsTieBreak(t *testing.T) {
}
}
func TestCollectDupeGroupsDeterministic(t *testing.T) {
func TestDupeGroupsDeterministic(t *testing.T) {
t.Parallel()
recs := []scanRec{
{size: 1, head: "a", tail: "a", path: "/p/1"},
{size: 1, head: "a", tail: "a", path: "/p/2"},
{size: 2, head: "b", tail: "b", path: "/q/1"},
{size: 2, head: "b", tail: "b", path: "/q/2"},
{size: 1, head: "a", tail: "a", content: "a", path: "/p/1"},
{size: 1, head: "a", tail: "a", content: "a", path: "/p/2"},
{size: 2, head: "b", tail: "b", content: "b", path: "/q/1"},
{size: 2, head: "b", tail: "b", content: "b", path: "/q/2"},
}
forward := collectDupeGroups(recs)
forward := dupeGroupsOf(t, recs)
reversed := slices.Clone(recs)
slices.Reverse(reversed)
backward := collectDupeGroups(reversed)
backward := dupeGroupsOf(t, reversed)
if !slices.EqualFunc(forward, backward, func(a, b dupeGroup) bool {
return a.size == b.size && slices.Equal(a.paths, b.paths)
}) {
+576 -113
View File
@@ -6,30 +6,70 @@ import (
"crypto/sha256"
"database/sql"
"encoding/hex"
"errors"
"fmt"
"io"
"io/fs"
"os"
"os/signal"
"path/filepath"
"slices"
"strings"
"sync"
"syscall"
"time"
)
// chunk is the number of bytes hashed from each end of a file.
const chunk = 1024
// The duplicate ladder (see hashSignature and README "Duplicate
// detection"). A same-size candidate below headTailMin is hashed in
// full and compared directly; a larger one is separated first by the
// hashes of its end windows, then by a content hash that is exact below
// wholeFileMax and deliberately sampled at or above it. The hash phase
// reads only the end windows of a larger file; the content phase reads
// it for its content hash only once its size, head, and tail match
// another file's.
// headTailMin is the size threshold for the end-window gate. A file
// smaller than this is hashed in full directly, with no separate head
// and tail step: its head, tail, and content all carry the whole-file
// hash. A file this size or larger is separated first by its end
// windows.
const headTailMin = 10 * 1024 * 1024
// headTailWindow is the number of bytes hashed from each end of a file
// at or above headTailMin (the head and tail rungs). Because
// headTailMin is far larger than two windows, the head and tail windows
// never overlap.
const headTailWindow = 64 * 1024
// wholeFileMax is the size boundary between the two content rungs: a
// file strictly smaller than this is content-hashed in full; a file
// this size or larger is content-hashed by sampling.
const wholeFileMax = 50 * 1024 * 1024
// sampleStride is the spacing between content samples for large files:
// one window is read at each gigabyte-aligned offset (0, 1 GiB, ...).
const sampleStride = 1024 * 1024 * 1024
// sampleWindow is the number of bytes read at each large-file sample
// offset, truncated at end of file.
const sampleWindow = 1024 * 1024
// workQueueDepth bounds the job and result channels feeding the walk
// and hash worker pools.
const workQueueDepth = 1024
// errInterrupted reports a scan stopped by SIGINT or SIGTERM. runScan
// has already printed its line, so run prints nothing more.
var errInterrupted = errors.New("scan interrupted")
// fileRec carries one statted file between the scan phases. dev and
// ino identify the underlying inode so hard-linked paths can share
// one read; both are zero when the platform exposes no inode.
type fileRec struct {
path string
size int64
mtime int64
mtime time.Time
dev uint64
ino uint64
}
@@ -40,27 +80,30 @@ type fileRec struct {
// they would dominate the scan's memory.
type fileMeta struct {
size int64
mtime int64
mtime time.Time
hashed bool
}
// runScan implements the scan subcommand: three sequential phases —
// walk (which stats each file as it is discovered), hash, update —
// that synchronize the persistent database with the filesystem state
// under the PATH operands. Only files whose size at least one other
// file shares are ever hashed: a size-unique file cannot be a
// duplicate. Flag parsing and the at-least-one-operand check are done
// by cobra. Errors are returned rather than exiting, so that the
// deferred close — which checkpoints the SQLite WAL — always runs.
// Cancelling ctx unwinds the worker pools and aborts the scan with the
// context's error.
// runScan implements the scan subcommand: four sequential phases —
// walk (which stats each file as it is discovered), hash, update,
// content — that synchronize the persistent database with the
// filesystem state under the PATH operands. Only files whose size at
// least one other file shares are ever hashed: a size-unique file
// cannot be a duplicate. A file of headTailMin or more gets its content
// hash only when its size, head, and tail match another file's. Flag
// parsing and the at-least-one-operand check are done by cobra. The
// scan holds the lock on the database for its whole run, so a second
// scan fails before it walks the filesystem or opens the database.
// Errors are returned rather than exiting, so that the deferred close —
// which takes the database out of WAL mode — always runs, and the lock
// is released after it. When ctx is cancelled, as by the SIGINT or
// SIGTERM that interruptContext catches, the scan keeps what it has
// hashed (see syncScan), prints how many files its walk reached, and
// returns errInterrupted. workers must be at least 1; the scan command
// rejects anything less.
func runScan(ctx context.Context, roots []string, workers int,
oneFS bool,
) error {
if workers < 1 {
workers = 1
}
roots, err := resolveRoots(roots)
if err != nil {
return err
@@ -68,14 +111,31 @@ func runScan(ctx context.Context, roots []string, workers int,
dbPath := databasePath()
db, err := openScanDatabase(ctx, dbPath)
lock, err := lockScanDatabase(dbPath)
if err != nil {
return err
}
defer func() { _ = db.Close() }()
defer func() { _ = lock.Close() }()
db, err := openScanDatabase(ctx, dbPath)
if err != nil && ctx.Err() != nil {
// Interrupted while opening; SQLite may report that with an
// error of its own rather than the context's.
return interrupted(0)
}
if err != nil {
return err
}
defer closeScanDatabase(ctx, db, dbPath)
st, err := syncScan(ctx, db, roots, workers, oneFS)
if errors.Is(err, context.Canceled) {
return interrupted(st.walked)
}
if err != nil {
return fmt.Errorf("update database %s: %w", dbPath, err)
}
@@ -89,6 +149,34 @@ func runScan(ctx context.Context, roots []string, workers int,
return nil
}
// interruptContext returns a copy of ctx that the first SIGINT or
// SIGTERM cancels; the scan command runs the scan under it. stop
// releases the signals.
func interruptContext(ctx context.Context) (context.Context, func()) {
// A SIGINT ignored from the start, as by a script's background job,
// stays ignored.
signals := []os.Signal{syscall.SIGTERM}
if !signal.Ignored(syscall.SIGINT) {
signals = append(signals, syscall.SIGINT)
}
ctx, stop := signal.NotifyContext(ctx, signals...)
// Stopping restores the default handling, so a second signal ends
// the process at once.
context.AfterFunc(ctx, stop)
return ctx, stop
}
// interrupted prints the line for a scan stopped by a signal after its
// walk reached walked files, and returns errInterrupted.
func interrupted(walked int) error {
fmt.Fprintf(os.Stderr, "scan: interrupted after %d files\n", walked)
return errInterrupted
}
// resolveRoots converts each PATH operand to an absolute, lexically
// cleaned path (symlinks are not resolved) and verifies that it
// exists. Database records are keyed by absolute path, so scan results
@@ -142,8 +230,9 @@ func pruneRoots(roots []string) []string {
}
// scanStats summarizes one scan's database synchronization for the
// final stderr summary.
// final stderr summary, or for the line an interrupted scan prints.
type scanStats struct {
walked int // files the walk reached
added int
updated int
removed int
@@ -165,49 +254,106 @@ type scanState struct {
st scanStats
}
// syncScan synchronizes the database with the filesystem under roots
// in three sequential phases: walk (enumerate and stat every file,
// building a complete size census), hash (read only the new or
// changed — or previously unhashed — files whose size at least one
// other file shares, committing results in batches as they arrive),
// and update (record the size-unique files without reading them, and
// delete the records the scan no longer verifies). Records outside
// the roots are never touched.
// syncScan synchronizes the database with the filesystem under roots;
// see runPhases. When ctx is cancelled, as by an interrupt, it commits
// the hashed records still waiting in the batch, starts no other write
// or deletion, and returns the cancellation.
func syncScan(ctx context.Context, db *sql.DB, roots []string,
workers int, oneFS bool,
) (scanStats, error) {
roots = pruneRoots(roots)
s := &scanState{db: db}
err := s.runPhases(ctx, roots, workers, oneFS)
if err == nil || ctx.Err() == nil {
return s.st, err
}
// The one write made after the cancellation, so it cannot use ctx.
err = applyChanges(context.WithoutCancel(ctx), db, s.batch, nil, nil)
if err != nil {
return s.st, err
}
return s.st, ctx.Err()
}
// runPhases synchronizes the database with the filesystem under roots
// in four sequential phases: walk (enumerate and stat every file,
// building a complete size census), hash (read only the new or
// changed — or previously unhashed — files whose size at least one
// other file shares, committing results in batches as they arrive),
// update (record the size-unique files without reading them, and
// delete the records the scan no longer verifies), and content (fill
// in the content hash of every record of headTailMin or more whose
// size, head, and tail match another record's). Records outside the
// roots are never touched, except that the content phase fills in
// their content hash. Operands the walk cannot start from are dropped
// first, so the records beneath them count as outside the roots unless
// they lie under another root.
func (s *scanState) runPhases(ctx context.Context, roots []string,
workers int, oneFS bool,
) error {
// Types are checked before pruning so that an operand under a
// dropped one is still scanned, not dropped as lying under it.
roots = pruneRoots(s.walkableRoots(roots))
err := s.loadIndex(ctx, roots)
if err != nil {
return s.st, err
return err
}
changed, unhashed := s.walkPhase(startWalk(ctx, roots, oneFS, workers))
// A cancelled walk stops early, so its size census covers only part
// of the roots, and every file it never reached looks vanished to
// the update phase. Defence in depth rather than the only barrier:
// that phase would today fail on its first BeginTx with the same
// cancelled context before deleting anything. But it is the barrier
// that survives a later decision to let an interrupted scan commit
// what it has, and it turns a confusing failure deep in the update
// phase into a clean abort at the phase boundary.
// of the roots, and every file it never reached would look vanished
// to the update phase. Stop before anything is written or deleted.
err = ctx.Err()
if err != nil {
return s.st, err
return err
}
s.partition(changed, unhashed)
err = s.hashPhase(ctx, workers)
if err != nil {
return s.st, err
return err
}
return s.st, s.updatePhase(ctx)
err = s.updatePhase(ctx)
if err != nil {
return err
}
return s.contentPhase(ctx, workers)
}
// walkableRoots returns the operands the walk can start from: regular
// files, and directories not named .zfs. Every other operand is warned
// about, counted as skipped, and dropped. A dropped operand is no
// longer a root, so the records stored beneath it count as outside the
// roots and are not deleted as unverified, unless it lies under another
// root. An operand that fails lstat here is kept, and the walk warns
// about it.
func (s *scanState) walkableRoots(roots []string) []string {
kept := make([]string, 0, len(roots))
for _, root := range roots {
fi, err := os.Lstat(root)
if err == nil {
warn := operandWarning(root, fi)
if warn != "" {
s.st.skipped++
fmt.Fprintln(os.Stderr, escapePath(warn))
continue
}
}
kept = append(kept, root)
}
return kept
}
// loadIndex indexes the database records under the scan roots for
@@ -224,7 +370,7 @@ func (s *scanState) loadIndex(ctx context.Context, roots []string) error {
s.existing = make(map[string]fileMeta)
return loadFileMeta(ctx, s.db,
func(path string, size, mtime int64, hashed bool) {
func(path string, size int64, mtime time.Time, hashed bool) {
prog.increment()
if underAnyRoot(path, roots) {
@@ -241,7 +387,8 @@ func (s *scanState) loadIndex(ctx context.Context, roots []string) error {
// walkPhase drains the walk, appending every walked file's size to
// the census and resolving what it can immediately: an unchanged file
// whose record already has hashes needs nothing further. It returns
// whose record already has hashes needs nothing from the hash phase
// (the content phase may still fill in its content hash). It returns
// the new-or-changed files and the unchanged files whose records lack
// hashes; both remain candidates until the census decides whether
// their sizes are shared.
@@ -262,13 +409,14 @@ func (s *scanState) walkPhase(
}
s.sizes = append(s.sizes, ev.rec.size)
s.st.walked++
prog.increment()
old, ok := s.existing[ev.rec.path]
switch {
case !ok || old.size != ev.rec.size || old.mtime < ev.rec.mtime:
case !ok || old.size != ev.rec.size || mtimeAfter(ev.rec.mtime, old.mtime):
changed = append(changed, ev.rec)
case old.hashed:
delete(s.existing, ev.rec.path)
@@ -380,27 +528,40 @@ func sameInode(a, b fileRec) bool {
return (a.dev != 0 || a.ino != 0) && a.dev == b.dev && a.ino == b.ino
}
// hashPhase hashes every queued file with the worker pool — one read
// per inode run, in inode order — committing completed records to the
// database in batches as results arrive, so a long scan persists its
// progress as it goes (an interrupted scan resumes cheaply: the next
// run skips everything already recorded). The total counts actual
// reads, so the bar shows a real ETA. A run that fails to hash is
// warned about and skipped; stale records for its paths, if any, are
// deleted by the update phase.
// hashPhase hashes every queued file with hashSignature — the head and
// tail of a file of headTailMin or more, the whole file below that —
// committing completed records to the database in batches as results
// arrive, so a long scan persists its progress as it goes (an
// interrupted scan resumes cheaply: the next run skips everything
// already recorded). A run that fails to hash is warned about and
// skipped; stale records for its paths, if any, are deleted by the
// update phase.
func (s *scanState) hashPhase(ctx context.Context, workers int) error {
runs := hashRuns(s.toHash)
s.toHash = nil
return s.readRuns(ctx, workers, "hash", runs, hashSignature, s.recordRun)
}
// readRuns reads runs with the worker pool, one read per inode run, in
// the order given, under a progress display named label. The workers
// compute each run's hashes with hash, and each result goes to record;
// a run that fails to read is warned about and counted as skipped
// instead. The total counts actual reads, so the bar shows a real ETA.
//
// Returning early — a failed database write, or a cancelled scan — must
// not strand the pool: the feeder would park forever on a full jobs
// channel and every worker on a full results channel. The deferred stop
// is what prevents that.
func (s *scanState) hashPhase(ctx context.Context, workers int) error {
runs := hashRuns(s.toHash)
s.toHash = nil
pool := startHashPool(ctx, runs, workers)
func (s *scanState) readRuns(ctx context.Context, workers int,
label string, runs [][]fileRec,
hash func(path string, size int64) (string, string, string, error),
record func(ctx context.Context, r hashResult) error,
) error {
pool := startHashPool(ctx, runs, workers, hash)
defer pool.stop()
prog := newProgress("hash", int64(len(runs)))
prog := newProgress(label, int64(len(runs)))
defer prog.finish()
for range runs {
@@ -417,12 +578,12 @@ func (s *scanState) hashPhase(ctx context.Context, workers int) error {
if r.err != nil {
s.st.skipped += len(r.run)
prog.warnf("hash %s: %v", r.run[0].path, r.err)
prog.warnf("%s %s: %v", label, r.run[0].path, r.err)
continue
}
err := s.recordRun(ctx, r)
err := record(ctx, r)
if err != nil {
return err
}
@@ -443,18 +604,31 @@ func (s *scanState) recordRun(ctx context.Context, r hashResult) error {
mtime: rec.mtime,
head: r.head,
tail: r.tail,
content: r.content,
path: rec.path,
})
}
return s.commitFullBatch(ctx)
}
// commitFullBatch commits the running batch once it holds
// updateBatchSize records. A batch that fails to commit is kept: the
// commit fails when the scan is interrupted, and syncScan then commits
// the batch itself.
func (s *scanState) commitFullBatch(ctx context.Context) error {
if len(s.batch) < updateBatchSize {
return nil
}
err := applyBatch(ctx, s.db, s.batch, nil, nil)
if err != nil {
return err
}
s.batch = s.batch[:0]
return err
return nil
}
// updatePhase writes the scan's tail under one progress display: the
@@ -504,6 +678,157 @@ func (s *scanState) updatePhase(ctx context.Context) error {
return applyChanges(ctx, s.db, nil, deletes, prog)
}
// contentPhase fills in the content hash of every record of headTailMin
// or more that lacks one and whose size, head, and tail equal another
// record's, anywhere in the database: records from this scan and
// records stored by earlier scans, inside or outside the roots. Only
// such a file can still be a duplicate, so no other file of headTailMin
// or more is read beyond its end windows. The files are read with the
// hash phase's worker pool and their records written back in batches. A
// failed read is warned about and counted as skipped; the record keeps
// its empty content, so it is never grouped, and a later scan tries
// again.
func (s *scanState) contentPhase(ctx context.Context, workers int) error {
toRead, recs, err := s.contentCandidates(ctx)
if err != nil {
return err
}
err = s.readRuns(ctx, workers, "content", hashRuns(toRead),
hashContentOnly, func(ctx context.Context, r hashResult) error {
// Every path in the run keeps its record's head and tail
// and gains the one content hash read for the run.
for _, f := range r.run {
rec := recs[f.path]
rec.content = r.content
s.batch = append(s.batch, rec)
}
return s.commitFullBatch(ctx)
})
if err != nil {
return err
}
return applyChanges(ctx, s.db, s.batch, nil, nil)
}
// contentCandidates returns the files the content phase reads, and
// their records by path. Every record contentCandidatesSQL returns has
// its file checked with lstat, whether or not it already has a content
// hash: a file that is gone, is no longer a regular file, or has
// changed by the walk's rule keeps its record as it is and does not
// count as a match for the others, and any other lstat error is warned
// about and counted as skipped, with the same result. If such a record
// has no content hash, it stays out of duplicate groups; if it has one,
// it is still reported until a scan covering its own tree updates or
// removes it. The files of a group that pass and have no content hash
// are read only if at least minGroupSize of the group's files pass, so
// a group whose other members are all stale costs no reads. Only the
// records to be read are kept.
func (s *scanState) contentCandidates(
ctx context.Context,
) ([]fileRec, map[string]scanRec, error) {
// The query and the checks take real time on a large database;
// without a display the scan looks hung before the reads begin.
prog := newProgress("content", -1)
defer prog.finish()
var (
toRead []fileRec
first scanRec // the current group's first record
passed int // the current group's files that passed the check
unread []fileRec // those of them without a content hash
)
recs := make(map[string]scanRec)
// endGroup queues the current group's files to read if at least
// minGroupSize of its files passed, and drops their records if not.
endGroup := func() {
if passed >= minGroupSize {
toRead = append(toRead, unread...)
} else {
for _, f := range unread {
delete(recs, f.path)
}
}
passed, unread = 0, nil
}
err := loadContentCandidates(ctx, s.db, func(r scanRec, hashed bool) {
prog.increment()
if r.size != first.size || r.head != first.head || r.tail != first.tail {
endGroup()
first = r
}
f, ok, err := unchangedFile(r)
if err != nil {
s.st.skipped++
prog.warnf("content %s: %v", r.path, err)
}
if !ok {
return
}
passed++
if !hashed {
unread = append(unread, f)
recs[r.path] = r
}
})
if err != nil {
return nil, nil, err
}
endGroup()
return toRead, recs, nil
}
// unchangedFile lstats the file r names and returns it for reading if
// it is still the regular file r records: the same size, and an mtime
// no newer than recorded (the walk's change rule). A file that is gone
// or has changed reports false; any other lstat error is returned.
func unchangedFile(r scanRec) (fileRec, bool, error) {
fi, err := os.Lstat(r.path)
if errors.Is(err, fs.ErrNotExist) {
return fileRec{}, false, nil
}
if err != nil {
return fileRec{}, false, err
}
if !fi.Mode().IsRegular() || fi.Size() != r.size ||
mtimeAfter(fi.ModTime(), r.mtime) {
return fileRec{}, false, nil
}
dev, ino := inodeOfInfo(fi)
return fileRec{
path: r.path, size: r.size, mtime: r.mtime, dev: dev, ino: ino,
}, true, nil
}
// mtimeAfter reports whether mtime a is later than mtime b.
// Not a.After(b): time.Time wraps an mtime past year 292 billion; Unix() undoes it.
func mtimeAfter(a, b time.Time) bool {
if a.Unix() != b.Unix() {
return a.Unix() > b.Unix()
}
return a.Nanosecond() > b.Nanosecond()
}
// underAnyRoot reports whether path is any of the roots or lies under
// one of them.
func underAnyRoot(path string, roots []string) bool {
@@ -584,11 +909,43 @@ func sendEvent(ctx context.Context, events chan<- walkEvent,
}
}
// operandWarning returns the one-line warning for an operand the walk
// does not start from, naming the path and what it is, or "" for one it
// does: a regular file, or a directory not named .zfs. Symlinks are
// never followed, including as operands.
func operandWarning(root string, fi fs.FileInfo) string {
var kind string
switch mode := fi.Mode(); {
case mode.IsRegular():
return ""
case mode.IsDir():
if filepath.Base(root) != ".zfs" {
return ""
}
kind = ".zfs directory"
case mode&fs.ModeSymlink != 0:
kind = "symlink"
case mode&fs.ModeSocket != 0:
kind = "socket"
case mode&fs.ModeNamedPipe != 0:
kind = "FIFO"
case mode&fs.ModeDevice != 0:
kind = "device node"
default:
kind = "non-regular file"
}
return fmt.Sprintf("walk %s: skipping %s operand", root, kind)
}
// seedRoot turns one PATH operand into the walk's starting state: a
// regular-file operand is statted and emitted directly, a directory
// operand becomes an initial job, and a symlink or other non-regular
// operand yields nothing (symlinks are never followed, including as
// operands).
// regular-file operand is statted and emitted directly, and a directory
// operand becomes an initial job. walkableRoots has already dropped
// every other operand. One that has changed into something else since
// is warned about and skipped here; it is still a root, so the records
// stored beneath it are deleted as unverified.
func seedRoot(ctx context.Context, root string,
events chan<- walkEvent,
) []dirJob {
@@ -602,30 +959,30 @@ func seedRoot(ctx context.Context, root string,
return nil
}
switch {
case fi.IsDir():
if filepath.Base(root) == ".zfs" {
warn := operandWarning(root, fi)
if warn != "" {
sendEvent(ctx, events, walkEvent{warn: warn, fail: true})
return nil
}
if fi.IsDir() {
dev, ok := deviceOfInfo(fi)
return []dirJob{{path: root, rootDev: dev, rootDevOK: ok}}
case fi.Mode().IsRegular():
}
dev, ino := inodeOfInfo(fi)
sendEvent(ctx, events, walkEvent{rec: fileRec{
path: root,
size: fi.Size(),
mtime: fi.ModTime().Unix(),
mtime: fi.ModTime(),
dev: dev,
ino: ino,
}})
return nil
default:
return nil
}
}
// startWalkWorkers starts the walk worker pool. Each worker processes
@@ -724,6 +1081,12 @@ func walkOneDir(ctx context.Context, job dirJob, oneFS bool,
var subs []dirJob
for _, e := range entries {
// A cancelled scan wants nothing more from this directory: stop
// rather than lstat the rest of a large one.
if ctx.Err() != nil {
return nil
}
p := filepath.Join(job.path, e.Name())
if e.IsDir() {
@@ -772,7 +1135,7 @@ func emitFile(ctx context.Context, p string, e fs.DirEntry,
sendEvent(ctx, events, walkEvent{rec: fileRec{
path: p,
size: info.Size(),
mtime: info.ModTime().Unix(),
mtime: info.ModTime(),
dev: dev,
ino: ino,
}})
@@ -834,12 +1197,15 @@ func inodeOfInfo(fi fs.FileInfo) (uint64, uint64) {
return statDev(st), st.Ino
}
// hashResult carries one inode run's head/tail hashes (or the error
// that prevented hashing it) from the hash workers to the hash phase.
// hashResult carries the hashes computed for one inode run (or the
// error that prevented computing them) from the pool's workers to the
// phase that started the pool: head, tail, and content from
// hashSignature, content alone from hashContentOnly.
type hashResult struct {
run []fileRec
head string
tail string
content string
err error
}
@@ -856,10 +1222,10 @@ type hashPool struct {
}
// startHashPool starts the feeder and the workers over runs. Workers
// hash each run's first path (all paths in a run are hard links to the
// same inode) and write one result per run.
func startHashPool(ctx context.Context, runs [][]fileRec,
workers int,
// hash each run's first path with hash (all paths in a run are hard
// links to the same inode) and write one result per run.
func startHashPool(ctx context.Context, runs [][]fileRec, workers int,
hash func(path string, size int64) (string, string, string, error),
) *hashPool {
ctx, cancel := context.WithCancel(ctx)
@@ -871,7 +1237,7 @@ func startHashPool(ctx context.Context, runs [][]fileRec,
wg.Go(func() { feedHashJobs(ctx, runs, jobs) })
for range workers {
wg.Go(func() { hashWorker(ctx, jobs, results) })
wg.Go(func() { hashWorker(ctx, jobs, results, hash) })
}
done := make(chan struct{})
@@ -917,24 +1283,25 @@ func feedHashJobs(ctx context.Context, runs [][]fileRec,
}
}
// hashWorker hashes one inode run at a time until jobs is closed or the
// scan is cancelled. A cancelled worker drops the runs still queued
// instead of stopping its reads of jobs: the range must run out for the
// pool to tear down, and reading a file nobody wants the hash of only
// delays that.
// hashWorker hashes one inode run at a time with hash until jobs is
// closed or the scan is cancelled. A cancelled worker drops the runs
// still queued instead of stopping its reads of jobs: the range must
// run out for the pool to tear down, and reading a file nobody wants
// the hash of only delays that.
func hashWorker(ctx context.Context, jobs <-chan []fileRec,
results chan<- hashResult,
hash func(path string, size int64) (string, string, string, error),
) {
for run := range jobs {
if ctx.Err() != nil {
continue
}
head, tail, err := hashHeadTail(run[0].path, run[0].size)
head, tail, content, err := hash(run[0].path, run[0].size)
select {
case results <- hashResult{
run: run, head: head, tail: tail, err: err,
run: run, head: head, tail: tail, content: content, err: err,
}:
case <-ctx.Done():
return
@@ -942,55 +1309,151 @@ func hashWorker(ctx context.Context, jobs <-chan []fileRec,
}
}
// emptyHash is the lowercase-hex SHA-256 of the empty input: the head
// and tail hash of every zero-length file.
// emptyHash is the lowercase-hex SHA-256 of the empty input: the head,
// tail, and content hash of every zero-length file.
const emptyHash = "e3b0c44298fc1c149afbf4c8996fb924" +
"27ae41e4649b934ca495991b7852b855"
// hashHeadTail returns the lowercase-hex SHA-256 of the first
// min(chunk, size) bytes and of the last min(chunk, size) bytes of the
// file at path. The two reads overlap when size < 2*chunk. size is the
// value recorded when the file was statted; a zero-length file's
// hashes are constant, so it is never even opened.
func hashHeadTail(path string, size int64) (string, string, error) {
// hashSignature computes the hashes the hash phase records for a file
// whose size is shared; with the file size they form its duplicate
// signature. A file below headTailMin is hashed in full and its
// whole-file SHA-256 is returned as head, tail, and content alike —
// that range takes no separate end-window step. For a file at or above
// headTailMin only the head and tail are computed, the SHA-256 of its
// first and last headTailWindow bytes, and content is returned empty:
// the content phase computes it with hashContentOnly once the file's
// size, head, and tail match another file's. Two files are duplicates
// only when all four agree; any mismatch means not a duplicate. size
// is the value recorded when the file was statted; a zero-length file
// has constant hashes and is never opened.
func hashSignature(path string, size int64) (string, string, string, error) {
if size == 0 {
return emptyHash, emptyHash, nil
return emptyHash, emptyHash, emptyHash, nil
}
//nolint:gosec // hashing operator-supplied paths is the tool's purpose
f, err := os.Open(path)
if err != nil {
return "", "", err
return "", "", "", err
}
defer func() { _ = f.Close() }()
n := min(int64(chunk), size)
// Below the threshold the whole file is hashed directly, with no
// end-window step: head and tail both carry the whole-file hash.
if size < int64(headTailMin) {
content, err := hashWhole(f, size)
if err != nil {
return "", "", "", err
}
buf := make([]byte, n)
return content, content, content, nil
}
_, err = f.ReadAt(buf, 0)
head, tail, err := hashEnds(f, size)
if err != nil {
return "", "", "", err
}
return head, tail, "", nil
}
// hashContentOnly returns the content hash of the file at path, which
// is at least headTailMin bytes: the content phase's read. head and
// tail are returned empty, because the content phase keeps the ones its
// records already hold.
func hashContentOnly(path string, size int64) (string, string, string, error) {
//nolint:gosec // hashing operator-supplied paths is the tool's purpose
f, err := os.Open(path)
if err != nil {
return "", "", "", err
}
defer func() { _ = f.Close() }()
content, err := hashContent(f, size)
return "", "", content, err
}
// hashEnds returns the SHA-256 of the first and last headTailWindow
// bytes of f. It is called only for files at least headTailMin, which
// is far larger than two windows, so the windows never overlap and both
// reads are always full.
func hashEnds(f *os.File, size int64) (string, string, error) {
buf := make([]byte, headTailWindow)
_, err := f.ReadAt(buf, 0)
if err != nil {
return "", "", err
}
h := sha256.Sum256(buf)
head := hex.EncodeToString(h[:])
// When the whole file fits in one chunk the tail window is exactly
// the bytes just read: reuse the head hash instead of issuing a
// second read for every small file.
if size <= int64(chunk) {
hh := hex.EncodeToString(h[:])
return hh, hh, nil
}
_, err = f.ReadAt(buf, size-n)
_, err = f.ReadAt(buf, size-int64(headTailWindow))
if err != nil {
return "", "", err
}
t := sha256.Sum256(buf)
return hex.EncodeToString(h[:]), hex.EncodeToString(t[:]), nil
return head, hex.EncodeToString(t[:]), nil
}
// hashContent returns the content-rung hash of f: the SHA-256 of the
// whole file when it is smaller than wholeFileMax, or of sampled
// windows when it is that size or larger.
func hashContent(f *os.File, size int64) (string, error) {
if size >= int64(wholeFileMax) {
return hashSamples(f, size)
}
return hashWhole(f, size)
}
// hashWhole returns the SHA-256 of the entire file. A SectionReader is
// used so the read is independent of the offset left by any end-window
// reads. Reading fewer than size bytes means the file shrank between
// the stat and the hash; that is an error rather than a hash of content
// that no longer matches the recorded size.
func hashWhole(f *os.File, size int64) (string, error) {
h := sha256.New()
n, err := io.Copy(h, io.NewSectionReader(f, 0, size))
if err != nil {
return "", err
}
if n != size {
return "", fmt.Errorf("read %d of %d bytes: %w", n, size,
io.ErrUnexpectedEOF)
}
return hex.EncodeToString(h.Sum(nil)), nil
}
// hashSamples feeds sampleWindow bytes at each gigabyte-aligned offset
// (0, sampleStride, 2*sampleStride, ... while inside the file), in
// order, into one hash, each window truncated at end of file. This is
// the probabilistic large-file rung: two files of equal size agreeing
// on every sample are reported as duplicates without every byte being
// read. Because size is part of the signature, files of different sizes
// never reach this comparison, so the sample boundaries always align.
func hashSamples(f *os.File, size int64) (string, error) {
h := sha256.New()
buf := make([]byte, sampleWindow)
for off := int64(0); off < size; off += int64(sampleStride) {
n := min(int64(sampleWindow), size-off)
_, err := f.ReadAt(buf[:n], off)
if err != nil {
return "", err
}
h.Write(buf[:n])
}
return hex.EncodeToString(h.Sum(nil)), nil
}
+1060 -76
View File
File diff suppressed because it is too large Load Diff
+66 -38
View File
@@ -2,19 +2,19 @@
# script/bootstrap: install all dependencies needed to build and develop
# this repo. Idempotent: every install is guarded by a check so already
# installed tools are skipped. Base tooling comes from nix, apt, brew,
# or apk (detected in that order); assumes nothing is present (not git,
# make, or go). The linter is NOT installed: golangci-lint runs via
# docker only (script/lint), pinned by image digest, so the only lint
# prerequisite is a working docker — which is warned about, not
# installed, because everything except linting works without it.
# or apk (detected in that order); assumes nothing is present. Node is
# used directly if installed; otherwise it is installed at a pinned
# version via nvm (installing nvm itself first, from a hash-verified
# release archive, never curl | sh).
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# yarn provides prettier, which formats Markdown. yarn is a tool, like
# node/git/make/go below; the reference that governs formatting output is
# prettier, pinned by yarn.lock's integrity hash and installed by
# `yarn install --frozen-lockfile`.
# Pinned versions, 2026-07-06
NODE_VERSION="22.17.0"
NVM_VERSION="0.40.3"
# sha256 of https://github.com/nvm-sh/nvm/archive/refs/tags/v0.40.3.tar.gz
NVM_SHA256="5f4d6aaa04a177dc93c985e31dbc411ab6b8c6e1e21d8015dbc1372625fcd1d0"
YARN_VERSION="1.22.22"
PKGMGR=""
@@ -49,6 +49,7 @@ pkg_install() {
case "$PKGMGR" in
nix) nix-env -iA "nixpkgs.$1" ;;
apt)
# Package lists may be empty (fresh images); refresh once per run.
if [ -z "$APT_UPDATED" ]; then
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get update
APT_UPDATED=1
@@ -64,55 +65,82 @@ missing() {
! command -v "$1" >/dev/null 2>&1
}
# verify_sha256 <file> <expected-hash>
verify_sha256() {
if command -v sha256sum >/dev/null 2>&1; then
actual="$(sha256sum "$1" | cut -d' ' -f1)"
else
actual="$(shasum -a 256 "$1" | cut -d' ' -f1)"
fi
if [ "$actual" != "$2" ]; then
echo "bootstrap: sha256 mismatch for $1" >&2
echo " expected: $2" >&2
echo " actual: $actual" >&2
exit 1
fi
}
# nvm is a bash script; run a command in a bash with nvm loaded
nvm_sh() {
bash -c ". \"\$HOME/.nvm/nvm.sh\" && $*"
}
ensure_nvm() {
[ -s "$HOME/.nvm/nvm.sh" ] && return 0
# nvm prerequisites; nvm itself requires bash
if missing bash; then pkg_install bash bash bash bash; fi
if missing curl; then pkg_install curl curl curl curl; fi
if missing git; then pkg_install git git git git; fi
tmp="$(mktemp -d)"
curl -fsSL -o "$tmp/nvm.tar.gz" \
"https://github.com/nvm-sh/nvm/archive/refs/tags/v${NVM_VERSION}.tar.gz"
verify_sha256 "$tmp/nvm.tar.gz" "$NVM_SHA256"
mkdir -p "$HOME/.nvm"
tar -xzf "$tmp/nvm.tar.gz" -C "$HOME/.nvm" --strip-components=1
rm -rf "$tmp"
}
ensure_node() {
if ! missing node; then return 0; fi
pkg_install nodejs nodejs node nodejs
ensure_nvm
nvm_sh "nvm install $NODE_VERSION"
}
ensure_yarn() {
if ! missing yarn; then return 0; fi
if ! missing corepack; then
corepack enable >/dev/null 2>&1 || true
corepack enable
corepack prepare "yarn@$YARN_VERSION" --activate
elif [ -s "$HOME/.nvm/nvm.sh" ]; then
nvm_sh "nvm use $NODE_VERSION >/dev/null && corepack enable && \
corepack prepare yarn@$YARN_VERSION --activate"
else
pkg_install yarn yarn yarn yarn
npm install -g "yarn@$YARN_VERSION"
fi
}
install_js_deps() {
if missing yarn && [ -s "$HOME/.nvm/nvm.sh" ]; then
nvm_sh "nvm use $NODE_VERSION >/dev/null && cd \"$ROOT\" && \
yarn install --frozen-lockfile"
else
yarn install --frozen-lockfile
fi
}
main() {
cd "$ROOT"
# System tooling, deliberately unpinned: these come from the host
# package manager and whatever version it ships is what the host
# gets, so a presence check is the right check. The repo pins no
# system toolchain versions — the Go language version is governed by
# go.mod, and builds that must be reproducible run in the Docker
# image, whose base images are pinned by digest.
if missing git; then pkg_install git git git git; fi
if missing make; then pkg_install gnumake make make make; fi
if missing git; then pkg_install git git git git; fi
# Go builds the binary and runs gofmt. Presence is the whole check:
# go.mod names the Go version, and the tests and the linter run in
# digest-pinned images.
if missing go; then pkg_install go golang go go; fi
# node runs prettier and is an unpinned host tool for the same reason
# git/make/go are: it comes from the host package manager, whatever
# version it ships. It is not installed via nvm the way the canonical
# template does, because nvm's prebuilt node is glibc-linked and does
# not run on this repo's musl/Alpine build image. prettier — the tool
# whose version affects formatting output — is pinned by yarn.lock.
ensure_node
ensure_yarn
yarn install --frozen-lockfile
# Linting runs via docker only (script/lint), so docker is a lint
# prerequisite rather than something bootstrap installs. Warn, do
# not fail: everything except `make lint` — and, through it,
# `make check`, `make docker` and the pre-commit hook — works
# without it.
if missing docker; then
echo "bootstrap: WARNING: docker not found; make lint, make check" >&2
echo "bootstrap: and make docker require it. Install docker to" >&2
echo "bootstrap: run the linter." >&2
fi
install_js_deps
go mod download
echo "bootstrap complete"
+3 -1
View File
@@ -1,6 +1,8 @@
#!/bin/sh
# script/check: run all checks (test, lint, fmt-check). Our own
# extension to scripts-to-rule-them-all. Must not modify any files.
# extension to scripts-to-rule-them-all. test and lint are Docker
# phases; fmt-check is native, because a formatter writes the working
# tree. Must not modify any files.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
+21 -25
View File
@@ -1,34 +1,30 @@
#!/bin/sh
# script/cibuild: run the CI build. The Gitea workflow runs this on
# push.
#
# The Dockerfile runs the gates individually as build steps, not the
# make check aggregate: the lint stage runs make fmt-check,
# script/verify-lint-image-pin, golangci-lint config verify and
# golangci-lint run; the build stage, dropped to an unprivileged user,
# runs make test and make fmt-check. Neither make lint nor make check
# appears, because both reach script/lint, which is itself a docker
# build, and a docker build cannot run inside one. Lint is not skipped
# by that — the linter is invoked directly in the lint stage, and the
# build stage's COPY --from=lint makes that stage a prerequisite, so
# BuildKit must finish it first. Between the two stages everything
# make check would run has run, which is why a successful build here
# implies the repo is green.
#
# That implication holds only because of CHECK_EPOCH. A COPY layer is
# invalidated by changed content, and a merge commit's tree is
# byte-identical to the branch head it merges, so without a fresh value
# here Docker serves the gate layers from cache and the build reports a
# green it never earned. Passing the current epoch invalidates the gate
# layers on every run while leaving the pinned base images and
# go mod download cached; see the Dockerfile for the placement.
# script/cibuild: run the CI build. It bootstraps first: a CI runner
# checks out and runs this and nothing else, and script/fmt-check runs
# the formatter on the host, which a pristine checkout cannot do.
# --no-cache for the same reason as script/docker: the gate phases the
# final stage depends on are RUN steps, and a cached one is a check that
# did not run.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
docker build --build-arg CHECK_EPOCH="$(date +%s)" .
"$SCRIPT_DIR/bootstrap"
"$SCRIPT_DIR/check"
# The version and the tag each get their own line: a failing
# command substitution inside an argument does not trip `set -e`,
# so the inline form degrades silently to an empty constant. The
# VERSION build argument takes precedence over the version a build
# stage derives from the .git in the context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
tag="$("$SCRIPT_DIR/projectname")"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$tag" .
}
main "$@"
+14 -13
View File
@@ -1,14 +1,8 @@
#!/bin/sh
# script/docker: build the Docker image tagged with the project name.
# The tag comes from script/projectname.
#
# CHECK_EPOCH is passed for the same reason script/cibuild passes it:
# without it Docker serves the Dockerfile's gate layers from cache on an
# unchanged tree and this exits 0 having run neither the lint stage's
# gates nor the builder stage's test and fmt-check gates. This is the
# set of gates a developer or reviewer runs by hand, so a cached pass
# here is the most misleading result the repo can produce. Dependency
# layers sit above the ARG and stay cached.
# Identical in all repos; the tag comes from script/projectname.
# --no-cache because the gate phases the final stage depends on are RUN
# steps, and a cached one is a check that did not run.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
@@ -16,10 +10,17 @@ ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
docker build \
--build-arg CHECK_EPOCH="$(date +%s)" \
-t "$("$SCRIPT_DIR/projectname")" \
.
# The version and the tag each get their own line: a failing
# command substitution inside an argument does not trip `set -e`,
# so the inline form degrades silently to an empty constant. The
# VERSION build argument takes precedence over the version a build
# stage derives from the .git in the context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
tag="$("$SCRIPT_DIR/projectname")"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$tag" .
}
main "$@"
+17 -8
View File
@@ -1,23 +1,32 @@
#!/bin/sh
# script/fmt: format all files (writes). gofmt for Go, prettier for
# Markdown. prettier is the pinned devDependency in package.json/
# yarn.lock; script/bootstrap installs it (see run_prettier).
# script/fmt: format all files (writes).
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
run_prettier() {
if ! command -v yarn >/dev/null 2>&1; then
echo "fmt: yarn not found; run script/bootstrap first" >&2
# Must match the pin in script/bootstrap.
NODE_VERSION="22.17.0"
# script/bootstrap installs node and yarn under nvm and leaves neither
# on the PATH of the shell that called it, so resolve the pinned
# toolchain here the way bootstrap's own install step does. nvm is a
# bash script, hence the subshell.
run_yarn() {
if command -v yarn >/dev/null 2>&1; then
exec yarn "$@"
fi
if [ ! -s "$HOME/.nvm/nvm.sh" ]; then
echo "fmt: no yarn; run script/bootstrap first" >&2
exit 1
fi
yarn run prettier "$@"
exec bash -c '. "$HOME/.nvm/nvm.sh" && nvm use "$1" >/dev/null &&
shift && exec yarn "$@"' bash "$NODE_VERSION" "$@"
}
main() {
cd "$ROOT"
gofmt -s -w .
run_prettier --write '**/*.md' --tab-width 4 --prose-wrap always
run_yarn run prettier --write '**/*.md' --tab-width 4 --prose-wrap always
}
main "$@"
+29 -16
View File
@@ -1,37 +1,50 @@
#!/bin/sh
# script/fmt-check: check formatting (read-only). Same scope as
# script/fmt: gofmt for Go, prettier for Markdown. Both run every time
# and each reports independently, so a failure names which formatter is
# unhappy; the script exits non-zero if either found unformatted files.
# script/fmt-check: check formatting (read-only).
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
run_prettier() {
if ! command -v yarn >/dev/null 2>&1; then
echo "fmt-check: yarn not found; run script/bootstrap first" >&2
# Must match the pin in script/bootstrap.
NODE_VERSION="22.17.0"
# script/bootstrap installs node and yarn under nvm and leaves neither
# on the PATH of the shell that called it, so resolve the pinned
# toolchain here the way bootstrap's own install step does. nvm is a
# bash script, hence the subshell.
run_yarn() {
if command -v yarn >/dev/null 2>&1; then
exec yarn "$@"
fi
if [ ! -s "$HOME/.nvm/nvm.sh" ]; then
echo "fmt-check: no yarn; run script/bootstrap first" >&2
exit 1
fi
yarn run prettier "$@"
exec bash -c '. "$HOME/.nvm/nvm.sh" && nvm use "$1" >/dev/null &&
shift && exec yarn "$@"' bash "$NODE_VERSION" "$@"
}
main() {
cd "$ROOT"
rc=0
status=0
files="$(gofmt -s -l .)"
# gofmt and prettier both run every time, so the output names each
# one that fails. Under set -e a bare assignment would end the
# script when gofmt fails (a Go file it cannot parse).
if ! files="$(gofmt -s -l .)"; then
echo "gofmt: failed; see its errors above" >&2
status=1
fi
if [ -n "$files" ]; then
echo "gofmt: files not formatted:" >&2
echo "$files" >&2
rc=1
status=1
fi
if ! run_prettier --check '**/*.md' --tab-width 4 --prose-wrap always; then
echo "prettier: Markdown not formatted; run make fmt" >&2
rc=1
fi
# run_yarn ends in exec; the subshell returns here afterwards.
(run_yarn run prettier --check '**/*.md' --tab-width 4 --prose-wrap always) ||
status=1
exit "$rc"
exit "$status"
}
main "$@"
+3 -7
View File
@@ -1,19 +1,15 @@
#!/bin/sh
# script/install-precommit: install the git pre-commit hook that runs
# script/precommit. Our own extension to scripts-to-rule-them-all.
# Hooks are shared between the main checkout and all worktrees, so
# resolve the common git dir instead of assuming .git is a directory.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
hooks_dir="$(git rev-parse --git-common-dir)/hooks"
mkdir -p "$hooks_dir"
hook="$hooks_dir/pre-commit"
printf '#!/bin/sh\nset -e\nscript/precommit\n' > "$hook"
chmod +x "$hook"
hook=".git/hooks/pre-commit"
printf '#!/bin/sh\nset -e\nscript/precommit\n' > .git/hooks/pre-commit
chmod +x .git/hooks/pre-commit
echo "pre-commit hook installed: runs script/precommit"
}
+13 -19
View File
@@ -1,29 +1,23 @@
#!/bin/sh
# script/lint: run the linter. golangci-lint is never installed on a
# host: it runs via docker only, one way, everywhere — this builds
# Dockerfile.lint, which COPYs the repo into the digest-pinned
# golangci-lint image and lints as a build step, so a successful build
# is a clean lint. The only prerequisite is a working docker. The gate
# steps make no network calls of their own, but Dockerfile.lint runs
# `go mod download` above them, so a cold cache does reach the network
# (as does pulling the pinned image); that layer stays cached, and once
# it is warm this runs offline until go.mod or go.sum changes.
# script/lint: run the linter. Linting is a phase of the Dockerfile and
# this builds that phase alone; the linter is never installed or run on
# a developer host, where a shared result cache and a host-global lock
# make its answer untrustworthy.
#
# CHECK_EPOCH is what makes the result mean anything. Without it docker
# serves the gate layers from cache on an unchanged tree and this exits
# 0 in well under a second having run no linter. The PID is in the value
# as well as the epoch because two lint runs land inside the same second
# easily, and `date +%s` alone would cache the second one.
# The phase is not the last stage in the file, so it is built only when
# --target names it. --no-cache because a cached lint layer is a lint
# that did not run. --output type=cacheonly writes no image, since
# nothing uses one.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
docker build \
--build-arg CHECK_EPOCH="$(date +%s)-$$" \
-f Dockerfile.lint \
.
docker build --no-cache \
--target lint \
--output type=cacheonly .
}
main "$@"
+2 -2
View File
@@ -1,13 +1,13 @@
#!/bin/sh
# script/precommit: run by the git pre-commit hook; fails the commit if
# checks fail. Our own extension to scripts-to-rule-them-all. Go extra:
# go mod tidy must be a no-op before the checks run.
# checks fail. Our own extension to scripts-to-rule-them-all.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
# Go extra: go mod tidy must be a no-op before the checks run.
cd "$ROOT"
go mod tidy
if ! git diff --exit-code -- go.mod go.sum; then
+1 -1
View File
@@ -1,6 +1,6 @@
#!/bin/sh
# script/setup: set up the repo for development after a fresh clone:
# installs dependencies (script/bootstrap) and the git pre-commit hook.
# installs dependencies and the git pre-commit hook.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
+10 -8
View File
@@ -1,17 +1,19 @@
#!/bin/sh
# script/test: run the test suite. Reruns verbosely on failure so CI
# logs show which test failed.
# script/test: run the test suite. Testing is a phase of the Dockerfile
# and this builds that phase alone, on the same terms as script/lint:
# --target because a phase that is not the last stage is built only when
# named, and --no-cache because a cached test layer is a test that did
# not run. --output type=cacheonly writes no image, since nothing uses one.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
go test -timeout 30s -cover ./... || {
echo "--- Rerunning with -v for details ---"
go test -timeout 30s -v ./...
exit 1
}
docker build --no-cache \
--target test \
--output type=cacheonly .
}
main "$@"
-84
View File
@@ -1,84 +0,0 @@
#!/bin/sh
# script/verify-lint-image-pin: fail unless the golangci-lint image
# referenced by Dockerfile.lint and the one referenced by the main
# Dockerfile's lint stage are the same image at the same digest. Our own
# extension to scripts-to-rule-them-all, not one of its entrypoints.
#
# The linter version is pinned in two independent files. That is the
# shape #42 turned into a build failure rather than tolerate: nothing
# else keeps the two in sync, and a bump applied to one file alone would
# leave `make lint` and the fail-fast lint stage of `make docker`
# linting the same tree against different rulesets, both green. This is
# the single guard that stops it, run as a gate in both files.
#
# It deliberately restates neither pin. A hardcoded expected digest here
# would be a third copy — one more thing to bump, and the same drift one
# file further out. It compares the two files to each other and knows
# nothing about which version is correct.
#
# A reference that cannot be read is a hard failure, not a skip: a
# comparison of two empty strings succeeds, which would turn this guard
# into exactly the unearned green it exists to prevent.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
LINT_DOCKERFILE="Dockerfile.lint"
MAIN_DOCKERFILE="Dockerfile"
# Echo the single golangci-lint image reference in the named Dockerfile.
# Scans every argument of every FROM instruction rather than assuming a
# field position, so `FROM --platform=... img AS stage` reads correctly.
# Exits non-zero, with a diagnosis, unless there is exactly one.
lint_image_ref() {
file="$1"
if [ ! -f "$file" ]; then
echo "verify-lint-image-pin: $file: not found" >&2
return 1
fi
refs="$(
awk '
toupper($1) == "FROM" {
for (i = 2; i <= NF; i++) {
if ($i ~ /^golangci\/golangci-lint[:@]/) {
print $i
}
}
}
' "$file"
)"
count="$(printf '%s' "$refs" | grep -c . || true)"
if [ "$count" -ne 1 ]; then
echo "verify-lint-image-pin: $file: expected exactly one" \
"golangci/golangci-lint FROM reference, found $count" >&2
return 1
fi
printf '%s\n' "$refs"
}
main() {
cd "$ROOT"
lint_ref="$(lint_image_ref "$LINT_DOCKERFILE")"
main_ref="$(lint_image_ref "$MAIN_DOCKERFILE")"
if [ "$lint_ref" != "$main_ref" ]; then
echo "verify-lint-image-pin: the linter image is pinned twice and" \
"the two pins disagree:" >&2
echo "verify-lint-image-pin: $LINT_DOCKERFILE: $lint_ref" >&2
echo "verify-lint-image-pin: $MAIN_DOCKERFILE: $main_ref" >&2
echo "verify-lint-image-pin: bump both FROM lines together so" \
"script/lint and the Dockerfile lint stage keep running the" \
"same linter" >&2
exit 1
fi
echo "verify-lint-image-pin: $LINT_DOCKERFILE and $MAIN_DOCKERFILE" \
"agree on $lint_ref"
}
main "$@"
+147 -86
View File
@@ -5,48 +5,59 @@ import (
"context"
"crypto/sha256"
"fmt"
"io"
"os"
"slices"
"strconv"
"strings"
)
// fileSig is a file's duplicate signature; mtime is excluded.
type fileSig struct {
size int64
head string
tail string
}
// treeNode is one directory reconstructed from the scan stream.
// treeNode is one directory reconstructed from the record paths.
type treeNode struct {
path string
parent *treeNode
dirs map[string]*treeNode
files map[string]fileSig
// entries holds the serialized child entries until the digest is
// computed from them, and is then dropped.
entries []string
digest [sha256.Size]byte
fileCount int64
totalSize int64
}
// runTrees implements the trees subcommand: it reads every record from
// the database, reconstructs the directory hierarchy from the record
// paths, computes a Merkle-style digest per directory, and prints
// maximal duplicate-tree groups as TSV on stdout. It never touches the
// scanned filesystem; its only I/O is the database, stdout, and
// stderr.
func runTrees(ctx context.Context) error {
recs, err := loadRecords(ctx)
// the database in path order, reconstructs the directory hierarchy from
// the record paths, computes a Merkle-style digest per directory, and
// prints maximal duplicate-tree groups as TSV on stdout. It never
// touches the scanned filesystem; its only I/O is the database, stdout,
// and stderr. Any database problem, including a missing database, is
// fatal.
func runTrees(ctx context.Context, stdout io.Writer) error {
dbPath := databasePath()
db, err := openReportDatabase(ctx, dbPath)
if err != nil {
return err
}
super, allDirs := buildHierarchy(recs)
super.compute()
defer func() { _ = db.Close() }()
records := 0
tree := newTreeBuilder()
err = loadFileRows(ctx, db, func(r scanRec) {
records++
tree.add(r)
})
if err != nil {
return fmt.Errorf("database %s: %w", dbPath, err)
}
super, allDirs := tree.finish()
dupes := collectTreeGroups(allDirs, super)
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
out := bufio.NewWriterSize(stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
if err != nil {
@@ -61,7 +72,8 @@ func runTrees(ctx context.Context) error {
first := g[0]
for _, n := range g[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n",
first.path, n.path, first.fileCount, first.totalSize)
escapePath(first.path), escapePath(n.path),
first.fileCount, first.totalSize)
if err != nil {
return fmt.Errorf("write stdout: %w", err)
}
@@ -79,63 +91,128 @@ func runTrees(ctx context.Context) error {
fmt.Fprintf(os.Stderr,
"trees: %d records read, %d duplicate tree groups, %d dupe trees, "+
"%s reclaimable\n",
len(recs), len(dupes), dupeTrees, humanBytes(reclaimable))
records, len(dupes), dupeTrees, humanBytes(reclaimable))
return nil
}
// buildHierarchy reconstructs the directory hierarchy from the record
// paths under a synthetic super-root. Paths are split on "/"; for
// absolute paths the first component is empty, which simply becomes a
// top-level node representing "/". It returns the super-root and every
// directory node created.
func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
// treeBuilder reconstructs the directory hierarchy from records added
// in path order, under a synthetic super-root. Paths are split on "/";
// for absolute paths the first component is empty, which becomes the
// top-level directory with path "/". In path order all the paths under
// one directory come together, so a directory is complete once a path
// outside it is added: its digest is computed then and its entries are
// dropped. Only the directories holding the latest path keep entries.
type treeBuilder struct {
super *treeNode
// open lists the directories holding the latest path, outermost
// first, starting with the super-root; names[i] is open[i]'s name.
open []*treeNode
names []string
// dirs lists every completed directory.
dirs []*treeNode
}
func newTreeBuilder() *treeBuilder {
super := &treeNode{}
var allDirs []*treeNode
return &treeBuilder{
super: super,
open: []*treeNode{super},
names: []string{""},
}
}
for _, r := range recs {
// add adds one record. Each record must come after the previous one in
// path order (byte order); otherwise a completed directory would be
// started again as a second directory with the same path.
func (b *treeBuilder) add(r scanRec) {
comps := strings.Split(r.path, "/")
dirNames, name := comps[:len(comps)-1], comps[len(comps)-1]
node := super
for _, c := range comps[:len(comps)-1] {
child := node.dirs[c]
if child == nil {
childPath := c
if node != super {
childPath = node.path + "/" + c
// Keep the open directories that hold this path; complete the rest.
depth := 1
for depth < len(b.open) && depth <= len(dirNames) &&
b.names[depth] == dirNames[depth-1] {
depth++
}
child = &treeNode{path: childPath, parent: node}
if node.dirs == nil {
node.dirs = make(map[string]*treeNode)
b.closeTo(depth)
for _, c := range dirNames[depth-1:] {
b.openDir(c)
}
node.dirs[c] = child
allDirs = append(allDirs, child)
dir := b.open[len(b.open)-1]
dir.entries = append(dir.entries, fileEntry(name, r))
dir.fileCount++
dir.totalSize += r.size
}
// openDir starts the directory called name inside the innermost open
// one.
func (b *treeBuilder) openDir(name string) {
parent := b.open[len(b.open)-1]
path := parent.path + "/" + name
// The root directory's path is "/", not empty, and its children's
// paths start with one slash, not two.
switch {
case parent == b.super && name == "":
path = "/"
case parent == b.super:
path = name
case parent.path == "/":
path = "/" + name
}
node = child
b.open = append(b.open, &treeNode{path: path, parent: parent})
b.names = append(b.names, name)
}
// closeTo completes the open directories after the first n, innermost
// first: each one's digest is computed and entered in its parent along
// with its totals.
func (b *treeBuilder) closeTo(n int) {
for len(b.open) > n {
last := len(b.open) - 1
dir, name := b.open[last], b.names[last]
b.open, b.names = b.open[:last], b.names[:last]
dir.computeDigest()
dir.parent.entries = append(dir.parent.entries,
"d\x00"+name+"\x00"+string(dir.digest[:]))
dir.parent.fileCount += dir.fileCount
dir.parent.totalSize += dir.totalSize
b.dirs = append(b.dirs, dir)
}
}
// finish completes every open directory and returns the super-root and
// every directory.
func (b *treeBuilder) finish() (*treeNode, []*treeNode) {
b.closeTo(1)
return b.super, b.dirs
}
// fileEntry serializes a file child for its directory's digest: its
// name and its signature (size, head, tail, content); mtime is
// excluded.
func fileEntry(name string, r scanRec) string {
content := r.content
// A record without a content hash has unknown content (README
// "Database"): give it a signature no other file can share, so
// trees containing it never compare equal. Real hashes are hex, so
// the NUL-prefixed form cannot collide.
if content == "" {
content = "unhashed\x00" + r.path
}
if node.files == nil {
node.files = make(map[string]fileSig)
}
sig := fileSig{size: r.size, head: r.head, tail: r.tail}
// An unhashed record (its size was unique when last scanned)
// has unknown content: give it a signature no other file can
// share, so trees containing it never compare equal. Real
// heads are hex, so the NUL-prefixed form cannot collide.
if sig.head == "" {
sig.head = "unhashed\x00" + r.path
}
node.files[comps[len(comps)-1]] = sig
}
return super, allDirs
return "f\x00" + name + "\x00" + strconv.FormatInt(r.size, 10) +
"\x00" + r.head + "\x00" + r.tail + "\x00" + content
}
// collectTreeGroups groups directories by digest and returns every
@@ -177,38 +254,22 @@ func collectTreeGroups(allDirs []*treeNode, super *treeNode) [][]*treeNode {
return dupes
}
// compute fills in digest, fileCount, and totalSize for n and all of
// its descendants. A directory's digest is the SHA-256 of its child
// entries — files serialized with name and signature, subdirectories
// with name and recursive digest — sorted byte-lexicographically.
// Filenames cannot contain NUL or "/", so NUL delimiters are
// unambiguous.
func (n *treeNode) compute() {
entries := make([]string, 0, len(n.dirs)+len(n.files))
for name, sig := range n.files {
entries = append(entries,
"f\x00"+name+"\x00"+strconv.FormatInt(sig.size, 10)+
"\x00"+sig.head+"\x00"+sig.tail)
n.fileCount++
n.totalSize += sig.size
}
for name, child := range n.dirs {
child.compute()
entries = append(entries, "d\x00"+name+"\x00"+string(child.digest[:]))
n.fileCount += child.fileCount
n.totalSize += child.totalSize
}
slices.Sort(entries)
// computeDigest sets n's digest and drops its entries. A directory's
// digest is the SHA-256 of its child entries — files serialized with
// name and signature, subdirectories with name and recursive digest —
// sorted byte-lexicographically. Filenames cannot contain NUL or "/",
// so NUL delimiters are unambiguous.
func (n *treeNode) computeDigest() {
slices.Sort(n.entries)
h := sha256.New()
for _, e := range entries {
for _, e := range n.entries {
h.Write([]byte(e))
h.Write([]byte{0})
}
copy(n.digest[:], h.Sum(nil))
n.entries = nil
}
// suppressed reports whether a duplicate-tree group is non-maximal: its
+145 -30
View File
@@ -1,6 +1,8 @@
package main
import (
"bytes"
"database/sql"
"slices"
"testing"
)
@@ -9,23 +11,56 @@ import (
const (
f1Head = "f1h"
f1Tail = "f1t"
f1Content = "f1c"
f2Head = "f2h"
f2Tail = "f2t"
f2Content = "f2c"
)
// smokeTreeRecs mirrors the README smoke-test tree layout: /d/t1 and
// /d/t2 are identical, /d/t3 differs from them only by one filename.
func smokeTreeRecs() []scanRec {
return []scanRec{
{size: 3000, head: f1Head, tail: f1Tail, path: "/d/t1/f1"},
{size: 100, head: f2Head, tail: f2Tail, path: "/d/t1/sub/f2"},
{size: 3000, head: f1Head, tail: f1Tail, path: "/d/t2/f1"},
{size: 100, head: f2Head, tail: f2Tail, path: "/d/t2/sub/f2"},
{size: 3000, head: f1Head, tail: f1Tail, path: "/d/t3/f1"},
{size: 100, head: f2Head, tail: f2Tail, path: "/d/t3/sub/f2renamed"},
{size: 3000, head: f1Head, tail: f1Tail, content: f1Content, path: "/d/t1/f1"},
{size: 100, head: f2Head, tail: f2Tail, content: f2Content, path: "/d/t1/sub/f2"},
{size: 3000, head: f1Head, tail: f1Tail, content: f1Content, path: "/d/t2/f1"},
{size: 100, head: f2Head, tail: f2Tail, content: f2Content, path: "/d/t2/sub/f2"},
{size: 3000, head: f1Head, tail: f1Tail, content: f1Content, path: "/d/t3/f1"},
{size: 100, head: f2Head, tail: f2Tail, content: f2Content,
path: "/d/t3/sub/f2renamed"},
}
}
// dbTree builds the directory hierarchy from the records in db the way
// trees does, and returns the super-root and every directory.
func dbTree(t *testing.T, db *sql.DB) (*treeNode, []*treeNode) {
t.Helper()
tree := newTreeBuilder()
err := loadFileRows(t.Context(), db, tree.add)
if err != nil {
t.Fatal(err)
}
return tree.finish()
}
// treeOf writes recs into a fresh database and builds the directory
// hierarchy from it the way trees does.
func treeOf(t *testing.T, recs []scanRec) (*treeNode, []*treeNode) {
t.Helper()
db := openTestDB(t)
err := applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
return dbTree(t, db)
}
// nodeByPath finds the directory node with the given path.
func nodeByPath(t *testing.T, dirs []*treeNode, path string) *treeNode {
t.Helper()
@@ -56,11 +91,10 @@ func groupPaths(groups [][]*treeNode) [][]string {
return out
}
func TestBuildHierarchyCounts(t *testing.T) {
func TestTreeCounts(t *testing.T) {
t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs())
super.compute()
_, dirs := treeOf(t, smokeTreeRecs())
d := nodeByPath(t, dirs, "/d")
if d.fileCount != 6 || d.totalSize != 9300 {
@@ -81,11 +115,98 @@ func TestBuildHierarchyCounts(t *testing.T) {
}
}
func TestTreeRootPath(t *testing.T) {
t.Parallel()
// The root directory's path is "/", never empty, and its
// children's paths start with a single slash.
_, dirs := treeOf(t, []scanRec{{path: "/f"}, {path: "/srv/g"}})
got := make([]string, 0, len(dirs))
for _, d := range dirs {
got = append(got, d.path)
}
slices.Sort(got)
want := []string{"/", "/srv"}
if !slices.Equal(got, want) {
t.Fatalf("directory paths = %q, want %q", got, want)
}
}
func TestTreeNamesSortingBeforeSlash(t *testing.T) {
t.Parallel()
// In path order "/a/b-x/f" and "/a/b.txt" come between the file
// "/a/b" and "/a/b/f", because "-" and "." sort before "/". Each
// directory must still be built once, whole, so /a matches /c.
recs := make([]scanRec, 0, 8)
for _, top := range []string{"/a", "/c"} {
for _, p := range []string{"/b", "/b-x/f", "/b.txt", "/b/f"} {
content := "c"
if p == "/b-x/f" {
content = "other"
}
recs = append(recs, scanRec{
size: 1, head: "h", tail: "t", content: content, path: top + p,
})
}
}
super, dirs := treeOf(t, recs)
got := make([]string, 0, len(dirs))
for _, d := range dirs {
got = append(got, d.path)
}
slices.Sort(got)
want := []string{"/", "/a", "/a/b", "/a/b-x", "/c", "/c/b", "/c/b-x"}
if !slices.Equal(got, want) {
t.Fatalf("directory paths = %q, want %q", got, want)
}
groups := collectTreeGroups(dirs, super)
gotGroups := groupPaths(groups)
wantGroups := [][]string{{"/a", "/c"}}
if !slices.EqualFunc(gotGroups, wantGroups, slices.Equal) {
t.Fatalf("groups = %v, want %v", gotGroups, wantGroups)
}
if groups[0][0].fileCount != 4 || groups[0][0].totalSize != 4 {
t.Errorf("group totals: %d files %d bytes, want 4 4",
groups[0][0].fileCount, groups[0][0].totalSize)
}
}
func TestRunTreesEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
code := run([]string{cmdTrees}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tfiles\tsize\n" +
`/d/\tone\ntwo\rthree\\four` + "\t/d/A\t1\t5\n"
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
func TestTreeDigests(t *testing.T) {
t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs())
super.compute()
_, dirs := treeOf(t, smokeTreeRecs())
t1 := nodeByPath(t, dirs, "/d/t1")
t2 := nodeByPath(t, dirs, "/d/t2")
@@ -114,12 +235,11 @@ func TestTreeDigestContentSensitivity(t *testing.T) {
const sharedTail = "same"
recs := []scanRec{
{size: 10, head: sharedTail, tail: sharedTail, path: "/r/a/f"},
{size: 10, head: "DIFF", tail: sharedTail, path: "/r/b/f"},
{size: 10, head: sharedTail, tail: sharedTail, content: "c", path: "/r/a/f"},
{size: 10, head: "DIFF", tail: sharedTail, content: "c", path: "/r/b/f"},
}
super, dirs := buildHierarchy(recs)
super.compute()
_, dirs := treeOf(t, recs)
a := nodeByPath(t, dirs, "/r/a")
b := nodeByPath(t, dirs, "/r/b")
@@ -132,8 +252,7 @@ func TestTreeDigestContentSensitivity(t *testing.T) {
func TestCollectTreeGroupsMaximal(t *testing.T) {
t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs())
super.compute()
super, dirs := treeOf(t, smokeTreeRecs())
groups := collectTreeGroups(dirs, super)
@@ -157,16 +276,14 @@ func TestCollectTreeGroupsDeterministic(t *testing.T) {
recs := smokeTreeRecs()
super, dirs := buildHierarchy(recs)
super.compute()
super, dirs := treeOf(t, recs)
forward := groupPaths(collectTreeGroups(dirs, super))
reversed := slices.Clone(recs)
slices.Reverse(reversed)
superR, dirsR := buildHierarchy(reversed)
superR.compute()
superR, dirsR := treeOf(t, reversed)
backward := groupPaths(collectTreeGroups(dirsR, superR))
if !slices.EqualFunc(forward, backward, slices.Equal) {
@@ -181,12 +298,11 @@ func TestCollectTreeGroupsSiblings(t *testing.T) {
// Identical sibling dirs share a parent, so their group cannot be
// implied by a parent group and must be reported.
recs := []scanRec{
{size: 10, head: "h", tail: "t", path: "/p/x1/f"},
{size: 10, head: "h", tail: "t", path: "/p/x2/f"},
{size: 10, head: "h", tail: "t", content: "c", path: "/p/x1/f"},
{size: 10, head: "h", tail: "t", content: "c", path: "/p/x2/f"},
}
super, dirs := buildHierarchy(recs)
super.compute()
super, dirs := treeOf(t, recs)
got := groupPaths(collectTreeGroups(dirs, super))
@@ -203,13 +319,12 @@ func TestCollectTreeGroupsDifferingParents(t *testing.T) {
// extra file, so the parents' digests differ and the x group must
// be reported.
recs := []scanRec{
{size: 10, head: "h", tail: "t", path: "/p/a/x/f"},
{size: 99, head: "e", tail: "e", path: "/p/a/extra"},
{size: 10, head: "h", tail: "t", path: "/q/b/x/f"},
{size: 10, head: "h", tail: "t", content: "c", path: "/p/a/x/f"},
{size: 99, head: "e", tail: "e", content: "e", path: "/p/a/extra"},
{size: 10, head: "h", tail: "t", content: "c", path: "/q/b/x/f"},
}
super, dirs := buildHierarchy(recs)
super.compute()
super, dirs := treeOf(t, recs)
got := groupPaths(collectTreeGroups(dirs, super))