Author SHA1 Message Date
sneak 56fd779f7a Format Markdown with prettier in make fmt and make fmt-check (closes #19)
check / check (push) Waiting to run
script/fmt and script/fmt-check run prettier over every Markdown file
again, next to gofmt. prettier is pinned by hash through package.json
and yarn.lock, copied from the prompts repo with .prettierrc and
.prettierignore, and is never installed on a host: a new prettier stage
of the Dockerfile installs it into a digest-pinned node image, and both
scripts build that stage and run it with the repository mounted. CI
checks the Markdown in a markdown stage that the build stage waits on.
Because make fmt-check now runs docker, the Dockerfile runs gofmt
directly in its lint stage instead. All Markdown is reformatted.

Model: opus-5-5
2026-10-04 12:55:03 +00:00
clawbot c9c8b1d06c Reject scan --workers below 1 as a usage error (closes #10)
check / check (push) Waiting to run
A --workers value of 0 or less used to be quietly raised to 1, so a
typo ran the whole scan on one worker with nothing on stderr to say
why. scan now refuses it before anything is scanned: one line on
stderr and exit 2, like the other usage errors. The clamp in runScan
is gone, and README states the rule and the default.

Model: opus-5-5
2026-10-04 14:30:21 +02:00
clawbot 7278c354f0 Test both walk cancellation checks on their own (closes #81)
check / check (push) Waiting to run
walkOneDir's check was hiding the worker's: a worker that walked a
queued directory on a cancelled scan still emitted nothing, because
walkOneDir stopped at its first entry. The worker test now queues a
missing directory, whose read fails and sends a warning before
walkOneDir's check is reached. A new test calls walkOneDir directly on
a cancelled scan and checks it returns no subdirectory to descend into.

Model: opus-5-5
2026-10-04 14:01:38 +02:00
clawbot 8032ea682b Test that scan refuses another schema version (closes #64)
check / check (push) Waiting to run
The version-mismatch test only opened its database through
openReportDatabase, so nothing exercised the branch of initSchema that
stops scan on a database stamped with an unknown schema version. The
test now opens the same database through openScanDatabase too and
requires errSchemaVersion, and is renamed to match the unversioned-file
test beside it, which also covers both paths.

Model: opus-5-5
2026-10-04 13:30:23 +02:00
clawbot 948fb03630 Correct four inaccurate comments in cancel_test.go and rename a constant (closes #33)
check / check (push) Waiting to run
Fix the comments the re-review of
#6 found misdescribing their
tests; no test's behaviour changes.

State the property the walkClock tests rely on, that the index load's
Done cost does not grow with the record count, instead of a wrong fixed
figure. Describe poolUnwind so it is true of every use, and mark the
tests that have no bound. Record that hashWorker's results-send exit is
reachable through stop but has no test that fails without it. Rename
walkCancelInFlightDirs to walkCancelInFlightFiles, since it counts
files. Note at the top why the file departs from the
one-test-file-per-source-file convention.

Model: opus-4-8 (implementation); opus-5-5 (rework)
2026-10-04 13:01:31 +02:00
clawbot 7ff0dbb6e1 Test the -x filesystem-boundary rejection branch (closes #17)
check / check (push) Waiting to run
No test ever ran the part of subdirJob that refuses to descend across
a filesystem boundary under -x. The new tests call subdirJob directly
with a made-up device for the operand, so no second real filesystem is
needed. They cover refusal across a boundary, skipping the check when
the operand's device is unknown, the warning when the subdirectory
cannot be statted, and crossing by default without -x. When descent is
accepted, the whole returned job is compared, so losing the operand's
device on the way down fails a test. scan.go is unchanged.

Model: opus-4-8 (implementation); opus-5-5 (rebase)
2026-10-04 12:13:23 +02:00
clawbot bccc14ffef Refuse an unversioned database that already has a files table (closes #11)
check / check (push) Waiting to run
A database at user_version 0 that already has a files table was made
by something else: scan used to run its CREATE TABLE on it and fail
with a raw SQLite error, and report and trees gave only a bare version
mismatch. All three now refuse such a database with the schema-version
error telling the operator to remove the file and rescan. scan creates
the table and index and sets the version in one transaction, so a first
scan stopped partway leaves an empty database the next scan sets up,
never a files table at version 0. A genuinely empty database is
unchanged.

Model: opus-4-8 (implementation); opus-5-5 (rebase)
2026-10-04 12:01:36 +02:00
clawbot 33f8607e17 Document install, a daily cron scan and reading the reports (closes #54)
check / check (push) Waiting to run
Getting Started gains three parts: installing with go install, from a
clone or as the Docker image; a crontab line for a daily root scan,
with where its stderr and failures go; and how to read the two reports,
why a row is a candidate rather than proof, and how to check a pair
with cmp before removing anything. The usage block names --help, and
the text below it --workers and -x.

The install line uses @main, not @latest: @latest resolves to the
v0.0.1 tag, which predates the database.

Model: opus-5-5
2026-10-04 11:47:20 +02:00
clawbot 2eeba3df6f Keep the module cache out of the build stage's chown (closes #43)
check / check (push) Waiting to run
The build stage handed /src and the whole Go module cache to the
unprivileged user with chown -R, in a layer that re-ran on every source
change. On this host that step took from about 80 s to over ten minutes,
depending on load.

The module cache now sits at /go/pkg/mod and stays root's: bootstrap
fills it as root. The sources are copied with --chown. A small chown,
cached with bootstrap, hands builder the /src directory itself, the
module cache's cache/download directory (where make build saves its
lookup of this module's own version), and the telemetry files root's go
commands left in its home. Tests and the build still run as builder.

Model: opus-5-5
2026-10-04 10:01:28 +02:00
clawbot cb5dda4f45 Print --version to stdout (closes #15)
check / check (push) Waiting to run
Cobra's built-in version flag prints through the writer that carries
help and usage, which is stderr here. The root command now defines
its own -v/--version flag and prints one line, "sfdupes VERSION", to
stdout; a failed write is a fatal error (exit 1). Help and usage stay
on stderr. README documents --version and --help, their streams and
exit codes. Tests cover both flags and the failed write.

Model: opus-5-5
2026-10-04 09:47:28 +02:00
clawbot 4847882e46 Stop scan cleanly on SIGINT or SIGTERM (closes #5)
check / check (push) Waiting to run
A first SIGINT or SIGTERM cancels the scan. It commits the hashed
records still in its batch, with a context that is not cancelled for
that one write, and starts no other write or deletion; deletions need
a complete walk, so records under paths an interrupted walk never
reached are kept. A batch whose commit failed is now kept for that
final commit instead of dropped. The progress display is finished (a
bar stopped short is no longer filled up), `scan: interrupted after N
files` goes to stderr, and the exit code is 1. A second signal ends the
process at once. A SIGINT inherited as ignored stays ignored.

Model: opus-5-5
2026-10-04 09:30:38 +02:00
clawbot 33cf3dd29a Stream report and trees instead of loading every record (closes #14)
check / check (push) Waiting to run
report now has SQLite group the records and put the rows in report
order, helped by a new files_signature index on (size, head, tail,
content), and writes each row as it reads it. trees reads the records
in path order, where all the paths under a directory come together, so
it computes each directory's digest as soon as the stream leaves it and
keeps only its path, parent, digest and totals. Output is unchanged.

The tests that called the removed in-memory grouping functions now group
records stored in a database. New tests check that both commands give
the same output whatever order the records were inserted in, and that a
stdout failure partway through a long report is reported as one.

Model: opus-5-5
2026-10-04 06:47:25 +02:00
clawbot 9abf81535a Print progress at once off a terminal, keep warnings out of redraws (closes #13)
check / check (push) Successful in 1m56s
When stderr is not a terminal, each phase prints its zero-state line
as it starts instead of after its first item. stderrIsTTY uses
term.IsTerminal from golang.org/x/term, now a direct dependency, so
/dev/null is no longer taken for a terminal.

A spinner keeps the library's background redraw, so its count and
elapsed time stay current while a phase waits for its next item. A
warning printed during a spinner phase goes through the bar
(progressbar.Bprintln), which prints it before its next redraw instead
of racing it. Bars with a total have no background redraw and still
print warnings directly.

Model: opus-5-5
2026-10-04 04:47:35 +02:00
clawbot 2dd1194f33 Warn about and skip non-regular and .zfs operands, keeping their records (closes #9)
check / check (push) Successful in 2m48s
A symlink, socket, FIFO or device-node operand, or a directory operand
named .zfs, was silently ignored yet stayed in the scanned operands, so
the update phase deleted every record stored beneath it. Such an
operand now gets a one-line warning, counts as skipped, and is dropped
before overlapping operands are pruned and the database index is
loaded: another operand beneath it is still scanned, and the records
beneath it count as outside the scanned operands and are not deleted,
unless it lies under another operand. The exit status stays 0. An
operand that turns into one of these after that check is warned about
and skipped by the walk instead. README "scan mode" and "Rules for the
walk" say so.

Model: opus-5-5
2026-10-04 04:01:43 +02:00
clawbot 705c8729ca Hold a lock so a second scan fails at once (closes #53)
check / check (push) Successful in 1m43s
scan takes an exclusive flock(2) on a lock file beside the database
(its path with .lock appended) before it walks anything or opens the
database, and holds it until it returns. A second scan against the
same database fails at once with a one-line error naming the lock
file and exits 1. report and trees never take the lock. The lock ends
with the process, so a fatal error or an interrupt releases it; the
file is never deleted. golang.org/x/sys becomes a direct dependency.

The README smoke test now keeps the database outside the scanned
tree, where its empty lock file would have joined the empty-file
group.

Model: opus-5-5
2026-10-04 02:30:21 +02:00
clawbot 01ff3bb5f0 Test report and trees stdout write failures (closes #30)
check / check (push) Successful in 1m34s
report and trees already checked every stdout write and the final
flush. run now takes the stdout it hands to them, so tests pass a
closed file or a failing writer instead of swapping os.Stdout: a
closed stdout exits 1 with a one-line diagnostic, and the writer's
error reaches the caller.

README "Error handling" now states the two cases that never reach
sfdupes as a failed write: a pipe reader that exits early ends the
process with SIGPIPE, as with cat; and stdout closed with >&- is
replaced by /dev/null by the Go runtime before main runs, so the run
succeeds.

Model: opus-5-5
2026-10-03 18:01:27 +02:00
clawbot d63d3cc7fc Open the database read-only for report and trees (closes #8)
check / check (push) Successful in 1m17s
report and trees now connect read-only (mode=ro, query_only, the same
busy timeout) and no longer set the journal mode, which is a write. A
read-only connection to a WAL database still needs its -wal and -shm
files, or write access to the directory to create them, so scan now
switches the database back to rollback-journal mode whenever it closes
it: between scans the file alone holds the database. If a report has
the database open at that moment the switch is refused; scan warns and
the database stays in WAL mode, with its -wal and -shm files, until the
next scan. README §Database states what readers need.

Model: opus-5-5
2026-10-03 16:30:19 +02:00
clawbot c887f80f57 Escape tab, newline, CR and backslash in report paths (closes #7)
check / check (push) Successful in 1m29s
A path holding a tab or newline split a row of the report or trees
output. The path columns of both now write a backslash, tab, newline
and carriage return as \\, \t, \n and \r; every other byte is written
unchanged. Grouping and sorting still use the stored path. Warnings
on stderr are escaped the same way in warnf, so each stays one line.

In trees, the root directory's node now has the path "/" instead of
an empty string, and its children's paths start with a single slash.

README states the rule under "Report output format".

Model: opus-5-5
2026-10-03 15:30:37 +02:00
clawbot c9bf22d483 Stamp the git tag or short commit in a plain docker build (closes #67)
check / check (push) Successful in 1m1s
.dockerignore now sends .git, without .git/config, which can hold a
credential. The build stage takes the VERSION build argument when one
is given, otherwise git describe --tags --always of that .git, and
fails if the context carries .git and still yields no version. A plain
docker build . used to stamp dev. The CI checkout fetches full history
so CI sees the tag and stamps the same value as make build.

Model: opus-5-5
2026-10-02 08:49:03 +02:00
26 changed files with 3797 additions and 1548 deletions
+2
View File
@@ -0,0 +1,2 @@
node_modules/
yarn.lock
+4
View File
@@ -0,0 +1,4 @@
{
"tabWidth": 4,
"proseWrap": "always"
}
+80 -29
View File
@@ -12,9 +12,9 @@ COPY . .
# build would exit 0 having run nothing. script/cibuild and # build would exit 0 having run nothing. script/cibuild and
# script/docker pass a fresh CHECK_EPOCH on every invocation. # script/docker pass a fresh CHECK_EPOCH on every invocation.
# #
# Two properties this depends on. ARG is per-stage, so the build stage # Two properties this depends on. ARG is per-stage, so the markdown and
# below declares it again; one declaration here would leave that # build stages below declare it again; one declaration here would leave
# stage's gate cacheable. And each gate RUN must reference the value, # their gates cacheable. And each gate RUN must reference the value,
# because BuildKit hashes the expanded command: a declared but # because BuildKit hashes the expanded command: a declared but
# unreferenced ARG invalidates nothing. # unreferenced ARG invalidates nothing.
# #
@@ -27,9 +27,16 @@ ARG CHECK_EPOCH
# target now runs `docker build -f Dockerfile.lint`, and a docker build # target now runs `docker build -f Dockerfile.lint`, and a docker build
# cannot run a docker build: routing the gate through make would mean # cannot run a docker build: routing the gate through make would mean
# nesting docker inside this image. Same reason `make check` is gone # nesting docker inside this image. Same reason `make check` is gone
# from the build stage below. `make fmt-check` stays as it is — it is a # from the build stage below, and `make fmt-check` from both stages: it
# gate, not the aggregate, and it shells out to nothing. # runs prettier through docker too. Its gofmt half is the step below,
RUN echo "gate fmt-check, epoch ${CHECK_EPOCH}" && make fmt-check # its Markdown half the markdown stage further down. gofmt's output is
# assigned to a variable first so that its own exit status, as when it
# cannot parse a file, still fails the step.
RUN echo "gate gofmt, epoch ${CHECK_EPOCH}" && \
files="$(gofmt -s -l .)" && \
if [ -n "$files" ]; then \
echo "gofmt: files not formatted:" >&2; echo "$files" >&2; exit 1; \
fi
# The FROM above and the one in Dockerfile.lint pin the same linter # The FROM above and the one in Dockerfile.lint pin the same linter
# twice, and nothing else keeps them in sync; this fails the build when # twice, and nothing else keeps them in sync; this fails the build when
@@ -46,32 +53,65 @@ RUN echo "gate config verify, epoch ${CHECK_EPOCH}" && \
RUN echo "gate lint, epoch ${CHECK_EPOCH}" && \ RUN echo "gate lint, epoch ${CHECK_EPOCH}" && \
golangci-lint run --config .golangci.yml ./... golangci-lint run --config .golangci.yml ./...
# Prettier stage: the prettier that formats this repository's Markdown,
# never installed on a host. script/fmt and script/fmt-check build this
# stage alone and run it with the repository mounted on /src. prettier
# is installed in /tools so that the repository, mounted or copied onto
# /src, cannot hide it.
# node:22-alpine, 2026-02-22
FROM node@sha256:e4bf2a82ad0a4037d28035ae71529873c069b13eb0455466ae0bc13363826e34 AS prettier
WORKDIR /tools
# yarn.lock pins prettier by hash, and --frozen-lockfile fails rather
# than install anything yarn.lock does not name.
COPY package.json yarn.lock ./
RUN yarn install --frozen-lockfile
ENV PATH=/tools/node_modules/.bin:$PATH
WORKDIR /src
# Markdown stage: the Markdown half of `make fmt-check`, as a gate.
FROM prettier AS markdown
COPY . .
# Second per-stage declaration of the gate cache-buster; see the lint
# stage above.
ARG CHECK_EPOCH
RUN echo "gate prettier, epoch ${CHECK_EPOCH}" && \
prettier --check '**/*.md' --tab-width 4 --prose-wrap always
# Build stage # Build stage
# golang:1.25-alpine, 2026-07-23 # golang:1.25-alpine, 2026-07-23
FROM golang@sha256:56961d79ea8129efddcc0b8643fd8a5416b4e6228cfd477e3fd61deb2672c587 AS builder FROM golang@sha256:56961d79ea8129efddcc0b8643fd8a5416b4e6228cfd477e3fd61deb2672c587 AS builder
# We never build or run as root. Create an unprivileged user and point # We never build or run as root. Create an unprivileged user and point
# HOME and the Go caches at its home so go build and go test can write # HOME and the build cache at its home so go build and go test can write
# their caches when we drop to it below. $GOPATH/bin is deliberately not # it when we drop to it below. $GOPATH/bin is deliberately not on PATH:
# on PATH: script/bootstrap no longer `go install`s anything (the linter # script/bootstrap no longer `go install`s anything (the linter runs
# runs from a pinned image, never from a host install), so nothing lands # from a pinned image, never from a host install), so nothing lands
# there and adding it would only widen what this image resolves. # there and adding it would only widen what this image resolves.
#
# The module cache is kept outside that home, at the base image's
# default /go/pkg/mod, and belongs to root: script/bootstrap fills it as
# root. Do not move it into the home and hand it over with `chown -R`:
# that walks every file in it, which took from about 80 s to over ten
# minutes on a shared host, depending on load.
RUN adduser -D -u 1000 builder RUN adduser -D -u 1000 builder
ENV HOME=/home/builder ENV HOME=/home/builder
ENV GOPATH=/home/builder/go ENV GOPATH=/home/builder/go
ENV GOMODCACHE=/go/pkg/mod
ENV GOCACHE=/home/builder/.cache/go-build ENV GOCACHE=/home/builder/.cache/go-build
WORKDIR /src WORKDIR /src
# No-op file copy whose only purpose is the build-graph edge: it is what # No-op file copies whose only purpose is the build-graph edge: they are
# makes this stage depend on the lint stage, and so what forces BuildKit # what make this stage depend on the lint and markdown stages, and so
# to finish fmt-check, the pin guard and lint before compilation and # what forces BuildKit to finish gofmt, the pin guard, lint and prettier
# tests start. Remove it and the fail-fast design dies silently — the # before compilation and tests start. Remove one and the fail-fast
# build stops gating on lint and still exits 0. It replaces a copy of # design dies silently — the build stops gating on that stage and still
# the linter binary itself, which is no longer wanted here: nothing in # exits 0. The first replaces a copy of the linter binary itself, which
# this stage runs the linter, because `make lint` is now a docker build # is no longer wanted here: nothing in this stage runs the linter,
# and a docker build cannot run inside one. # because `make lint` is now a docker build and a docker build cannot
# run inside one.
COPY --from=lint /src/go.sum /dev/null COPY --from=lint /src/go.sum /dev/null
COPY --from=markdown /src/go.sum /dev/null
# Install development prerequisites the same way a developer does, # Install development prerequisites the same way a developer does,
# rather than duplicating the installs inline. Only script/ and the # rather than duplicating the installs inline. Only script/ and the
@@ -83,30 +123,41 @@ COPY script/ script/
COPY go.mod go.sum ./ COPY go.mod go.sum ./
RUN script/bootstrap RUN script/bootstrap
COPY . . # Hand builder only what it writes to, without walking the module cache.
# This layer stays cached with bootstrap.
# - /src itself: make build writes the binary into it, and git refuses
# a repository whose top directory belongs to another user.
# - the module cache's cache/download directory itself, not what is in
# it: Go only reads the downloaded modules, but make build saves its
# lookup of this module's own version from git there, in a new
# directory named after the module path.
# - builder's home: the go commands bootstrap ran as root left Go's
# telemetry files there, a few small files.
RUN chown builder:builder /src /go/pkg/mod/cache/download && \
chown -R builder:builder /home/builder
# Hand the sources and caches to the unprivileged user, then drop root # The sources are handed to builder as they are copied, so no layer has
# before running any checks or builds. # to walk them. Then drop root before running any checks or builds.
RUN chown -R builder:builder /src /home/builder COPY --chown=builder:builder . .
USER builder USER builder
# Fail the build unless the branch is green. Runs as non-root so the # Fail the build unless the branch is green. Runs as non-root so the
# permission-denied test paths are exercised legitimately (root would # permission-denied test paths are exercised legitimately (root would
# bypass the chmod(0) the tests rely on). # bypass the chmod(0) the tests rely on).
# #
# The gates are the individual targets, not `make check`: that aggregate # The gate is `make test`, not `make check`: that aggregate runs
# runs `script/lint`, which is now a docker build, and nothing inside an # `script/lint` and `script/fmt-check`, which both run docker, and
# image build may shell out to docker. Lint is not skipped by this — it # nothing inside an image build may shell out to docker. Lint and the
# ran in the lint stage above, which this stage's COPY --from makes a # format checks are not skipped by this — they ran in the lint and
# prerequisite. `make`, not the scripts directly, because the Makefile's # markdown stages above, which this stage's COPY --from lines make
# prerequisites. `make`, not the script directly, because the Makefile's
# `export CGO_ENABLED = 0` applies only to what it invokes. # `export CGO_ENABLED = 0` applies only to what it invokes.
# #
# Second per-stage declaration of the gate cache-buster; see the lint # Third per-stage declaration of the gate cache-buster; see the lint
# stage above for why one is not enough. It is placed after USER so the # stage above for why one is not enough. It is placed after USER so the
# drop to the unprivileged user still happens before the checks run. # drop to the unprivileged user still happens before the checks run.
ARG CHECK_EPOCH ARG CHECK_EPOCH
RUN echo "gate test, epoch ${CHECK_EPOCH}" && make test RUN echo "gate test, epoch ${CHECK_EPOCH}" && make test
RUN echo "gate fmt-check, epoch ${CHECK_EPOCH}" && make fmt-check
# The version stamped into the binary: the VERSION build argument when # The version stamped into the binary: the VERSION build argument when
# one is given, otherwise `git describe --tags --always` of the .git in # one is given, otherwise `git describe --tags --always` of the .git in
+709 -540
View File
File diff suppressed because it is too large Load Diff
+424 -409
View File
@@ -1,464 +1,479 @@
# Workflow # Workflow
- take an issue from the `1.0.0` milestone on the tracker; work not - take an issue from the `1.0.0` milestone on the tracker; work not yet on the
yet on the tracker gets filed as an issue first tracker gets filed as an issue first
- branch (from `main`) - branch (from `main`)
- do the work, with tests, in small focused commits - do the work, with tests, in small focused commits
- record it at the top of Completed Steps (`TODO.md` changes in the - record it at the top of Completed Steps (`TODO.md` changes in the same commit
same commit as the work) as the work)
- push the branch and open a PR whose title ends with - push the branch and open a PR whose title ends with ` (closes #N)`
` (closes #N)` - an independent review gates the merge; every finding is addressed or
- an independent review gates the merge; every finding is addressed explicitly rebutted on the PR
or explicitly rebutted on the PR
- merge to `main` once the review passes - merge to `main` once the review passes
# Status # Status
- pre-1.0 - pre-1.0
- the Gitea tracker is authoritative for the pre-1.0 backlog: the - the Gitea tracker is authoritative for the pre-1.0 backlog: the open issues
open issues under the `1.0.0` milestone are what remains before under the `1.0.0` milestone are what remains before the tag, and this file
the tag, and this file records history and process, not the queue records history and process, not the queue
# Next Step # Next Step
- take the next issue from the `1.0.0` milestone on the tracker: - take the next issue from the `1.0.0` milestone on the tracker:
https://git.eeqj.de/sneak/sfdupes/milestone/17 — the milestone is https://git.eeqj.de/sneak/sfdupes/milestone/17 — the milestone is the source
the source of truth for what is left before 1.0.0. Individual of truth for what is left before 1.0.0. Individual issues are deliberately not
issues are deliberately not restated here; a copy in this file restated here; a copy in this file drifts out of date the moment the tracker
drifts out of date the moment the tracker moves moves
# Completed Steps # Completed Steps
- stamp the git tag or short commit in a plain `docker build .` - `make fmt` and `make fmt-check` run prettier over all Markdown, in Docker, and
instead of `dev` (2026-10-02, branch `next`, closes CI checks it; all Markdown reformatted (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/67): `.dockerignore` now https://git.eeqj.de/sneak/sfdupes/issues/19)
sends `.git`, without `.git/config`, and the `Dockerfile` build
stage takes the `VERSION` build argument when one is given,
otherwise `git describe --tags --always` of that `.git`. The build
fails if the context carries `.git` and the version still comes out
empty, `dev` or `unknown`. The CI checkout step fetches the full
history (`fetch-depth: 0`) so CI sees the tag and stamps the same
value as `make build`.
- replace the 1 KiB end-window sampling with the head/tail plus - `scan` rejects `--workers` below 1 as a usage error instead of running
content-hash ladder (2026-09-22, branch `next`, closes single-threaded (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/10)
https://git.eeqj.de/sneak/sfdupes/issues/61): a file under 10 MiB is
hashed in full and compared directly, with no end-window step — its
`head`, `tail`, and `content` all hold the whole-file hash. A file at
10 MiB or above gets only the 64 KiB `head` and `tail` in the hash
phase; a new content phase, after the update phase, reads it for its
`content` hash — the whole file below 50 MiB, gigabyte-spaced 1 MiB
samples at or above — only when its size, `head`, and `tail` match
another record's, from the same scan or stored by an earlier one, so
a stored file gains its content hash when it gains a match. A file
that is gone or has changed since its record was written is not
read. The `content` column is part of the version 1 schema. `report`
and `trees` group by the extended signature and leave out any record
without a `content` hash, so the ladder is applied across the whole
database. README "Duplicate detection" documents every rung including
the probabilistic large-file path.
- remove the dead `files.dat` references from `Makefile`, `.gitignore` - a test fails when either walk cancellation check in `scan.go` is removed
and `.dockerignore` (2026-09-21, branch `next`, closes (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/81)
- test that `scan` refuses a database with another schema version (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/64)
- correct four inaccurate comments in `cancel_test.go` and rename
`walkCancelInFlightDirs` to `walkCancelInFlightFiles` (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/33)
- test the `-x` filesystem-boundary rules in `subdirJob` (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/17)
- `scan` creates the schema in one transaction; a version-0 database with a
`files` table is refused with a clear schema-version error (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/11)
- README documents install, Docker, a daily cron scan and how to read and check
the reports (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/54)
- the `Dockerfile` build stage keeps the Go module cache out of `builder`'s home
and copies the sources with `--chown`, so no `chown -R` walks them
(2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/43)
- `--version` prints `sfdupes VERSION` to stdout; README documents it and
`--help` (2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/15)
- `scan` stops cleanly on `SIGINT` or `SIGTERM`: commits what it has hashed,
deletes nothing more, exits 1 (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/5)
- `report` and `trees` stream the records instead of holding them all in memory;
the schema gains the `files_signature` index (2026-10-04,
https://git.eeqj.de/sneak/sfdupes/issues/14)
- progress prints at once on a non-terminal, uses a real terminal test, and
prints warnings through a spinner instead of racing its redraw (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/13)
- warn about and skip symlink, socket, FIFO, device and `.zfs` operands, keeping
the records beneath them (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/9)
- `scan` holds a lock on a lock file beside the database for its whole run, so a
second `scan` fails at once with exit 1 (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/53)
- test stdout write failures in `report` and `trees`; README states that
`| head` ends sfdupes by `SIGPIPE` and `>&-` writes to `/dev/null`
(2026-10-03, https://git.eeqj.de/sneak/sfdupes/issues/30)
- `report` and `trees` open the database read-only, and `scan` leaves it out of
WAL mode, so reading needs only read access (2026-10-03, closes
https://git.eeqj.de/sneak/sfdupes/issues/8)
- escape tabs, newlines, carriage returns and backslashes in report, trees and
warning paths; the root directory's path is `/` (2026-10-03,
https://git.eeqj.de/sneak/sfdupes/issues/7)
- stamp the git tag or short commit in a plain `docker build .` instead of `dev`
(2026-10-02, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/67): `.dockerignore` now sends
`.git`, without `.git/config`, and the `Dockerfile` build stage takes the
`VERSION` build argument when one is given, otherwise
`git describe --tags --always` of that `.git`. The build fails if the context
carries `.git` and the version still comes out empty, `dev` or `unknown`. The
CI checkout step fetches the full history (`fetch-depth: 0`) so CI sees the
tag and stamps the same value as `make build`.
- replace the 1 KiB end-window sampling with the head/tail plus content-hash
ladder (2026-09-22, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/61): a file under 10 MiB is hashed in
full and compared directly, with no end-window step — its `head`, `tail`, and
`content` all hold the whole-file hash. A file at 10 MiB or above gets only
the 64 KiB `head` and `tail` in the hash phase; a new content phase, after the
update phase, reads it for its `content` hash — the whole file below 50 MiB,
gigabyte-spaced 1 MiB samples at or above — only when its size, `head`, and
`tail` match another record's, from the same scan or stored by an earlier one,
so a stored file gains its content hash when it gains a match. A file that is
gone or has changed since its record was written is not read. The `content`
column is part of the version 1 schema. `report` and `trees` group by the
extended signature and leave out any record without a `content` hash, so the
ladder is applied across the whole database. README "Duplicate detection"
documents every rung including the probabilistic large-file path.
- remove the dead `files.dat` references from `Makefile`, `.gitignore` and
`.dockerignore` (2026-09-21, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/22) https://git.eeqj.de/sneak/sfdupes/issues/22)
- fix the lint-image pin comments and `FROM` form in `Dockerfile` and - fix the lint-image pin comments and `FROM` form in `Dockerfile` and
`Dockerfile.lint` (2026-08-10, branch `next`, closes `Dockerfile.lint` (2026-08-10, branch `next`, closes
https://git.eeqj.de/sneak/sfdupes/issues/25): dropped the false https://git.eeqj.de/sneak/sfdupes/issues/25): dropped the false
`(Debian-based)` parenthetical (v2.12.1 was Debian too) and the `(Debian-based)` parenthetical (v2.12.1 was Debian too) and the redundant tag,
redundant tag, so both pins are the policy `# image:vX.Y.Z, so both pins are the policy `# image:vX.Y.Z, YYYY-MM-DD` comment over a bare
YYYY-MM-DD` comment over a bare `FROM image@sha256:...`. Digest `FROM image@sha256:...`. Digest unchanged. `script/verify-lint-image-pin`
unchanged. `script/verify-lint-image-pin` parses those `FROM` lines parses those `FROM` lines and still matches the tagless form; its advice line
and still matches the tagless form; its advice line lost the now lost the now meaningless "tag and digest". With no tag in either reference, a
meaningless "tag and digest". With no tag in either reference, a tag-only disagreement no longer exists — a one-sided tag is caught as a plain
tag-only disagreement no longer exists — a one-sided tag is caught as mismatch.
a plain mismatch.
- run all linting in Docker via `Dockerfile.lint` and `script/lint` - run all linting in Docker via `Dockerfile.lint` and `script/lint` (2026-08-10,
(2026-08-10, branch `next`, closes branch `next`, closes https://git.eeqj.de/sneak/sfdupes/issues/46): per the
https://git.eeqj.de/sneak/sfdupes/issues/46): per the owner ruling, the owner ruling, the linter runs inside a container invoked through the `script/`
linter runs inside a container invoked through the `script/` entrypoint and is never installed on a host. New root `Dockerfile.lint` COPYs
entrypoint and is never installed on a host. New root the repo into the digest-pinned `golangci/golangci-lint:v2.12.2` image and
`Dockerfile.lint` COPYs the repo into the digest-pinned runs `golangci-lint config verify` and `golangci-lint run` as build steps, so
`golangci/golangci-lint:v2.12.2` image and runs a successful build IS a clean lint; `script/lint` is reduced to building it.
`golangci-lint config verify` and `golangci-lint run` as build `script/bootstrap` loses the `go install`, the pin constants, the version
steps, so a successful build IS a clean lint; `script/lint` is parser and `verify_golangci_lint` outright rather than hardening them — with
reduced to building it. `script/bootstrap` loses the `go install`, nothing linting on the host, the `$GOPATH/bin` versus `PATH` problem that
the pin constants, the version parser and `verify_golangci_lint` motivated them has no subject — and now warns rather than fails when `docker`
outright rather than hardening them — with nothing linting on the is absent. Two traps handled. A lint build on an unchanged tree returns
host, the `$GOPATH/bin` versus `PATH` problem that motivated them has success in well under a second having run no linter, which is
no subject — and now warns rather than fails when `docker` is absent.
Two traps handled. A lint build on an unchanged tree returns success
in well under a second having run no linter, which is
https://git.eeqj.de/sneak/sfdupes/issues/32 and https://git.eeqj.de/sneak/sfdupes/issues/32 and
https://git.eeqj.de/sneak/sfdupes/issues/39 again, so https://git.eeqj.de/sneak/sfdupes/issues/39 again, so `Dockerfile.lint`
`Dockerfile.lint` carries `ARG CHECK_EPOCH` referenced carries `ARG CHECK_EPOCH` referenced inside every gate `RUN` (BuildKit hashes
inside every gate `RUN` (BuildKit hashes the expanded command, not the expanded command, not the declaration) and `script/lint` passes
the declaration) and `script/lint` passes `"$(date +%s)-$$"` — the `"$(date +%s)-$$"` — the PID matters because two lint runs land inside the
PID matters because two lint runs land inside the same second easily. same second easily. And nothing inside an image build may shell out to docker,
And nothing inside an image build may shell out to docker, so the so the main `Dockerfile`'s lint stage now invokes `golangci-lint` directly
main `Dockerfile`'s lint stage now invokes `golangci-lint` directly
instead of `make lint`, and its build stage runs `make test` and instead of `make lint`, and its build stage runs `make test` and
`make fmt-check` instead of the `make check` aggregate (`make`, not `make fmt-check` instead of the `make check` aggregate (`make`, not the
the scripts bare, because the Makefile's `export CGO_ENABLED = 0` scripts bare, because the Makefile's `export CGO_ENABLED = 0` only reaches
only reaches what it invokes). `COPY --from=lint` what it invokes). `COPY --from=lint` `/usr/bin/golangci-lint` is replaced by
`/usr/bin/golangci-lint` is replaced by `COPY --from=lint /src/go.sum /dev/null`: the copied binary was the only edge
`COPY --from=lint /src/go.sum /dev/null`: the copied binary was the forcing BuildKit to finish linting before the build stage starts, and dropping
only edge forcing BuildKit to finish linting before the build stage it without replacing the edge would have ended fail-fast linting silently
starts, and dropping it without replacing the edge would have ended under a still-green build. That is canonical `REPO_POLICIES.md:107`'s ordering
fail-fast linting silently under a still-green build. That is edge, restored. `ENV PATH=/home/builder/go/bin:$PATH` is gone with the
canonical `REPO_POLICIES.md:107`'s ordering edge, restored. `go install` that justified it. `script/verify-linter-pin` is retired, deleted
`ENV PATH=/home/builder/go/bin:$PATH` is gone with the `go install` along with its README entry, because both of its subjects ceased to exist in
that justified it. `script/verify-linter-pin` is retired, deleted the same change: it compared a linter binary against `GOLANGCI_LINT_VERSION`
along with its README entry, because both of its subjects ceased to in `script/bootstrap`, and there is now neither a binary crossing between
exist in the same change: it compared a linter binary against stages nor a version pin in bootstrap. The drift it guarded has not gone away,
`GOLANGCI_LINT_VERSION` in `script/bootstrap`, and there is now it has moved — the linter is still pinned twice, now as the `FROM` line of
neither a binary crossing between stages nor a version pin in `Dockerfile.lint` and the `FROM` line of the `Dockerfile` lint stage, with
bootstrap. The drift it guarded has not gone away, it has moved — the nothing syncing them, which is exactly what
linter is still pinned twice, now as the `FROM` line of
`Dockerfile.lint` and the `FROM` line of the `Dockerfile` lint stage,
with nothing syncing them, which is exactly what
https://git.eeqj.de/sneak/sfdupes/issues/42 made a build failure. Its https://git.eeqj.de/sneak/sfdupes/issues/42 made a build failure. Its
replacement is one new `script/verify-lint-image-pin`, replacement is one new `script/verify-lint-image-pin`, run as a gate in both
run as a gate in both files, which compares the two references to files, which compares the two references to each other and deliberately
each other and deliberately restates neither: a hardcoded expected restates neither: a hardcoded expected digest would be a third copy and the
digest would be a third copy and the same drift one file further out. same drift one file further out. `golangci-lint config verify` is included per
`golangci-lint config verify` is included per the ruling, and the the ruling, and the concern about its unpinned live HTTPS schema fetch was
concern about its unpinned live HTTPS schema fetch was measured measured rather than assumed — under `--network none` the pinned binary both
rather than assumed — under `--network none` the pinned binary both passes a valid config and rejects an invalid one with the jsonschema error, so
passes a valid config and rejects an invalid one with the jsonschema it validates from an embedded schema and makes no network call of its own. The
error, so it validates from an embedded schema and makes no network README scopes that to the gate steps rather than to linting as a whole:
call of its own. The README scopes that to the gate steps rather `Dockerfile.lint` runs `go mod download` above them, so a cold cache still
than to linting as a whole: `Dockerfile.lint` runs `go mod download` needs the network and only a warm one lints offline. Verified: `make lint`
above them, so a cold cache still needs the network and only a warm green with every `PATH` directory containing a `golangci-lint` removed
one lints offline. Verified: `make lint` green with every `PATH`
directory containing a `golangci-lint` removed
(`/home/user/go/bin`, `/home/user/.local/bin`, `/usr/local/bin`; (`/home/user/go/bin`, `/home/user/.local/bin`, `/usr/local/bin`;
`command -v golangci-lint` empty); two consecutive `script/lint` runs `command -v golangci-lint` empty); two consecutive `script/lint` runs on an
on an untouched tree both executed the linter, 27.7s and 28.7s in the untouched tree both executed the linter, 27.7s and 28.7s in the lint step
lint step under distinct epochs with the `COPY . .` layer `CACHED` under distinct epochs with the `COPY . .` layer `CACHED` above them, at 42.2s
above them, at 42.2s and 41.8s wall clock — the no-cache rule was not and 41.8s wall clock — the no-cache rule was not weakened to shorten that.
weakened to shorten that. Negative control: a planted Negative control: a planted `var unusedIssue46Sentinel = 1` failed
`var unusedIssue46Sentinel = 1` failed `script/lint` with `script/lint` with
`report.go:173:5: var unusedIssue46Sentinel is unused (unused)`, and `report.go:173:5: var unusedIssue46Sentinel is unused (unused)`, and failed
failed `make docker` at `[lint 9/9]` with the build stage stopped at `make docker` at `[lint 9/9]` with the build stage stopped at `[builder 3/12]`
`[builder 3/12]` — `COPY --from=lint`, `script/bootstrap`, the test — `COPY --from=lint`, `script/bootstrap`, the test gate and `make build` all
gate and `make build` all zero occurrences — then reverted clean. The zero occurrences — then reverted clean. The drift guard fails on a tag-only
drift guard fails on a tag-only disagreement, on a digest-only disagreement, on a digest-only disagreement, and on an unreadable reference,
disagreement, and on an unreadable reference, naming both sides. naming both sides. `make docker` green in 5m35s with all six gates executing
`make docker` green in 5m35s with all six gates executing under one under one epoch (lint 37.6s, test 25.2s reporting
epoch (lint 37.6s, test 25.2s reporting `ok sneak.berlin/go/sfdupes 1.938s coverage: 88.5%`, not `(cached)`). The
`ok sneak.berlin/go/sfdupes 1.938s coverage: 88.5%`, not `(cached)`). non-root quirk still holds: in the builder image with the Go test cache off,
The non-root quirk still holds: in the builder image with the Go test `--user 0:0` fails `TestScanHardlinkRunFailsTogether` (exit 1) where the
cache off, `--user 0:0` fails `TestScanHardlinkRunFailsTogether` unprivileged user passes (exit 0). Noted for follow-up, not fixed here:
(exit 1) where the unprivileged user passes (exit 0). Noted for `golangci-lint` warns that the `gomodguard` linter is deprecated since v2.12.0
follow-up, not fixed here: `golangci-lint` warns that the in favour of `gomodguard_v2`.
`gomodguard` linter is deprecated since v2.12.0 in favour of
`gomodguard_v2`.
- install the Docker build stage's prerequisites by running - install the Docker build stage's prerequisites by running `script/bootstrap`
`script/bootstrap` instead of `apk add --no-cache make` inline instead of `apk add --no-cache make` inline (2026-08-09, branch
(2026-08-09, branch `dockerfile-bootstrap`, closes #42): canonical `dockerfile-bootstrap`, closes #42): canonical `REPO_POLICIES.md:97` requires
`REPO_POLICIES.md:97` requires it, and the inline install left the it, and the inline install left the build stage maintaining its own notion of
build stage maintaining its own notion of the toolchain — exactly the toolchain — exactly the divergence #24 exists to close, one layer down.
the divergence #24 exists to close, one layer down. The stage now The stage now copies `script/` plus `go.mod`/`go.sum` and runs
copies `script/` plus `go.mod`/`go.sum` and runs `script/bootstrap`, `script/bootstrap`, which ends in `go mod download`, so the separate
which ends in `go mod download`, so the separate invocation of that invocation of that is gone. `COPY --from=lint /usr/bin/golangci-lint` stays,
is gone. `COPY --from=lint /usr/bin/golangci-lint` stays, and moves and moves above the bootstrap layer. It is the only edge making this stage
above the bootstrap layer. It is the only edge making this stage depend on the lint stage, so deleting it as redundant would end fail-fast
depend on the lint stage, so deleting it as redundant would end linting silently. Letting bootstrap install its own linter here would have
fail-fast linting silently. Letting bootstrap install its own linter reintroduced the second toolchain and paid for a from-source build of it. What
here would have reintroduced the second toolchain and paid for a makes the two stages provably one toolchain rather than two that happen to
from-source build of it. What makes the two stages provably one agree is a new `script/verify-linter-pin`, run in the build stage on the
toolchain rather than two that happen to agree is a new binary that arrives from the lint stage, before bootstrap: it fails the build
`script/verify-linter-pin`, run in the build stage on the binary naming both versions unless that binary is the version `script/bootstrap`
that arrives from the lint stage, before bootstrap: it fails the pins. Bootstrap's own check could not serve that purpose — it reinstalls its
build naming both versions unless that binary is the version pin from source and then verifies whatever `PATH` resolves, so drift
`script/bootstrap` pins. Bootstrap's own check could not serve that self-heals silently and a lint stage image bumped on its own would lint at the
purpose — it reinstalls its pin from source and then verifies new version while `make check` ran at the old one, green. The linter version
whatever `PATH` resolves, so drift self-heals silently and a lint is pinned in two independent places (the lint stage image digest and
stage image bumped on its own would lint at the new version while
`make check` ran at the old one, green. The linter version is pinned
in two independent places (the lint stage image digest and
`GOLANGCI_LINT_VERSION`) and nothing else keeps them in sync, so a `GOLANGCI_LINT_VERSION`) and nothing else keeps them in sync, so a
half-applied bump is now a build failure. The pin is read out of half-applied bump is now a build failure. The pin is read out of
`script/bootstrap`, which stays the single source of truth; a pin `script/bootstrap`, which stays the single source of truth; a pin that cannot
that cannot be read is a hard failure, not a skip. The check needs be read is a hard failure, not a skip. The check needs no `CHECK_EPOCH`: its
no `CHECK_EPOCH`: its only inputs are the copied binary and only inputs are the copied binary and `script/`, so Docker invalidates the
`script/`, so Docker invalidates the layer exactly when a cached layer exactly when a cached result would stop being true, and it is documented
result would stop being true, and it is documented with the other with the other entrypoints in the README. `$GOPATH/bin` joins `PATH` because
entrypoints in the README. `$GOPATH/bin` joins `PATH` because that is where bootstrap's `go install` lands and bootstrap verifies its
that is where bootstrap's `go install` lands and bootstrap verifies installs against what `PATH` resolves — nothing in the image is shadowed by
its installs against what `PATH` resolves — nothing in the image is it, the directory does not exist until bootstrap runs. Everything added sits
shadowed by it, the directory does not exist until bootstrap runs. above `ARG CHECK_EPOCH`, and the `chown` and `USER builder` still precede
Everything added sits above `ARG CHECK_EPOCH`, and the `chown` and `make check`. Verified: the guard fails the build with both versions named
`USER builder` still precede `make check`. Verified: the guard fails when the lint stage's linter is faked to a different version, and an
the build with both versions named when the lint stage's linter is unmodified build still passes it; bootstrap runs clean under Alpine's `sh` and
faked to a different version, and an unmodified build still passes its `apk` branch, installing `git` and `make` and finding the copied linter
it; bootstrap runs clean under Alpine's `sh` and its `apk` branch, already at the pin; a second build served the bootstrap and dependency layers
installing `git` and `make` and finding the copied `CACHED` while both gates ran with a fresh epoch; a planted `unused` finding
linter already at the pin; a second build served the bootstrap and failed the build at the lint gate in 48.9s with the build stage's `make check`
dependency layers `CACHED` while both gates ran with a fresh epoch; never starting; and the suite run in the image as `--user 0:0` fails
a planted `unused` finding failed the build at the lint gate in `TestScanHardlinkRunFailsTogether`, so the drop to the unprivileged user is
48.9s with the build stage's `make check` never starting; and the still load-bearing. That last check needs the Go test cache disabled — the
suite run in the image as `--user 0:0` fails first attempt reported `ok ... (cached)` as root, reusing the result the
`TestScanHardlinkRunFailsTogether`, so the drop to the unprivileged build-time run had left in the shared cache, which would have read as a pass.
user is still load-bearing. That last check needs the Go test cache Build wall time, on a shared host running many concurrent builds and so noisy:
disabled — the first attempt reported `ok ... (cached)` as root, 2m13s on an unchanged tree, 2m17s and 4m29s for two builds after a source
reusing the result the build-time run had left in the shared cache, change, 5m14s cold. Only the cold one breaches the policy ceiling, and not
which would have read as a pass. Build wall time, on a shared host because of this change — `chown -R builder:builder /src /home/builder` walks
running many concurrent builds and so noisy: 2m13s on an unchanged the module cache and re-runs on every source change, and it alone varied
tree, 2m17s and 4m29s for two builds after a source change, 5m14s between 77s and 210s across those four builds, which is also the whole spread
cold. Only the cold one breaches the policy ceiling, and not because in the totals. The same cold measurement against `main` is 5m03s with a 209s
of this change — `chown -R builder:builder /src /home/builder` walks `chown`. Filed as #43
the module cache and re-runs on every source change, and it alone - bust the Docker layer cache for the gate steps, so `script/cibuild` and
varied between 77s and 210s across those four builds, which is also `script/docker` cannot report a green they did not earn (2026-08-09, branch
the whole spread in the totals. The same cold measurement against `cibuild-cache-bust`, closes #32): both scripts were bare `docker build`
`main` is 5m03s with a 209s `chown`. Filed as #43 invocations with no cache control, and the `Dockerfile` copies the tree before
- bust the Docker layer cache for the gate steps, so `script/cibuild` running its gates, so on an unchanged tree Docker served those layers from
and `script/docker` cannot report a green they did not earn cache and the build exited 0 having executed nothing. That is not hypothetical
(2026-08-09, branch `cibuild-cache-bust`, closes #32): both scripts here — every merge this repo has done is a non-fast-forward merge of an
were bare `docker build` invocations with no cache control, and the undiverged branch, so each merge commit's tree is byte-identical to the branch
`Dockerfile` copies the tree before running its gates, so on an head's and each merge CI run was almost certainly a full cache hit; and PR
unchanged tree Docker served those layers from cache and the build #31's reviewer found `make docker` returning success as a 17-layer cache hit,
exited 0 having executed nothing. That is not hypothetical here — catching it only by being suspicious. The fix is `ARG CHECK_EPOCH` with the
every merge this repo has done is a non-fast-forward merge of an scripts passing `--build-arg CHECK_EPOCH="$(date +%s)"`. Two details make or
undiverged branch, so each merge commit's tree is byte-identical to break it. `ARG` is scoped per stage and this `Dockerfile` has three gates
the branch head's and each merge CI run was almost certainly a full across two — `make fmt-check` and `make lint` in the lint stage, `make check`
cache hit; and PR #31's reviewer found `make docker` returning in the build stage — so a single declaration would have left one stage
success as a 17-layer cache hit, catching it only by being silently cacheable; it is declared in both. And BuildKit hashes the expanded
suspicious. The fix is `ARG CHECK_EPOCH` with the scripts passing command, not the declaration, so a declared-but-unreferenced `ARG` invalidates
`--build-arg CHECK_EPOCH="$(date +%s)"`. Two details make or break nothing: each gate `RUN` echoes the epoch, which also puts the value in the
it. `ARG` is scoped per stage and this `Dockerfile` has three gates build log as evidence the layer really ran. Placement is below the dependency
across two — `make fmt-check` and `make lint` in the lint stage, layers on purpose — a build that goes cold every time would be a different
`make check` in the build stage — so a single declaration would have bug, not a fix. Verified by running each script twice back to back on an
left one stage silently cacheable; it is declared in both. And unchanged tree under `BUILDKIT_PROGRESS=plain`: all three gates executed on
BuildKit hashes the expanded command, not the declaration, so a all four runs, each with a fresh epoch in the log (`script/cibuild` 78.8s then
declared-but-unreferenced `ARG` invalidates nothing: each gate `RUN` 61.1s; `script/docker` 61.1s then 53.4s), and twelve steps were still served
echoes the epoch, which also puts the value in the build log as `CACHED` in the steady state — both `go mod download`s, `apk add`, `adduser`,
evidence the layer really ran. Placement is below the dependency the `chown`, every `go.mod`/`go.sum` and source copy, the linter copy out of
layers on purpose — a build that goes cold every time would be a the lint stage, and the binary copy into the runtime stage. The lint stage
different bug, not a fix. Verified by running each script twice back still gates the build stage: with a deliberate `unused` finding planted in the
to back on an unchanged tree under `BUILDKIT_PROGRESS=plain`: all tree, the build failed at `make lint` in 36.1s and the build-stage
three gates executed on all four runs, each with a fresh epoch in `make check` never started. The build stage also still drops to the
the log (`script/cibuild` 78.8s then 61.1s; `script/docker` 61.1s unprivileged `builder` user before `make check`, which the suite depends on
then 53.4s), and twelve steps were still served `CACHED` in the rather than merely prefers: forcing the same image to run the tests as root
steady state — both `go mod download`s, `apk add`, `adduser`, the fails `TestScanHardlinkRunFailsTogether`, because root reads straight through
`chown`, every `go.mod`/`go.sum` and source copy, the linter copy the `chmod(0)` the test uses to prove hard links are read once. This is the
out of the lint stage, and the binary copy into the runtime stage. local fix only; propagating it to the canonical templates is `prompts` #26
The lint stage still gates the build stage: with a deliberate - check the installed golangci-lint version in `script/bootstrap` instead of
`unused` finding planted in the tree, the build failed at only its presence (2026-08-09, branch `bootstrap-version-check`, closes #24):
`make lint` in 36.1s and the build-stage `make check` never started. `missing golangci-lint` meant any linter already on `PATH` satisfied the
The build stage also still drops to the unprivileged `builder` user check, so the pin was never consulted and the v2.12.2 bump from #3 was inert
before `make check`, which the suite depends on rather than merely on every host that already had one — this host ran v2.10.1 against a v2.12.2
prefers: forcing the same image to run the tests as root fails pin, `make check` went green, and `make docker` then rejected the same commit
`TestScanHardlinkRunFailsTogether`, because root reads straight with findings the local gate never saw. The version now lives in one place,
through the `chmod(0)` the test uses to prove hard links are read `GOLANGCI_LINT_VERSION`, with the `go install` module ref derived from it so a
once. This is the local fix only; propagating it to the canonical bump cannot half-apply; a `golangci_lint_version` helper parses
templates is `prompts` #26 `golangci-lint --version` (taking the field after the word `version` and
- check the installed golangci-lint version in `script/bootstrap` tolerating an optional leading `v`, which the module ref carries and the
instead of only its presence (2026-08-09, branch binary's output does not), and any version that is not the pin — older, newer,
`bootstrap-version-check`, closes #24): `missing golangci-lint` meant absent or unparseable — is reinstalled. The install is then verified against
any linter already on `PATH` satisfied the check, so the pin was never the binary `PATH` actually resolves: `go install` writes into `GOBIN` (or
consulted and the v2.12.2 bump from #3 was inert on every host that `GOPATH/bin`) while `make lint` runs whichever `golangci-lint` comes first on
already had one — this host ran v2.10.1 against a v2.12.2 pin, `PATH`, so a wrong-version one sitting ahead of it — nix, apt, brew, apk, or
`make check` went green, and `make docker` then rejected the same the `/usr/local/bin` copy the `Dockerfile` builder stage makes — would swallow
commit with findings the local gate never saw. The version now lives the install and leave the local gate disagreeing with CI under an affirmative
in one place, `GOLANGCI_LINT_VERSION`, with the `go install` module `bootstrap complete`. Bootstrap now re-reads the effective version after
ref derived from it so a bump cannot half-apply; a installing and, on a mismatch, prints both paths and both versions to stderr
`golangci_lint_version` helper parses `golangci-lint --version` and exits non-zero instead of claiming success; it does not reorder anyone's
(taking the field after the word `version` and tolerating an optional
leading `v`, which the module ref carries and the binary's output does
not), and any version that is not the pin — older, newer, absent or
unparseable — is reinstalled. The install is then verified against the
binary `PATH` actually resolves: `go install` writes into `GOBIN` (or
`GOPATH/bin`) while `make lint` runs whichever `golangci-lint` comes
first on `PATH`, so a wrong-version one sitting ahead of it — nix,
apt, brew, apk, or the `/usr/local/bin` copy the `Dockerfile` builder
stage makes — would swallow the install and leave the local gate
disagreeing with CI under an affirmative `bootstrap complete`.
Bootstrap now re-reads the effective version after installing and, on
a mismatch, prints both paths and both versions to stderr and exits
non-zero instead of claiming success; it does not reorder anyone's
`PATH` or delete their binary. The `--version` call keeps its stderr `PATH` or delete their binary. The `--version` call keeps its stderr
connected, so a present-but-broken binary says why rather than connected, so a present-but-broken binary says why rather than reinstalling
reinstalling forever in silence, and is bounded by `timeout(1)` where forever in silence, and is bounded by `timeout(1)` where that exists, so a
that exists, so a wedged binary cannot hang bootstrap. `git`, `make` wedged binary cannot hang bootstrap. `git`, `make` and `go` keep their
and `go` keep their presence-only checks and now say why in a presence-only checks and now say why in a comment: they are host
comment: they are host package-manager tools the repo deliberately package-manager tools the repo deliberately does not pin, with `go.mod`
does not pin, with `go.mod` governing the language version and the governing the language version and the digest-pinned images covering
digest-pinned images covering reproducible builds. Verified on this reproducible builds. Verified on this host by bootstrapping from v2.10.1 to
host by bootstrapping from v2.10.1 to v2.12.2 and running it again to v2.12.2 and running it again to a no-op, plus stub runs of the real script
a no-op, plus stub runs of the real script under `dash` covering a under `dash` covering a thirteen-input parse matrix (absent, older, newer,
thirteen-input parse matrix (absent, older, newer, host-style, host-style, image-style, leading-`v`, stderr-only, empty, non-zero exit,
image-style, leading-`v`, stderr-only, empty, non-zero exit, impostor impostor binary, `(devel)`, trailing `version`), a shadowed install that must
binary, `(devel)`, trailing `version`), a shadowed install that must exit non-zero, an install destination not on `PATH` at all, `GOBIN` set, and a
exit non-zero, an install destination not on `PATH` at all, `GOBIN` wedged binary that must hit the timeout; `make check` and `make lint` are
set, and a wedged binary that must hit the timeout; `make check` and clean at v2.12.2, so v2.10.1 was not hiding any findings on `main`
`make lint` are clean at v2.12.2, so v2.10.1 was not hiding any
findings on `main`
- unwind the hash worker pool on the error path (2026-08-09, branch - unwind the hash worker pool on the error path (2026-08-09, branch
`hash-pool-cleanup`, closes #6): `hashPhase` used to return the `hash-pool-cleanup`, closes #6): `hashPhase` used to return the moment
moment `recordRun` failed and abandon the pool — the feeder parked `recordRun` failed and abandon the pool — the feeder parked forever on a full
forever on a full `jobs` channel and every worker on a full `jobs` channel and every worker on a full `results` channel. That only stopped
`results` channel. That only stopped being invisible when #4 landed being invisible when #4 landed and `runScan` began unwinding instead of
and `runScan` began unwinding instead of calling `os.Exit`. The calling `os.Exit`. The pool is now an owned, context-aware `hashPool`: every
pool is now an owned, context-aware `hashPool`: every blocking send blocking send in the feeder and the workers selects on `ctx.Done()`, `jobs` is
in the feeder and the workers selects on `ctx.Done()`, `jobs` is closed on every path out, and `hashPhase` defers `pool.stop()`, which cancels
closed on every path out, and `hashPhase` defers `pool.stop()`, and then drains `results` until the last goroutine has exited — draining is
which cancels and then drains `results` until the last goroutine what frees a worker already parked on a send. `ctx` is threaded from
has exited — draining is what frees a worker already parked on a `cmd.Context()` through `runScan`, `syncScan`, both worker pools and the whole
send. `ctx` is threaded from `cmd.Context()` through `runScan`, database layer (it is the first parameter everywhere), so #5 can hand this
`syncScan`, both worker pools and the whole database layer (it is path a signal and needs to add nothing else. The walk pool never leaked,
the first parameter everywhere), so #5 can hand this path a signal because `walkPhase` always drains its events to close, but it has the same
and needs to add nothing else. The walk pool never leaked, because unbounded-send shape and #5 will give it an early return, so it gets the same
`walkPhase` always drains its events to close, but it has the same treatment plus a `ctx.Err()` guard after the walk: a cancelled walk yields a
unbounded-send shape and #5 will give it an early return, so it partial size census, and every file it never reached looks vanished to the
gets the same treatment plus a `ctx.Err()` guard after the walk: a update phase. That phase's own `BeginTx` fails on the same cancelled context
cancelled walk yields a partial size census, and every file it never before deleting anything, so the guard is defence in depth rather than the
reached looks vanished to the update phase. That phase's own only barrier — but it is the one that survives #5 deciding an interrupted scan
`BeginTx` fails on the same cancelled context before deleting may commit what it has. Tests drive `run(scan)` against a database whose
anything, so the guard is defence in depth rather than the only insert trigger aborts, and assert both that the scan fails instead of hanging
barrier — but it is the one that survives #5 deciding an interrupted and that `runtime.NumGoroutine()` polls back to its pre-scan baseline; a
scan may commit what it has. Tests drive `run(scan)` against a second set cancels a scan part-way through the walk — deterministically, by
database whose insert trigger aborts, and assert both that the scan counting the scan's own consultations of `ctx.Done()` rather than racing a
fails instead of hanging and that `runtime.NumGoroutine()` polls timer — and asserts that it stops at the guard holding a partial census and a
back to its pre-scan baseline; a second set cancels a scan part-way still-populated record index, with every record intact. The remaining
through the walk — deterministically, by counting the scan's own cancellation branches of both pools are covered by direct tests of
consultations of `ctx.Done()` rather than racing a timer — and `sendEvent`, the walk workers, `dispatchDirs`, `feedHashJobs`, `hashWorker`
asserts that it stops at the guard holding a partial census and a and `hashPhase`
still-populated record index, with every record intact. The - guarantee the database is closed on every fatal exit path (2026-08-09, branch
remaining cancellation branches of both pools are covered by direct `db-close-on-fatal`, closes #4): `fatalf` and its `os.Exit(1)` are gone, so
tests of `sendEvent`, the walk workers, `dispatchDirs`, the deferred `db.Close()` — and with it the SQLite WAL checkpoint — now
`feedHashJobs`, `hashWorker` and `hashPhase` actually runs when a subcommand fails; `runScan`, `runReport`, `runTrees`,
- guarantee the database is closed on every fatal exit path `loadRecords` and `resolveRoots` return errors instead. The single exit point
(2026-08-09, branch `db-close-on-fatal`, closes #4): `fatalf` and is `run` in `main.go`: it maps a `fatalError` (anything a subcommand returned)
its `os.Exit(1)` are gone, so the deferred `db.Close()` — and with to exit 1 and cobra's own argument and flag errors to exit 2, which keeps a
it the SQLite WAL checkpoint — now actually runs when a subcommand runtime failure from being reported as a usage error or printing the usage
fails; `runScan`, `runReport`, `runTrees`, `loadRecords` and text. New `main_test.go` drives the CLI in-process and asserts the exit codes
`resolveRoots` return errors instead. The single exit point is `run` from README §Error handling plus the stdout/stderr split, including that a
in `main.go`: it maps a `fatalError` (anything a subcommand fatal error raised after the database is open leaves no `-wal`/`-shm` sidecar
returned) to exit 1 and cobra's own argument and flag errors to exit behind for `scan`, `report` or `trees`
2, which keeps a runtime failure from being reported as a usage - update golangci-lint to v2.12.2 with the canonical config (2026-08-09, branch
error or printing the usage text. New `main_test.go` drives the CLI `golangci-v2.12.2`, merged as `38a01bd`, closes #3): bumped the pinned linter
in-process and asserts the exit codes from README §Error handling in the `Dockerfile` lint stage and `script/bootstrap` from v2.12.1 to v2.12.2,
plus the stdout/stderr split, including that a fatal error raised and replaced `.golangci.yml` with the canonical file — the linter settings
after the database is open leaves no `-wal`/`-shm` sidecar behind
for `scan`, `report` or `trees`
- update golangci-lint to v2.12.2 with the canonical config
(2026-08-09, branch `golangci-v2.12.2`, merged as `38a01bd`,
closes #3): bumped the pinned linter in the `Dockerfile` lint
stage and `script/bootstrap` from v2.12.1 to v2.12.2, and replaced
`.golangci.yml` with the canonical file — the linter settings
(`lll`, `funlen`, `cyclop`, `dupl` thresholds) now live under (`lll`, `funlen`, `cyclop`, `dupl` thresholds) now live under
`linters.settings` per the v2 schema, so they are actually `linters.settings` per the v2 schema, so they are actually applied; no new
applied; no new lint findings surfaced lint findings surfaced
- convert Makefile targets to scripts-to-rule-them-all `script/` - convert Makefile targets to scripts-to-rule-them-all `script/` entrypoints
entrypoints like the other managed repos (2026-07-26, commit like the other managed repos (2026-07-26, commit `3abeacf`, closes #1): all 12
`3abeacf`, closes #1): all 12 `script/` entrypoints exist `script/` entrypoints exist (`bootstrap`, `setup`, `projectname`, `test`,
(`bootstrap`, `setup`, `projectname`, `test`, `lint`, `fmt`, `lint`, `fmt`, `fmt-check`, `check`, `docker`, `cibuild`, `precommit`,
`fmt-check`, `check`, `docker`, `cibuild`, `precommit`, `install-precommit`) and every Makefile target is now a thin shim over them,
`install-precommit`) and every Makefile target is now a thin shim matching the other managed repos
over them, matching the other managed repos
- make the binary the default Make target (2026-07-24, branch - make the binary the default Make target (2026-07-24, branch
`make-default-target`): plain `make` now builds `sfdupes` `make-default-target`): plain `make` now builds `sfdupes` (previously it ran
(previously it ran `check` plus `build`); `make build` remains as `check` plus `build`); `make build` remains as an alias
an alias - scan-wide phases, concurrent operands, batched updates (2026-07-24, branch
- scan-wide phases, concurrent operands, batched updates (2026-07-24, `scan-wide-phases`): all operands seed the shared walk pool and every pass
branch `scan-wide-phases`): all operands seed the shared walk pool runs once over the whole scan, so totals and ETAs are scan-global; the
and every pass runs once over the whole scan, so totals and ETAs per-operand walk/hash/update cycles and their stderr announcements are gone;
are scan-global; the per-operand walk/hash/update cycles and their the update pass commits in batched transactions — the filesystem is
stderr announcements are gone; the update pass commits in batched authoritative and the database an eventually-consistent reflection, so
transactions — the filesystem is authoritative and the database an scan-level atomicity is not required
eventually-consistent reflection, so scan-level atomicity is not
required
- split the stat pass back out of the walk (2026-07-24, branch - split the stat pass back out of the walk (2026-07-24, branch
`parallel-phases`): phases are strictly sequential again — walk, `parallel-phases`): phases are strictly sequential again — walk, stat, hash,
stat, hash, update per operand — with parallelism only inside each update per operand — with parallelism only inside each phase; the walk
phase; the walk enumerates paths with per-directory workers and the enumerates paths with per-directory workers and the stat pass lstats them with
stat pass lstats them with per-file workers, restoring the exact per-file workers, restoring the exact total/ETA stat bar
total/ETA stat bar - announce each operand on stderr before its passes (2026-07-24, branch
- announce each operand on stderr before its passes (2026-07-24, `scan-operand-progress`): with per-operand walk/hash/update cycles, a
branch `scan-operand-progress`): with per-operand walk/hash/update multi-operand run (e.g. `scan /srv/*`) showed pass totals that looked like the
cycles, a multi-operand run (e.g. `scan /srv/*`) showed pass totals whole run's — an operator watching operand 3 of 14 hash 300k files concluded
that looked like the whole run's — an operator watching operand 3 of 20M files were being skipped
14 hash 300k files concluded 20M files were being skipped - parallel walk (2026-07-24, branch `parallel-walk`): the walk pass was a single
- parallel walk (2026-07-24, branch `parallel-walk`): the walk pass goroutine and took hours at ~20M files on a busy pool (observed: 22M files in
was a single goroutine and took hours at ~20M files on a busy pool 4h on a ZFS server); it is now a per-directory worker-pool traversal that
(observed: 22M files in 4h on a ZFS server); it is now a records size/mtime during the walk (folding away the separate stat pass,
per-directory worker-pool traversal that records size/mtime during halving metadata I/O), and each `PATH` operand commits in its own transaction
the walk (folding away the separate stat pass, halving metadata so an interrupted scan keeps completed operands
I/O), and each `PATH` operand commits in its own transaction so an
interrupted scan keeps completed operands
- persistent scan database (2026-07-24, branch `persistent-database`): - persistent scan database (2026-07-24, branch `persistent-database`): `scan`
`scan` now maintains a SQLite database (`modernc.org/sqlite`, pure now maintains a SQLite database (`modernc.org/sqlite`, pure Go, cgo stays
Go, cgo stays disabled) keyed by absolute path that survives between disabled) keyed by absolute path that survives between runs — a rescan hashes
runs — a rescan hashes only new or changed files (by mtime/size), only new or changed files (by mtime/size), deletes records for files vanished
deletes records for files vanished from under the scanned operands, from under the scanned operands, and leaves records outside them untouched, so
and leaves records outside them untouched, so `scan` can be cronned `scan` can be cronned daily; `report` and `trees` read the database (no
daily; `report` and `trees` read the database (no positional positional arguments) instead of a scan stream. Database at
arguments) instead of a scan stream. Database at `/var/lib/sfdupes/db.sqlite`, overridable via `SFDUPES_DATABASE`; WAL
`/var/lib/sfdupes/db.sqlite`, overridable via `SFDUPES_DATABASE`; journaling plus a single-transaction update keep a report run during a scan
WAL journaling plus a single-transaction update keep a report run safe
during a scan safe - add the `origin` remote (`git@git.eeqj.de:sneak/sfdupes.git`), tag `v0.0.1`,
- add the `origin` remote (`git@git.eeqj.de:sneak/sfdupes.git`), tag and push `main` plus tags (2026-07-23)
`v0.0.1`, and push `main` plus tags (2026-07-23)
- `scan` CLI rework (2026-07-23, branch `scan-required-paths`): required - `scan` CLI rework (2026-07-23, branch `scan-required-paths`): required
`PATH...` operands via cobra flags replacing the `/srv` `-root` `PATH...` operands via cobra flags replacing the `/srv` `-root` default; new
default; new `-x`/`--one-file-system` flag (GNU convention) to stop `-x`/`--one-file-system` flag (GNU convention) to stop at filesystem
at filesystem boundaries, which are crossed by default boundaries, which are crossed by default
- bring the repo into full policy compliance (2026-07-23, branch - bring the repo into full policy compliance (2026-07-23, branch
`repo-policy-compliance`; checklist below) `repo-policy-compliance`; checklist below)
- `git init` with README-only first commit; code baseline committed on - `git init` with README-only first commit; code baseline committed on `main`
`main` (2026-07-22) (2026-07-22)
- implement `scan`, `report`, and `trees` subcommands (pre-git history) - implement `scan`, `report`, and `trees` subcommands (pre-git history)
# Future Steps # Future Steps
- possible later features (explicitly out of scope per README): - possible later features (explicitly out of scope per README): full-content
full-content verification of candidates, removal-script helpers verification of candidates, removal-script helpers
# Repo Policy Compliance # Repo Policy Compliance
Audited 2026-07-22 against `REPO_POLICIES.md` (2026-07-06), the existing Audited 2026-07-22 against `REPO_POLICIES.md` (2026-07-06), the existing repo
repo checklist, and the Go styleguide. Code is already gofmt-clean, so no checklist, and the Go styleguide. Code is already gofmt-clean, so no standalone
standalone formatting commit is needed. formatting commit is needed.
- [x] `.gitignore` missing — the compiled `sfdupes` binary and - [x] `.gitignore` missing — the compiled `sfdupes` binary and `files.dat` sit
`files.dat` sit untracked in the tree; needs OS/editor/Go untracked in the tree; needs OS/editor/Go artifacts plus secrets patterns
artifacts plus secrets patterns
- [x] `.editorconfig` missing - [x] `.editorconfig` missing
- [x] `LICENSE` missing and README has no License section (MIT assumed - [x] `LICENSE` missing and README has no License section (MIT assumed from
from house convention — user to confirm) house convention — user to confirm)
- [x] `REPO_POLICIES.md` missing from repo root - [x] `REPO_POLICIES.md` missing from repo root
- [x] `.golangci.yml` missing (install canonical copy); code must then - [x] `.golangci.yml` missing (install canonical copy); code must then pass
pass `make lint` (150 findings fixed; `make lint` is clean) `make lint` (150 findings fixed; `make lint` is clean)
- [x] `Makefile` lacks required targets `test`, `lint`, `fmt`, - [x] `Makefile` lacks required targets `test`, `lint`, `fmt`, `fmt-check`,
`fmt-check`, `docker`, `hooks`; `check` currently depends on `docker`, `hooks`; `check` currently depends on `build`, which writes the
`build`, which writes the binary (`make check` must not modify binary (`make check` must not modify files)
files) - [x] no tests — `go test ./...` has nothing to run; policy requires real tests
- [x] no tests — `go test ./...` has nothing to run; policy requires with a 30-second timeout and the conditional `-v` rerun pattern (suite
real tests with a 30-second timeout and the conditional `-v` covers parsing, grouping, digests, suppression, hashing, and the scan
rerun pattern (suite covers parsing, grouping, digests, pipeline; 64% coverage)
suppression, hashing, and the scan pipeline; 64% coverage) - [x] `Dockerfile` missing — Go multistage with hash-pinned images: fail-fast
- [x] `Dockerfile` missing — Go multistage with hash-pinned images: lint stage, build stage running `make check`
fail-fast lint stage, build stage running `make check`
- [x] `.dockerignore` missing - [x] `.dockerignore` missing
- [x] `.gitea/workflows/check.yml` missing (`docker build .` on push, - [x] `.gitea/workflows/check.yml` missing (`docker build .` on push, checkout
checkout action pinned by commit SHA) action pinned by commit SHA)
- [x] README lacks required sections: Description first line - [x] README lacks required sections: Description first line
(name/purpose/category/license/author), Getting Started, (name/purpose/category/license/author), Getting Started, Rationale, TODO,
Rationale, TODO, License, Author License, Author
- [x] README non-goal "no git repository setup and no CI" is stale now - [x] README non-goal "no git repository setup and no CI" is stale now that the
that the repo is under git with CI repo is under git with CI
- [x] pre-commit hook not installed (`make hooks` once the target - [x] pre-commit hook not installed (`make hooks` once the target exists)
exists)
Accepted divergences (no action): Accepted divergences (no action):
- flat single-package layout with `.go` files in the repo root — fine - flat single-package layout with `.go` files in the repo root — fine for a
for a small single-binary tool per the Go styleguide; the tracker small single-binary tool per the Go styleguide; the tracker audit agrees
audit agrees - `go test` runs without `-race` — the repo mandates `CGO_ENABLED=0` (pure-Go
- `go test` runs without `-race` — the repo mandates `CGO_ENABLED=0` builds) and the race detector requires cgo
(pure-Go builds) and the race detector requires cgo
+340 -36
View File
@@ -4,20 +4,35 @@ import (
"context" "context"
"database/sql" "database/sql"
"errors" "errors"
"fmt"
"os" "os"
"os/signal"
"path/filepath" "path/filepath"
"slices"
"strconv" "strconv"
"strings"
"sync" "sync"
"sync/atomic" "sync/atomic"
"syscall"
"testing" "testing"
"time" "time"
) )
// poolUnwind bounds how long a goroutine is given to leave a pool // This file gathers the tests for scan cancellation and worker-pool
// after its context is cancelled. Only a failing run ever waits this // unwinding. Everything it exercises lives in scan.go, so by the repo's
// long: a pool that ignored its cancellation parks forever, and this // convention of one test file per source file it would belong in
// is what turns that into a failed assertion instead of a suite that // scan_test.go. It is kept separate on purpose: cancellation behaviour
// hangs until the test binary's own timeout. // cuts across both the walk pool and the hash pool as a single concern,
// and scan_test.go is already over 1,600 lines. That is the deliberate
// exception the convention otherwise expects to be stated.
// poolUnwind bounds how long a test waits for a cancellation to take
// effect: for a goroutine to return or a channel to close once its
// context is cancelled, or for a signal to cancel the scan's context.
// Only a failing run waits this long, and the bound is what makes that
// failure an assertion instead of a hang. A call made without it, as
// most of this file's scans are, has no bound: a regression that parks
// it is caught only as the test binary's own timeout.
const poolUnwind = 2 * time.Second const poolUnwind = 2 * time.Second
// walkClock is a context whose cancellation is driven by the scan's // walkClock is a context whose cancellation is driven by the scan's
@@ -28,9 +43,14 @@ const poolUnwind = 2 * time.Second
// //
// The accounting behind the n chosen by each test: every blocking // The accounting behind the n chosen by each test: every blocking
// channel operation in the walk selects on Done, so the walk spends // channel operation in the walk selects on Done, so the walk spends
// one consultation per file event plus a couple per directory, while // one consultation per file event plus a couple per directory. The
// the index load that runs ahead of it spends a small fixed number // index load that runs ahead of it also consults Done, but a bounded
// (three) whatever the record count. // number of times that does not grow with the record count. The tests
// depend on that property, not on the bound's exact value: each test
// sets n from the consultations of the walk, plus those of the hash
// phase when it cancels mid-hash, far from both ends of the phase it
// interrupts, so the cancellation lands inside that phase whatever the
// record count.
type walkClock struct { type walkClock struct {
n int64 n int64
seen atomic.Int64 seen atomic.Int64
@@ -90,7 +110,10 @@ const (
walkCancelFilesPerDir = 20 walkCancelFilesPerDir = 20
walkCancelFiles = walkCancelDirs * walkCancelFilesPerDir walkCancelFiles = walkCancelDirs * walkCancelFilesPerDir
walkCancelWorkers = 4 walkCancelWorkers = 4
walkCancelInFlightDirs = walkCancelWorkers * walkCancelFilesPerDir // The most files the walkCancelWorkers directories already in
// flight when the scan is cancelled can still emit, at
// walkCancelFilesPerDir each. A file count, not a directory count.
walkCancelInFlightFiles = walkCancelWorkers * walkCancelFilesPerDir
) )
// walkCancelAtDone is the consultation on which the fixture's context // walkCancelAtDone is the consultation on which the fixture's context
@@ -150,9 +173,13 @@ func assertRecordsIntact(t *testing.T, db *sql.DB, before []string) {
// Every one of those records would look vanished to the update phase. // Every one of those records would look vanished to the update phase.
// The guard is what stops the scan there, and this test is what // The guard is what stops the scan there, and this test is what
// notices if it stops doing so: deleting the guard, or making it // notices if it stops doing so: deleting the guard, or making it
// unreachable, makes the scan carry its truncated view into a later // unreachable, makes the scan carry its truncated view into the update
// phase and fail there instead, with a wrapped error rather than the // phase, which counts every record the walk never reached for removal.
// bare cancellation. //
// The syncScan call here is not bounded by poolUnwind: a regression
// that left a worker pool parked would hang it, and that regression is
// caught only by the test binary's own timeout, not by a quick
// assertion.
// //
//nolint:paralleltest // counts goroutines: must not run beside others //nolint:paralleltest // counts goroutines: must not run beside others
func TestSyncScanCancelledMidWalkKeepsRecords(t *testing.T) { func TestSyncScanCancelledMidWalkKeepsRecords(t *testing.T) {
@@ -182,10 +209,10 @@ func TestSyncScanCancelledMidWalkKeepsRecords(t *testing.T) {
// assertWalkGuardAborted checks that the scan stopped at the post-walk // assertWalkGuardAborted checks that the scan stopped at the post-walk
// guard: with a census that is neither empty (the walk really ran) // guard: with a census that is neither empty (the walk really ran)
// nor complete (it really was cut short), and with the guard's own // nor complete (it really was cut short), and with no record counted
// bare cancellation as the error. A wrapped error means the partial // for removal. A removal count means the partial census was carried
// census was carried past the guard into the hash or update phase, // past the guard into the update phase, which is the failure this test
// which is the failure this test exists to catch. // exists to catch.
func assertWalkGuardAborted(t *testing.T, st scanStats, err error) { func assertWalkGuardAborted(t *testing.T, st scanStats, err error) {
t.Helper() t.Helper()
@@ -194,12 +221,6 @@ func assertWalkGuardAborted(t *testing.T, st scanStats, err error) {
err, context.Canceled) err, context.Canceled)
} }
if errors.Unwrap(err) != nil {
t.Errorf("syncScan reported %q, want the guard's bare "+
"cancellation: a wrapped error means the truncated census "+
"reached a later phase", err)
}
if st.unchanged == 0 { if st.unchanged == 0 {
t.Fatalf("stats = %+v: the census is empty, so the walk never "+ t.Fatalf("stats = %+v: the census is empty, so the walk never "+
"ran and the guard was reached for the wrong reason", st) "ran and the guard was reached for the wrong reason", st)
@@ -211,10 +232,11 @@ func assertWalkGuardAborted(t *testing.T, st scanStats, err error) {
} }
// The workers drop every directory still queued once the scan is // The workers drop every directory still queued once the scan is
// cancelled, so only the directories already in flight can add to // cancelled, so only the files in the directories already in flight
// the census after the fact. A census beyond that bound would mean // can add to the census after the fact. A census beyond that bound
// the cancellation was not observed where it should have been. // would mean the cancellation was not observed where it should have
limit := walkCancelAtDone + walkCancelInFlightDirs // been.
limit := walkCancelAtDone + walkCancelInFlightFiles
if st.unchanged > limit { if st.unchanged > limit {
t.Errorf("census covers %d files, want at most %d: the walk kept "+ t.Errorf("census covers %d files, want at most %d: the walk kept "+
"taking directories off the queue after cancellation", "taking directories off the queue after cancellation",
@@ -260,6 +282,242 @@ func TestSyncScanCancelledBeforeLoadIndex(t *testing.T) {
assertRecordsIntact(t, db, before) assertRecordsIntact(t, db, before)
} }
// hashCancelAtDone is the consultation on which the mid-hash test's
// context cancels itself. The walk of buildWalkCancelTree spends about
// one per file and three per directory, and the hash phase then one per
// file hashed, so this lands about half way through the hash phase.
const hashCancelAtDone = walkCancelFiles + 3*walkCancelDirs +
walkCancelFiles/2
// TestSyncScanCancelledMidHashKeepsHashedRecords cancels a first scan
// part-way through its hash phase. The fixture holds fewer files than a
// batch, so every file hashed is still waiting to be committed: the scan
// must commit them all before it returns, and the next scan must hash
// only the rest.
func TestSyncScanCancelledMidHashKeepsHashedRecords(t *testing.T) {
t.Parallel()
dir := buildWalkCancelTree(t)
db := openTestDB(t)
st, err := syncScan(newWalkClock(hashCancelAtDone), db,
[]string{dir}, walkCancelWorkers, false)
if !errors.Is(err, context.Canceled) {
t.Fatalf("syncScan cancelled mid-hash = %v, want %v",
err, context.Canceled)
}
if st.walked != walkCancelFiles || st.added == 0 ||
st.added >= walkCancelFiles {
t.Fatalf("stats = %+v: want the walk complete and the hash phase "+
"cut short", st)
}
if got := len(dbRecords(t, db)); got != st.added {
t.Errorf("%d records after the cancelled scan, want the %d it hashed",
got, st.added)
}
hashed := st.added
st = syncTree(t, db, dir)
if st.added != walkCancelFiles-hashed || st.unchanged != hashed {
t.Errorf("next scan stats = %+v, want %d added %d unchanged",
st, walkCancelFiles-hashed, hashed)
}
}
// storedPaths opens the database at path as report does, which fails
// unless it is a valid database, and returns its records' paths.
func storedPaths(t *testing.T, path string) []string {
t.Helper()
db, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
defer func() { _ = db.Close() }()
return recordPaths(dbRecords(t, db))
}
// TestRunScanInterrupted calls the scan entrypoint with a context that
// is already cancelled, as when a signal arrives at once. It must return
// errInterrupted promptly with its one line on stderr, leave the
// database valid and as it was, and leave nothing in the way of the
// next scan, which must bring the database up to date.
func TestRunScanInterrupted(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
stderr := captureStderr(t)
dir := buildSmokeTree(t)
err := runScan(t.Context(), []string{dir}, walkCancelWorkers, false)
if err != nil {
t.Fatal(err)
}
before := storedPaths(t, path)
// A vanished file and a new one: the interrupted scan records
// neither.
gone := filepath.Join(dir, "a", "unique.bin")
err = os.Remove(gone)
if err != nil {
t.Fatal(err)
}
added := writeFile(t, dir, "a/new.bin", pattern(50, 10))
shown := len(stderr())
done := make(chan struct{})
go func() {
defer close(done)
err = runScan(cancelledContext(t), []string{dir}, walkCancelWorkers,
false)
}()
awaitReturn(t, done, "runScan")
if !errors.Is(err, errInterrupted) {
t.Fatalf("runScan on a cancelled context = %v, want %v",
err, errInterrupted)
}
want := "scan: interrupted after 0 files\n"
if got := stderr()[shown:]; got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
assertNoSidecars(t, path)
if got := storedPaths(t, path); !slices.Equal(got, before) {
t.Errorf("records = %q after the interrupted scan, want %q",
got, before)
}
err = runScan(t.Context(), []string{dir}, walkCancelWorkers, false)
if err != nil {
t.Fatal(err)
}
got := storedPaths(t, path)
if slices.Contains(got, gone) || !slices.Contains(got, added) {
t.Errorf("records = %q after the next scan, want %q gone and %q "+
"added", got, gone, added)
}
}
// TestRunScanInterruptedMidHash interrupts the scan entrypoint part-way
// through its hash phase, after the database is open. It must return
// errInterrupted, release the lock, end stderr with its line counting
// every file the walk reached, close the database out of WAL mode, and
// keep the records it hashed.
func TestRunScanInterruptedMidHash(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
stderr := captureStderr(t)
dir := buildWalkCancelTree(t)
err := runScan(newWalkClock(hashCancelAtDone), []string{dir},
walkCancelWorkers, false)
if !errors.Is(err, errInterrupted) {
t.Fatalf("runScan interrupted mid-hash = %v, want %v",
err, errInterrupted)
}
holdScanLock(t, path)
want := fmt.Sprintf("scan: interrupted after %d files\n", walkCancelFiles)
if got := stderr(); !strings.HasSuffix(got, want) {
t.Errorf("stderr = %q, want it to end with %q", got, want)
}
assertNoSidecars(t, path)
db, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
defer func() { _ = db.Close() }()
// A plain close also removes the sidecars, but leaves WAL mode on.
var mode string
err = db.QueryRowContext(t.Context(), "PRAGMA journal_mode").Scan(&mode)
if err != nil {
t.Fatal(err)
}
if mode != "delete" {
t.Errorf("journal mode = %q after the interrupted scan, want %q",
mode, "delete")
}
kept := len(dbRecords(t, db))
if kept == 0 || kept >= walkCancelFiles {
t.Errorf("%d records after the interrupted scan, want those it "+
"hashed: some but not all of the %d files", kept, walkCancelFiles)
}
}
// TestInterruptContextCatchesSIGTERM sends SIGTERM to the test process
// while the scan's handler is installed, and checks that it cancels the
// scan's context.
//
//nolint:paralleltest // signals the whole process: must not run beside a scan
func TestInterruptContextCatchesSIGTERM(t *testing.T) {
// Caught here as well, so that a handler that misses SIGTERM fails
// this test instead of ending the test process.
caught := make(chan os.Signal, 1)
signal.Notify(caught, syscall.SIGTERM)
defer signal.Stop(caught)
ctx, stop := interruptContext(t.Context())
defer stop()
err := syscall.Kill(os.Getpid(), syscall.SIGTERM)
if err != nil {
t.Fatal(err)
}
select {
case <-ctx.Done():
case <-time.After(poolUnwind):
t.Fatal("SIGTERM did not cancel the scan's context")
}
}
// TestCommitFullBatchKeepsFailedBatch checks that a full batch whose
// commit fails, as it does once the scan is interrupted, stays in the
// batch, so that syncScan's final commit saves it.
func TestCommitFullBatchKeepsFailedBatch(t *testing.T) {
t.Parallel()
s := &scanState{db: openTestDB(t)}
for i := range updateBatchSize {
s.batch = append(s.batch, scanRec{path: "/f" + strconv.Itoa(i)})
}
err := s.commitFullBatch(cancelledContext(t))
if !errors.Is(err, context.Canceled) {
t.Fatalf("commitFullBatch on a cancelled context = %v, want %v",
err, context.Canceled)
}
if len(s.batch) != updateBatchSize {
t.Errorf("batch holds %d records after the failed commit, want %d",
len(s.batch), updateBatchSize)
}
}
// drainClosed counts the values received from ch until it closes, // drainClosed counts the values received from ch until it closes,
// failing the test if it does not close within poolUnwind. A pool that // failing the test if it does not close within poolUnwind. A pool that
// ignored its cancellation leaves its channel open with its goroutines // ignored its cancellation leaves its channel open with its goroutines
@@ -328,20 +586,49 @@ func TestSendEventAbandonsBlockedSend(t *testing.T) {
awaitReturn(t, done, "sendEvent") awaitReturn(t, done, "sendEvent")
} }
// TestWalkWorkersDropQueuedDirs checks that cancelled walk workers keep // TestWalkOneDirStopsWhenCancelled checks that a cancelled scan stops
// reading jobs and drop the directories rather than stopping their // reading a directory instead of going through the rest of its
// read: the range over jobs has to run out for the pool to tear down // entries. A walk that kept going would return the subdirectory below
// and close its event stream. // to descend into. Unlike a file event, that return is not a send the
func TestWalkWorkersDropQueuedDirs(t *testing.T) { // cancellation can abandon, so the test catches the regression every
// time.
func TestWalkOneDirStopsWhenCancelled(t *testing.T) {
t.Parallel() t.Parallel()
dir := t.TempDir() dir := t.TempDir()
writeEmptyFiles(t, dir, walkCancelFilesPerDir)
err := os.Mkdir(filepath.Join(dir, "sub"), 0o750)
if err != nil {
t.Fatal(err)
}
// Unbuffered and unread: on a cancelled scan every send gives up.
events := make(chan walkEvent)
subs := walkOneDir(cancelledContext(t), dirJob{path: dir}, false, events)
if len(subs) != 0 {
t.Errorf("cancelled walkOneDir returned %+v to descend into, "+
"want none", subs)
}
}
// TestWalkWorkersDropQueuedDirs checks that cancelled walk workers keep
// reading jobs and drop the directories rather than stopping their
// read: the range over jobs has to run out for the pool to tear down
// and close its event stream. The queued directory does not exist, so
// a worker that walked it anyway would send a warning before
// walkOneDir's own cancellation check could stop it. On a cancelled
// scan that send delivers or gives up at random, so with 64 jobs
// queued the regression has a one in 2^64 chance of passing.
func TestWalkWorkersDropQueuedDirs(t *testing.T) {
t.Parallel()
missing := filepath.Join(t.TempDir(), "missing")
jobs, _, events := startWalkWorkers(cancelledContext(t), 2, false) jobs, _, events := startWalkWorkers(cancelledContext(t), 2, false)
for range 4 { for range 64 {
jobs <- dirJob{path: dir} jobs <- dirJob{path: missing}
} }
close(jobs) close(jobs)
@@ -412,7 +699,11 @@ func TestDispatchDirsClosesJobsWhenCancelled(t *testing.T) {
// TestFeedHashJobsClosesJobsWhenCancelled checks that the hash feeder // TestFeedHashJobsClosesJobsWhenCancelled checks that the hash feeder
// abandons the runs it has not queued yet and still closes the job // abandons the runs it has not queued yet and still closes the job
// channel, which is what lets the workers' range terminate. // channel, which is what lets the workers' range terminate. The
// receive on jobs below is not bounded: a feeder that returned without
// closing jobs would leave that receive with no sender and no close, so
// this regression is caught by the test binary's timeout rather than by
// a bounded assertion.
func TestFeedHashJobsClosesJobsWhenCancelled(t *testing.T) { func TestFeedHashJobsClosesJobsWhenCancelled(t *testing.T) {
t.Parallel() t.Parallel()
@@ -440,6 +731,16 @@ func TestFeedHashJobsClosesJobsWhenCancelled(t *testing.T) {
// nobody wants the hashes of — while still letting the range run out // nobody wants the hashes of — while still letting the range run out
// so the pool tears down. The queued run names a file that does not // so the pool tears down. The queued run names a file that does not
// exist, so a worker that hashed it anyway would produce a result. // exist, so a worker that hashed it anyway would produce a result.
//
// hashWorker's other cancellation exit, abandoning the send of a
// result, is reachable from the scan: stop cancels the pool before it
// drains results, so a worker waiting on that send can leave through
// it. The tests that stop a scan mid-hash, among them
// TestScanHashWriteFailureUnwindsPool, reach it in some runs only,
// depending on timing, and no test fails without it, since stop's
// drain frees a waiting worker anyway. This test, for its part, catches
// a removed drop check in some runs only: a worker that hashes the run
// anyway then picks at random between sending the result and leaving.
func TestHashWorkerDropsQueuedRuns(t *testing.T) { func TestHashWorkerDropsQueuedRuns(t *testing.T) {
t.Parallel() t.Parallel()
@@ -472,7 +773,10 @@ func TestHashWorkerDropsQueuedRuns(t *testing.T) {
// TestHashPhaseCancelledReturnsContextError checks the result loop's // TestHashPhaseCancelledReturnsContextError checks the result loop's
// own exit: with the pool cancelled, no result will ever arrive, and // own exit: with the pool cancelled, no result will ever arrive, and
// the loop must leave through the cancellation rather than wait for a // the loop must leave through the cancellation rather than wait for a
// receive that cannot happen. // receive that cannot happen. This call is not bounded by poolUnwind: a
// loop that dropped its cancellation case would block on that receive,
// so the regression surfaces as the test binary's timeout rather than
// as a bounded assertion.
func TestHashPhaseCancelledReturnsContextError(t *testing.T) { func TestHashPhaseCancelledReturnsContextError(t *testing.T) {
t.Parallel() t.Parallel()
+230 -22
View File
@@ -11,6 +11,7 @@ import (
"slices" "slices"
"strconv" "strconv"
"golang.org/x/sys/unix"
// The pure-Go SQLite driver, registered as "sqlite"; keeps cgo // The pure-Go SQLite driver, registered as "sqlite"; keeps cgo
// disabled. // disabled.
_ "modernc.org/sqlite" _ "modernc.org/sqlite"
@@ -32,6 +33,11 @@ const schemaVersion = 1
// scan. // scan.
const dbDirPerm = 0o755 const dbDirPerm = 0o755
// lockFilePerm is the mode for the scan lock file. Anyone who can open
// the file can hold the lock and keep every scan from running, so it
// is open to its owner only.
const lockFilePerm = 0o600
// createTableSQL is the schema applied to a fresh database. Paths are // createTableSQL is the schema applied to a fresh database. Paths are
// BLOBs because Unix paths are raw bytes, not guaranteed UTF-8. // BLOBs because Unix paths are raw bytes, not guaranteed UTF-8.
const createTableSQL = ` const createTableSQL = `
@@ -45,6 +51,12 @@ CREATE TABLE files (
) WITHOUT ROWID ) WITHOUT ROWID
` `
// createIndexSQL indexes the records by signature, so report can have
// SQLite group them without sorting the whole table.
const createIndexSQL = `
CREATE INDEX files_signature ON files (size, head, tail, content)
`
// upsertSQL inserts one file record, replacing any existing record for // upsertSQL inserts one file record, replacing any existing record for
// the same path. // the same path.
const upsertSQL = ` const upsertSQL = `
@@ -64,6 +76,10 @@ var errNoDatabase = errors.New(
// does not understand. // does not understand.
var errSchemaVersion = errors.New("unsupported database schema version") var errSchemaVersion = errors.New("unsupported database schema version")
// errScanRunning reports that another scan holds the lock on the
// database.
var errScanRunning = errors.New("another scan is running")
// databasePath resolves the database location: SFDUPES_DATABASE when // databasePath resolves the database location: SFDUPES_DATABASE when
// set and non-empty, the compiled-in default otherwise. // set and non-empty, the compiled-in default otherwise.
func databasePath() string { func databasePath() string {
@@ -74,16 +90,24 @@ func databasePath() string {
return defaultDatabasePath return defaultDatabasePath
} }
// openDB opens the SQLite database at path with WAL journaling and a // scanParams are the connection parameters for scan: read-write, with
// busy timeout, so a report can run while a cron scan is in progress. // WAL journaling and a busy timeout, so a report can run while a cron
// It does not create or verify the schema. // scan is in progress. closeScanDatabase leaves WAL mode again.
func openDB(path string) (*sql.DB, error) { const scanParams = "_pragma=busy_timeout(10000)" +
dsn := "file:" + path +
"?_pragma=busy_timeout(10000)" +
"&_pragma=journal_mode(WAL)" + "&_pragma=journal_mode(WAL)" +
"&_pragma=synchronous(NORMAL)" "&_pragma=synchronous(NORMAL)"
db, err := sql.Open("sqlite", dsn) // reportParams are the connection parameters for report and trees:
// read-only, with the same busy timeout. They set no journal mode,
// because setting one is a write.
const reportParams = "mode=ro" +
"&_pragma=busy_timeout(10000)" +
"&_pragma=query_only(1)"
// openDB opens the SQLite database at path with the connection
// parameters params. It does not create or verify the schema.
func openDB(path, params string) (*sql.DB, error) {
db, err := sql.Open("sqlite", "file:"+path+"?"+params)
if err != nil { if err != nil {
return nil, fmt.Errorf("open database %s: %w", path, err) return nil, fmt.Errorf("open database %s: %w", path, err)
} }
@@ -96,6 +120,43 @@ func openDB(path string) (*sql.DB, error) {
return db, nil return db, nil
} }
// lockScanDatabase takes the lock that keeps a second scan off the
// database at path: an exclusive flock(2) on the file beside it named
// path with ".lock" appended, created along with the database's parent
// directory if missing. A lock held by another scan fails at once
// instead of waiting. The lock lasts until the returned file is closed
// or the process ends. The file is never deleted: a scan that deleted
// it would let the next scan lock a new file while another still holds
// the old one.
func lockScanDatabase(path string) (*os.File, error) {
err := os.MkdirAll(filepath.Dir(path), dbDirPerm)
if err != nil {
return nil, fmt.Errorf("create database directory: %w", err)
}
lockPath := path + ".lock"
//nolint:gosec // the operator chooses the database path
f, err := os.OpenFile(lockPath, os.O_RDWR|os.O_CREATE, lockFilePerm)
if err != nil {
return nil, err
}
err = unix.Flock(int(f.Fd()), unix.LOCK_EX|unix.LOCK_NB)
if err != nil {
_ = f.Close()
if errors.Is(err, unix.EWOULDBLOCK) {
return nil, fmt.Errorf("%w (lock held on %s)",
errScanRunning, lockPath)
}
return nil, fmt.Errorf("lock %s: %w", lockPath, err)
}
return f, nil
}
// openScanDatabase opens the database for the scan subcommand, creating // openScanDatabase opens the database for the scan subcommand, creating
// the file, its parent directory, and the schema as needed. // the file, its parent directory, and the schema as needed.
func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) { func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
@@ -104,7 +165,7 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return nil, fmt.Errorf("create database directory: %w", err) return nil, fmt.Errorf("create database directory: %w", err)
} }
db, err := openDB(path) db, err := openDB(path, scanParams)
if err != nil { if err != nil {
return nil, err return nil, err
} }
@@ -119,6 +180,24 @@ func openScanDatabase(ctx context.Context, path string) (*sql.DB, error) {
return db, nil return db, nil
} }
// closeScanDatabase switches the database at path from WAL back to
// rollback-journal mode and closes it. Out of WAL mode the database
// file alone holds the whole database, so a reader needs no -wal or
// -shm file beside it, nor write access to create them. The switch
// fails while a report has the database open; the database then stays
// in WAL mode, still readable, until a later scan closes it.
func closeScanDatabase(ctx context.Context, db *sql.DB, path string) {
// Runs on the way out of a cancelled scan too.
_, err := db.ExecContext(context.WithoutCancel(ctx),
"PRAGMA journal_mode = DELETE")
if err != nil {
fmt.Fprintf(os.Stderr, "scan: database %s left in WAL mode: %v\n",
path, err)
}
_ = db.Close()
}
// openReportDatabase opens an existing database for the report and // openReportDatabase opens an existing database for the report and
// trees subcommands. A missing database file is an error directing the // trees subcommands. A missing database file is an error directing the
// user to run scan first; the schema version must match exactly. // user to run scan first; the schema version must match exactly.
@@ -134,12 +213,18 @@ func openReportDatabase(ctx context.Context,
return nil, fmt.Errorf("database: %w", err) return nil, fmt.Errorf("database: %w", err)
} }
db, err := openDB(path) db, err := openDB(path, reportParams)
if err != nil { if err != nil {
return nil, err return nil, err
} }
v, err := userVersion(ctx, db) v, err := userVersion(ctx, db)
if err == nil && v == 0 {
// An empty database passes this check and fails the version
// check below.
err = checkUnversioned(ctx, db)
}
if err != nil { if err != nil {
_ = db.Close() _ = db.Close()
@@ -166,6 +251,11 @@ func initSchema(ctx context.Context, db *sql.DB) error {
switch v { switch v {
case 0: case 0:
err = checkUnversioned(ctx, db)
if err != nil {
return err
}
return createSchema(ctx, db) return createSchema(ctx, db)
case schemaVersion: case schemaVersion:
return nil return nil
@@ -175,20 +265,65 @@ func initSchema(ctx context.Context, db *sql.DB) error {
} }
} }
// checkUnversioned checks a database at user_version 0 before it is
// taken for an empty one. createSchema creates the files table and
// sets the version together, so a files table at version 0 was made by
// something else. Adopting it could corrupt unrelated data, so that is
// a schema-version error telling the operator to remove the file and
// rescan.
func checkUnversioned(ctx context.Context, db *sql.DB) error {
var name string
err := db.QueryRowContext(ctx,
"SELECT name FROM sqlite_master "+
"WHERE type = 'table' AND name = 'files'").Scan(&name)
switch {
case err == nil:
return fmt.Errorf(
"has a files table but no schema version; "+
"remove the file and rescan: %w", errSchemaVersion)
case errors.Is(err, sql.ErrNoRows):
return nil
default:
return fmt.Errorf("check for files table: %w", err)
}
}
// createSchema applies the schema to a fresh database and stamps the // createSchema applies the schema to a fresh database and stamps the
// schema version. // schema version in one transaction, so a creation stopped partway, by
// an interrupt or an error, leaves an empty database the next scan
// sets up, never a files table at version 0, which checkUnversioned
// refuses.
func createSchema(ctx context.Context, db *sql.DB) error { func createSchema(ctx context.Context, db *sql.DB) error {
_, err := db.ExecContext(ctx, createTableSQL) tx, err := db.BeginTx(ctx, nil)
if err != nil { if err != nil {
return fmt.Errorf("create schema: %w", err) return fmt.Errorf("create schema: %w", err)
} }
_, err = db.ExecContext(ctx, defer func() { _ = tx.Rollback() }()
_, err = tx.ExecContext(ctx, createTableSQL)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
_, err = tx.ExecContext(ctx, createIndexSQL)
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
_, err = tx.ExecContext(ctx,
"PRAGMA user_version = "+strconv.Itoa(schemaVersion)) "PRAGMA user_version = "+strconv.Itoa(schemaVersion))
if err != nil { if err != nil {
return fmt.Errorf("set schema version: %w", err) return fmt.Errorf("set schema version: %w", err)
} }
err = tx.Commit()
if err != nil {
return fmt.Errorf("create schema: %w", err)
}
return nil return nil
} }
@@ -204,18 +339,18 @@ func userVersion(ctx context.Context, db *sql.DB) (int, error) {
return v, nil return v, nil
} }
// loadFileRows reads every record from the files table. // loadFileRows streams every record to fn in path order: byte order,
func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) { // which is the order of the primary key, so SQLite does not sort.
func loadFileRows(ctx context.Context, db *sql.DB, fn func(r scanRec)) error {
rows, err := db.QueryContext(ctx, rows, err := db.QueryContext(ctx,
"SELECT path, size, mtime, head, tail, content FROM files") "SELECT path, size, mtime, head, tail, content FROM files "+
"ORDER BY path")
if err != nil { if err != nil {
return nil, fmt.Errorf("read records: %w", err) return fmt.Errorf("read records: %w", err)
} }
defer func() { _ = rows.Close() }() defer func() { _ = rows.Close() }()
var recs []scanRec
for rows.Next() { for rows.Next() {
var ( var (
path []byte path []byte
@@ -225,19 +360,92 @@ func loadFileRows(ctx context.Context, db *sql.DB) ([]scanRec, error) {
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail, err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail,
&r.content) &r.content)
if err != nil { if err != nil {
return nil, fmt.Errorf("read record: %w", err) return fmt.Errorf("read record: %w", err)
} }
r.path = string(path) r.path = string(path)
recs = append(recs, r) fn(r)
} }
err = rows.Err() err = rows.Err()
if err != nil { if err != nil {
return nil, fmt.Errorf("read records: %w", err) return fmt.Errorf("read records: %w", err)
} }
return recs, nil return nil
}
// dupeRowsSQL selects every record in a duplicate group, with the
// group's first path. A group is the records with a content hash that
// share a size, head, tail, and content, when there are two or more of
// them. The rows come in report order: groups by size descending, then
// by first path, and each group's paths ascending.
const dupeRowsSQL = `
SELECT g.first, f.path, f.size
FROM files AS f
JOIN (
SELECT size, head, tail, content, MIN(path) AS first
FROM files
WHERE content <> ''
GROUP BY size, head, tail, content
HAVING COUNT(*) > 1
) AS g USING (size, head, tail, content)
ORDER BY f.size DESC, g.first, f.path
`
// loadDupeRows streams the rows of dupeRowsSQL to fn and returns the
// number of records in the database. The count and the rows are read
// in one transaction, so they agree while a scan is committing. An
// error from fn stops the reading and is returned as it is.
func loadDupeRows(ctx context.Context, db *sql.DB,
fn func(first, path string, size int64) error,
) (int, error) {
// Everything goes through tx: the report connection is the only
// one, so a query on db would wait for tx forever.
tx, err := db.BeginTx(ctx, &sql.TxOptions{ReadOnly: true})
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
defer func() { _ = tx.Rollback() }()
var records int
err = tx.QueryRowContext(ctx, "SELECT COUNT(*) FROM files").Scan(&records)
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
rows, err := tx.QueryContext(ctx, dupeRowsSQL)
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
defer func() { _ = rows.Close() }()
for rows.Next() {
var (
first, path []byte
size int64
)
err = rows.Scan(&first, &path, &size)
if err != nil {
return 0, fmt.Errorf("read record: %w", err)
}
err = fn(string(first), string(path), size)
if err != nil {
return 0, err
}
}
err = rows.Err()
if err != nil {
return 0, fmt.Errorf("read records: %w", err)
}
return records, nil
} }
// loadFileMeta streams every record's path, size, mtime, and whether // loadFileMeta streams every record's path, size, mtime, and whether
+132 -25
View File
@@ -5,6 +5,7 @@ import (
"database/sql" "database/sql"
"errors" "errors"
"fmt" "fmt"
"os"
"path/filepath" "path/filepath"
"slices" "slices"
"strings" "strings"
@@ -72,9 +73,78 @@ func TestOpenScanDatabaseCreates(t *testing.T) {
defer func() { _ = db.Close() }() defer func() { _ = db.Close() }()
recs, err := loadFileRows(t.Context(), db) if recs := dbRecords(t, db); len(recs) != 0 {
if err != nil || len(recs) != 0 { t.Fatalf("records = %v, want none", recs)
t.Fatalf("loadFileRows = %v, %v; want empty, nil", recs, err) }
}
func TestOpenDatabaseUnversionedForeign(t *testing.T) {
t.Parallel()
// A database that has a files table but user_version 0, written by
// some other tool. report, trees and scan must refuse it with the
// schema-version error, not adopt it and not emit a raw SQLite
// "table files already exists".
path := testDBPath(t)
db, err := sql.Open("sqlite", path)
if err != nil {
t.Fatal(err)
}
_, err = db.ExecContext(t.Context(), "CREATE TABLE files (x INTEGER)")
if err != nil {
t.Fatal(err)
}
_ = db.Close()
_, err = openReportDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) ||
!strings.Contains(err.Error(), "remove the file and rescan") {
t.Fatalf("report: err = %v, want errSchemaVersion telling the "+
"operator to remove the file and rescan", err)
}
_, err = openScanDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) ||
!strings.Contains(err.Error(), "remove the file and rescan") {
t.Fatalf("scan: err = %v, want errSchemaVersion telling the "+
"operator to remove the file and rescan", err)
}
}
func TestSchemaCreationStoppedPartway(t *testing.T) {
t.Parallel()
// A first scan stopped while creating the schema must leave a
// database the next scan accepts. max_page_count(2) leaves room for
// the files table but not its index, so schema creation fails right
// after CREATE TABLE, a point an interrupt could also stop it at.
path := testDBPath(t)
db, err := openDB(path, scanParams+"&_pragma=max_page_count(2)")
if err != nil {
t.Fatal(err)
}
err = initSchema(t.Context(), db)
_ = db.Close()
if err == nil {
t.Fatal("initSchema with no room for the index succeeded")
}
db, err = openScanDatabase(t.Context(), path)
if err != nil {
t.Fatalf("next scan: %v", err)
}
defer func() { _ = db.Close() }()
v, err := userVersion(t.Context(), db)
if err != nil || v != schemaVersion {
t.Fatalf("userVersion = %d, %v; want %d, nil", v, err, schemaVersion)
} }
} }
@@ -87,9 +157,11 @@ func TestOpenReportDatabaseMissing(t *testing.T) {
} }
} }
func TestOpenReportDatabaseVersionMismatch(t *testing.T) { func TestOpenDatabaseVersionMismatch(t *testing.T) {
t.Parallel() t.Parallel()
// A database stamped with a schema version other than 0 and
// schemaVersion. report, trees and scan must all refuse it.
path := testDBPath(t) path := testDBPath(t)
db, err := openScanDatabase(t.Context(), path) db, err := openScanDatabase(t.Context(), path)
@@ -106,7 +178,12 @@ func TestOpenReportDatabaseVersionMismatch(t *testing.T) {
_, err = openReportDatabase(t.Context(), path) _, err = openReportDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) { if !errors.Is(err, errSchemaVersion) {
t.Fatalf("err = %v, want errSchemaVersion", err) t.Fatalf("report: err = %v, want errSchemaVersion", err)
}
_, err = openScanDatabase(t.Context(), path)
if !errors.Is(err, errSchemaVersion) {
t.Fatalf("scan: err = %v, want errSchemaVersion", err)
} }
} }
@@ -130,6 +207,49 @@ func TestOpenReportDatabaseOK(t *testing.T) {
_ = db.Close() _ = db.Close()
} }
func TestCloseScanDatabaseWhileReportOpen(t *testing.T) {
t.Parallel()
// A report holding the database open stops scan from taking it out
// of WAL mode. The -wal and -shm files must then stay beside it, so
// that a later report still needs only read access.
path := testDBPath(t)
scanDB, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
reportDB, err := openReportDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
closeScanDatabase(t.Context(), scanDB, path)
_ = reportDB.Close()
_, err = os.Stat(path + "-wal")
if err != nil {
t.Fatalf("no -wal left: the switch out of WAL mode was not "+
"stopped: %v", err)
}
makeReadOnly(t, path)
reportDB, err = openReportDatabase(t.Context(), path)
if err != nil {
t.Fatalf("openReportDatabase: %v", err)
}
defer func() { _ = reportDB.Close() }()
err = loadFileRows(t.Context(), reportDB, func(scanRec) {})
if err != nil {
t.Fatalf("loadFileRows: %v", err)
}
}
func TestApplyChangesRoundTrip(t *testing.T) { func TestApplyChangesRoundTrip(t *testing.T) {
t.Parallel() t.Parallel()
@@ -152,15 +272,8 @@ func TestApplyChangesRoundTrip(t *testing.T) {
t.Fatalf("applyChanges: %v", err) t.Fatalf("applyChanges: %v", err)
} }
got, err := loadFileRows(t.Context(), db) // The records come back in path order, which is the order of recs.
if err != nil { got := dbRecords(t, db)
t.Fatal(err)
}
slices.SortFunc(got, func(a, b scanRec) int {
return strings.Compare(a.path, b.path)
})
if !slices.Equal(got, recs) { if !slices.Equal(got, recs) {
t.Fatalf("rows = %+v, want %+v", got, recs) t.Fatalf("rows = %+v, want %+v", got, recs)
} }
@@ -177,11 +290,7 @@ func TestApplyChangesRoundTrip(t *testing.T) {
t.Fatalf("applyChanges: %v", err) t.Fatalf("applyChanges: %v", err)
} }
got, err = loadFileRows(t.Context(), db) got = dbRecords(t, db)
if err != nil {
t.Fatal(err)
}
if len(got) != 1 || got[0] != upd { if len(got) != 1 || got[0] != upd {
t.Fatalf("rows = %+v, want just %+v", got, upd) t.Fatalf("rows = %+v, want just %+v", got, upd)
} }
@@ -210,9 +319,8 @@ func TestApplyChangesBatching(t *testing.T) {
t.Fatalf("applyChanges: %v", err) t.Fatalf("applyChanges: %v", err)
} }
got, err := loadFileRows(t.Context(), db) if got := dbRecords(t, db); len(got) != n {
if err != nil || len(got) != n { t.Fatalf("records = %d, want %d", len(got), n)
t.Fatalf("loadFileRows = %d rows, %v; want %d", len(got), err, n)
} }
deletes := make([]string, 0, n) deletes := make([]string, 0, n)
@@ -226,8 +334,7 @@ func TestApplyChangesBatching(t *testing.T) {
t.Fatalf("applyChanges deletes: %v", err) t.Fatalf("applyChanges deletes: %v", err)
} }
got, err = loadFileRows(t.Context(), db) if got := dbRecords(t, db); len(got) != 0 {
if err != nil || len(got) != 0 { t.Fatalf("records = %d, want 0", len(got))
t.Fatalf("loadFileRows = %d rows, %v; want 0", len(got), err)
} }
} }
+2 -2
View File
@@ -5,6 +5,8 @@ go 1.25.7
require ( require (
github.com/schollz/progressbar/v3 v3.19.1 github.com/schollz/progressbar/v3 v3.19.1
github.com/spf13/cobra v1.10.2 github.com/spf13/cobra v1.10.2
golang.org/x/sys v0.46.0
golang.org/x/term v0.44.0
modernc.org/sqlite v1.54.0 modernc.org/sqlite v1.54.0
) )
@@ -18,8 +20,6 @@ require (
github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect github.com/remyoudompheng/bigfft v0.0.0-20230129092748-24d4a6f8daec // indirect
github.com/rivo/uniseg v0.4.7 // indirect github.com/rivo/uniseg v0.4.7 // indirect
github.com/spf13/pflag v1.0.9 // indirect github.com/spf13/pflag v1.0.9 // indirect
golang.org/x/sys v0.46.0 // indirect
golang.org/x/term v0.44.0 // indirect
modernc.org/libc v1.74.1 // indirect modernc.org/libc v1.74.1 // indirect
modernc.org/mathutil v1.7.1 // indirect modernc.org/mathutil v1.7.1 // indirect
modernc.org/memory v1.11.0 // indirect modernc.org/memory v1.11.0 // indirect
+70 -20
View File
@@ -15,6 +15,7 @@
// sfdupes scan [--workers N] [-x] PATH... // sfdupes scan [--workers N] [-x] PATH...
// sfdupes report > dupes.tsv // sfdupes report > dupes.tsv
// sfdupes trees > dupetrees.tsv // sfdupes trees > dupetrees.tsv
// sfdupes --version
// //
// See README.md for the complete specification. // See README.md for the complete specification.
package main package main
@@ -50,6 +51,10 @@ const (
// cobra prints for it is the whole message. // cobra prints for it is the whole message.
var errNoSubcommand = errors.New("no subcommand") var errNoSubcommand = errors.New("no subcommand")
// errWorkersBelowOne is the usage error for a scan --workers value
// below 1.
var errWorkersBelowOne = errors.New("--workers must be at least 1")
// Version is the build version, injected at link time via -ldflags // Version is the build version, injected at link time via -ldflags
// (see the Makefile); "dev" for a plain go build. // (see the Makefile); "dev" for a plain go build.
// //
@@ -57,22 +62,27 @@ var errNoSubcommand = errors.New("no subcommand")
var Version = "dev" var Version = "dev"
func main() { func main() {
os.Exit(run(os.Args[1:], os.Stderr)) // Once the reader of a stdout pipe has gone, as in "sfdupes report |
// head", the Go runtime ends the process with SIGPIPE on the next
// write instead of returning an error (README "Error handling").
// Registering for SIGPIPE with os/signal would change that.
os.Exit(run(os.Args[1:], os.Stdout, os.Stderr))
} }
// run executes args against the command tree and returns the process // run executes args against the command tree and returns the process
// exit code. It is the program's single exit point: the subcommands // exit code. It is the program's single exit point: the subcommands
// return their errors instead of exiting, so every deferred cleanup — // return their errors instead of exiting, so every deferred cleanup —
// above all closing the database, which checkpoints the SQLite WAL — // above all closing the database, which checkpoints the SQLite WAL —
// runs before the process ends. // runs before the process ends. The report and trees subcommands write
func run(args []string, stderr io.Writer) int { // their data to stdout.
func run(args []string, stdout, stderr io.Writer) int {
// A nil slice makes cobra fall back to os.Args, which would let a // A nil slice makes cobra fall back to os.Args, which would let a
// test binary's own flags reach the command tree. // test binary's own flags reach the command tree.
if args == nil { if args == nil {
args = []string{} args = []string{}
} }
root := newRootCommand(stderr) root := newRootCommand(stdout, stderr)
root.SetArgs(args) root.SetArgs(args)
err := root.Execute() err := root.Execute()
@@ -82,6 +92,9 @@ func run(args []string, stderr io.Writer) int {
switch { switch {
case err == nil: case err == nil:
return exitOK return exitOK
case errors.Is(err, errInterrupted):
// The interrupted scan has printed its own line.
return exitFatal
case errors.As(err, &fatal): case errors.As(err, &fatal):
// The command ran and failed: a runtime error, reported // The command ran and failed: a runtime error, reported
// without the usage text that a usage error gets. // without the usage text that a usage error gets.
@@ -89,22 +102,35 @@ func run(args []string, stderr io.Writer) int {
return exitFatal return exitFatal
default: default:
// A usage error: cobra has already printed the message and // A usage error, which cobra has already reported on stderr.
// the usage text.
return exitUsage return exitUsage
} }
} }
// newRootCommand builds the command tree. Everything on stdout is // newRootCommand builds the command tree. Everything on stdout is
// machine-readable data; all human-facing output (help, usage, errors) // machine-readable data, the version line included; all human-facing
// goes to stderr. // output (help, usage, errors) goes to stderr.
func newRootCommand(stderr io.Writer) *cobra.Command { func newRootCommand(stdout, stderr io.Writer) *cobra.Command {
var showVersion bool
printVersion := runE(func(context.Context, []string) error {
_, err := fmt.Fprintf(stdout, "sfdupes %s\n", Version)
if err != nil {
return fmt.Errorf("write stdout: %w", err)
}
return nil
})
root := &cobra.Command{ root := &cobra.Command{
Use: "sfdupes", Use: "sfdupes",
Short: "Find candidate duplicate files by size and head/tail/content SHA-256", Short: "Find candidate duplicate files by size and head/tail/content SHA-256",
Version: Version,
Args: cobra.NoArgs, Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, _ []string) error { RunE: func(cmd *cobra.Command, args []string) error {
if showVersion {
return printVersion(cmd, args)
}
// A missing subcommand prints usage and exits 2: cobra // A missing subcommand prints usage and exits 2: cobra
// prints the usage text for the returned error, and run // prints the usage text for the returned error, and run
// maps everything that is not a fatal error to exit 2. // maps everything that is not a fatal error to exit 2.
@@ -117,6 +143,11 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
root.SetErr(stderr) root.SetErr(stderr)
root.CompletionOptions.DisableDefaultCmd = true root.CompletionOptions.DisableDefaultCmd = true
// Cobra's built-in version flag prints through the help writer,
// stderr; this one prints to stdout.
root.Flags().BoolVarP(&showVersion, "version", "v", false,
"print the version to stdout")
var ( var (
scanWorkers int scanWorkers int
scanOneFS bool scanOneFS bool
@@ -126,7 +157,13 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Use: cmdScan + " [--workers N] [-x] PATH...", Use: cmdScan + " [--workers N] [-x] PATH...",
Short: "Walk trees and synchronize the scan database", Short: "Walk trees and synchronize the scan database",
Args: cobra.MinimumNArgs(1), Args: cobra.MinimumNArgs(1),
PreRunE: func(cmd *cobra.Command, _ []string) error {
return checkScanWorkers(cmd, scanWorkers)
},
RunE: runE(func(ctx context.Context, args []string) error { RunE: runE(func(ctx context.Context, args []string) error {
ctx, stop := interruptContext(ctx)
defer stop()
return runScan(ctx, args, scanWorkers, scanOneFS) return runScan(ctx, args, scanWorkers, scanOneFS)
}), }),
} }
@@ -140,7 +177,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the file-level duplicates report", Short: "Read the scan database and print the file-level duplicates report",
Args: cobra.NoArgs, Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error { RunE: runE(func(ctx context.Context, _ []string) error {
return runReport(ctx) return runReport(ctx, stdout)
}), }),
} }
@@ -149,7 +186,7 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
Short: "Read the scan database and print the duplicate-tree report", Short: "Read the scan database and print the duplicate-tree report",
Args: cobra.NoArgs, Args: cobra.NoArgs,
RunE: runE(func(ctx context.Context, _ []string) error { RunE: runE(func(ctx context.Context, _ []string) error {
return runTrees(ctx) return runTrees(ctx, stdout)
}), }),
} }
@@ -158,13 +195,26 @@ func newRootCommand(stderr io.Writer) *cobra.Command {
return root return root
} }
// runE adapts a subcommand implementation to cobra's RunE. Cobra // checkScanWorkers rejects a scan --workers value below 1. That is a
// prints the error and the command's usage text for every error RunE // usage error reported in one line: cobra prints the returned message
// returns, but a subcommand that ran and failed has no usage problem // without the usage text, and run exits 2.
// to report: both are silenced here, and the error is marked fatal so func checkScanWorkers(cmd *cobra.Command, workers int) error {
// that run reports it on stderr and exits 1 rather than 2. The command's if workers >= 1 {
// context is handed to the implementation: cancelling it unwinds the return nil
// scan's worker pools. }
cmd.SilenceUsage = true
return fmt.Errorf("%w, got %d", errWorkersBelowOne, workers)
}
// runE adapts a subcommand implementation, or the version print, to
// cobra's RunE. Cobra prints the error and the command's usage text for
// every error RunE returns, but a subcommand that ran and failed has no
// usage problem to report: both are silenced here, and the error is
// marked fatal so that run reports it on stderr and exits 1 rather than
// 2. The command's context is handed to the implementation: cancelling
// it unwinds the scan's worker pools.
func runE( func runE(
fn func(ctx context.Context, args []string) error, fn func(ctx context.Context, args []string) error,
) func(*cobra.Command, []string) error { ) func(*cobra.Command, []string) error {
+464 -69
View File
@@ -44,23 +44,61 @@ func assertNoSidecars(t *testing.T, path string) {
} }
} }
// captureStdout redirects os.Stdout to a file for the rest of the test // makeReadOnly takes write permission away from the database at path,
// and returns a function reading back everything written to it. Only // from any WAL sidecar beside it, and from their directory, as for a
// machine-readable data belongs on stdout (README design goal 4), so // user reading a database that a root cron scan keeps. Root ignores
// the tests assert on it directly. // file permissions, so it skips the test when run as root.
func captureStdout(t *testing.T) func() string { func makeReadOnly(t *testing.T, path string) {
t.Helper() t.Helper()
f, err := os.Create(filepath.Join(t.TempDir(), "stdout")) if os.Geteuid() == 0 {
t.Skip("root ignores file permissions")
}
err := os.Chmod(path, 0o400)
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
saved := os.Stdout for _, suffix := range walSuffixes {
os.Stdout = f err = os.Chmod(path+suffix, 0o400)
if err != nil && !errors.Is(err, fs.ErrNotExist) {
t.Fatal(err)
}
}
dir := filepath.Dir(path)
//nolint:gosec // reaching the database needs the search bit
err = os.Chmod(dir, 0o500)
if err != nil {
t.Fatal(err)
}
// Runs before t.TempDir's own cleanup, which must delete the files.
t.Cleanup(func() {
//nolint:gosec // removing the directory needs its search bit back
_ = os.Chmod(dir, 0o700)
})
}
// captureStderr redirects os.Stderr to a file for the rest of the test
// and returns a function reading back everything written to it. scan
// writes its warnings and summary straight to os.Stderr, not to the
// stderr writer run is given.
func captureStderr(t *testing.T) func() string {
t.Helper()
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
os.Stderr = f
t.Cleanup(func() { t.Cleanup(func() {
os.Stdout = saved os.Stderr = saved
_ = f.Close() _ = f.Close()
}) })
@@ -91,13 +129,13 @@ func captureStdout(t *testing.T) func() string {
// brokenDatabase writes a database that opens cleanly and passes the // brokenDatabase writes a database that opens cleanly and passes the
// schema-version check but has no files table, so the first query // schema-version check but has no files table, so the first query
// fails with the database already open: a fatal error on a path that // fails with the database already open: a fatal error on a path that
// owns an open database. // owns an open database. It closes the database the way scan does.
func brokenDatabase(t *testing.T) string { func brokenDatabase(t *testing.T) string {
t.Helper() t.Helper()
path := testDBPath(t) path := testDBPath(t)
db, err := openDB(path) db, err := openDB(path, scanParams)
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -108,10 +146,7 @@ func brokenDatabase(t *testing.T) string {
t.Fatal(err) t.Fatal(err)
} }
err = db.Close() closeScanDatabase(t.Context(), db, path)
if err != nil {
t.Fatal(err)
}
return path return path
} }
@@ -144,7 +179,10 @@ func TestOpenDatabaseKeepsWALWhileOpen(t *testing.T) {
func TestRunFatalAfterOpenClosesDatabase(t *testing.T) { func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
// Every subcommand that owns an open database must close it when // Every subcommand that owns an open database must close it when
// it fails: no os.Exit between the open and the return. // it fails: no os.Exit between the open and the return. The
// sidecar check is evidence of the close only for scan: report and
// trees only read a database that is out of WAL mode, which leaves
// nothing on disk whether they close it or not.
cases := map[string][]string{ cases := map[string][]string{
cmdScan: {cmdScan}, cmdScan: {cmdScan},
cmdReport: {cmdReport}, cmdReport: {cmdReport},
@@ -160,17 +198,15 @@ func TestRunFatalAfterOpenClosesDatabase(t *testing.T) {
args = append(args, t.TempDir()) args = append(args, t.TempDir())
} }
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run(args, &stdout, &stderr)
code := run(args, &stderr)
if code != exitFatal { if code != exitFatal {
t.Errorf("run(%v) = %d, want %d", args, code, exitFatal) t.Errorf("run(%v) = %d, want %d", args, code, exitFatal)
} }
assertNoSidecars(t, path) assertNoSidecars(t, path)
assertFatalOutput(t, stderr.String(), stdout()) assertFatalOutput(t, stderr.String(), stdout.String())
// Proof that the failure happened after the open: only a // Proof that the failure happened after the open: only a
// query against the opened database can report this. // query against the opened database can report this.
@@ -188,18 +224,16 @@ func TestRunMissingOperandIsFatalNotUsage(t *testing.T) {
// must not dump the usage text. // must not dump the usage text.
t.Setenv(databaseEnv, testDBPath(t)) t.Setenv(databaseEnv, testDBPath(t))
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t)
missing := filepath.Join(t.TempDir(), "nope") missing := filepath.Join(t.TempDir(), "nope")
code := run([]string{cmdScan, missing}, &stderr) code := run([]string{cmdScan, missing}, &stdout, &stderr)
if code != exitFatal { if code != exitFatal {
t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal) t.Errorf("run(scan %s) = %d, want %d", missing, code, exitFatal)
} }
assertFatalOutput(t, stderr.String(), stdout()) assertFatalOutput(t, stderr.String(), stdout.String())
} }
// assertFatalOutput checks that a fatal error was reported the way // assertFatalOutput checks that a fatal error was reported the way
@@ -243,11 +277,9 @@ func TestRunUsageErrors(t *testing.T) {
// path that does not exist. // path that does not exist.
t.Setenv(databaseEnv, testDBPath(t)) t.Setenv(databaseEnv, testDBPath(t))
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run(tc.args, &stdout, &stderr)
code := run(tc.args, &stderr)
if code != exitUsage { if code != exitUsage {
t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage) t.Errorf("run(%v) = %d, want %d", tc.args, code, exitUsage)
} }
@@ -256,43 +288,115 @@ func TestRunUsageErrors(t *testing.T) {
t.Errorf("stderr = %q, want %q", stderr.String(), tc.want) t.Errorf("stderr = %q, want %q", stderr.String(), tc.want)
} }
if got := stdout(); got != "" { if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got) t.Errorf("stdout = %q, want nothing (data only)", got)
} }
}) })
} }
} }
// TestRunHelpAndVersionSucceed checks that the two informational flags func TestRunScanRejectsWorkersBelowOne(t *testing.T) {
// exit 0 and keep their human-facing output on stderr. // README §scan mode: --workers below 1 is a usage error reported in
// // one line on stderr, before the scan opens the database.
//nolint:paralleltest // captureStdout replaces the process-wide os.Stdout for _, workers := range []string{"0", "-1"} {
func TestRunHelpAndVersionSucceed(t *testing.T) { t.Run(workers, func(t *testing.T) {
assertHumanOutput(t, "--help") dbPath := testDBPath(t)
assertHumanOutput(t, "--version") t.Setenv(databaseEnv, dbPath)
var stdout, stderr bytes.Buffer
args := []string{cmdScan, "--workers", workers, t.TempDir()}
code := run(args, &stdout, &stderr)
if code != exitUsage {
t.Errorf("run(%v) = %d, want %d", args, code, exitUsage)
} }
// assertHumanOutput runs sfdupes with one informational flag and checks want := "Error: --workers must be at least 1, got " + workers +
// that it succeeds with its output on stderr and stdout untouched "\n"
// (README design goal 4). if got := stderr.String(); got != want {
func assertHumanOutput(t *testing.T, arg string) { t.Errorf("stderr = %q, want %q", got, want)
t.Helper() }
var stderr bytes.Buffer if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
stdout := captureStdout(t) _, err := os.Stat(dbPath)
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("stat %s: %v, want the database never created",
dbPath, err)
}
})
}
}
code := run([]string{arg}, &stderr) func TestRunHelp(t *testing.T) {
t.Parallel()
// README §Subcommands: help goes to stderr, exits 0, and leaves
// stdout empty.
cases := [][]string{{"--help"}, {"-h"}, {cmdScan, "--help"}}
for _, args := range cases {
var stdout, stderr bytes.Buffer
code := run(args, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%v) = %d, want %d", args, code, exitOK)
}
if !strings.Contains(stderr.String(), usageMarker) {
t.Errorf("run(%v) stderr = %q, want the help text",
args, stderr.String())
}
if got := stdout.String(); got != "" {
t.Errorf("run(%v) stdout = %q, want nothing (data only)",
args, got)
}
}
}
func TestRunVersion(t *testing.T) {
t.Parallel()
// README §Subcommands: the version is one line on stdout, with
// nothing on stderr, and exits 0.
for _, arg := range []string{"--version", "-v"} {
var stdout, stderr bytes.Buffer
code := run([]string{arg}, &stdout, &stderr)
if code != exitOK { if code != exitOK {
t.Errorf("run(%s) = %d, want %d", arg, code, exitOK) t.Errorf("run(%s) = %d, want %d", arg, code, exitOK)
} }
if stderr.Len() == 0 { want := "sfdupes " + Version + "\n"
t.Errorf("run(%s) wrote nothing to stderr", arg) if got := stdout.String(); got != want {
t.Errorf("run(%s) stdout = %q, want %q", arg, got, want)
} }
if got := stdout(); got != "" { if got := stderr.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got) t.Errorf("run(%s) stderr = %q, want nothing", arg, got)
}
}
}
func TestRunVersionWriteFailureIsFatal(t *testing.T) {
t.Parallel()
// README §Error handling: a stdout write failure exits 1, reported
// in one line on stderr.
var stderr bytes.Buffer
code := run([]string{"--version"}, failingWriter{}, &stderr)
if code != exitFatal {
t.Errorf("run(--version) = %d, want %d", code, exitFatal)
}
want := "sfdupes: write stdout: " + errWriteFailed.Error() + "\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
} }
} }
@@ -320,21 +424,31 @@ func scanFixture(t *testing.T) []string {
t.Fatal(err) t.Fatal(err)
} }
var stderr bytes.Buffer scanOK(t, dir)
stdout := captureStdout(t) return dupes
code := run([]string{cmdScan, dir}, &stderr)
if code != exitOK {
t.Fatalf("run(scan) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
} }
if got := stdout(); got != "" { // scanOK runs scan over operands, fails the test unless it exits 0 with
// nothing on stdout, and returns everything it printed to stderr.
func scanOK(t *testing.T, operands ...string) string {
t.Helper()
var stdout bytes.Buffer
stderr := captureStderr(t)
code := run(append([]string{cmdScan}, operands...), &stdout, os.Stderr)
if code != exitOK {
t.Fatalf("run(scan %q) = %d, want %d; stderr: %s",
operands, code, exitOK, stderr())
}
if got := stdout.String(); got != "" {
t.Errorf("scan stdout = %q, want nothing (data only)", got) t.Errorf("scan stdout = %q, want nothing (data only)", got)
} }
return dupes return stderr()
} }
func TestRunScanSucceedsDespiteWarnings(t *testing.T) { func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
@@ -345,24 +459,119 @@ func TestRunScanSucceedsDespiteWarnings(t *testing.T) {
assertNoSidecars(t, path) assertNoSidecars(t, path)
} }
func TestRunScanSkipsSymlinkOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dir := t.TempDir()
writeFile(t, dir, "target/sub/f", pattern(1, 10))
link := filepath.Join(dir, "link")
err := os.Symlink(filepath.Join(dir, "target"), link)
if err != nil {
t.Fatal(err)
}
// Scanning a directory through the symlink stores a record beneath
// the symlink's own path for a file beneath its target.
scanOK(t, filepath.Join(link, "sub"))
assertOperandSkipped(t, path, link, "symlink",
filepath.Join(link, "sub", "f"))
}
func TestRunScanWalksOperandUnderSymlinkOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dir := t.TempDir()
writeFile(t, dir, "target/sub/f", pattern(1, 10))
link := filepath.Join(dir, "link")
err := os.Symlink(filepath.Join(dir, "target"), link)
if err != nil {
t.Fatal(err)
}
// link is dropped as a symlink, but link/sub must still be scanned,
// not dropped as lying under link.
scanOK(t, link, filepath.Join(link, "sub"))
db, err := openDB(path, reportParams)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
recordByPath(t, dbRecords(t, db), filepath.Join(link, "sub", "f"))
}
func TestRunScanSkipsZFSOperand(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
zfs := filepath.Join(t.TempDir(), ".zfs")
snapshot := filepath.Join(zfs, "snapshot", "hourly")
f := writeFile(t, snapshot, "f", pattern(1, 10))
// An operand beneath a .zfs directory is walked, because it is not
// itself named .zfs.
scanOK(t, snapshot)
assertOperandSkipped(t, path, zfs, ".zfs directory", f)
}
// assertOperandSkipped scans operand alone and checks that it is skipped
// as kind: a warning naming it, one skip in the summary, exit 0, and the
// record for kept, which an earlier scan stored beneath operand, still
// in the database at dbPath.
func assertOperandSkipped(t *testing.T, dbPath, operand, kind,
kept string,
) {
t.Helper()
stderr := scanOK(t, operand)
warning := "walk " + operand + ": skipping " + kind + " operand\n"
if !strings.Contains(stderr, warning) {
t.Errorf("stderr = %q, want %q", stderr, warning)
}
summary := "scan: 0 files seen (0 added, 0 updated, 0 removed, " +
"0 unchanged), 1 skipped\n"
if !strings.Contains(stderr, summary) {
t.Errorf("stderr = %q, want %q", stderr, summary)
}
db, err := openDB(dbPath, reportParams)
if err != nil {
t.Fatal(err)
}
t.Cleanup(func() { _ = db.Close() })
recordByPath(t, dbRecords(t, db), kept)
}
func TestRunReportSucceeds(t *testing.T) { func TestRunReportSucceeds(t *testing.T) {
path := testDBPath(t) path := testDBPath(t)
t.Setenv(databaseEnv, path) t.Setenv(databaseEnv, path)
dupes := scanFixture(t) dupes := scanFixture(t)
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run([]string{cmdReport}, &stdout, &stderr)
code := run([]string{cmdReport}, &stderr)
if code != exitOK { if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s", t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String()) code, exitOK, stderr.String())
} }
want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n" want := "first\tdupe\tsize\n" + dupes[0] + "\t" + dupes[1] + "\t300\n"
if got := stdout(); got != want { if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want) t.Errorf("stdout = %q, want %q", got, want)
} }
@@ -375,11 +584,9 @@ func TestRunTreesSucceeds(t *testing.T) {
dupes := scanFixture(t) dupes := scanFixture(t)
var stderr bytes.Buffer var stdout, stderr bytes.Buffer
stdout := captureStdout(t) code := run([]string{cmdTrees}, &stdout, &stderr)
code := run([]string{cmdTrees}, &stderr)
if code != exitOK { if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s", t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String()) code, exitOK, stderr.String())
@@ -389,9 +596,197 @@ func TestRunTreesSucceeds(t *testing.T) {
// trees of each other. // trees of each other.
want := "first\tdupe\tfiles\tsize\n" + want := "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n" filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) + "\t1\t300\n"
if got := stdout(); got != want { if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want) t.Errorf("stdout = %q, want %q", got, want)
} }
assertNoSidecars(t, path) assertNoSidecars(t, path)
} }
func TestRunReportsNeedOnlyReadAccess(t *testing.T) {
// README §Database: report and trees need only read access to the
// database file. With its directory read-only as well, SQLite
// cannot create any file beside it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
dupes := scanFixture(t)
assertNoSidecars(t, path)
makeReadOnly(t, path)
cases := map[string]string{
cmdReport: "first\tdupe\tsize\n" +
dupes[0] + "\t" + dupes[1] + "\t300\n",
cmdTrees: "first\tdupe\tfiles\tsize\n" +
filepath.Dir(dupes[0]) + "\t" + filepath.Dir(dupes[1]) +
"\t1\t300\n",
}
for name, want := range cases {
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
continue
}
if got := stdout.String(); got != want {
t.Errorf("%s stdout = %q, want %q", name, got, want)
}
}
}
// holdScanLock takes the lock on the database at path, as a running
// scan does, and holds it until the test ends. It fails the test when
// the lock is already held.
func holdScanLock(t *testing.T, path string) {
t.Helper()
lock, err := lockScanDatabase(path)
if err != nil {
t.Fatalf("lock %s: %v", path, err)
}
t.Cleanup(func() { _ = lock.Close() })
}
func TestRunSecondScanFails(t *testing.T) {
// README §Database: while one scan holds the lock, a second scan
// fails at once, naming the lock file, without creating the
// database.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
holdScanLock(t, path)
var stdout, stderr bytes.Buffer
code := run([]string{cmdScan, t.TempDir()}, &stdout, &stderr)
if code != exitFatal {
t.Errorf("run(scan) = %d, want %d", code, exitFatal)
}
want := "sfdupes: another scan is running (lock held on " +
path + ".lock)\n"
if got := stderr.String(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
if got := stdout.String(); got != "" {
t.Errorf("stdout = %q, want nothing (data only)", got)
}
_, err := os.Stat(path)
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("stat %s = %v, want the database not created", path, err)
}
}
func TestRunScanReleasesLock(t *testing.T) {
// README §Database: a scan releases the lock however it ends.
t.Run("success", func(t *testing.T) {
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
})
t.Run("fatal error", func(t *testing.T) {
path := brokenDatabase(t)
t.Setenv(databaseEnv, path)
code := run([]string{cmdScan, t.TempDir()}, io.Discard, io.Discard)
if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d", code, exitFatal)
}
holdScanLock(t, path)
})
}
func TestRunReportsDuringScan(t *testing.T) {
// README §Database: report and trees never take the lock, so they
// run while a scan holds it.
path := testDBPath(t)
t.Setenv(databaseEnv, path)
scanFixture(t)
holdScanLock(t, path)
for _, name := range []string{cmdReport, cmdTrees} {
var stderr bytes.Buffer
code := run([]string{name}, io.Discard, &stderr)
if code != exitOK {
t.Errorf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
}
}
}
func TestRunStdoutClosedIsFatal(t *testing.T) {
// README §Error handling: a stdout write failure exits 1, reported
// in one line on stderr.
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
stdout, err := os.Create(filepath.Join(t.TempDir(), "stdout"))
if err != nil {
t.Fatal(err)
}
err = stdout.Close()
if err != nil {
t.Fatal(err)
}
var stderr bytes.Buffer
code := run([]string{name}, stdout, &stderr)
if code != exitFatal {
t.Errorf("run(%s) = %d, want %d", name, code, exitFatal)
}
got := stderr.String()
if !strings.HasPrefix(got, "sfdupes: write stdout: ") ||
!strings.Contains(got, os.ErrClosed.Error()) ||
strings.Count(got, "\n") != 1 {
t.Errorf("stderr = %q, want one line reporting the "+
"failed stdout write", got)
}
})
}
}
// errWriteFailed is the error failingWriter returns.
var errWriteFailed = errors.New("write failed")
// failingWriter is a stdout that fails every write.
type failingWriter struct{}
func (failingWriter) Write([]byte) (int, error) { return 0, errWriteFailed }
func TestStdoutWriteErrorPropagates(t *testing.T) {
t.Setenv(databaseEnv, testDBPath(t))
scanFixture(t)
cases := map[string]func(context.Context, io.Writer) error{
cmdReport: runReport,
cmdTrees: runTrees,
}
for name, fn := range cases {
err := fn(t.Context(), failingWriter{})
if !errors.Is(err, errWriteFailed) {
t.Errorf("%s: error = %v, want %v", name, err, errWriteFailed)
}
}
}
+5
View File
@@ -0,0 +1,5 @@
{
"devDependencies": {
"prettier": "3.8.1"
}
}
+46 -17
View File
@@ -6,6 +6,7 @@ import (
"time" "time"
"github.com/schollz/progressbar/v3" "github.com/schollz/progressbar/v3"
"golang.org/x/term"
) )
// plainInterval is the minimum time between progress lines when stderr // plainInterval is the minimum time between progress lines when stderr
@@ -24,24 +25,23 @@ const percentScale = 100
// stderrIsTTY reports whether stderr is attached to a terminal. // stderrIsTTY reports whether stderr is attached to a terminal.
func stderrIsTTY() bool { func stderrIsTTY() bool {
fi, err := os.Stderr.Stat() return term.IsTerminal(int(os.Stderr.Fd()))
if err != nil {
return false
}
return fi.Mode()&os.ModeCharDevice != 0
} }
// progress renders one scan pass's progress on stderr. On a TTY it // progress renders one scan pass's progress on stderr. On a TTY it
// delegates to the progressbar library (spinner style when the total is // delegates to the progressbar library (spinner style when the total is
// unknown, full bar with count/percent/rate/elapsed/ETA otherwise). When // unknown, full bar with count/percent/rate/elapsed/ETA otherwise). When
// stderr is not a TTY it emits no ANSI redraws: it prints a plain // stderr is not a TTY it emits no ANSI redraws: it prints a plain
// one-line update no more often than every plainInterval. // one-line update as the pass starts, then no more often than every
// plainInterval.
// //
// All methods must be called from the main goroutine only. A nil // All methods must be called from the main goroutine only. On a TTY
// *progress is a valid no-display receiver: every method is a no-op, // the library also redraws a spinner from its own goroutine, several
// so batched database flushes during the streaming pass can reuse the // times a second, so its count and elapsed time stay current while a
// update-pass helpers without rendering anything. // pass waits for its next item. A nil *progress is a valid
// no-display receiver: every method is a no-op, so batched database
// flushes during the streaming pass can reuse the update-pass helpers
// without rendering anything.
type progress struct { type progress struct {
label string label string
total int64 // -1 when unknown (walk pass) total int64 // -1 when unknown (walk pass)
@@ -53,10 +53,22 @@ type progress struct {
func newProgress(label string, total int64) *progress { func newProgress(label string, total int64) *progress {
p := &progress{label: label, total: total, start: time.Now()} p := &progress{label: label, total: total, start: time.Now()}
if !stderrIsTTY() { if stderrIsTTY() {
p.bar = newBar(label, total)
return p return p
} }
// Print the zero state at once: the first item may take minutes,
// and a pass must never look hung.
p.last = p.start
fmt.Fprintln(os.Stderr, p.plainLine())
return p
}
// newBar builds the TTY display for newProgress.
func newBar(label string, total int64) *progressbar.ProgressBar {
opts := []progressbar.Option{ opts := []progressbar.Option{
progressbar.OptionSetWriter(os.Stderr), progressbar.OptionSetWriter(os.Stderr),
progressbar.OptionSetDescription(label), progressbar.OptionSetDescription(label),
@@ -81,9 +93,7 @@ func newProgress(label string, total int64) *progress {
) )
} }
p.bar = progressbar.NewOptions64(total, opts...) return progressbar.NewOptions64(total, opts...)
return p
} }
// increment records one completed item and refreshes the display. // increment records one completed item and refreshes the display.
@@ -106,26 +116,45 @@ func (p *progress) increment() {
} }
// warnf prints a one-line warning to stderr without corrupting the bar. // warnf prints a one-line warning to stderr without corrupting the bar.
// The whole message is escaped like a report's path columns, so a path
// holding a newline cannot split the warning.
func (p *progress) warnf(format string, args ...any) { func (p *progress) warnf(format string, args ...any) {
if p == nil { if p == nil {
return return
} }
msg := escapePath(fmt.Sprintf(format, args...))
if p.bar != nil && p.total < 0 {
// The library also redraws a spinner from its own goroutine, so
// a direct write could land inside a redraw. The bar prints the
// warning itself, just before its next redraw.
_, _ = progressbar.Bprintln(p.bar, msg)
return
}
if p.bar != nil { if p.bar != nil {
_ = p.bar.Clear() _ = p.bar.Clear()
} }
fmt.Fprintf(os.Stderr, format+"\n", args...) fmt.Fprintln(os.Stderr, msg)
} }
// finish terminates the pass's display. // finish terminates the pass's display. A bar whose pass stopped short
// of its total, as an interrupted one does, is left as last drawn; the
// library's Finish would fill it up.
func (p *progress) finish() { func (p *progress) finish() {
if p == nil { if p == nil {
return return
} }
if p.bar != nil { if p.bar != nil {
if p.total >= 0 && p.count < p.total {
_ = p.bar.Exit()
} else {
_ = p.bar.Finish() _ = p.bar.Finish()
}
fmt.Fprintln(os.Stderr) fmt.Fprintln(os.Stderr)
+189
View File
@@ -0,0 +1,189 @@
package main
import (
"os"
"path/filepath"
"strings"
"testing"
"time"
)
// spinnerIdle comfortably outlasts the 100ms interval at which the
// progressbar library redraws a spinner from its own goroutine.
const spinnerIdle = 500 * time.Millisecond
//nolint:paralleltest // replaces the process-wide os.Stderr
func TestStderrIsTTYFalseForNonTerminals(t *testing.T) {
r, pipe, err := os.Pipe()
if err != nil {
t.Fatal(err)
}
regular, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
devNull, err := os.OpenFile(os.DevNull, os.O_WRONLY, 0)
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
t.Cleanup(func() {
os.Stderr = saved
for _, f := range []*os.File{r, pipe, regular, devNull} {
_ = f.Close()
}
})
cases := map[string]*os.File{
"a pipe": pipe,
"a regular file": regular,
os.DevNull: devNull,
}
for name, f := range cases {
os.Stderr = f
if stderrIsTTY() {
t.Errorf("stderrIsTTY() = true with stderr on %s", name)
}
}
}
// TestNewProgressPrintsBeforeFirstItem checks that each pass shows its
// zero state the moment it starts when stderr is not a terminal, and
// that the next line still waits for plainInterval.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestNewProgressPrintsBeforeFirstItem(t *testing.T) {
stderr := captureStderr(t)
newProgress("walk", -1).increment()
newProgress("hash", 10).increment()
want := "walk: 0 files, elapsed 0s\n" +
"hash: [0/10] 0% 0 files/s elapsed 0s eta ?\n"
if got := stderr(); got != want {
t.Errorf("stderr = %q, want %q", got, want)
}
}
// newWalkSpinner returns the walk pass's terminal display, writing to
// os.Stderr whether or not it is a terminal, and stops the library's
// redraws when the test ends.
func newWalkSpinner(t *testing.T) *progress {
t.Helper()
p := &progress{
label: "walk", total: -1, start: time.Now(),
bar: newBar("walk", -1),
}
t.Cleanup(p.finish)
return p
}
// TestProgressWarningsOnOwnLines drives the terminal display of the walk
// pass through a run of warnings with no items between them, as when the
// walk meets many unreadable paths, for several of the spinner's
// redraws: every warning must land on a line of its own, never inside a
// redraw.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestProgressWarningsOnOwnLines(t *testing.T) {
stderr := captureStderr(t)
p := newWalkSpinner(t)
// No pause between warnings: one written straight to stderr is
// garbled only if a redraw lands while it is being written.
issued := 0
for start := time.Now(); time.Since(start) < spinnerIdle; issued++ {
p.warnf("warning")
}
// The spinner prints the warnings at its next redraw.
time.Sleep(spinnerIdle)
// A terminal shows each line as the text after its last carriage
// return.
shown := 0
for line := range strings.SplitSeq(stderr(), "\n") {
if !strings.Contains(line, "warning") {
continue
}
shown++
if text := line[strings.LastIndex(line, "\r")+1:]; text != "warning" {
t.Errorf("terminal shows %q, want %q", text, "warning")
}
}
if shown != issued {
t.Errorf("%d warning lines, want %d", shown, issued)
}
}
// TestSpinnerShowsCountAfterBurst checks that once a burst of items
// faster than the redraw limit is over, the walk display shows every
// item completed while it waits for the next one.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestSpinnerShowsCountAfterBurst(t *testing.T) {
stderr := captureStderr(t)
p := newWalkSpinner(t)
for range 50 {
p.increment()
}
time.Sleep(spinnerIdle)
if shown := lastFrame(stderr()); !strings.Contains(shown, "(50/-,") {
t.Errorf("terminal shows %q, want a count of 50", shown)
}
}
// lastFrame returns what a terminal shows of the frames a bar drew: the
// last one. The library starts each frame with a carriage return and
// erases the previous one with spaces first.
func lastFrame(out string) string {
var shown string
for frame := range strings.SplitSeq(out, "\r") {
if strings.TrimSpace(frame) != "" {
shown = frame
}
}
return shown
}
// TestBarStoppedShortKeepsCount checks that the terminal display of a
// pass that stops before its total, as an interrupted one does, is left
// as last drawn instead of being filled up.
//
//nolint:paralleltest // captureStderr replaces the process-wide os.Stderr
func TestBarStoppedShortKeepsCount(t *testing.T) {
stderr := captureStderr(t)
p := &progress{
label: "hash", total: 10, start: time.Now(),
bar: newBar("hash", 10),
}
p.increment()
// Past the redraw limit, so the bar draws the next count.
time.Sleep(2 * barThrottle)
p.increment()
p.finish()
if shown := lastFrame(stderr()); !strings.Contains(shown, "(2/10,") {
t.Errorf("terminal shows %q, want a count of 2 of 10", shown)
}
}
+46 -91
View File
@@ -4,8 +4,8 @@ import (
"bufio" "bufio"
"context" "context"
"fmt" "fmt"
"io"
"os" "os"
"slices"
"strings" "strings"
) )
@@ -28,73 +28,59 @@ type scanRec struct {
path string path string
} }
// loadRecords opens the database and reads every file record for the // runReport implements the report subcommand: it prints the file-level
// report and trees subcommands. Any database problem — including a // duplicates report as TSV on stdout. SQLite groups and orders the
// missing database — is fatal. The error is returned rather than // records, and each row is written as it is read, so no group is held
// exiting, so that the deferred close — which checkpoints the SQLite // in memory. It never touches the scanned filesystem; its only I/O is
// WAL — always runs; the database is closed before the caller formats // the database (with SQLite's temporary sort file), stdout, and stderr.
// its output, so it stays closed even if that output fails. // Any database problem, including a missing database, is fatal.
func loadRecords(ctx context.Context) ([]scanRec, error) { func runReport(ctx context.Context, stdout io.Writer) error {
dbPath := databasePath() dbPath := databasePath()
db, err := openReportDatabase(ctx, dbPath) db, err := openReportDatabase(ctx, dbPath)
if err != nil { if err != nil {
return nil, err return err
} }
defer func() { _ = db.Close() }() defer func() { _ = db.Close() }()
recs, err := loadFileRows(ctx, db) out := bufio.NewWriterSize(stdout, ioBufSize)
if err != nil {
return nil, fmt.Errorf("database %s: %w", dbPath, err)
}
return recs, nil
}
// dupeGroup is one set of candidate-duplicate files: identical size,
// head hash, tail hash, and content hash. paths is sorted
// lexicographically; the first entry is the group's "first", the rest
// are dupes.
type dupeGroup struct {
size int64
paths []string
}
// runReport implements the report subcommand: it reads every record
// from the database and prints the file-level duplicates report as TSV
// on stdout. It never touches the scanned filesystem; its only I/O is
// the database, stdout, and stderr.
func runReport(ctx context.Context) error {
recs, err := loadRecords(ctx)
if err != nil {
return err
}
dupes := collectDupeGroups(recs)
out := bufio.NewWriterSize(os.Stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tsize") _, err = fmt.Fprintln(out, "first\tdupe\tsize")
if err != nil { if err != nil {
return fmt.Errorf("write stdout: %w", err) return fmt.Errorf("write stdout: %w", err)
} }
dupeFiles := 0 var (
groups, dupeFiles int
reclaimable int64
writeErr error
)
var reclaimable int64 records, err := loadDupeRows(ctx, db,
func(first, path string, size int64) error {
// A group's first path is its first row; every other
// path is a dupe.
if path == first {
groups++
for _, g := range dupes { return nil
for _, p := range g.paths[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\n",
g.paths[0], p, g.size)
if err != nil {
return fmt.Errorf("write stdout: %w", err)
} }
_, writeErr = fmt.Fprintf(out, "%s\t%s\t%d\n",
escapePath(first), escapePath(path), size)
dupeFiles++ dupeFiles++
reclaimable += g.size reclaimable += size
return writeErr
})
if writeErr != nil {
return fmt.Errorf("write stdout: %w", writeErr)
} }
if err != nil {
return fmt.Errorf("database %s: %w", dbPath, err)
} }
err = out.Flush() err = out.Flush()
@@ -105,55 +91,24 @@ func runReport(ctx context.Context) error {
fmt.Fprintf(os.Stderr, fmt.Fprintf(os.Stderr,
"report: %d records read, %d duplicate groups, %d dupe files, "+ "report: %d records read, %d duplicate groups, %d dupe files, "+
"%s reclaimable\n", "%s reclaimable\n",
len(recs), len(dupes), dupeFiles, humanBytes(reclaimable)) records, groups, dupeFiles, humanBytes(reclaimable))
return nil return nil
} }
// collectDupeGroups groups records by signature and returns every group // escapePath returns a path as it is written in a report column (README
// with two or more paths, each group's paths sorted lexicographically, // "Report output format"): a backslash, tab, newline or carriage return
// groups ordered by size descending then by first path ascending. // becomes \\, \t, \n or \r, and every other byte is kept as it is.
func collectDupeGroups(recs []scanRec) []dupeGroup { // Grouping and sorting use the raw path, never this form.
groups := make(map[fileSig][]string) func escapePath(p string) string {
// Most paths need no escaping; skip building a replacer for them.
for _, r := range recs { if !strings.ContainsAny(p, "\\\t\n\r") {
// A record without a content hash has unknown content and is return p
// never reported as a duplicate (README "Database").
if r.content == "" {
continue
} }
k := fileSig{ return strings.NewReplacer(
size: r.size, head: r.head, tail: r.tail, content: r.content, `\`, `\\`, "\t", `\t`, "\n", `\n`, "\r", `\r`,
} ).Replace(p)
groups[k] = append(groups[k], r.path)
}
var dupes []dupeGroup
for k, paths := range groups {
if len(paths) < minGroupSize {
continue
}
slices.Sort(paths)
dupes = append(dupes, dupeGroup{size: k.size, paths: paths})
}
// Biggest reclaimable space first; ties broken by first path.
slices.SortFunc(dupes, func(a, b dupeGroup) int {
if a.size != b.size {
if a.size > b.size {
return -1
}
return 1
}
return strings.Compare(a.paths[0], b.paths[0])
})
return dupes
} }
// humanBytes formats a byte count in human units (binary prefixes). // humanBytes formats a byte count in human units (binary prefixes).
+260 -15
View File
@@ -1,11 +1,254 @@
package main package main
import ( import (
"bytes"
"database/sql"
"errors"
"fmt"
"io"
"os"
"path/filepath"
"slices" "slices"
"strings"
"testing" "testing"
) )
func TestCollectDupeGroups(t *testing.T) { // awkwardDir is a directory name holding every byte the reports escape.
const awkwardDir = "/d/\tone\ntwo\rthree\\four"
// awkwardPairRecs is a duplicate pair in sibling directories /d/A and
// awkwardDir. A raw tab sorts before "A" but its escaped form `\t`
// sorts after it, so awkwardDir coming first shows that sorting uses
// the raw path.
func awkwardPairRecs() []scanRec {
return []scanRec{
{size: 5, head: "h", tail: "t", content: "c", path: "/d/A/f"},
{size: 5, head: "h", tail: "t", content: "c", path: awkwardDir + "/f"},
}
}
// seedDatabase writes recs into a fresh database and returns its path.
func seedDatabase(t *testing.T, recs []scanRec) string {
t.Helper()
path := testDBPath(t)
db, err := openScanDatabase(t.Context(), path)
if err != nil {
t.Fatal(err)
}
err = applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
err = db.Close()
if err != nil {
t.Fatal(err)
}
return path
}
// dupeGroup is one duplicate group as report reads it: the size, and
// the paths in report order, first path first.
type dupeGroup struct {
size int64
paths []string
}
// dupeGroups returns the duplicate groups report reads from db, in
// report order.
func dupeGroups(t *testing.T, db *sql.DB) []dupeGroup {
t.Helper()
var groups []dupeGroup
_, err := loadDupeRows(t.Context(), db,
func(first, path string, size int64) error {
if path == first {
groups = append(groups, dupeGroup{size: size})
}
g := &groups[len(groups)-1]
g.paths = append(g.paths, path)
return nil
})
if err != nil {
t.Fatal(err)
}
return groups
}
// dupeGroupsOf writes recs into a fresh database and returns the
// duplicate groups report reads from it.
func dupeGroupsOf(t *testing.T, recs []scanRec) []dupeGroup {
t.Helper()
db := openTestDB(t)
err := applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
return dupeGroups(t, db)
}
func TestRunReportEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
code := run([]string{cmdReport}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(report) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tsize\n" +
`/d/\tone\ntwo\rthree\\four/f` + "\t/d/A/f\t5\n"
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
func TestReportStdoutFailsWhileReading(t *testing.T) {
// Each row holds two paths longer than dir, so the report is more
// than twice the stdout buffer and stdout fails while rows are
// still being read, not at the final flush.
dir := "/" + strings.Repeat("d", 4096)
recs := make([]scanRec, ioBufSize/len(dir))
for i := range recs {
recs[i] = scanRec{
size: 1, head: "h", tail: "t", content: "c",
path: fmt.Sprintf("%s/%d", dir, i),
}
}
t.Setenv(databaseEnv, seedDatabase(t, recs))
err := runReport(t.Context(), failingWriter{})
if !errors.Is(err, errWriteFailed) ||
!strings.HasPrefix(err.Error(), "write stdout: ") {
t.Errorf("error = %v, want write stdout: %v", err, errWriteFailed)
}
}
func TestRunReportsIgnoreInsertionOrder(t *testing.T) {
// README §Constraints: identical database contents give identical
// output, whatever order the records were inserted in.
recs := append(smokeTreeRecs(), awkwardPairRecs()...)
recs = append(recs,
scanRec{size: 50, head: "b", tail: "b", content: "b", path: "/y/2"},
scanRec{size: 50, head: "b", tail: "b", content: "b", path: "/y/1"},
scanRec{size: 50, head: "a", tail: "a", content: "a", path: "/x/2"},
scanRec{size: 50, head: "a", tail: "a", content: "a", path: "/x/1"},
scanRec{size: 50, path: "/x/unhashed"},
)
reversed := slices.Clone(recs)
slices.Reverse(reversed)
for _, name := range []string{cmdReport, cmdTrees} {
t.Run(name, func(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, recs))
forward := runStdout(t, name)
t.Setenv(databaseEnv, seedDatabase(t, reversed))
backward := runStdout(t, name)
if strings.Count(forward, "\n") < 3 {
t.Errorf("stdout = %q, want at least two rows", forward)
}
if forward != backward {
t.Errorf("stdout depends on insertion order: %q vs %q",
forward, backward)
}
})
}
}
// runStdout runs the subcommand name and returns its stdout, failing
// the test unless it succeeds.
func runStdout(t *testing.T, name string) string {
t.Helper()
var stdout, stderr bytes.Buffer
code := run([]string{name}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(%s) = %d, want %d; stderr: %s",
name, code, exitOK, stderr.String())
}
return stdout.String()
}
func TestEscapePath(t *testing.T) {
t.Parallel()
cases := map[string]string{
"/srv/plain": "/srv/plain",
"/a\tb": `/a\tb`,
"/a\nb": `/a\nb`,
"/a\rb": `/a\rb`,
`/a\b`: `/a\\b`,
`/a\tb`: `/a\\tb`,
"/not-utf8\xff": "/not-utf8\xff",
}
for in, want := range cases {
if got := escapePath(in); got != want {
t.Errorf("escapePath(%q) = %q, want %q", in, got, want)
}
}
}
// TestWarnfEscapes checks that a warning naming a path that holds a
// newline is still one line.
//
//nolint:paralleltest // replaces the process-wide os.Stderr
func TestWarnfEscapes(t *testing.T) {
f, err := os.Create(filepath.Join(t.TempDir(), "stderr"))
if err != nil {
t.Fatal(err)
}
saved := os.Stderr
os.Stderr = f
t.Cleanup(func() {
os.Stderr = saved
_ = f.Close()
})
(&progress{}).warnf("stat %s: %s", "/d/a\nb", "gone")
_, err = f.Seek(0, io.SeekStart)
if err != nil {
t.Fatal(err)
}
got, err := io.ReadAll(f)
if err != nil {
t.Fatal(err)
}
want := `stat /d/a\nb: gone` + "\n"
if string(got) != want {
t.Errorf("warning = %q, want %q", got, want)
}
}
func TestDupeGroups(t *testing.T) {
t.Parallel() t.Parallel()
recs := []scanRec{ recs := []scanRec{
@@ -20,7 +263,7 @@ func TestCollectDupeGroups(t *testing.T) {
{size: 7, head: "u", tail: "u", content: "u", path: "/lonely"}, {size: 7, head: "u", tail: "u", content: "u", path: "/lonely"},
} }
groups := collectDupeGroups(recs) groups := dupeGroupsOf(t, recs)
if len(groups) != 2 { if len(groups) != 2 {
t.Fatalf("len(groups) = %d, want 2", len(groups)) t.Fatalf("len(groups) = %d, want 2", len(groups))
} }
@@ -38,7 +281,7 @@ func TestCollectDupeGroups(t *testing.T) {
} }
} }
func TestCollectDupeGroupsContentSeparates(t *testing.T) { func TestDupeGroupsContentSeparates(t *testing.T) {
t.Parallel() t.Parallel()
// Same size, head, and tail, but different content hashes: the final // Same size, head, and tail, but different content hashes: the final
@@ -53,7 +296,7 @@ func TestCollectDupeGroupsContentSeparates(t *testing.T) {
{size: 100, head: "h", tail: "t", path: "/e"}, {size: 100, head: "h", tail: "t", path: "/e"},
} }
groups := collectDupeGroups(recs) groups := dupeGroupsOf(t, recs)
if len(groups) != 1 { if len(groups) != 1 {
t.Fatalf("len(groups) = %d, want 1 (only the matching content)", t.Fatalf("len(groups) = %d, want 1 (only the matching content)",
len(groups)) len(groups))
@@ -64,7 +307,7 @@ func TestCollectDupeGroupsContentSeparates(t *testing.T) {
} }
} }
func TestCollectDupeGroupsMtimeExcluded(t *testing.T) { func TestDupeGroupsMtimeExcluded(t *testing.T) {
t.Parallel() t.Parallel()
// mtime is informational only; records differing only in mtime // mtime is informational only; records differing only in mtime
@@ -74,23 +317,25 @@ func TestCollectDupeGroupsMtimeExcluded(t *testing.T) {
{size: 9, mtime: 200, head: "h", tail: "t", content: "c", path: "/m/2"}, {size: 9, mtime: 200, head: "h", tail: "t", content: "c", path: "/m/2"},
} }
groups := collectDupeGroups(recs) groups := dupeGroupsOf(t, recs)
if len(groups) != 1 { if len(groups) != 1 {
t.Fatalf("len(groups) = %d, want 1", len(groups)) t.Fatalf("len(groups) = %d, want 1", len(groups))
} }
} }
func TestCollectDupeGroupsTieBreak(t *testing.T) { func TestDupeGroupsTieBreak(t *testing.T) {
t.Parallel() t.Parallel()
// The hashes sort opposite to the first paths, so ordering the
// groups by hash instead of by first path fails this test.
recs := []scanRec{ recs := []scanRec{
{size: 50, head: "b", tail: "b", content: "b", path: "/beta/2"}, {size: 50, head: "a", tail: "a", content: "a", path: "/beta/2"},
{size: 50, head: "b", tail: "b", content: "b", path: "/beta/1"}, {size: 50, head: "a", tail: "a", content: "a", path: "/beta/1"},
{size: 50, head: "a", tail: "a", content: "a", path: "/alpha/2"}, {size: 50, head: "b", tail: "b", content: "b", path: "/alpha/2"},
{size: 50, head: "a", tail: "a", content: "a", path: "/alpha/1"}, {size: 50, head: "b", tail: "b", content: "b", path: "/alpha/1"},
} }
groups := collectDupeGroups(recs) groups := dupeGroupsOf(t, recs)
if len(groups) != 2 { if len(groups) != 2 {
t.Fatalf("len(groups) = %d, want 2", len(groups)) t.Fatalf("len(groups) = %d, want 2", len(groups))
} }
@@ -102,7 +347,7 @@ func TestCollectDupeGroupsTieBreak(t *testing.T) {
} }
} }
func TestCollectDupeGroupsDeterministic(t *testing.T) { func TestDupeGroupsDeterministic(t *testing.T) {
t.Parallel() t.Parallel()
recs := []scanRec{ recs := []scanRec{
@@ -112,12 +357,12 @@ func TestCollectDupeGroupsDeterministic(t *testing.T) {
{size: 2, head: "b", tail: "b", content: "b", path: "/q/2"}, {size: 2, head: "b", tail: "b", content: "b", path: "/q/2"},
} }
forward := collectDupeGroups(recs) forward := dupeGroupsOf(t, recs)
reversed := slices.Clone(recs) reversed := slices.Clone(recs)
slices.Reverse(reversed) slices.Reverse(reversed)
backward := collectDupeGroups(reversed) backward := dupeGroupsOf(t, reversed)
if !slices.EqualFunc(forward, backward, func(a, b dupeGroup) bool { if !slices.EqualFunc(forward, backward, func(a, b dupeGroup) bool {
return a.size == b.size && slices.Equal(a.paths, b.paths) return a.size == b.size && slices.Equal(a.paths, b.paths)
}) { }) {
+190 -43
View File
@@ -11,6 +11,7 @@ import (
"io" "io"
"io/fs" "io/fs"
"os" "os"
"os/signal"
"path/filepath" "path/filepath"
"slices" "slices"
"strings" "strings"
@@ -57,6 +58,10 @@ const sampleWindow = 1024 * 1024
// and hash worker pools. // and hash worker pools.
const workQueueDepth = 1024 const workQueueDepth = 1024
// errInterrupted reports a scan stopped by SIGINT or SIGTERM. runScan
// has already printed its line, so run prints nothing more.
var errInterrupted = errors.New("scan interrupted")
// fileRec carries one statted file between the scan phases. dev and // fileRec carries one statted file between the scan phases. dev and
// ino identify the underlying inode so hard-linked paths can share // ino identify the underlying inode so hard-linked paths can share
// one read; both are zero when the platform exposes no inode. // one read; both are zero when the platform exposes no inode.
@@ -85,17 +90,19 @@ type fileMeta struct {
// least one other file shares are ever hashed: a size-unique file // least one other file shares are ever hashed: a size-unique file
// cannot be a duplicate. A file of headTailMin or more gets its content // cannot be a duplicate. A file of headTailMin or more gets its content
// hash only when its size, head, and tail match another file's. Flag // hash only when its size, head, and tail match another file's. Flag
// parsing and the at-least-one-operand check are done by cobra. Errors // parsing and the at-least-one-operand check are done by cobra. The
// are returned rather than exiting, so that the deferred close — which // scan holds the lock on the database for its whole run, so a second
// checkpoints the SQLite WAL — always runs. Cancelling ctx unwinds the // scan fails before it walks the filesystem or opens the database.
// worker pools and aborts the scan with the context's error. // Errors are returned rather than exiting, so that the deferred close —
// which takes the database out of WAL mode — always runs, and the lock
// is released after it. When ctx is cancelled, as by the SIGINT or
// SIGTERM that interruptContext catches, the scan keeps what it has
// hashed (see syncScan), prints how many files its walk reached, and
// returns errInterrupted. workers must be at least 1; the scan command
// rejects anything less.
func runScan(ctx context.Context, roots []string, workers int, func runScan(ctx context.Context, roots []string, workers int,
oneFS bool, oneFS bool,
) error { ) error {
if workers < 1 {
workers = 1
}
roots, err := resolveRoots(roots) roots, err := resolveRoots(roots)
if err != nil { if err != nil {
return err return err
@@ -103,14 +110,31 @@ func runScan(ctx context.Context, roots []string, workers int,
dbPath := databasePath() dbPath := databasePath()
db, err := openScanDatabase(ctx, dbPath) lock, err := lockScanDatabase(dbPath)
if err != nil { if err != nil {
return err return err
} }
defer func() { _ = db.Close() }() defer func() { _ = lock.Close() }()
db, err := openScanDatabase(ctx, dbPath)
if err != nil && ctx.Err() != nil {
// Interrupted while opening; SQLite may report that with an
// error of its own rather than the context's.
return interrupted(0)
}
if err != nil {
return err
}
defer closeScanDatabase(ctx, db, dbPath)
st, err := syncScan(ctx, db, roots, workers, oneFS) st, err := syncScan(ctx, db, roots, workers, oneFS)
if errors.Is(err, context.Canceled) {
return interrupted(st.walked)
}
if err != nil { if err != nil {
return fmt.Errorf("update database %s: %w", dbPath, err) return fmt.Errorf("update database %s: %w", dbPath, err)
} }
@@ -124,6 +148,34 @@ func runScan(ctx context.Context, roots []string, workers int,
return nil return nil
} }
// interruptContext returns a copy of ctx that the first SIGINT or
// SIGTERM cancels; the scan command runs the scan under it. stop
// releases the signals.
func interruptContext(ctx context.Context) (context.Context, func()) {
// A SIGINT ignored from the start, as by a script's background job,
// stays ignored.
signals := []os.Signal{syscall.SIGTERM}
if !signal.Ignored(syscall.SIGINT) {
signals = append(signals, syscall.SIGINT)
}
ctx, stop := signal.NotifyContext(ctx, signals...)
// Stopping restores the default handling, so a second signal ends
// the process at once.
context.AfterFunc(ctx, stop)
return ctx, stop
}
// interrupted prints the line for a scan stopped by a signal after its
// walk reached walked files, and returns errInterrupted.
func interrupted(walked int) error {
fmt.Fprintf(os.Stderr, "scan: interrupted after %d files\n", walked)
return errInterrupted
}
// resolveRoots converts each PATH operand to an absolute, lexically // resolveRoots converts each PATH operand to an absolute, lexically
// cleaned path (symlinks are not resolved) and verifies that it // cleaned path (symlinks are not resolved) and verifies that it
// exists. Database records are keyed by absolute path, so scan results // exists. Database records are keyed by absolute path, so scan results
@@ -177,8 +229,9 @@ func pruneRoots(roots []string) []string {
} }
// scanStats summarizes one scan's database synchronization for the // scanStats summarizes one scan's database synchronization for the
// final stderr summary. // final stderr summary, or for the line an interrupted scan prints.
type scanStats struct { type scanStats struct {
walked int // files the walk reached
added int added int
updated int updated int
removed int removed int
@@ -200,7 +253,30 @@ type scanState struct {
st scanStats st scanStats
} }
// syncScan synchronizes the database with the filesystem under roots // syncScan synchronizes the database with the filesystem under roots;
// see runPhases. When ctx is cancelled, as by an interrupt, it commits
// the hashed records still waiting in the batch, starts no other write
// or deletion, and returns the cancellation.
func syncScan(ctx context.Context, db *sql.DB, roots []string,
workers int, oneFS bool,
) (scanStats, error) {
s := &scanState{db: db}
err := s.runPhases(ctx, roots, workers, oneFS)
if err == nil || ctx.Err() == nil {
return s.st, err
}
// The one write made after the cancellation, so it cannot use ctx.
err = applyChanges(context.WithoutCancel(ctx), db, s.batch, nil, nil)
if err != nil {
return s.st, err
}
return s.st, ctx.Err()
}
// runPhases synchronizes the database with the filesystem under roots
// in four sequential phases: walk (enumerate and stat every file, // in four sequential phases: walk (enumerate and stat every file,
// building a complete size census), hash (read only the new or // building a complete size census), hash (read only the new or
// changed — or previously unhashed — files whose size at least one // changed — or previously unhashed — files whose size at least one
@@ -210,47 +286,73 @@ type scanState struct {
// in the content hash of every record of headTailMin or more whose // in the content hash of every record of headTailMin or more whose
// size, head, and tail match another record's). Records outside the // size, head, and tail match another record's). Records outside the
// roots are never touched, except that the content phase fills in // roots are never touched, except that the content phase fills in
// their content hash. // their content hash. Operands the walk cannot start from are dropped
func syncScan(ctx context.Context, db *sql.DB, roots []string, // first, so the records beneath them count as outside the roots unless
// they lie under another root.
func (s *scanState) runPhases(ctx context.Context, roots []string,
workers int, oneFS bool, workers int, oneFS bool,
) (scanStats, error) { ) error {
roots = pruneRoots(roots) // Types are checked before pruning so that an operand under a
// dropped one is still scanned, not dropped as lying under it.
s := &scanState{db: db} roots = pruneRoots(s.walkableRoots(roots))
err := s.loadIndex(ctx, roots) err := s.loadIndex(ctx, roots)
if err != nil { if err != nil {
return s.st, err return err
} }
changed, unhashed := s.walkPhase(startWalk(ctx, roots, oneFS, workers)) changed, unhashed := s.walkPhase(startWalk(ctx, roots, oneFS, workers))
// A cancelled walk stops early, so its size census covers only part // A cancelled walk stops early, so its size census covers only part
// of the roots, and every file it never reached looks vanished to // of the roots, and every file it never reached would look vanished
// the update phase. Defence in depth rather than the only barrier: // to the update phase. Stop before anything is written or deleted.
// that phase would today fail on its first BeginTx with the same
// cancelled context before deleting anything. But it is the barrier
// that survives a later decision to let an interrupted scan commit
// what it has, and it turns a confusing failure deep in the update
// phase into a clean abort at the phase boundary.
err = ctx.Err() err = ctx.Err()
if err != nil { if err != nil {
return s.st, err return err
} }
s.partition(changed, unhashed) s.partition(changed, unhashed)
err = s.hashPhase(ctx, workers) err = s.hashPhase(ctx, workers)
if err != nil { if err != nil {
return s.st, err return err
} }
err = s.updatePhase(ctx) err = s.updatePhase(ctx)
if err != nil { if err != nil {
return s.st, err return err
} }
return s.st, s.contentPhase(ctx, workers) return s.contentPhase(ctx, workers)
}
// walkableRoots returns the operands the walk can start from: regular
// files, and directories not named .zfs. Every other operand is warned
// about, counted as skipped, and dropped. A dropped operand is no
// longer a root, so the records stored beneath it count as outside the
// roots and are not deleted as unverified, unless it lies under another
// root. An operand that fails lstat here is kept, and the walk warns
// about it.
func (s *scanState) walkableRoots(roots []string) []string {
kept := make([]string, 0, len(roots))
for _, root := range roots {
fi, err := os.Lstat(root)
if err == nil {
warn := operandWarning(root, fi)
if warn != "" {
s.st.skipped++
fmt.Fprintln(os.Stderr, escapePath(warn))
continue
}
}
kept = append(kept, root)
}
return kept
} }
// loadIndex indexes the database records under the scan roots for // loadIndex indexes the database records under the scan roots for
@@ -306,6 +408,7 @@ func (s *scanState) walkPhase(
} }
s.sizes = append(s.sizes, ev.rec.size) s.sizes = append(s.sizes, ev.rec.size)
s.st.walked++
prog.increment() prog.increment()
@@ -509,16 +612,22 @@ func (s *scanState) recordRun(ctx context.Context, r hashResult) error {
} }
// commitFullBatch commits the running batch once it holds // commitFullBatch commits the running batch once it holds
// updateBatchSize records. // updateBatchSize records. A batch that fails to commit is kept: the
// commit fails when the scan is interrupted, and syncScan then commits
// the batch itself.
func (s *scanState) commitFullBatch(ctx context.Context) error { func (s *scanState) commitFullBatch(ctx context.Context) error {
if len(s.batch) < updateBatchSize { if len(s.batch) < updateBatchSize {
return nil return nil
} }
err := applyBatch(ctx, s.db, s.batch, nil, nil) err := applyBatch(ctx, s.db, s.batch, nil, nil)
if err != nil {
return err
}
s.batch = s.batch[:0] s.batch = s.batch[:0]
return err return nil
} }
// updatePhase writes the scan's tail under one progress display: the // updatePhase writes the scan's tail under one progress display: the
@@ -789,11 +898,43 @@ func sendEvent(ctx context.Context, events chan<- walkEvent,
} }
} }
// operandWarning returns the one-line warning for an operand the walk
// does not start from, naming the path and what it is, or "" for one it
// does: a regular file, or a directory not named .zfs. Symlinks are
// never followed, including as operands.
func operandWarning(root string, fi fs.FileInfo) string {
var kind string
switch mode := fi.Mode(); {
case mode.IsRegular():
return ""
case mode.IsDir():
if filepath.Base(root) != ".zfs" {
return ""
}
kind = ".zfs directory"
case mode&fs.ModeSymlink != 0:
kind = "symlink"
case mode&fs.ModeSocket != 0:
kind = "socket"
case mode&fs.ModeNamedPipe != 0:
kind = "FIFO"
case mode&fs.ModeDevice != 0:
kind = "device node"
default:
kind = "non-regular file"
}
return fmt.Sprintf("walk %s: skipping %s operand", root, kind)
}
// seedRoot turns one PATH operand into the walk's starting state: a // seedRoot turns one PATH operand into the walk's starting state: a
// regular-file operand is statted and emitted directly, a directory // regular-file operand is statted and emitted directly, and a directory
// operand becomes an initial job, and a symlink or other non-regular // operand becomes an initial job. walkableRoots has already dropped
// operand yields nothing (symlinks are never followed, including as // every other operand. One that has changed into something else since
// operands). // is warned about and skipped here; it is still a root, so the records
// stored beneath it are deleted as unverified.
func seedRoot(ctx context.Context, root string, func seedRoot(ctx context.Context, root string,
events chan<- walkEvent, events chan<- walkEvent,
) []dirJob { ) []dirJob {
@@ -807,16 +948,19 @@ func seedRoot(ctx context.Context, root string,
return nil return nil
} }
switch { warn := operandWarning(root, fi)
case fi.IsDir(): if warn != "" {
if filepath.Base(root) == ".zfs" { sendEvent(ctx, events, walkEvent{warn: warn, fail: true})
return nil return nil
} }
if fi.IsDir() {
dev, ok := deviceOfInfo(fi) dev, ok := deviceOfInfo(fi)
return []dirJob{{path: root, rootDev: dev, rootDevOK: ok}} return []dirJob{{path: root, rootDev: dev, rootDevOK: ok}}
case fi.Mode().IsRegular(): }
dev, ino := inodeOfInfo(fi) dev, ino := inodeOfInfo(fi)
sendEvent(ctx, events, walkEvent{rec: fileRec{ sendEvent(ctx, events, walkEvent{rec: fileRec{
@@ -828,9 +972,6 @@ func seedRoot(ctx context.Context, root string,
}}) }})
return nil return nil
default:
return nil
}
} }
// startWalkWorkers starts the walk worker pool. Each worker processes // startWalkWorkers starts the walk worker pool. Each worker processes
@@ -929,6 +1070,12 @@ func walkOneDir(ctx context.Context, job dirJob, oneFS bool,
var subs []dirJob var subs []dirJob
for _, e := range entries { for _, e := range entries {
// A cancelled scan wants nothing more from this directory: stop
// rather than lstat the rest of a large one.
if ctx.Err() != nil {
return nil
}
p := filepath.Join(job.path, e.Name()) p := filepath.Join(job.path, e.Name())
if e.IsDir() { if e.IsDir() {
+211 -43
View File
@@ -6,13 +6,18 @@ import (
"crypto/sha256" "crypto/sha256"
"database/sql" "database/sql"
"encoding/hex" "encoding/hex"
"errors"
"fmt" "fmt"
"io"
"io/fs"
"os" "os"
"os/signal"
"path/filepath" "path/filepath"
"runtime" "runtime"
"slices" "slices"
"strconv" "strconv"
"strings" "strings"
"syscall"
"testing" "testing"
"time" "time"
) )
@@ -399,7 +404,7 @@ func TestScanContentGate(t *testing.T) {
} }
} }
groups := collectDupeGroups(recs) groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, same) { if len(groups) != 1 || !slices.Equal(groups[0].paths, same) {
t.Fatalf("groups = %+v, want only the identical pair %q", t.Fatalf("groups = %+v, want only the identical pair %q",
groups, same) groups, same)
@@ -434,7 +439,7 @@ func TestScanContentAcrossOperands(t *testing.T) {
want := []string{a, b} want := []string{a, b}
slices.Sort(want) slices.Sort(want)
groups := collectDupeGroups(recs) groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) { if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
t.Fatalf("groups = %+v, want the pair %q", groups, want) t.Fatalf("groups = %+v, want the pair %q", groups, want)
} }
@@ -455,13 +460,13 @@ func TestScanContentWithinOperand(t *testing.T) {
added := sparseFile(t, dir, "d2", headTailMin) added := sparseFile(t, dir, "d2", headTailMin)
st := syncTree(t, db, dir) st := syncTree(t, db, dir)
if st != (scanStats{added: 1, unchanged: 2}) { if st != (scanStats{walked: 3, added: 1, unchanged: 2}) {
t.Fatalf("rescan stats = %+v, want 1 added 2 unchanged", st) t.Fatalf("rescan stats = %+v, want 1 added 2 unchanged", st)
} }
want := []string{stored, added} want := []string{stored, added}
groups := collectDupeGroups(dbRecords(t, db)) groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) { if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
t.Fatalf("groups = %+v, want the pair %q", groups, want) t.Fatalf("groups = %+v, want the pair %q", groups, want)
} }
@@ -501,7 +506,7 @@ func TestScanContentStalePartners(t *testing.T) {
sparseFile(t, dirB, "changed-copy", headTailMin+1) sparseFile(t, dirB, "changed-copy", headTailMin+1)
st := syncTree(t, db, dirB) st := syncTree(t, db, dirB)
if st != (scanStats{added: 2}) { if st != (scanStats{walked: 2, added: 2}) {
t.Errorf("stats = %+v, want 2 added and nothing skipped", st) t.Errorf("stats = %+v, want 2 added and nothing skipped", st)
} }
@@ -519,7 +524,7 @@ func TestScanContentStalePartners(t *testing.T) {
} }
} }
if groups := collectDupeGroups(recs); len(groups) != 0 { if groups := dupeGroups(t, db); len(groups) != 0 {
t.Errorf("groups = %+v, want none", groups) t.Errorf("groups = %+v, want none", groups)
} }
} }
@@ -558,7 +563,7 @@ func TestScanContentHashedStalePartners(t *testing.T) {
b := sparseFile(t, t.TempDir(), "copy", headTailMin) b := sparseFile(t, t.TempDir(), "copy", headTailMin)
st := syncTree(t, db, filepath.Dir(b)) st := syncTree(t, db, filepath.Dir(b))
if st != (scanStats{added: 1}) { if st != (scanStats{walked: 1, added: 1}) {
t.Errorf("stats = %+v, want 1 added and nothing skipped", st) t.Errorf("stats = %+v, want 1 added and nothing skipped", st)
} }
@@ -570,7 +575,7 @@ func TestScanContentHashedStalePartners(t *testing.T) {
// The stored records lie outside the operand and are left as they // The stored records lie outside the operand and are left as they
// are, so they still group with each other, but not with the copy. // are, so they still group with each other, but not with the copy.
groups := collectDupeGroups(recs) groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, stored) { if len(groups) != 1 || !slices.Equal(groups[0].paths, stored) {
t.Errorf("groups = %+v, want only the stored pair %q", groups, stored) t.Errorf("groups = %+v, want only the stored pair %q", groups, stored)
} }
@@ -598,7 +603,7 @@ func TestScanContentReadFailure(t *testing.T) {
b := sparseFile(t, dirB, "b", headTailMin) b := sparseFile(t, dirB, "b", headTailMin)
st := syncTree(t, db, dirB) st := syncTree(t, db, dirB)
if st != (scanStats{added: 1, skipped: 1}) { if st != (scanStats{walked: 1, added: 1, skipped: 1}) {
t.Fatalf("stats = %+v, want 1 added 1 skipped", st) t.Fatalf("stats = %+v, want 1 added 1 skipped", st)
} }
@@ -612,14 +617,14 @@ func TestScanContentReadFailure(t *testing.T) {
} }
st = syncTree(t, db, dirB) st = syncTree(t, db, dirB)
if st != (scanStats{unchanged: 1}) { if st != (scanStats{walked: 1, unchanged: 1}) {
t.Fatalf("rescan stats = %+v, want 1 unchanged", st) t.Fatalf("rescan stats = %+v, want 1 unchanged", st)
} }
want := []string{a, b} want := []string{a, b}
slices.Sort(want) slices.Sort(want)
groups := collectDupeGroups(dbRecords(t, db)) groups := dupeGroups(t, db)
if len(groups) != 1 || !slices.Equal(groups[0].paths, want) { if len(groups) != 1 || !slices.Equal(groups[0].paths, want) {
t.Fatalf("groups = %+v, want the pair %q after the retry", t.Fatalf("groups = %+v, want the pair %q after the retry",
groups, want) groups, want)
@@ -658,7 +663,7 @@ func TestScanContentCheckError(t *testing.T) {
b := sparseFile(t, t.TempDir(), "b", headTailMin) b := sparseFile(t, t.TempDir(), "b", headTailMin)
st := syncTree(t, db, filepath.Dir(b)) st := syncTree(t, db, filepath.Dir(b))
if st != (scanStats{added: 1, skipped: 1}) { if st != (scanStats{walked: 1, added: 1, skipped: 1}) {
t.Fatalf("stats = %+v, want 1 added 1 skipped", st) t.Fatalf("stats = %+v, want 1 added 1 skipped", st)
} }
@@ -686,7 +691,7 @@ func TestScanContentHardlinks(t *testing.T) {
c := sparseFile(t, dir, "copy", headTailMin) c := sparseFile(t, dir, "copy", headTailMin)
st := syncTree(t, db, dir) st := syncTree(t, db, dir)
if st != (scanStats{added: 3}) { if st != (scanStats{walked: 3, added: 3}) {
t.Fatalf("stats = %+v, want 3 added", st) t.Fatalf("stats = %+v, want 3 added", st)
} }
@@ -856,9 +861,11 @@ func TestWalkFileAndSymlinkOperands(t *testing.T) {
t.Fatalf("file operand: recs = %+v, errs = %d", recs, errs) t.Fatalf("file operand: recs = %+v, errs = %d", recs, errs)
} }
// A symlink operand is not followed and yields nothing. // A symlink operand that reaches the walk (it became one after
// walkableRoots checked it) is not followed: it yields a warning and
// no records.
recs, errs = collectWalk(t, []string{link}, false, 2) recs, errs = collectWalk(t, []string{link}, false, 2)
if errs != 0 || len(recs) != 0 { if errs != 1 || len(recs) != 0 {
t.Fatalf("symlink operand: recs = %+v, errs = %d", recs, errs) t.Fatalf("symlink operand: recs = %+v, errs = %d", recs, errs)
} }
} }
@@ -907,6 +914,159 @@ func TestDeviceOfInfo(t *testing.T) {
} }
} }
// dirEntryFor returns the fs.DirEntry for name within dir, obtained via
// the same os.ReadDir the walk uses, so it carries a real Info().
func dirEntryFor(t *testing.T, dir, name string) fs.DirEntry {
t.Helper()
entries, err := os.ReadDir(dir)
if err != nil {
t.Fatal(err)
}
for _, e := range entries {
if e.Name() == name {
return e
}
}
t.Fatalf("entry %q not found in %q", name, dir)
return nil
}
// callSubdirJob runs subdirJob against e under parent, collecting any
// warning events it emits (subdirJob emits at most one).
func callSubdirJob(t *testing.T, p string, e fs.DirEntry,
parent dirJob, oneFS bool,
) (dirJob, bool, []walkEvent) {
t.Helper()
events := make(chan walkEvent, 1)
job, ok := subdirJob(t.Context(), p, e, parent, oneFS, events)
close(events)
var evs []walkEvent
for ev := range events {
evs = append(evs, ev)
}
return job, ok, evs
}
// TestSubdirJobOneFilesystem exercises the -x boundary check in
// subdirJob directly, so no second real filesystem is needed. The
// subdirectory's real device is compared against a fabricated operand
// device.
func TestSubdirJobOneFilesystem(t *testing.T) {
t.Parallel()
dir := t.TempDir()
sub := filepath.Join(dir, "sub")
err := os.Mkdir(sub, 0o750)
if err != nil {
t.Fatal(err)
}
info, err := os.Lstat(sub)
if err != nil {
t.Fatal(err)
}
dev, ok := deviceOfInfo(info)
if !ok {
t.Skip("platform exposes no device id")
}
// A device the subdirectory is not on, standing in for an operand
// rooted on a different filesystem.
otherDev := dev + 1
e := dirEntryFor(t, dir, "sub")
cases := []struct {
name string
oneFS bool
parent dirJob
wantOK bool
}{
// -x on, subdirectory on a different device than its operand:
// descent is refused.
{"reject across boundary", true,
dirJob{rootDev: otherDev, rootDevOK: true}, false},
// -x on but the operand's own device is unknown: the boundary
// check is bypassed and descent proceeds.
{"bypass when root device unknown", true,
dirJob{rootDev: otherDev, rootDevOK: false}, true},
// Default (no -x): boundaries are crossed even onto a different
// device.
{"cross by default", false,
dirJob{rootDev: otherDev, rootDevOK: true}, true},
}
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
job, ok, evs := callSubdirJob(t, sub, e, tc.parent, tc.oneFS)
if ok != tc.wantOK {
t.Fatalf("accepted = %v, want %v", ok, tc.wantOK)
}
if len(evs) != 0 {
t.Fatalf("unexpected events: %+v", evs)
}
// The accepted job must carry the operand's device down, or -x
// stops checking below the first level.
want := dirJob{
path: sub,
rootDev: tc.parent.rootDev,
rootDevOK: tc.parent.rootDevOK,
}
if ok && job != want {
t.Fatalf("job = %+v, want %+v", job, want)
}
})
}
}
// errInfoUnavailable is returned by errDirEntry.Info().
var errInfoUnavailable = errors.New("info unavailable")
// errDirEntry is a directory entry whose Info() always fails, driving
// subdirJob's stat-error branch deterministically.
type errDirEntry struct{ name string }
func (e errDirEntry) Name() string { return e.name }
func (errDirEntry) IsDir() bool { return true }
func (errDirEntry) Type() fs.FileMode {
return fs.ModeDir
}
func (errDirEntry) Info() (fs.FileInfo, error) {
return nil, errInfoUnavailable
}
// TestSubdirJobStatError asserts that when a subdirectory's Info()
// fails under -x, subdirJob warns and refuses descent.
func TestSubdirJobStatError(t *testing.T) {
t.Parallel()
p := "/does/not/matter/sub"
_, ok, evs := callSubdirJob(t, p, errDirEntry{name: "sub"},
dirJob{rootDev: 1, rootDevOK: true}, true)
if ok {
t.Fatal("descent accepted after stat error, want refused")
}
if len(evs) != 1 || !evs[0].fail || !strings.Contains(evs[0].warn, p) {
t.Fatalf("want one warning naming the path, got %+v", evs)
}
}
// buildSmokeTree recreates the README smoke-test filesystem layout // buildSmokeTree recreates the README smoke-test filesystem layout
// with deterministic content and returns the tree root. // with deterministic content and returns the tree root.
func buildSmokeTree(t *testing.T) string { func buildSmokeTree(t *testing.T) string {
@@ -953,11 +1113,16 @@ func syncTree(t *testing.T, db *sql.DB, roots ...string) scanStats {
return st return st
} }
// dbRecords returns every record currently in the database. // dbRecords returns every record currently in the database, in path
// order.
func dbRecords(t *testing.T, db *sql.DB) []scanRec { func dbRecords(t *testing.T, db *sql.DB) []scanRec {
t.Helper() t.Helper()
recs, err := loadFileRows(t.Context(), db) var recs []scanRec
err := loadFileRows(t.Context(), db, func(r scanRec) {
recs = append(recs, r)
})
if err != nil { if err != nil {
t.Fatal(err) t.Fatal(err)
} }
@@ -994,10 +1159,10 @@ func recordPaths(recs []scanRec) []string {
// assertSmokeDupeGroups checks the file-level duplicate groups for the // assertSmokeDupeGroups checks the file-level duplicate groups for the
// smoke tree rooted at dir. // smoke tree rooted at dir.
func assertSmokeDupeGroups(t *testing.T, dir string, parsed []scanRec) { func assertSmokeDupeGroups(t *testing.T, dir string, db *sql.DB) {
t.Helper() t.Helper()
groups := collectDupeGroups(parsed) groups := dupeGroups(t, db)
if len(groups) != 5 { if len(groups) != 5 {
t.Fatalf("len(groups) = %d, want 5", len(groups)) t.Fatalf("len(groups) = %d, want 5", len(groups))
} }
@@ -1022,11 +1187,10 @@ func assertSmokeDupeGroups(t *testing.T, dir string, parsed []scanRec) {
// assertSmokeTreeGroups checks the duplicate-tree groups for the smoke // assertSmokeTreeGroups checks the duplicate-tree groups for the smoke
// tree rooted at dir. // tree rooted at dir.
func assertSmokeTreeGroups(t *testing.T, dir string, parsed []scanRec) { func assertSmokeTreeGroups(t *testing.T, dir string, db *sql.DB) {
t.Helper() t.Helper()
super, dirs := buildHierarchy(parsed) super, dirs := dbTree(t, db)
super.compute()
tg := collectTreeGroups(dirs, super) tg := collectTreeGroups(dirs, super)
if len(tg) != 1 { if len(tg) != 1 {
@@ -1051,7 +1215,7 @@ func TestScanPipeline(t *testing.T) {
db := openTestDB(t) db := openTestDB(t)
st := syncTree(t, db, dir) st := syncTree(t, db, dir)
if st != (scanStats{added: smokeTreeFiles}) { if st != (scanStats{walked: smokeTreeFiles, added: smokeTreeFiles}) {
t.Fatalf("stats = %+v, want %d added only", st, smokeTreeFiles) t.Fatalf("stats = %+v, want %d added only", st, smokeTreeFiles)
} }
@@ -1060,8 +1224,8 @@ func TestScanPipeline(t *testing.T) {
t.Fatalf("len(records) = %d, want %d", len(parsed), smokeTreeFiles) t.Fatalf("len(records) = %d, want %d", len(parsed), smokeTreeFiles)
} }
assertSmokeDupeGroups(t, dir, parsed) assertSmokeDupeGroups(t, dir, db)
assertSmokeTreeGroups(t, dir, parsed) assertSmokeTreeGroups(t, dir, db)
} }
func TestSyncScanUnchangedReuse(t *testing.T) { func TestSyncScanUnchangedReuse(t *testing.T) {
@@ -1074,7 +1238,7 @@ func TestSyncScanUnchangedReuse(t *testing.T) {
writeFile(t, dir, "b.bin", pattern(2, 600)) writeFile(t, dir, "b.bin", pattern(2, 600))
st := syncTree(t, db, dir) st := syncTree(t, db, dir)
if st != (scanStats{added: 2}) { if st != (scanStats{walked: 2, added: 2}) {
t.Fatalf("first scan stats = %+v, want 2 added", st) t.Fatalf("first scan stats = %+v, want 2 added", st)
} }
@@ -1088,7 +1252,7 @@ func TestSyncScanUnchangedReuse(t *testing.T) {
} }
st = syncTree(t, db, dir) st = syncTree(t, db, dir)
if st != (scanStats{unchanged: 2}) { if st != (scanStats{walked: 2, unchanged: 2}) {
t.Fatalf("rescan stats = %+v, want 2 unchanged", st) t.Fatalf("rescan stats = %+v, want 2 unchanged", st)
} }
@@ -1117,7 +1281,7 @@ func TestSyncScanMtimeBump(t *testing.T) {
} }
st := syncTree(t, db, dir) st := syncTree(t, db, dir)
if st != (scanStats{updated: 1}) { if st != (scanStats{walked: 1, updated: 1}) {
t.Fatalf("mtime-bump stats = %+v, want 1 updated", st) t.Fatalf("mtime-bump stats = %+v, want 1 updated", st)
} }
@@ -1145,7 +1309,7 @@ func TestSyncScanAddRemove(t *testing.T) {
} }
st := syncTree(t, db, dir) st := syncTree(t, db, dir)
if st != (scanStats{added: 1, removed: 1, unchanged: 1}) { if st != (scanStats{walked: 2, added: 1, removed: 1, unchanged: 1}) {
t.Fatalf("add/remove stats = %+v, want 1 added 1 removed 1 unchanged", t.Fatalf("add/remove stats = %+v, want 1 added 1 removed 1 unchanged",
st) st)
} }
@@ -1262,7 +1426,7 @@ func TestSyncScanOverlappingRoots(t *testing.T) {
// A file reachable via two overlapping operands is deduplicated // A file reachable via two overlapping operands is deduplicated
// by path in the shared walk and processed once. // by path in the shared walk and processed once.
st := syncTree(t, db, dir, filepath.Join(dir, "sub")) st := syncTree(t, db, dir, filepath.Join(dir, "sub"))
if st != (scanStats{added: 1}) { if st != (scanStats{walked: 1, added: 1}) {
t.Fatalf("stats = %+v, want 1 added", st) t.Fatalf("stats = %+v, want 1 added", st)
} }
@@ -1283,7 +1447,7 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
// Neither size is shared, so neither file is read: both records // Neither size is shared, so neither file is read: both records
// are written without hashes and no duplicates are reported. // are written without hashes and no duplicates are reported.
st := syncTree(t, db, dir) st := syncTree(t, db, dir)
if st != (scanStats{added: 2}) { if st != (scanStats{walked: 2, added: 2}) {
t.Fatalf("stats = %+v, want 2 added", st) t.Fatalf("stats = %+v, want 2 added", st)
} }
@@ -1295,7 +1459,7 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
} }
} }
if groups := collectDupeGroups(recs); len(groups) != 0 { if groups := dupeGroups(t, db); len(groups) != 0 {
t.Fatalf("groups = %+v, want none from unhashed records", groups) t.Fatalf("groups = %+v, want none from unhashed records", groups)
} }
@@ -1305,12 +1469,12 @@ func TestScanSkipsUniqueSizes(t *testing.T) {
c := writeFile(t, dir, "c.bin", pattern(1, 500)) c := writeFile(t, dir, "c.bin", pattern(1, 500))
st = syncTree(t, db, dir) st = syncTree(t, db, dir)
if st != (scanStats{added: 1, updated: 1, unchanged: 1}) { if st != (scanStats{walked: 3, added: 1, updated: 1, unchanged: 1}) {
t.Fatalf("rescan stats = %+v, want 1 added 1 updated 1 unchanged", t.Fatalf("rescan stats = %+v, want 1 added 1 updated 1 unchanged",
st) st)
} }
groups := collectDupeGroups(dbRecords(t, db)) groups := dupeGroups(t, db)
if len(groups) != 1 { if len(groups) != 1 {
t.Fatalf("groups = %+v, want the a/c pair", groups) t.Fatalf("groups = %+v, want the a/c pair", groups)
} }
@@ -1346,8 +1510,7 @@ func TestTreesUnhashedNeverEqual(t *testing.T) {
} }
for name, unknown := range cases { for name, unknown := range cases {
super, dirs := buildHierarchy(append(slices.Clone(shared), unknown...)) super, dirs := treeOf(t, append(slices.Clone(shared), unknown...))
super.compute()
if tg := collectTreeGroups(dirs, super); len(tg) != 0 { if tg := collectTreeGroups(dirs, super); len(tg) != 0 {
t.Errorf("%s: tree groups = %d, want 0 (the files may differ)", t.Errorf("%s: tree groups = %d, want 0 (the files may differ)",
@@ -1370,7 +1533,7 @@ func TestScanHardlinksReadOnce(t *testing.T) {
} }
st := syncTree(t, db, dir) st := syncTree(t, db, dir)
if st != (scanStats{added: 2}) { if st != (scanStats{walked: 2, added: 2}) {
t.Fatalf("stats = %+v, want 2 added", st) t.Fatalf("stats = %+v, want 2 added", st)
} }
@@ -1384,7 +1547,7 @@ func TestScanHardlinksReadOnce(t *testing.T) {
t.Fatalf("hardlink hashes differ: %+v vs %+v", ra, rb) t.Fatalf("hardlink hashes differ: %+v vs %+v", ra, rb)
} }
if groups := collectDupeGroups(recs); len(groups) != 1 { if groups := dupeGroups(t, db); len(groups) != 1 {
t.Fatalf("groups = %+v, want the hardlink pair", groups) t.Fatalf("groups = %+v, want the hardlink pair", groups)
} }
} }
@@ -1492,6 +1655,13 @@ func injectWriteFailure(t *testing.T, path string) {
func baselineGoroutines(t *testing.T) int { func baselineGoroutines(t *testing.T) int {
t.Helper() t.Helper()
// The first scan command in a process starts os/signal's goroutine,
// which never exits. Start it now, so the baseline counts it
// instead of the scan seeming to leave it behind.
ch := make(chan os.Signal, 1)
signal.Notify(ch, syscall.SIGINT)
signal.Stop(ch)
deadline := time.Now().Add(goroutineSettle) deadline := time.Now().Add(goroutineSettle)
last := runtime.NumGoroutine() last := runtime.NumGoroutine()
@@ -1548,7 +1718,7 @@ func TestScanHashWriteFailureUnwindsPool(t *testing.T) {
code := run([]string{ code := run([]string{
cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir, cmdScan, "--workers", strconv.Itoa(hashLeakWorkers), dir,
}, &stderr) }, io.Discard, &stderr)
if code != exitFatal { if code != exitFatal {
t.Fatalf("run(scan) = %d, want %d; stderr: %s", t.Fatalf("run(scan) = %d, want %d; stderr: %s",
code, exitFatal, stderr.String()) code, exitFatal, stderr.String())
@@ -1625,10 +1795,8 @@ func TestReportsNeverTouchFilesystem(t *testing.T) {
t.Fatal(err) t.Fatal(err)
} }
recs := dbRecords(t, db) assertSmokeDupeGroups(t, dir, db)
assertSmokeTreeGroups(t, dir, db)
assertSmokeDupeGroups(t, dir, recs)
assertSmokeTreeGroups(t, dir, recs)
} }
func TestUnderRoot(t *testing.T) { func TestUnderRoot(t *testing.T) {
+13 -11
View File
@@ -3,10 +3,12 @@
# this repo. Idempotent: every install is guarded by a check so already # this repo. Idempotent: every install is guarded by a check so already
# installed tools are skipped. Base tooling comes from nix, apt, brew, # installed tools are skipped. Base tooling comes from nix, apt, brew,
# or apk (detected in that order); assumes nothing is present (not git, # or apk (detected in that order); assumes nothing is present (not git,
# make, or go). The linter is NOT installed: golangci-lint runs via # make, or go). Neither the linter nor the Markdown formatter is
# docker only (script/lint), pinned by image digest, so the only lint # installed: golangci-lint (script/lint) and prettier (script/fmt,
# script/fmt-check) run via docker only, pinned by hash, so their only
# prerequisite is a working docker — which is warned about, not # prerequisite is a working docker — which is warned about, not
# installed, because everything except linting works without it. # installed, because everything except linting and formatting works
# without it.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
@@ -71,15 +73,15 @@ main() {
if missing make; then pkg_install gnumake make make make; fi if missing make; then pkg_install gnumake make make make; fi
if missing go; then pkg_install go golang go go; fi if missing go; then pkg_install go golang go go; fi
# Linting runs via docker only (script/lint), so docker is a lint # Linting and Markdown formatting run via docker only, so docker is
# prerequisite rather than something bootstrap installs. Warn, do # their prerequisite rather than something bootstrap installs. Warn,
# not fail: everything except `make lint` — and, through it, # do not fail: everything except `make lint`, `make fmt` and
# `make check`, `make docker` and the pre-commit hook — works # `make fmt-check` — and, through them, `make check`, `make docker`
# without it. # and the pre-commit hook — works without it.
if missing docker; then if missing docker; then
echo "bootstrap: WARNING: docker not found; make lint, make check" >&2 echo "bootstrap: WARNING: docker not found; make lint, make fmt," >&2
echo "bootstrap: and make docker require it. Install docker to" >&2 echo "bootstrap: make fmt-check, make check and make docker" >&2
echo "bootstrap: run the linter." >&2 echo "bootstrap: require it." >&2
fi fi
go mod download go mod download
+10 -10
View File
@@ -3,17 +3,17 @@
# push. # push.
# #
# The Dockerfile runs the gates individually as build steps, not the # The Dockerfile runs the gates individually as build steps, not the
# make check aggregate: the lint stage runs make fmt-check, # make check aggregate: the lint stage runs the gofmt check,
# script/verify-lint-image-pin, golangci-lint config verify and # script/verify-lint-image-pin, golangci-lint config verify and
# golangci-lint run; the build stage, dropped to an unprivileged user, # golangci-lint run; the markdown stage runs the prettier check; the
# runs make test and make fmt-check. Neither make lint nor make check # build stage, dropped to an unprivileged user, runs make test. None of
# appears, because both reach script/lint, which is itself a docker # make lint, make fmt-check or make check appears, because each runs
# build, and a docker build cannot run inside one. Lint is not skipped # docker, and docker cannot run inside a docker build. Nothing is
# by that — the linter is invoked directly in the lint stage, and the # skipped by that — the linter, gofmt and prettier are invoked directly
# build stage's COPY --from=lint makes that stage a prerequisite, so # in their stages, and the build stage's COPY --from lines make those
# BuildKit must finish it first. Between the two stages everything # stages prerequisites, so BuildKit must finish them first. Between the
# make check would run has run, which is why a successful build here # three stages everything make check would run has run, which is why a
# implies the repo is green. # successful build here implies the repo is green.
# #
# That implication holds only because of CHECK_EPOCH. A COPY layer is # That implication holds only because of CHECK_EPOCH. A COPY layer is
# invalidated only by changed content, and a rebuild of an unchanged # invalidated only by changed content, and a rebuild of an unchanged
+5 -5
View File
@@ -4,11 +4,11 @@
# #
# CHECK_EPOCH is passed for the same reason script/cibuild passes it: # CHECK_EPOCH is passed for the same reason script/cibuild passes it:
# without it Docker serves the Dockerfile's gate layers from cache on an # without it Docker serves the Dockerfile's gate layers from cache on an
# unchanged tree and this exits 0 having run neither the lint stage's # unchanged tree and this exits 0 having run none of the lint stage's
# gates nor the builder stage's test and fmt-check gates. This is the # gates, the markdown stage's prettier gate or the builder stage's test
# set of gates a developer or reviewer runs by hand, so a cached pass # gate. This is the set of gates a developer or reviewer runs by hand,
# here is the most misleading result the repo can produce. Dependency # so a cached pass here is the most misleading result the repo can
# layers sit above the ARG and stay cached. # produce. Dependency layers sit above the ARG and stay cached.
set -eu set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
+12 -2
View File
@@ -1,12 +1,22 @@
#!/bin/sh #!/bin/sh
# script/fmt: format all files (writes). # script/fmt: format all files (writes): the Go sources with gofmt, the
# Markdown with prettier. prettier is never installed on the host: it
# runs from the Dockerfile's prettier stage with the repository mounted,
# as the calling user so the files it rewrites keep their owner. The tag
# makes each build replace the previous image instead of leaving another
# one behind.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
gofmt -s -w . gofmt -s -w .
image="$("$SCRIPT_DIR/projectname")-prettier"
docker build -q --target prettier -t "$image" . >/dev/null
docker run --rm --user "$(id -u):$(id -g)" -v "$ROOT:/src" "$image" \
prettier --write '**/*.md' --tab-width 4 --prose-wrap always
} }
main "$@" main "$@"
+19 -3
View File
@@ -1,18 +1,34 @@
#!/bin/sh #!/bin/sh
# script/fmt-check: check formatting (read-only). Same scope as # script/fmt-check: check formatting (read-only). Same scope as
# script/fmt, but fails instead of writing. # script/fmt, but fails instead of writing. gofmt and prettier both run
# every time and each reports its own failure, so the output says which
# one failed.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
status=0
files="$(gofmt -s -l .)" files="$(gofmt -s -l .)"
if [ -n "$files" ]; then if [ -n "$files" ]; then
echo "gofmt: files not formatted:" >&2 echo "gofmt: files not formatted:" >&2
echo "$files" >&2 echo "$files" >&2
exit 1 status=1
fi fi
# Same image as script/fmt; see there.
image="$("$SCRIPT_DIR/projectname")-prettier"
docker build -q --target prettier -t "$image" . >/dev/null
if ! docker run --rm -v "$ROOT:/src:ro" "$image" \
prettier --check '**/*.md' --tab-width 4 --prose-wrap always; then
echo "prettier: Markdown not formatted; run make fmt" >&2
status=1
fi
exit "$status"
} }
main "$@" main "$@"
+139 -81
View File
@@ -5,49 +5,59 @@ import (
"context" "context"
"crypto/sha256" "crypto/sha256"
"fmt" "fmt"
"io"
"os" "os"
"slices" "slices"
"strconv" "strconv"
"strings" "strings"
) )
// fileSig is a file's duplicate signature; mtime is excluded. // treeNode is one directory reconstructed from the record paths.
type fileSig struct {
size int64
head string
tail string
content string
}
// treeNode is one directory reconstructed from the scan stream.
type treeNode struct { type treeNode struct {
path string path string
parent *treeNode parent *treeNode
dirs map[string]*treeNode // entries holds the serialized child entries until the digest is
files map[string]fileSig // computed from them, and is then dropped.
entries []string
digest [sha256.Size]byte digest [sha256.Size]byte
fileCount int64 fileCount int64
totalSize int64 totalSize int64
} }
// runTrees implements the trees subcommand: it reads every record from // runTrees implements the trees subcommand: it reads every record from
// the database, reconstructs the directory hierarchy from the record // the database in path order, reconstructs the directory hierarchy from
// paths, computes a Merkle-style digest per directory, and prints // the record paths, computes a Merkle-style digest per directory, and
// maximal duplicate-tree groups as TSV on stdout. It never touches the // prints maximal duplicate-tree groups as TSV on stdout. It never
// scanned filesystem; its only I/O is the database, stdout, and // touches the scanned filesystem; its only I/O is the database, stdout,
// stderr. // and stderr. Any database problem, including a missing database, is
func runTrees(ctx context.Context) error { // fatal.
recs, err := loadRecords(ctx) func runTrees(ctx context.Context, stdout io.Writer) error {
dbPath := databasePath()
db, err := openReportDatabase(ctx, dbPath)
if err != nil { if err != nil {
return err return err
} }
super, allDirs := buildHierarchy(recs) defer func() { _ = db.Close() }()
super.compute()
records := 0
tree := newTreeBuilder()
err = loadFileRows(ctx, db, func(r scanRec) {
records++
tree.add(r)
})
if err != nil {
return fmt.Errorf("database %s: %w", dbPath, err)
}
super, allDirs := tree.finish()
dupes := collectTreeGroups(allDirs, super) dupes := collectTreeGroups(allDirs, super)
out := bufio.NewWriterSize(os.Stdout, ioBufSize) out := bufio.NewWriterSize(stdout, ioBufSize)
_, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize") _, err = fmt.Fprintln(out, "first\tdupe\tfiles\tsize")
if err != nil { if err != nil {
@@ -62,7 +72,8 @@ func runTrees(ctx context.Context) error {
first := g[0] first := g[0]
for _, n := range g[1:] { for _, n := range g[1:] {
_, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n", _, err = fmt.Fprintf(out, "%s\t%s\t%d\t%d\n",
first.path, n.path, first.fileCount, first.totalSize) escapePath(first.path), escapePath(n.path),
first.fileCount, first.totalSize)
if err != nil { if err != nil {
return fmt.Errorf("write stdout: %w", err) return fmt.Errorf("write stdout: %w", err)
} }
@@ -80,65 +91,128 @@ func runTrees(ctx context.Context) error {
fmt.Fprintf(os.Stderr, fmt.Fprintf(os.Stderr,
"trees: %d records read, %d duplicate tree groups, %d dupe trees, "+ "trees: %d records read, %d duplicate tree groups, %d dupe trees, "+
"%s reclaimable\n", "%s reclaimable\n",
len(recs), len(dupes), dupeTrees, humanBytes(reclaimable)) records, len(dupes), dupeTrees, humanBytes(reclaimable))
return nil return nil
} }
// buildHierarchy reconstructs the directory hierarchy from the record // treeBuilder reconstructs the directory hierarchy from records added
// paths under a synthetic super-root. Paths are split on "/"; for // in path order, under a synthetic super-root. Paths are split on "/";
// absolute paths the first component is empty, which simply becomes a // for absolute paths the first component is empty, which becomes the
// top-level node representing "/". It returns the super-root and every // top-level directory with path "/". In path order all the paths under
// directory node created. // one directory come together, so a directory is complete once a path
func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) { // outside it is added: its digest is computed then and its entries are
// dropped. Only the directories holding the latest path keep entries.
type treeBuilder struct {
super *treeNode
// open lists the directories holding the latest path, outermost
// first, starting with the super-root; names[i] is open[i]'s name.
open []*treeNode
names []string
// dirs lists every completed directory.
dirs []*treeNode
}
func newTreeBuilder() *treeBuilder {
super := &treeNode{} super := &treeNode{}
var allDirs []*treeNode return &treeBuilder{
super: super,
open: []*treeNode{super},
names: []string{""},
}
}
for _, r := range recs { // add adds one record. Each record must come after the previous one in
// path order (byte order); otherwise a completed directory would be
// started again as a second directory with the same path.
func (b *treeBuilder) add(r scanRec) {
comps := strings.Split(r.path, "/") comps := strings.Split(r.path, "/")
dirNames, name := comps[:len(comps)-1], comps[len(comps)-1]
node := super // Keep the open directories that hold this path; complete the rest.
for _, c := range comps[:len(comps)-1] { depth := 1
child := node.dirs[c] for depth < len(b.open) && depth <= len(dirNames) &&
if child == nil { b.names[depth] == dirNames[depth-1] {
childPath := c depth++
if node != super {
childPath = node.path + "/" + c
} }
child = &treeNode{path: childPath, parent: node} b.closeTo(depth)
if node.dirs == nil {
node.dirs = make(map[string]*treeNode) for _, c := range dirNames[depth-1:] {
b.openDir(c)
} }
node.dirs[c] = child dir := b.open[len(b.open)-1]
allDirs = append(allDirs, child) dir.entries = append(dir.entries, fileEntry(name, r))
dir.fileCount++
dir.totalSize += r.size
} }
node = child // openDir starts the directory called name inside the innermost open
// one.
func (b *treeBuilder) openDir(name string) {
parent := b.open[len(b.open)-1]
path := parent.path + "/" + name
// The root directory's path is "/", not empty, and its children's
// paths start with one slash, not two.
switch {
case parent == b.super && name == "":
path = "/"
case parent == b.super:
path = name
case parent.path == "/":
path = "/" + name
} }
if node.files == nil { b.open = append(b.open, &treeNode{path: path, parent: parent})
node.files = make(map[string]fileSig) b.names = append(b.names, name)
} }
sig := fileSig{ // closeTo completes the open directories after the first n, innermost
size: r.size, head: r.head, tail: r.tail, content: r.content, // first: each one's digest is computed and entered in its parent along
// with its totals.
func (b *treeBuilder) closeTo(n int) {
for len(b.open) > n {
last := len(b.open) - 1
dir, name := b.open[last], b.names[last]
b.open, b.names = b.open[:last], b.names[:last]
dir.computeDigest()
dir.parent.entries = append(dir.parent.entries,
"d\x00"+name+"\x00"+string(dir.digest[:]))
dir.parent.fileCount += dir.fileCount
dir.parent.totalSize += dir.totalSize
b.dirs = append(b.dirs, dir)
} }
}
// finish completes every open directory and returns the super-root and
// every directory.
func (b *treeBuilder) finish() (*treeNode, []*treeNode) {
b.closeTo(1)
return b.super, b.dirs
}
// fileEntry serializes a file child for its directory's digest: its
// name and its signature (size, head, tail, content); mtime is
// excluded.
func fileEntry(name string, r scanRec) string {
content := r.content
// A record without a content hash has unknown content (README // A record without a content hash has unknown content (README
// "Database"): give it a signature no other file can share, so // "Database"): give it a signature no other file can share, so
// trees containing it never compare equal. Real hashes are // trees containing it never compare equal. Real hashes are hex, so
// hex, so the NUL-prefixed form cannot collide. // the NUL-prefixed form cannot collide.
if sig.content == "" { if content == "" {
sig.content = "unhashed\x00" + r.path content = "unhashed\x00" + r.path
} }
node.files[comps[len(comps)-1]] = sig return "f\x00" + name + "\x00" + strconv.FormatInt(r.size, 10) +
} "\x00" + r.head + "\x00" + r.tail + "\x00" + content
return super, allDirs
} }
// collectTreeGroups groups directories by digest and returns every // collectTreeGroups groups directories by digest and returns every
@@ -180,38 +254,22 @@ func collectTreeGroups(allDirs []*treeNode, super *treeNode) [][]*treeNode {
return dupes return dupes
} }
// compute fills in digest, fileCount, and totalSize for n and all of // computeDigest sets n's digest and drops its entries. A directory's
// its descendants. A directory's digest is the SHA-256 of its child // digest is the SHA-256 of its child entries — files serialized with
// entries — files serialized with name and signature, subdirectories // name and signature, subdirectories with name and recursive digest —
// with name and recursive digest — sorted byte-lexicographically. // sorted byte-lexicographically. Filenames cannot contain NUL or "/",
// Filenames cannot contain NUL or "/", so NUL delimiters are // so NUL delimiters are unambiguous.
// unambiguous. func (n *treeNode) computeDigest() {
func (n *treeNode) compute() { slices.Sort(n.entries)
entries := make([]string, 0, len(n.dirs)+len(n.files))
for name, sig := range n.files {
entries = append(entries,
"f\x00"+name+"\x00"+strconv.FormatInt(sig.size, 10)+
"\x00"+sig.head+"\x00"+sig.tail+"\x00"+sig.content)
n.fileCount++
n.totalSize += sig.size
}
for name, child := range n.dirs {
child.compute()
entries = append(entries, "d\x00"+name+"\x00"+string(child.digest[:]))
n.fileCount += child.fileCount
n.totalSize += child.totalSize
}
slices.Sort(entries)
h := sha256.New() h := sha256.New()
for _, e := range entries { for _, e := range n.entries {
h.Write([]byte(e)) h.Write([]byte(e))
h.Write([]byte{0}) h.Write([]byte{0})
} }
copy(n.digest[:], h.Sum(nil)) copy(n.digest[:], h.Sum(nil))
n.entries = nil
} }
// suppressed reports whether a duplicate-tree group is non-maximal: its // suppressed reports whether a duplicate-tree group is non-maximal: its
+129 -17
View File
@@ -1,6 +1,8 @@
package main package main
import ( import (
"bytes"
"database/sql"
"slices" "slices"
"testing" "testing"
) )
@@ -29,6 +31,36 @@ func smokeTreeRecs() []scanRec {
} }
} }
// dbTree builds the directory hierarchy from the records in db the way
// trees does, and returns the super-root and every directory.
func dbTree(t *testing.T, db *sql.DB) (*treeNode, []*treeNode) {
t.Helper()
tree := newTreeBuilder()
err := loadFileRows(t.Context(), db, tree.add)
if err != nil {
t.Fatal(err)
}
return tree.finish()
}
// treeOf writes recs into a fresh database and builds the directory
// hierarchy from it the way trees does.
func treeOf(t *testing.T, recs []scanRec) (*treeNode, []*treeNode) {
t.Helper()
db := openTestDB(t)
err := applyChanges(t.Context(), db, recs, nil, nil)
if err != nil {
t.Fatal(err)
}
return dbTree(t, db)
}
// nodeByPath finds the directory node with the given path. // nodeByPath finds the directory node with the given path.
func nodeByPath(t *testing.T, dirs []*treeNode, path string) *treeNode { func nodeByPath(t *testing.T, dirs []*treeNode, path string) *treeNode {
t.Helper() t.Helper()
@@ -59,11 +91,10 @@ func groupPaths(groups [][]*treeNode) [][]string {
return out return out
} }
func TestBuildHierarchyCounts(t *testing.T) { func TestTreeCounts(t *testing.T) {
t.Parallel() t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs()) _, dirs := treeOf(t, smokeTreeRecs())
super.compute()
d := nodeByPath(t, dirs, "/d") d := nodeByPath(t, dirs, "/d")
if d.fileCount != 6 || d.totalSize != 9300 { if d.fileCount != 6 || d.totalSize != 9300 {
@@ -84,11 +115,98 @@ func TestBuildHierarchyCounts(t *testing.T) {
} }
} }
func TestTreeRootPath(t *testing.T) {
t.Parallel()
// The root directory's path is "/", never empty, and its
// children's paths start with a single slash.
_, dirs := treeOf(t, []scanRec{{path: "/f"}, {path: "/srv/g"}})
got := make([]string, 0, len(dirs))
for _, d := range dirs {
got = append(got, d.path)
}
slices.Sort(got)
want := []string{"/", "/srv"}
if !slices.Equal(got, want) {
t.Fatalf("directory paths = %q, want %q", got, want)
}
}
func TestTreeNamesSortingBeforeSlash(t *testing.T) {
t.Parallel()
// In path order "/a/b-x/f" and "/a/b.txt" come between the file
// "/a/b" and "/a/b/f", because "-" and "." sort before "/". Each
// directory must still be built once, whole, so /a matches /c.
recs := make([]scanRec, 0, 8)
for _, top := range []string{"/a", "/c"} {
for _, p := range []string{"/b", "/b-x/f", "/b.txt", "/b/f"} {
content := "c"
if p == "/b-x/f" {
content = "other"
}
recs = append(recs, scanRec{
size: 1, head: "h", tail: "t", content: content, path: top + p,
})
}
}
super, dirs := treeOf(t, recs)
got := make([]string, 0, len(dirs))
for _, d := range dirs {
got = append(got, d.path)
}
slices.Sort(got)
want := []string{"/", "/a", "/a/b", "/a/b-x", "/c", "/c/b", "/c/b-x"}
if !slices.Equal(got, want) {
t.Fatalf("directory paths = %q, want %q", got, want)
}
groups := collectTreeGroups(dirs, super)
gotGroups := groupPaths(groups)
wantGroups := [][]string{{"/a", "/c"}}
if !slices.EqualFunc(gotGroups, wantGroups, slices.Equal) {
t.Fatalf("groups = %v, want %v", gotGroups, wantGroups)
}
if groups[0][0].fileCount != 4 || groups[0][0].totalSize != 4 {
t.Errorf("group totals: %d files %d bytes, want 4 4",
groups[0][0].fileCount, groups[0][0].totalSize)
}
}
func TestRunTreesEscapesPaths(t *testing.T) {
t.Setenv(databaseEnv, seedDatabase(t, awkwardPairRecs()))
var stdout, stderr bytes.Buffer
code := run([]string{cmdTrees}, &stdout, &stderr)
if code != exitOK {
t.Fatalf("run(trees) = %d, want %d; stderr: %s",
code, exitOK, stderr.String())
}
want := "first\tdupe\tfiles\tsize\n" +
`/d/\tone\ntwo\rthree\\four` + "\t/d/A\t1\t5\n"
if got := stdout.String(); got != want {
t.Errorf("stdout = %q, want %q", got, want)
}
}
func TestTreeDigests(t *testing.T) { func TestTreeDigests(t *testing.T) {
t.Parallel() t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs()) _, dirs := treeOf(t, smokeTreeRecs())
super.compute()
t1 := nodeByPath(t, dirs, "/d/t1") t1 := nodeByPath(t, dirs, "/d/t1")
t2 := nodeByPath(t, dirs, "/d/t2") t2 := nodeByPath(t, dirs, "/d/t2")
@@ -121,8 +239,7 @@ func TestTreeDigestContentSensitivity(t *testing.T) {
{size: 10, head: "DIFF", tail: sharedTail, content: "c", path: "/r/b/f"}, {size: 10, head: "DIFF", tail: sharedTail, content: "c", path: "/r/b/f"},
} }
super, dirs := buildHierarchy(recs) _, dirs := treeOf(t, recs)
super.compute()
a := nodeByPath(t, dirs, "/r/a") a := nodeByPath(t, dirs, "/r/a")
b := nodeByPath(t, dirs, "/r/b") b := nodeByPath(t, dirs, "/r/b")
@@ -135,8 +252,7 @@ func TestTreeDigestContentSensitivity(t *testing.T) {
func TestCollectTreeGroupsMaximal(t *testing.T) { func TestCollectTreeGroupsMaximal(t *testing.T) {
t.Parallel() t.Parallel()
super, dirs := buildHierarchy(smokeTreeRecs()) super, dirs := treeOf(t, smokeTreeRecs())
super.compute()
groups := collectTreeGroups(dirs, super) groups := collectTreeGroups(dirs, super)
@@ -160,16 +276,14 @@ func TestCollectTreeGroupsDeterministic(t *testing.T) {
recs := smokeTreeRecs() recs := smokeTreeRecs()
super, dirs := buildHierarchy(recs) super, dirs := treeOf(t, recs)
super.compute()
forward := groupPaths(collectTreeGroups(dirs, super)) forward := groupPaths(collectTreeGroups(dirs, super))
reversed := slices.Clone(recs) reversed := slices.Clone(recs)
slices.Reverse(reversed) slices.Reverse(reversed)
superR, dirsR := buildHierarchy(reversed) superR, dirsR := treeOf(t, reversed)
superR.compute()
backward := groupPaths(collectTreeGroups(dirsR, superR)) backward := groupPaths(collectTreeGroups(dirsR, superR))
if !slices.EqualFunc(forward, backward, slices.Equal) { if !slices.EqualFunc(forward, backward, slices.Equal) {
@@ -188,8 +302,7 @@ func TestCollectTreeGroupsSiblings(t *testing.T) {
{size: 10, head: "h", tail: "t", content: "c", path: "/p/x2/f"}, {size: 10, head: "h", tail: "t", content: "c", path: "/p/x2/f"},
} }
super, dirs := buildHierarchy(recs) super, dirs := treeOf(t, recs)
super.compute()
got := groupPaths(collectTreeGroups(dirs, super)) got := groupPaths(collectTreeGroups(dirs, super))
@@ -211,8 +324,7 @@ func TestCollectTreeGroupsDifferingParents(t *testing.T) {
{size: 10, head: "h", tail: "t", content: "c", path: "/q/b/x/f"}, {size: 10, head: "h", tail: "t", content: "c", path: "/q/b/x/f"},
} }
super, dirs := buildHierarchy(recs) super, dirs := treeOf(t, recs)
super.compute()
got := groupPaths(collectTreeGroups(dirs, super)) got := groupPaths(collectTreeGroups(dirs, super))
+8
View File
@@ -0,0 +1,8 @@
# THIS IS AN AUTOGENERATED FILE. DO NOT EDIT THIS FILE DIRECTLY.
# yarn lockfile v1
prettier@3.8.1:
version "3.8.1"
resolved "https://registry.yarnpkg.com/prettier/-/prettier-3.8.1.tgz#edf48977cf991558f4fcbd8a3ba6015ba2a3a173"
integrity sha512-UOnG6LftzbdaHZcKoPFtOcCKztrQ57WkHDeRD9t/PTQtmT0NHSeWWepj6pS0z/N7+08BHFDQVUrfmfMRcZwbMg==