Document install, a daily cron scan and reading the reports (closes #54)
check / check (push) Successful in 1m16s

Getting Started gains three parts: installing with go install, from a
clone or as the Docker image; a crontab line for a daily root scan,
with where its stderr and failures go; and how to read the two reports,
why a row is a candidate rather than proof, and how to check a pair
with cmp before removing anything. The usage block names --help, and
the text below it --workers and -x.

The install line uses @main, not @latest: @latest resolves to the
v0.0.1 tag, which predates the database.

Model: opus-5-5
This commit is contained in:
2026-10-04 09:03:10 +00:00
parent 2eeba3df6f
commit 3f0d2e6a60
2 changed files with 128 additions and 4 deletions
+124 -4
View File
@@ -51,6 +51,122 @@ daily `sfdupes scan` cron job, with the reporting commands run
interactively whenever needed; their results are as fresh as the last
completed scan.
### Install
With Go installed, this builds and installs the current `main` branch:
```sh
go install sneak.berlin/go/sfdupes@main
```
The binary goes to `$(go env GOPATH)/bin`, or to `$GOBIN` when that is
set. A binary installed this way reports its version as `dev`; one
built from a clone or into the Docker image carries the git tag or
commit it was built from.
From a clone, `make build` writes the binary to `./sfdupes`:
```sh
git clone https://git.eeqj.de/sneak/sfdupes.git
cd sfdupes
make build
```
Copy the binary to `/usr/local/bin` for the cron job below.
`make docker` builds the Docker image, tagged `sfdupes`, after running
the tests and the linter (see "Build"). The image runs `sfdupes` as
root with the database at its default path, so a bind mount of
`/var/lib/sfdupes` keeps the database between runs. Mount the scanned
tree at the same path inside the container as on the host; read-only
is enough. The database records paths as the container sees them, so
the reports then name the host's paths.
```sh
make docker
docker run --rm -v /srv:/srv:ro -v /var/lib/sfdupes:/var/lib/sfdupes \
sfdupes scan /srv
docker run --rm -v /var/lib/sfdupes:/var/lib/sfdupes sfdupes report > dupes.tsv
```
### Daily scan from cron
Run `scan` as root, so that it can read every file: a path it cannot
read is skipped with a warning and loses its database record (see
"Rules for the walk"). As a file `/etc/cron.d/sfdupes`:
```
30 3 * * * root /usr/local/bin/sfdupes scan /srv 2>>/var/log/sfdupes.log || tail -n 3 /var/log/sfdupes.log
```
- The database is `/var/lib/sfdupes/db.sqlite`, created with its
directory by the first scan. To keep it elsewhere, set
`SFDUPES_DATABASE=/path/to/db.sqlite` before the command on the same
line.
- `scan` writes nothing to stdout. Its stderr, appended here to
`/var/log/sfdupes.log`, holds a plain progress line as each phase
starts and then at most every 5 seconds, a warning for each path it
skips, and the summary line (see "Progress" and "`scan` mode"). The
log grows with every scan; rotate it like any other.
- Skipped paths do not fail a scan: it still exits 0, and cron sends
nothing. A scan that fails, or is stopped by `SIGINT` or `SIGTERM`,
exits 1 with the reason among the last lines of the log; `tail`
prints them, and cron mails them to root if the host can send mail.
- A scan still running when the next one starts carries on. The new
one fails at once, and the lines cron mails include
`sfdupes: another scan is running (lock held on /var/lib/sfdupes/db.sqlite.lock)`.
- `report` and `trees` need only read access to the database (see
"Database"). Under the usual umask of `022` the first scan creates
it readable by every user, so an unprivileged user can run them
against root's database.
### Reading the reports
Each row of `report` names two copies of one file, and each row of
`trees` two copies of one directory tree (see "Report output format"
and "Trees output format"). In a group of copies, the path that sorts
first byte by byte is `first` and every other path is a `dupe` of it.
`first` says nothing about which copy is the original or the oldest;
which copy to keep is your choice.
A row is a candidate, not proof:
- The reports read only the database, so they show the files as of
the last scan; a file may have changed or gone since.
- A file of 50 MiB or more is compared only on samples of its content
(see "Duplicate detection").
- Paths that are hard links to one file are listed as duplicates, but
they share their data, so removing one frees nothing.
Compare a pair byte for byte before removing either copy. For the row
`/srv/a/big.iso`, `/srv/b/big-copy.iso`, `4294967296`:
```sh
cmp /srv/a/big.iso /srv/b/big-copy.iso && echo identical
[ /srv/a/big.iso -ef /srv/b/big-copy.iso ] && echo "hard links"
```
`cmp` prints nothing and exits 0 only when every byte matches, and
otherwise reports where the files differ. The second line prints
`hard links` when the two paths are the same file, so removing either
frees nothing.
A path holding a backslash, tab, newline or carriage return is escaped
in the reports (see "Report output format"). Undo the escapes before
using it. `printf '%b'` does exactly that, because every backslash in
an escaped path starts one of the four escapes. Command substitution
drops trailing newlines, so print an `x` after the path and remove it
afterwards, or a path that ends in a newline names a different file:
```sh
p="$(printf '%bx' '/srv/a/tab\tname.txt')"; p="${p%x}"
cmp "$p" /srv/b/tab-copy.txt
```
Check a `trees` row with `diff -r`, which compares the two trees file
by file and also names anything present in only one of them, such as
an empty directory or a symlink, which `trees` does not see.
## Rationale
Duplicate finders that hash entire files do not scale to the target
@@ -150,12 +266,16 @@ sfdupes scan [--workers N] [-x] PATH...
sfdupes report > dupes.tsv
sfdupes trees > dupetrees.tsv
sfdupes --version
sfdupes [command] --help
```
`sfdupes --version` (or `-v`) prints one line, `sfdupes VERSION`, to
stdout and exits 0, writing nothing to stderr. `-h` or `--help`, alone
or after a subcommand, prints the help text to stderr and exits 0,
writing nothing to stdout.
`--workers N` sets the size of each `scan` worker pool (default: the
number of CPUs), and `-x` (`--one-file-system`) keeps the walk of each
operand on that operand's filesystem; both are described under
"`scan` mode". `sfdupes --version` (or `-v`) prints one line,
`sfdupes VERSION`, to stdout and exits 0, writing nothing to stderr.
`-h` or `--help`, alone or after a subcommand, prints the help text to
stderr and exits 0, writing nothing to stdout.
### Database