Document install, a daily cron scan and reading the reports (closes #54) #80
@@ -51,6 +51,122 @@ daily `sfdupes scan` cron job, with the reporting commands run
|
||||
interactively whenever needed; their results are as fresh as the last
|
||||
completed scan.
|
||||
|
||||
### Install
|
||||
|
||||
With Go installed, this builds and installs the current `main` branch:
|
||||
|
||||
```sh
|
||||
go install sneak.berlin/go/sfdupes@main
|
||||
```
|
||||
|
||||
The binary goes to `$(go env GOPATH)/bin`, or to `$GOBIN` when that is
|
||||
set. A binary installed this way reports its version as `dev`; one
|
||||
built from a clone or into the Docker image carries the git tag or
|
||||
commit it was built from.
|
||||
|
||||
From a clone, `make build` writes the binary to `./sfdupes`:
|
||||
|
||||
```sh
|
||||
git clone https://git.eeqj.de/sneak/sfdupes.git
|
||||
cd sfdupes
|
||||
make build
|
||||
```
|
||||
|
||||
Copy the binary to `/usr/local/bin` for the cron job below.
|
||||
|
||||
`make docker` builds the Docker image, tagged `sfdupes`, after running
|
||||
the tests and the linter (see "Build"). The image runs `sfdupes` as
|
||||
root with the database at its default path, so a bind mount of
|
||||
`/var/lib/sfdupes` keeps the database between runs. Mount the scanned
|
||||
tree at the same path inside the container as on the host; read-only
|
||||
is enough. The database records paths as the container sees them, so
|
||||
the reports then name the host's paths.
|
||||
|
||||
```sh
|
||||
make docker
|
||||
docker run --rm -v /srv:/srv:ro -v /var/lib/sfdupes:/var/lib/sfdupes \
|
||||
sfdupes scan /srv
|
||||
docker run --rm -v /var/lib/sfdupes:/var/lib/sfdupes sfdupes report > dupes.tsv
|
||||
```
|
||||
|
||||
### Daily scan from cron
|
||||
|
||||
Run `scan` as root, so that it can read every file: a path it cannot
|
||||
read is skipped with a warning and loses its database record (see
|
||||
"Rules for the walk"). As a file `/etc/cron.d/sfdupes`:
|
||||
|
||||
```
|
||||
30 3 * * * root /usr/local/bin/sfdupes scan /srv 2>>/var/log/sfdupes.log || tail -n 3 /var/log/sfdupes.log
|
||||
```
|
||||
|
||||
- The database is `/var/lib/sfdupes/db.sqlite`, created with its
|
||||
directory by the first scan. To keep it elsewhere, set
|
||||
`SFDUPES_DATABASE=/path/to/db.sqlite` before the command on the same
|
||||
line.
|
||||
- `scan` writes nothing to stdout. Its stderr, appended here to
|
||||
`/var/log/sfdupes.log`, holds a plain progress line as each phase
|
||||
starts and then at most every 5 seconds, a warning for each path it
|
||||
skips, and the summary line (see "Progress" and "`scan` mode"). The
|
||||
log grows with every scan; rotate it like any other.
|
||||
- Skipped paths do not fail a scan: it still exits 0, and cron sends
|
||||
nothing. A scan that fails, or is stopped by `SIGINT` or `SIGTERM`,
|
||||
exits 1 with the reason among the last lines of the log; `tail`
|
||||
prints them, and cron mails them to root if the host can send mail.
|
||||
- A scan still running when the next one starts carries on. The new
|
||||
one fails at once, and the lines cron mails include
|
||||
`sfdupes: another scan is running (lock held on /var/lib/sfdupes/db.sqlite.lock)`.
|
||||
- `report` and `trees` need only read access to the database (see
|
||||
"Database"). Under the usual umask of `022` the first scan creates
|
||||
it readable by every user, so an unprivileged user can run them
|
||||
against root's database.
|
||||
|
||||
### Reading the reports
|
||||
|
||||
Each row of `report` names two copies of one file, and each row of
|
||||
`trees` two copies of one directory tree (see "Report output format"
|
||||
and "Trees output format"). In a group of copies, the path that sorts
|
||||
first byte by byte is `first` and every other path is a `dupe` of it.
|
||||
`first` says nothing about which copy is the original or the oldest;
|
||||
which copy to keep is your choice.
|
||||
|
||||
A row is a candidate, not proof:
|
||||
|
||||
- The reports read only the database, so they show the files as of
|
||||
the last scan; a file may have changed or gone since.
|
||||
- A file of 50 MiB or more is compared only on samples of its content
|
||||
(see "Duplicate detection").
|
||||
- Paths that are hard links to one file are listed as duplicates, but
|
||||
they share their data, so removing one frees nothing.
|
||||
|
||||
Compare a pair byte for byte before removing either copy. For the row
|
||||
`/srv/a/big.iso`, `/srv/b/big-copy.iso`, `4294967296`:
|
||||
|
||||
```sh
|
||||
cmp /srv/a/big.iso /srv/b/big-copy.iso && echo identical
|
||||
[ /srv/a/big.iso -ef /srv/b/big-copy.iso ] && echo "hard links"
|
||||
```
|
||||
|
||||
`cmp` prints nothing and exits 0 only when every byte matches, and
|
||||
otherwise reports where the files differ. The second line prints
|
||||
`hard links` when the two paths are the same file, so removing either
|
||||
frees nothing.
|
||||
|
||||
A path holding a backslash, tab, newline or carriage return is escaped
|
||||
in the reports (see "Report output format"). Undo the escapes before
|
||||
using it. `printf '%b'` does exactly that, because every backslash in
|
||||
an escaped path starts one of the four escapes. Command substitution
|
||||
drops trailing newlines, so print an `x` after the path and remove it
|
||||
afterwards, or a path that ends in a newline names a different file:
|
||||
|
||||
```sh
|
||||
p="$(printf '%bx' '/srv/a/tab\tname.txt')"; p="${p%x}"
|
||||
cmp "$p" /srv/b/tab-copy.txt
|
||||
```
|
||||
|
||||
Check a `trees` row with `diff -r`, which compares the two trees file
|
||||
by file and also names anything present in only one of them, such as
|
||||
an empty directory or a symlink, which `trees` does not see.
|
||||
|
||||
## Rationale
|
||||
|
||||
Duplicate finders that hash entire files do not scale to the target
|
||||
@@ -150,12 +266,16 @@ sfdupes scan [--workers N] [-x] PATH...
|
||||
sfdupes report > dupes.tsv
|
||||
sfdupes trees > dupetrees.tsv
|
||||
sfdupes --version
|
||||
sfdupes [command] --help
|
||||
```
|
||||
|
||||
`sfdupes --version` (or `-v`) prints one line, `sfdupes VERSION`, to
|
||||
stdout and exits 0, writing nothing to stderr. `-h` or `--help`, alone
|
||||
or after a subcommand, prints the help text to stderr and exits 0,
|
||||
writing nothing to stdout.
|
||||
`--workers N` sets the size of each `scan` worker pool (default: the
|
||||
number of CPUs), and `-x` (`--one-file-system`) keeps the walk of each
|
||||
operand on that operand's filesystem; both are described under
|
||||
"`scan` mode". `sfdupes --version` (or `-v`) prints one line,
|
||||
`sfdupes VERSION`, to stdout and exits 0, writing nothing to stderr.
|
||||
`-h` or `--help`, alone or after a subcommand, prints the help text to
|
||||
stderr and exits 0, writing nothing to stdout.
|
||||
|
||||
### Database
|
||||
|
||||
|
||||
@@ -29,6 +29,10 @@
|
||||
|
||||
# Completed Steps
|
||||
|
||||
- README documents install, Docker, a daily cron scan and how to read
|
||||
and check the reports (2026-10-04,
|
||||
https://git.eeqj.de/sneak/sfdupes/issues/54)
|
||||
|
||||
- the `Dockerfile` build stage keeps the Go module cache out of `builder`'s
|
||||
home and copies the sources with `--chown`, so no `chown -R` walks them
|
||||
(2026-10-04, https://git.eeqj.de/sneak/sfdupes/issues/43)
|
||||
|
||||
Reference in New Issue
Block a user