Check every group member and show the content phase from its start
check / check (push) Successful in 57s
check / check (push) Successful in 57s
The content phase now checks every record sharing a size, head and tail with lstat, including those that already have a content hash, and reads those without one only when at least two pass. Only a missing file is passed over silently; any other lstat error is warned about and counted as skipped. The phase shows a running count while it queries and checks, then its bar. The README states once when content is empty. New tests cover these and hard links. Model: opus-5-5
This commit is contained in:
@@ -184,11 +184,11 @@ All three subcommands operate on a single SQLite database file:
|
|||||||
the first- and last-64 KiB hashes and `content` the whole-file or
|
the first- and last-64 KiB hashes and `content` the whole-file or
|
||||||
sampled hash. All three are empty strings when the file has never
|
sampled hash. All three are empty strings when the file has never
|
||||||
been hashed because its size was unique as of the last scan that
|
been hashed because its size was unique as of the last scan that
|
||||||
covered it. `content` alone is empty for a file of 10 MiB or more
|
covered it. For a file of 10 MiB or more, `content` stays empty
|
||||||
that matches no other record on size, `head`, and `tail` yet, or
|
until the content phase of a scan (see "`scan` mode" below) has
|
||||||
whose content read failed. A record with an empty `content` still
|
read the file. A record with an empty `content` is never part of a
|
||||||
defines the file for tree reconstruction but is not a duplicate
|
duplicate group, though it still defines the file for tree
|
||||||
until a later scan fills it in.
|
reconstruction.
|
||||||
|
|
||||||
### Duplicate detection
|
### Duplicate detection
|
||||||
|
|
||||||
@@ -280,9 +280,9 @@ scanned operands:
|
|||||||
removes records for deleted files. It also removes records for
|
removes records for deleted files. It also removes records for
|
||||||
paths that failed to stat or hash this run: the database only ever
|
paths that failed to stat or hash this run: the database only ever
|
||||||
contains signatures verified by the most recent scan that covered
|
contains signatures verified by the most recent scan that covered
|
||||||
them (a subsequent successful scan re-adds such files). A failed
|
them (a subsequent successful scan re-adds such files). A failure
|
||||||
content read in the content phase below removes nothing: the
|
in the content phase below removes nothing: the record is left as
|
||||||
record keeps its `head` and `tail`, with `content` empty.
|
it is.
|
||||||
- Database records outside the scanned operands are untouched, so
|
- Database records outside the scanned operands are untouched, so
|
||||||
disjoint trees can be scanned on different schedules into the same
|
disjoint trees can be scanned on different schedules into the same
|
||||||
database. The one exception is the content phase below: a stored
|
database. The one exception is the content phase below: a stored
|
||||||
@@ -333,16 +333,19 @@ during the hash phase:
|
|||||||
`content` hash whose size, `head`, and `tail` equal another
|
`content` hash whose size, `head`, and `tail` equal another
|
||||||
record's, anywhere in the database: records from this scan and
|
record's, anywhere in the database: records from this scan and
|
||||||
records stored by earlier scans, inside or outside the scanned
|
records stored by earlier scans, inside or outside the scanned
|
||||||
operands. SQLite finds them, so only their records are loaded into
|
operands. SQLite finds them, so only the records to be read are
|
||||||
memory, never every file's hashes. Each such file is checked with
|
kept in memory, never every file's hashes. Every record sharing
|
||||||
`lstat` first; one that is gone, is no longer a regular file, or
|
their size, `head`, and `tail`, including one that already has a
|
||||||
has changed (a different size, or an mtime newer than recorded)
|
`content` hash, has its file checked with `lstat` first. A file
|
||||||
keeps its record as it is and is not a duplicate. The files that
|
that is gone, is no longer a regular file, or has changed (a
|
||||||
pass are read only if at least two records sharing their size,
|
different size, or an mtime newer than recorded) keeps its record
|
||||||
`head`, and `tail` remain, counting those that already have a
|
as it is and is not a duplicate. Any other `lstat` error is warned
|
||||||
`content` hash, so a file whose only matches are stale costs no
|
about and counted as skipped, with the same result. The files that
|
||||||
read. They are read by a worker pool as in the hash phase, in inode
|
pass and have no `content` hash are read only if at least two of
|
||||||
order and once per inode, and their content hashes are committed in
|
those records pass, so a file whose only matches are stale costs no
|
||||||
|
read; a file that already has a `content` hash is never read again.
|
||||||
|
They are read by a worker pool as in the hash phase, in inode order
|
||||||
|
and once per inode, and their content hashes are committed in
|
||||||
batches. A failed read is warned about and counted as skipped; its
|
batches. A failed read is warned about and counted as skipped; its
|
||||||
record keeps an empty `content`, so it is not a duplicate, and a
|
record keeps an empty `content`, so it is not a duplicate, and a
|
||||||
later scan tries again.
|
later scan tries again.
|
||||||
@@ -363,9 +366,9 @@ Rules for the walk:
|
|||||||
path, and continue. Per-file errors never abort the run; the final
|
path, and continue. Per-file errors never abort the run; the final
|
||||||
summary reports how many were skipped. As specified above, a
|
summary reports how many were skipped. As specified above, a
|
||||||
skipped path that has a database record from an earlier scan loses
|
skipped path that has a database record from an earlier scan loses
|
||||||
that record, unless only its content read failed; an unreadable
|
that record, unless it failed only in the content phase; an
|
||||||
directory subtree likewise loses its records (accepted: the
|
unreadable directory subtree likewise loses its records (accepted:
|
||||||
database mirrors what the latest scan could actually verify).
|
the database mirrors what the latest scan could actually verify).
|
||||||
|
|
||||||
Concurrency: the walk phase (which also stats files), the hash phase,
|
Concurrency: the walk phase (which also stats files), the hash phase,
|
||||||
and the content phase each use a worker pool of `--workers` workers
|
and the content phase each use a worker pool of `--workers` workers
|
||||||
@@ -399,10 +402,9 @@ mounted.
|
|||||||
|
|
||||||
Processing:
|
Processing:
|
||||||
|
|
||||||
- Records without a `content` hash (size-unique when last scanned,
|
- Records without a `content` hash (see "Database" above) are
|
||||||
or 10 MiB or more and not yet matched on size, `head`, and `tail`)
|
excluded: their content is unknown, so they are never reported as
|
||||||
are excluded: their content is unknown, so they are never reported
|
duplicates.
|
||||||
as duplicates.
|
|
||||||
- Group the remaining records by the key
|
- Group the remaining records by the key
|
||||||
`(size, head, tail, content)`.
|
`(size, head, tail, content)`.
|
||||||
- Every group with two or more paths is a duplicate group.
|
- Every group with two or more paths is a duplicate group.
|
||||||
@@ -507,10 +509,13 @@ Each phase gets its own display, rendered the moment the phase
|
|||||||
starts — a scan must never look hung. Loading the existing-record
|
starts — a scan must never look hung. Loading the existing-record
|
||||||
index (`load`) and the walk have no known totals while running: show
|
index (`load`) and the walk have no known totals while running: show
|
||||||
a live count, rate, and elapsed time (spinner-style, no percentage or
|
a live count, rate, and elapsed time (spinner-style, no percentage or
|
||||||
ETA). The hash, update, and content phases
|
ETA). The content phase's display (`content`) starts the same way,
|
||||||
have exact totals — only files that actually need hashing appear in
|
counting the records checked while SQLite finds the files to read and
|
||||||
the hash and content totals, so their ETAs are meaningful. Required
|
`lstat` checks them, then shows a bar once reading starts. The hash
|
||||||
elements for the bars with known totals:
|
and update phases, and the content phase's reads, have exact totals —
|
||||||
|
only files that actually need hashing appear in the hash and content
|
||||||
|
totals, so their ETAs are meaningful. Required elements for the bars
|
||||||
|
with known totals:
|
||||||
|
|
||||||
- elapsed time
|
- elapsed time
|
||||||
- estimated time remaining
|
- estimated time remaining
|
||||||
|
|||||||
@@ -279,31 +279,29 @@ func loadFileMeta(ctx context.Context, db *sql.DB,
|
|||||||
}
|
}
|
||||||
|
|
||||||
// contentCandidatesSQL selects every record of at least headTailMin
|
// contentCandidatesSQL selects every record of at least headTailMin
|
||||||
// bytes that has no content hash but whose size, head, and tail equal
|
// bytes whose size, head, and tail equal another record's, in each
|
||||||
// another record's, with the number of records in its group (the
|
// group (the records sharing a size, head, and tail) where at least one
|
||||||
// records sharing that size, head, and tail) that already have a
|
// record has no content hash, with whether each record has one. SQLite
|
||||||
// content hash. SQLite does the grouping, so no other record's hashes
|
// does the grouping, so no other record's hashes are loaded into
|
||||||
// are loaded into memory; the rows come ordered by size, head, and
|
// memory; the rows come ordered by size, head, and tail, so each
|
||||||
// tail, so each group's rows arrive together.
|
// group's rows arrive together.
|
||||||
const contentCandidatesSQL = `
|
const contentCandidatesSQL = `
|
||||||
SELECT f.path, f.size, f.mtime, f.head, f.tail, g.hashed
|
SELECT f.path, f.size, f.mtime, f.head, f.tail, f.content <> ''
|
||||||
FROM files AS f
|
FROM files AS f
|
||||||
JOIN (
|
JOIN (
|
||||||
SELECT size, head, tail, SUM(content <> '') AS hashed
|
SELECT size, head, tail
|
||||||
FROM files
|
FROM files
|
||||||
WHERE size >= ? AND head <> ''
|
WHERE size >= ? AND head <> ''
|
||||||
GROUP BY size, head, tail
|
GROUP BY size, head, tail
|
||||||
HAVING COUNT(*) > 1
|
HAVING COUNT(*) > 1 AND SUM(content = '') > 0
|
||||||
) AS g USING (size, head, tail)
|
) AS g USING (size, head, tail)
|
||||||
WHERE f.content = ''
|
|
||||||
ORDER BY size, head, tail
|
ORDER BY size, head, tail
|
||||||
`
|
`
|
||||||
|
|
||||||
// loadContentCandidates streams the rows of contentCandidatesSQL to fn:
|
// loadContentCandidates streams the rows of contentCandidatesSQL to fn:
|
||||||
// each record, without its content hash, and the number of records in
|
// each record, without its content hash, and whether it has one.
|
||||||
// its group that already have one.
|
|
||||||
func loadContentCandidates(ctx context.Context, db *sql.DB,
|
func loadContentCandidates(ctx context.Context, db *sql.DB,
|
||||||
fn func(r scanRec, hashed int),
|
fn func(r scanRec, hashed bool),
|
||||||
) error {
|
) error {
|
||||||
rows, err := db.QueryContext(ctx, contentCandidatesSQL, headTailMin)
|
rows, err := db.QueryContext(ctx, contentCandidatesSQL, headTailMin)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
@@ -316,7 +314,7 @@ func loadContentCandidates(ctx context.Context, db *sql.DB,
|
|||||||
var (
|
var (
|
||||||
path []byte
|
path []byte
|
||||||
r scanRec
|
r scanRec
|
||||||
hashed int
|
hashed int64
|
||||||
)
|
)
|
||||||
|
|
||||||
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail, &hashed)
|
err = rows.Scan(&path, &r.size, &r.mtime, &r.head, &r.tail, &hashed)
|
||||||
@@ -325,7 +323,7 @@ func loadContentCandidates(ctx context.Context, db *sql.DB,
|
|||||||
}
|
}
|
||||||
|
|
||||||
r.path = string(path)
|
r.path = string(path)
|
||||||
fn(r, hashed)
|
fn(r, hashed != 0)
|
||||||
}
|
}
|
||||||
|
|
||||||
err = rows.Err()
|
err = rows.Err()
|
||||||
|
|||||||
@@ -118,9 +118,7 @@ func collectDupeGroups(recs []scanRec) []dupeGroup {
|
|||||||
|
|
||||||
for _, r := range recs {
|
for _, r := range recs {
|
||||||
// A record without a content hash has unknown content and is
|
// A record without a content hash has unknown content and is
|
||||||
// never reported as a duplicate: its size was unique when last
|
// never reported as a duplicate (README "Database").
|
||||||
// scanned, or it is headTailMin or more and has not yet matched
|
|
||||||
// another record on size, head, and tail.
|
|
||||||
if r.content == "" {
|
if r.content == "" {
|
||||||
continue
|
continue
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -6,6 +6,7 @@ import (
|
|||||||
"crypto/sha256"
|
"crypto/sha256"
|
||||||
"database/sql"
|
"database/sql"
|
||||||
"encoding/hex"
|
"encoding/hex"
|
||||||
|
"errors"
|
||||||
"fmt"
|
"fmt"
|
||||||
"io"
|
"io"
|
||||||
"io/fs"
|
"io/fs"
|
||||||
@@ -578,7 +579,7 @@ func (s *scanState) updatePhase(ctx context.Context) error {
|
|||||||
// its empty content, so it is never grouped, and a later scan tries
|
// its empty content, so it is never grouped, and a later scan tries
|
||||||
// again.
|
// again.
|
||||||
func (s *scanState) contentPhase(ctx context.Context, workers int) error {
|
func (s *scanState) contentPhase(ctx context.Context, workers int) error {
|
||||||
toRead, recs, err := contentCandidates(ctx, s.db)
|
toRead, recs, err := s.contentCandidates(ctx)
|
||||||
if err != nil {
|
if err != nil {
|
||||||
return err
|
return err
|
||||||
}
|
}
|
||||||
@@ -602,47 +603,71 @@ func (s *scanState) contentPhase(ctx context.Context, workers int) error {
|
|||||||
return applyChanges(ctx, s.db, s.batch, nil, nil)
|
return applyChanges(ctx, s.db, s.batch, nil, nil)
|
||||||
}
|
}
|
||||||
|
|
||||||
// contentCandidates returns the files the content phase reads, and the
|
// contentCandidates returns the files the content phase reads, and
|
||||||
// records of the files that passed the check, by path. Every record
|
// their records by path. Every record contentCandidatesSQL returns has
|
||||||
// contentCandidatesSQL returns has its file checked with lstat: a file
|
// its file checked with lstat, whether or not it already has a content
|
||||||
// that is gone, is no longer a regular file, or has changed by the
|
// hash: a file that is gone, is no longer a regular file, or has
|
||||||
// walk's rule keeps its record as it is and is not a duplicate. The
|
// changed by the walk's rule keeps its record as it is and is not a
|
||||||
// files of a group that pass are read only if the group still has at
|
// duplicate, and any other lstat error is warned about and counted as
|
||||||
// least minGroupSize members, counting its records that already have a
|
// skipped. The files of a group that pass and have no content hash are
|
||||||
// content hash, so a group whose other members are all stale costs no
|
// read only if at least minGroupSize of the group's files pass, so a
|
||||||
// reads.
|
// group whose other members are all stale costs no reads. Only the
|
||||||
func contentCandidates(ctx context.Context,
|
// records to be read are kept.
|
||||||
db *sql.DB,
|
func (s *scanState) contentCandidates(
|
||||||
|
ctx context.Context,
|
||||||
) ([]fileRec, map[string]scanRec, error) {
|
) ([]fileRec, map[string]scanRec, error) {
|
||||||
|
// The query and the checks take real time on a large database;
|
||||||
|
// without a display the scan looks hung before the reads begin.
|
||||||
|
prog := newProgress("content", -1)
|
||||||
|
defer prog.finish()
|
||||||
|
|
||||||
var (
|
var (
|
||||||
toRead []fileRec
|
toRead []fileRec
|
||||||
passed []fileRec // the current group's files that passed the check
|
|
||||||
first scanRec // the current group's first record
|
first scanRec // the current group's first record
|
||||||
hashed int // the current group's records with a content hash
|
passed int // the current group's files that passed the check
|
||||||
|
unread []fileRec // those of them without a content hash
|
||||||
)
|
)
|
||||||
|
|
||||||
recs := make(map[string]scanRec)
|
recs := make(map[string]scanRec)
|
||||||
|
|
||||||
// endGroup queues the current group's files that passed the check,
|
// endGroup queues the current group's files to read if at least
|
||||||
// if the group still has at least minGroupSize members.
|
// minGroupSize of its files passed, and drops their records if not.
|
||||||
endGroup := func() {
|
endGroup := func() {
|
||||||
if len(passed)+hashed >= minGroupSize {
|
if passed >= minGroupSize {
|
||||||
toRead = append(toRead, passed...)
|
toRead = append(toRead, unread...)
|
||||||
|
} else {
|
||||||
|
for _, f := range unread {
|
||||||
|
delete(recs, f.path)
|
||||||
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
passed = nil
|
passed, unread = 0, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
err := loadContentCandidates(ctx, db, func(r scanRec, groupHashed int) {
|
err := loadContentCandidates(ctx, s.db, func(r scanRec, hashed bool) {
|
||||||
|
prog.increment()
|
||||||
|
|
||||||
if r.size != first.size || r.head != first.head || r.tail != first.tail {
|
if r.size != first.size || r.head != first.head || r.tail != first.tail {
|
||||||
endGroup()
|
endGroup()
|
||||||
|
|
||||||
first, hashed = r, groupHashed
|
first = r
|
||||||
}
|
}
|
||||||
|
|
||||||
f, ok := unchangedFile(r)
|
f, ok, err := unchangedFile(r)
|
||||||
if ok {
|
if err != nil {
|
||||||
passed = append(passed, f)
|
s.st.skipped++
|
||||||
|
|
||||||
|
prog.warnf("content %s: %v", r.path, err)
|
||||||
|
}
|
||||||
|
|
||||||
|
if !ok {
|
||||||
|
return
|
||||||
|
}
|
||||||
|
|
||||||
|
passed++
|
||||||
|
|
||||||
|
if !hashed {
|
||||||
|
unread = append(unread, f)
|
||||||
recs[r.path] = r
|
recs[r.path] = r
|
||||||
}
|
}
|
||||||
})
|
})
|
||||||
@@ -657,20 +682,28 @@ func contentCandidates(ctx context.Context,
|
|||||||
|
|
||||||
// unchangedFile lstats the file r names and returns it for reading if
|
// unchangedFile lstats the file r names and returns it for reading if
|
||||||
// it is still the regular file r records: the same size, and an mtime
|
// it is still the regular file r records: the same size, and an mtime
|
||||||
// no newer than recorded (the walk's change rule). Otherwise it
|
// no newer than recorded (the walk's change rule). A file that is gone
|
||||||
// reports false.
|
// or has changed reports false; any other lstat error is returned.
|
||||||
func unchangedFile(r scanRec) (fileRec, bool) {
|
func unchangedFile(r scanRec) (fileRec, bool, error) {
|
||||||
fi, err := os.Lstat(r.path)
|
fi, err := os.Lstat(r.path)
|
||||||
if err != nil || !fi.Mode().IsRegular() || fi.Size() != r.size ||
|
if errors.Is(err, fs.ErrNotExist) {
|
||||||
|
return fileRec{}, false, nil
|
||||||
|
}
|
||||||
|
|
||||||
|
if err != nil {
|
||||||
|
return fileRec{}, false, err
|
||||||
|
}
|
||||||
|
|
||||||
|
if !fi.Mode().IsRegular() || fi.Size() != r.size ||
|
||||||
fi.ModTime().Unix() > r.mtime {
|
fi.ModTime().Unix() > r.mtime {
|
||||||
return fileRec{}, false
|
return fileRec{}, false, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
dev, ino := inodeOfInfo(fi)
|
dev, ino := inodeOfInfo(fi)
|
||||||
|
|
||||||
return fileRec{
|
return fileRec{
|
||||||
path: r.path, size: r.size, mtime: r.mtime, dev: dev, ino: ino,
|
path: r.path, size: r.size, mtime: r.mtime, dev: dev, ino: ino,
|
||||||
}, true
|
}, true, nil
|
||||||
}
|
}
|
||||||
|
|
||||||
// underAnyRoot reports whether path is any of the roots or lies under
|
// underAnyRoot reports whether path is any of the roots or lies under
|
||||||
|
|||||||
+130
@@ -524,6 +524,58 @@ func TestScanContentStalePartners(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestScanContentHashedStalePartners checks that stored matches outside
|
||||||
|
// the operand that already have a content hash are checked like any
|
||||||
|
// other: once one has vanished and the other has changed, a copy of
|
||||||
|
// them scanned in another tree has no match left, so it is not read and
|
||||||
|
// is not reported as their duplicate.
|
||||||
|
func TestScanContentHashedStalePartners(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
db := openTestDB(t)
|
||||||
|
dirA := t.TempDir()
|
||||||
|
stored := []string{
|
||||||
|
sparseFile(t, dirA, "changed", headTailMin),
|
||||||
|
sparseFile(t, dirA, "gone", headTailMin),
|
||||||
|
}
|
||||||
|
|
||||||
|
// The two stored files match, so this scan gives both a content
|
||||||
|
// hash.
|
||||||
|
syncTree(t, db, dirA)
|
||||||
|
|
||||||
|
err := os.Remove(stored[1])
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
future := time.Now().Add(time.Hour)
|
||||||
|
|
||||||
|
err = os.Chtimes(stored[0], future, future)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
b := sparseFile(t, t.TempDir(), "copy", headTailMin)
|
||||||
|
|
||||||
|
st := syncTree(t, db, filepath.Dir(b))
|
||||||
|
if st != (scanStats{added: 1}) {
|
||||||
|
t.Errorf("stats = %+v, want 1 added and nothing skipped", st)
|
||||||
|
}
|
||||||
|
|
||||||
|
recs := dbRecords(t, db)
|
||||||
|
if r := recordByPath(t, recs, b); r.content != "" {
|
||||||
|
t.Errorf("copy: content = %q, want none: its only matches are stale",
|
||||||
|
r.content)
|
||||||
|
}
|
||||||
|
|
||||||
|
// The stored records lie outside the operand and are left as they
|
||||||
|
// are, so they still group with each other, but not with the copy.
|
||||||
|
groups := collectDupeGroups(recs)
|
||||||
|
if len(groups) != 1 || !slices.Equal(groups[0].paths, stored) {
|
||||||
|
t.Errorf("groups = %+v, want only the stored pair %q", groups, stored)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// TestScanContentReadFailure checks that a failed content read is
|
// TestScanContentReadFailure checks that a failed content read is
|
||||||
// counted as skipped and leaves the record without a content hash, and
|
// counted as skipped and leaves the record without a content hash, and
|
||||||
// that a later scan tries the read again.
|
// that a later scan tries the read again.
|
||||||
@@ -574,6 +626,84 @@ func TestScanContentReadFailure(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
|
// TestScanContentCheckError checks that a stored file the content phase
|
||||||
|
// cannot lstat, for a reason other than its being gone, is counted as
|
||||||
|
// skipped and does not count as a match.
|
||||||
|
func TestScanContentCheckError(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
db := openTestDB(t)
|
||||||
|
sub := filepath.Join(t.TempDir(), "sub")
|
||||||
|
|
||||||
|
err := os.Mkdir(sub, 0o700)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
sparseFileWithoutMatch(t, sub, "a", headTailMin)
|
||||||
|
syncTree(t, db, sub)
|
||||||
|
|
||||||
|
// Without search permission on its directory, the stored file's
|
||||||
|
// lstat fails with permission denied.
|
||||||
|
err = os.Chmod(sub, 0)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
t.Cleanup(func() {
|
||||||
|
//nolint:gosec // removing the directory needs its search bit back
|
||||||
|
_ = os.Chmod(sub, 0o700)
|
||||||
|
})
|
||||||
|
|
||||||
|
b := sparseFile(t, t.TempDir(), "b", headTailMin)
|
||||||
|
|
||||||
|
st := syncTree(t, db, filepath.Dir(b))
|
||||||
|
if st != (scanStats{added: 1, skipped: 1}) {
|
||||||
|
t.Fatalf("stats = %+v, want 1 added 1 skipped", st)
|
||||||
|
}
|
||||||
|
|
||||||
|
if r := recordByPath(t, dbRecords(t, db), b); r.content != "" {
|
||||||
|
t.Errorf("b: content = %q, want none: its only match could not be "+
|
||||||
|
"checked", r.content)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
// TestScanContentHardlinks checks that the content phase stores the
|
||||||
|
// content hash of a hard-linked file on every one of its links.
|
||||||
|
func TestScanContentHardlinks(t *testing.T) {
|
||||||
|
t.Parallel()
|
||||||
|
|
||||||
|
dir := t.TempDir()
|
||||||
|
db := openTestDB(t)
|
||||||
|
a := sparseFile(t, dir, "a", headTailMin)
|
||||||
|
b := filepath.Join(dir, "b")
|
||||||
|
|
||||||
|
err := os.Link(a, b)
|
||||||
|
if err != nil {
|
||||||
|
t.Fatal(err)
|
||||||
|
}
|
||||||
|
|
||||||
|
c := sparseFile(t, dir, "copy", headTailMin)
|
||||||
|
|
||||||
|
st := syncTree(t, db, dir)
|
||||||
|
if st != (scanStats{added: 3}) {
|
||||||
|
t.Fatalf("stats = %+v, want 3 added", st)
|
||||||
|
}
|
||||||
|
|
||||||
|
recs := dbRecords(t, db)
|
||||||
|
|
||||||
|
want := recordByPath(t, recs, c).content
|
||||||
|
if want == "" {
|
||||||
|
t.Fatal("the copy has no content hash")
|
||||||
|
}
|
||||||
|
|
||||||
|
for _, p := range []string{a, b} {
|
||||||
|
if got := recordByPath(t, recs, p).content; got != want {
|
||||||
|
t.Errorf("%s: content = %q, want %q", p, got, want)
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
// collectWalk runs a walk over roots and returns the emitted records
|
// collectWalk runs a walk over roots and returns the emitted records
|
||||||
// and the number of warning events.
|
// and the number of warning events.
|
||||||
func collectWalk(t *testing.T, roots []string, oneFS bool,
|
func collectWalk(t *testing.T, roots []string, oneFS bool,
|
||||||
|
|||||||
@@ -127,10 +127,8 @@ func buildHierarchy(recs []scanRec) (*treeNode, []*treeNode) {
|
|||||||
size: r.size, head: r.head, tail: r.tail, content: r.content,
|
size: r.size, head: r.head, tail: r.tail, content: r.content,
|
||||||
}
|
}
|
||||||
|
|
||||||
// A record without a content hash (its size was unique when
|
// A record without a content hash has unknown content (README
|
||||||
// last scanned, or it is headTailMin or more and has not yet
|
// "Database"): give it a signature no other file can share, so
|
||||||
// matched another record on size, head, and tail) has unknown
|
|
||||||
// content: give it a signature no other file can share, so
|
|
||||||
// trees containing it never compare equal. Real hashes are
|
// trees containing it never compare equal. Real hashes are
|
||||||
// hex, so the NUL-prefixed form cannot collide.
|
// hex, so the NUL-prefixed form cannot collide.
|
||||||
if sig.content == "" {
|
if sig.content == "" {
|
||||||
|
|||||||
Reference in New Issue
Block a user