Store mtime to the nanosecond so a same-second rewrite is re-hashed (closes #12)
check / check (push) Waiting to run
check / check (push) Waiting to run
scan recorded mtime in whole seconds, so a file rewritten in place at the same size within the same second as its recorded mtime was classed unchanged and kept its old hashes. The files table keeps mtime as whole Unix seconds and gains mtime_nsec, the nanoseconds within that second. scan holds the mtime as a time.Time and decides "newer" by comparing Unix() and then Nanosecond(), so any time a filesystem can record compares in the right order; After would misorder one too late for a time.Time to hold without wrapping. The walk, a file given as an operand, and the content phase's recheck all move over. PRAGMA user_version stays 1, per the owner's ruling. README states what both columns hold. Model: opus-5-5
This commit is contained in:
@@ -312,29 +312,34 @@ All three subcommands operate on a single SQLite database file:
|
||||
|
||||
```sql
|
||||
CREATE TABLE files (
|
||||
path BLOB PRIMARY KEY, -- absolute path, raw bytes
|
||||
size INTEGER NOT NULL, -- bytes, from lstat
|
||||
mtime INTEGER NOT NULL, -- Unix seconds, from lstat
|
||||
head TEXT NOT NULL, -- lowercase-hex SHA-256; first 64 KiB, or whole file under 10 MiB
|
||||
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
|
||||
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
|
||||
path BLOB PRIMARY KEY, -- absolute path, raw bytes
|
||||
size INTEGER NOT NULL, -- bytes, from lstat
|
||||
mtime INTEGER NOT NULL, -- whole Unix seconds of the mtime, from lstat
|
||||
mtime_nsec INTEGER NOT NULL, -- nanoseconds within that second, 0 to 999999999
|
||||
head TEXT NOT NULL, -- lowercase-hex SHA-256; first 64 KiB, or whole file under 10 MiB
|
||||
tail TEXT NOT NULL, -- lowercase-hex SHA-256; last 64 KiB, or whole file under 10 MiB
|
||||
content TEXT NOT NULL -- lowercase-hex SHA-256, whole file or samples
|
||||
) WITHOUT ROWID;
|
||||
CREATE INDEX files_signature ON files (size, head, tail, content);
|
||||
```
|
||||
|
||||
Paths are stored as BLOBs because Unix paths are raw bytes, not guaranteed
|
||||
UTF-8. `mtime` is used only for change detection; it is not part of the
|
||||
duplicate key. For a file under 10 MiB `head`, `tail`, and `content` all
|
||||
hold the whole-file hash (that range is hashed in full, with no end
|
||||
windows); for a larger file `head` and `tail` hold the first- and last-64
|
||||
KiB hashes and `content` the whole-file or sampled hash. All three are empty
|
||||
strings when the file has never been hashed because its size was unique as
|
||||
of the last scan that covered it. For a file of 10 MiB or more, `content`
|
||||
stays empty until the content phase of a scan (see "`scan` mode" below) has
|
||||
read the file. A record with an empty `content` is never part of a duplicate
|
||||
group, though it still defines the file for tree reconstruction. The
|
||||
`files_signature` index lets SQLite group the records by signature for
|
||||
`report` without sorting the whole table.
|
||||
UTF-8. `mtime` and `mtime_nsec` hold the file's mtime to the nanosecond:
|
||||
`mtime` the whole Unix seconds, rounded down, and `mtime_nsec` the
|
||||
nanoseconds past that second. Split this way they hold any mtime a
|
||||
filesystem can record, one before 1678 or after 2262 included, which a
|
||||
single 64-bit count of nanoseconds cannot. They are used only for change
|
||||
detection and are not part of the duplicate key. For a file under 10 MiB
|
||||
`head`, `tail`, and `content` all hold the whole-file hash (that range is
|
||||
hashed in full, with no end windows); for a larger file `head` and `tail`
|
||||
hold the first- and last-64 KiB hashes and `content` the whole-file or
|
||||
sampled hash. All three are empty strings when the file has never been
|
||||
hashed because its size was unique as of the last scan that covered it. For
|
||||
a file of 10 MiB or more, `content` stays empty until the content phase of a
|
||||
scan (see "`scan` mode" below) has read the file. A record with an empty
|
||||
`content` is never part of a duplicate group, though it still defines the
|
||||
file for tree reconstruction. The `files_signature` index lets SQLite group
|
||||
the records by signature for `report` without sorting the whole table.
|
||||
|
||||
### Duplicate detection
|
||||
|
||||
@@ -421,6 +426,9 @@ operands:
|
||||
once its size, `head`, and `tail` match another record's.
|
||||
- A file whose mtime is newer than recorded, or whose size differs, is processed
|
||||
as if new: re-hashed, or recorded without hashes, per the shared-size rule.
|
||||
Change detection compares the mtime to the nanosecond, as finely as the
|
||||
filesystem records it, so a same-size rewrite counts as a change whenever the
|
||||
filesystem gives it a later mtime than recorded, even within the same second.
|
||||
- A database record whose path lies under one of the scanned operands but was
|
||||
not successfully processed this run is deleted. This removes records for
|
||||
deleted files. It also removes records for paths that failed to stat or hash
|
||||
|
||||
Reference in New Issue
Block a user