Re-chunk a known file whose chunks no uploaded blob holds (closes #214)
File rows are shared by every snapshot and updated in place, while a blob row is deleted once no snapshot references it. Removing the newest snapshot, or the prune after an interrupted run, could drop the only blob holding a changed file's current chunks while an older snapshot kept the file row. The next backup compared metadata only, skipped the file, and completed a snapshot that could not restore it. The scanner now loads the IDs of known files that list a chunk no uploaded blob holds and re-chunks them even when their metadata is unchanged. The tests append to a file, so the file keeps its first chunk in a blob the first snapshot still references. Each backup run gets its own snapshot name, so the second-precision snapshot IDs differ without sleeping. Model: opus-5-5
This commit was merged in pull request #234.
This commit is contained in:
+1
-1
@@ -353,7 +353,7 @@ CreateSnapshot(opts)
|
||||
|
||||
## Deduplication Strategy
|
||||
|
||||
1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid)
|
||||
1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid), unless the file lists a chunk that no uploaded blob holds; such a file is re-chunked
|
||||
|
||||
2. **Chunk-level**: Chunks are content-addressed by SHA256 hash. If a chunk hash already exists in the database, the chunk data is not re-uploaded.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user