Re-chunk a known file whose chunks no uploaded blob holds (closes #214)
check / check (pull_request) Successful in 8m2s

File rows are shared by every snapshot and updated in place, while a
blob row is deleted once no snapshot references it. Removing the newest
snapshot, or the prune after an interrupted run, could drop the only
blob holding a changed file's current chunks while an older snapshot
kept the file row. The next backup compared metadata only, skipped the
file, and completed a snapshot that could not restore it.

The scanner now loads the IDs of known files that list a chunk no
uploaded blob holds and re-chunks them even when their metadata is
unchanged.

The tests give each backup run its own snapshot name, so the
second-precision snapshot IDs differ without sleeping.

Model: opus-5-5
This commit is contained in:
2026-10-06 00:12:56 +00:00
parent 070090124a
commit 28cf89c548
6 changed files with 315 additions and 1 deletions
+1 -1
View File
@@ -353,7 +353,7 @@ CreateSnapshot(opts)
## Deduplication Strategy
1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid)
1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid), unless the file lists a chunk that no uploaded blob holds; such a file is re-chunked
2. **Chunk-level**: Chunks are content-addressed by SHA256 hash. If a chunk hash already exists in the database, the chunk data is not re-uploaded.