Re-chunk a known file whose chunks no uploaded blob holds (closes #214)
check / check (push) Successful in 6m53s
check / check (pull_request) Successful in 6m17s

File rows are shared by every snapshot and updated in place, while a
blob row is deleted once no snapshot references it. Removing the newest
snapshot, or the prune after an interrupted run, could drop the only
blob holding a changed file's current chunks while an older snapshot
kept the file row. The next backup compared metadata only, skipped the
file, and completed a snapshot that could not restore it.

The scanner now loads the IDs of known files that list a chunk no
uploaded blob holds and re-chunks them even when their metadata is
unchanged.

The tests append to a file, so the file keeps its first chunk in a blob
the first snapshot still references. Each backup run gets its own
snapshot name, so the second-precision snapshot IDs differ without
sleeping.

Model: opus-5-5
This commit was merged in pull request #234.
This commit is contained in:
2026-10-06 04:46:16 +02:00
parent 070090124a
commit 35cf985c18
7 changed files with 343 additions and 4 deletions
+1 -1
View File
@@ -353,7 +353,7 @@ CreateSnapshot(opts)
## Deduplication Strategy
1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid)
1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid), unless the file lists a chunk that no uploaded blob holds; such a file is re-chunked
2. **Chunk-level**: Chunks are content-addressed by SHA256 hash. If a chunk hash already exists in the database, the chunk data is not re-uploaded.