1 Commits
Author SHA1 Message Date
sneak 28cf89c548 Re-chunk a known file whose chunks no uploaded blob holds (closes #214)
check / check (pull_request) Successful in 8m2s
File rows are shared by every snapshot and updated in place, while a
blob row is deleted once no snapshot references it. Removing the newest
snapshot, or the prune after an interrupted run, could drop the only
blob holding a changed file's current chunks while an older snapshot
kept the file row. The next backup compared metadata only, skipped the
file, and completed a snapshot that could not restore it.

The scanner now loads the IDs of known files that list a chunk no
uploaded blob holds and re-chunks them even when their metadata is
unchanged.

The tests give each backup run its own snapshot name, so the
second-precision snapshot IDs differ without sleeping.

Model: opus-5-5
2026-10-06 00:12:56 +00:00
3 changed files with 8 additions and 12 deletions
+3 -5
View File
@@ -40,8 +40,7 @@ Features:
* modern encryption ([age](https://age-encryption.org/), X25519 + ChaCha20-Poly1305) * modern encryption ([age](https://age-encryption.org/), X25519 + ChaCha20-Poly1305)
* content-defined chunking with deduplication (FastCDC) * content-defined chunking with deduplication (FastCDC)
* incremental backups (a file is re-chunked only when it changed or a * incremental backups (only changed files are re-chunked)
chunk it lists is held by no uploaded blob)
* multithreaded zstd compression at configurable levels * multithreaded zstd compression at configurable levels
* content-addressed immutable storage * content-addressed immutable storage
* local state tracking in SQLite (enables write-only incremental backups) * local state tracking in SQLite (enables write-only incremental backups)
@@ -500,9 +499,8 @@ format does and does not protect.
* Content-defined chunking using the FastCDC algorithm * Content-defined chunking using the FastCDC algorithm
* Average chunk size: configurable (default 10MB) * Average chunk size: configurable (default 10MB)
* Deduplication at file level (unchanged files skipped, unless a chunk * Deduplication at file level (unchanged files skipped) and chunk level
the file lists is held by no uploaded blob) and chunk level (identical (identical chunks across files stored once)
chunks across files stored once)
* Multiple chunks packed into blobs to reduce object count * Multiple chunks packed into blobs to reduce object count
### encryption ### encryption
+3 -3
View File
@@ -268,9 +268,9 @@ func (r *FileRepository) ListByPrefix(
// ListIDsWithChunksNotInUploadedBlobs returns the IDs of the files whose // ListIDsWithChunksNotInUploadedBlobs returns the IDs of the files whose
// path starts with prefix and that list at least one chunk held by no // path starts with prefix and that list at least one chunk held by no
// blob whose upload has completed (uploaded_ts set). A new snapshot // blob whose upload has completed (uploaded_ts set). Such a file's
// cannot reference such a chunk, so a backup must not treat the file as // metadata can still match the file on disk while part of its data is
// unchanged even when its metadata matches the file on disk. // not in remote storage, so a backup must not treat it as unchanged.
func (r *FileRepository) ListIDsWithChunksNotInUploadedBlobs( func (r *FileRepository) ListIDsWithChunksNotInUploadedBlobs(
ctx context.Context, prefix string, ctx context.Context, prefix string,
) ([]types.FileID, error) { ) ([]types.FileID, error) {
+2 -4
View File
@@ -344,9 +344,7 @@ func (s *Scanner) loadDatabaseState(
// while a blob row is deleted once no snapshot references it. Removing // while a blob row is deleted once no snapshot references it. Removing
// the only snapshot that references a changed file's current blob // the only snapshot that references a changed file's current blob
// therefore leaves an older snapshot keeping a file row whose metadata // therefore leaves an older snapshot keeping a file row whose metadata
// matches the disk while no blob the local index records as uploaded // matches the disk but whose data is not in remote storage.
// holds its chunks, so a new snapshot cannot reference them. The dropped
// blob can still be in remote storage until prune removes it.
func (s *Scanner) loadFilesToRechunk(ctx context.Context, path string) error { func (s *Scanner) loadFilesToRechunk(ctx context.Context, path string) error {
ids, err := s.repos.Files.ListIDsWithChunksNotInUploadedBlobs(ctx, path) ids, err := s.repos.Files.ListIDsWithChunksNotInUploadedBlobs(ctx, path)
if err != nil { if err != nil {
@@ -1246,7 +1244,7 @@ func (s *Scanner) checkFileInMemory(
return file, true return file, true
} }
// No uploaded blob holds one of its chunks (see loadFilesToRechunk) // Part of its data is not in remote storage (see loadFilesToRechunk)
if _, rechunk := s.filesToRechunk[fileID]; rechunk { if _, rechunk := s.filesToRechunk[fileID]; rechunk {
return file, true return file, true
} }