Compare commits

1 Commits
Author SHA1 Message Date
sneak fb746ab5e5 Re-chunk a known file whose chunks no uploaded blob holds (closes #214)
check / check (pull_request) Successful in 8m16s
File rows are shared by every snapshot and updated in place, while a
blob row is deleted once no snapshot references it. Removing the newest
snapshot, or the prune after an interrupted run, could drop the only
blob holding a changed file's current chunks while an older snapshot
kept the file row. The next backup compared metadata only, skipped the
file, and completed a snapshot that could not restore it.

The scanner now loads the IDs of known files that list a chunk no
uploaded blob holds and re-chunks them even when their metadata is
unchanged.

The tests give each backup run its own snapshot name, so the
second-precision snapshot IDs differ without sleeping.

Model: opus-5-5
2026-10-06 00:47:31 +00:00
3 changed files with 12 additions and 8 deletions
+5 -3
View File
@@ -40,7 +40,8 @@ Features:
* modern encryption ([age](https://age-encryption.org/), X25519 + ChaCha20-Poly1305) * modern encryption ([age](https://age-encryption.org/), X25519 + ChaCha20-Poly1305)
* content-defined chunking with deduplication (FastCDC) * content-defined chunking with deduplication (FastCDC)
* incremental backups (only changed files are re-chunked) * incremental backups (a file is re-chunked only when it changed or a
chunk it lists is held by no uploaded blob)
* multithreaded zstd compression at configurable levels * multithreaded zstd compression at configurable levels
* content-addressed immutable storage * content-addressed immutable storage
* local state tracking in SQLite (enables write-only incremental backups) * local state tracking in SQLite (enables write-only incremental backups)
@@ -499,8 +500,9 @@ format does and does not protect.
* Content-defined chunking using the FastCDC algorithm * Content-defined chunking using the FastCDC algorithm
* Average chunk size: configurable (default 10MB) * Average chunk size: configurable (default 10MB)
* Deduplication at file level (unchanged files skipped) and chunk level * Deduplication at file level (unchanged files skipped, unless a chunk
(identical chunks across files stored once) the file lists is held by no uploaded blob) and chunk level (identical
chunks across files stored once)
* Multiple chunks packed into blobs to reduce object count * Multiple chunks packed into blobs to reduce object count
### encryption ### encryption
+3 -3
View File
@@ -268,9 +268,9 @@ func (r *FileRepository) ListByPrefix(
// ListIDsWithChunksNotInUploadedBlobs returns the IDs of the files whose // ListIDsWithChunksNotInUploadedBlobs returns the IDs of the files whose
// path starts with prefix and that list at least one chunk held by no // path starts with prefix and that list at least one chunk held by no
// blob whose upload has completed (uploaded_ts set). Such a file's // blob whose upload has completed (uploaded_ts set). A new snapshot
// metadata can still match the file on disk while part of its data is // cannot reference such a chunk, so a backup must not treat the file as
// not in remote storage, so a backup must not treat it as unchanged. // unchanged even when its metadata matches the file on disk.
func (r *FileRepository) ListIDsWithChunksNotInUploadedBlobs( func (r *FileRepository) ListIDsWithChunksNotInUploadedBlobs(
ctx context.Context, prefix string, ctx context.Context, prefix string,
) ([]types.FileID, error) { ) ([]types.FileID, error) {
+4 -2
View File
@@ -344,7 +344,9 @@ func (s *Scanner) loadDatabaseState(
// while a blob row is deleted once no snapshot references it. Removing // while a blob row is deleted once no snapshot references it. Removing
// the only snapshot that references a changed file's current blob // the only snapshot that references a changed file's current blob
// therefore leaves an older snapshot keeping a file row whose metadata // therefore leaves an older snapshot keeping a file row whose metadata
// matches the disk but whose data is not in remote storage. // matches the disk while no blob the local index records as uploaded
// holds its chunks, so a new snapshot cannot reference them. The dropped
// blob can still be in remote storage until prune removes it.
func (s *Scanner) loadFilesToRechunk(ctx context.Context, path string) error { func (s *Scanner) loadFilesToRechunk(ctx context.Context, path string) error {
ids, err := s.repos.Files.ListIDsWithChunksNotInUploadedBlobs(ctx, path) ids, err := s.repos.Files.ListIDsWithChunksNotInUploadedBlobs(ctx, path)
if err != nil { if err != nil {
@@ -1244,7 +1246,7 @@ func (s *Scanner) checkFileInMemory(
return file, true return file, true
} }
// Part of its data is not in remote storage (see loadFilesToRechunk) // No uploaded blob holds one of its chunks (see loadFilesToRechunk)
if _, rechunk := s.filesToRechunk[fileID]; rechunk { if _, rechunk := s.filesToRechunk[fileID]; rechunk {
return file, true return file, true
} }