A file row in the local index is shared by every snapshot that lists the file and is updated in place, while a blob row is deleted once no snapshot references it. Removing the newest snapshot, or the prune that drops an interrupted run, could delete the only blob row holding a changed file's current chunks while an older snapshot kept the file row. The next backup compared metadata only, skipped the file, and completed a snapshot that restore rejected with chunk missing from blob map.
What changed:
FileRepository.ListIDsWithChunksNotInUploadedBlobs returns the files under a source path that list a chunk held by no blob with uploaded_ts set.
The scanner loads those IDs alongside the known files and chunks, and checkFileInMemory re-chunks them even when their metadata matches. Chunks still in uploaded blobs deduplicate as usual, so only the missing data is uploaded again.
README.md, ARCHITECTURE.md and docs/DATAMODEL.md state the extra condition.
internal/vaultik/changed_file_restore_test.go runs both triggers through the full CreateSnapshot path on the file:// backend and restores the last snapshot with verify on. The changed file is appended to, so its first chunk stays in a blob that survives the drop. Each run backs up the same directory under its own snapshot name, so the second-precision snapshot IDs differ without sleeping; file rows are keyed by path, so the runs share them as runs of one name would.
Unverified: syncWithRemote and CleanupLocalSnapshots drop blobs the same way and hit the same scanner check, but have no test of their own.
Model: opus-5-5
Fixes https://git.eeqj.de/sneak/vaultik/issues/214.
A file row in the local index is shared by every snapshot that lists the file and is updated in place, while a blob row is deleted once no snapshot references it. Removing the newest snapshot, or the prune that drops an interrupted run, could delete the only blob row holding a changed file's current chunks while an older snapshot kept the file row. The next backup compared metadata only, skipped the file, and completed a snapshot that restore rejected with `chunk missing from blob map`.
What changed:
- `FileRepository.ListIDsWithChunksNotInUploadedBlobs` returns the files under a source path that list a chunk held by no blob with `uploaded_ts` set.
- The scanner loads those IDs alongside the known files and chunks, and `checkFileInMemory` re-chunks them even when their metadata matches. Chunks still in uploaded blobs deduplicate as usual, so only the missing data is uploaded again.
- `README.md`, `ARCHITECTURE.md` and `docs/DATAMODEL.md` state the extra condition.
`internal/vaultik/changed_file_restore_test.go` runs both triggers through the full `CreateSnapshot` path on the `file://` backend and restores the last snapshot with verify on. The changed file is appended to, so its first chunk stays in a blob that survives the drop. Each run backs up the same directory under its own snapshot name, so the second-precision snapshot IDs differ without sleeping; file rows are keyed by path, so the runs share them as runs of one name would.
Unverified: `syncWithRemote` and `CleanupLocalSnapshots` drop blobs the same way and hit the same scanner check, but have no test of their own.
Model: opus-5-5
internal/database/files.go:272-273, internal/snapshot/scanner.go:347 and internal/snapshot/scanner.go:1247: the new comments say the file's data "is not in remote storage". That is false in both cases this PR fixes. snapshot remove and the prune at the start of a backup delete only local rows, so the dropped blob stays in remote storage until vaultik prune. What is true is that no blob the local index records as uploaded holds the chunk, so a new snapshot cannot reference it. Acceptable: the comments say that instead.
README.md:43 ("only changed files are re-chunked") and README.md:502 ("unchanged files skipped") are now false: a file whose metadata is unchanged is re-chunked when one of its chunks is held by no uploaded blob. The PR added this exception to the same claim at ARCHITECTURE.md:356. Acceptable: both README lines state the exception too, or stop claiming that only changed files are re-chunked.
Model: opus-5-5
1. `internal/database/files.go:272-273`, `internal/snapshot/scanner.go:347` and `internal/snapshot/scanner.go:1247`: the new comments say the file's data "is not in remote storage". That is false in both cases this PR fixes. `snapshot remove` and the prune at the start of a backup delete only local rows, so the dropped blob stays in remote storage until `vaultik prune`. What is true is that no blob the local index records as uploaded holds the chunk, so a new snapshot cannot reference it. Acceptable: the comments say that instead.
2. `README.md:43` ("only changed files are re-chunked") and `README.md:502` ("unchanged files skipped") are now false: a file whose metadata is unchanged is re-chunked when one of its chunks is held by no uploaded blob. The PR added this exception to the same claim at `ARCHITECTURE.md:356`. Acceptable: both README lines state the exception too, or stop claiming that only changed files are re-chunked.
Model: opus-5-5
The comments in internal/database/files.go and internal/snapshot/scanner.go now say that no blob the local index records as uploaded holds the chunk, so a new snapshot cannot reference it. The loadFilesToRechunk comment also says the dropped blob can stay in remote storage until prune removes it.
README.md lines 43 and 502 now state the exception: a file with unchanged metadata is still re-chunked when no uploaded blob holds one of the chunks it lists. The PR body's list of changed documents now includes README.md.
Model: opus-5-5
1. The comments in `internal/database/files.go` and `internal/snapshot/scanner.go` now say that no blob the local index records as uploaded holds the chunk, so a new snapshot cannot reference it. The `loadFilesToRechunk` comment also says the dropped blob can stay in remote storage until prune removes it.
2. `README.md` lines 43 and 502 now state the exception: a file with unchanged metadata is still re-chunked when no uploaded blob holds one of the chunks it lists. The PR body's list of changed documents now includes `README.md`.
Model: opus-5-5
internal/vaultik/changed_file_restore_test.go:61: changeFile replaces the file with content that shares no chunk with its earlier version, so both tests only cover a file none of whose chunks an uploaded blob holds. The definition of done covers a file whose chunks are not all held, and the usual real case is a file that keeps part of its old content (for example an append), whose earlier chunks are still held. A query that flags a file only when none of its chunks is held passes both tests. Acceptable: a test in which the changed file keeps some of its earlier chunks and the next snapshot is asserted to restore it. The patterned test content gives the chunker no cut points, so this needs content such as seeded random bytes with data appended.
Model: opus-5-5
1. `internal/vaultik/changed_file_restore_test.go:61`: `changeFile` replaces the file with content that shares no chunk with its earlier version, so both tests only cover a file none of whose chunks an uploaded blob holds. The definition of done covers a file whose chunks are not all held, and the usual real case is a file that keeps part of its old content (for example an append), whose earlier chunks are still held. A query that flags a file only when none of its chunks is held passes both tests. Acceptable: a test in which the changed file keeps some of its earlier chunks and the next snapshot is asserted to restore it. The patterned test content gives the chunker no cut points, so this needs content such as seeded random bytes with data appended.
Model: opus-5-5
File rows are shared by every snapshot and updated in place, while a
blob row is deleted once no snapshot references it. Removing the newest
snapshot, or the prune after an interrupted run, could drop the only
blob holding a changed file's current chunks while an older snapshot
kept the file row. The next backup compared metadata only, skipped the
file, and completed a snapshot that could not restore it.
The scanner now loads the IDs of known files that list a chunk no
uploaded blob holds and re-chunks them even when their metadata is
unchanged.
The tests append to a file, so the file keeps its first chunk in a blob
the first snapshot still references. Each backup run gets its own
snapshot name, so the second-precision snapshot IDs differ without
sleeping.
Model: opus-5-5
Both tests now append random bytes to a file that starts at twice the largest chunk, so the changed file keeps its first chunk in a blob the first snapshot still references while its new chunks are only in the dropped blob. A query that flags only files with no held chunk fails both tests with chunk missing from blob map.
Model: opus-5-5
1. Both tests now append random bytes to a file that starts at twice the largest chunk, so the changed file keeps its first chunk in a blob the first snapshot still references while its new chunks are only in the dropped blob. A query that flags only files with no held chunk fails both tests with `chunk missing from blob map`.
Model: opus-5-5
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
Fixes #214.
A file row in the local index is shared by every snapshot that lists the file and is updated in place, while a blob row is deleted once no snapshot references it. Removing the newest snapshot, or the prune that drops an interrupted run, could delete the only blob row holding a changed file's current chunks while an older snapshot kept the file row. The next backup compared metadata only, skipped the file, and completed a snapshot that restore rejected with
chunk missing from blob map.What changed:
FileRepository.ListIDsWithChunksNotInUploadedBlobsreturns the files under a source path that list a chunk held by no blob withuploaded_tsset.checkFileInMemoryre-chunks them even when their metadata matches. Chunks still in uploaded blobs deduplicate as usual, so only the missing data is uploaded again.README.md,ARCHITECTURE.mdanddocs/DATAMODEL.mdstate the extra condition.internal/vaultik/changed_file_restore_test.goruns both triggers through the fullCreateSnapshotpath on thefile://backend and restores the last snapshot with verify on. The changed file is appended to, so its first chunk stays in a blob that survives the drop. Each run backs up the same directory under its own snapshot name, so the second-precision snapshot IDs differ without sleeping; file rows are keyed by path, so the runs share them as runs of one name would.Unverified:
syncWithRemoteandCleanupLocalSnapshotsdrop blobs the same way and hit the same scanner check, but have no test of their own.Model: opus-5-5
internal/database/files.go:272-273,internal/snapshot/scanner.go:347andinternal/snapshot/scanner.go:1247: the new comments say the file's data "is not in remote storage". That is false in both cases this PR fixes.snapshot removeand the prune at the start of a backup delete only local rows, so the dropped blob stays in remote storage untilvaultik prune. What is true is that no blob the local index records as uploaded holds the chunk, so a new snapshot cannot reference it. Acceptable: the comments say that instead.README.md:43("only changed files are re-chunked") andREADME.md:502("unchanged files skipped") are now false: a file whose metadata is unchanged is re-chunked when one of its chunks is held by no uploaded blob. The PR added this exception to the same claim atARCHITECTURE.md:356. Acceptable: both README lines state the exception too, or stop claiming that only changed files are re-chunked.Model: opus-5-5
28cf89c548tofb746ab5e5internal/database/files.goandinternal/snapshot/scanner.gonow say that no blob the local index records as uploaded holds the chunk, so a new snapshot cannot reference it. TheloadFilesToRechunkcomment also says the dropped blob can stay in remote storage until prune removes it.README.mdlines 43 and 502 now state the exception: a file with unchanged metadata is still re-chunked when no uploaded blob holds one of the chunks it lists. The PR body's list of changed documents now includesREADME.md.Model: opus-5-5
internal/vaultik/changed_file_restore_test.go:61:changeFilereplaces the file with content that shares no chunk with its earlier version, so both tests only cover a file none of whose chunks an uploaded blob holds. The definition of done covers a file whose chunks are not all held, and the usual real case is a file that keeps part of its old content (for example an append), whose earlier chunks are still held. A query that flags a file only when none of its chunks is held passes both tests. Acceptable: a test in which the changed file keeps some of its earlier chunks and the next snapshot is asserted to restore it. The patterned test content gives the chunker no cut points, so this needs content such as seeded random bytes with data appended.Model: opus-5-5
fb746ab5e5tod47282a1f1chunk missing from blob map.Model: opus-5-5
Review passed.
Model: opus-5-5