A backup reports success but cannot restore a changed file after the snapshot holding its current version is dropped #214

Closed
opened 2026-10-06 01:49:40 +02:00 by clawbot · 1 comment
Collaborator

File rows in the local index are shared by every snapshot and are updated in place. A changed file keeps its files.id; its old file_chunks rows are deleted, and the row is upserted with the new metadata and chunk list (internal/snapshot/scanner.go:711-717, :1187-1190, internal/database/files.go:346-353). Orphan cleanup works differently for each table. CleanupOrphanedData (internal/snapshot/snapshot.go:337-380) deletes blobs that no snapshot_blobs row references. It keeps files that any snapshot still lists, and chunks that any file_chunks row references.

The failure: an older snapshot still lists a file, and the only snapshot referencing that file's current blob is dropped. The file row survives with the new size, mtime and chunk hashes, but no blob holds those chunks any more. The next run compares only size, mtime, mode, uid and gid (scanner.go:1212-1216), so it treats the file as unchanged and attaches it without re-chunking (scanner.go:1006-1009). PopulateReferencedBlobs finds no blob for it, and the snapshot is still marked complete. Restore then fails with chunk missing from blob map (internal/vaultik/restore_plan.go:75). Both shallow and deep verify pass, and every later snapshot carries the same hole until the file changes on disk again.

Two triggers, both reproduced on next at 0700901 with the file:// backend:

  • snapshot remove of the newest snapshot after a file changed (internal/vaultik/snapshot.go:1136-1151), then a new backup.
  • An interrupted backup after an earlier completed one (for example, the manifest upload fails). The next run's PruneDatabase drops the incomplete snapshot and its blobs (snapshot.go:1678-1698), while the older snapshot keeps the file rows alive. syncWithRemote and CleanupLocalSnapshots dropping a local record have the same effect.

Acceptable: a file is treated as unchanged only if an uploaded blob holds every chunk it lists. #148 established the same rule for chunks.

Definition of done

  1. A known file whose chunks are not all held by uploaded blobs is re-chunked on the next run, whichever snapshots list it.
  2. Tests cover both triggers end to end (remove the newest snapshot and then back up; an interrupted run after a completed snapshot and then a backup), and each test asserts that the new snapshot restores the changed file.
  3. make check passes.

Model: fable-5-1 (audit); opus-5-5 (issue)

File rows in the local index are shared by every snapshot and are updated in place. A changed file keeps its `files.id`; its old `file_chunks` rows are deleted, and the row is upserted with the new metadata and chunk list (`internal/snapshot/scanner.go:711-717`, `:1187-1190`, `internal/database/files.go:346-353`). Orphan cleanup works differently for each table. `CleanupOrphanedData` (`internal/snapshot/snapshot.go:337-380`) deletes blobs that no `snapshot_blobs` row references. It keeps files that any snapshot still lists, and chunks that any `file_chunks` row references. The failure: an older snapshot still lists a file, and the only snapshot referencing that file's current blob is dropped. The file row survives with the new size, mtime and chunk hashes, but no blob holds those chunks any more. The next run compares only size, mtime, mode, uid and gid (`scanner.go:1212-1216`), so it treats the file as unchanged and attaches it without re-chunking (`scanner.go:1006-1009`). `PopulateReferencedBlobs` finds no blob for it, and the snapshot is still marked complete. Restore then fails with `chunk missing from blob map` (`internal/vaultik/restore_plan.go:75`). Both shallow and deep verify pass, and every later snapshot carries the same hole until the file changes on disk again. Two triggers, both reproduced on `next` at `0700901` with the `file://` backend: - `snapshot remove` of the newest snapshot after a file changed (`internal/vaultik/snapshot.go:1136-1151`), then a new backup. - An interrupted backup after an earlier completed one (for example, the manifest upload fails). The next run's `PruneDatabase` drops the incomplete snapshot and its blobs (`snapshot.go:1678-1698`), while the older snapshot keeps the file rows alive. `syncWithRemote` and `CleanupLocalSnapshots` dropping a local record have the same effect. Acceptable: a file is treated as unchanged only if an uploaded blob holds every chunk it lists. https://git.eeqj.de/sneak/vaultik/issues/148 established the same rule for chunks. ## Definition of done 1. A known file whose chunks are not all held by uploaded blobs is re-chunked on the next run, whichever snapshots list it. 2. Tests cover both triggers end to end (remove the newest snapshot and then back up; an interrupted run after a completed snapshot and then a backup), and each test asserts that the new snapshot restores the changed file. 3. `make check` passes. Model: fable-5-1 (audit); opus-5-5 (issue)
clawbot self-assigned this 2026-10-06 01:49:40 +02:00
Author
Collaborator

Fixed in #234: a backup now re-chunks a known file that lists a chunk no uploaded blob holds, even when its metadata is unchanged. Both triggers have end-to-end tests that restore the changed file from the next snapshot.

Model: opus-5-5

Fixed in https://git.eeqj.de/sneak/vaultik/pulls/234: a backup now re-chunks a known file that lists a chunk no uploaded blob holds, even when its metadata is unchanged. Both triggers have end-to-end tests that restore the changed file from the next snapshot. Model: opus-5-5
Sign in to join this conversation.