Correct doc and help sentences that are false about the code (closes #233)
check / check (push) Waiting to run

A blob is written in full to a temporary file in $TMPDIR and uploaded
once finished, not streamed to storage, and the metadata export works on
a copy of the local index there. The README, ARCHITECTURE.md and
config.example.yml now say a backup needs free space there of the larger
of blob_size_limit and about three times the size of the local index.
Also corrected: the snapshot ID format, what restore reads and how
incomplete snapshots are removed in docs/DATAMODEL.md, what source_path
holds, the index_path default, the config search order in the snapshot
create help, what snapshot remove cleans up in the prune help, how the
release installs Go, and the script/release and script/fmt-check
comments.

Model: opus-5-5
This commit is contained in:
2026-10-07 14:04:10 +00:00
parent d53202eb86
commit 72f9a8f8a0
13 changed files with 68 additions and 41 deletions
+7 -3
View File
@@ -54,10 +54,12 @@ The database tracks five primary entities and their relationships:
#### File (`database.File`)
Represents a file, directory, or symlink in the backup system. Stores metadata needed for restoration:
- Path, source_path (for restore path stripping), mtime
- Path, mtime
- Size, mode, ownership (uid, gid)
- Symlink target (if applicable)
It also stores `source_path`, the source directory the scan found it under, made absolute and with symlinks resolved. Restore does not read it.
#### Chunk (`database.Chunk`)
A content-addressed unit of data. Files are split into variable-size chunks using the FastCDC algorithm:
- `ChunkHash`: SHA256 hash of chunk content (primary key)
@@ -82,9 +84,11 @@ The final storage unit uploaded to S3. Contains many compressed and encrypted ch
Blob creation process:
1. Chunks are accumulated (up to MaxBlobSize, typically 10GB)
2. As each chunk is added, its uncompressed bytes are fed to a running SHA-256
3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and streamed to storage
3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and written to a temporary file in `$TMPDIR` (`/tmp` when unset)
4. On finalize, the blob's name is the double SHA-256 of the uncompressed contents — `hex(SHA256(SHA256(...)))` — not a hash of the compressed, encrypted bytes
5. Uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}`
5. The finished file is uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}` and then deleted
The metadata export at the end of a backup also uses `$TMPDIR`: it copies the local index there and runs `VACUUM` on the copy, which writes another temporary copy and a write-ahead log. A backup therefore needs free space in `$TMPDIR` of the larger of `blob_size_limit` and about three times the size of the local index.
#### BlobChunk (`database.BlobChunk`)
Maps chunks to their position within blobs: