Correct doc and help sentences that are false about the code (closes #233)
check / check (push) Waiting to run

A blob is written whole to a temporary file and uploaded once finished,
not streamed to storage. The README, ARCHITECTURE.md and
config.example.yml now say a backup needs free temporary space for each
blob (twice that for an rclone destination that cannot stream uploads)
and for the metadata export's copies of the local index, in $TMPDIR, or
partly in /var/tmp when TMPDIR is unset. Also corrected: the snapshot ID
format, what restore reads and how incomplete snapshots are removed in
docs/DATAMODEL.md, what source_path holds, the index_path default, the
config search order in the snapshot create help, what snapshot remove
cleans up in the prune help, how the release installs Go, and the
script/release and script/fmt-check comments.

Model: opus-5-5
This commit is contained in:
2026-10-07 15:07:02 +00:00
parent d53202eb86
commit c2ce69d258
13 changed files with 71 additions and 41 deletions
+7 -3
View File
@@ -54,10 +54,12 @@ The database tracks five primary entities and their relationships:
#### File (`database.File`)
Represents a file, directory, or symlink in the backup system. Stores metadata needed for restoration:
- Path, source_path (for restore path stripping), mtime
- Path, mtime
- Size, mode, ownership (uid, gid)
- Symlink target (if applicable)
It also stores `source_path`, the source directory the scan found it under, made absolute and with symlinks resolved. Restore does not read it.
#### Chunk (`database.Chunk`)
A content-addressed unit of data. Files are split into variable-size chunks using the FastCDC algorithm:
- `ChunkHash`: SHA256 hash of chunk content (primary key)
@@ -82,9 +84,11 @@ The final storage unit uploaded to S3. Contains many compressed and encrypted ch
Blob creation process:
1. Chunks are accumulated (up to MaxBlobSize, typically 10GB)
2. As each chunk is added, its uncompressed bytes are fed to a running SHA-256
3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and streamed to storage
3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and written to a temporary file
4. On finalize, the blob's name is the double SHA-256 of the uncompressed contents — `hex(SHA256(SHA256(...)))` — not a hash of the compressed, encrypted bytes
5. Uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}`
5. The finished file is uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}` and then deleted
A backup needs free temporary space, because each blob is written whole to a temporary file before it is uploaded (up to about `blob_size_limit`; an rclone destination that cannot stream uploads needs about twice that) and the metadata export writes copies of the local index. Temporary files go to `$TMPDIR` (default `/tmp`); with `TMPDIR` unset, SQLite writes one of those copies to `/var/tmp`.
#### BlobChunk (`database.BlobChunk`)
Maps chunks to their position within blobs: