From 72f9a8f8a0a9368b362a5f1579850c38d8473eb6 Mon Sep 17 00:00:00 2001 From: sneak Date: Wed, 7 Oct 2026 12:35:07 +0000 Subject: [PATCH] Correct doc and help sentences that are false about the code (closes #233) A blob is written in full to a temporary file in $TMPDIR and uploaded once finished, not streamed to storage, and the metadata export works on a copy of the local index there. The README, ARCHITECTURE.md and config.example.yml now say a backup needs free space there of the larger of blob_size_limit and about three times the size of the local index. Also corrected: the snapshot ID format, what restore reads and how incomplete snapshots are removed in docs/DATAMODEL.md, what source_path holds, the index_path default, the config search order in the snapshot create help, what snapshot remove cleans up in the prune help, how the release installs Go, and the script/release and script/fmt-check comments. Model: opus-5-5 --- ARCHITECTURE.md | 10 +++++++--- README.md | 17 +++++++++-------- TODO.md | 14 ++++++++++++++ config.example.yml | 11 +++++++++-- docs/DATAMODEL.md | 11 +++++------ internal/cli/prune.go | 8 +++++--- internal/cli/snapshot.go | 5 +++-- internal/database/models.go | 3 +-- internal/database/schema/001.sql | 2 +- internal/snapshot/scanner.go | 10 ++++------ internal/types/types.go | 4 ++-- script/fmt-check | 7 +++++-- script/release | 7 +++---- 13 files changed, 68 insertions(+), 41 deletions(-) diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 0541199..f312100 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -54,10 +54,12 @@ The database tracks five primary entities and their relationships: #### File (`database.File`) Represents a file, directory, or symlink in the backup system. Stores metadata needed for restoration: -- Path, source_path (for restore path stripping), mtime +- Path, mtime - Size, mode, ownership (uid, gid) - Symlink target (if applicable) +It also stores `source_path`, the source directory the scan found it under, made absolute and with symlinks resolved. Restore does not read it. + #### Chunk (`database.Chunk`) A content-addressed unit of data. Files are split into variable-size chunks using the FastCDC algorithm: - `ChunkHash`: SHA256 hash of chunk content (primary key) @@ -82,9 +84,11 @@ The final storage unit uploaded to S3. Contains many compressed and encrypted ch Blob creation process: 1. Chunks are accumulated (up to MaxBlobSize, typically 10GB) 2. As each chunk is added, its uncompressed bytes are fed to a running SHA-256 -3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and streamed to storage +3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and written to a temporary file in `$TMPDIR` (`/tmp` when unset) 4. On finalize, the blob's name is the double SHA-256 of the uncompressed contents — `hex(SHA256(SHA256(...)))` — not a hash of the compressed, encrypted bytes -5. Uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}` +5. The finished file is uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}` and then deleted + +The metadata export at the end of a backup also uses `$TMPDIR`: it copies the local index there and runs `VACUUM` on the copy, which writes another temporary copy and a write-ahead log. A backup therefore needs free space in `$TMPDIR` of the larger of `blob_size_limit` and about three times the size of the local index. #### BlobChunk (`database.BlobChunk`) Maps chunks to their position within blobs: diff --git a/README.md b/README.md index cf9e39d..6198f1e 100644 --- a/README.md +++ b/README.md @@ -546,7 +546,7 @@ complete annotated example also lives in | `s3.*` | | Legacy S3 configuration (endpoint, bucket, credentials) | | `exclude` | | Global exclude patterns (applied to all snapshots) | | `chunk_size` | `10MB` | Average chunk size for content-defined chunking | -| `blob_size_limit` | `10GB` | Maximum blob size before splitting. Must be at least four times `chunk_size` (the largest chunk the chunker can emit), otherwise a single-chunk blob could exceed the limit | +| `blob_size_limit` | `10GB` | Maximum blob size before splitting. Must be at least four times `chunk_size` (the largest chunk the chunker can emit), otherwise a single-chunk blob could exceed the limit. Each blob is written in full to a temporary file in `$TMPDIR` (`/tmp` when unset) before it is uploaded, and the metadata export at the end of a backup works on a copy of the local index there. A backup needs free space there of the larger of this limit and about three times the size of the local index | | `compression_level` | `3` | zstd compression level (1-19) | | `hostname` | system hostname | Hostname used in snapshot IDs | | `index_path` | platform data dir | Local SQLite index path | @@ -805,7 +805,8 @@ them. We provide: in the `Dockerfile` together. * `script/release` — cross-compile and publish the release artifacts with the pinned `goreleaser`. Refuses a `goreleaser` on `PATH` whose - version is not the pinned one, on the same reasoning as `script/lint`. + version is not the pinned one, because a different version would build + a different release from the same tag. * `script/release-snapshot` — the same build with no publishing and no tagging, into `./dist` * `script/test` — run the test suite by building the `test` phase of @@ -915,14 +916,14 @@ It is passed to `goreleaser` as `GITEA_TOKEN`. The runner's automatic token is deliberately not used: it is not guaranteed to carry release write access. -The Go toolchain that compiles the released binaries comes from an -`actions/setup-go` step pinned by commit sha, reading its version from -`go.mod` (currently `1.26.1`, the same version the `Dockerfile` builder -stage pins by digest). `goreleaser` shells out to `go` for every +The Go toolchain that compiles the released binaries is installed by +`script/install-go`, which downloads the version named by `go.mod` +(currently `1.26.1`, the same version the `Dockerfile` builder stage +pins by digest) and refuses the archive unless its sha256 matches the +value committed in the script. `goreleaser` shells out to `go` for every cross-compile, so without that step the release would either fail outright or ship binaries built by whatever unpinned toolchain the -runner happened to carry — the one unpinned thing in an otherwise -hash-pinned release path. +runner happened to carry. To rehearse the whole build without publishing or tagging anything: diff --git a/TODO.md b/TODO.md index 2f33a00..edb3aeb 100644 --- a/TODO.md +++ b/TODO.md @@ -22,6 +22,20 @@ the tag exists and is exercised; what is left is merging `next` to # Completed Steps +- 2026-10-07: Corrected documentation, help text and comments that were + false about the code + ([issue #233](https://git.eeqj.de/sneak/vaultik/issues/233)). A blob + is not streamed to storage; it is written in full to a temporary file + in `$TMPDIR` and uploaded once finished, and the metadata export works + on a copy of the local index there. The README and + `config.example.yml` now say a backup needs free space there of the + larger of `blob_size_limit` and about three times the size of the + local index. Also corrected: the snapshot ID format, what restore + reads and how incomplete snapshots are removed in `docs/DATAMODEL.md`, + what `source_path` holds, the `index_path` and config file defaults, + what `snapshot remove` cleans up, how the release gets its Go + toolchain, and the `script/release` and `script/fmt-check` comments. + - 2026-10-07: Made two messages say only what is true ([issue #240](https://git.eeqj.de/sneak/vaultik/issues/240)). A config file that others can read was warned about as containing S3 diff --git a/config.example.yml b/config.example.yml index 3ad5870..ff96fff 100644 --- a/config.example.yml +++ b/config.example.yml @@ -295,8 +295,10 @@ storage_url: "rclone://myremote/path/to/backups" # Path to local SQLite index database # This database tracks file state for incremental backups -# Default: /var/lib/vaultik/index.sqlite -#index_path: /var/lib/vaultik/index.sqlite +# Default: the platform data directory, e.g. +# macOS: ~/Library/Application Support/vaultik/index.sqlite +# Linux: ~/.local/share/vaultik/index.sqlite +#index_path: /path/to/index.sqlite # Average chunk size for content-defined chunking # Smaller chunks = better deduplication but more metadata @@ -311,6 +313,11 @@ storage_url: "rclone://myremote/path/to/backups" # Chunking uses no secret (the FastCDC parameters are fixed and public). At a # large limit a blob holds hundreds of chunks, so individual chunk lengths are # not visible in its size; lowering the limit toward chunk_size exposes them. +# Each blob is written in full to a temporary file in $TMPDIR (/tmp when +# unset) before it is uploaded, and the metadata export at the end of a +# backup works on a copy of the local index there. A backup needs free space +# there of the larger of this limit and about three times the size of the +# local index. # Supports: 1GB, 10G, 500MB, 1GiB, etc. # Default: 10GB #blob_size_limit: 10GB diff --git a/docs/DATAMODEL.md b/docs/DATAMODEL.md index 285c52e..bcb9d1c 100644 --- a/docs/DATAMODEL.md +++ b/docs/DATAMODEL.md @@ -111,7 +111,7 @@ Maps chunks to the blobs that contain them. Tracks backup snapshots. **Columns:** -- `id` (TEXT PRIMARY KEY) - Snapshot ID (format: hostname-YYYYMMDD-HHMMSSZ) +- `id` (TEXT PRIMARY KEY) - Snapshot ID (format: `hostname_name_timestamp`, e.g. `server1_home_2025-06-01T12:00:00Z`: the hostname up to its first `.`, the snapshot name, and an RFC 3339 UTC timestamp) - `hostname` (TEXT) - Hostname where backup was created - `vaultik_version` (TEXT) - Version of Vaultik used - `vaultik_git_revision` (TEXT) - Git revision of Vaultik used @@ -218,8 +218,8 @@ The `{remote-key}` directory name is a one-way hash of the human snapshot ID, so ### 4. Restore Process The restore process doesn't use the local database. Instead: -1. Downloads snapshot metadata from S3 -2. Downloads required blobs based on manifest +1. Downloads and decrypts the snapshot's metadata database (`db.zst.age`) from S3 +2. Downloads the blobs holding the chunks of the files being restored, found through that database's `blob_chunks` table; the manifest is not read 3. Reconstructs files from decrypted and decompressed chunks ### 5. Pruning @@ -232,9 +232,8 @@ The restore process doesn't use the local database. Instead: Before each backup: 1. Query incomplete snapshots (where `completed_at IS NULL`) -2. Check if metadata exists in S3 -3. If no metadata, delete snapshot and all associations -4. Clean up orphaned files, chunks, and blobs +2. Delete each one and all its associations, without checking S3 for its metadata +3. Clean up orphaned files, chunks, and blobs ## Repository Pattern diff --git a/internal/cli/prune.go b/internal/cli/prune.go index ef3cef6..e73d45d 100644 --- a/internal/cli/prune.go +++ b/internal/cli/prune.go @@ -22,9 +22,11 @@ scans every snapshot manifest in the destination store, builds the set of still-referenced blob hashes, and deletes any blob not in that set. -Snapshot create --prune and snapshot remove run the same cleanup -automatically; this command is the manual entry point for the same -work (e.g. after a crashed backup or to reclaim storage).`, +Snapshot create --prune runs the same cleanup automatically; this +command is the manual entry point for the same work (e.g. after a +crashed backup or to reclaim storage). Snapshot remove leaves blobs in +place; run this command afterwards to delete the ones no longer +referenced.`, Args: cobra.NoArgs, RunE: func(cmd *cobra.Command, _ []string) error { // Use unified config resolution diff --git a/internal/cli/snapshot.go b/internal/cli/snapshot.go index fb2dc4d..a9c6e5a 100644 --- a/internal/cli/snapshot.go +++ b/internal/cli/snapshot.go @@ -66,8 +66,9 @@ func newSnapshotCreateCommand() *cobra.Command { If snapshot names are provided, only those snapshots are created. If no names are provided, all configured snapshots are created. -Config is located at /etc/vaultik/config.yml by default, but can be overridden by -specifying a path using --config or by setting VAULTIK_CONFIG to a path.`, +The config is read from the path given by --config or VAULTIK_CONFIG; +otherwise from the platform config directory (~/.config/vaultik/config.yml +on Linux), then /etc/vaultik/config.yml.`, Args: cobra.ArbitraryArgs, RunE: func(cmd *cobra.Command, args []string) error { // Pass snapshot names from args diff --git a/internal/database/models.go b/internal/database/models.go index 9c23c59..c71b1cd 100644 --- a/internal/database/models.go +++ b/internal/database/models.go @@ -14,8 +14,7 @@ type File struct { ID types.FileID // UUID primary key Path types.FilePath // Absolute path of the file - // SourcePath is the source directory this file came from (used for - // restore path stripping). + // SourcePath is the source directory this file came from. SourcePath types.SourcePath MTime time.Time Size int64 diff --git a/internal/database/schema/001.sql b/internal/database/schema/001.sql index 243c0f8..c07d978 100644 --- a/internal/database/schema/001.sql +++ b/internal/database/schema/001.sql @@ -5,7 +5,7 @@ CREATE TABLE IF NOT EXISTS files ( id TEXT PRIMARY KEY, -- UUID path TEXT NOT NULL UNIQUE, - source_path TEXT NOT NULL DEFAULT '', -- The source directory this file came from (for restore path stripping) + source_path TEXT NOT NULL DEFAULT '', -- The source directory this file came from mtime INTEGER NOT NULL, -- whole seconds since the Unix epoch mtime_nsec INTEGER NOT NULL, -- nanoseconds within that second, 0 to 999999999 size INTEGER NOT NULL, diff --git a/internal/snapshot/scanner.go b/internal/snapshot/scanner.go index 24fcaa7..b611857 100644 --- a/internal/snapshot/scanner.go +++ b/internal/snapshot/scanner.go @@ -58,8 +58,8 @@ type Scanner struct { compressionLevel int ageRecipient string snapshotID string // Current snapshot being processed - // currentSourcePath is the source directory being scanned (used for - // restore path stripping). + // currentSourcePath is the source directory being scanned, stored with + // each file record. currentSourcePath string exclude []string // Glob patterns for files/directories to exclude compiledExclude []compiledPattern // Compiled glob patterns @@ -211,7 +211,6 @@ func (s *Scanner) Scan( ctx context.Context, path string, snapshotID string, ) (*ScanResult, error) { s.snapshotID = snapshotID - // Store source path for file records (used during restore) s.currentSourcePath = path result := &ScanResult{ StartTime: time.Now().UTC(), @@ -1198,9 +1197,8 @@ func (s *Scanner) checkFileInMemory( } file := &database.File{ - ID: fileID, - Path: types.FilePath(path), - // Store source directory for restore path stripping + ID: fileID, + Path: types.FilePath(path), SourcePath: types.SourcePath(s.currentSourcePath), MTime: info.ModTime(), Size: info.Size(), diff --git a/internal/types/types.go b/internal/types/types.go index d4dc997..e9fdd72 100644 --- a/internal/types/types.go +++ b/internal/types/types.go @@ -155,8 +155,8 @@ type BlobHash string // FilePath represents an absolute path to a file or directory. type FilePath string -// SourcePath represents the root directory from which files are backed up. -// Used during restore to strip the source prefix from paths. +// SourcePath is the source directory a scan found a file under, made +// absolute and with symlinks resolved. type SourcePath string // Hostname identifies a host machine. diff --git a/script/fmt-check b/script/fmt-check index 84c23df..9a9a333 100755 --- a/script/fmt-check +++ b/script/fmt-check @@ -1,6 +1,9 @@ #!/bin/sh -# script/fmt-check: check formatting (read-only). Same scope as -# script/fmt, but fails instead of writing. +# script/fmt-check: check formatting (read-only). Fails instead of +# writing. It checks every Go file outside .tool, which is more than +# script/fmt formats: `go fmt ./...` skips `testdata` directories and +# files and directories whose names start with `.` or `_`. Fix a file +# only this reports with `gofmt -w`. set -eu ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" diff --git a/script/release b/script/release index 05ec0ae..cdc1b1c 100755 --- a/script/release +++ b/script/release @@ -21,10 +21,9 @@ goreleaser_version() { head -n 1 } -# Resolve the goreleaser to run, on the same rule script/lint uses for -# golangci-lint: a binary on PATH is accepted only when it is exactly -# the pinned version, because a differently versioned tool would -# produce a differently built release from the same tag. Anything else +# Resolve the goreleaser to run. A binary on PATH is accepted only when +# it is exactly the pinned version, because a differently versioned tool +# would produce a differently built release from the same tag. Anything else # comes from .tool/bin, and a missing one is a loud failure naming the # script that installs it rather than a silent fallback. resolve_goreleaser() {