Correct doc and help sentences that are false about the code (closes #233)
check / check (push) Waiting to run

A blob is written in full to a temporary file in $TMPDIR and uploaded
once finished, not streamed to storage, and the metadata export works on
a copy of the local index there. The README, ARCHITECTURE.md and
config.example.yml now say a backup needs free space there of the larger
of blob_size_limit and about three times the size of the local index.
Also corrected: the snapshot ID format, what restore reads and how
incomplete snapshots are removed in docs/DATAMODEL.md, what source_path
holds, the index_path default, the config search order in the snapshot
create help, what snapshot remove cleans up in the prune help, how the
release installs Go, and the script/release and script/fmt-check
comments.

Model: opus-5-5
This commit is contained in:
2026-10-07 14:04:10 +00:00
parent d53202eb86
commit 72f9a8f8a0
13 changed files with 68 additions and 41 deletions
+7 -3
View File
@@ -54,10 +54,12 @@ The database tracks five primary entities and their relationships:
#### File (`database.File`)
Represents a file, directory, or symlink in the backup system. Stores metadata needed for restoration:
- Path, source_path (for restore path stripping), mtime
- Path, mtime
- Size, mode, ownership (uid, gid)
- Symlink target (if applicable)
It also stores `source_path`, the source directory the scan found it under, made absolute and with symlinks resolved. Restore does not read it.
#### Chunk (`database.Chunk`)
A content-addressed unit of data. Files are split into variable-size chunks using the FastCDC algorithm:
- `ChunkHash`: SHA256 hash of chunk content (primary key)
@@ -82,9 +84,11 @@ The final storage unit uploaded to S3. Contains many compressed and encrypted ch
Blob creation process:
1. Chunks are accumulated (up to MaxBlobSize, typically 10GB)
2. As each chunk is added, its uncompressed bytes are fed to a running SHA-256
3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and streamed to storage
3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and written to a temporary file in `$TMPDIR` (`/tmp` when unset)
4. On finalize, the blob's name is the double SHA-256 of the uncompressed contents — `hex(SHA256(SHA256(...)))` — not a hash of the compressed, encrypted bytes
5. Uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}`
5. The finished file is uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}` and then deleted
The metadata export at the end of a backup also uses `$TMPDIR`: it copies the local index there and runs `VACUUM` on the copy, which writes another temporary copy and a write-ahead log. A backup therefore needs free space in `$TMPDIR` of the larger of `blob_size_limit` and about three times the size of the local index.
#### BlobChunk (`database.BlobChunk`)
Maps chunks to their position within blobs:
+9 -8
View File
@@ -546,7 +546,7 @@ complete annotated example also lives in
| `s3.*` | | Legacy S3 configuration (endpoint, bucket, credentials) |
| `exclude` | | Global exclude patterns (applied to all snapshots) |
| `chunk_size` | `10MB` | Average chunk size for content-defined chunking |
| `blob_size_limit` | `10GB` | Maximum blob size before splitting. Must be at least four times `chunk_size` (the largest chunk the chunker can emit), otherwise a single-chunk blob could exceed the limit |
| `blob_size_limit` | `10GB` | Maximum blob size before splitting. Must be at least four times `chunk_size` (the largest chunk the chunker can emit), otherwise a single-chunk blob could exceed the limit. Each blob is written in full to a temporary file in `$TMPDIR` (`/tmp` when unset) before it is uploaded, and the metadata export at the end of a backup works on a copy of the local index there. A backup needs free space there of the larger of this limit and about three times the size of the local index |
| `compression_level` | `3` | zstd compression level (1-19) |
| `hostname` | system hostname | Hostname used in snapshot IDs |
| `index_path` | platform data dir | Local SQLite index path |
@@ -805,7 +805,8 @@ them. We provide:
in the `Dockerfile` together.
* `script/release` — cross-compile and publish the release artifacts
with the pinned `goreleaser`. Refuses a `goreleaser` on `PATH` whose
version is not the pinned one, on the same reasoning as `script/lint`.
version is not the pinned one, because a different version would build
a different release from the same tag.
* `script/release-snapshot` — the same build with no publishing and no
tagging, into `./dist`
* `script/test` — run the test suite by building the `test` phase of
@@ -915,14 +916,14 @@ It is passed to `goreleaser` as `GITEA_TOKEN`. The runner's automatic
token is deliberately not used: it is not guaranteed to carry release
write access.
The Go toolchain that compiles the released binaries comes from an
`actions/setup-go` step pinned by commit sha, reading its version from
`go.mod` (currently `1.26.1`, the same version the `Dockerfile` builder
stage pins by digest). `goreleaser` shells out to `go` for every
The Go toolchain that compiles the released binaries is installed by
`script/install-go`, which downloads the version named by `go.mod`
(currently `1.26.1`, the same version the `Dockerfile` builder stage
pins by digest) and refuses the archive unless its sha256 matches the
value committed in the script. `goreleaser` shells out to `go` for every
cross-compile, so without that step the release would either fail
outright or ship binaries built by whatever unpinned toolchain the
runner happened to carry — the one unpinned thing in an otherwise
hash-pinned release path.
runner happened to carry.
To rehearse the whole build without publishing or tagging anything:
+14
View File
@@ -22,6 +22,20 @@ the tag exists and is exercised; what is left is merging `next` to
# Completed Steps
- 2026-10-07: Corrected documentation, help text and comments that were
false about the code
([issue #233](https://git.eeqj.de/sneak/vaultik/issues/233)). A blob
is not streamed to storage; it is written in full to a temporary file
in `$TMPDIR` and uploaded once finished, and the metadata export works
on a copy of the local index there. The README and
`config.example.yml` now say a backup needs free space there of the
larger of `blob_size_limit` and about three times the size of the
local index. Also corrected: the snapshot ID format, what restore
reads and how incomplete snapshots are removed in `docs/DATAMODEL.md`,
what `source_path` holds, the `index_path` and config file defaults,
what `snapshot remove` cleans up, how the release gets its Go
toolchain, and the `script/release` and `script/fmt-check` comments.
- 2026-10-07: Made two messages say only what is true
([issue #240](https://git.eeqj.de/sneak/vaultik/issues/240)). A config
file that others can read was warned about as containing S3
+9 -2
View File
@@ -295,8 +295,10 @@ storage_url: "rclone://myremote/path/to/backups"
# Path to local SQLite index database
# This database tracks file state for incremental backups
# Default: /var/lib/vaultik/index.sqlite
#index_path: /var/lib/vaultik/index.sqlite
# Default: the platform data directory, e.g.
# macOS: ~/Library/Application Support/vaultik/index.sqlite
# Linux: ~/.local/share/vaultik/index.sqlite
#index_path: /path/to/index.sqlite
# Average chunk size for content-defined chunking
# Smaller chunks = better deduplication but more metadata
@@ -311,6 +313,11 @@ storage_url: "rclone://myremote/path/to/backups"
# Chunking uses no secret (the FastCDC parameters are fixed and public). At a
# large limit a blob holds hundreds of chunks, so individual chunk lengths are
# not visible in its size; lowering the limit toward chunk_size exposes them.
# Each blob is written in full to a temporary file in $TMPDIR (/tmp when
# unset) before it is uploaded, and the metadata export at the end of a
# backup works on a copy of the local index there. A backup needs free space
# there of the larger of this limit and about three times the size of the
# local index.
# Supports: 1GB, 10G, 500MB, 1GiB, etc.
# Default: 10GB
#blob_size_limit: 10GB
+5 -6
View File
@@ -111,7 +111,7 @@ Maps chunks to the blobs that contain them.
Tracks backup snapshots.
**Columns:**
- `id` (TEXT PRIMARY KEY) - Snapshot ID (format: hostname-YYYYMMDD-HHMMSSZ)
- `id` (TEXT PRIMARY KEY) - Snapshot ID (format: `hostname_name_timestamp`, e.g. `server1_home_2025-06-01T12:00:00Z`: the hostname up to its first `.`, the snapshot name, and an RFC 3339 UTC timestamp)
- `hostname` (TEXT) - Hostname where backup was created
- `vaultik_version` (TEXT) - Version of Vaultik used
- `vaultik_git_revision` (TEXT) - Git revision of Vaultik used
@@ -218,8 +218,8 @@ The `{remote-key}` directory name is a one-way hash of the human snapshot ID, so
### 4. Restore Process
The restore process doesn't use the local database. Instead:
1. Downloads snapshot metadata from S3
2. Downloads required blobs based on manifest
1. Downloads and decrypts the snapshot's metadata database (`db.zst.age`) from S3
2. Downloads the blobs holding the chunks of the files being restored, found through that database's `blob_chunks` table; the manifest is not read
3. Reconstructs files from decrypted and decompressed chunks
### 5. Pruning
@@ -232,9 +232,8 @@ The restore process doesn't use the local database. Instead:
Before each backup:
1. Query incomplete snapshots (where `completed_at IS NULL`)
2. Check if metadata exists in S3
3. If no metadata, delete snapshot and all associations
4. Clean up orphaned files, chunks, and blobs
2. Delete each one and all its associations, without checking S3 for its metadata
3. Clean up orphaned files, chunks, and blobs
## Repository Pattern
+5 -3
View File
@@ -22,9 +22,11 @@ scans every snapshot manifest in the destination store, builds the
set of still-referenced blob hashes, and deletes any blob not in that
set.
Snapshot create --prune and snapshot remove run the same cleanup
automatically; this command is the manual entry point for the same
work (e.g. after a crashed backup or to reclaim storage).`,
Snapshot create --prune runs the same cleanup automatically; this
command is the manual entry point for the same work (e.g. after a
crashed backup or to reclaim storage). Snapshot remove leaves blobs in
place; run this command afterwards to delete the ones no longer
referenced.`,
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, _ []string) error {
// Use unified config resolution
+3 -2
View File
@@ -66,8 +66,9 @@ func newSnapshotCreateCommand() *cobra.Command {
If snapshot names are provided, only those snapshots are created.
If no names are provided, all configured snapshots are created.
Config is located at /etc/vaultik/config.yml by default, but can be overridden by
specifying a path using --config or by setting VAULTIK_CONFIG to a path.`,
The config is read from the path given by --config or VAULTIK_CONFIG;
otherwise from the platform config directory (~/.config/vaultik/config.yml
on Linux), then /etc/vaultik/config.yml.`,
Args: cobra.ArbitraryArgs,
RunE: func(cmd *cobra.Command, args []string) error {
// Pass snapshot names from args
+1 -2
View File
@@ -14,8 +14,7 @@ type File struct {
ID types.FileID // UUID primary key
Path types.FilePath // Absolute path of the file
// SourcePath is the source directory this file came from (used for
// restore path stripping).
// SourcePath is the source directory this file came from.
SourcePath types.SourcePath
MTime time.Time
Size int64
+1 -1
View File
@@ -5,7 +5,7 @@
CREATE TABLE IF NOT EXISTS files (
id TEXT PRIMARY KEY, -- UUID
path TEXT NOT NULL UNIQUE,
source_path TEXT NOT NULL DEFAULT '', -- The source directory this file came from (for restore path stripping)
source_path TEXT NOT NULL DEFAULT '', -- The source directory this file came from
mtime INTEGER NOT NULL, -- whole seconds since the Unix epoch
mtime_nsec INTEGER NOT NULL, -- nanoseconds within that second, 0 to 999999999
size INTEGER NOT NULL,
+4 -6
View File
@@ -58,8 +58,8 @@ type Scanner struct {
compressionLevel int
ageRecipient string
snapshotID string // Current snapshot being processed
// currentSourcePath is the source directory being scanned (used for
// restore path stripping).
// currentSourcePath is the source directory being scanned, stored with
// each file record.
currentSourcePath string
exclude []string // Glob patterns for files/directories to exclude
compiledExclude []compiledPattern // Compiled glob patterns
@@ -211,7 +211,6 @@ func (s *Scanner) Scan(
ctx context.Context, path string, snapshotID string,
) (*ScanResult, error) {
s.snapshotID = snapshotID
// Store source path for file records (used during restore)
s.currentSourcePath = path
result := &ScanResult{
StartTime: time.Now().UTC(),
@@ -1198,9 +1197,8 @@ func (s *Scanner) checkFileInMemory(
}
file := &database.File{
ID: fileID,
Path: types.FilePath(path),
// Store source directory for restore path stripping
ID: fileID,
Path: types.FilePath(path),
SourcePath: types.SourcePath(s.currentSourcePath),
MTime: info.ModTime(),
Size: info.Size(),
+2 -2
View File
@@ -155,8 +155,8 @@ type BlobHash string
// FilePath represents an absolute path to a file or directory.
type FilePath string
// SourcePath represents the root directory from which files are backed up.
// Used during restore to strip the source prefix from paths.
// SourcePath is the source directory a scan found a file under, made
// absolute and with symlinks resolved.
type SourcePath string
// Hostname identifies a host machine.
+5 -2
View File
@@ -1,6 +1,9 @@
#!/bin/sh
# script/fmt-check: check formatting (read-only). Same scope as
# script/fmt, but fails instead of writing.
# script/fmt-check: check formatting (read-only). Fails instead of
# writing. It checks every Go file outside .tool, which is more than
# script/fmt formats: `go fmt ./...` skips `testdata` directories and
# files and directories whose names start with `.` or `_`. Fix a file
# only this reports with `gofmt -w`.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
+3 -4
View File
@@ -21,10 +21,9 @@ goreleaser_version() {
head -n 1
}
# Resolve the goreleaser to run, on the same rule script/lint uses for
# golangci-lint: a binary on PATH is accepted only when it is exactly
# the pinned version, because a differently versioned tool would
# produce a differently built release from the same tag. Anything else
# Resolve the goreleaser to run. A binary on PATH is accepted only when
# it is exactly the pinned version, because a differently versioned tool
# would produce a differently built release from the same tag. Anything else
# comes from .tool/bin, and a missing one is a loud failure naming the
# script that installs it rather than a silent fallback.
resolve_goreleaser() {