Correct doc and help sentences that are false about the code (closes #233)
check / check (push) Waiting to run
check / check (push) Waiting to run
A blob is written whole to a temporary file and uploaded once finished, not streamed to storage. The README, ARCHITECTURE.md and config.example.yml now say a backup needs free temporary space for each blob (twice that for an rclone destination that cannot stream uploads) and for the metadata export's copies of the local index, in $TMPDIR, or partly in /var/tmp when TMPDIR is unset. Also corrected: the snapshot ID format, what restore reads and how incomplete snapshots are removed in docs/DATAMODEL.md, what source_path holds, the index_path default, the config search order in the snapshot create help, what snapshot remove cleans up in the prune help, how the release installs Go, and the script/release and script/fmt-check comments. Model: opus-5-5
This commit is contained in:
+7
-3
@@ -54,10 +54,12 @@ The database tracks five primary entities and their relationships:
|
||||
|
||||
#### File (`database.File`)
|
||||
Represents a file, directory, or symlink in the backup system. Stores metadata needed for restoration:
|
||||
- Path, source_path (for restore path stripping), mtime
|
||||
- Path, mtime
|
||||
- Size, mode, ownership (uid, gid)
|
||||
- Symlink target (if applicable)
|
||||
|
||||
It also stores `source_path`, the source directory the scan found it under, made absolute and with symlinks resolved. Restore does not read it.
|
||||
|
||||
#### Chunk (`database.Chunk`)
|
||||
A content-addressed unit of data. Files are split into variable-size chunks using the FastCDC algorithm:
|
||||
- `ChunkHash`: SHA256 hash of chunk content (primary key)
|
||||
@@ -82,9 +84,11 @@ The final storage unit uploaded to S3. Contains many compressed and encrypted ch
|
||||
Blob creation process:
|
||||
1. Chunks are accumulated (up to MaxBlobSize, typically 10GB)
|
||||
2. As each chunk is added, its uncompressed bytes are fed to a running SHA-256
|
||||
3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and streamed to storage
|
||||
3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and written to a temporary file
|
||||
4. On finalize, the blob's name is the double SHA-256 of the uncompressed contents — `hex(SHA256(SHA256(...)))` — not a hash of the compressed, encrypted bytes
|
||||
5. Uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}`
|
||||
5. The finished file is uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}` and then deleted
|
||||
|
||||
A backup needs free temporary space, because each blob is written whole to a temporary file before it is uploaded (up to about `blob_size_limit`; an rclone destination that cannot stream uploads needs about twice that) and the metadata export writes copies of the local index. Temporary files go to `$TMPDIR` (default `/tmp`); with `TMPDIR` unset, SQLite writes one of those copies to `/var/tmp`.
|
||||
|
||||
#### BlobChunk (`database.BlobChunk`)
|
||||
Maps chunks to their position within blobs:
|
||||
|
||||
@@ -546,7 +546,7 @@ complete annotated example also lives in
|
||||
| `s3.*` | | Legacy S3 configuration (endpoint, bucket, credentials) |
|
||||
| `exclude` | | Global exclude patterns (applied to all snapshots) |
|
||||
| `chunk_size` | `10MB` | Average chunk size for content-defined chunking |
|
||||
| `blob_size_limit` | `10GB` | Maximum blob size before splitting. Must be at least four times `chunk_size` (the largest chunk the chunker can emit), otherwise a single-chunk blob could exceed the limit |
|
||||
| `blob_size_limit` | `10GB` | Maximum blob size before splitting. Must be at least four times `chunk_size` (the largest chunk the chunker can emit), otherwise a single-chunk blob could exceed the limit. A backup needs free temporary space, because each blob is written whole to a temporary file before it is uploaded (up to about `blob_size_limit`; an rclone destination that cannot stream uploads needs about twice that) and the metadata export writes copies of the local index. Temporary files go to `$TMPDIR` (default `/tmp`); with `TMPDIR` unset, SQLite writes one of those copies to `/var/tmp` |
|
||||
| `compression_level` | `3` | zstd compression level (1-19) |
|
||||
| `hostname` | system hostname | Hostname used in snapshot IDs |
|
||||
| `index_path` | platform data dir | Local SQLite index path |
|
||||
@@ -805,7 +805,8 @@ them. We provide:
|
||||
in the `Dockerfile` together.
|
||||
* `script/release` — cross-compile and publish the release artifacts
|
||||
with the pinned `goreleaser`. Refuses a `goreleaser` on `PATH` whose
|
||||
version is not the pinned one, on the same reasoning as `script/lint`.
|
||||
version is not the pinned one, because a different version would build
|
||||
a different release from the same tag.
|
||||
* `script/release-snapshot` — the same build with no publishing and no
|
||||
tagging, into `./dist`
|
||||
* `script/test` — run the test suite by building the `test` phase of
|
||||
@@ -915,14 +916,14 @@ It is passed to `goreleaser` as `GITEA_TOKEN`. The runner's automatic
|
||||
token is deliberately not used: it is not guaranteed to carry release
|
||||
write access.
|
||||
|
||||
The Go toolchain that compiles the released binaries comes from an
|
||||
`actions/setup-go` step pinned by commit sha, reading its version from
|
||||
`go.mod` (currently `1.26.1`, the same version the `Dockerfile` builder
|
||||
stage pins by digest). `goreleaser` shells out to `go` for every
|
||||
The Go toolchain that compiles the released binaries is installed by
|
||||
`script/install-go`, which downloads the version named by `go.mod`
|
||||
(currently `1.26.1`, the same version the `Dockerfile` builder stage
|
||||
pins by digest) and refuses the archive unless its sha256 matches the
|
||||
value committed in the script. `goreleaser` shells out to `go` for every
|
||||
cross-compile, so without that step the release would either fail
|
||||
outright or ship binaries built by whatever unpinned toolchain the
|
||||
runner happened to carry — the one unpinned thing in an otherwise
|
||||
hash-pinned release path.
|
||||
runner happened to carry.
|
||||
|
||||
To rehearse the whole build without publishing or tagging anything:
|
||||
|
||||
|
||||
@@ -22,6 +22,22 @@ the tag exists and is exercised; what is left is merging `next` to
|
||||
|
||||
# Completed Steps
|
||||
|
||||
- 2026-10-07: Corrected documentation, help text and comments that were
|
||||
false about the code
|
||||
([issue #233](https://git.eeqj.de/sneak/vaultik/issues/233)). A blob
|
||||
is not streamed to storage. The README, `ARCHITECTURE.md` and
|
||||
`config.example.yml` now say a backup needs free temporary space,
|
||||
because each blob is written whole to a temporary file before it is
|
||||
uploaded (up to about `blob_size_limit`; an rclone destination that
|
||||
cannot stream uploads needs about twice that) and the metadata export
|
||||
writes copies of the local index. Temporary files go to `$TMPDIR`
|
||||
(default `/tmp`); with `TMPDIR` unset, SQLite writes one of those
|
||||
copies to `/var/tmp`. Also corrected: the snapshot ID format, what restore
|
||||
reads and how incomplete snapshots are removed in `docs/DATAMODEL.md`,
|
||||
what `source_path` holds, the `index_path` and config file defaults,
|
||||
what `snapshot remove` cleans up, how the release gets its Go
|
||||
toolchain, and the `script/release` and `script/fmt-check` comments.
|
||||
|
||||
- 2026-10-07: Cut the time the `internal/vaultik` and `internal/database`
|
||||
tests take ([issue #235](https://git.eeqj.de/sneak/vaultik/issues/235)).
|
||||
Most of the `internal/vaultik` time went to 24 tests that ran one at a
|
||||
|
||||
+10
-2
@@ -295,8 +295,10 @@ storage_url: "rclone://myremote/path/to/backups"
|
||||
|
||||
# Path to local SQLite index database
|
||||
# This database tracks file state for incremental backups
|
||||
# Default: /var/lib/vaultik/index.sqlite
|
||||
#index_path: /var/lib/vaultik/index.sqlite
|
||||
# Default: the platform data directory, e.g.
|
||||
# macOS: ~/Library/Application Support/vaultik/index.sqlite
|
||||
# Linux: ~/.local/share/vaultik/index.sqlite
|
||||
#index_path: /path/to/index.sqlite
|
||||
|
||||
# Average chunk size for content-defined chunking
|
||||
# Smaller chunks = better deduplication but more metadata
|
||||
@@ -311,6 +313,12 @@ storage_url: "rclone://myremote/path/to/backups"
|
||||
# Chunking uses no secret (the FastCDC parameters are fixed and public). At a
|
||||
# large limit a blob holds hundreds of chunks, so individual chunk lengths are
|
||||
# not visible in its size; lowering the limit toward chunk_size exposes them.
|
||||
# A backup needs free temporary space, because each blob is written whole to
|
||||
# a temporary file before it is uploaded (up to about blob_size_limit; an
|
||||
# rclone destination that cannot stream uploads needs about twice that) and
|
||||
# the metadata export writes copies of the local index. Temporary files go to
|
||||
# $TMPDIR (default /tmp); with TMPDIR unset, SQLite writes one of those copies
|
||||
# to /var/tmp.
|
||||
# Supports: 1GB, 10G, 500MB, 1GiB, etc.
|
||||
# Default: 10GB
|
||||
#blob_size_limit: 10GB
|
||||
|
||||
+5
-6
@@ -111,7 +111,7 @@ Maps chunks to the blobs that contain them.
|
||||
Tracks backup snapshots.
|
||||
|
||||
**Columns:**
|
||||
- `id` (TEXT PRIMARY KEY) - Snapshot ID (format: hostname-YYYYMMDD-HHMMSSZ)
|
||||
- `id` (TEXT PRIMARY KEY) - Snapshot ID (format: `hostname_name_timestamp`, e.g. `server1_home_2025-06-01T12:00:00Z`: the hostname up to its first `.`, the snapshot name, and an RFC 3339 UTC timestamp)
|
||||
- `hostname` (TEXT) - Hostname where backup was created
|
||||
- `vaultik_version` (TEXT) - Version of Vaultik used
|
||||
- `vaultik_git_revision` (TEXT) - Git revision of Vaultik used
|
||||
@@ -218,8 +218,8 @@ The `{remote-key}` directory name is a one-way hash of the human snapshot ID, so
|
||||
### 4. Restore Process
|
||||
|
||||
The restore process doesn't use the local database. Instead:
|
||||
1. Downloads snapshot metadata from S3
|
||||
2. Downloads required blobs based on manifest
|
||||
1. Downloads and decrypts the snapshot's metadata database (`db.zst.age`) from S3
|
||||
2. Downloads the blobs holding the chunks of the files being restored, found through that database's `blob_chunks` table; the manifest is not read
|
||||
3. Reconstructs files from decrypted and decompressed chunks
|
||||
|
||||
### 5. Pruning
|
||||
@@ -232,9 +232,8 @@ The restore process doesn't use the local database. Instead:
|
||||
|
||||
Before each backup:
|
||||
1. Query incomplete snapshots (where `completed_at IS NULL`)
|
||||
2. Check if metadata exists in S3
|
||||
3. If no metadata, delete snapshot and all associations
|
||||
4. Clean up orphaned files, chunks, and blobs
|
||||
2. Delete each one and all its associations, without checking S3 for its metadata
|
||||
3. Clean up orphaned files, chunks, and blobs
|
||||
|
||||
## Repository Pattern
|
||||
|
||||
|
||||
@@ -22,9 +22,11 @@ scans every snapshot manifest in the destination store, builds the
|
||||
set of still-referenced blob hashes, and deletes any blob not in that
|
||||
set.
|
||||
|
||||
Snapshot create --prune and snapshot remove run the same cleanup
|
||||
automatically; this command is the manual entry point for the same
|
||||
work (e.g. after a crashed backup or to reclaim storage).`,
|
||||
Snapshot create --prune runs the same cleanup automatically; this
|
||||
command is the manual entry point for the same work (e.g. after a
|
||||
crashed backup or to reclaim storage). Snapshot remove leaves blobs in
|
||||
place; run this command afterwards to delete the ones no longer
|
||||
referenced.`,
|
||||
Args: cobra.NoArgs,
|
||||
RunE: func(cmd *cobra.Command, _ []string) error {
|
||||
// Use unified config resolution
|
||||
|
||||
@@ -66,8 +66,9 @@ func newSnapshotCreateCommand() *cobra.Command {
|
||||
If snapshot names are provided, only those snapshots are created.
|
||||
If no names are provided, all configured snapshots are created.
|
||||
|
||||
Config is located at /etc/vaultik/config.yml by default, but can be overridden by
|
||||
specifying a path using --config or by setting VAULTIK_CONFIG to a path.`,
|
||||
The config is read from the path given by --config or VAULTIK_CONFIG;
|
||||
otherwise from the platform config directory (~/.config/vaultik/config.yml
|
||||
on Linux), then /etc/vaultik/config.yml.`,
|
||||
Args: cobra.ArbitraryArgs,
|
||||
RunE: func(cmd *cobra.Command, args []string) error {
|
||||
// Pass snapshot names from args
|
||||
|
||||
@@ -14,8 +14,7 @@ type File struct {
|
||||
ID types.FileID // UUID primary key
|
||||
Path types.FilePath // Absolute path of the file
|
||||
|
||||
// SourcePath is the source directory this file came from (used for
|
||||
// restore path stripping).
|
||||
// SourcePath is the source directory this file came from.
|
||||
SourcePath types.SourcePath
|
||||
MTime time.Time
|
||||
Size int64
|
||||
|
||||
@@ -5,7 +5,7 @@
|
||||
CREATE TABLE IF NOT EXISTS files (
|
||||
id TEXT PRIMARY KEY, -- UUID
|
||||
path TEXT NOT NULL UNIQUE,
|
||||
source_path TEXT NOT NULL DEFAULT '', -- The source directory this file came from (for restore path stripping)
|
||||
source_path TEXT NOT NULL DEFAULT '', -- The source directory this file came from
|
||||
mtime INTEGER NOT NULL, -- whole seconds since the Unix epoch
|
||||
mtime_nsec INTEGER NOT NULL, -- nanoseconds within that second, 0 to 999999999
|
||||
size INTEGER NOT NULL,
|
||||
|
||||
@@ -58,8 +58,8 @@ type Scanner struct {
|
||||
compressionLevel int
|
||||
ageRecipient string
|
||||
snapshotID string // Current snapshot being processed
|
||||
// currentSourcePath is the source directory being scanned (used for
|
||||
// restore path stripping).
|
||||
// currentSourcePath is the source directory being scanned, stored with
|
||||
// each file record.
|
||||
currentSourcePath string
|
||||
exclude []string // Glob patterns for files/directories to exclude
|
||||
compiledExclude []compiledPattern // Compiled glob patterns
|
||||
@@ -211,7 +211,6 @@ func (s *Scanner) Scan(
|
||||
ctx context.Context, path string, snapshotID string,
|
||||
) (*ScanResult, error) {
|
||||
s.snapshotID = snapshotID
|
||||
// Store source path for file records (used during restore)
|
||||
s.currentSourcePath = path
|
||||
result := &ScanResult{
|
||||
StartTime: time.Now().UTC(),
|
||||
@@ -1198,9 +1197,8 @@ func (s *Scanner) checkFileInMemory(
|
||||
}
|
||||
|
||||
file := &database.File{
|
||||
ID: fileID,
|
||||
Path: types.FilePath(path),
|
||||
// Store source directory for restore path stripping
|
||||
ID: fileID,
|
||||
Path: types.FilePath(path),
|
||||
SourcePath: types.SourcePath(s.currentSourcePath),
|
||||
MTime: info.ModTime(),
|
||||
Size: info.Size(),
|
||||
|
||||
@@ -155,8 +155,8 @@ type BlobHash string
|
||||
// FilePath represents an absolute path to a file or directory.
|
||||
type FilePath string
|
||||
|
||||
// SourcePath represents the root directory from which files are backed up.
|
||||
// Used during restore to strip the source prefix from paths.
|
||||
// SourcePath is the source directory a scan found a file under, made
|
||||
// absolute and with symlinks resolved.
|
||||
type SourcePath string
|
||||
|
||||
// Hostname identifies a host machine.
|
||||
|
||||
+5
-2
@@ -1,6 +1,9 @@
|
||||
#!/bin/sh
|
||||
# script/fmt-check: check formatting (read-only). Same scope as
|
||||
# script/fmt, but fails instead of writing.
|
||||
# script/fmt-check: check formatting (read-only). Fails instead of
|
||||
# writing. It checks every Go file outside .tool, which is more than
|
||||
# script/fmt formats: `go fmt ./...` skips `testdata` directories and
|
||||
# files and directories whose names start with `.` or `_`. Fix a file
|
||||
# only this reports with `gofmt -w`.
|
||||
set -eu
|
||||
|
||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||
|
||||
+3
-4
@@ -21,10 +21,9 @@ goreleaser_version() {
|
||||
head -n 1
|
||||
}
|
||||
|
||||
# Resolve the goreleaser to run, on the same rule script/lint uses for
|
||||
# golangci-lint: a binary on PATH is accepted only when it is exactly
|
||||
# the pinned version, because a differently versioned tool would
|
||||
# produce a differently built release from the same tag. Anything else
|
||||
# Resolve the goreleaser to run. A binary on PATH is accepted only when
|
||||
# it is exactly the pinned version, because a differently versioned tool
|
||||
# would produce a differently built release from the same tag. Anything else
|
||||
# comes from .tool/bin, and a missing one is a loud failure naming the
|
||||
# script that installs it rather than a silent fallback.
|
||||
resolve_goreleaser() {
|
||||
|
||||
Reference in New Issue
Block a user