Author SHA1 Message Date
sneak 3543f45ba8 Add storage tests for URL parsing and every in-process backend (closes #66)
check / check (pull_request) Failing after 1s
internal/storage had no tests. This adds table-driven ParseStorageURL
coverage and a shared Storer conformance suite, defined once in
conformance_test.go and run against every backend that can run
in-process: file:// over a temp dir and s3:// over the same in-memory
S3 harness (gofakes3 + s3mem) that internal/s3 and the not-found test
use, so a new backend inherits the contract by passing its constructor.
The rclone backend gets construction and argument-shaping tests via
rclone's in-process ":local" backend, plus the missing-remote error
path; a comment records that its data-plane operations need a configured
remote and so cannot run in a unit test.

Tests only; no production code changed and no defect surfaced.

Model: opus-4-8
2026-09-21 19:55:20 +00:00
25 changed files with 378 additions and 855 deletions
+3 -8
View File
@@ -104,12 +104,7 @@ Version: 2025-06-08
13. Pre-1.0: NEVER write database migrations. There are no live databases 13. Pre-1.0: NEVER write database migrations. There are no live databases
anywhere — every user's local index can be rebuilt from a fresh full anywhere — every user's local index can be rebuilt from a fresh full
backup. To change the schema, edit `internal/database/schema/001.sql` backup. When the schema changes, just change `schema.sql` (and any code
(and any code that touches the affected tables) directly; do not add new that touches the affected tables). The local index is disposable until
numbered schema files. Those numbered files and the `schema_migrations` 1.0 ships and is tagged.
table they populate only bootstrap a fresh database — they are not an
upgrade path. The local index is disposable until 1.0 ships and is
tagged; once 1.0 is tagged that clause expires and the question of
upgrading existing indexes returns. See [`docs/DATAMODEL.md`](docs/DATAMODEL.md)
for the full explanation.
+3 -3
View File
@@ -63,7 +63,7 @@ A content-addressed unit of data. Files are split into variable-size chunks usin
- `ChunkHash`: SHA256 hash of chunk content (primary key) - `ChunkHash`: SHA256 hash of chunk content (primary key)
- `Size`: Chunk size in bytes - `Size`: Chunk size in bytes
Chunk sizes vary between `avgChunkSize/4` and `avgChunkSize*4` (2.5MB-40MB for the 10MB default average). Chunk sizes vary between `avgChunkSize/4` and `avgChunkSize*4` (typically 16KB-256KB for 64KB average).
#### FileChunk (`database.FileChunk`) #### FileChunk (`database.FileChunk`)
Maps files to their constituent chunks: Maps files to their constituent chunks:
@@ -120,7 +120,7 @@ The CLI uses fx for dependency injection. Here's the instantiation order:
```go ```go
// cli/app.go: NewApp() // cli/app.go: NewApp()
fx.New( fx.New(
fx.Supply(config.Path(opts.ConfigPath)), // 1. Config path fx.Supply(config.ConfigPath(opts.ConfigPath)), // 1. Config path
fx.Supply(opts.LogOptions), // 2. Log options fx.Supply(opts.LogOptions), // 2. Log options
fx.Provide(globals.New), // 3. Globals fx.Provide(globals.New), // 3. Globals
fx.Provide(log.New), // 4. Logger config fx.Provide(log.New), // 4. Logger config
@@ -193,7 +193,7 @@ scanner := v.ScannerFactory(snapshot.ScannerParams{
- **Created by**: `chunker.NewChunker(avgChunkSize)` - **Created by**: `chunker.NewChunker(avgChunkSize)`
- **When**: Inside `snapshot.NewScanner()` - **When**: Inside `snapshot.NewScanner()`
- **Configuration**: - **Configuration**:
- `avgChunkSize`: From config (default 10MB) - `avgChunkSize`: From config (typically 64KB)
- `minChunkSize`: avgChunkSize / 4 - `minChunkSize`: avgChunkSize / 4
- `maxChunkSize`: avgChunkSize * 4 - `maxChunkSize`: avgChunkSize * 4
+3 -24
View File
@@ -20,6 +20,8 @@
# golang:1.26.1-alpine, 2026-03-17 # golang:1.26.1-alpine, 2026-03-17
FROM golang:1.26.1-alpine@sha256:2389ebfa5b7f43eeafbd6be0c3700cc46690ef842ad962f6c5bd6be49ed82039 AS builder FROM golang:1.26.1-alpine@sha256:2389ebfa5b7f43eeafbd6be0c3700cc46690ef842ad962f6c5bd6be49ed82039 AS builder
ARG VERSION=dev
# Build tooling: make, plus a C toolchain because `go test -race` needs cgo. # Build tooling: make, plus a C toolchain because `go test -race` needs cgo.
# The sqlite driver is pure Go (modernc.org/sqlite), so no sqlite library or # The sqlite driver is pure Go (modernc.org/sqlite), so no sqlite library or
# CLI is required. # CLI is required.
@@ -64,31 +66,8 @@ RUN [ -n "$CHECK_EPOCH" ] || exit 1
RUN echo "check epoch: ${CHECK_EPOCH}" && make fmt-check RUN echo "check epoch: ${CHECK_EPOCH}" && make fmt-check
RUN echo "check epoch: ${CHECK_EPOCH}" && make test RUN echo "check epoch: ${CHECK_EPOCH}" && make test
# Version, commit and build date are computed on the host by
# script/docker and script/cibuild (where .git exists) and passed in as
# build args. The build context excludes .git (see .dockerignore), so
# the build cannot derive them itself: it used to try, with `git
# rev-parse` inside this stage, and always got "unknown". VERSION comes
# from script/version, the source of truth shared with the Makefile, so
# it carries the same tag / dev-<sha> / -dirty rules and a Docker image
# reports the same string a local build of the same tree would.
#
# The defaults are the fallback for a bare `docker build .` that passes
# none of them: an unset arg would otherwise stamp an empty string and
# produce an image that cannot report its own version, commit or date.
# They match what an out-of-git build reports elsewhere.
#
# These ARGs sit here, after the checks, rather than at the top of the
# stage: every commit changes their values, and a value change
# invalidates all layers below the ARG. Declared up top they would bust
# `go mod download`; here they only rekey this build layer, which the
# COPY of the sources above already rebuilds on any change anyway.
ARG VERSION=dev
ARG COMMIT=unknown
ARG COMMIT_DATE=unknown
# Build (pure Go, no CGO required since we use modernc.org/sqlite) # Build (pure Go, no CGO required since we use modernc.org/sqlite)
RUN CGO_ENABLED=0 go build -ldflags "-X 'sneak.berlin/go/vaultik/internal/globals.Version=${VERSION}' -X 'sneak.berlin/go/vaultik/internal/globals.Commit=${COMMIT}' -X 'sneak.berlin/go/vaultik/internal/globals.CommitDate=${COMMIT_DATE}'" -o /vaultik ./cmd/vaultik RUN CGO_ENABLED=0 go build -ldflags "-X 'sneak.berlin/go/vaultik/internal/globals.Version=${VERSION}' -X 'sneak.berlin/go/vaultik/internal/globals.Commit=$(git rev-parse HEAD 2>/dev/null || echo unknown)' -X 'sneak.berlin/go/vaultik/internal/globals.CommitDate=$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)'" -o /vaultik ./cmd/vaultik
# Runtime stage # Runtime stage
# alpine:3.21, 2026-02-25 # alpine:3.21, 2026-02-25
+19 -112
View File
@@ -84,57 +84,6 @@ VAULTIK_AGE_SECRET_KEY='AGE-SECRET-KEY-...' vaultik snapshot restore <snapshot-i
# 0 3 * * * vaultik snapshot create --cron --prune --keep-newer-than 4w # 0 3 * * * vaultik snapshot create --cron --prune --keep-newer-than 4w
``` ```
## restoring on another machine
Restoring on a host that never ran the backup — a replacement machine
after the original is gone — is the case vaultik is built for. That host
needs only three things: the `vaultik` binary, the age **private** key,
and the storage credentials for the destination. It does **not** need the
local index, the original config file, or the original hostname.
```sh
# install
go install sneak.berlin/go/vaultik/cmd/vaultik@latest
# create a config and point it at the ORIGINAL backup destination
vaultik config init
vaultik config set storage_url "s3://bucket/prefix?endpoint=https://s3.example.com"
vaultik config set s3.access_key_id "..."
vaultik config set s3.secret_access_key "..."
# see what is on the destination store
vaultik snapshot list
```
`snapshot list` reads the destination store without the private key. A
snapshot that is not in this host's (empty) local index is shown as
remote-only: its row is identified by `<remote only:...>` rather than by
a `hostname_name_timestamp` name, because the name lives only in the
local index and the encrypted database and cannot be recovered from the
store. Its timestamp and compressed size are real. (See the `snapshot
list` description under [command details](#command-details) for the full
explanation.)
Use that remote key — the hex printed inside `<remote only:...>`, or the
full `remote_key` from `snapshot list --json` — to restore and verify:
```sh
# restore everything to /tmp/restored, then check every restored file's
# chunk hashes
VAULTIK_AGE_SECRET_KEY='AGE-SECRET-KEY-...' \
vaultik snapshot restore --verify <remote-key> /tmp/restored
# optionally, deep-verify the snapshot against the store (downloads and
# cryptographically checks every blob)
VAULTIK_AGE_SECRET_KEY='AGE-SECRET-KEY-...' \
vaultik snapshot verify --deep <remote-key>
```
`age_recipients` (the public key) is not needed to restore — only the
private key in `VAULTIK_AGE_SECRET_KEY`. Both the abbreviated key printed
in the table and the full 64-character key from `--json` are accepted; a
leading part of the key is enough as long as it is unambiguous.
--- ---
## cli ## cli
@@ -147,10 +96,10 @@ vaultik [--config <path>] config edit
vaultik [--config <path>] config get <key> vaultik [--config <path>] config get <key>
vaultik [--config <path>] config set <key> <value> vaultik [--config <path>] config set <key> <value>
vaultik [--config <path>] snapshot create [snapshot-names...] [--cron] [--prune] [--keep-newer-than <duration>] vaultik [--config <path>] snapshot create [snapshot-names...] [--cron] [--prune] [--keep-newer-than <duration>]
vaultik [--config <path>] snapshot list [--json] # alias: ls vaultik [--config <path>] snapshot list [--json]
vaultik [--config <path>] snapshot verify <snapshot-id> [--deep] [--json] vaultik [--config <path>] snapshot verify <snapshot-id> [--deep] [--json]
vaultik [--config <path>] snapshot purge [--keep-latest | --older-than <duration>] [--snapshot <name>...] [--force] vaultik [--config <path>] snapshot purge [--keep-latest | --older-than <duration>] [--snapshot <name>...] [--force]
vaultik [--config <path>] snapshot remove <snapshot-id> [--dry-run] [--force] [--local-only] [--json] # alias: rm vaultik [--config <path>] snapshot remove <snapshot-id> [--dry-run] [--force] [--local-only] [--json]
vaultik [--config <path>] snapshot restore <snapshot-id> <target-dir> [paths...] [--verify] vaultik [--config <path>] snapshot restore <snapshot-id> <target-dir> [paths...] [--verify]
vaultik [--config <path>] prune [--force] [--json] vaultik [--config <path>] prune [--force] [--json]
vaultik [--config <path>] info vaultik [--config <path>] info
@@ -169,21 +118,6 @@ vaultik version
* `--quiet`, `-q`: Suppress non-error output (also suppresses startup banner) * `--quiet`, `-q`: Suppress non-error output (also suppresses startup banner)
* `--skip-errors`: Continue past per-file errors instead of aborting (applies to `snapshot create` and `restore`) * `--skip-errors`: Continue past per-file errors instead of aborting (applies to `snapshot create` and `restore`)
### locking
Every command that opens the local index — `snapshot create`, `snapshot
list`, `snapshot verify`, `snapshot purge`, `snapshot remove`, `snapshot
restore`, `prune`, `info`, and `remote info`/`remote nuke` — takes a
process-wide lock at `$XDG_DATA_HOME/vaultik/vaultik.pid`
(`~/.local/share/vaultik/vaultik.pid` on Linux) for the whole run. Only
one such command runs at a time: a second one exits immediately with an
"already running" error rather than waiting. The lock is not scoped to
mutating commands, so read-only commands are affected too — `vaultik
snapshot list` fails while a backup is in progress; scoping it so
read-only commands run during a backup is tracked in
[issue #150](https://git.eeqj.de/sneak/vaultik/issues/150). `config`,
`database delete`, `completion`, and `version` do not take the lock.
### stdout and stderr ### stdout and stderr
Log output — everything from `--verbose` and `--debug`, and every Log output — everything from `--verbose` and `--debug`, and every
@@ -218,8 +152,6 @@ and `vaultik prune --json | jq .` both work as written.
* `VAULTIK_AGE_SECRET_KEY`: Age private key for decryption (required for `snapshot restore` and `snapshot verify --deep`) * `VAULTIK_AGE_SECRET_KEY`: Age private key for decryption (required for `snapshot restore` and `snapshot verify --deep`)
* `VAULTIK_CONFIG`: Path to config file (overridden by `--config`) * `VAULTIK_CONFIG`: Path to config file (overridden by `--config`)
* `VAULTIK_INDEX_PATH`: Override local SQLite index path * `VAULTIK_INDEX_PATH`: Override local SQLite index path
* `VAULTIK_CPUPROFILE`: Write a CPU profile to this path for the duration of the run (development/debugging)
* `VAULTIK_MEMPROFILE`: Write a heap profile to this path when the run exits (development/debugging)
### shell completion ### shell completion
@@ -313,8 +245,6 @@ local index alone, and still exits zero.
* Default (shallow): checks that all blobs referenced in the manifest exist in storage * Default (shallow): checks that all blobs referenced in the manifest exist in storage
* `--deep`: Downloads and decrypts each blob, verifies chunk hashes against the * `--deep`: Downloads and decrypts each blob, verifies chunk hashes against the
encrypted metadata database encrypted metadata database
* Accepts the same identifiers as `snapshot restore`: a snapshot ID, or a
remote-only snapshot's remote key (or an unambiguous leading part of it)
* `--json`: Output results as JSON * `--json`: Output results as JSON
**`snapshot purge`**: Remove old snapshots based on criteria. Retention is **`snapshot purge`**: Remove old snapshots based on criteria. Retention is
@@ -345,10 +275,6 @@ on the destination in one go, use `vaultik remote nuke --force`.
**`snapshot restore`**: Restore files from a backup snapshot. **`snapshot restore`**: Restore files from a backup snapshot.
* Requires `VAULTIK_AGE_SECRET_KEY` environment variable * Requires `VAULTIK_AGE_SECRET_KEY` environment variable
* Accepts a snapshot ID, or — for a snapshot only on the destination
store — its remote key (or an unambiguous leading part of it) as shown
by `snapshot list`. See
[restoring on another machine](#restoring-on-another-machine).
* Optional path arguments to restore specific files/directories (default: all) * Optional path arguments to restore specific files/directories (default: all)
* Preserves file permissions, timestamps, ownership (ownership requires root), * Preserves file permissions, timestamps, ownership (ownership requires root),
symlinks, and empty directories symlinks, and empty directories
@@ -412,10 +338,6 @@ both are set.
## architecture ## architecture
For an implementation-level view of the internals — the data model, the
`fx` dependency-injection wiring, and the scanner — see
[`ARCHITECTURE.md`](ARCHITECTURE.md).
### remote storage layout ### remote storage layout
``` ```
@@ -493,24 +415,19 @@ derivation.
### compression ### compression
* zstd compression at configurable level (1-19, default 3). The level is * zstd compression at configurable level (1-19, default 3)
accepted as 1-19 but maps onto zstd's four internal speed presets:
1-2 fastest, 3-5 default, 6-9 better, 10-19 best. Levels within the
same band compress identically.
* Applied before encryption at the blob level * Applied before encryption at the blob level
--- ---
## configuration reference ## configuration reference
Run `vaultik config init` to generate a fully commented config file; a Run `vaultik config init` to generate a fully commented config file.
complete annotated example also lives in Key fields:
[`config.example.yml`](config.example.yml). Key fields:
| Field | Default | Description | | Field | Default | Description |
|-------|---------|-------------| |-------|---------|-------------|
| `age_recipients` | (required) | Age public keys for encryption | | `age_recipients` | (required) | Age public keys for encryption |
| `age_secret_key` | (unset) | Age private key for decryption (`snapshot restore`, `snapshot verify --deep`). Setting it in the config file places the private key on the backed-up host, defeating the public-key-only design (see "why" above). Prefer the `VAULTIK_AGE_SECRET_KEY` environment variable, supplied only on the machine you restore from. |
| `snapshots` | (required) | Named snapshot definitions with paths and excludes | | `snapshots` | (required) | Named snapshot definitions with paths and excludes |
| `storage_url` | | Storage backend URL (`s3://`, `file://`, `rclone://`) | | `storage_url` | | Storage backend URL (`s3://`, `file://`, `rclone://`) |
| `s3.*` | | Legacy S3 configuration (endpoint, bucket, credentials) | | `s3.*` | | Legacy S3 configuration (endpoint, bucket, credentials) |
@@ -540,13 +457,9 @@ complete annotated example also lives in
sequentially. Restore speed is bound by single-stream throughput. sequentially. Restore speed is bound by single-stream throughput.
* **Device nodes, named pipes, and sockets are silently skipped.** Only * **Device nodes, named pipes, and sockets are silently skipped.** Only
regular files, directories, and symlinks are backed up. regular files, directories, and symlinks are backed up.
* **No upgrade path between versions.** There is no supported way to carry * **No database migrations.** If the local SQLite schema changes between
an existing local index across a schema change; if the local SQLite versions, delete the local database (`vaultik database delete`) and run
schema changes between versions, delete the local database (`vaultik a full backup. Remote storage is unaffected.
database delete`) and run a full backup. Remote storage is unaffected.
(The binary does embed numbered schema files and a `schema_migrations`
table to bootstrap a fresh database — see [`docs/DATAMODEL.md`](docs/DATAMODEL.md)
— but that is not an upgrade path.)
* **Files that change during backup may be inconsistent.** There is no * **Files that change during backup may be inconsistent.** There is no
filesystem snapshot or freeze. If a file is modified between the scan filesystem snapshot or freeze. If a file is modified between the scan
and chunk phases, the backed-up copy may reflect a partial write. and chunk phases, the backed-up copy may reflect a partial write.
@@ -612,12 +525,14 @@ priority.
### infrastructure ### infrastructure
* **Cross-version schema upgrades.** There is no upgrade path between * **Cross-machine restore documentation.** The "restore from
released versions — pre-1.0 schema changes are handled by `vaultik another host" workflow works but isn't documented as a
database delete` plus a full re-scan (see first-class operation in this README. Worth a dedicated section
[`docs/DATAMODEL.md`](docs/DATAMODEL.md)). Post-1.0 we'll need a once it's settled.
migration story to keep existing index databases usable across * **Schema migrations.** Currently nonexistent — pre-1.0 schema
upgrades. changes are handled by `vaultik database delete` plus a full
re-scan. Post-1.0 we'll need a migration story to keep existing
index databases usable across upgrades.
* **Storage backend coverage tests.** S3, file://, and rclone:// * **Storage backend coverage tests.** S3, file://, and rclone://
all share the Storer interface but the rclone path is the least all share the Storer interface but the rclone path is the least
exercised in CI. exercised in CI.
@@ -626,17 +541,9 @@ priority.
## output style ## output style
The operational narration of the long-running commands — the Begin, All user-facing output goes through helpers in `internal/ui` and conforms
Complete, Progress, and status lines of `snapshot create`, `prune`, to a uniform style. Color is enabled when stdout is a TTY and the
`snapshot restore`, and the like — goes through helpers in `internal/ui` `NO_COLOR` environment variable is unset (https://no-color.org/).
and conforms to the uniform style below. Some commands instead write
plain text straight to stdout (`version`, `info`, `config`, the
`database delete` prompt, and the `snapshot list` table); that output is
unstyled and does not honor `--quiet`. Routing it through `internal/ui`
is tracked in
[issue #149](https://git.eeqj.de/sneak/vaultik/issues/149). Color is
enabled when stdout is a TTY and the `NO_COLOR` environment variable is
unset (https://no-color.org/).
`internal/ui` writes to stdout; it is the output the user asked for. `internal/ui` writes to stdout; it is the output the user asked for.
Structured log records are a different thing and go through Structured log records are a different thing and go through
-102
View File
@@ -1,102 +0,0 @@
package main_test
import (
"strings"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// This file guards the version stamping of the product image (issue
// #75). The failure it protects against is silent: the image still
// builds and runs, but `vaultik version` inside it reports "commit:
// unknown", so an operator cannot tell which source produced a given
// backup. .dockerignore excludes .git, so the build cannot derive the
// commit itself; the values must be computed on the host and passed in.
//
// These are parses of the committed files, for the same reason the lint
// guards next door are: shelling out to docker would nest a build
// inside `make test`. That `vaultik version` in the built image really
// prints the host's version is verified by hand and recorded on the
// pull request.
// dockerScript is script/docker, relative to the repository root.
const dockerScript = "script/docker"
// versionArgs are the ldflag targets the build stamps and, matching
// them, the build args the host must supply. The names line up so the
// same list checks both files.
func versionArgs() []string {
return []string{"VERSION", "COMMIT", "COMMIT_DATE"}
}
// TestProductDockerfileTakesVersionAsBuildArgs fails unless the build
// declares each version arg and stamps it into the binary by ldflag
// reference, rather than computing it in the container.
func TestProductDockerfileTakesVersionAsBuildArgs(t *testing.T) {
t.Parallel()
found := instructions(t, productDockerfile)
for _, arg := range versionArgs() {
require.GreaterOrEqual(t, indexOf(found, "ARG "+arg), 0,
"%s must declare `ARG %s` so the host can pass it in",
productDockerfile, arg)
assertLdflagReferences(t, found, arg)
}
}
// TestProductDockerfileDoesNotDeriveVersionItself is the anti-regression
// for the original defect: the container ran `git rev-parse`, but .git
// is not in the build context, so it always resolved to "unknown". No
// git command may reach into a build that cannot see the history.
func TestProductDockerfileDoesNotDeriveVersionItself(t *testing.T) {
t.Parallel()
text := instructionText(readRepoFile(t, productDockerfile))
assert.NotContains(t, text, "git ",
"%s must not run git: .git is excluded from the build context, so"+
" any value it derives is wrong. Pass version, commit and date"+
" in as build args instead.", productDockerfile)
}
// TestDockerScriptComputesVersionOnTheHost fails unless script/docker
// derives each value where .git exists and passes it as a build arg,
// with VERSION coming from script/version so a Docker build reports the
// same string a local build of the same tree would.
func TestDockerScriptComputesVersionOnTheHost(t *testing.T) {
t.Parallel()
script := readRepoFile(t, dockerScript)
for _, arg := range versionArgs() {
assert.Contains(t, script, "--build-arg "+arg+"=",
"%s must pass --build-arg %s to the build", dockerScript, arg)
}
assert.Contains(t, script, "/version",
"%s must take VERSION from script/version, the source of truth"+
" shared with the Makefile", dockerScript)
}
// assertLdflagReferences fails unless some build instruction stamps the
// named variable from the ARG (a ${arg} reference), not from a value
// computed inside the container.
func assertLdflagReferences(t *testing.T, found []string, arg string) {
t.Helper()
for _, instruction := range found {
if strings.HasPrefix(instruction, "RUN ") &&
strings.Contains(instruction, "go build") &&
strings.Contains(instruction, "${"+arg+"}") {
return
}
}
assert.Fail(t, "version arg is declared but never stamped",
"the go build in %s must reference ${%s} in its ldflags, or the"+
" arg is passed and discarded", productDockerfile, arg)
}
+2 -6
View File
@@ -304,14 +304,10 @@ func instructionText(contents string) string {
} }
// indexOf returns the position of the first instruction equal to, or // indexOf returns the position of the first instruction equal to, or
// beginning with, want; -1 if there is none. An `ARG NAME=default` // beginning with, want; -1 if there is none.
// counts as beginning with `ARG NAME`, so a declared arg is found
// whether or not it carries a default.
func indexOf(found []string, want string) int { func indexOf(found []string, want string) int {
for i, instruction := range found { for i, instruction := range found {
if instruction == want || if instruction == want || strings.HasPrefix(instruction, want+" ") {
strings.HasPrefix(instruction, want+" ") ||
strings.HasPrefix(instruction, want+"=") {
return i return i
} }
} }
+1 -11
View File
@@ -10,16 +10,6 @@ import (
) )
func main() { func main() {
os.Exit(run())
}
// run sets up optional profiling, runs the CLI, and returns the process
// exit code. os.Exit lives in main so it fires only after run's deferred
// profile writers have flushed. cli.Entry returns a status code rather
// than calling os.Exit itself: an os.Exit from inside it would skip
// these defers and truncate the profile of a failing command -- exactly
// the command one most often wants to profile.
func run() int {
// CPU profiling: set VAULTIK_CPUPROFILE=/path/to/cpu.prof // CPU profiling: set VAULTIK_CPUPROFILE=/path/to/cpu.prof
if cpuProfile := os.Getenv("VAULTIK_CPUPROFILE"); cpuProfile != "" { if cpuProfile := os.Getenv("VAULTIK_CPUPROFILE"); cpuProfile != "" {
f, err := os.Create(cpuProfile) //nolint:gosec // G304: operator-set path f, err := os.Create(cpuProfile) //nolint:gosec // G304: operator-set path
@@ -56,5 +46,5 @@ func run() int {
}() }()
} }
return cli.Entry() cli.Entry()
} }
+5 -24
View File
@@ -5,30 +5,11 @@
Vaultik uses a local SQLite database to track file metadata, chunk mappings, and blob associations during the backup process. This database serves as an index for incremental backups and enables efficient deduplication. Vaultik uses a local SQLite database to track file metadata, chunk mappings, and blob associations during the backup process. This database serves as an index for incremental backups and enables efficient deduplication.
**Important Notes:** **Important Notes:**
- **No Migration Support (pre-1.0)**: Vaultik does not support database schema
This section is the authoritative explanation of the schema/migration story; migrations. The local index is treated as disposable — if the schema changes,
other documents (the README and `AGENTS.md`) link here. delete the local SQLite database (`vaultik database delete`) and run a full
backup. The remote storage is unaffected; the new index will re-deduplicate
- **No upgrade path between versions (pre-1.0)**: Vaultik has no supported way to against existing remote blobs.
carry an existing local index across a schema change. The index is disposable
— if the on-disk schema changes between versions, delete the local SQLite
database (`vaultik database delete`) and run a full backup. Remote storage is
unaffected; the new index re-deduplicates against existing remote blobs. This
is the standing project policy, and it is separate from the schema bootstrap
described next.
- **Schema bootstrap**: a fresh database is populated from numbered SQL files
embedded in the binary under `internal/database/schema/`. `000.sql` creates the
`schema_migrations` table; `001.sql` creates the application tables. On opening
a database the code applies each numbered file that has not yet run and records
its version in `schema_migrations`. This bootstraps a new database; it does not
upgrade an existing one between released versions.
- **Changing the schema (pre-1.0)**: edit `internal/database/schema/001.sql` (and
the code that touches the affected tables) directly. Do not add new numbered
files — there is no installed base to migrate.
- **Disposability expires at 1.0**: the index is treated as disposable only until
1.0 ships and is tagged. Once 1.0 is tagged that clause expires and the
question of upgrading existing indexes returns. It is deliberately left open
here.
- **Version Compatibility**: In rare cases, you may need to use the same version - **Version Compatibility**: In rare cases, you may need to use the same version
of Vaultik to restore a backup as was used to create it. This ensures of Vaultik to restore a backup as was used to create it. This ensures
compatibility with the metadata format stored in S3. compatibility with the metadata format stored in S3.
+38 -89
View File
@@ -11,7 +11,6 @@ import (
"os/signal" "os/signal"
"path/filepath" "path/filepath"
"strings" "strings"
"sync"
"syscall" "syscall"
"time" "time"
@@ -197,90 +196,13 @@ func RunApp(ctx context.Context, app *fx.App) error {
} }
} }
// errReported marks a failure the operation has already shown the user
// (and deliberately withheld under --json). Entry turns it into a
// non-zero exit status without printing anything further, so the error
// line is not doubled. It flows up from RunOperation through cobra to
// Entry.
var errReported = errors.New("operation failed")
// RunOperation runs op against the Vaultik instance inside the fx app
// and turns a failure into a returned error rather than an os.Exit from
// within the goroutine. An os.Exit there skipped main's deferred
// profile writers -- so profiling a failing command yielded a truncated
// profile (issue #75) -- and RunWithApp's PID-lock release, and denied
// the app any graceful shutdown; returning the error to the top runs
// all three.
//
// op runs in a goroutine so OnStart returns promptly and an interrupt
// can still cancel through OnStop; when it finishes, success or failure,
// it triggers shutdown, which is what lets RunWithApp return. report is
// called with a non-canceled failure so the caller can log it (and
// suppress it under --json) before it becomes errReported. A context
// cancellation is the interrupt path, not a failure: it is neither
// reported nor counted as one.
func RunOperation(
ctx context.Context, opts AppOptions,
op func(v *vaultik.Vaultik) error, report func(err error),
) error {
var (
mu sync.Mutex
failed bool
)
opts.Invokes = append(opts.Invokes,
fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
lc.Append(fx.Hook{
OnStart: func(_ context.Context) error {
go func() {
err := op(v)
if err != nil && !errors.Is(err, context.Canceled) {
report(err)
mu.Lock()
failed = true
mu.Unlock()
}
stopErr := v.Shutdowner.Shutdown()
if stopErr != nil {
log.Error("Failed to shutdown", "error", stopErr)
}
}()
return nil
},
OnStop: func(_ context.Context) error {
v.Cancel()
return nil
},
})
}))
err := RunWithApp(ctx, opts)
if err != nil {
return err
}
// The goroutine sets failed before triggering the shutdown that lets
// RunWithApp return, so the write is in place by the time we read it.
mu.Lock()
defer mu.Unlock()
if failed {
return errReported
}
return nil
}
// runVaultikApp runs the standard single-operation command lifecycle // runVaultikApp runs the standard single-operation command lifecycle
// shared by the list/purge/verify/remove/remote-info subcommands: // shared by the list/purge/verify/remove/remote-info subcommands:
// resolve the config, then run op against the Vaultik instance through // resolve the config, start the fx app, run op against the Vaultik
// RunOperation, reporting a failure prefixed with failMsg (suppressed // instance in a goroutine, report a failure prefixed with failMsg
// while suppressErrors is true, e.g. under --json). extraQuiet is OR-ed // (suppressed while suppressErrors is true, e.g. under --json), then
// into LogOptions.Quiet (e.g. --json output modes). // trigger shutdown. The operation is cancelled when the app stops.
// extraQuiet is OR-ed into LogOptions.Quiet (e.g. --json output modes).
func runVaultikApp( func runVaultikApp(
cmd *cobra.Command, extraQuiet, suppressErrors bool, cmd *cobra.Command, extraQuiet, suppressErrors bool,
failMsg string, op func(v *vaultik.Vaultik) error, failMsg string, op func(v *vaultik.Vaultik) error,
@@ -292,20 +214,47 @@ func runVaultikApp(
rootFlags := GetRootFlags() rootFlags := GetRootFlags()
return RunOperation(cmd.Context(), AppOptions{ return RunWithApp(cmd.Context(), AppOptions{
ConfigPath: configPath, ConfigPath: configPath,
LogOptions: log.Options{ LogOptions: log.Options{
Verbose: rootFlags.Verbose, Verbose: rootFlags.Verbose,
Debug: rootFlags.Debug, Debug: rootFlags.Debug,
Quiet: rootFlags.Quiet || extraQuiet, Quiet: rootFlags.Quiet || extraQuiet,
}, },
}, op, func(err error) { Modules: []fx.Option{},
if suppressErrors { Invokes: []fx.Option{
return fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
} lc.Append(fx.Hook{
OnStart: func(_ context.Context) error {
go func() {
err := op(v)
if err != nil {
if !errors.Is(err, context.Canceled) {
if !suppressErrors {
log.Error(failMsg, "error", err) log.Error(failMsg, "error", err)
ReportErrorf("%s: %v", failMsg, err) ReportErrorf("%s: %v", failMsg, err)
}
os.Exit(1)
}
}
err = v.Shutdowner.Shutdown()
if err != nil {
log.Error("Failed to shutdown", "error", err)
}
}()
return nil
},
OnStop: func(_ context.Context) error {
v.Cancel()
return nil
},
})
}),
},
}) })
} }
+2 -17
View File
@@ -1,7 +1,6 @@
package cli package cli
import ( import (
"errors"
"io" "io"
"os" "os"
"strings" "strings"
@@ -20,11 +19,7 @@ const shortCommitLen = 12
// flag is present in os.Args — see bannerSuppressedInArgs), executes the // flag is present in os.Args — see bannerSuppressedInArgs), executes the
// root cobra command, and routes any returned error through the // root cobra command, and routes any returned error through the
// ui.Writer so the user sees a properly formatted "🛑 ERROR:" line. // ui.Writer so the user sees a properly formatted "🛑 ERROR:" line.
// func Entry() {
// It returns the process exit code (0 on success, 1 on error) rather
// than calling os.Exit, so that main's deferred profile writers run
// before the process ends. See run in cmd/vaultik/main.go.
func Entry() int {
emitStartupBanner(os.Args[1:], os.Stdout) emitStartupBanner(os.Args[1:], os.Stdout)
rootCmd := NewRootCommand() rootCmd := NewRootCommand()
@@ -32,19 +27,9 @@ func Entry() int {
err := rootCmd.Execute() err := rootCmd.Execute()
if err != nil { if err != nil {
// An operation that ran inside the fx app has already reported
// its own failure (and suppressed it under --json); errReported
// says so. Printing it again here would double the error line.
// Every other error — bad arguments, a config that would not
// load — reaches Entry unreported, so it is shown here.
if !errors.Is(err, errReported) {
ReportErrorf("%s", err.Error()) ReportErrorf("%s", err.Error())
os.Exit(1)
} }
return 1
}
return 0
} }
// emitStartupBanner writes the startup banner to w unless args (the // emitStartupBanner writes the startup banner to w unless args (the
+1 -1
View File
@@ -230,7 +230,7 @@ func TestEntryJSONStdoutIsExactlyOneDocument(t *testing.T) {
programName, flagConfig, configPath, cmdSnapshot, cmdList, flagJSON, programName, flagConfig, configPath, cmdSnapshot, cmdList, flagJSON,
} }
stdout := captureProcessStdout(t, func() { _ = Entry() }) stdout := captureProcessStdout(t, Entry)
requireExactlyOneJSONDocument(t, stdout) requireExactlyOneJSONDocument(t, stdout)
+1 -1
View File
@@ -81,7 +81,7 @@ func TestEntryPruneJSONStdoutIsExactlyOneDocument(t *testing.T) {
programName, flagConfig, configPath, cmdPrune, flagJSON, programName, flagConfig, configPath, cmdPrune, flagJSON,
} }
stdout := captureProcessStdout(t, func() { _ = Entry() }) stdout := captureProcessStdout(t, Entry)
requireExactlyOneJSONDocument(t, stdout) requireExactlyOneJSONDocument(t, stdout)
-58
View File
@@ -1,58 +0,0 @@
package cli //nolint:testpackage // shares programName and the capture helpers
import (
"os"
"testing"
"github.com/stretchr/testify/assert"
)
// TestEntryReturnsStatusCode pins the contract main() relies on for
// issue #75: Entry reports success or failure through its return value
// and never calls os.Exit. An os.Exit from inside Entry would skip
// main's deferred profile writers and truncate the profile of a failing
// command. main turns this code into os.Exit only after those defers
// run, so a failing command must come back with a non-zero code rather
// than ending the process here.
//
// Stdout is captured only to keep the banner and command output off the
// test log; the assertion is on the returned code.
//
//nolint:paralleltest // replaces os.Args and rootFlags
func TestEntryReturnsStatusCode(t *testing.T) {
for _, testCase := range []struct {
name string
args []string
want int
}{
{
// version is self-contained: it needs no config and no
// destination store, so it exercises the success path.
name: "successful command returns zero",
args: []string{programName, "version"},
want: 0,
},
{
name: "unknown command returns one",
args: []string{programName, "no-such-command"},
want: 1,
},
} {
t.Run(testCase.name, func(t *testing.T) {
previousArgs := os.Args
t.Cleanup(func() {
os.Args = previousArgs
rootFlags = RootFlags{}
})
os.Args = testCase.args
var code int
_ = captureProcessStdout(t, func() { code = Entry() })
assert.Equal(t, testCase.want, code)
})
}
}
+35 -4
View File
@@ -1,7 +1,12 @@
package cli package cli
import ( import (
"context"
"errors"
"os"
"github.com/spf13/cobra" "github.com/spf13/cobra"
"go.uber.org/fx"
"sneak.berlin/go/vaultik/internal/log" "sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/vaultik" "sneak.berlin/go/vaultik/internal/vaultik"
) )
@@ -28,18 +33,44 @@ func NewInfoCommand() *cobra.Command {
// Use the app framework // Use the app framework
rootFlags := GetRootFlags() rootFlags := GetRootFlags()
return RunOperation(cmd.Context(), AppOptions{ return RunWithApp(cmd.Context(), AppOptions{
ConfigPath: configPath, ConfigPath: configPath,
LogOptions: log.Options{ LogOptions: log.Options{
Verbose: rootFlags.Verbose, Verbose: rootFlags.Verbose,
Debug: rootFlags.Debug, Debug: rootFlags.Debug,
Quiet: rootFlags.Quiet, Quiet: rootFlags.Quiet,
}, },
}, func(v *vaultik.Vaultik) error { Modules: []fx.Option{},
return v.ShowInfo() Invokes: []fx.Option{
}, func(err error) { fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
lc.Append(fx.Hook{
OnStart: func(_ context.Context) error {
go func() {
err := v.ShowInfo()
if err != nil {
if !errors.Is(err, context.Canceled) {
log.Error("Failed to show info", "error", err) log.Error("Failed to show info", "error", err)
ReportErrorf("Failed to show info: %v", err) ReportErrorf("Failed to show info: %v", err)
os.Exit(1)
}
}
err = v.Shutdowner.Shutdown()
if err != nil {
log.Error("Failed to shutdown", "error", err)
}
}()
return nil
},
OnStop: func(_ context.Context) error {
v.Cancel()
return nil
},
})
}),
},
}) })
}, },
} }
+42 -8
View File
@@ -1,7 +1,12 @@
package cli package cli
import ( import (
"context"
"errors"
"os"
"github.com/spf13/cobra" "github.com/spf13/cobra"
"go.uber.org/fx"
"sneak.berlin/go/vaultik/internal/log" "sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/vaultik" "sneak.berlin/go/vaultik/internal/vaultik"
) )
@@ -36,22 +41,51 @@ work (e.g. after a crashed backup or to reclaim storage).`,
// Use the app framework like other commands // Use the app framework like other commands
rootFlags := GetRootFlags() rootFlags := GetRootFlags()
return RunOperation(cmd.Context(), AppOptions{ return RunWithApp(cmd.Context(), AppOptions{
ConfigPath: configPath, ConfigPath: configPath,
LogOptions: log.Options{ LogOptions: log.Options{
Verbose: rootFlags.Verbose, Verbose: rootFlags.Verbose,
Debug: rootFlags.Debug, Debug: rootFlags.Debug,
Quiet: rootFlags.Quiet || opts.JSON, Quiet: rootFlags.Quiet || opts.JSON,
}, },
}, func(v *vaultik.Vaultik) error { Modules: []fx.Option{},
return v.Prune(opts) Invokes: []fx.Option{
}, func(err error) { fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
if opts.JSON { lc.Append(fx.Hook{
return OnStart: func(_ context.Context) error {
} // Start the prune operation in a goroutine
go func() {
// Run the prune operation
err := v.Prune(opts)
if err != nil {
if !errors.Is(err, context.Canceled) {
if !opts.JSON {
log.Error("Prune operation failed", "error", err) log.Error("Prune operation failed", "error", err)
ReportErrorf("Prune failed: %v", err) ReportErrorf("Prune failed: %v", err)
}
os.Exit(1)
}
}
// Shutdown the app when prune completes
err = v.Shutdowner.Shutdown()
if err != nil {
log.Error("Failed to shutdown", "error", err)
}
}()
return nil
},
OnStop: func(_ context.Context) error {
log.Debug("Stopping prune operation")
v.Cancel()
return nil
},
})
}),
},
}) })
}, },
} }
+36 -8
View File
@@ -1,9 +1,12 @@
package cli package cli
import ( import (
"context"
"errors" "errors"
"os"
"github.com/spf13/cobra" "github.com/spf13/cobra"
"go.uber.org/fx"
"sneak.berlin/go/vaultik/internal/log" "sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/vaultik" "sneak.berlin/go/vaultik/internal/vaultik"
) )
@@ -80,22 +83,47 @@ func newRemoteInfoCommand() *cobra.Command {
rootFlags := GetRootFlags() rootFlags := GetRootFlags()
return RunOperation(cmd.Context(), AppOptions{ return RunWithApp(cmd.Context(), AppOptions{
ConfigPath: configPath, ConfigPath: configPath,
LogOptions: log.Options{ LogOptions: log.Options{
Verbose: rootFlags.Verbose, Verbose: rootFlags.Verbose,
Debug: rootFlags.Debug, Debug: rootFlags.Debug,
Quiet: rootFlags.Quiet || jsonOutput, Quiet: rootFlags.Quiet || jsonOutput,
}, },
}, func(v *vaultik.Vaultik) error { Modules: []fx.Option{},
return v.RemoteInfo(jsonOutput) Invokes: []fx.Option{
}, func(err error) { fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
if jsonOutput { lc.Append(fx.Hook{
return OnStart: func(_ context.Context) error {
} go func() {
err := v.RemoteInfo(jsonOutput)
if err != nil {
if !errors.Is(err, context.Canceled) {
if !jsonOutput {
log.Error("Failed to get remote info", "error", err) log.Error("Failed to get remote info", "error", err)
ReportErrorf("Failed to get remote info: %v", err) ReportErrorf("Failed to get remote info: %v", err)
}
os.Exit(1)
}
}
err = v.Shutdowner.Shutdown()
if err != nil {
log.Error("Failed to shutdown", "error", err)
}
}()
return nil
},
OnStop: func(_ context.Context) error {
v.Cancel()
return nil
},
})
}),
},
}) })
}, },
} }
+72 -17
View File
@@ -1,10 +1,13 @@
package cli package cli
import ( import (
"context"
"errors" "errors"
"fmt" "fmt"
"os"
"github.com/spf13/cobra" "github.com/spf13/cobra"
"go.uber.org/fx"
"sneak.berlin/go/vaultik/internal/log" "sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/vaultik" "sneak.berlin/go/vaultik/internal/vaultik"
) )
@@ -83,8 +86,7 @@ specifying a path using --config or by setting VAULTIK_CONFIG to a path.`,
// Use the backup functionality from cli package // Use the backup functionality from cli package
rootFlags := GetRootFlags() rootFlags := GetRootFlags()
// --cron suppression is wired through v.UI by setupGlobals. return RunWithApp(cmd.Context(), AppOptions{
return RunOperation(cmd.Context(), AppOptions{
ConfigPath: configPath, ConfigPath: configPath,
LogOptions: log.Options{ LogOptions: log.Options{
Verbose: rootFlags.Verbose, Verbose: rootFlags.Verbose,
@@ -92,11 +94,42 @@ specifying a path using --config or by setting VAULTIK_CONFIG to a path.`,
Cron: opts.Cron, Cron: opts.Cron,
Quiet: rootFlags.Quiet, Quiet: rootFlags.Quiet,
}, },
}, func(v *vaultik.Vaultik) error { Modules: []fx.Option{},
return v.CreateSnapshot(opts) Invokes: []fx.Option{
}, func(err error) { fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
lc.Append(fx.Hook{
OnStart: func(_ context.Context) error {
// Start the snapshot creation in a goroutine
go func() {
// --cron suppression is wired through v.UI by setupGlobals.
err := v.CreateSnapshot(opts)
if err != nil {
if !errors.Is(err, context.Canceled) {
log.Error("Snapshot creation failed", "error", err) log.Error("Snapshot creation failed", "error", err)
ReportErrorf("Snapshot creation failed: %v", err) ReportErrorf("Snapshot creation failed: %v", err)
os.Exit(1)
}
}
// Shutdown the app when snapshot completes
err = v.Shutdowner.Shutdown()
if err != nil {
log.Error("Failed to shutdown", "error", err)
}
}()
return nil
},
OnStop: func(_ context.Context) error {
log.Debug("Stopping snapshot creation")
// Cancel the Vaultik context
v.Cancel()
return nil
},
})
}),
},
}) })
}, },
} }
@@ -188,10 +221,7 @@ func newSnapshotVerifyCommand() *cobra.Command {
cmd := &cobra.Command{ cmd := &cobra.Command{
Use: "verify <snapshot-id>", Use: "verify <snapshot-id>",
Short: "Verify snapshot integrity", Short: "Verify snapshot integrity",
Long: "Verifies that all blobs referenced in a snapshot exist.\n\n" + Long: "Verifies that all blobs referenced in a snapshot exist",
"The snapshot may be named by its ID or, on a host with no local\n" +
"index, by the remote key that 'snapshot list' prints for a\n" +
"remote-only snapshot (an unambiguous leading part is enough).",
Args: requireSnapshotIDArg, Args: requireSnapshotIDArg,
RunE: func(cmd *cobra.Command, args []string) error { RunE: func(cmd *cobra.Command, args []string) error {
snapshotID := args[0] snapshotID := args[0]
@@ -204,22 +234,47 @@ func newSnapshotVerifyCommand() *cobra.Command {
rootFlags := GetRootFlags() rootFlags := GetRootFlags()
return RunOperation(cmd.Context(), AppOptions{ return RunWithApp(cmd.Context(), AppOptions{
ConfigPath: configPath, ConfigPath: configPath,
LogOptions: log.Options{ LogOptions: log.Options{
Verbose: rootFlags.Verbose, Verbose: rootFlags.Verbose,
Debug: rootFlags.Debug, Debug: rootFlags.Debug,
Quiet: rootFlags.Quiet || opts.JSON, Quiet: rootFlags.Quiet || opts.JSON,
}, },
}, func(v *vaultik.Vaultik) error { Modules: []fx.Option{},
return v.VerifySnapshotWithOptions(snapshotID, opts) Invokes: []fx.Option{
}, func(err error) { fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
if opts.JSON { lc.Append(fx.Hook{
return OnStart: func(_ context.Context) error {
} go func() {
err := v.VerifySnapshotWithOptions(snapshotID, opts)
if err != nil {
if !errors.Is(err, context.Canceled) {
if !opts.JSON {
log.Error("Verification failed", "error", err) log.Error("Verification failed", "error", err)
ReportErrorf("Verification failed: %v", err) ReportErrorf("Verification failed: %v", err)
}
os.Exit(1)
}
}
err = v.Shutdowner.Shutdown()
if err != nil {
log.Error("Failed to shutdown", "error", err)
}
}()
return nil
},
OnStop: func(_ context.Context) error {
v.Cancel()
return nil
},
})
}),
},
}) })
}, },
} }
+81 -12
View File
@@ -1,8 +1,16 @@
package cli package cli
import ( import (
"context"
"errors"
"os"
"github.com/spf13/cobra" "github.com/spf13/cobra"
"go.uber.org/fx"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/globals"
"sneak.berlin/go/vaultik/internal/log" "sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/vaultik" "sneak.berlin/go/vaultik/internal/vaultik"
) )
@@ -17,6 +25,15 @@ type RestoreOptions struct {
Verify bool // Verify restored files after restore Verify bool // Verify restored files after restore
} }
// RestoreApp contains all dependencies needed for restore
type RestoreApp struct {
Globals *globals.Globals
Config *config.Config
Storage storage.Storer
Vaultik *vaultik.Vaultik
Shutdowner fx.Shutdowner
}
// newSnapshotRestoreCommand creates the 'snapshot restore' subcommand // newSnapshotRestoreCommand creates the 'snapshot restore' subcommand
func newSnapshotRestoreCommand() *cobra.Command { func newSnapshotRestoreCommand() *cobra.Command {
opts := &RestoreOptions{} opts := &RestoreOptions{}
@@ -31,10 +48,6 @@ target directory.
If no paths are specified, all files are restored. If no paths are specified, all files are restored.
If paths are specified, only matching files/directories are restored. If paths are specified, only matching files/directories are restored.
The snapshot may be named by its ID or, when restoring on a host with no
local index, by the remote key that 'snapshot list' prints for a
remote-only snapshot (an unambiguous leading part is enough).
Requires the VAULTIK_AGE_SECRET_KEY environment variable to be set with Requires the VAULTIK_AGE_SECRET_KEY environment variable to be set with
the age private key. the age private key.
@@ -64,8 +77,7 @@ Examples:
return cmd return cmd
} }
// runRestore parses arguments and runs the restore operation through the // runRestore parses arguments and runs the restore operation through the app framework
// app framework.
func runRestore(cmd *cobra.Command, args []string, opts *RestoreOptions) error { func runRestore(cmd *cobra.Command, args []string, opts *RestoreOptions) error {
snapshotID := args[0] snapshotID := args[0]
@@ -74,30 +86,87 @@ func runRestore(cmd *cobra.Command, args []string, opts *RestoreOptions) error {
opts.Paths = args[restoreMinArgs:] opts.Paths = args[restoreMinArgs:]
} }
// Use unified config resolution
configPath, err := ResolveConfigPath() configPath, err := ResolveConfigPath()
if err != nil { if err != nil {
return err return err
} }
// Use the app framework like other commands
rootFlags := GetRootFlags() rootFlags := GetRootFlags()
return RunOperation(cmd.Context(), AppOptions{ return RunWithApp(cmd.Context(), AppOptions{
ConfigPath: configPath, ConfigPath: configPath,
LogOptions: log.Options{ LogOptions: log.Options{
Verbose: rootFlags.Verbose, Verbose: rootFlags.Verbose,
Debug: rootFlags.Debug, Debug: rootFlags.Debug,
Quiet: rootFlags.Quiet, Quiet: rootFlags.Quiet,
}, },
}, func(v *vaultik.Vaultik) error { Modules: buildRestoreModules(),
return v.Restore(&vaultik.RestoreOptions{ Invokes: buildRestoreInvokes(snapshotID, opts),
})
}
// buildRestoreModules returns the fx.Options for dependency injection in restore
func buildRestoreModules() []fx.Option {
return []fx.Option{
fx.Provide(fx.Annotate(
func(g *globals.Globals, cfg *config.Config,
storer storage.Storer, v *vaultik.Vaultik, shutdowner fx.Shutdowner) *RestoreApp {
return &RestoreApp{
Globals: g,
Config: cfg,
Storage: storer,
Vaultik: v,
Shutdowner: shutdowner,
}
},
)),
}
}
// buildRestoreInvokes returns the fx.Options that wire up the restore lifecycle
func buildRestoreInvokes(snapshotID string, opts *RestoreOptions) []fx.Option {
return []fx.Option{
fx.Invoke(func(app *RestoreApp, lc fx.Lifecycle) {
lc.Append(fx.Hook{
OnStart: func(_ context.Context) error {
// Start the restore operation in a goroutine
go func() {
// Run the restore operation
restoreOpts := &vaultik.RestoreOptions{
SnapshotID: snapshotID, SnapshotID: snapshotID,
TargetDir: opts.TargetDir, TargetDir: opts.TargetDir,
Paths: opts.Paths, Paths: opts.Paths,
Verify: opts.Verify, Verify: opts.Verify,
SkipErrors: rootFlags.SkipErrors, SkipErrors: GetRootFlags().SkipErrors,
}) }
}, func(err error) {
err := app.Vaultik.Restore(restoreOpts)
if err != nil {
if !errors.Is(err, context.Canceled) {
log.Error("Restore operation failed", "error", err) log.Error("Restore operation failed", "error", err)
ReportErrorf("Restore failed: %v", err) ReportErrorf("Restore failed: %v", err)
os.Exit(1)
}
}
// Shutdown the app when restore completes
err = app.Shutdowner.Shutdown()
if err != nil {
log.Error("Failed to shutdown", "error", err)
}
}()
return nil
},
OnStop: func(_ context.Context) error {
log.Debug("Stopping restore operation")
app.Vaultik.Cancel()
return nil
},
}) })
}),
}
} }
+5 -10
View File
@@ -18,6 +18,7 @@ import (
"sneak.berlin/go/vaultik/internal/blobgen" "sneak.berlin/go/vaultik/internal/blobgen"
"sneak.berlin/go/vaultik/internal/database" "sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log" "sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/snapshot"
"sneak.berlin/go/vaultik/internal/types" "sneak.berlin/go/vaultik/internal/types"
) )
@@ -576,20 +577,14 @@ func (v *Vaultik) handleRestoreVerification(
} }
// downloadSnapshotDB downloads and decrypts the snapshot metadata // downloadSnapshotDB downloads and decrypts the snapshot metadata
// database. The identifier is resolved to the snapshot's remote key: a // database. The snapshotID is the human ID; we hash it to the remote
// human ID is hashed, and a remote key (or its abbreviation, as printed // key for the storage path.
// for a remote-only snapshot) is used as-is, so a host with no local
// index can restore the snapshots it can only see on the store.
func (v *Vaultik) downloadSnapshotDB( func (v *Vaultik) downloadSnapshotDB(
snapshotID string, identity age.Identity, snapshotID string, identity age.Identity,
) (*database.DB, error) { ) (*database.DB, error) {
remoteKey, err := v.resolveSnapshotRemoteKey(snapshotID)
if err != nil {
return nil, err
}
// Download encrypted database from storage // Download encrypted database from storage
dbKey := fmt.Sprintf("metadata/%s/db.zst.age", remoteKey) dbKey := fmt.Sprintf("metadata/%s/db.zst.age",
snapshot.RemoteSnapshotKey(snapshotID))
reader, err := v.Storage.Get(v.ctx, dbKey) reader, err := v.Storage.Get(v.ctx, dbKey)
if err != nil { if err != nil {
@@ -1,167 +0,0 @@
package vaultik_test
import (
"bytes"
"context"
"io"
"path/filepath"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/snapshot"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/ui"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// TestRestoreOnAnotherMachine proves the disaster-recovery path: a host
// that has only the vaultik binary, the age secret key, and the storage
// credentials — no local index, a different hostname, and no
// age_recipients configured — can list, restore, and verify a snapshot
// straight from the destination store.
//
// The backup half writes a snapshot with one index and hostname. The
// restore half throws that index away entirely: a fresh, empty index and
// a config that shares nothing with the original but the storage location
// and the secret key. If restore or verify needed the original local
// index — or the human snapshot ID that only that index holds — this test
// could not run, because the recovery host can know neither.
func TestRestoreOnAnotherMachine(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "source")
storeDir := filepath.Join(tempDir, "remote")
restoreDir := filepath.Join(tempDir, "restored")
dbPath := filepath.Join(tempDir, "index.sqlite")
chunkSize := int64(64 * 1024)
maxBlobSize := int64(512 * 1024)
sourceFiles := writeRecoverySourceTree(t, fs, dataDir, chunkSize)
ctx := context.Background()
// Backup host: one index, hostname test-host, age_recipients set.
// runFileStorageBackup closes the index before returning, so nothing
// below can lean on it.
_, storer, originalID := runFileStorageBackup(
ctx, t, fs, dataDir, storeDir, dbPath, chunkSize, maxBlobSize)
// Recovery host: a fresh empty index, a different hostname, and no
// age_recipients — only the secret key and the same storage location.
recovery, stdout := newRecoveryHost(ctx, t, fs, storer)
// The recovery index really is empty. This is the assertion that makes
// the test a guard against restore quietly depending on the original
// index: if it did, an empty index would make restore fail.
localSnaps, err := recovery.Repositories.Snapshots.ListRecent(ctx, 100)
require.NoError(t, err)
require.Empty(t, localSnaps, "recovery host must start with no local index")
// List: the snapshot shows up as remote-only, identified by its remote
// key, with no recoverable human ID.
require.NoError(t, recovery.ListSnapshots(true))
rows := decodeListJSON(t, stdout.String())
require.Len(t, rows, 1)
remote := rows[0]
assert.False(t, remote.LocallyTracked, "snapshot must be remote-only here")
assert.Empty(t, remote.ID, "the human ID is unknown to the recovery host")
require.Len(t, remote.RemoteKey, 64)
assert.Equal(t, snapshot.RemoteSnapshotKey(originalID), remote.RemoteKey,
"the listed key is the hashed snapshot ID")
// Restore driven by the abbreviated identifier the table prints (the
// first 12 hex of the remote key), then deep-verify from the store
// keyed by the full remote key. Both are what a recovery host can know.
require.NoError(t, recovery.Restore(&vaultik.RestoreOptions{
SnapshotID: remote.RemoteKey[:12],
TargetDir: restoreDir,
Verify: true,
}))
require.NoError(t, recovery.RunDeepVerify(
remote.RemoteKey, &vaultik.VerifyOptions{Deep: true}))
assertRestoredTreeMatches(t, fs, restoreDir, sourceFiles)
}
// writeRecoverySourceTree writes a small source tree spanning several
// chunks (so restore reassembles real multi-chunk files) and returns the
// content keyed by absolute path.
func writeRecoverySourceTree(
t *testing.T, fs afero.Fs, dataDir string, chunkSize int64,
) map[string][]byte {
t.Helper()
sourceFiles := map[string][]byte{
filepath.Join(dataDir, "notes.txt"): []byte("recover me"),
filepath.Join(dataDir, "sub", "big.bin"): bytesPattern("big-", int(chunkSize*3)),
filepath.Join(dataDir, "sub", "small.bin"): bytesPattern("small-", 128),
}
for path, content := range sourceFiles {
require.NoError(t, fs.MkdirAll(filepath.Dir(path), 0o755))
require.NoError(t, afero.WriteFile(fs, path, content, 0o644))
}
return sourceFiles
}
// newRecoveryHost builds the Vaultik a replacement machine would run: an
// empty in-memory index, a hostname different from the backup host, no
// age_recipients, and only the secret key plus the shared storer. It
// returns the instance and the buffer its stdout is wired to.
func newRecoveryHost(
ctx context.Context, t *testing.T, fs afero.Fs, storer storage.Storer,
) (*vaultik.Vaultik, *bytes.Buffer) {
t.Helper()
recoveryDB, err := database.New(ctx, ":memory:")
require.NoError(t, err)
t.Cleanup(func() { _ = recoveryDB.Close() })
stdout := &bytes.Buffer{}
recovery := &vaultik.Vaultik{
Config: &config.Config{
AgeSecretKey: testAgeSecretKey,
Hostname: "recovery-host",
},
Storage: storer,
Fs: fs,
Repositories: database.NewRepositories(recoveryDB),
DB: recoveryDB,
Stdout: stdout,
Stderr: io.Discard,
UI: ui.NewWithColor(io.Discard, false),
}
recovery.SetContext(ctx)
return recovery, stdout
}
// assertRestoredTreeMatches byte-compares every restored file against its
// source content.
func assertRestoredTreeMatches(
t *testing.T, fs afero.Fs, restoreDir string, sourceFiles map[string][]byte,
) {
t.Helper()
for origPath, expected := range sourceFiles {
restored := filepath.Join(restoreDir, origPath)
got, err := afero.ReadFile(fs, restored)
require.NoErrorf(t, err, "restored file missing: %s", restored)
require.Truef(t, bytes.Equal(got, expected),
"byte mismatch for %s", origPath)
}
}
+3 -5
View File
@@ -670,11 +670,9 @@ func (v *Vaultik) VerifySnapshotWithOptions(
v.printVerifyHeader(snapshotID, opts) v.printVerifyHeader(snapshotID, opts)
// Resolve the identifier to the snapshot's remote key and download the // Download and parse manifest. The caller supplies a human
// manifest. A human ID is hashed; a remote key (or its abbreviation, // snapshot ID; we hash it to address remote storage.
// as printed for a remote-only snapshot) is used as-is, so a host with manifest, err := v.downloadManifestByKey(snapshot.RemoteSnapshotKey(snapshotID))
// no local index can verify a snapshot it can only see on the store.
manifest, err := v.resolveAndDownloadManifest(snapshotID)
if err != nil { if err != nil {
if opts.JSON { if opts.JSON {
result.Status = verifyStatusFailed result.Status = verifyStatusFailed
-101
View File
@@ -1,101 +0,0 @@
package vaultik
import (
"errors"
"fmt"
"strings"
"sneak.berlin/go/vaultik/internal/snapshot"
)
// remoteKeyHexLen is the length of a full remote snapshot key: a SHA256
// digest rendered as lowercase hex.
const remoteKeyHexLen = 64
// Sentinel errors for resolving a snapshot identifier against the store.
var (
errSnapshotKeyNotFound = errors.New(
"no snapshot on the destination store matches this identifier")
errSnapshotKeyAmbiguous = errors.New(
"identifier matches more than one snapshot on the destination store")
)
// resolveSnapshotRemoteKey turns a snapshot identifier supplied on the
// command line into the remote key that names the snapshot's metadata
// directory on the destination store. Every remote path a restore or
// verify reads is built from that key.
//
// Two forms are accepted, matching the two things a host can know:
//
// - A human snapshot ID (hostname_name_timestamp), which a host holding
// the local index has. It is hashed to its remote key; the store is
// not consulted.
// - A remote key, or the leading part of one, which is all a host with
// no local index can know — it is exactly what `snapshot list` prints
// for a remote-only snapshot (see formatRemoteOnlyID). It is resolved
// against the destination store's metadata listing; an identifier that
// matches no snapshot, or more than one, is an error.
//
// The two are told apart by shape: a remote key is lowercase hex, and a
// human snapshot ID never is (it carries a hostname, underscores, and an
// RFC3339 timestamp).
func (v *Vaultik) resolveSnapshotRemoteKey(identifier string) (string, error) {
if !isRemoteKeyOrPrefix(identifier) {
return snapshot.RemoteSnapshotKey(identifier), nil
}
keys, err := v.listAllRemoteSnapshotKeys()
if err != nil {
return "", fmt.Errorf(
"listing destination store to resolve %q: %w", identifier, err)
}
var matches []string
for _, key := range keys {
if strings.HasPrefix(key, identifier) {
matches = append(matches, key)
}
}
switch len(matches) {
case 1:
return matches[0], nil
case 0:
return "", fmt.Errorf("%w: %s", errSnapshotKeyNotFound, identifier)
default:
return "", fmt.Errorf("%w: %s (%d matches)",
errSnapshotKeyAmbiguous, identifier, len(matches))
}
}
// resolveAndDownloadManifest resolves a snapshot identifier to its remote
// key (see resolveSnapshotRemoteKey) and downloads that snapshot's
// manifest.
func (v *Vaultik) resolveAndDownloadManifest(
identifier string,
) (*snapshot.Manifest, error) {
remoteKey, err := v.resolveSnapshotRemoteKey(identifier)
if err != nil {
return nil, err
}
return v.downloadManifestByKey(remoteKey)
}
// isRemoteKeyOrPrefix reports whether s is a full remote key or the
// leading part of one: 1 to 64 lowercase hex characters. A human snapshot
// ID is never all hex, so this shape test is enough to tell the two apart.
func isRemoteKeyOrPrefix(s string) bool {
if s == "" || len(s) > remoteKeyHexLen {
return false
}
for _, r := range s {
if (r < '0' || r > '9') && (r < 'a' || r > 'f') {
return false
}
}
return true
}
+9 -18
View File
@@ -138,15 +138,8 @@ func (v *Vaultik) RunDeepVerify(snapshotID string, opts *VerifyOptions) error {
func (v *Vaultik) loadVerificationData( func (v *Vaultik) loadVerificationData(
snapshotID string, opts *VerifyOptions, result *VerifyResult, snapshotID string, opts *VerifyOptions, result *VerifyResult,
) (*snapshot.Manifest, *tempDB, []snapshot.BlobInfo, error) { ) (*snapshot.Manifest, *tempDB, []snapshot.BlobInfo, error) {
// Resolve the identifier to the snapshot's remote key. A human ID is // All remote paths use the hashed key derived from the human ID.
// hashed; a remote key (or its abbreviation, as printed for a remoteKey := snapshot.RemoteSnapshotKey(snapshotID)
// remote-only snapshot) is used as-is, so a host with no local index
// can verify a snapshot it can only see on the store.
remoteKey, err := v.resolveSnapshotRemoteKey(snapshotID)
if err != nil {
return nil, nil, nil, v.deepVerifyFailure(result, opts,
fmt.Sprintf("resolving snapshot identifier: %v", err), err)
}
// Download manifest. downloadManifestByKey is the single reader for // Download manifest. downloadManifestByKey is the single reader for
// remote manifests; see its doc comment. // remote manifests; see its doc comment.
@@ -193,7 +186,7 @@ func (v *Vaultik) loadVerificationData(
fmt.Errorf("failed to decrypt database: %w", err)) fmt.Errorf("failed to decrypt database: %w", err))
} }
dbBlobs, err := v.getBlobsFromDatabase(tdb.DB) dbBlobs, err := v.getBlobsFromDatabase(snapshotID, tdb.DB)
if err != nil { if err != nil {
_ = tdb.Close() _ = tdb.Close()
@@ -508,21 +501,19 @@ func (v *Vaultik) verifyBlobFinalIntegrity(
return nil return nil
} }
// getBlobsFromDatabase gets all blobs for the snapshot from the database. // getBlobsFromDatabase gets all blobs for the snapshot from the database
// func (v *Vaultik) getBlobsFromDatabase(
// The exported per-snapshot database holds exactly one snapshot's data snapshotID string, db *sql.DB,
// (see cleanSnapshotDB), so every row in snapshot_blobs belongs to it. ) ([]snapshot.BlobInfo, error) {
// We select them directly rather than filtering by the human snapshot ID,
// which a host restoring from the store alone does not have.
func (v *Vaultik) getBlobsFromDatabase(db *sql.DB) ([]snapshot.BlobInfo, error) {
query := ` query := `
SELECT b.blob_hash, b.compressed_size SELECT b.blob_hash, b.compressed_size
FROM snapshot_blobs sb FROM snapshot_blobs sb
JOIN blobs b ON sb.blob_hash = b.blob_hash JOIN blobs b ON sb.blob_hash = b.blob_hash
WHERE sb.snapshot_id = ?
ORDER BY b.blob_hash ORDER BY b.blob_hash
` `
rows, err := db.QueryContext(v.ctx, query) rows, err := db.QueryContext(v.ctx, query, snapshotID)
if err != nil { if err != nil {
return nil, fmt.Errorf("failed to query snapshot blobs: %w", err) return nil, fmt.Errorf("failed to query snapshot blobs: %w", err)
} }
+1 -16
View File
@@ -56,23 +56,8 @@ main() {
docker build --output=type=cacheonly \ docker build --output=type=cacheonly \
--build-arg CHECK_EPOCH="$epoch" -f Dockerfile.lint . --build-arg CHECK_EPOCH="$epoch" -f Dockerfile.lint .
# Version, commit and build date are computed here on the host, the
# same way script/docker does, and passed into the product build so
# the CI-built image reports its real source. The build context
# excludes .git (see .dockerignore), so the build cannot derive them
# itself; without these it would stamp the Dockerfile's dev/unknown
# fallbacks. VERSION comes from script/version, the source of truth
# shared with the Makefile.
version="$("$ROOT/script/version")"
commit="$(git rev-parse HEAD 2>/dev/null || echo unknown)"
commit_date="$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)"
epoch="$(date +%s%N)$$" epoch="$(date +%s%N)$$"
docker build --build-arg CHECK_EPOCH="$epoch" \ docker build --build-arg CHECK_EPOCH="$epoch" .
--build-arg VERSION="$version" \
--build-arg COMMIT="$commit" \
--build-arg COMMIT_DATE="$commit_date" \
.
} }
main "$@" main "$@"
-17
View File
@@ -24,24 +24,7 @@ main() {
# whether the tree is clean. The Dockerfile now refuses to build # whether the tree is clean. The Dockerfile now refuses to build
# without a non-empty value, so this is required, not optional. # without a non-empty value, so this is required, not optional.
epoch="$(date +%s%N)$$" epoch="$(date +%s%N)$$"
# Version, commit and build date are computed here on the host,
# where .git exists, and passed into the build. The build context
# excludes .git (see .dockerignore), so the container cannot derive
# them itself -- it used to try and always got "unknown", giving
# every image a "commit: unknown" it could not be traced from.
# VERSION comes from script/version, the source of truth shared with
# the Makefile, so a Docker build reports the same string (tag,
# dev-<sha>, or a -dirty variant) that a local build of the same
# tree would.
version="$("$SCRIPT_DIR/version")"
commit="$(git rev-parse HEAD 2>/dev/null || echo unknown)"
commit_date="$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)"
docker build --build-arg CHECK_EPOCH="$epoch" \ docker build --build-arg CHECK_EPOCH="$epoch" \
--build-arg VERSION="$version" \
--build-arg COMMIT="$commit" \
--build-arg COMMIT_DATE="$commit_date" \
-t "$("$SCRIPT_DIR/projectname")" . -t "$("$SCRIPT_DIR/projectname")" .
} }