Compare commits
31
Commits
main
..
46c295acf3
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
46c295acf3 | ||
|
|
b4654f8e52 | ||
|
|
39aef1c47c | ||
|
|
96ebcd40d7 | ||
|
|
d9f0220f94 | ||
|
|
4c83e82543 | ||
|
|
86361c8b50 | ||
|
|
d77663d039 | ||
|
|
3abe9cbd9e | ||
|
|
76a6917a35 | ||
|
|
38ebfd843a | ||
|
|
6b7517a4dc | ||
|
|
994e5de613 | ||
|
|
42f4e648d7 | ||
|
|
343129f891 | ||
|
|
a50e3fa038 | ||
|
|
6fcd8e1668 | ||
|
|
aab6a87f8c | ||
|
|
c355ef4d25 | ||
|
|
5927e1aa3d | ||
|
|
9ca962969a | ||
|
|
89ebfc78e2 | ||
|
|
07ef3a1c78 | ||
|
|
c423d13191 | ||
|
|
75a10d3a22 | ||
|
|
3d56dd7eb0 | ||
|
|
bdce350041 | ||
|
|
753bc3ef60 | ||
|
|
d2a0510cb4 | ||
|
|
583f65040a | ||
|
|
d257f8f658 |
@@ -1,9 +1,9 @@
|
||||
name: check
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
branches: [main, next]
|
||||
pull_request:
|
||||
branches: [main]
|
||||
branches: [main, next]
|
||||
jobs:
|
||||
check:
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
@@ -20,33 +20,21 @@ jobs:
|
||||
# check.yml runs script/cibuild, which does all of its work inside
|
||||
# the digest-pinned Dockerfile images -- so without this step the
|
||||
# release either fails at the before-hook or, worse, ships binaries
|
||||
# built by whatever unpinned Go the runner happens to carry.
|
||||
# REPO_POLICIES.md requires every external reference to be pinned,
|
||||
# and script/release already refuses a goreleaser that is not the
|
||||
# pinned build; the compiler that actually produces the artifacts
|
||||
# is the last thing that should be exempt from that.
|
||||
# built by whatever Go the runner happens to carry.
|
||||
#
|
||||
# go-version-file rather than a literal: go.mod's `go 1.26.1` is
|
||||
# the single source of truth for the toolchain, the same way the
|
||||
# Dockerfile FROM line is the single source of truth for the
|
||||
# linter version that script/lint enforces. It is a three-component
|
||||
# version, so setup-go resolves it exactly -- no silent drift onto
|
||||
# a newer patch release.
|
||||
#
|
||||
# actions/setup-go v5.6.0, 2025-12-15. Pinned by commit sha, like
|
||||
# the checkout above. v5.x is a node20 action, matching the node20
|
||||
# actions/checkout v4 already in use here; the v6/v7 line requires
|
||||
# a node24 runner, which this Gitea runner has never been asked
|
||||
# for and cannot be assumed to provide.
|
||||
# actions/setup-go would pin the action by commit sha, but the Go
|
||||
# tarball it downloads at runtime is verified against no value in
|
||||
# this repo, and the action exposes no checksum input.
|
||||
# REPO_POLICIES.md requires every external reference to be pinned
|
||||
# by hash with no exceptions, and this is the compiler that
|
||||
# produces the published binaries -- the input where a substituted
|
||||
# artifact matters most. So Go is installed the way goreleaser is:
|
||||
# script/install-go downloads the exact archive for go.mod's `go`
|
||||
# directive and refuses it unless its sha256 matches the value
|
||||
# committed in the script, then puts .tool/go/bin on PATH for the
|
||||
# steps below.
|
||||
- name: Install Go
|
||||
uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff
|
||||
with:
|
||||
go-version-file: go.mod
|
||||
# setup-go's module cache needs a runner-side cache backend.
|
||||
# A release is cut rarely and a cold module download costs
|
||||
# seconds; a release failing because a cache service is absent
|
||||
# costs a re-tag. Off, deliberately.
|
||||
cache: false
|
||||
run: script/install-go
|
||||
- name: Install goreleaser
|
||||
run: script/install-goreleaser
|
||||
- name: Release
|
||||
@@ -58,3 +46,8 @@ jobs:
|
||||
# It is deliberately not the runner's automatic token, which is
|
||||
# not guaranteed to carry that scope.
|
||||
GITEA_TOKEN: ${{ secrets.RELEASE_TOKEN }}
|
||||
# Build with the toolchain install-go just verified, never a
|
||||
# different one auto-downloaded from a `toolchain` directive:
|
||||
# the point of the hash pin is that this exact compiler makes
|
||||
# the release.
|
||||
GOTOOLCHAIN: local
|
||||
|
||||
@@ -104,7 +104,12 @@ Version: 2025-06-08
|
||||
|
||||
13. Pre-1.0: NEVER write database migrations. There are no live databases
|
||||
anywhere — every user's local index can be rebuilt from a fresh full
|
||||
backup. When the schema changes, just change `schema.sql` (and any code
|
||||
that touches the affected tables). The local index is disposable until
|
||||
1.0 ships and is tagged.
|
||||
backup. To change the schema, edit `internal/database/schema/001.sql`
|
||||
(and any code that touches the affected tables) directly; do not add new
|
||||
numbered schema files. Those numbered files and the `schema_migrations`
|
||||
table they populate only bootstrap a fresh database — they are not an
|
||||
upgrade path. The local index is disposable until 1.0 ships and is
|
||||
tagged; once 1.0 is tagged that clause expires and the question of
|
||||
upgrading existing indexes returns. See [`docs/DATAMODEL.md`](docs/DATAMODEL.md)
|
||||
for the full explanation.
|
||||
|
||||
|
||||
+10
-4
@@ -63,7 +63,7 @@ A content-addressed unit of data. Files are split into variable-size chunks usin
|
||||
- `ChunkHash`: SHA256 hash of chunk content (primary key)
|
||||
- `Size`: Chunk size in bytes
|
||||
|
||||
Chunk sizes vary between `avgChunkSize/4` and `avgChunkSize*4` (typically 16KB-256KB for 64KB average).
|
||||
Chunk sizes vary between `avgChunkSize/4` and `avgChunkSize*4` (2.5MB-40MB for the 10MB default average).
|
||||
|
||||
#### FileChunk (`database.FileChunk`)
|
||||
Maps files to their constituent chunks:
|
||||
@@ -120,7 +120,7 @@ The CLI uses fx for dependency injection. Here's the instantiation order:
|
||||
```go
|
||||
// cli/app.go: NewApp()
|
||||
fx.New(
|
||||
fx.Supply(config.ConfigPath(opts.ConfigPath)), // 1. Config path
|
||||
fx.Supply(config.Path(opts.ConfigPath)), // 1. Config path
|
||||
fx.Supply(opts.LogOptions), // 2. Log options
|
||||
fx.Provide(globals.New), // 3. Globals
|
||||
fx.Provide(log.New), // 4. Logger config
|
||||
@@ -193,7 +193,7 @@ scanner := v.ScannerFactory(snapshot.ScannerParams{
|
||||
- **Created by**: `chunker.NewChunker(avgChunkSize)`
|
||||
- **When**: Inside `snapshot.NewScanner()`
|
||||
- **Configuration**:
|
||||
- `avgChunkSize`: From config (typically 64KB)
|
||||
- `avgChunkSize`: From config (default 10MB)
|
||||
- `minChunkSize`: avgChunkSize / 4
|
||||
- `maxChunkSize`: avgChunkSize * 4
|
||||
|
||||
@@ -366,11 +366,17 @@ bucket/
|
||||
│ └── {full-hash} # Compressed+encrypted blob
|
||||
│
|
||||
└── metadata/
|
||||
└── {snapshot-id}/
|
||||
└── {remote-key}/
|
||||
├── db.zst.age # Encrypted binary SQLite database
|
||||
└── manifest.json.zst # Blob list (for pruning/verification)
|
||||
```
|
||||
|
||||
The `{remote-key}` directory name is a one-way double SHA-256 hash of the human
|
||||
snapshot ID, so the human ID (hostname, snapshot name, timestamp) is never
|
||||
written to the store as a directory name. See
|
||||
[docs/REPOSTRUCTURE.md](docs/REPOSTRUCTURE.md#remote-key-derivation) for the
|
||||
derivation and a worked example.
|
||||
|
||||
## Thread Safety
|
||||
|
||||
- `Packer`: Thread-safe via mutex. Multiple goroutines can call `AddChunk()`.
|
||||
|
||||
+49
-50
@@ -1,24 +1,30 @@
|
||||
# Lint stage
|
||||
# This file has no lint stage, deliberately.
|
||||
#
|
||||
# This FROM line is the single source of truth for the linter version:
|
||||
# script/lint parses the image reference out of it and runs that exact
|
||||
# image, so a local `make lint` and CI use the same linter. Bump the
|
||||
# linter here (tag AND digest) and nowhere else.
|
||||
# Linting lives in Dockerfile.lint, built by script/lint, and
|
||||
# script/cibuild builds both. A lint stage here would have to either
|
||||
# shell out to `make lint` -- which is now `docker build`, so
|
||||
# docker-in-docker inside a BuildKit step with no daemon -- or call
|
||||
# golangci-lint directly, which would mean a second, independently
|
||||
# bumpable digest pin for the linter alongside the one in
|
||||
# Dockerfile.lint. Two pins for one tool is the drift that
|
||||
# https://git.eeqj.de/sneak/vaultik/issues/78 was filed over. See
|
||||
# https://git.eeqj.de/sneak/vaultik/issues/113 for the ruling.
|
||||
#
|
||||
# golangci/golangci-lint:v2.12.2-alpine, 2026-08-07
|
||||
FROM golangci/golangci-lint:v2.12.2-alpine@sha256:91b27804074a0bacea298707f016911e60cf0cdbc6c7bf5ccacb5f0606d18d60 AS lint
|
||||
# Consequence, stated rather than left to be discovered: script/docker
|
||||
# builds this file only and therefore does not lint. `make fmt-check`
|
||||
# and `make test` still run here, so what a green build of this file
|
||||
# means is "formatted, tested, and it compiles" -- the lint verdict
|
||||
# comes from script/lint or script/cibuild.
|
||||
|
||||
# Build stage
|
||||
# golang:1.26.1-alpine, 2026-03-17
|
||||
FROM golang:1.26.1-alpine@sha256:2389ebfa5b7f43eeafbd6be0c3700cc46690ef842ad962f6c5bd6be49ed82039 AS builder
|
||||
|
||||
# Build tooling: make, plus a C toolchain because `go test -race` needs cgo.
|
||||
# The sqlite driver is pure Go (modernc.org/sqlite), so no sqlite library or
|
||||
# CLI is required.
|
||||
RUN apk add --no-cache make build-base
|
||||
|
||||
# The context signal for script/lint's native path. This stage runs
|
||||
# `make lint` with no docker daemon available, so it is the one place
|
||||
# that must run the golangci-lint on PATH directly. script/lint takes
|
||||
# that path only when this is set AND the version matches the pin above;
|
||||
# version equality alone would also admit a developer's locally
|
||||
# installed copy on a host, bypassing the digest pin (issue #80).
|
||||
# Nothing outside this stage sets it.
|
||||
ENV VAULTIK_LINT_IN_CONTAINER=1
|
||||
|
||||
WORKDIR /src
|
||||
|
||||
# Copy go mod files first for better layer caching
|
||||
@@ -28,7 +34,7 @@ RUN go mod download
|
||||
# Copy source code
|
||||
COPY . .
|
||||
|
||||
# Run formatting check and linter.
|
||||
# Run the format check and the tests.
|
||||
#
|
||||
# CHECK_EPOCH must stay immediately above these RUNs. These layers are
|
||||
# keyed on its value, so they are cache-eligible only for a value
|
||||
@@ -47,55 +53,48 @@ COPY . .
|
||||
# runs the checks and every one after it on an unchanged tree replays
|
||||
# these layers from cache, executes nothing, and still exits 0. Failed
|
||||
# steps are never cached, so the guard fails on EVERY invocation rather
|
||||
# than once -- a bare `docker build .` is now a loud error, not a quiet
|
||||
# than once -- a bare `docker build .` is a loud error, not a quiet
|
||||
# green. Do not give CHECK_EPOCH a default value; a default would
|
||||
# satisfy the guard with a constant and restore the hole.
|
||||
#
|
||||
# ARG scope is per-stage, so the builder stage declares its own.
|
||||
# Everything above this line (apk, go.mod, `go mod download`) is
|
||||
# deliberately outside the busted range and keeps caching.
|
||||
ARG CHECK_EPOCH
|
||||
RUN [ -n "$CHECK_EPOCH" ] || exit 1
|
||||
RUN echo "check epoch: ${CHECK_EPOCH}" && make fmt-check
|
||||
RUN echo "check epoch: ${CHECK_EPOCH}" && make lint
|
||||
|
||||
# Build stage
|
||||
# golang:1.26.1-alpine, 2026-03-17
|
||||
FROM golang:1.26.1-alpine@sha256:2389ebfa5b7f43eeafbd6be0c3700cc46690ef842ad962f6c5bd6be49ed82039 AS builder
|
||||
|
||||
# Depend on lint stage passing
|
||||
COPY --from=lint /src/go.sum /dev/null
|
||||
|
||||
ARG VERSION=dev
|
||||
|
||||
# Install build dependencies for CGO (mattn/go-sqlite3) and sqlite3 CLI (tests)
|
||||
RUN apk add --no-cache make build-base sqlite
|
||||
|
||||
WORKDIR /src
|
||||
|
||||
# Copy go mod files first for better layer caching
|
||||
COPY go.mod go.sum ./
|
||||
RUN go mod download
|
||||
|
||||
# Copy source code
|
||||
COPY . .
|
||||
|
||||
# Run tests. See the CHECK_EPOCH comment in the lint stage for the
|
||||
# mechanism; ARG scope is per-stage, so this stage needs its own
|
||||
# declaration, its own guard, and its own expansion, and they must stay
|
||||
# immediately above the check RUN.
|
||||
ARG CHECK_EPOCH
|
||||
RUN [ -n "$CHECK_EPOCH" ] || exit 1
|
||||
RUN echo "check epoch: ${CHECK_EPOCH}" && make test
|
||||
|
||||
# Version, commit and build date are computed on the host by
|
||||
# script/docker and script/cibuild (where .git exists) and passed in as
|
||||
# build args. The build context excludes .git (see .dockerignore), so
|
||||
# the build cannot derive them itself: it used to try, with `git
|
||||
# rev-parse` inside this stage, and always got "unknown". VERSION comes
|
||||
# from script/version, the source of truth shared with the Makefile, so
|
||||
# it carries the same tag / dev-<sha> / -dirty rules and a Docker image
|
||||
# reports the same string a local build of the same tree would.
|
||||
#
|
||||
# The defaults are the fallback for a bare `docker build .` that passes
|
||||
# none of them: an unset arg would otherwise stamp an empty string and
|
||||
# produce an image that cannot report its own version, commit or date.
|
||||
# They match what an out-of-git build reports elsewhere.
|
||||
#
|
||||
# These ARGs sit here, after the checks, rather than at the top of the
|
||||
# stage: every commit changes their values, and a value change
|
||||
# invalidates all layers below the ARG. Declared up top they would bust
|
||||
# `go mod download`; here they only rekey this build layer, which the
|
||||
# COPY of the sources above already rebuilds on any change anyway.
|
||||
ARG VERSION=dev
|
||||
ARG COMMIT=unknown
|
||||
ARG COMMIT_DATE=unknown
|
||||
|
||||
# Build (pure Go, no CGO required since we use modernc.org/sqlite)
|
||||
RUN CGO_ENABLED=0 go build -ldflags "-X 'sneak.berlin/go/vaultik/internal/globals.Version=${VERSION}' -X 'sneak.berlin/go/vaultik/internal/globals.Commit=$(git rev-parse HEAD 2>/dev/null || echo unknown)' -X 'sneak.berlin/go/vaultik/internal/globals.CommitDate=$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)'" -o /vaultik ./cmd/vaultik
|
||||
RUN CGO_ENABLED=0 go build -ldflags "-X 'sneak.berlin/go/vaultik/internal/globals.Version=${VERSION}' -X 'sneak.berlin/go/vaultik/internal/globals.Commit=${COMMIT}' -X 'sneak.berlin/go/vaultik/internal/globals.CommitDate=${COMMIT_DATE}'" -o /vaultik ./cmd/vaultik
|
||||
|
||||
# Runtime stage
|
||||
# alpine:3.21, 2026-02-25
|
||||
FROM alpine:3.21@sha256:c3f8e73fdb79deaebaa2037150150191b9dcbfba68b4a46d70103204c53f4709
|
||||
|
||||
RUN apk add --no-cache ca-certificates sqlite
|
||||
RUN apk add --no-cache ca-certificates
|
||||
|
||||
# Copy binary from builder
|
||||
COPY --from=builder /vaultik /usr/local/bin/vaultik
|
||||
|
||||
+104
@@ -0,0 +1,104 @@
|
||||
# Lint image.
|
||||
#
|
||||
# Every lint run in this repo happens inside this image, invoked through
|
||||
# script/lint, and linting is a BUILD STEP rather than a container
|
||||
# command: a successful build of this file IS a clean lint. That shape
|
||||
# also works where the docker daemon is remote and bind mounts are
|
||||
# impossible, which `docker run` against a mounted worktree does not.
|
||||
#
|
||||
# This FROM line is the single source of truth for the linter version in
|
||||
# this repo. Nothing else pins golangci-lint: the product Dockerfile has
|
||||
# no lint stage, deliberately, so there is no second digest to bump and
|
||||
# no pair of pins that can drift apart. Bump the tag AND the digest here
|
||||
# and nowhere else.
|
||||
#
|
||||
# Note for readers coming from REPO_POLICIES.md: that document still
|
||||
# describes the older pattern, a lint stage inside the product
|
||||
# Dockerfile wired up with `COPY --from=lint /src/go.sum /dev/null`.
|
||||
# That pattern is superseded here by the owner's ruling recorded in
|
||||
# https://git.eeqj.de/sneak/vaultik/issues/113 -- lint runs in its own
|
||||
# image, per run, with its own cache and its own lock, which is what
|
||||
# makes concurrent runs on one host safe. The policy text is org-wide
|
||||
# and is being amended separately; this file is what this repo does.
|
||||
#
|
||||
# golangci/golangci-lint:v2.12.2, 2026-08-10
|
||||
FROM golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240
|
||||
|
||||
WORKDIR /src
|
||||
|
||||
# Copy the dependency manifests first so the module download layer stays
|
||||
# cached until they change. Everything above the ARG below is cacheable
|
||||
# on purpose; a cold module download on every lint would make the inner
|
||||
# loop unusable and buys nothing, because it is not what the gate is
|
||||
# asserting.
|
||||
COPY go.mod go.sum ./
|
||||
RUN go mod download
|
||||
|
||||
COPY . .
|
||||
|
||||
# Force the check layers to execute on every invocation.
|
||||
#
|
||||
# CHECK_EPOCH must stay immediately above the RUNs below. Those layers
|
||||
# are keyed on its value, so they are cache-eligible only for a value
|
||||
# already built against this same tree; script/lint and script/cibuild
|
||||
# each pass a fresh value on every invocation, which is what makes their
|
||||
# green mean the linter really ran. Without it, `docker build -f
|
||||
# Dockerfile.lint .` on an unchanged tree exits 0 in well under a second
|
||||
# having linted nothing.
|
||||
#
|
||||
# The value is expanded into each check command itself rather than left
|
||||
# to a bare declaration, so the cache miss does not depend on BuildKit's
|
||||
# unreferenced-ARG handling staying as it is. It also puts the epoch in
|
||||
# the build log, where a reader can see the layer was keyed fresh.
|
||||
#
|
||||
# The guard is what makes a build that omits --build-arg fail instead of
|
||||
# lie. An unset ARG is an empty string, and an empty string is a
|
||||
# perfectly stable cache key: without the guard the first such build
|
||||
# lints and every one after it on an unchanged tree replays this layer,
|
||||
# executes nothing, and still exits 0. Failed steps are never cached, so
|
||||
# the guard fails on EVERY invocation rather than once. Do not give
|
||||
# CHECK_EPOCH a default value; a default would satisfy the guard with a
|
||||
# constant and restore the hole.
|
||||
ARG CHECK_EPOCH
|
||||
RUN [ -n "$CHECK_EPOCH" ] || exit 1
|
||||
|
||||
# Validate .golangci.yml before linting with it.
|
||||
#
|
||||
# This is not belt-and-braces; it closes a hole that `golangci-lint run`
|
||||
# leaves wide open. `run` rejects YAML it cannot PARSE, but it silently
|
||||
# IGNORES an unknown top-level KEY. Renaming `linters:` to `linterz:` --
|
||||
# one character -- discards `default: all`, the whole disable list and
|
||||
# every threshold, leaves only golangci-lint's small default linter set
|
||||
# running, and exits 0 reporting `0 issues.` on a tree the real config
|
||||
# fails. Demonstrated on this repo at this pin, recorded on
|
||||
# https://git.eeqj.de/sneak/vaultik/pulls/114: with a planted
|
||||
# over-length line, `script/lint` exits 1 naming the `revive` finding
|
||||
# with `linters:` and exits 0 with `linterz:`. A set-but-ineffective
|
||||
# config quietly falling back to defaults is precisely the false-green
|
||||
# class this gate exists to eliminate, so it must not sit in the gate's
|
||||
# own configuration.
|
||||
#
|
||||
# `config verify` catches it, and it does so OFFLINE at this pinned
|
||||
# version -- verified, not assumed. Under `docker run --network none`
|
||||
# against the pinned digest it exits 0 on this repo's config and exits 3
|
||||
# on the `linterz:` variant with `additional properties 'linterz' not
|
||||
# allowed`. An earlier revision of this file asserted the opposite, that
|
||||
# the schema is fetched over live HTTPS from an unpinned URL, and used
|
||||
# that to justify omitting this line. That claim was false at v2.12.2;
|
||||
# the schema is embedded. If a future bump reintroduces a network fetch
|
||||
# the failure is loud and this comment is where to record it.
|
||||
#
|
||||
# It is keyed on CHECK_EPOCH, like the lint run below, so it executes on
|
||||
# every invocation. Content-addressing alone would arguably be enough --
|
||||
# .golangci.yml arrives through `COPY . .`, so a cache hit here implies
|
||||
# a byte-identical config was validated when the layer really ran. That
|
||||
# argument is exactly the one that would also excuse caching the lint
|
||||
# layer, and this repo has ruled it insufficient: a cached check layer
|
||||
# checks nothing, and the cost of being wrong is silent. Forcing it costs
|
||||
# milliseconds and puts the epoch in the log, where a reader can see that
|
||||
# this validation ran rather than being replayed.
|
||||
RUN echo "check epoch: ${CHECK_EPOCH}" && \
|
||||
golangci-lint config verify --config .golangci.yml
|
||||
|
||||
RUN echo "check epoch: ${CHECK_EPOCH}" && \
|
||||
golangci-lint run --config .golangci.yml ./...
|
||||
@@ -87,10 +87,10 @@ clean:
|
||||
go clean
|
||||
|
||||
# Install dependencies. The linter is deliberately not installed here:
|
||||
# script/lint runs the digest-pinned golangci-lint image declared by the
|
||||
# Dockerfile's lint stage, which is the single source of truth for the
|
||||
# linter version. A second, separately pinned copy on PATH could drift
|
||||
# from it and make a local `make lint` disagree with CI.
|
||||
# script/lint lints by building Dockerfile.lint, whose FROM line is the
|
||||
# single source of truth for the linter version. A second, separately
|
||||
# pinned copy on PATH could drift from it and make a local `make lint`
|
||||
# disagree with CI.
|
||||
deps:
|
||||
go mod download
|
||||
|
||||
|
||||
@@ -84,6 +84,57 @@ VAULTIK_AGE_SECRET_KEY='AGE-SECRET-KEY-...' vaultik snapshot restore <snapshot-i
|
||||
# 0 3 * * * vaultik snapshot create --cron --prune --keep-newer-than 4w
|
||||
```
|
||||
|
||||
## restoring on another machine
|
||||
|
||||
Restoring on a host that never ran the backup — a replacement machine
|
||||
after the original is gone — is the case vaultik is built for. That host
|
||||
needs only three things: the `vaultik` binary, the age **private** key,
|
||||
and the storage credentials for the destination. It does **not** need the
|
||||
local index, the original config file, or the original hostname.
|
||||
|
||||
```sh
|
||||
# install
|
||||
go install sneak.berlin/go/vaultik/cmd/vaultik@latest
|
||||
|
||||
# create a config and point it at the ORIGINAL backup destination
|
||||
vaultik config init
|
||||
vaultik config set storage_url "s3://bucket/prefix?endpoint=https://s3.example.com"
|
||||
vaultik config set s3.access_key_id "..."
|
||||
vaultik config set s3.secret_access_key "..."
|
||||
|
||||
# see what is on the destination store
|
||||
vaultik snapshot list
|
||||
```
|
||||
|
||||
`snapshot list` reads the destination store without the private key. A
|
||||
snapshot that is not in this host's (empty) local index is shown as
|
||||
remote-only: its row is identified by `<remote only:...>` rather than by
|
||||
a `hostname_name_timestamp` name, because the name lives only in the
|
||||
local index and the encrypted database and cannot be recovered from the
|
||||
store. Its timestamp and compressed size are real. (See the `snapshot
|
||||
list` description under [command details](#command-details) for the full
|
||||
explanation.)
|
||||
|
||||
Use that remote key — the hex printed inside `<remote only:...>`, or the
|
||||
full `remote_key` from `snapshot list --json` — to restore and verify:
|
||||
|
||||
```sh
|
||||
# restore everything to /tmp/restored, then check every restored file's
|
||||
# chunk hashes
|
||||
VAULTIK_AGE_SECRET_KEY='AGE-SECRET-KEY-...' \
|
||||
vaultik snapshot restore --verify <remote-key> /tmp/restored
|
||||
|
||||
# optionally, deep-verify the snapshot against the store (downloads and
|
||||
# cryptographically checks every blob)
|
||||
VAULTIK_AGE_SECRET_KEY='AGE-SECRET-KEY-...' \
|
||||
vaultik snapshot verify --deep <remote-key>
|
||||
```
|
||||
|
||||
`age_recipients` (the public key) is not needed to restore — only the
|
||||
private key in `VAULTIK_AGE_SECRET_KEY`. Both the abbreviated key printed
|
||||
in the table and the full 64-character key from `--json` are accepted; a
|
||||
leading part of the key is enough as long as it is unambiguous.
|
||||
|
||||
---
|
||||
|
||||
## cli
|
||||
@@ -96,10 +147,10 @@ vaultik [--config <path>] config edit
|
||||
vaultik [--config <path>] config get <key>
|
||||
vaultik [--config <path>] config set <key> <value>
|
||||
vaultik [--config <path>] snapshot create [snapshot-names...] [--cron] [--prune] [--keep-newer-than <duration>]
|
||||
vaultik [--config <path>] snapshot list [--json]
|
||||
vaultik [--config <path>] snapshot list [--json] # alias: ls
|
||||
vaultik [--config <path>] snapshot verify <snapshot-id> [--deep] [--json]
|
||||
vaultik [--config <path>] snapshot purge [--keep-latest | --older-than <duration>] [--snapshot <name>...] [--force]
|
||||
vaultik [--config <path>] snapshot remove <snapshot-id> [--dry-run] [--force] [--local-only] [--json]
|
||||
vaultik [--config <path>] snapshot remove <snapshot-id> [--dry-run] [--force] [--local-only] [--json] # alias: rm
|
||||
vaultik [--config <path>] snapshot restore <snapshot-id> <target-dir> [paths...] [--verify]
|
||||
vaultik [--config <path>] prune [--force] [--json]
|
||||
vaultik [--config <path>] info
|
||||
@@ -116,7 +167,24 @@ vaultik version
|
||||
* `--verbose`, `-v`: Enable verbose output (on stderr — see below)
|
||||
* `--debug`: Enable debug output (on stderr — see below)
|
||||
* `--quiet`, `-q`: Suppress non-error output (also suppresses startup banner)
|
||||
* `--skip-errors`: Continue past per-file errors instead of aborting (applies to `snapshot create` and `restore`)
|
||||
* `--skip-errors`: Skip files that cannot be read when creating a snapshot, or that cannot be restored when restoring, instead of aborting. Packing and storage errors (which would leave a chunk recorded but not stored) still abort the run.
|
||||
|
||||
### locking
|
||||
|
||||
Commands that write persistent state — `snapshot create`, `snapshot
|
||||
purge`, `snapshot remove`, `prune`, and `remote nuke` — take a
|
||||
process-wide lock at `$XDG_DATA_HOME/vaultik/vaultik.pid`
|
||||
(`~/.local/share/vaultik/vaultik.pid` on Linux) for the whole run. Only
|
||||
one of them runs at a time: a second one exits immediately with an
|
||||
"already running" error rather than waiting, so two writers can never
|
||||
corrupt the local index or the destination store.
|
||||
|
||||
Read-only commands — `info`, `snapshot list`, `snapshot verify`, and
|
||||
`remote info` — do not take the lock and are never blocked, so they run
|
||||
even while a backup is in progress. `snapshot restore` does not take the
|
||||
lock either: it writes only to the target directory you name, not the
|
||||
local index or the destination store. `config`, `database delete`,
|
||||
`completion`, and `version` do not take the lock.
|
||||
|
||||
### stdout and stderr
|
||||
|
||||
@@ -152,6 +220,8 @@ and `vaultik prune --json | jq .` both work as written.
|
||||
* `VAULTIK_AGE_SECRET_KEY`: Age private key for decryption (required for `snapshot restore` and `snapshot verify --deep`)
|
||||
* `VAULTIK_CONFIG`: Path to config file (overridden by `--config`)
|
||||
* `VAULTIK_INDEX_PATH`: Override local SQLite index path
|
||||
* `VAULTIK_CPUPROFILE`: Write a CPU profile to this path for the duration of the run (development/debugging)
|
||||
* `VAULTIK_MEMPROFILE`: Write a heap profile to this path when the run exits (development/debugging)
|
||||
|
||||
### shell completion
|
||||
|
||||
@@ -245,13 +315,16 @@ local index alone, and still exits zero.
|
||||
* Default (shallow): checks that all blobs referenced in the manifest exist in storage
|
||||
* `--deep`: Downloads and decrypts each blob, verifies chunk hashes against the
|
||||
encrypted metadata database
|
||||
* Accepts the same identifiers as `snapshot restore`: a snapshot ID, or a
|
||||
remote-only snapshot's remote key (or an unambiguous leading part of it)
|
||||
* `--json`: Output results as JSON
|
||||
|
||||
**`snapshot purge`**: Remove old snapshots based on criteria. Retention is
|
||||
per-snapshot-name (`--keep-latest` keeps the latest of each name, not the
|
||||
latest globally).
|
||||
* `--keep-latest`: Keep only the most recent snapshot of each name
|
||||
* `--older-than <duration>`: Remove snapshots older than duration (e.g. `30d`, `6m`, `1y`)
|
||||
* `--older-than <duration>`: Remove snapshots older than duration (e.g. `30d`,
|
||||
`4w`, `6mo`, `1y`; `m` is minutes, `mo` is months)
|
||||
* `--snapshot <name>`: Restrict to specific snapshot names (repeat for multiple)
|
||||
* `--force`: Skip confirmation prompt
|
||||
|
||||
@@ -274,6 +347,10 @@ on the destination in one go, use `vaultik remote nuke --force`.
|
||||
|
||||
**`snapshot restore`**: Restore files from a backup snapshot.
|
||||
* Requires `VAULTIK_AGE_SECRET_KEY` environment variable
|
||||
* Accepts a snapshot ID, or — for a snapshot only on the destination
|
||||
store — its remote key (or an unambiguous leading part of it) as shown
|
||||
by `snapshot list`. See
|
||||
[restoring on another machine](#restoring-on-another-machine).
|
||||
* Optional path arguments to restore specific files/directories (default: all)
|
||||
* Preserves file permissions, timestamps, ownership (ownership requires root),
|
||||
symlinks, and empty directories
|
||||
@@ -337,6 +414,10 @@ both are set.
|
||||
|
||||
## architecture
|
||||
|
||||
For an implementation-level view of the internals — the data model, the
|
||||
`fx` dependency-injection wiring, and the scanner — see
|
||||
[`ARCHITECTURE.md`](ARCHITECTURE.md).
|
||||
|
||||
### remote storage layout
|
||||
|
||||
```
|
||||
@@ -344,7 +425,7 @@ both are set.
|
||||
├── blobs/
|
||||
│ └── <aa>/<bb>/<full_blob_hash>
|
||||
└── metadata/
|
||||
└── <snapshot_id>/
|
||||
└── <remote-key>/
|
||||
├── db.zst.age # Encrypted binary SQLite database
|
||||
└── manifest.json.zst # Unencrypted blob list (for pruning)
|
||||
```
|
||||
@@ -355,8 +436,18 @@ both are set.
|
||||
* `manifest.json.zst` is an unencrypted compressed JSON blob list, enabling
|
||||
pruning without the private key
|
||||
|
||||
Snapshot IDs follow the format `<hostname>_<snapshot-name>_<RFC3339-timestamp>`
|
||||
(e.g. `server1_home_2025-06-01T12:00:00Z`).
|
||||
Snapshot IDs follow the human-readable format
|
||||
`<hostname>_<snapshot-name>_<RFC3339-timestamp>` (e.g.
|
||||
`server1_home_2025-06-01T12:00:00Z`), but this ID is never written to the
|
||||
destination store in plaintext. Each snapshot's metadata directory is named
|
||||
with its `<remote-key>`, a one-way double SHA-256 hash of the ID, so a listing
|
||||
of the store reveals no hostname or snapshot name. The backup time is not
|
||||
hidden: manifest.json.zst carries a plaintext timestamp, and object
|
||||
modification times are visible at the storage layer regardless. For example,
|
||||
`server1_home_2025-06-01T12:00:00Z` is stored under
|
||||
`metadata/17f97bcde958748af076b926af59823943db59e80ce7170b40f124dfa28f64aa/`.
|
||||
See [docs/REPOSTRUCTURE.md](docs/REPOSTRUCTURE.md#remote-key-derivation) for the
|
||||
derivation.
|
||||
|
||||
### data flow
|
||||
|
||||
@@ -373,7 +464,7 @@ Snapshot IDs follow the format `<hostname>_<snapshot-name>_<RFC3339-timestamp>`
|
||||
|
||||
**restore:**
|
||||
|
||||
1. Download and decrypt `metadata/<snapshot_id>/db.zst.age`
|
||||
1. Download and decrypt `metadata/<remote-key>/db.zst.age`
|
||||
2. Open the binary SQLite database
|
||||
3. Query files (optionally filtered by paths)
|
||||
4. Download and decrypt required blobs
|
||||
@@ -404,25 +495,30 @@ Snapshot IDs follow the format `<hostname>_<snapshot-name>_<RFC3339-timestamp>`
|
||||
|
||||
### compression
|
||||
|
||||
* zstd compression at configurable level (1-19, default 3)
|
||||
* zstd compression at configurable level (1-19, default 3). The level is
|
||||
accepted as 1-19 but maps onto zstd's four internal speed presets:
|
||||
1-2 fastest, 3-5 default, 6-9 better, 10-19 best. Levels within the
|
||||
same band compress identically.
|
||||
* Applied before encryption at the blob level
|
||||
|
||||
---
|
||||
|
||||
## configuration reference
|
||||
|
||||
Run `vaultik config init` to generate a fully commented config file.
|
||||
Key fields:
|
||||
Run `vaultik config init` to generate a fully commented config file; a
|
||||
complete annotated example also lives in
|
||||
[`config.example.yml`](config.example.yml). Key fields:
|
||||
|
||||
| Field | Default | Description |
|
||||
|-------|---------|-------------|
|
||||
| `age_recipients` | (required) | Age public keys for encryption |
|
||||
| `age_secret_key` | (unset) | Age private key for decryption (`snapshot restore`, `snapshot verify --deep`). Setting it in the config file places the private key on the backed-up host, defeating the public-key-only design (see "why" above). Prefer the `VAULTIK_AGE_SECRET_KEY` environment variable, supplied only on the machine you restore from. |
|
||||
| `snapshots` | (required) | Named snapshot definitions with paths and excludes |
|
||||
| `storage_url` | | Storage backend URL (`s3://`, `file://`, `rclone://`) |
|
||||
| `s3.*` | | Legacy S3 configuration (endpoint, bucket, credentials) |
|
||||
| `exclude` | | Global exclude patterns (applied to all snapshots) |
|
||||
| `chunk_size` | `10MB` | Average chunk size for content-defined chunking |
|
||||
| `blob_size_limit` | `10GB` | Maximum blob size before splitting |
|
||||
| `blob_size_limit` | `10GB` | Maximum blob size before splitting. Must be at least four times `chunk_size` (the largest chunk the chunker can emit), otherwise a single-chunk blob could exceed the limit |
|
||||
| `compression_level` | `3` | zstd compression level (1-19) |
|
||||
| `hostname` | system hostname | Hostname used in snapshot IDs |
|
||||
| `index_path` | platform data dir | Local SQLite index path |
|
||||
@@ -446,9 +542,13 @@ Key fields:
|
||||
sequentially. Restore speed is bound by single-stream throughput.
|
||||
* **Device nodes, named pipes, and sockets are silently skipped.** Only
|
||||
regular files, directories, and symlinks are backed up.
|
||||
* **No database migrations.** If the local SQLite schema changes between
|
||||
versions, delete the local database (`vaultik database delete`) and run
|
||||
a full backup. Remote storage is unaffected.
|
||||
* **No upgrade path between versions.** There is no supported way to carry
|
||||
an existing local index across a schema change; if the local SQLite
|
||||
schema changes between versions, delete the local database (`vaultik
|
||||
database delete`) and run a full backup. Remote storage is unaffected.
|
||||
(The binary does embed numbered schema files and a `schema_migrations`
|
||||
table to bootstrap a fresh database — see [`docs/DATAMODEL.md`](docs/DATAMODEL.md)
|
||||
— but that is not an upgrade path.)
|
||||
* **Files that change during backup may be inconsistent.** There is no
|
||||
filesystem snapshot or freeze. If a file is modified between the scan
|
||||
and chunk phases, the backed-up copy may reflect a partial write.
|
||||
@@ -514,14 +614,12 @@ priority.
|
||||
|
||||
### infrastructure
|
||||
|
||||
* **Cross-machine restore documentation.** The "restore from
|
||||
another host" workflow works but isn't documented as a
|
||||
first-class operation in this README. Worth a dedicated section
|
||||
once it's settled.
|
||||
* **Schema migrations.** Currently nonexistent — pre-1.0 schema
|
||||
changes are handled by `vaultik database delete` plus a full
|
||||
re-scan. Post-1.0 we'll need a migration story to keep existing
|
||||
index databases usable across upgrades.
|
||||
* **Cross-version schema upgrades.** There is no upgrade path between
|
||||
released versions — pre-1.0 schema changes are handled by `vaultik
|
||||
database delete` plus a full re-scan (see
|
||||
[`docs/DATAMODEL.md`](docs/DATAMODEL.md)). Post-1.0 we'll need a
|
||||
migration story to keep existing index databases usable across
|
||||
upgrades.
|
||||
* **Storage backend coverage tests.** S3, file://, and rclone://
|
||||
all share the Storer interface but the rclone path is the least
|
||||
exercised in CI.
|
||||
@@ -530,9 +628,17 @@ priority.
|
||||
|
||||
## output style
|
||||
|
||||
All user-facing output goes through helpers in `internal/ui` and conforms
|
||||
to a uniform style. Color is enabled when stdout is a TTY and the
|
||||
`NO_COLOR` environment variable is unset (https://no-color.org/).
|
||||
The operational narration of the long-running commands — the Begin,
|
||||
Complete, Progress, and status lines of `snapshot create`, `prune`,
|
||||
`snapshot restore`, and the like — goes through helpers in `internal/ui`
|
||||
and conforms to the uniform style below. Some commands instead write
|
||||
plain text straight to stdout (`version`, `info`, `config`, the
|
||||
`database delete` prompt, and the `snapshot list` table); that output is
|
||||
unstyled and does not honor `--quiet`. Routing it through `internal/ui`
|
||||
is tracked in
|
||||
[issue #149](https://git.eeqj.de/sneak/vaultik/issues/149). Color is
|
||||
enabled when stdout is a TTY and the `NO_COLOR` environment variable is
|
||||
unset (https://no-color.org/).
|
||||
|
||||
`internal/ui` writes to stdout; it is the output the user asked for.
|
||||
Structured log records are a different thing and go through
|
||||
@@ -598,11 +704,11 @@ regardless of color setting (emoji are not color).
|
||||
|
||||
* Go 1.26 or later
|
||||
* Docker, with a reachable daemon, to lint, check, or commit:
|
||||
`script/lint` runs the digest-pinned `golangci-lint` image declared by
|
||||
the `Dockerfile` lint stage, and `make check` and the pre-commit hook
|
||||
both run it. A `golangci-lint` installed on `PATH` is not a substitute
|
||||
and is never used on a host, whatever its version.
|
||||
* `sqlite3` CLI, which the test suite shells out to
|
||||
`script/lint` lints by building `Dockerfile.lint`, which runs the
|
||||
digest-pinned `golangci-lint` image as a build step, and `make check`
|
||||
and the pre-commit hook both run it. A `golangci-lint` installed on
|
||||
`PATH` is not a substitute and is never used on a host, whatever its
|
||||
version.
|
||||
* S3-compatible object storage (or local filesystem, or rclone remote)
|
||||
|
||||
## development workflow
|
||||
@@ -633,8 +739,8 @@ standard: normalized scripts in `script/` are the entrypoints for the
|
||||
development workflow, and the Makefile targets are thin shims that call
|
||||
them. We provide:
|
||||
|
||||
* `script/bootstrap` — install all development dependencies (go, sqlite3,
|
||||
Go module download). It deliberately does not install `golangci-lint`;
|
||||
* `script/bootstrap` — install all development dependencies (go, Go
|
||||
module download). It deliberately does not install `golangci-lint`;
|
||||
see `script/lint` below.
|
||||
* `script/setup` — make a fresh clone ready for development: runs
|
||||
`script/bootstrap`, then `script/install-precommit`
|
||||
@@ -648,6 +754,14 @@ them. We provide:
|
||||
called by `script/bootstrap`; the release workflow calls it directly
|
||||
because it needs `goreleaser` but not the Docker daemon
|
||||
`script/bootstrap` insists on.
|
||||
* `script/install-go` — install the Go toolchain named by `go.mod`'s
|
||||
`go` directive into `.tool/go` from a sha256-verified `go.dev`
|
||||
archive, and put it on `PATH`. Idempotent. Called only by the release
|
||||
workflow, which needs a host Go for `goreleaser` to shell out to;
|
||||
nothing else on the release runner does. `actions/setup-go` is not
|
||||
used because it verifies the downloaded toolchain against no value in
|
||||
this repo. Bumping Go edits `go.mod`, the checksum in this script, and
|
||||
the `Dockerfile` `golang` digest together.
|
||||
* `script/release` — cross-compile and publish the release artifacts
|
||||
with the pinned `goreleaser`. Refuses a `goreleaser` on `PATH` whose
|
||||
version is not the pinned one, on the same reasoning as `script/lint`.
|
||||
@@ -671,46 +785,71 @@ them. We provide:
|
||||
diverges from the 30s `REPO_POLICIES.md` mandates; the reasoning is in
|
||||
the comment in the script, and issue #101 proposes amending the policy
|
||||
text.
|
||||
* `script/lint` — run `golangci-lint run ./...` at the exact version CI
|
||||
uses, by running the digest-pinned `golangci-lint` image declared by
|
||||
the `Dockerfile` lint stage (requires Docker; it fails loudly rather
|
||||
than falling back to a differently versioned `golangci-lint` on
|
||||
`PATH`). That `FROM` line is the single source of truth for the linter
|
||||
version — bump it there and nowhere else.
|
||||
* `script/lint` — lint by building `Dockerfile.lint`, which runs
|
||||
`golangci-lint run --config .golangci.yml ./...` as a build step
|
||||
inside the digest-pinned `golangci-lint` image, so a successful build
|
||||
*is* a clean lint. Nothing lints on the host, at any version, ever;
|
||||
the script requires Docker and fails loudly rather than falling back
|
||||
to a `golangci-lint` on `PATH`. That `FROM` line is the single source
|
||||
of truth for the linter version — bump it there and nowhere else.
|
||||
|
||||
It takes no arguments, because a build step has no command line to
|
||||
pass flags to, and it passes a fresh `--build-arg CHECK_EPOCH` on
|
||||
every invocation so the lint layer cannot be replayed from cache (see
|
||||
`script/cibuild` below for what that mechanism defends against). To
|
||||
watch the linter execute, run it as
|
||||
`BUILDKIT_PROGRESS=plain script/lint` and check that the lint layer
|
||||
says `RUN … golangci-lint` rather than `CACHED`.
|
||||
|
||||
One container per run means one lint cache and one `golangci-lint`
|
||||
lock per run, both private to it and discarded with it, so concurrent
|
||||
runs on one host cannot contaminate or block each other.
|
||||
* `script/lint-fix` — apply the linter's autofixes (rewrites files),
|
||||
using the same pinned linter
|
||||
using the same pinned image, parsed out of `Dockerfile.lint`. It
|
||||
cannot be a build step, because fixes have to land in the worktree, so
|
||||
it bind-mounts the tree into a `docker run` and therefore needs a
|
||||
*local* daemon. It is a developer convenience and never a gate: no
|
||||
gate reads its exit status. Run `make lint` afterwards to find out
|
||||
whether the tree is clean.
|
||||
* `script/fmt` — format all code (writes)
|
||||
* `script/fmt-check` — check formatting (read-only)
|
||||
* `script/check` — run `script/test`, `script/lint`, and
|
||||
`script/fmt-check`. This is authoritative *because* `script/lint` uses
|
||||
the pinned linter: a local `make check` and CI cannot disagree about
|
||||
lint findings.
|
||||
`script/fmt-check`. This is authoritative *because* `script/lint`
|
||||
builds `Dockerfile.lint`: a local `make check` and CI cannot disagree
|
||||
about lint findings.
|
||||
* `script/docker` — build the Docker image tagged via
|
||||
`script/projectname`. Passes a fresh `--build-arg CHECK_EPOCH` for the
|
||||
same reason `script/cibuild` does, so a local image build cannot be
|
||||
green on checks it replayed from cache.
|
||||
* `script/cibuild` — CI entrypoint: `docker build` (the `Dockerfile`
|
||||
runs `make fmt-check` and `make lint` in its lint stage and `make
|
||||
test` in its builder stage). This is the full CI-equivalent gate — it
|
||||
runs the checks in the same containers CI does, from a clean copy of
|
||||
the tree, so it also catches anything that depends on host state. It
|
||||
passes a fresh `--build-arg CHECK_EPOCH`, unique per invocation, which
|
||||
the `Dockerfile` declares immediately above the check `RUN`s in both
|
||||
stages and expands into each check command. Those layers are keyed on
|
||||
green on checks it replayed from cache. It builds the *product* image
|
||||
only, and the product `Dockerfile` has no lint stage, so it does not
|
||||
lint: a green here means formatted, tested, and it compiles.
|
||||
* `script/cibuild` — CI entrypoint, and the full gate. Two builds, in
|
||||
order: `Dockerfile.lint` (the linter, as a build step) and then
|
||||
`Dockerfile` (`make fmt-check` and `make test` in its builder stage,
|
||||
then the product image). Either failing fails the script. It runs the
|
||||
checks in the same containers CI does, from a clean copy of the tree,
|
||||
so it also catches anything that depends on host state.
|
||||
`.gitea/workflows/check.yml` runs it on every push to `main` and
|
||||
`next` and on every pull request against either.
|
||||
|
||||
It passes a fresh `--build-arg CHECK_EPOCH` to each build, unique per
|
||||
invocation, which both files declare immediately above their check
|
||||
`RUN`s and expand into each check command. Those layers are keyed on
|
||||
that value, so a new value re-runs them even on a byte-identical tree,
|
||||
and a green from this script means the checks executed. Dependency and
|
||||
module layers sit above the `ARG` and still cache, so a build is not
|
||||
cold.
|
||||
|
||||
A build that supplies no `CHECK_EPOCH` — a bare `docker build .` —
|
||||
fails rather than lying. An unset `ARG` is an empty string and an
|
||||
empty string is a stable cache key, so without a guard such a build
|
||||
would serve all three check layers from cache, execute nothing, and
|
||||
still exit 0. Each check stage therefore asserts the value is
|
||||
non-empty before running anything, and because failed steps are never
|
||||
cached that assertion fires on every invocation rather than once. Use
|
||||
`script/cibuild` (or `script/docker`, which passes the same arg); a
|
||||
bare `docker build .` is now a loud error.
|
||||
A build that supplies no `CHECK_EPOCH` — a bare `docker build .` or
|
||||
`docker build -f Dockerfile.lint .` — fails rather than lying. An
|
||||
unset `ARG` is an empty string and an empty string is a stable cache
|
||||
key, so without a guard such a build would serve every check layer
|
||||
from cache, execute nothing, and still exit 0. Each file therefore
|
||||
asserts the value is non-empty before running anything, and because
|
||||
failed steps are never cached that assertion fires on every
|
||||
invocation rather than once. Use `script/lint`, `script/docker` or
|
||||
`script/cibuild`, which pass the arg; a bare `docker build` is a loud
|
||||
error.
|
||||
* `script/precommit` — pre-commit gate: `go mod tidy` + `go fmt` (must
|
||||
not change files), then `script/check`
|
||||
* `script/install-precommit` — install the git pre-commit hook that
|
||||
|
||||
@@ -25,6 +25,175 @@ release" is exactly the contradiction
|
||||
|
||||
# Completed Steps
|
||||
|
||||
- 2026-09-21: Stopped an interrupted blob upload from making a later
|
||||
backup deduplicate against data that was never stored
|
||||
([issue #148](https://git.eeqj.de/sneak/vaultik/issues/148)). The
|
||||
packer commits a blob's `chunks`, `blob_chunks`, and `blobs` rows
|
||||
before the upload is attempted, so a failed upload left chunk rows
|
||||
behind and the next run skipped re-uploading them, producing a
|
||||
snapshot that reported success but could not be restored. A run now
|
||||
deduplicates only against chunks held by a blob whose `uploaded_ts` is
|
||||
set, and at startup drops any un-uploaded blob rows (and the chunks
|
||||
they orphan) so the affected data is re-chunked and re-uploaded. Blobs
|
||||
recorded with no remote backend are marked uploaded so this invariant
|
||||
holds uniformly.
|
||||
|
||||
- 2026-09-22: Made restore refuse any snapshot path that would write
|
||||
outside the target directory
|
||||
([issue #154](https://git.eeqj.de/sneak/vaultik/issues/154)).
|
||||
`restoreFile` and `verifyRestoredFiles` joined the stored path onto the
|
||||
target with no containment check, so a `..` segment or an absolute path
|
||||
escaped the target and a restored symlink could redirect a later child
|
||||
write anywhere on disk. Every stored path is now rejected unless
|
||||
`filepath.IsLocal` accepts it with the leading separator removed, and
|
||||
each existing ancestor directory below the target is `Lstat`ed to refuse
|
||||
descending through a symlink; honest symlinks pointing outside the tree
|
||||
are still written verbatim. age decryption proves a snapshot is
|
||||
readable, not honest, and restore usually runs as root.
|
||||
|
||||
- 2026-09-21: Stopped `--json` from silencing stderr diagnostics
|
||||
([issue #112](https://git.eeqj.de/sneak/vaultik/issues/112)). `--json`
|
||||
used to be folded into `Quiet`, which pinned the log level to `WARN`,
|
||||
so `prune --json` gave a machine consumer no record of the local index
|
||||
rows it deleted even under `--verbose`. `--json` now quiets only the
|
||||
stdout UI (the JSON document must stay clean, per
|
||||
[issue #108](https://git.eeqj.de/sneak/vaultik/issues/108)); the stderr
|
||||
log level follows `--verbose`/`--debug` again. The coupling was
|
||||
removed the same way for `snapshot verify`, `snapshot remove`, and
|
||||
`remote info`, which carried it for the same outdated reason.
|
||||
|
||||
- 2026-09-21: Stopped `prune` from reporting a failed row count as 0
|
||||
([issue #96](https://git.eeqj.de/sneak/vaultik/issues/96)). The seven
|
||||
`getTableCount` reads in `PruneDatabase` discarded their error, so a
|
||||
query that could not run became a plausible `0` and the before/after
|
||||
delta computed from it looked like real work. Each read now logs at
|
||||
warn on failure and renders as `unknown`, never `0`, so an empty table
|
||||
is distinguishable from one that could not be queried. The counts have
|
||||
no `--json` representation — under `--json` the summary is suppressed
|
||||
entirely — so nothing there can show a false `0`.
|
||||
|
||||
- 2026-09-21: Made the s3 storage backend report a missing object as
|
||||
`storage.ErrNotFound`, like the `file` and `rclone` backends and as the
|
||||
`Storer` interface documents. `S3Storer.Get` and `Stat` returned the raw
|
||||
AWS SDK error, so `errors.Is(err, storage.ErrNotFound)` was false on s3
|
||||
and callers branched differently per backend. Added a small `s3.IsNotFound`
|
||||
helper (reused by `HeadObject`) and a test that a missing key maps to
|
||||
`ErrNotFound`
|
||||
([issue #129](https://git.eeqj.de/sneak/vaultik/issues/129)).
|
||||
- 2026-09-21: Fixed `verify --deep` reporting healthy snapshots as
|
||||
corrupt. Its final blob-integrity check hashed the encrypted
|
||||
downloaded bytes with a single SHA256 and compared that to the blob
|
||||
ID, which is the double SHA256 of the plaintext, so the two could
|
||||
never match. It now hashes the decompressed plaintext and compares the
|
||||
double SHA256. Added a test that backs up a real snapshot, deep-verifies
|
||||
it, then flips a byte in one stored blob and confirms deep verification
|
||||
then fails
|
||||
([issue #131](https://git.eeqj.de/sneak/vaultik/issues/131)).
|
||||
|
||||
- 2026-09-21: Made `snapshot create` VACUUM the per-snapshot metadata
|
||||
database through the `modernc.org/sqlite` driver instead of shelling
|
||||
out to the external `sqlite` command-line binary (issue #120). A
|
||||
backup no longer needs that binary on `PATH`, so `make check` passes
|
||||
on a stock `go install` host; `script/bootstrap` and the `Dockerfile`
|
||||
(both the test-build and the shipped runtime stage) no longer install
|
||||
it, and a new test asserts the uploaded database keeps no pages from
|
||||
deleted rows. Dropped the now-false note on the 2026-08-07 entry below
|
||||
that said bootstrap installs it.
|
||||
- 2026-09-21: Made `.gitea/workflows/check.yml` run on pushes to `main`
|
||||
and `next` and on pull requests against either, so unit PRs (whose
|
||||
base is `next`) and `next` itself get a CI run instead of relying on a
|
||||
local `make check`
|
||||
([issue #122](https://git.eeqj.de/sneak/vaultik/issues/122)).
|
||||
|
||||
- 2026-09-21: Hash-verified the Go toolchain in the release workflow
|
||||
([issue #105](https://git.eeqj.de/sneak/vaultik/issues/105)). New
|
||||
`script/install-go` downloads the exact `go.dev` archive for `go.mod`'s
|
||||
`go` directive and refuses it unless its sha256 matches a value
|
||||
committed in the script; `.gitea/workflows/release.yml` calls it
|
||||
instead of `actions/setup-go`, which verified the downloaded toolchain
|
||||
against nothing in the repo. `GOTOOLCHAIN: local` on the release step
|
||||
keeps that exact compiler from auto-switching. Bumping Go now touches
|
||||
`go.mod`, the checksum, and the `Dockerfile` `golang` digest together.
|
||||
|
||||
- 2026-09-21: Collapsed the two duration parsers into one and fixed the
|
||||
`--older-than` months example
|
||||
([issue #123](https://git.eeqj.de/sneak/vaultik/issues/123)). Two
|
||||
functions named `parseDuration` existed with different grammars;
|
||||
`snapshot purge --older-than` and `--keep-newer-than` both already went
|
||||
through the one in `internal/vaultik`, while the richer copy in
|
||||
`internal/cli/duration.go` was reachable only from its own test. Kept
|
||||
the live-path parser and deleted the unused one, so no flag's accepted
|
||||
grammar changes. The trap the issue was filed over: `README.md`
|
||||
documented `6m` as the months example for `--older-than`, but `m` is
|
||||
minutes, so the documented command deleted every snapshot older than
|
||||
six minutes on a destructive flag. Corrected the doc to `6mo` and put
|
||||
both flags' help text on one example list that states `m` is minutes
|
||||
and `mo` is months. The surviving parser now rejects negatives, which
|
||||
it previously accepted (`-5h`) or silently made positive (`-5d`).
|
||||
Table-driven tests cover every unit, `6m` as six minutes, `6mo` as 180
|
||||
days, and rejection of a bare number, an unknown unit, and a negative.
|
||||
|
||||
- 2026-08-10: Moved every lint run into its own container, as a build
|
||||
step ([issue #113](https://git.eeqj.de/sneak/vaultik/issues/113)).
|
||||
New root `Dockerfile.lint`, built by `script/lint`, runs
|
||||
`golangci-lint run --config .golangci.yml ./...` as a `RUN`
|
||||
instruction in the digest-pinned `golangci/golangci-lint` image: a
|
||||
successful build of that file *is* a clean lint, and it works even
|
||||
where the daemon is remote and bind mounts are impossible. That
|
||||
`FROM` line is now the only pin of the linter version in the repo.
|
||||
|
||||
This supersedes the per-worktree cache isolation landed for
|
||||
[issue #99](https://git.eeqj.de/sneak/vaultik/issues/99). Isolation
|
||||
fixed cross-worktree contamination but not lock contention — two
|
||||
concurrent runs with entirely separate cache directories still
|
||||
collided. A container per run has its own cache and its own lock, so
|
||||
the whole class is gone, and with it the per-worktree cache
|
||||
machinery, the lock-retry loop, and `script/lint-audit`, which
|
||||
existed to catch replayed findings from a cache that no longer
|
||||
exists. The host lint path went too: no escape hatch, no
|
||||
`VAULTIK_LINT_IN_CONTAINER`, no version detection. Nothing lints on
|
||||
the host at any version.
|
||||
|
||||
A cached build lints nothing, so the same `CHECK_EPOCH` mechanism the
|
||||
product `Dockerfile` already used is what makes the green mean
|
||||
something: `ARG CHECK_EPOCH` with no default below the module layers,
|
||||
a `RUN [ -n "$CHECK_EPOCH" ] || exit 1` guard, the value expanded
|
||||
into each check command, and a fresh `$(date +%s%N)$$` per invocation
|
||||
computed as a bare assignment. `cmd/vaultik/lintdocker_test.go`
|
||||
parses both Dockerfiles and both scripts and fails if any part of
|
||||
that is dropped, because every way of losing it is silent. No test
|
||||
asserts that no script runs the host linter: `script/lint` is the one
|
||||
lint entry point and runs `golangci-lint` only inside the container,
|
||||
and keeping it that way is a review matter, not something a test
|
||||
proves.
|
||||
|
||||
The product `Dockerfile` lost its lint stage rather than gaining a
|
||||
second linter pin: `make lint` is now `docker build`, so the stage
|
||||
would have been docker-in-docker with no daemon, and calling
|
||||
`golangci-lint` directly there would have restored the two-pins drift
|
||||
of [issue #78](https://git.eeqj.de/sneak/vaultik/issues/78).
|
||||
`make fmt-check` moved beside `make test` in the builder stage, and
|
||||
`script/cibuild` now builds `Dockerfile.lint` and then `Dockerfile`,
|
||||
each with its own fresh epoch. Consequence, stated rather than left
|
||||
to be found: `script/docker` builds the product image only and no
|
||||
longer lints; the gates are `script/check` and `script/cibuild`.
|
||||
|
||||
`golangci-lint config verify` runs as its own epoch-keyed layer,
|
||||
above the lint. `golangci-lint run` rejects a config it cannot parse
|
||||
but silently ignores an unknown top-level *key*: renaming `linters:`
|
||||
to `linterz:` discarded `default: all` and every threshold and still
|
||||
exited 0 on a tree the real config fails. `config verify` catches
|
||||
that, and it does so with the network off at this pin — checked under
|
||||
`docker run --network none`, not assumed. An earlier revision omitted
|
||||
it on the claim that it fetches its schema over live HTTPS; that
|
||||
claim was false at v2.12.2.
|
||||
|
||||
`script/lint-fix` is kept, reimplemented as a
|
||||
bind-mounted `docker run` against the image parsed out of
|
||||
`Dockerfile.lint` — it cannot be a build step, because fixes have to
|
||||
land in the worktree — and marked in its header as a developer
|
||||
convenience that no gate reads.
|
||||
|
||||
- 2026-08-09: Finished the `--json` stdout contract and gave `make build`
|
||||
a rule ([issue #108](https://git.eeqj.de/sneak/vaultik/issues/108),
|
||||
[issue #110](https://git.eeqj.de/sneak/vaultik/issues/110)). Two
|
||||
@@ -463,7 +632,7 @@ release" is exactly the contradiction
|
||||
was green was wrong.
|
||||
- 2026-08-07: Added the standard `.golangci.yml` and `.editorconfig`
|
||||
(issue #59); lint findings under the new config are tracked in issue
|
||||
#61. `script/bootstrap` now installs sqlite3 (needed by tests).
|
||||
#61.
|
||||
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints,
|
||||
Makefile shims, README Entrypoints section
|
||||
- 2026-07-02: Consolidated CLI verbs, retired overlapping commands; bound
|
||||
|
||||
@@ -0,0 +1,102 @@
|
||||
package main_test
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
// This file guards the version stamping of the product image (issue
|
||||
// #75). The failure it protects against is silent: the image still
|
||||
// builds and runs, but `vaultik version` inside it reports "commit:
|
||||
// unknown", so an operator cannot tell which source produced a given
|
||||
// backup. .dockerignore excludes .git, so the build cannot derive the
|
||||
// commit itself; the values must be computed on the host and passed in.
|
||||
//
|
||||
// These are parses of the committed files, for the same reason the lint
|
||||
// guards next door are: shelling out to docker would nest a build
|
||||
// inside `make test`. That `vaultik version` in the built image really
|
||||
// prints the host's version is verified by hand and recorded on the
|
||||
// pull request.
|
||||
|
||||
// dockerScript is script/docker, relative to the repository root.
|
||||
const dockerScript = "script/docker"
|
||||
|
||||
// versionArgs are the ldflag targets the build stamps and, matching
|
||||
// them, the build args the host must supply. The names line up so the
|
||||
// same list checks both files.
|
||||
func versionArgs() []string {
|
||||
return []string{"VERSION", "COMMIT", "COMMIT_DATE"}
|
||||
}
|
||||
|
||||
// TestProductDockerfileTakesVersionAsBuildArgs fails unless the build
|
||||
// declares each version arg and stamps it into the binary by ldflag
|
||||
// reference, rather than computing it in the container.
|
||||
func TestProductDockerfileTakesVersionAsBuildArgs(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
found := instructions(t, productDockerfile)
|
||||
|
||||
for _, arg := range versionArgs() {
|
||||
require.GreaterOrEqual(t, indexOf(found, "ARG "+arg), 0,
|
||||
"%s must declare `ARG %s` so the host can pass it in",
|
||||
productDockerfile, arg)
|
||||
|
||||
assertLdflagReferences(t, found, arg)
|
||||
}
|
||||
}
|
||||
|
||||
// TestProductDockerfileDoesNotDeriveVersionItself is the anti-regression
|
||||
// for the original defect: the container ran `git rev-parse`, but .git
|
||||
// is not in the build context, so it always resolved to "unknown". No
|
||||
// git command may reach into a build that cannot see the history.
|
||||
func TestProductDockerfileDoesNotDeriveVersionItself(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
text := instructionText(readRepoFile(t, productDockerfile))
|
||||
|
||||
assert.NotContains(t, text, "git ",
|
||||
"%s must not run git: .git is excluded from the build context, so"+
|
||||
" any value it derives is wrong. Pass version, commit and date"+
|
||||
" in as build args instead.", productDockerfile)
|
||||
}
|
||||
|
||||
// TestDockerScriptComputesVersionOnTheHost fails unless script/docker
|
||||
// derives each value where .git exists and passes it as a build arg,
|
||||
// with VERSION coming from script/version so a Docker build reports the
|
||||
// same string a local build of the same tree would.
|
||||
func TestDockerScriptComputesVersionOnTheHost(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
script := readRepoFile(t, dockerScript)
|
||||
|
||||
for _, arg := range versionArgs() {
|
||||
assert.Contains(t, script, "--build-arg "+arg+"=",
|
||||
"%s must pass --build-arg %s to the build", dockerScript, arg)
|
||||
}
|
||||
|
||||
assert.Contains(t, script, "/version",
|
||||
"%s must take VERSION from script/version, the source of truth"+
|
||||
" shared with the Makefile", dockerScript)
|
||||
}
|
||||
|
||||
// assertLdflagReferences fails unless some build instruction stamps the
|
||||
// named variable from the ARG (a ${arg} reference), not from a value
|
||||
// computed inside the container.
|
||||
func assertLdflagReferences(t *testing.T, found []string, arg string) {
|
||||
t.Helper()
|
||||
|
||||
for _, instruction := range found {
|
||||
if strings.HasPrefix(instruction, "RUN ") &&
|
||||
strings.Contains(instruction, "go build") &&
|
||||
strings.Contains(instruction, "${"+arg+"}") {
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
assert.Fail(t, "version arg is declared but never stamped",
|
||||
"the go build in %s must reference ${%s} in its ldflags, or the"+
|
||||
" arg is passed and discarded", productDockerfile, arg)
|
||||
}
|
||||
@@ -0,0 +1,367 @@
|
||||
package main_test
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
// This file guards the shape of the lint gate. Every property asserted
|
||||
// here is one whose loss is SILENT: the build still exits 0, the gate
|
||||
// still looks green, and nothing was linted or tested.
|
||||
//
|
||||
// The gate is a build step. script/lint builds Dockerfile.lint, which
|
||||
// runs golangci-lint as a RUN instruction, so a successful build is a
|
||||
// clean lint. BuildKit will happily replay that RUN from cache on an
|
||||
// unchanged tree in well under a second, which is why the check layers
|
||||
// are keyed on a CHECK_EPOCH build arg that the calling script
|
||||
// regenerates per invocation, and why an empty value is a hard error
|
||||
// rather than a stable cache key.
|
||||
//
|
||||
// These are parses rather than invocations. Shelling out to docker from
|
||||
// the test suite would nest a build inside `make test`, which itself
|
||||
// runs inside a build in CI. The one property a parse cannot establish
|
||||
// -- that a real finding actually fails the build -- is verified by
|
||||
// hand against a deliberately broken tree, recorded on the pull
|
||||
// request.
|
||||
//
|
||||
// One property is deliberately NOT tested here: that no script runs the
|
||||
// linter on the host. script/lint is the only lint entry point, and it
|
||||
// runs golangci-lint only inside the container; keeping it that way is a
|
||||
// review matter, not something a test in this file establishes.
|
||||
|
||||
// The files under guard, relative to the repository root.
|
||||
const (
|
||||
lintDockerfile = "Dockerfile.lint"
|
||||
productDockerfile = "Dockerfile"
|
||||
lintScript = "script/lint"
|
||||
cibuildScript = "script/cibuild"
|
||||
)
|
||||
|
||||
// linterBinary is the linter's command name, used to locate the
|
||||
// config-verify and lint steps in Dockerfile.lint.
|
||||
const linterBinary = "golangci-lint"
|
||||
|
||||
// checkEpochARG is the declaration, with no default value. A default
|
||||
// would satisfy the non-empty guard with a constant, and a constant is
|
||||
// a stable cache key: the checks would be replayed from cache forever
|
||||
// after the first build.
|
||||
const checkEpochARG = "ARG CHECK_EPOCH"
|
||||
|
||||
// checkEpochGuard is what turns a build that omits --build-arg into a
|
||||
// loud failure instead of a quiet green. Failed steps are never cached,
|
||||
// so it fires on every such invocation rather than once.
|
||||
const checkEpochGuard = `RUN [ -n "$CHECK_EPOCH" ] || exit 1`
|
||||
|
||||
// freshEpoch is the epoch computation the calling scripts must use, as
|
||||
// a bare assignment on its own line. Inline in an argument, a failing
|
||||
// `date` would not abort under `set -eu`; CHECK_EPOCH would become the
|
||||
// empty string, and the guard above would be the only thing standing
|
||||
// between that and a permanently cached green. `$$` is required because
|
||||
// `date +%s` is second-granular and busybox silently drops `%N`, so
|
||||
// without the pid two concurrent runs in one second can collide.
|
||||
const freshEpoch = `epoch="$(date +%s%N)$$"`
|
||||
|
||||
// TestLintDockerfilePinsTheLinterByDigest fails if the lint image stops
|
||||
// being pinned. An unpinned tag makes the gate's verdict depend on
|
||||
// whatever the registry currently serves under that name.
|
||||
func TestLintDockerfilePinsTheLinterByDigest(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
from := ""
|
||||
|
||||
for _, instruction := range instructions(t, lintDockerfile) {
|
||||
if strings.HasPrefix(instruction, "FROM ") {
|
||||
from = instruction
|
||||
|
||||
break
|
||||
}
|
||||
}
|
||||
|
||||
require.NotEmpty(t, from, "%s declares no FROM", lintDockerfile)
|
||||
assert.Contains(t, from, "golangci/golangci-lint",
|
||||
"the lint image must be the golangci-lint image")
|
||||
assert.Contains(t, from, "@sha256:",
|
||||
"the lint image must be pinned by digest, not by tag alone")
|
||||
}
|
||||
|
||||
// TestLintDockerfileCannotBeCachedGreen pins the whole cache-busting
|
||||
// mechanism in the file that lints: the declaration with no default,
|
||||
// the non-empty guard, and the value expanded into the lint command
|
||||
// itself rather than merely declared.
|
||||
func TestLintDockerfileCannotBeCachedGreen(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
found := instructions(t, lintDockerfile)
|
||||
|
||||
argAt := indexOf(found, checkEpochARG)
|
||||
require.GreaterOrEqual(t, argAt, 0,
|
||||
"%s must declare `%s` with no default value",
|
||||
lintDockerfile, checkEpochARG)
|
||||
|
||||
assert.GreaterOrEqual(t, indexOf(found, checkEpochGuard), argAt,
|
||||
"%s must guard against an empty CHECK_EPOCH with `%s`",
|
||||
lintDockerfile, checkEpochGuard)
|
||||
|
||||
assertEpochExpandedInto(t, found[argAt:], "golangci-lint run")
|
||||
|
||||
// Dependency layers must stay above the ARG, or every lint run
|
||||
// re-downloads the module cache and the inner loop becomes
|
||||
// unusable.
|
||||
download := indexOf(found, "RUN go mod download")
|
||||
require.GreaterOrEqual(t, download, 0,
|
||||
"%s must download modules in their own layer", lintDockerfile)
|
||||
assert.Less(t, download, argAt,
|
||||
"`%s` must come after `go mod download` so dependency layers"+
|
||||
" still cache", checkEpochARG)
|
||||
}
|
||||
|
||||
// TestLintDockerfileVerifiesTheLinterConfig guards the validation of
|
||||
// .golangci.yml itself. `golangci-lint run` rejects a config it cannot
|
||||
// parse but silently IGNORES an unknown top-level key, so renaming
|
||||
// `linters:` to `linterz:` discards `default: all` and every threshold
|
||||
// and still exits 0 reporting no issues. `config verify` is what turns
|
||||
// that into a failure, and it has to run BEFORE the lint, or the lint
|
||||
// spends a minute reporting a verdict from a config already known to be
|
||||
// wrong.
|
||||
func TestLintDockerfileVerifiesTheLinterConfig(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
found := instructions(t, lintDockerfile)
|
||||
verify := linterBinary + " config verify"
|
||||
|
||||
verifyAt := indexContaining(found, verify)
|
||||
require.GreaterOrEqual(t, verifyAt, 0,
|
||||
"%s must run `%s --config .golangci.yml`: without it a typo'd"+
|
||||
" top-level key in .golangci.yml is silently ignored and the"+
|
||||
" gate passes with only the default linter set", lintDockerfile,
|
||||
verify)
|
||||
|
||||
runAt := indexContaining(found, linterBinary+" run")
|
||||
require.GreaterOrEqual(t, runAt, 0, "%s must lint", lintDockerfile)
|
||||
assert.Less(t, verifyAt, runAt,
|
||||
"%s must verify the config before linting with it", lintDockerfile)
|
||||
|
||||
// Keyed on the epoch like every other check layer, so it executes
|
||||
// per invocation rather than being replayed. A cached validation
|
||||
// validates nothing.
|
||||
assertEpochExpandedInto(t, found, verify)
|
||||
}
|
||||
|
||||
// TestProductDockerfileCannotBeCachedGreen holds the same line for the
|
||||
// checks that remain in the product image build.
|
||||
func TestProductDockerfileCannotBeCachedGreen(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
found := instructions(t, productDockerfile)
|
||||
|
||||
argAt := indexOf(found, checkEpochARG)
|
||||
require.GreaterOrEqual(t, argAt, 0,
|
||||
"%s must declare `%s` with no default value",
|
||||
productDockerfile, checkEpochARG)
|
||||
|
||||
assert.GreaterOrEqual(t, indexOf(found, checkEpochGuard), argAt,
|
||||
"%s must guard against an empty CHECK_EPOCH", productDockerfile)
|
||||
|
||||
assertEpochExpandedInto(t, found[argAt:], "make fmt-check")
|
||||
assertEpochExpandedInto(t, found[argAt:], "make test")
|
||||
}
|
||||
|
||||
// TestProductDockerfileDoesNotLint records the split deliberately: the
|
||||
// linter lives in Dockerfile.lint and nowhere else, so there is exactly
|
||||
// one digest pinning it. A lint stage reintroduced here would either be
|
||||
// docker-in-docker (`make lint` is now `docker build`) or a second,
|
||||
// independently bumpable pin.
|
||||
func TestProductDockerfileDoesNotLint(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
contents := readRepoFile(t, productDockerfile)
|
||||
|
||||
for _, forbidden := range []string{"golangci", "make lint"} {
|
||||
assert.NotContains(t, instructionText(contents), forbidden,
|
||||
"%s must not lint: the linter is pinned once, in %s",
|
||||
productDockerfile, lintDockerfile)
|
||||
}
|
||||
}
|
||||
|
||||
// TestLintScriptBuildsTheLintDockerfileWithAFreshEpoch is the other
|
||||
// half of the mechanism. The Dockerfile's guard only rejects an EMPTY
|
||||
// epoch; a constant non-empty one would satisfy it and still be served
|
||||
// from cache forever.
|
||||
func TestLintScriptBuildsTheLintDockerfileWithAFreshEpoch(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
script := readRepoFile(t, lintScript)
|
||||
|
||||
assertBareEpochAssignment(t, script, lintScript)
|
||||
assert.Contains(t, script, `--build-arg CHECK_EPOCH="$epoch"`,
|
||||
"%s must pass the fresh epoch to the build", lintScript)
|
||||
assert.Contains(t, script, lintDockerfile,
|
||||
"%s must build %s", lintScript, lintDockerfile)
|
||||
}
|
||||
|
||||
// TestCibuildBuildsBothDockerfilesWithFreshEpochs guards the CI gate:
|
||||
// dropping either build silently removes a whole class of check from
|
||||
// CI while leaving it green.
|
||||
func TestCibuildBuildsBothDockerfilesWithFreshEpochs(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
script := readRepoFile(t, cibuildScript)
|
||||
|
||||
assertBareEpochAssignment(t, script, cibuildScript)
|
||||
assert.Equal(t, 2, strings.Count(script, freshEpoch),
|
||||
"%s must compute a fresh epoch for each of its two builds",
|
||||
cibuildScript)
|
||||
assert.Equal(t, 2,
|
||||
strings.Count(script, `--build-arg CHECK_EPOCH="$epoch"`),
|
||||
"%s must pass a fresh epoch to both builds", cibuildScript)
|
||||
assert.Contains(t, script, "-f Dockerfile.lint",
|
||||
"%s must build %s", cibuildScript, lintDockerfile)
|
||||
}
|
||||
|
||||
// assertEpochExpandedInto fails unless some instruction runs the named
|
||||
// command with the epoch expanded into it. Expansion, not mere
|
||||
// declaration: an ARG that no instruction references is not guaranteed
|
||||
// to key the layer, and the expansion also puts the value in the build
|
||||
// log where a reader can see the layer was keyed fresh.
|
||||
func assertEpochExpandedInto(t *testing.T, found []string, command string) {
|
||||
t.Helper()
|
||||
|
||||
for _, instruction := range found {
|
||||
if !strings.HasPrefix(instruction, "RUN ") {
|
||||
continue
|
||||
}
|
||||
|
||||
if strings.Contains(instruction, command) &&
|
||||
strings.Contains(instruction, "${CHECK_EPOCH}") {
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
assert.Fail(t, "no epoch-keyed layer runs the command",
|
||||
"`%s` must run in a layer that expands ${CHECK_EPOCH}, or it"+
|
||||
" will be replayed from cache without executing", command)
|
||||
}
|
||||
|
||||
// assertBareEpochAssignment fails unless the script computes the epoch
|
||||
// as a bare assignment on its own line.
|
||||
func assertBareEpochAssignment(t *testing.T, script, name string) {
|
||||
t.Helper()
|
||||
|
||||
for line := range strings.SplitSeq(script, "\n") {
|
||||
if strings.TrimSpace(line) == freshEpoch {
|
||||
return
|
||||
}
|
||||
}
|
||||
|
||||
assert.Fail(t, "no bare epoch assignment",
|
||||
"%s must compute `%s` as a bare assignment on its own line, so"+
|
||||
" `set -e` catches a failing date instead of quietly"+
|
||||
" building with an empty epoch", name, freshEpoch)
|
||||
}
|
||||
|
||||
// instructions returns the Dockerfile's instructions, one per element,
|
||||
// with comments and blank lines dropped and continuation lines joined,
|
||||
// so a multi-line RUN is one string.
|
||||
func instructions(t *testing.T, name string) []string {
|
||||
t.Helper()
|
||||
|
||||
return strings.Split(instructionText(readRepoFile(t, name)), "\n")
|
||||
}
|
||||
|
||||
// instructionText is instructions' parse, before splitting: it is also
|
||||
// what a "must not contain" assertion should look at, so that a word
|
||||
// appearing only in a comment is not mistaken for behaviour.
|
||||
func instructionText(contents string) string {
|
||||
var (
|
||||
out []string
|
||||
continued string
|
||||
isContinued bool
|
||||
)
|
||||
|
||||
for line := range strings.SplitSeq(contents, "\n") {
|
||||
trimmed := strings.TrimSpace(line)
|
||||
if !isContinued && (trimmed == "" || strings.HasPrefix(trimmed, "#")) {
|
||||
continue
|
||||
}
|
||||
|
||||
isContinued = strings.HasSuffix(trimmed, `\`)
|
||||
continued += strings.TrimSuffix(trimmed, `\`)
|
||||
|
||||
if isContinued {
|
||||
continue
|
||||
}
|
||||
|
||||
out = append(out, strings.Join(strings.Fields(continued), " "))
|
||||
continued = ""
|
||||
}
|
||||
|
||||
return strings.Join(out, "\n")
|
||||
}
|
||||
|
||||
// indexOf returns the position of the first instruction equal to, or
|
||||
// beginning with, want; -1 if there is none. An `ARG NAME=default`
|
||||
// counts as beginning with `ARG NAME`, so a declared arg is found
|
||||
// whether or not it carries a default.
|
||||
func indexOf(found []string, want string) int {
|
||||
for i, instruction := range found {
|
||||
if instruction == want ||
|
||||
strings.HasPrefix(instruction, want+" ") ||
|
||||
strings.HasPrefix(instruction, want+"=") {
|
||||
return i
|
||||
}
|
||||
}
|
||||
|
||||
return -1
|
||||
}
|
||||
|
||||
// indexContaining returns the position of the first instruction
|
||||
// containing want; -1 if there is none.
|
||||
func indexContaining(found []string, want string) int {
|
||||
for i, instruction := range found {
|
||||
if strings.Contains(instruction, want) {
|
||||
return i
|
||||
}
|
||||
}
|
||||
|
||||
return -1
|
||||
}
|
||||
|
||||
// readRepoFile reads a file by its path relative to the repository
|
||||
// root.
|
||||
func readRepoFile(t *testing.T, name string) string {
|
||||
t.Helper()
|
||||
|
||||
//nolint:gosec // G304: the path is a constant relative to this repo
|
||||
contents, err := os.ReadFile(filepath.Join(repoRoot(t), name))
|
||||
require.NoError(t, err)
|
||||
|
||||
return string(contents)
|
||||
}
|
||||
|
||||
// repoRoot returns the repository root. The test binary runs with its
|
||||
// package directory as the working directory, so the root is found by
|
||||
// walking up until the module file appears.
|
||||
func repoRoot(t *testing.T) string {
|
||||
t.Helper()
|
||||
|
||||
dir, err := os.Getwd()
|
||||
require.NoError(t, err)
|
||||
|
||||
for {
|
||||
_, err = os.Stat(filepath.Join(dir, "go.mod"))
|
||||
if err == nil {
|
||||
return dir
|
||||
}
|
||||
|
||||
parent := filepath.Dir(dir)
|
||||
require.NotEqual(t, dir, parent,
|
||||
"walked to the filesystem root without finding a go.mod")
|
||||
|
||||
dir = parent
|
||||
}
|
||||
}
|
||||
+11
-1
@@ -10,6 +10,16 @@ import (
|
||||
)
|
||||
|
||||
func main() {
|
||||
os.Exit(run())
|
||||
}
|
||||
|
||||
// run sets up optional profiling, runs the CLI, and returns the process
|
||||
// exit code. os.Exit lives in main so it fires only after run's deferred
|
||||
// profile writers have flushed. cli.Entry returns a status code rather
|
||||
// than calling os.Exit itself: an os.Exit from inside it would skip
|
||||
// these defers and truncate the profile of a failing command -- exactly
|
||||
// the command one most often wants to profile.
|
||||
func run() int {
|
||||
// CPU profiling: set VAULTIK_CPUPROFILE=/path/to/cpu.prof
|
||||
if cpuProfile := os.Getenv("VAULTIK_CPUPROFILE"); cpuProfile != "" {
|
||||
f, err := os.Create(cpuProfile) //nolint:gosec // G304: operator-set path
|
||||
@@ -46,5 +56,5 @@ func main() {
|
||||
}()
|
||||
}
|
||||
|
||||
cli.Entry()
|
||||
return cli.Entry()
|
||||
}
|
||||
|
||||
@@ -1,8 +1,6 @@
|
||||
package main_test
|
||||
|
||||
import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"slices"
|
||||
"strings"
|
||||
@@ -87,27 +85,11 @@ func TestBuildTargetBuildsTheBinary(t *testing.T) {
|
||||
}
|
||||
|
||||
// readMakefile returns the contents of the repository's Makefile. The
|
||||
// test binary runs with its package directory as the working directory,
|
||||
// so the root is found by walking up until the Makefile appears.
|
||||
// root is located by the shared walk in lintdocker_test.go.
|
||||
func readMakefile(t *testing.T) string {
|
||||
t.Helper()
|
||||
|
||||
dir, err := os.Getwd()
|
||||
require.NoError(t, err)
|
||||
|
||||
for {
|
||||
//nolint:gosec // G304: the path is this test's own directory walk
|
||||
contents, err := os.ReadFile(filepath.Join(dir, "Makefile"))
|
||||
if err == nil {
|
||||
return string(contents)
|
||||
}
|
||||
|
||||
parent := filepath.Dir(dir)
|
||||
require.NotEqual(t, dir, parent,
|
||||
"walked to the filesystem root without finding a Makefile")
|
||||
|
||||
dir = parent
|
||||
}
|
||||
return readRepoFile(t, "Makefile")
|
||||
}
|
||||
|
||||
// phonyTargets returns every name declared phony, across all .PHONY
|
||||
|
||||
@@ -304,6 +304,8 @@ storage_url: "rclone://las1stor1//srv/pool.2024.04/backups/heraklion"
|
||||
|
||||
# Maximum blob size
|
||||
# Multiple chunks are packed into blobs up to this size
|
||||
# Must be at least four times chunk_size (the largest chunk the chunker can
|
||||
# emit); a smaller limit would let a single-chunk blob exceed it.
|
||||
# Supports: 1GB, 10G, 500MB, 1GiB, etc.
|
||||
# Default: 10GB
|
||||
#blob_size_limit: 10GB
|
||||
|
||||
+29
-8
@@ -5,11 +5,30 @@
|
||||
Vaultik uses a local SQLite database to track file metadata, chunk mappings, and blob associations during the backup process. This database serves as an index for incremental backups and enables efficient deduplication.
|
||||
|
||||
**Important Notes:**
|
||||
- **No Migration Support (pre-1.0)**: Vaultik does not support database schema
|
||||
migrations. The local index is treated as disposable — if the schema changes,
|
||||
delete the local SQLite database (`vaultik database delete`) and run a full
|
||||
backup. The remote storage is unaffected; the new index will re-deduplicate
|
||||
against existing remote blobs.
|
||||
|
||||
This section is the authoritative explanation of the schema/migration story;
|
||||
other documents (the README and `AGENTS.md`) link here.
|
||||
|
||||
- **No upgrade path between versions (pre-1.0)**: Vaultik has no supported way to
|
||||
carry an existing local index across a schema change. The index is disposable
|
||||
— if the on-disk schema changes between versions, delete the local SQLite
|
||||
database (`vaultik database delete`) and run a full backup. Remote storage is
|
||||
unaffected; the new index re-deduplicates against existing remote blobs. This
|
||||
is the standing project policy, and it is separate from the schema bootstrap
|
||||
described next.
|
||||
- **Schema bootstrap**: a fresh database is populated from numbered SQL files
|
||||
embedded in the binary under `internal/database/schema/`. `000.sql` creates the
|
||||
`schema_migrations` table; `001.sql` creates the application tables. On opening
|
||||
a database the code applies each numbered file that has not yet run and records
|
||||
its version in `schema_migrations`. This bootstraps a new database; it does not
|
||||
upgrade an existing one between released versions.
|
||||
- **Changing the schema (pre-1.0)**: edit `internal/database/schema/001.sql` (and
|
||||
the code that touches the affected tables) directly. Do not add new numbered
|
||||
files — there is no installed base to migrate.
|
||||
- **Disposability expires at 1.0**: the index is treated as disposable only until
|
||||
1.0 ships and is tagged. Once 1.0 is tagged that clause expires and the
|
||||
question of upgrading existing indexes returns. It is deliberately left open
|
||||
here.
|
||||
- **Version Compatibility**: In rare cases, you may need to use the same version
|
||||
of Vaultik to restore a backup as was used to create it. This ensures
|
||||
compatibility with the metadata format stored in S3.
|
||||
@@ -192,10 +211,12 @@ Tracks blob upload metrics.
|
||||
After a snapshot is completed:
|
||||
1. Copy database to temporary file
|
||||
2. Clean temporary database to contain only current snapshot data
|
||||
3. Export to SQL dump using sqlite3
|
||||
3. VACUUM the trimmed database so deleted rows leave no pages behind
|
||||
4. Compress with zstd and encrypt with age
|
||||
5. Upload to S3 as `metadata/{snapshot-id}/db.zst.age`
|
||||
6. Generate blob manifest and upload as `metadata/{snapshot-id}/manifest.json.zst`
|
||||
5. Upload to S3 as `metadata/{remote-key}/db.zst.age`
|
||||
6. Generate blob manifest and upload as `metadata/{remote-key}/manifest.json.zst`
|
||||
|
||||
The `{remote-key}` directory name is a one-way hash of the human snapshot ID, so the ID is never written to the store in plaintext; see [REPOSTRUCTURE.md](REPOSTRUCTURE.md#remote-key-derivation).
|
||||
|
||||
### 4. Restore Process
|
||||
|
||||
|
||||
+40
-20
@@ -17,11 +17,13 @@ Vaultik stores all backup data in an S3-compatible object store. The repository
|
||||
│ └── <hash[2:4]>/
|
||||
│ └── <full-hash>
|
||||
└── metadata/
|
||||
└── <snapshot-id>/
|
||||
└── <remote-key>/
|
||||
├── db.zst.age
|
||||
└── manifest.json.zst
|
||||
```
|
||||
|
||||
The metadata subdirectory is named with the **remote key**, a one-way hash of the snapshot ID, not with the human-readable snapshot ID itself. See [Remote Key Derivation](#remote-key-derivation).
|
||||
|
||||
## Blobs Directory (`blobs/`)
|
||||
|
||||
### Structure
|
||||
@@ -40,9 +42,11 @@ Blobs contain the actual file data from backups and must be encrypted for securi
|
||||
|
||||
## Metadata Directory (`metadata/`)
|
||||
|
||||
Each snapshot has its own subdirectory named with the snapshot ID.
|
||||
Each snapshot has its own subdirectory. The directory is **not** named with the human-readable snapshot ID; it is named with the remote key — a one-way hash of that ID. The human ID is never written to the destination store as a directory name (see [Remote Key Derivation](#remote-key-derivation)).
|
||||
|
||||
### Snapshot ID Format
|
||||
|
||||
The human-readable snapshot ID is used in CLI arguments, log lines, and the local database. It is not written to the destination store.
|
||||
- **Format**: `<hostname>_<snapshot-name>_<RFC3339>` (or `<hostname>_<RFC3339>` if no
|
||||
name was specified)
|
||||
- **Example**: `laptop_home_2024-01-15T14:30:52Z`
|
||||
@@ -51,6 +55,19 @@ Each snapshot has its own subdirectory named with the snapshot ID.
|
||||
- Snapshot name from the configured `snapshots:` map (optional)
|
||||
- RFC3339 UTC timestamp
|
||||
|
||||
This ID reveals the hostname, the configured snapshot name, and the backup time, so it is never used as the on-disk directory name — the remote key is used instead.
|
||||
|
||||
### Remote Key Derivation
|
||||
|
||||
The remote key is `hex(SHA256(SHA256("vaultik|" + snapshot-id)))`: a double SHA-256 over the snapshot ID, with a `vaultik|` domain-separation prefix. The result is a 64-character hex string with no structure a remote observer can reverse. Implemented in `internal/snapshot/remotekey.go`.
|
||||
|
||||
Worked example:
|
||||
- Snapshot ID: `server1_home_2025-06-01T12:00:00Z`
|
||||
- Remote key: `17f97bcde958748af076b926af59823943db59e80ce7170b40f124dfa28f64aa`
|
||||
- Directory: `metadata/17f97bcde958748af076b926af59823943db59e80ce7170b40f124dfa28f64aa/`
|
||||
|
||||
Because the hash is one-way, a listing of the destination store reveals neither the hostname nor the snapshot name of any backup. The same remote key is stored in the manifest's `snapshot_id` field.
|
||||
|
||||
### Files in Each Snapshot Directory
|
||||
|
||||
#### `db.zst.age` - Encrypted Database
|
||||
@@ -68,16 +85,17 @@ Each snapshot has its own subdirectory named with the snapshot ID.
|
||||
- **Structure**:
|
||||
```json
|
||||
{
|
||||
"snapshot_id": "laptop_home_2024-01-15T14:30:52Z",
|
||||
"timestamp": "2024-01-15T14:30:52Z",
|
||||
"snapshot_id": "17f97bcde958748af076b926af59823943db59e80ce7170b40f124dfa28f64aa",
|
||||
"timestamp": "2025-06-01T12:00:00Z",
|
||||
"blob_count": 42,
|
||||
"total_compressed_size": 1048576,
|
||||
"blobs": [
|
||||
"cafebabe1234567890abcdef1234567890abcdef1234567890abcdef12345678",
|
||||
"deadbeef1234567890abcdef1234567890abcdef1234567890abcdef12345678",
|
||||
...
|
||||
{ "hash": "cafebabe1234567890abcdef1234567890abcdef1234567890abcdef12345678", "compressed_size": 24576 },
|
||||
{ "hash": "deadbeef1234567890abcdef1234567890abcdef1234567890abcdef12345678", "compressed_size": 32768 }
|
||||
]
|
||||
}
|
||||
```
|
||||
`snapshot_id` is the remote key (a hash), not the human ID; `timestamp` is written in the clear.
|
||||
|
||||
### Why Manifest is Unencrypted
|
||||
The manifest must be readable without the private key to enable:
|
||||
@@ -86,7 +104,7 @@ The manifest must be readable without the private key to enable:
|
||||
3. **Verification** - Checking blob existence without decryption
|
||||
4. **Cross-snapshot deduplication analysis** - Finding shared blobs between snapshots
|
||||
|
||||
The manifest only contains blob hashes, not file names or any other sensitive information.
|
||||
The manifest contains the remote key, the backup timestamp, the blob count and total compressed size, and each blob's hash and compressed size. It contains no file names, paths, or other decrypted metadata.
|
||||
|
||||
## Security Considerations
|
||||
|
||||
@@ -96,19 +114,21 @@ The manifest only contains blob hashes, not file names or any other sensitive in
|
||||
- **File-to-chunk mappings** (in db.zst.age)
|
||||
|
||||
### What's Not Encrypted
|
||||
- **Blob hashes** (in manifest.json.zst)
|
||||
- **Snapshot IDs** (directory names)
|
||||
- **Blob count per snapshot** (in manifest.json.zst)
|
||||
- **The remote key** — directory names and the manifest `snapshot_id`, a one-way hash of the snapshot ID (see [Remote Key Derivation](#remote-key-derivation))
|
||||
- **The backup timestamp** (in manifest.json.zst)
|
||||
- **Blob hashes and their compressed sizes** (in manifest.json.zst)
|
||||
- **Blob count and total compressed size per snapshot** (in manifest.json.zst)
|
||||
|
||||
### Privacy Implications
|
||||
From the unencrypted data, an observer can determine:
|
||||
- When backups were taken (from snapshot IDs)
|
||||
- Which hostname created backups (from snapshot IDs)
|
||||
- How many blobs each snapshot references
|
||||
- Which blobs are shared between snapshots (deduplication patterns)
|
||||
- The size of each encrypted blob
|
||||
From the unencrypted data, an observer of the destination store can determine:
|
||||
- **When each backup was taken** — not from the directory name, which is a one-way hash, but from the plaintext `timestamp` field in manifest.json.zst, which is published in the clear
|
||||
- How many blobs each snapshot references, and the total compressed size
|
||||
- The compressed size of each blob, and which blobs are shared between snapshots (deduplication patterns)
|
||||
|
||||
Together these give an observer a timing-and-size profile of every snapshot. This is an accepted, documented property of the format, not a defect: the manifest is unencrypted so that pruning can run without the private key, and the timing channel could not be closed by encrypting it anyway — object creation times and per-object sizes stay visible at the storage layer on both `s3://` and `file://` destinations regardless.
|
||||
|
||||
An observer cannot determine:
|
||||
- The hostname or snapshot name of any backup (the directory name and the manifest `snapshot_id` are one-way hashes of the human ID)
|
||||
- File names or paths
|
||||
- File contents
|
||||
- File permissions or ownership
|
||||
@@ -125,10 +145,10 @@ An observer cannot determine:
|
||||
## Pruning Safety
|
||||
|
||||
The prune operation is safe because:
|
||||
1. It only deletes blobs not referenced in any manifest
|
||||
1. It keeps every blob listed in any snapshot's manifest and deletes only blobs that no manifest references
|
||||
2. Manifests are unencrypted and can be read without keys
|
||||
3. The operation compares the latest local DB snapshot with the latest S3 snapshot to ensure consistency
|
||||
4. Pruning will fail if these don't match, preventing accidental deletion of needed blobs
|
||||
3. If any manifest cannot be downloaded or decoded, prune deletes nothing and exits with an error, rather than treating that snapshot's blobs as unreferenced
|
||||
4. Prune requires exclusive access to the destination: running it during a concurrent backup can race a snapshot whose manifest is not yet written, so do not prune while a backup is in progress
|
||||
|
||||
## Restoration Requirements
|
||||
|
||||
|
||||
@@ -33,9 +33,10 @@ type Chunker struct {
|
||||
maxChunkSize int
|
||||
}
|
||||
|
||||
// chunkSizeSpread is the FastCDC-recommended factor between the average
|
||||
// ChunkSizeSpread is the FastCDC-recommended factor between the average
|
||||
// chunk size and the minimum (avg/spread) and maximum (avg*spread) sizes.
|
||||
const chunkSizeSpread = 4
|
||||
// The largest chunk the chunker can emit is therefore avg*ChunkSizeSpread.
|
||||
const ChunkSizeSpread = 4
|
||||
|
||||
// NewChunker creates a new chunker with the specified average chunk size.
|
||||
// The actual chunk sizes will vary between avgChunkSize/4 and avgChunkSize*4
|
||||
@@ -45,8 +46,8 @@ func NewChunker(avgChunkSize int64) *Chunker {
|
||||
// FastCDC recommends min = avg/4 and max = avg*4
|
||||
return &Chunker{
|
||||
avgChunkSize: int(avgChunkSize),
|
||||
minChunkSize: int(avgChunkSize / chunkSizeSpread),
|
||||
maxChunkSize: int(avgChunkSize * chunkSizeSpread),
|
||||
minChunkSize: int(avgChunkSize / ChunkSizeSpread),
|
||||
maxChunkSize: int(avgChunkSize * ChunkSizeSpread),
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
+155
-58
@@ -11,6 +11,7 @@ import (
|
||||
"os/signal"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync"
|
||||
"syscall"
|
||||
"time"
|
||||
|
||||
@@ -32,14 +33,33 @@ import (
|
||||
// may take before we give up.
|
||||
const shutdownTimeout = 30 * time.Second
|
||||
|
||||
// AppOptions contains common options for creating the fx application.
|
||||
// It includes the configuration file path, logging options, and additional
|
||||
// fx modules and invocations that should be included in the application.
|
||||
// lockMode says whether a command mutates persistent state — the local
|
||||
// index database or the remote store — and so must hold the process-wide
|
||||
// PID lock, or only reads that state and may run alongside a mutator.
|
||||
type lockMode int
|
||||
|
||||
const (
|
||||
// mutating commands (snapshot create, snapshot purge, snapshot remove,
|
||||
// prune, remote nuke) write the local index or the remote store. They
|
||||
// hold the PID lock so that at most one runs at a time.
|
||||
mutating lockMode = iota
|
||||
// readOnly commands (info, snapshot list, snapshot verify, remote info,
|
||||
// snapshot restore) do not write the local index or the remote store,
|
||||
// so they run without the lock and are never blocked by a running
|
||||
// mutator. restore writes only to the target directory it is given.
|
||||
readOnly
|
||||
)
|
||||
|
||||
// AppOptions contains common options for creating and running the fx
|
||||
// application: the configuration file path, logging options, additional fx
|
||||
// modules and invocations, and whether the command mutates persistent
|
||||
// state (which decides whether it takes the PID lock).
|
||||
type AppOptions struct {
|
||||
ConfigPath string
|
||||
LogOptions log.Options
|
||||
Modules []fx.Option
|
||||
Invokes []fx.Option
|
||||
Mode lockMode
|
||||
}
|
||||
|
||||
// setupGlobals records the startup time and, when an output-suppression
|
||||
@@ -48,6 +68,11 @@ type AppOptions struct {
|
||||
// silenced — per the documented convention that --quiet suppresses
|
||||
// non-error output only. The startup banner is printed by Entry
|
||||
// before cobra parses arguments, gated by the same arg-level check.
|
||||
//
|
||||
// --json quiets the UI here too, because stdout then carries a JSON
|
||||
// document and human narration would corrupt it. Unlike Quiet it does
|
||||
// not lower the stderr log level (issue #112), so --verbose/--debug
|
||||
// still surface diagnostics alongside the document.
|
||||
func setupGlobals(
|
||||
lc fx.Lifecycle, g *globals.Globals, v *vaultik.Vaultik, opts log.Options,
|
||||
) {
|
||||
@@ -55,7 +80,7 @@ func setupGlobals(
|
||||
OnStart: func(_ context.Context) error {
|
||||
g.StartTime = time.Now().UTC()
|
||||
|
||||
if opts.Cron || opts.Quiet {
|
||||
if opts.Cron || opts.Quiet || opts.JSON {
|
||||
v.UI.SetQuiet(true)
|
||||
}
|
||||
|
||||
@@ -196,52 +221,54 @@ func RunApp(ctx context.Context, app *fx.App) error {
|
||||
}
|
||||
}
|
||||
|
||||
// runVaultikApp runs the standard single-operation command lifecycle
|
||||
// shared by the list/purge/verify/remove/remote-info subcommands:
|
||||
// resolve the config, start the fx app, run op against the Vaultik
|
||||
// instance in a goroutine, report a failure prefixed with failMsg
|
||||
// (suppressed while suppressErrors is true, e.g. under --json), then
|
||||
// trigger shutdown. The operation is cancelled when the app stops.
|
||||
// extraQuiet is OR-ed into LogOptions.Quiet (e.g. --json output modes).
|
||||
func runVaultikApp(
|
||||
cmd *cobra.Command, extraQuiet, suppressErrors bool,
|
||||
failMsg string, op func(v *vaultik.Vaultik) error,
|
||||
// errReported marks a failure the operation has already shown the user
|
||||
// (and deliberately withheld under --json). Entry turns it into a
|
||||
// non-zero exit status without printing anything further, so the error
|
||||
// line is not doubled. It flows up from RunOperation through cobra to
|
||||
// Entry.
|
||||
var errReported = errors.New("operation failed")
|
||||
|
||||
// RunOperation runs op against the Vaultik instance inside the fx app
|
||||
// and turns a failure into a returned error rather than an os.Exit from
|
||||
// within the goroutine. An os.Exit there skipped main's deferred
|
||||
// profile writers -- so profiling a failing command yielded a truncated
|
||||
// profile (issue #75) -- and RunWithApp's PID-lock release, and denied
|
||||
// the app any graceful shutdown; returning the error to the top runs
|
||||
// all three.
|
||||
//
|
||||
// op runs in a goroutine so OnStart returns promptly and an interrupt
|
||||
// can still cancel through OnStop; when it finishes, success or failure,
|
||||
// it triggers shutdown, which is what lets RunWithApp return. report is
|
||||
// called with a non-canceled failure so the caller can log it (and
|
||||
// suppress it under --json) before it becomes errReported. A context
|
||||
// cancellation is the interrupt path, not a failure: it is neither
|
||||
// reported nor counted as one.
|
||||
func RunOperation(
|
||||
ctx context.Context, opts AppOptions,
|
||||
op func(v *vaultik.Vaultik) error, report func(err error),
|
||||
) error {
|
||||
configPath, err := ResolveConfigPath()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
var (
|
||||
mu sync.Mutex
|
||||
failed bool
|
||||
)
|
||||
|
||||
rootFlags := GetRootFlags()
|
||||
|
||||
return RunWithApp(cmd.Context(), AppOptions{
|
||||
ConfigPath: configPath,
|
||||
LogOptions: log.Options{
|
||||
Verbose: rootFlags.Verbose,
|
||||
Debug: rootFlags.Debug,
|
||||
Quiet: rootFlags.Quiet || extraQuiet,
|
||||
},
|
||||
Modules: []fx.Option{},
|
||||
Invokes: []fx.Option{
|
||||
opts.Invokes = append(opts.Invokes,
|
||||
fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
|
||||
lc.Append(fx.Hook{
|
||||
OnStart: func(_ context.Context) error {
|
||||
go func() {
|
||||
err := op(v)
|
||||
if err != nil {
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
if !suppressErrors {
|
||||
log.Error(failMsg, "error", err)
|
||||
ReportErrorf("%s: %v", failMsg, err)
|
||||
if err != nil && !errors.Is(err, context.Canceled) {
|
||||
report(err)
|
||||
|
||||
mu.Lock()
|
||||
failed = true
|
||||
mu.Unlock()
|
||||
}
|
||||
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
err = v.Shutdowner.Shutdown()
|
||||
if err != nil {
|
||||
log.Error("Failed to shutdown", "error", err)
|
||||
stopErr := v.Shutdowner.Shutdown()
|
||||
if stopErr != nil {
|
||||
log.Error("Failed to shutdown", "error", stopErr)
|
||||
}
|
||||
}()
|
||||
|
||||
@@ -253,36 +280,106 @@ func runVaultikApp(
|
||||
return nil
|
||||
},
|
||||
})
|
||||
}),
|
||||
}))
|
||||
|
||||
err := RunWithApp(ctx, opts)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
// The goroutine sets failed before triggering the shutdown that lets
|
||||
// RunWithApp return, so the write is in place by the time we read it.
|
||||
mu.Lock()
|
||||
defer mu.Unlock()
|
||||
|
||||
if failed {
|
||||
return errReported
|
||||
}
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// runVaultikApp runs the standard single-operation command lifecycle
|
||||
// shared by the snapshot list/purge/remove and remote nuke subcommands:
|
||||
// resolve the config, then run op against the Vaultik instance through
|
||||
// RunOperation, reporting a failure prefixed with failMsg (suppressed
|
||||
// while suppressErrors is true, e.g. under --json). mode says whether the
|
||||
// command takes the PID lock. jsonOutput marks a command whose stdout is a
|
||||
// JSON document: it quiets the UI but, unlike Quiet, leaves the stderr log
|
||||
// level alone.
|
||||
func runVaultikApp(
|
||||
cmd *cobra.Command, mode lockMode, jsonOutput, suppressErrors bool,
|
||||
failMsg string, op func(v *vaultik.Vaultik) error,
|
||||
) error {
|
||||
configPath, err := ResolveConfigPath()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
rootFlags := GetRootFlags()
|
||||
|
||||
return RunOperation(cmd.Context(), AppOptions{
|
||||
ConfigPath: configPath,
|
||||
LogOptions: log.Options{
|
||||
Verbose: rootFlags.Verbose,
|
||||
Debug: rootFlags.Debug,
|
||||
Quiet: rootFlags.Quiet,
|
||||
JSON: jsonOutput,
|
||||
},
|
||||
Mode: mode,
|
||||
}, op, func(err error) {
|
||||
if suppressErrors {
|
||||
return
|
||||
}
|
||||
|
||||
log.Error(failMsg, "error", err)
|
||||
ReportErrorf("%s: %v", failMsg, err)
|
||||
})
|
||||
}
|
||||
|
||||
// RunWithApp is a helper that creates and runs an fx app with the given options.
|
||||
// It combines NewApp and RunApp into a single convenient function. This is the
|
||||
// preferred way to run CLI commands that need the full application context.
|
||||
// It acquires a PID lock before starting to prevent concurrent instances.
|
||||
// A mutating command takes the process-wide PID lock before starting so that
|
||||
// only one runs at a time; a read-only command runs without it and is not
|
||||
// blocked while a mutator holds the lock (opts.Mode).
|
||||
func RunWithApp(ctx context.Context, opts AppOptions) error {
|
||||
// Acquire PID lock to prevent concurrent instances
|
||||
lockDir := filepath.Join(xdg.DataHome, "vaultik")
|
||||
|
||||
lock, err := pidlock.Acquire(lockDir)
|
||||
release, err := acquireLockIfMutating(opts.Mode,
|
||||
filepath.Join(xdg.DataHome, "vaultik"))
|
||||
if err != nil {
|
||||
if errors.Is(err, pidlock.ErrAlreadyRunning) {
|
||||
return fmt.Errorf("cannot start: %w", err)
|
||||
return err
|
||||
}
|
||||
|
||||
return fmt.Errorf("failed to acquire lock: %w", err)
|
||||
}
|
||||
|
||||
defer func() {
|
||||
err := lock.Release()
|
||||
if err != nil {
|
||||
log.Warn("Failed to release PID lock", "error", err)
|
||||
}
|
||||
}()
|
||||
defer release()
|
||||
|
||||
app := NewApp(opts)
|
||||
|
||||
return RunApp(ctx, app)
|
||||
}
|
||||
|
||||
// acquireLockIfMutating takes the process-wide PID lock in lockDir for a
|
||||
// mutating command and returns a function that releases it. A read-only
|
||||
// command takes no lock, so it returns a no-op release and is never blocked
|
||||
// while a mutator holds the lock. ErrAlreadyRunning (another mutator holds
|
||||
// the lock) is surfaced as a "cannot start" error.
|
||||
func acquireLockIfMutating(mode lockMode, lockDir string) (func(), error) {
|
||||
if mode != mutating {
|
||||
return func() {}, nil
|
||||
}
|
||||
|
||||
lock, err := pidlock.Acquire(lockDir)
|
||||
if err != nil {
|
||||
if errors.Is(err, pidlock.ErrAlreadyRunning) {
|
||||
return nil, fmt.Errorf("cannot start: %w", err)
|
||||
}
|
||||
|
||||
return nil, fmt.Errorf("failed to acquire lock: %w", err)
|
||||
}
|
||||
|
||||
return func() {
|
||||
err := lock.Release()
|
||||
if err != nil {
|
||||
log.Warn("Failed to release PID lock", "error", err)
|
||||
}
|
||||
}, nil
|
||||
}
|
||||
|
||||
@@ -2,7 +2,10 @@ package cli //nolint:testpackage // needs access to unexported cleanStartupError
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/pidlock"
|
||||
)
|
||||
|
||||
func TestCleanStartupError(t *testing.T) {
|
||||
@@ -53,3 +56,42 @@ func TestCleanStartupError(t *testing.T) {
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestLockScopedToMutatingCommands proves the partition the PID lock now
|
||||
// enforces: a read-only command runs while a mutator holds the lock, and
|
||||
// two mutating commands still mutually exclude.
|
||||
func TestLockScopedToMutatingCommands(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
lockDir := filepath.Join(t.TempDir(), "vaultik")
|
||||
|
||||
// A mutating command takes the process-wide lock.
|
||||
releaseMutator, err := acquireLockIfMutating(mutating, lockDir)
|
||||
if err != nil {
|
||||
t.Fatalf("mutating command could not acquire lock: %v", err)
|
||||
}
|
||||
|
||||
// A read-only command runs to completion even while the lock is held.
|
||||
releaseReader, err := acquireLockIfMutating(readOnly, lockDir)
|
||||
if err != nil {
|
||||
t.Fatalf("read-only command was blocked by held lock: %v", err)
|
||||
}
|
||||
|
||||
releaseReader()
|
||||
|
||||
// A second mutating command is refused while the first holds the lock.
|
||||
_, err = acquireLockIfMutating(mutating, lockDir)
|
||||
if !errors.Is(err, pidlock.ErrAlreadyRunning) {
|
||||
t.Fatalf("second mutating command was not excluded, got: %v", err)
|
||||
}
|
||||
|
||||
// Once the first mutator releases, another mutating command may run.
|
||||
releaseMutator()
|
||||
|
||||
release, err := acquireLockIfMutating(mutating, lockDir)
|
||||
if err != nil {
|
||||
t.Fatalf("mutating command could not acquire released lock: %v", err)
|
||||
}
|
||||
|
||||
release()
|
||||
}
|
||||
|
||||
+58
-13
@@ -1,8 +1,10 @@
|
||||
package cli
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
@@ -24,6 +26,11 @@ const configSetArgs = 2
|
||||
// parent config dirs (e.g. ~/.config) are conventionally traversable.
|
||||
const configDirMode = 0o755
|
||||
|
||||
// configYAMLIndent matches the 2-space indentation of defaultConfigTemplate,
|
||||
// so `config set` writes the file back with the same indentation rather than
|
||||
// yaml.Marshal's 4-space default.
|
||||
const configYAMLIndent = 2
|
||||
|
||||
var (
|
||||
errConfigExists = errors.New("config file already exists")
|
||||
errEmptyConfig = errors.New("empty config file")
|
||||
@@ -186,8 +193,8 @@ storage_url: ""
|
||||
# access_key_id: YOUR_ACCESS_KEY
|
||||
# secret_access_key: YOUR_SECRET_KEY
|
||||
# # region: us-east-1 # Default: us-east-1
|
||||
# # use_ssl: true # Default: true
|
||||
# # part_size: 5MB # Multipart upload part size. Default: 5MB
|
||||
# # For the s3:// form, disable TLS with ?ssl=false in the URL, not use_ssl.
|
||||
|
||||
# ─── OPTIONAL ────────────────────────────────────────────────────────────────
|
||||
|
||||
@@ -206,6 +213,8 @@ storage_url: ""
|
||||
# chunk_size: 10MB
|
||||
|
||||
# Maximum blob size before splitting into a new blob.
|
||||
# Must be at least four times chunk_size (the largest chunk the chunker can
|
||||
# emit); a smaller limit would let a single-chunk blob exceed it.
|
||||
# Accepts: 1GB, 10G, 500MB, etc.
|
||||
# Default: 10GB
|
||||
# blob_size_limit: 10GB
|
||||
@@ -371,38 +380,74 @@ Examples:
|
||||
return err
|
||||
}
|
||||
|
||||
return writeConfigSet(os.Stdout, path, args[0], args[1])
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// writeConfigSet applies key=value to the config at path, writes it back
|
||||
// owner-only, and confirms the write by printing just the key name to w.
|
||||
// The value is never echoed: it may be a secret such as
|
||||
// s3.secret_access_key, and captured stdout or a pasted terminal would
|
||||
// then leak it.
|
||||
func writeConfigSet(w io.Writer, path, key, value string) error {
|
||||
root, err := loadYAMLFile(path)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
err = yamlPathSet(root, strings.Split(args[0], "."), args[1])
|
||||
err = yamlPathSet(root, strings.Split(key, "."), value)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
out, err := yaml.Marshal(root)
|
||||
out, err := marshalConfigYAML(root)
|
||||
if err != nil {
|
||||
return fmt.Errorf("marshaling config: %w", err)
|
||||
}
|
||||
|
||||
mode := os.FileMode(configFileMode)
|
||||
|
||||
info, statErr := os.Stat(path)
|
||||
if statErr == nil {
|
||||
mode = info.Mode().Perm()
|
||||
}
|
||||
|
||||
err = os.WriteFile(path, out, mode)
|
||||
err = os.WriteFile(path, out, configFileMode)
|
||||
if err != nil {
|
||||
return fmt.Errorf("writing config file: %w", err)
|
||||
}
|
||||
|
||||
_, _ = fmt.Fprintf(os.Stdout, "%s = %s\n", args[0], args[1])
|
||||
// os.WriteFile does not change the mode of a file that already exists,
|
||||
// so a config that was group- or world-readable stays that way. As it
|
||||
// may hold S3 credentials, tighten it to owner-only after writing.
|
||||
info, statErr := os.Stat(path)
|
||||
if statErr == nil && info.Mode().Perm()&0o044 != 0 {
|
||||
err = os.Chmod(path, configFileMode)
|
||||
if err != nil {
|
||||
return fmt.Errorf("tightening config file permissions: %w", err)
|
||||
}
|
||||
}
|
||||
|
||||
_, _ = fmt.Fprintln(w, key)
|
||||
|
||||
return nil
|
||||
},
|
||||
}
|
||||
|
||||
// marshalConfigYAML renders a config document tree with 2-space indentation,
|
||||
// matching defaultConfigTemplate. yaml.Marshal defaults to 4 spaces, which
|
||||
// would reindent the whole file on the first `config set` despite the promise
|
||||
// to preserve formatting.
|
||||
func marshalConfigYAML(root *yaml.Node) ([]byte, error) {
|
||||
var buf bytes.Buffer
|
||||
|
||||
enc := yaml.NewEncoder(&buf)
|
||||
enc.SetIndent(configYAMLIndent)
|
||||
|
||||
err := enc.Encode(root)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
err = enc.Close()
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
return buf.Bytes(), nil
|
||||
}
|
||||
|
||||
// loadYAMLFile parses a YAML file into a yaml.Node document tree,
|
||||
|
||||
@@ -1,6 +1,9 @@
|
||||
package cli //nolint:testpackage // exercises unexported yamlPathGet/yamlPathSet
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
@@ -188,6 +191,109 @@ func TestYAMLPathSet(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestConfigSetPreservesFormatting asserts the `config set` write path
|
||||
// (marshalConfigYAML) round-trips a 2-space-indented file without reindenting
|
||||
// it to yaml.Marshal's 4-space default, and keeps comments.
|
||||
func TestConfigSetPreservesFormatting(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
root := parseTestYAML(t)
|
||||
|
||||
err := yamlPathSet(root, splitPath("s3.bucket"), "newbucket")
|
||||
if err != nil {
|
||||
t.Fatalf("set s3.bucket: %v", err)
|
||||
}
|
||||
|
||||
out, err := marshalConfigYAML(root)
|
||||
if err != nil {
|
||||
t.Fatalf("marshal: %v", err)
|
||||
}
|
||||
|
||||
text := string(out)
|
||||
|
||||
for _, want := range []string{"# top comment", "# inline comment"} {
|
||||
if !contains(text, want) {
|
||||
t.Errorf("round-tripped YAML dropped comment %q:\n%s", want, text)
|
||||
}
|
||||
}
|
||||
|
||||
// Nested map keys stay at 2-space indent; the bug reindented them to 4.
|
||||
if !contains(text, "\n bucket: newbucket") {
|
||||
t.Errorf("expected 2-space indent for s3.bucket, got:\n%s", text)
|
||||
}
|
||||
|
||||
if contains(text, "\n bucket:") {
|
||||
t.Errorf("s3.bucket reindented to 4 spaces:\n%s", text)
|
||||
}
|
||||
|
||||
// Sequence items under a key also stay at 2 spaces.
|
||||
if !contains(text, "\n - age1aaa") {
|
||||
t.Errorf("expected 2-space indent for sequence item, got:\n%s", text)
|
||||
}
|
||||
}
|
||||
|
||||
// TestWriteConfigSetHidesSecret checks that setting a secret key prints
|
||||
// only the key name, never the value, to the confirmation output.
|
||||
func TestWriteConfigSetHidesSecret(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const secret = "SUPERSECRETVALUE"
|
||||
|
||||
path := filepath.Join(t.TempDir(), "config.yaml")
|
||||
|
||||
err := os.WriteFile(path, []byte("version: 1\n"), 0o600)
|
||||
if err != nil {
|
||||
t.Fatalf("seed config: %v", err)
|
||||
}
|
||||
|
||||
var out bytes.Buffer
|
||||
|
||||
err = writeConfigSet(&out, path, "s3.secret_access_key", secret)
|
||||
if err != nil {
|
||||
t.Fatalf("writeConfigSet: %v", err)
|
||||
}
|
||||
|
||||
if strings.Contains(out.String(), secret) {
|
||||
t.Errorf("output echoed the secret value: %q", out.String())
|
||||
}
|
||||
|
||||
if !strings.Contains(out.String(), "s3.secret_access_key") {
|
||||
t.Errorf("output did not confirm the key name: %q", out.String())
|
||||
}
|
||||
}
|
||||
|
||||
// TestWriteConfigSetTightensMode checks that a pre-existing group- or
|
||||
// world-readable config is tightened to owner-only after a set, since
|
||||
// os.WriteFile leaves an existing file's mode untouched.
|
||||
func TestWriteConfigSetTightensMode(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
path := filepath.Join(t.TempDir(), "config.yaml")
|
||||
|
||||
// Seed a world-readable config; the loose mode is the condition under
|
||||
// test, so gosec's G306 is expected here.
|
||||
err := os.WriteFile(path, []byte("version: 1\n"), 0o644) //nolint:gosec // G306
|
||||
if err != nil {
|
||||
t.Fatalf("seed config: %v", err)
|
||||
}
|
||||
|
||||
var out bytes.Buffer
|
||||
|
||||
err = writeConfigSet(&out, path, "compression_level", "9")
|
||||
if err != nil {
|
||||
t.Fatalf("writeConfigSet: %v", err)
|
||||
}
|
||||
|
||||
info, err := os.Stat(path)
|
||||
if err != nil {
|
||||
t.Fatalf("stat config: %v", err)
|
||||
}
|
||||
|
||||
if info.Mode().Perm() != 0o600 {
|
||||
t.Errorf("config mode = %04o, want 0600", info.Mode().Perm())
|
||||
}
|
||||
}
|
||||
|
||||
func splitPath(s string) []string {
|
||||
return strings.Split(s, ".")
|
||||
}
|
||||
|
||||
@@ -1,126 +0,0 @@
|
||||
package cli
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
)
|
||||
|
||||
// Approximate lengths of the extended calendar units accepted by
|
||||
// parseDuration.
|
||||
const (
|
||||
durationDay = 24 * time.Hour
|
||||
durationWeek = 7 * durationDay
|
||||
durationMonth = 30 * durationDay
|
||||
durationYear = 365 * durationDay
|
||||
)
|
||||
|
||||
var (
|
||||
errNegativeDuration = errors.New("negative durations are not supported")
|
||||
errInvalidDuration = errors.New("invalid duration format")
|
||||
errUnknownTimeUnit = errors.New("unknown time unit")
|
||||
)
|
||||
|
||||
// parseDuration parses duration strings. Supports standard Go duration format
|
||||
// (e.g., "3h30m", "1h45m30s") as well as extended units:
|
||||
// - d: days (e.g., "30d", "7d")
|
||||
// - w: weeks (e.g., "2w", "4w")
|
||||
// - mo: months (30 days) (e.g., "6mo", "1mo")
|
||||
// - y: years (365 days) (e.g., "1y", "2y")
|
||||
//
|
||||
// Can combine units: "1y6mo", "2w3d", "1d12h30m"
|
||||
func parseDuration(s string) (time.Duration, error) {
|
||||
// First try standard Go duration parsing
|
||||
d, err := time.ParseDuration(s)
|
||||
if err == nil {
|
||||
return d, nil
|
||||
}
|
||||
|
||||
// Extended duration parsing
|
||||
// Check for negative values
|
||||
if strings.HasPrefix(strings.TrimSpace(s), "-") {
|
||||
return 0, errNegativeDuration
|
||||
}
|
||||
|
||||
// Pattern matches: number + unit, repeated
|
||||
re := regexp.MustCompile(`(\d+(?:\.\d+)?)\s*([a-zA-Z]+)`)
|
||||
matches := re.FindAllStringSubmatch(s, -1)
|
||||
|
||||
if len(matches) == 0 {
|
||||
return 0, fmt.Errorf("%w: %q", errInvalidDuration, s)
|
||||
}
|
||||
|
||||
var total time.Duration
|
||||
|
||||
for _, match := range matches {
|
||||
valueStr := match[1]
|
||||
unit := strings.ToLower(match[2])
|
||||
|
||||
value, err := strconv.ParseFloat(valueStr, 64)
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("invalid number %q: %w", valueStr, err)
|
||||
}
|
||||
|
||||
d, err := durationForUnit(value, unit)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
|
||||
total += d
|
||||
}
|
||||
|
||||
return total, nil
|
||||
}
|
||||
|
||||
// durationForUnit converts a value with a (case-normalized) unit suffix
|
||||
// into a time.Duration, accepting Go's standard units plus the extended
|
||||
// calendar units.
|
||||
func durationForUnit(value float64, unit string) (time.Duration, error) {
|
||||
switch unit {
|
||||
// Standard time units
|
||||
case "ns", "nanosecond", "nanoseconds":
|
||||
return time.Duration(value), nil
|
||||
case "us", "µs", "microsecond", "microseconds":
|
||||
return time.Duration(value * float64(time.Microsecond)), nil
|
||||
case "ms", "millisecond", "milliseconds":
|
||||
return time.Duration(value * float64(time.Millisecond)), nil
|
||||
case "s", "sec", "second", "seconds":
|
||||
return time.Duration(value * float64(time.Second)), nil
|
||||
case "m", "min", "minute", "minutes":
|
||||
return time.Duration(value * float64(time.Minute)), nil
|
||||
case "h", "hr", "hour", "hours":
|
||||
return time.Duration(value * float64(time.Hour)), nil
|
||||
// Extended units
|
||||
case "d", "day", "days":
|
||||
return time.Duration(value * float64(durationDay)), nil
|
||||
case "w", "week", "weeks":
|
||||
return time.Duration(value * float64(durationWeek)), nil
|
||||
case "mo", "month", "months":
|
||||
// Using 30 days as approximation
|
||||
return time.Duration(value * float64(durationMonth)), nil
|
||||
case "y", "year", "years":
|
||||
// Using 365 days as approximation
|
||||
return time.Duration(value * float64(durationYear)), nil
|
||||
default:
|
||||
// Try parsing as standard Go duration unit
|
||||
testStr := "1" + unit
|
||||
|
||||
_, err := time.ParseDuration(testStr)
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("%w: %q", errUnknownTimeUnit, unit)
|
||||
}
|
||||
|
||||
// It's a valid Go duration unit, parse the full value
|
||||
fullStr := fmt.Sprintf("%g%s", value, unit)
|
||||
|
||||
d, err := time.ParseDuration(fullStr)
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("invalid duration %q: %w", fullStr, err)
|
||||
}
|
||||
|
||||
return d, nil
|
||||
}
|
||||
}
|
||||
@@ -1,299 +0,0 @@
|
||||
package cli //nolint:testpackage // needs access to unexported parseDuration
|
||||
|
||||
import (
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
type parseDurationCase struct {
|
||||
name string
|
||||
input string
|
||||
expected time.Duration
|
||||
wantErr bool
|
||||
}
|
||||
|
||||
// runParseDurationCases executes a table of parseDuration cases as
|
||||
// parallel subtests.
|
||||
func runParseDurationCases(t *testing.T, tests []parseDurationCase) {
|
||||
t.Helper()
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
got, err := parseDuration(tt.input)
|
||||
|
||||
if tt.wantErr {
|
||||
require.Error(t, err, "expected error for input %q", tt.input)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
require.NoError(t, err, "unexpected error for input %q", tt.input)
|
||||
assert.Equal(t, tt.expected, got, "duration mismatch for input %q", tt.input)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseDurationStandard(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
runParseDurationCases(t, []parseDurationCase{
|
||||
{
|
||||
name: "standard seconds",
|
||||
input: "30s",
|
||||
expected: 30 * time.Second,
|
||||
},
|
||||
{
|
||||
name: "standard minutes",
|
||||
input: "45m",
|
||||
expected: 45 * time.Minute,
|
||||
},
|
||||
{
|
||||
name: "standard hours",
|
||||
input: "2h",
|
||||
expected: 2 * time.Hour,
|
||||
},
|
||||
{
|
||||
name: "standard combined",
|
||||
input: "3h30m",
|
||||
expected: 3*time.Hour + 30*time.Minute,
|
||||
},
|
||||
{
|
||||
name: "standard complex",
|
||||
input: "1h45m30s",
|
||||
expected: 1*time.Hour + 45*time.Minute + 30*time.Second,
|
||||
},
|
||||
{
|
||||
name: "standard with milliseconds",
|
||||
input: "1s500ms",
|
||||
expected: 1*time.Second + 500*time.Millisecond,
|
||||
},
|
||||
})
|
||||
}
|
||||
|
||||
func TestParseDurationExtendedUnits(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
runParseDurationCases(t, []parseDurationCase{
|
||||
// Extended units - days
|
||||
{
|
||||
name: "single day",
|
||||
input: "1d",
|
||||
expected: 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
name: "multiple days",
|
||||
input: "7d",
|
||||
expected: 7 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
name: "fractional days",
|
||||
input: "1.5d",
|
||||
expected: 36 * time.Hour,
|
||||
},
|
||||
{
|
||||
name: "days spelled out",
|
||||
input: "3days",
|
||||
expected: 3 * 24 * time.Hour,
|
||||
},
|
||||
// Extended units - weeks
|
||||
{
|
||||
name: "single week",
|
||||
input: "1w",
|
||||
expected: 7 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
name: "multiple weeks",
|
||||
input: "4w",
|
||||
expected: 4 * 7 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
name: "weeks spelled out",
|
||||
input: "2weeks",
|
||||
expected: 2 * 7 * 24 * time.Hour,
|
||||
},
|
||||
// Extended units - months
|
||||
{
|
||||
name: "single month",
|
||||
input: "1mo",
|
||||
expected: 30 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
name: "multiple months",
|
||||
input: "6mo",
|
||||
expected: 6 * 30 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
name: "months spelled out",
|
||||
input: "3months",
|
||||
expected: 3 * 30 * 24 * time.Hour,
|
||||
},
|
||||
// Extended units - years
|
||||
{
|
||||
name: "single year",
|
||||
input: "1y",
|
||||
expected: 365 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
name: "multiple years",
|
||||
input: "2y",
|
||||
expected: 2 * 365 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
name: "years spelled out",
|
||||
input: "1year",
|
||||
expected: 365 * 24 * time.Hour,
|
||||
},
|
||||
})
|
||||
}
|
||||
|
||||
func TestParseDurationCombinedAndErrors(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
runParseDurationCases(t, []parseDurationCase{
|
||||
// Combined extended units
|
||||
{
|
||||
name: "weeks and days",
|
||||
input: "2w3d",
|
||||
expected: 2*7*24*time.Hour + 3*24*time.Hour,
|
||||
},
|
||||
{
|
||||
name: "years and months",
|
||||
input: "1y6mo",
|
||||
expected: 365*24*time.Hour + 6*30*24*time.Hour,
|
||||
},
|
||||
{
|
||||
name: "days and hours",
|
||||
input: "1d12h",
|
||||
expected: 24*time.Hour + 12*time.Hour,
|
||||
},
|
||||
{
|
||||
name: "complex combination",
|
||||
input: "1y2mo3w4d5h6m7s",
|
||||
expected: 365*24*time.Hour + 2*30*24*time.Hour +
|
||||
3*7*24*time.Hour + 4*24*time.Hour +
|
||||
5*time.Hour + 6*time.Minute + 7*time.Second,
|
||||
},
|
||||
{
|
||||
name: "with spaces",
|
||||
input: "1d 12h 30m",
|
||||
expected: 24*time.Hour + 12*time.Hour + 30*time.Minute,
|
||||
},
|
||||
// Edge cases
|
||||
{
|
||||
name: "zero duration",
|
||||
input: "0s",
|
||||
expected: 0,
|
||||
},
|
||||
{
|
||||
name: "large duration",
|
||||
input: "10y",
|
||||
expected: 10 * 365 * 24 * time.Hour,
|
||||
},
|
||||
// Error cases
|
||||
{
|
||||
name: "empty string",
|
||||
input: "",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "invalid format",
|
||||
input: "abc",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "unknown unit",
|
||||
input: "5x",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "invalid number",
|
||||
input: "xyzd",
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "negative not supported",
|
||||
input: "-5d",
|
||||
wantErr: true,
|
||||
},
|
||||
})
|
||||
}
|
||||
|
||||
func TestParseDurationSpecialCases(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// Test that standard Go durations work exactly as expected
|
||||
standardDurations := []string{
|
||||
"300ms",
|
||||
"1.5h",
|
||||
"2h45m",
|
||||
"72h",
|
||||
"1us",
|
||||
"1µs",
|
||||
"1ns",
|
||||
}
|
||||
|
||||
for _, d := range standardDurations {
|
||||
expected, err := time.ParseDuration(d)
|
||||
require.NoError(t, err)
|
||||
|
||||
got, err := parseDuration(d)
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, expected, got, "standard duration %q should parse identically", d)
|
||||
}
|
||||
}
|
||||
|
||||
func TestParseDurationRealWorldExamples(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// Test real-world snapshot purge scenarios
|
||||
tests := []struct {
|
||||
description string
|
||||
input string
|
||||
olderThan time.Duration
|
||||
}{
|
||||
{
|
||||
description: "keep snapshots from last 30 days",
|
||||
input: "30d",
|
||||
olderThan: 30 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
description: "keep snapshots from last 6 months",
|
||||
input: "6mo",
|
||||
olderThan: 6 * 30 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
description: "keep snapshots from last year",
|
||||
input: "1y",
|
||||
olderThan: 365 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
description: "keep snapshots from last week and a half",
|
||||
input: "1w3d",
|
||||
olderThan: 10 * 24 * time.Hour,
|
||||
},
|
||||
{
|
||||
description: "keep snapshots from last 90 days",
|
||||
input: "90d",
|
||||
olderThan: 90 * 24 * time.Hour,
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.description, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
got, err := parseDuration(tt.input)
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, tt.olderThan, got)
|
||||
|
||||
// Verify the duration makes sense for snapshot purging
|
||||
assert.Greater(t, got, time.Hour,
|
||||
"snapshot purge duration should be at least an hour")
|
||||
})
|
||||
}
|
||||
}
|
||||
+17
-2
@@ -1,6 +1,7 @@
|
||||
package cli
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"io"
|
||||
"os"
|
||||
"strings"
|
||||
@@ -19,7 +20,11 @@ const shortCommitLen = 12
|
||||
// flag is present in os.Args — see bannerSuppressedInArgs), executes the
|
||||
// root cobra command, and routes any returned error through the
|
||||
// ui.Writer so the user sees a properly formatted "🛑 ERROR:" line.
|
||||
func Entry() {
|
||||
//
|
||||
// It returns the process exit code (0 on success, 1 on error) rather
|
||||
// than calling os.Exit, so that main's deferred profile writers run
|
||||
// before the process ends. See run in cmd/vaultik/main.go.
|
||||
func Entry() int {
|
||||
emitStartupBanner(os.Args[1:], os.Stdout)
|
||||
|
||||
rootCmd := NewRootCommand()
|
||||
@@ -27,9 +32,19 @@ func Entry() {
|
||||
|
||||
err := rootCmd.Execute()
|
||||
if err != nil {
|
||||
// An operation that ran inside the fx app has already reported
|
||||
// its own failure (and suppressed it under --json); errReported
|
||||
// says so. Printing it again here would double the error line.
|
||||
// Every other error — bad arguments, a config that would not
|
||||
// load — reaches Entry unreported, so it is shown here.
|
||||
if !errors.Is(err, errReported) {
|
||||
ReportErrorf("%s", err.Error())
|
||||
os.Exit(1)
|
||||
}
|
||||
|
||||
return 1
|
||||
}
|
||||
|
||||
return 0
|
||||
}
|
||||
|
||||
// emitStartupBanner writes the startup banner to w unless args (the
|
||||
|
||||
@@ -230,7 +230,7 @@ func TestEntryJSONStdoutIsExactlyOneDocument(t *testing.T) {
|
||||
programName, flagConfig, configPath, cmdSnapshot, cmdList, flagJSON,
|
||||
}
|
||||
|
||||
stdout := captureProcessStdout(t, Entry)
|
||||
stdout := captureProcessStdout(t, func() { _ = Entry() })
|
||||
|
||||
requireExactlyOneJSONDocument(t, stdout)
|
||||
|
||||
|
||||
@@ -0,0 +1,140 @@
|
||||
package cli //nolint:testpackage // shares the prune fixtures and capture helpers
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"io"
|
||||
"os"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
)
|
||||
|
||||
// staleRecordLogMessage is the local-cleanup audit line CleanupLocalSnapshots
|
||||
// logs for each stale record. It is exactly the signal issue #112 says a
|
||||
// machine consumer lost under --json: gated off stdout, and pinned below
|
||||
// the log level on stderr because --json used to force Quiet.
|
||||
const staleRecordLogMessage = "Removing stale local snapshot record"
|
||||
|
||||
// TestEntryPruneJSONStderrHonoursVerbosity is the end-to-end regression
|
||||
// guard for issue #112. Under --json the log level must still follow
|
||||
// --verbose/--debug rather than being pinned to WARN, so the
|
||||
// local-cleanup records reach stderr under --verbose while stdout stays
|
||||
// exactly one JSON document; without --verbose they stay below the
|
||||
// level, as they do without --json.
|
||||
//
|
||||
// Both halves are asserted together on the same run, because the fix has
|
||||
// to keep the document clean (issue #108) while freeing stderr.
|
||||
//
|
||||
// Not parallel: it replaces os.Args, os.Stdout, os.Stderr and the xdg
|
||||
// globals.
|
||||
//
|
||||
//nolint:paralleltest // replaces os.Args, os.Stdout, os.Stderr and the xdg globals
|
||||
func TestEntryPruneJSONStderrHonoursVerbosity(t *testing.T) {
|
||||
for _, testCase := range []struct {
|
||||
name string
|
||||
verbose bool
|
||||
wantOnStderr bool
|
||||
}{
|
||||
{
|
||||
name: "verbose json surfaces the cleanup record on stderr",
|
||||
verbose: true,
|
||||
wantOnStderr: true,
|
||||
},
|
||||
{
|
||||
name: "json alone keeps the cleanup record below the level",
|
||||
verbose: false,
|
||||
wantOnStderr: false,
|
||||
},
|
||||
} {
|
||||
t.Run(testCase.name, func(t *testing.T) {
|
||||
configPath := writeHermeticPruneConfig(t, true)
|
||||
|
||||
previousArgs := os.Args
|
||||
|
||||
t.Cleanup(func() {
|
||||
os.Args = previousArgs
|
||||
rootFlags = RootFlags{}
|
||||
})
|
||||
|
||||
args := []string{
|
||||
programName, flagConfig, configPath, cmdPrune, flagJSON,
|
||||
}
|
||||
if testCase.verbose {
|
||||
args = append(args, "--verbose")
|
||||
}
|
||||
|
||||
os.Args = args
|
||||
|
||||
stdout, stderr := captureProcessStdoutAndStderr(t,
|
||||
func() { _ = Entry() })
|
||||
|
||||
// The document stays clean in both cases: freeing stderr must
|
||||
// not regress issue #108.
|
||||
requireExactlyOneJSONDocument(t, stdout)
|
||||
|
||||
if testCase.wantOnStderr {
|
||||
assert.Contains(t, stderr, staleRecordLogMessage,
|
||||
"--verbose --json must emit the cleanup record on stderr")
|
||||
assert.Contains(t, stderr, stalePruneSnapshotID,
|
||||
"the record must name the snapshot it removed")
|
||||
} else {
|
||||
assert.NotContains(t, stderr, staleRecordLogMessage,
|
||||
"without --verbose the record stays below the log level")
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// captureProcessStdoutAndStderr redirects both of the process's own
|
||||
// standard streams to pipes for the duration of fn and returns what was
|
||||
// written to each. The redirection is at the file-descriptor level
|
||||
// because the logger binds os.Stderr when it initializes inside fn, and
|
||||
// the JSON document reaches os.Stdout independently; the point is to see
|
||||
// where each actually lands.
|
||||
//
|
||||
// Not parallel-safe: os.Stdout and os.Stderr are process-global.
|
||||
func captureProcessStdoutAndStderr(t *testing.T, fn func()) (string, string) {
|
||||
t.Helper()
|
||||
|
||||
outReader, outWriter, err := os.Pipe()
|
||||
require.NoError(t, err)
|
||||
|
||||
errReader, errWriter, err := os.Pipe()
|
||||
require.NoError(t, err)
|
||||
|
||||
previousOut, previousErr := os.Stdout, os.Stderr
|
||||
os.Stdout, os.Stderr = outWriter, errWriter
|
||||
|
||||
capturedOut := drain(outReader)
|
||||
capturedErr := drain(errReader)
|
||||
|
||||
fn()
|
||||
|
||||
os.Stdout, os.Stderr = previousOut, previousErr
|
||||
|
||||
require.NoError(t, outWriter.Close())
|
||||
require.NoError(t, errWriter.Close())
|
||||
|
||||
out, errOut := <-capturedOut, <-capturedErr
|
||||
|
||||
require.NoError(t, outReader.Close())
|
||||
require.NoError(t, errReader.Close())
|
||||
|
||||
return out, errOut
|
||||
}
|
||||
|
||||
// drain copies a reader to a string on a goroutine and delivers the
|
||||
// result once the writer end is closed.
|
||||
func drain(reader io.Reader) <-chan string {
|
||||
captured := make(chan string, 1)
|
||||
|
||||
go func() {
|
||||
var buf bytes.Buffer
|
||||
|
||||
_, _ = io.Copy(&buf, reader)
|
||||
captured <- buf.String()
|
||||
}()
|
||||
|
||||
return captured
|
||||
}
|
||||
@@ -81,7 +81,7 @@ func TestEntryPruneJSONStdoutIsExactlyOneDocument(t *testing.T) {
|
||||
programName, flagConfig, configPath, cmdPrune, flagJSON,
|
||||
}
|
||||
|
||||
stdout := captureProcessStdout(t, Entry)
|
||||
stdout := captureProcessStdout(t, func() { _ = Entry() })
|
||||
|
||||
requireExactlyOneJSONDocument(t, stdout)
|
||||
|
||||
|
||||
@@ -0,0 +1,58 @@
|
||||
package cli //nolint:testpackage // shares programName and the capture helpers
|
||||
|
||||
import (
|
||||
"os"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
)
|
||||
|
||||
// TestEntryReturnsStatusCode pins the contract main() relies on for
|
||||
// issue #75: Entry reports success or failure through its return value
|
||||
// and never calls os.Exit. An os.Exit from inside Entry would skip
|
||||
// main's deferred profile writers and truncate the profile of a failing
|
||||
// command. main turns this code into os.Exit only after those defers
|
||||
// run, so a failing command must come back with a non-zero code rather
|
||||
// than ending the process here.
|
||||
//
|
||||
// Stdout is captured only to keep the banner and command output off the
|
||||
// test log; the assertion is on the returned code.
|
||||
//
|
||||
//nolint:paralleltest // replaces os.Args and rootFlags
|
||||
func TestEntryReturnsStatusCode(t *testing.T) {
|
||||
for _, testCase := range []struct {
|
||||
name string
|
||||
args []string
|
||||
want int
|
||||
}{
|
||||
{
|
||||
// version is self-contained: it needs no config and no
|
||||
// destination store, so it exercises the success path.
|
||||
name: "successful command returns zero",
|
||||
args: []string{programName, "version"},
|
||||
want: 0,
|
||||
},
|
||||
{
|
||||
name: "unknown command returns one",
|
||||
args: []string{programName, "no-such-command"},
|
||||
want: 1,
|
||||
},
|
||||
} {
|
||||
t.Run(testCase.name, func(t *testing.T) {
|
||||
previousArgs := os.Args
|
||||
|
||||
t.Cleanup(func() {
|
||||
os.Args = previousArgs
|
||||
rootFlags = RootFlags{}
|
||||
})
|
||||
|
||||
os.Args = testCase.args
|
||||
|
||||
var code int
|
||||
|
||||
_ = captureProcessStdout(t, func() { code = Entry() })
|
||||
|
||||
assert.Equal(t, testCase.want, code)
|
||||
})
|
||||
}
|
||||
}
|
||||
+5
-35
@@ -1,12 +1,7 @@
|
||||
package cli
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
|
||||
"github.com/spf13/cobra"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
@@ -33,44 +28,19 @@ func NewInfoCommand() *cobra.Command {
|
||||
// Use the app framework
|
||||
rootFlags := GetRootFlags()
|
||||
|
||||
return RunWithApp(cmd.Context(), AppOptions{
|
||||
return RunOperation(cmd.Context(), AppOptions{
|
||||
ConfigPath: configPath,
|
||||
LogOptions: log.Options{
|
||||
Verbose: rootFlags.Verbose,
|
||||
Debug: rootFlags.Debug,
|
||||
Quiet: rootFlags.Quiet,
|
||||
},
|
||||
Modules: []fx.Option{},
|
||||
Invokes: []fx.Option{
|
||||
fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
|
||||
lc.Append(fx.Hook{
|
||||
OnStart: func(_ context.Context) error {
|
||||
go func() {
|
||||
err := v.ShowInfo()
|
||||
if err != nil {
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
Mode: readOnly,
|
||||
}, func(v *vaultik.Vaultik) error {
|
||||
return v.ShowInfo()
|
||||
}, func(err error) {
|
||||
log.Error("Failed to show info", "error", err)
|
||||
ReportErrorf("Failed to show info: %v", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
err = v.Shutdowner.Shutdown()
|
||||
if err != nil {
|
||||
log.Error("Failed to shutdown", "error", err)
|
||||
}
|
||||
}()
|
||||
|
||||
return nil
|
||||
},
|
||||
OnStop: func(_ context.Context) error {
|
||||
v.Cancel()
|
||||
|
||||
return nil
|
||||
},
|
||||
})
|
||||
}),
|
||||
},
|
||||
})
|
||||
},
|
||||
}
|
||||
|
||||
+11
-43
@@ -1,12 +1,7 @@
|
||||
package cli
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
|
||||
"github.com/spf13/cobra"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
@@ -41,51 +36,24 @@ work (e.g. after a crashed backup or to reclaim storage).`,
|
||||
// Use the app framework like other commands
|
||||
rootFlags := GetRootFlags()
|
||||
|
||||
return RunWithApp(cmd.Context(), AppOptions{
|
||||
return RunOperation(cmd.Context(), AppOptions{
|
||||
ConfigPath: configPath,
|
||||
LogOptions: log.Options{
|
||||
Verbose: rootFlags.Verbose,
|
||||
Debug: rootFlags.Debug,
|
||||
Quiet: rootFlags.Quiet || opts.JSON,
|
||||
Quiet: rootFlags.Quiet,
|
||||
JSON: opts.JSON,
|
||||
},
|
||||
Modules: []fx.Option{},
|
||||
Invokes: []fx.Option{
|
||||
fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
|
||||
lc.Append(fx.Hook{
|
||||
OnStart: func(_ context.Context) error {
|
||||
// Start the prune operation in a goroutine
|
||||
go func() {
|
||||
// Run the prune operation
|
||||
err := v.Prune(opts)
|
||||
if err != nil {
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
if !opts.JSON {
|
||||
Mode: mutating,
|
||||
}, func(v *vaultik.Vaultik) error {
|
||||
return v.Prune(opts)
|
||||
}, func(err error) {
|
||||
if opts.JSON {
|
||||
return
|
||||
}
|
||||
|
||||
log.Error("Prune operation failed", "error", err)
|
||||
ReportErrorf("Prune failed: %v", err)
|
||||
}
|
||||
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
// Shutdown the app when prune completes
|
||||
err = v.Shutdowner.Shutdown()
|
||||
if err != nil {
|
||||
log.Error("Failed to shutdown", "error", err)
|
||||
}
|
||||
}()
|
||||
|
||||
return nil
|
||||
},
|
||||
OnStop: func(_ context.Context) error {
|
||||
log.Debug("Stopping prune operation")
|
||||
v.Cancel()
|
||||
|
||||
return nil
|
||||
},
|
||||
})
|
||||
}),
|
||||
},
|
||||
})
|
||||
},
|
||||
}
|
||||
|
||||
+12
-38
@@ -1,12 +1,9 @@
|
||||
package cli
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
|
||||
"github.com/spf13/cobra"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
@@ -48,7 +45,7 @@ This is destructive and irreversible. Requires --force.`,
|
||||
return errNukeNeedsForce
|
||||
}
|
||||
|
||||
return runVaultikApp(cmd, false, false, "Remote nuke failed",
|
||||
return runVaultikApp(cmd, mutating, false, false, "Remote nuke failed",
|
||||
func(v *vaultik.Vaultik) error {
|
||||
return v.NukeRemote(true)
|
||||
})
|
||||
@@ -83,47 +80,24 @@ func newRemoteInfoCommand() *cobra.Command {
|
||||
|
||||
rootFlags := GetRootFlags()
|
||||
|
||||
return RunWithApp(cmd.Context(), AppOptions{
|
||||
return RunOperation(cmd.Context(), AppOptions{
|
||||
ConfigPath: configPath,
|
||||
LogOptions: log.Options{
|
||||
Verbose: rootFlags.Verbose,
|
||||
Debug: rootFlags.Debug,
|
||||
Quiet: rootFlags.Quiet || jsonOutput,
|
||||
Quiet: rootFlags.Quiet,
|
||||
JSON: jsonOutput,
|
||||
},
|
||||
Modules: []fx.Option{},
|
||||
Invokes: []fx.Option{
|
||||
fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
|
||||
lc.Append(fx.Hook{
|
||||
OnStart: func(_ context.Context) error {
|
||||
go func() {
|
||||
err := v.RemoteInfo(jsonOutput)
|
||||
if err != nil {
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
if !jsonOutput {
|
||||
Mode: readOnly,
|
||||
}, func(v *vaultik.Vaultik) error {
|
||||
return v.RemoteInfo(jsonOutput)
|
||||
}, func(err error) {
|
||||
if jsonOutput {
|
||||
return
|
||||
}
|
||||
|
||||
log.Error("Failed to get remote info", "error", err)
|
||||
ReportErrorf("Failed to get remote info: %v", err)
|
||||
}
|
||||
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
err = v.Shutdowner.Shutdown()
|
||||
if err != nil {
|
||||
log.Error("Failed to shutdown", "error", err)
|
||||
}
|
||||
}()
|
||||
|
||||
return nil
|
||||
},
|
||||
OnStop: func(_ context.Context) error {
|
||||
v.Cancel()
|
||||
|
||||
return nil
|
||||
},
|
||||
})
|
||||
}),
|
||||
},
|
||||
})
|
||||
},
|
||||
}
|
||||
|
||||
@@ -57,8 +57,9 @@ on the source system.`,
|
||||
cmd.PersistentFlags().BoolVarP(&rootFlags.Quiet, "quiet", "q", false,
|
||||
"Suppress non-error output")
|
||||
cmd.PersistentFlags().BoolVar(&rootFlags.SkipErrors, "skip-errors", false,
|
||||
"Continue past per-file errors instead of aborting "+
|
||||
"(applies to snapshot create and restore)")
|
||||
"Skip files that cannot be read when creating a snapshot, or "+
|
||||
"that cannot be restored when restoring, instead of aborting "+
|
||||
"(packing and storage errors still abort)")
|
||||
|
||||
// Add subcommands
|
||||
cmd.AddCommand(
|
||||
|
||||
+29
-79
@@ -1,13 +1,10 @@
|
||||
package cli
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"os"
|
||||
|
||||
"github.com/spf13/cobra"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
@@ -86,7 +83,8 @@ specifying a path using --config or by setting VAULTIK_CONFIG to a path.`,
|
||||
// Use the backup functionality from cli package
|
||||
rootFlags := GetRootFlags()
|
||||
|
||||
return RunWithApp(cmd.Context(), AppOptions{
|
||||
// --cron suppression is wired through v.UI by setupGlobals.
|
||||
return RunOperation(cmd.Context(), AppOptions{
|
||||
ConfigPath: configPath,
|
||||
LogOptions: log.Options{
|
||||
Verbose: rootFlags.Verbose,
|
||||
@@ -94,54 +92,25 @@ specifying a path using --config or by setting VAULTIK_CONFIG to a path.`,
|
||||
Cron: opts.Cron,
|
||||
Quiet: rootFlags.Quiet,
|
||||
},
|
||||
Modules: []fx.Option{},
|
||||
Invokes: []fx.Option{
|
||||
fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
|
||||
lc.Append(fx.Hook{
|
||||
OnStart: func(_ context.Context) error {
|
||||
// Start the snapshot creation in a goroutine
|
||||
go func() {
|
||||
// --cron suppression is wired through v.UI by setupGlobals.
|
||||
err := v.CreateSnapshot(opts)
|
||||
if err != nil {
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
Mode: mutating,
|
||||
}, func(v *vaultik.Vaultik) error {
|
||||
return v.CreateSnapshot(opts)
|
||||
}, func(err error) {
|
||||
log.Error("Snapshot creation failed", "error", err)
|
||||
ReportErrorf("Snapshot creation failed: %v", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
// Shutdown the app when snapshot completes
|
||||
err = v.Shutdowner.Shutdown()
|
||||
if err != nil {
|
||||
log.Error("Failed to shutdown", "error", err)
|
||||
}
|
||||
}()
|
||||
|
||||
return nil
|
||||
},
|
||||
OnStop: func(_ context.Context) error {
|
||||
log.Debug("Stopping snapshot creation")
|
||||
// Cancel the Vaultik context
|
||||
v.Cancel()
|
||||
|
||||
return nil
|
||||
},
|
||||
})
|
||||
}),
|
||||
},
|
||||
})
|
||||
},
|
||||
}
|
||||
|
||||
cmd.Flags().BoolVar(&opts.Cron, "cron", false,
|
||||
"Run in cron mode (silent unless error)")
|
||||
"Run in cron mode (silent unless warning or error)")
|
||||
cmd.Flags().BoolVar(&opts.Prune, "prune", false,
|
||||
"After backup, drop older snapshots of the same name and remove "+
|
||||
"orphaned blobs")
|
||||
cmd.Flags().StringVar(&opts.KeepNewerThan, "keep-newer-than", "",
|
||||
"With --prune: keep snapshots newer than this duration "+
|
||||
"(e.g. 4w, 30d, 6mo) instead of only the latest")
|
||||
"(e.g. 30d, 4w, 6mo, 1y; m is minutes, mo is months) "+
|
||||
"instead of only the latest")
|
||||
|
||||
return cmd
|
||||
}
|
||||
@@ -157,7 +126,7 @@ func newSnapshotListCommand() *cobra.Command {
|
||||
Long: "Lists all snapshots with their ID, timestamp, and compressed size",
|
||||
Args: cobra.NoArgs,
|
||||
RunE: func(cmd *cobra.Command, _ []string) error {
|
||||
return runVaultikApp(cmd, false, false,
|
||||
return runVaultikApp(cmd, readOnly, false, false,
|
||||
"Failed to list snapshots",
|
||||
func(v *vaultik.Vaultik) error {
|
||||
return v.ListSnapshots(jsonOutput)
|
||||
@@ -193,7 +162,7 @@ restrict the operation to specific snapshot names.`,
|
||||
return errPurgeCriteriaBoth
|
||||
}
|
||||
|
||||
return runVaultikApp(cmd, false, false,
|
||||
return runVaultikApp(cmd, mutating, false, false,
|
||||
"Failed to purge snapshots",
|
||||
func(v *vaultik.Vaultik) error {
|
||||
return v.PurgeSnapshotsWithOptions(opts)
|
||||
@@ -204,7 +173,8 @@ restrict the operation to specific snapshot names.`,
|
||||
cmd.Flags().BoolVar(&opts.KeepLatest, "keep-latest", false,
|
||||
"Keep only the latest snapshot of each name")
|
||||
cmd.Flags().StringVar(&opts.OlderThan, "older-than", "",
|
||||
"Remove snapshots older than duration (e.g., 30d, 6m, 1y)")
|
||||
"Remove snapshots older than duration "+
|
||||
"(e.g. 30d, 4w, 6mo, 1y; m is minutes, mo is months)")
|
||||
cmd.Flags().BoolVar(&opts.Force, "force", false, "Skip confirmation prompt")
|
||||
cmd.Flags().StringArrayVar(&opts.Names, "snapshot", nil,
|
||||
"Restrict to snapshots with these names (repeat for multiple)")
|
||||
@@ -219,7 +189,10 @@ func newSnapshotVerifyCommand() *cobra.Command {
|
||||
cmd := &cobra.Command{
|
||||
Use: "verify <snapshot-id>",
|
||||
Short: "Verify snapshot integrity",
|
||||
Long: "Verifies that all blobs referenced in a snapshot exist",
|
||||
Long: "Verifies that all blobs referenced in a snapshot exist.\n\n" +
|
||||
"The snapshot may be named by its ID or, on a host with no local\n" +
|
||||
"index, by the remote key that 'snapshot list' prints for a\n" +
|
||||
"remote-only snapshot (an unambiguous leading part is enough).",
|
||||
Args: requireSnapshotIDArg,
|
||||
RunE: func(cmd *cobra.Command, args []string) error {
|
||||
snapshotID := args[0]
|
||||
@@ -232,47 +205,24 @@ func newSnapshotVerifyCommand() *cobra.Command {
|
||||
|
||||
rootFlags := GetRootFlags()
|
||||
|
||||
return RunWithApp(cmd.Context(), AppOptions{
|
||||
return RunOperation(cmd.Context(), AppOptions{
|
||||
ConfigPath: configPath,
|
||||
LogOptions: log.Options{
|
||||
Verbose: rootFlags.Verbose,
|
||||
Debug: rootFlags.Debug,
|
||||
Quiet: rootFlags.Quiet || opts.JSON,
|
||||
Quiet: rootFlags.Quiet,
|
||||
JSON: opts.JSON,
|
||||
},
|
||||
Modules: []fx.Option{},
|
||||
Invokes: []fx.Option{
|
||||
fx.Invoke(func(v *vaultik.Vaultik, lc fx.Lifecycle) {
|
||||
lc.Append(fx.Hook{
|
||||
OnStart: func(_ context.Context) error {
|
||||
go func() {
|
||||
err := v.VerifySnapshotWithOptions(snapshotID, opts)
|
||||
if err != nil {
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
if !opts.JSON {
|
||||
Mode: readOnly,
|
||||
}, func(v *vaultik.Vaultik) error {
|
||||
return v.VerifySnapshotWithOptions(snapshotID, opts)
|
||||
}, func(err error) {
|
||||
if opts.JSON {
|
||||
return
|
||||
}
|
||||
|
||||
log.Error("Verification failed", "error", err)
|
||||
ReportErrorf("Verification failed: %v", err)
|
||||
}
|
||||
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
err = v.Shutdowner.Shutdown()
|
||||
if err != nil {
|
||||
log.Error("Failed to shutdown", "error", err)
|
||||
}
|
||||
}()
|
||||
|
||||
return nil
|
||||
},
|
||||
OnStop: func(_ context.Context) error {
|
||||
v.Cancel()
|
||||
|
||||
return nil
|
||||
},
|
||||
})
|
||||
}),
|
||||
},
|
||||
})
|
||||
},
|
||||
}
|
||||
@@ -311,7 +261,7 @@ To wipe the entire destination store and start over, use 'vaultik remote
|
||||
nuke --force' — it is the single supported entry point for that.`,
|
||||
Args: requireSnapshotIDArg,
|
||||
RunE: func(cmd *cobra.Command, args []string) error {
|
||||
return runVaultikApp(cmd, opts.JSON, opts.JSON,
|
||||
return runVaultikApp(cmd, mutating, opts.JSON, opts.JSON,
|
||||
"Failed to remove snapshot",
|
||||
func(v *vaultik.Vaultik) error {
|
||||
_, err := v.RemoveSnapshot(args[0], opts)
|
||||
|
||||
@@ -1,16 +1,8 @@
|
||||
package cli
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
|
||||
"github.com/spf13/cobra"
|
||||
"go.uber.org/fx"
|
||||
"sneak.berlin/go/vaultik/internal/config"
|
||||
"sneak.berlin/go/vaultik/internal/globals"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
|
||||
@@ -25,15 +17,6 @@ type RestoreOptions struct {
|
||||
Verify bool // Verify restored files after restore
|
||||
}
|
||||
|
||||
// RestoreApp contains all dependencies needed for restore
|
||||
type RestoreApp struct {
|
||||
Globals *globals.Globals
|
||||
Config *config.Config
|
||||
Storage storage.Storer
|
||||
Vaultik *vaultik.Vaultik
|
||||
Shutdowner fx.Shutdowner
|
||||
}
|
||||
|
||||
// newSnapshotRestoreCommand creates the 'snapshot restore' subcommand
|
||||
func newSnapshotRestoreCommand() *cobra.Command {
|
||||
opts := &RestoreOptions{}
|
||||
@@ -48,6 +31,10 @@ target directory.
|
||||
If no paths are specified, all files are restored.
|
||||
If paths are specified, only matching files/directories are restored.
|
||||
|
||||
The snapshot may be named by its ID or, when restoring on a host with no
|
||||
local index, by the remote key that 'snapshot list' prints for a
|
||||
remote-only snapshot (an unambiguous leading part is enough).
|
||||
|
||||
Requires the VAULTIK_AGE_SECRET_KEY environment variable to be set with
|
||||
the age private key.
|
||||
|
||||
@@ -77,7 +64,8 @@ Examples:
|
||||
return cmd
|
||||
}
|
||||
|
||||
// runRestore parses arguments and runs the restore operation through the app framework
|
||||
// runRestore parses arguments and runs the restore operation through the
|
||||
// app framework.
|
||||
func runRestore(cmd *cobra.Command, args []string, opts *RestoreOptions) error {
|
||||
snapshotID := args[0]
|
||||
|
||||
@@ -86,87 +74,31 @@ func runRestore(cmd *cobra.Command, args []string, opts *RestoreOptions) error {
|
||||
opts.Paths = args[restoreMinArgs:]
|
||||
}
|
||||
|
||||
// Use unified config resolution
|
||||
configPath, err := ResolveConfigPath()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
// Use the app framework like other commands
|
||||
rootFlags := GetRootFlags()
|
||||
|
||||
return RunWithApp(cmd.Context(), AppOptions{
|
||||
return RunOperation(cmd.Context(), AppOptions{
|
||||
ConfigPath: configPath,
|
||||
LogOptions: log.Options{
|
||||
Verbose: rootFlags.Verbose,
|
||||
Debug: rootFlags.Debug,
|
||||
Quiet: rootFlags.Quiet,
|
||||
},
|
||||
Modules: buildRestoreModules(),
|
||||
Invokes: buildRestoreInvokes(snapshotID, opts),
|
||||
})
|
||||
}
|
||||
|
||||
// buildRestoreModules returns the fx.Options for dependency injection in restore
|
||||
func buildRestoreModules() []fx.Option {
|
||||
return []fx.Option{
|
||||
fx.Provide(fx.Annotate(
|
||||
func(g *globals.Globals, cfg *config.Config,
|
||||
storer storage.Storer, v *vaultik.Vaultik, shutdowner fx.Shutdowner) *RestoreApp {
|
||||
return &RestoreApp{
|
||||
Globals: g,
|
||||
Config: cfg,
|
||||
Storage: storer,
|
||||
Vaultik: v,
|
||||
Shutdowner: shutdowner,
|
||||
}
|
||||
},
|
||||
)),
|
||||
}
|
||||
}
|
||||
|
||||
// buildRestoreInvokes returns the fx.Options that wire up the restore lifecycle
|
||||
func buildRestoreInvokes(snapshotID string, opts *RestoreOptions) []fx.Option {
|
||||
return []fx.Option{
|
||||
fx.Invoke(func(app *RestoreApp, lc fx.Lifecycle) {
|
||||
lc.Append(fx.Hook{
|
||||
OnStart: func(_ context.Context) error {
|
||||
// Start the restore operation in a goroutine
|
||||
go func() {
|
||||
// Run the restore operation
|
||||
restoreOpts := &vaultik.RestoreOptions{
|
||||
Mode: readOnly,
|
||||
}, func(v *vaultik.Vaultik) error {
|
||||
return v.Restore(&vaultik.RestoreOptions{
|
||||
SnapshotID: snapshotID,
|
||||
TargetDir: opts.TargetDir,
|
||||
Paths: opts.Paths,
|
||||
Verify: opts.Verify,
|
||||
SkipErrors: GetRootFlags().SkipErrors,
|
||||
}
|
||||
|
||||
err := app.Vaultik.Restore(restoreOpts)
|
||||
if err != nil {
|
||||
if !errors.Is(err, context.Canceled) {
|
||||
SkipErrors: rootFlags.SkipErrors,
|
||||
})
|
||||
}, func(err error) {
|
||||
log.Error("Restore operation failed", "error", err)
|
||||
ReportErrorf("Restore failed: %v", err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
|
||||
// Shutdown the app when restore completes
|
||||
err = app.Shutdowner.Shutdown()
|
||||
if err != nil {
|
||||
log.Error("Failed to shutdown", "error", err)
|
||||
}
|
||||
}()
|
||||
|
||||
return nil
|
||||
},
|
||||
OnStop: func(_ context.Context) error {
|
||||
log.Debug("Stopping restore operation")
|
||||
app.Vaultik.Cancel()
|
||||
|
||||
return nil
|
||||
},
|
||||
})
|
||||
}),
|
||||
}
|
||||
}
|
||||
|
||||
@@ -1,12 +0,0 @@
|
||||
package cli
|
||||
|
||||
import "time"
|
||||
|
||||
// SnapshotInfo represents snapshot information for listing
|
||||
//
|
||||
//nolint:tagliatelle // snake_case is the established output format
|
||||
type SnapshotInfo struct {
|
||||
ID string `json:"id"`
|
||||
Timestamp time.Time `json:"timestamp"`
|
||||
CompressedSize int64 `json:"compressed_size"`
|
||||
}
|
||||
@@ -16,6 +16,7 @@ import (
|
||||
"github.com/adrg/xdg"
|
||||
"go.uber.org/fx"
|
||||
"gopkg.in/yaml.v3"
|
||||
"sneak.berlin/go/vaultik/internal/chunker"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
)
|
||||
|
||||
@@ -41,7 +42,9 @@ var (
|
||||
"at least one snapshot must be configured (see config.example.yml)")
|
||||
errSnapshotNoPaths = errors.New("snapshot must have at least one path")
|
||||
errChunkSizeTooSmall = errors.New("chunk_size must be at least 1MB")
|
||||
errBlobSizeTooSmall = errors.New("blob_size_limit must be at least chunk_size")
|
||||
errBlobSizeTooSmall = errors.New(
|
||||
"blob_size_limit must be at least the largest chunk the chunker can " +
|
||||
"emit (chunk_size times the FastCDC size spread)")
|
||||
errBadCompression = errors.New("compression_level must be between 1 and 19")
|
||||
errBadStorageScheme = errors.New(
|
||||
"storage_url must start with s3://, file://, or rclone://")
|
||||
@@ -162,7 +165,9 @@ type S3Config struct {
|
||||
AccessKeyID string `yaml:"access_key_id"`
|
||||
SecretAccessKey string `yaml:"secret_access_key"`
|
||||
Region string `yaml:"region"`
|
||||
UseSSL bool `yaml:"use_ssl"`
|
||||
// UseSSL selects HTTPS for a scheme-less endpoint. Omitted (nil) means
|
||||
// the default, TLS; set it to false only to force plain HTTP.
|
||||
UseSSL *bool `yaml:"use_ssl"`
|
||||
PartSize Size `yaml:"part_size"`
|
||||
}
|
||||
|
||||
@@ -289,8 +294,11 @@ func Load(path string) (*Config, error) {
|
||||
// - At least one snapshot must be configured with at least one path
|
||||
// - Storage must be configured (either storage_url or s3.* fields)
|
||||
// - Chunk size must be at least 1MB
|
||||
// - Blob size limit must be at least the chunk size
|
||||
// - Blob size limit must be at least the largest chunk the chunker can emit
|
||||
// (chunk_size times chunker.ChunkSizeSpread), so a single-chunk blob never
|
||||
// exceeds the configured limit
|
||||
// - Compression level must be between 1 and 19
|
||||
//
|
||||
// Returns an error describing the first validation failure encountered.
|
||||
func (c *Config) Validate() error {
|
||||
if len(c.AgeRecipients) == 0 {
|
||||
@@ -317,8 +325,13 @@ func (c *Config) Validate() error {
|
||||
return errChunkSizeTooSmall
|
||||
}
|
||||
|
||||
if c.BlobSizeLimit.Int64() < c.ChunkSize.Int64() {
|
||||
return errBlobSizeTooSmall
|
||||
// The chunker can emit chunks up to chunk_size * ChunkSizeSpread, and the
|
||||
// packer places a single such chunk into an otherwise empty blob. A limit
|
||||
// below that bound would let a blob exceed it, so reject it.
|
||||
largestChunk := c.ChunkSize.Int64() * chunker.ChunkSizeSpread
|
||||
if c.BlobSizeLimit.Int64() < largestChunk {
|
||||
return fmt.Errorf("%w: need at least %d bytes",
|
||||
errBlobSizeTooSmall, largestChunk)
|
||||
}
|
||||
|
||||
if c.CompressionLevel < minCompressionLevel ||
|
||||
|
||||
@@ -1,9 +1,12 @@
|
||||
package config //nolint:testpackage // exercises unexported extractAgeSecretKey
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/chunker"
|
||||
)
|
||||
|
||||
const (
|
||||
@@ -101,6 +104,80 @@ func TestConfigFromEnv(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestValidateBlobSizeLimit checks the blob_size_limit boundary: it must be at
|
||||
// least the largest chunk the chunker can emit (chunk_size times
|
||||
// chunker.ChunkSizeSpread), because the packer places a single such chunk into
|
||||
// an otherwise empty blob. A limit between chunk_size and that bound is rejected.
|
||||
func TestValidateBlobSizeLimit(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const chunkSize = Size(10 * 1024 * 1024) // 10MB
|
||||
|
||||
largestChunk := chunkSize.Int64() * chunker.ChunkSizeSpread
|
||||
|
||||
newConfig := func(blobLimit Size) *Config {
|
||||
return &Config{
|
||||
AgeRecipients: []string{testSneakAgePublicKey},
|
||||
Snapshots: map[string]SnapshotConfig{"test": {Paths: []string{"/tmp/src"}}},
|
||||
StorageURL: "file:///tmp/vaultik-test-store",
|
||||
ChunkSize: chunkSize,
|
||||
BlobSizeLimit: blobLimit,
|
||||
CompressionLevel: 3,
|
||||
}
|
||||
}
|
||||
|
||||
tests := []struct {
|
||||
name string
|
||||
blobLimit Size
|
||||
wantErr bool
|
||||
}{
|
||||
{
|
||||
name: "at chunk_size but below largest chunk is rejected",
|
||||
blobLimit: chunkSize,
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "between chunk_size and largest chunk is rejected",
|
||||
blobLimit: Size(chunkSize.Int64() * 2),
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "one byte below largest chunk is rejected",
|
||||
blobLimit: Size(largestChunk - 1),
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "exactly at largest chunk is accepted",
|
||||
blobLimit: Size(largestChunk),
|
||||
wantErr: false,
|
||||
},
|
||||
{
|
||||
name: "above largest chunk is accepted",
|
||||
blobLimit: Size(largestChunk * 100),
|
||||
wantErr: false,
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
err := newConfig(tt.blobLimit).Validate()
|
||||
if tt.wantErr {
|
||||
if !errors.Is(err, errBlobSizeTooSmall) {
|
||||
t.Fatalf("Validate() error = %v, want errBlobSizeTooSmall", err)
|
||||
}
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
if err != nil {
|
||||
t.Fatalf("Validate() unexpected error: %v", err)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestExtractAgeSecretKey tests extraction of AGE-SECRET-KEY from various inputs
|
||||
func TestExtractAgeSecretKey(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
@@ -208,6 +208,30 @@ func (r *BlobRepository) DeleteOrphaned(ctx context.Context) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// DeleteUnuploaded deletes blob rows whose upload never completed
|
||||
// (uploaded_ts IS NULL) and returns how many were removed. Their
|
||||
// blob_chunks rows are removed by the ON DELETE CASCADE foreign key.
|
||||
// A blob is only ever attached to a snapshot once its upload has been
|
||||
// recorded, so an un-uploaded blob is never referenced by a completed
|
||||
// snapshot: dropping it discards chunk rows that point at data which
|
||||
// was never stored remotely, so the affected content is re-chunked and
|
||||
// re-uploaded on the next run.
|
||||
func (r *BlobRepository) DeleteUnuploaded(ctx context.Context) (int64, error) {
|
||||
query := `DELETE FROM blobs WHERE uploaded_ts IS NULL`
|
||||
|
||||
result, err := r.db.ExecWithLog(ctx, query)
|
||||
if err != nil {
|
||||
return 0, fmt.Errorf("deleting un-uploaded blobs: %w", err)
|
||||
}
|
||||
|
||||
rowsAffected, _ := result.RowsAffected()
|
||||
if rowsAffected > 0 {
|
||||
log.Debug("Deleted un-uploaded blobs", "count", rowsAffected)
|
||||
}
|
||||
|
||||
return rowsAffected, nil
|
||||
}
|
||||
|
||||
// getOne fetches a single blob row matched on the given column, or
|
||||
// (nil, nil) when no row matches.
|
||||
func (r *BlobRepository) getOne(
|
||||
|
||||
@@ -7,12 +7,32 @@ import (
|
||||
|
||||
// List returns every chunk in the index, ordered by chunk hash.
|
||||
func (r *ChunkRepository) List(ctx context.Context) ([]*Chunk, error) {
|
||||
query := `
|
||||
return r.list(ctx, `
|
||||
SELECT chunk_hash, size
|
||||
FROM chunks
|
||||
ORDER BY chunk_hash
|
||||
`
|
||||
`)
|
||||
}
|
||||
|
||||
// ListInUploadedBlobs returns the chunks that are stored in a blob whose
|
||||
// upload has completed (uploaded_ts set), ordered by chunk hash. These
|
||||
// are the only chunks a backup may safely deduplicate against: a chunk
|
||||
// recorded solely in a blob that was never uploaded refers to data that
|
||||
// is not in remote storage, so trusting it would silently drop that data
|
||||
// from later snapshots.
|
||||
func (r *ChunkRepository) ListInUploadedBlobs(ctx context.Context) ([]*Chunk, error) {
|
||||
return r.list(ctx, `
|
||||
SELECT DISTINCT c.chunk_hash, c.size
|
||||
FROM chunks c
|
||||
JOIN blob_chunks bc ON c.chunk_hash = bc.chunk_hash
|
||||
JOIN blobs b ON bc.blob_id = b.id
|
||||
WHERE b.uploaded_ts IS NOT NULL
|
||||
ORDER BY c.chunk_hash
|
||||
`)
|
||||
}
|
||||
|
||||
// list runs a chunk-selecting query and scans the (chunk_hash, size) rows.
|
||||
func (r *ChunkRepository) list(ctx context.Context, query string) ([]*Chunk, error) {
|
||||
rows, err := r.db.conn.QueryContext(ctx, query)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("querying chunks: %w", err)
|
||||
|
||||
@@ -17,6 +17,7 @@ import (
|
||||
"embed"
|
||||
"errors"
|
||||
"fmt"
|
||||
"net/url"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"sort"
|
||||
@@ -219,6 +220,135 @@ func openWithRecovery(ctx context.Context, path string) (*DB, error) {
|
||||
return db, nil
|
||||
}
|
||||
|
||||
// errUntrustedSnapshotSchema is returned when a downloaded snapshot
|
||||
// database carries schema objects the real schema never defines, or is
|
||||
// missing a table the restore and deep-verify queries read.
|
||||
var errUntrustedSnapshotSchema = errors.New(
|
||||
"downloaded snapshot database has an untrusted schema")
|
||||
|
||||
// snapshotReadOnlyDSN builds the driver DSN that opens a materialized
|
||||
// snapshot database file read-only. mode=ro opens the file read-only at
|
||||
// the OS level, query_only rejects any write the engine is asked to make,
|
||||
// and trusted_schema=OFF refuses to run application code named in the
|
||||
// schema. The file: URI form is required for the driver to honour the
|
||||
// mode parameter.
|
||||
func snapshotReadOnlyDSN(path string) string {
|
||||
u := url.URL{
|
||||
Scheme: "file",
|
||||
Path: path,
|
||||
RawQuery: "mode=ro&_pragma=query_only(true)&_pragma=trusted_schema(false)",
|
||||
}
|
||||
|
||||
return u.String()
|
||||
}
|
||||
|
||||
// OpenReadOnly opens an already-materialized SQLite file for read-only
|
||||
// querying of a snapshot database downloaded from the store, used by
|
||||
// restore and deep verify. Unlike New it never applies schema migrations
|
||||
// and never writes: the connection is opened read-only with query_only
|
||||
// and trusted_schema=OFF. It refuses any file whose schema carries a
|
||||
// trigger, view or virtual table, or lacks an expected table, so a forged
|
||||
// file cannot redefine what the restore queries return. The caller owns
|
||||
// the file and must remove it.
|
||||
func OpenReadOnly(ctx context.Context, path string) (*DB, error) {
|
||||
conn, err := sql.Open("sqlite", snapshotReadOnlyDSN(path))
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("opening read-only database: %w", err)
|
||||
}
|
||||
|
||||
configureConnPool(conn)
|
||||
|
||||
err = conn.PingContext(ctx)
|
||||
if err != nil {
|
||||
_ = conn.Close()
|
||||
|
||||
return nil, fmt.Errorf("opening read-only database: %w", err)
|
||||
}
|
||||
|
||||
err = verifySnapshotSchema(ctx, conn)
|
||||
if err != nil {
|
||||
_ = conn.Close()
|
||||
|
||||
return nil, err
|
||||
}
|
||||
|
||||
return &DB{conn: conn, path: path}, nil
|
||||
}
|
||||
|
||||
// verifySnapshotSchema rejects a downloaded database whose schema is not
|
||||
// the plain table set the real schema defines. Any trigger, view or
|
||||
// virtual table, or a missing expected table, fails the open.
|
||||
func verifySnapshotSchema(ctx context.Context, conn *sql.DB) error {
|
||||
// expectedSnapshotTables are the tables the restore and deep-verify
|
||||
// queries read. A downloaded database missing any of them is not a
|
||||
// genuine snapshot database and is refused.
|
||||
expectedSnapshotTables := []string{
|
||||
"blob_chunks",
|
||||
"blobs",
|
||||
"chunks",
|
||||
"file_chunks",
|
||||
"files",
|
||||
}
|
||||
|
||||
rows, err := conn.QueryContext(
|
||||
ctx, "SELECT type, name, sql FROM sqlite_master")
|
||||
if err != nil {
|
||||
return fmt.Errorf("reading snapshot schema: %w", err)
|
||||
}
|
||||
|
||||
defer func() { _ = rows.Close() }()
|
||||
|
||||
present := make(map[string]struct{})
|
||||
|
||||
for rows.Next() {
|
||||
var objType, name string
|
||||
|
||||
var objSQL sql.NullString
|
||||
|
||||
err = rows.Scan(&objType, &name, &objSQL)
|
||||
if err != nil {
|
||||
return fmt.Errorf("reading snapshot schema: %w", err)
|
||||
}
|
||||
|
||||
switch objType {
|
||||
case "trigger", "view":
|
||||
return fmt.Errorf(
|
||||
"%w: unexpected %s %q", errUntrustedSnapshotSchema, objType, name)
|
||||
case "table":
|
||||
if isVirtualTableSQL(objSQL.String) {
|
||||
return fmt.Errorf(
|
||||
"%w: unexpected virtual table %q",
|
||||
errUntrustedSnapshotSchema, name)
|
||||
}
|
||||
|
||||
present[name] = struct{}{}
|
||||
}
|
||||
}
|
||||
|
||||
err = rows.Err()
|
||||
if err != nil {
|
||||
return fmt.Errorf("reading snapshot schema: %w", err)
|
||||
}
|
||||
|
||||
for _, table := range expectedSnapshotTables {
|
||||
if _, ok := present[table]; !ok {
|
||||
return fmt.Errorf(
|
||||
"%w: missing table %q", errUntrustedSnapshotSchema, table)
|
||||
}
|
||||
}
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// isVirtualTableSQL reports whether a sqlite_master row's SQL defines a
|
||||
// virtual table. Virtual tables are recorded with type 'table' but a
|
||||
// "CREATE VIRTUAL TABLE" definition and can run module code, so they are
|
||||
// refused alongside triggers and views.
|
||||
func isVirtualTableSQL(createSQL string) bool {
|
||||
return strings.HasPrefix(
|
||||
strings.ToUpper(strings.TrimSpace(createSQL)), "CREATE VIRTUAL TABLE")
|
||||
}
|
||||
|
||||
// NewTestDB creates an in-memory SQLite database for testing purposes.
|
||||
// The database is automatically initialized with the schema and is ready
|
||||
// for use. Each call creates a new independent database instance.
|
||||
|
||||
@@ -0,0 +1,145 @@
|
||||
//nolint:testpackage // exercises unexported read-only open internals
|
||||
package database
|
||||
|
||||
import (
|
||||
"context"
|
||||
"database/sql"
|
||||
"errors"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
)
|
||||
|
||||
// genuineSnapshotDB writes a real snapshot database (the full schema
|
||||
// applied) to a fresh file and returns its path.
|
||||
func genuineSnapshotDB(t *testing.T) string {
|
||||
t.Helper()
|
||||
|
||||
path := filepath.Join(t.TempDir(), "snapshot.db")
|
||||
|
||||
db, err := New(context.Background(), path)
|
||||
if err != nil {
|
||||
t.Fatalf("creating snapshot database: %v", err)
|
||||
}
|
||||
|
||||
err = db.Close()
|
||||
if err != nil {
|
||||
t.Fatalf("closing snapshot database: %v", err)
|
||||
}
|
||||
|
||||
return path
|
||||
}
|
||||
|
||||
// forgedDB creates an empty database file and runs the given statements
|
||||
// against it read-write, so a test can plant schema objects the real
|
||||
// schema never defines.
|
||||
func forgedDB(t *testing.T, stmts ...string) string {
|
||||
t.Helper()
|
||||
|
||||
path := filepath.Join(t.TempDir(), "forged.db")
|
||||
|
||||
db, err := sql.Open("sqlite", path)
|
||||
if err != nil {
|
||||
t.Fatalf("opening forged database: %v", err)
|
||||
}
|
||||
|
||||
for _, stmt := range stmts {
|
||||
_, err = db.ExecContext(context.Background(), stmt)
|
||||
if err != nil {
|
||||
t.Fatalf("executing %q: %v", stmt, err)
|
||||
}
|
||||
}
|
||||
|
||||
err = db.Close()
|
||||
if err != nil {
|
||||
t.Fatalf("closing forged database: %v", err)
|
||||
}
|
||||
|
||||
return path
|
||||
}
|
||||
|
||||
func TestOpenReadOnlyAcceptsGenuineSnapshot(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
db, err := OpenReadOnly(context.Background(), genuineSnapshotDB(t))
|
||||
if err != nil {
|
||||
t.Fatalf("OpenReadOnly refused a genuine snapshot database: %v", err)
|
||||
}
|
||||
|
||||
t.Cleanup(func() { _ = db.Close() })
|
||||
}
|
||||
|
||||
func TestOpenReadOnlyRefusesWrites(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
db, err := OpenReadOnly(context.Background(), genuineSnapshotDB(t))
|
||||
if err != nil {
|
||||
t.Fatalf("OpenReadOnly: %v", err)
|
||||
}
|
||||
|
||||
t.Cleanup(func() { _ = db.Close() })
|
||||
|
||||
// A schema write depends on no table columns, so the only reason it
|
||||
// can fail is that the database is open read-only.
|
||||
_, err = db.Conn().ExecContext(context.Background(),
|
||||
"CREATE TABLE probe_readonly (x)")
|
||||
if err == nil {
|
||||
t.Fatal("expected a write to a read-only snapshot database to fail")
|
||||
}
|
||||
}
|
||||
|
||||
func TestOpenReadOnlyRejectsView(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
path := forgedDB(t, "CREATE VIEW files AS SELECT 1 AS path")
|
||||
|
||||
_, err := OpenReadOnly(context.Background(), path)
|
||||
if !errors.Is(err, errUntrustedSnapshotSchema) {
|
||||
t.Fatalf("expected a view named files to be refused, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestOpenReadOnlyRejectsTrigger(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
path := forgedDB(t,
|
||||
"CREATE TABLE files (path TEXT)",
|
||||
"CREATE TRIGGER t AFTER INSERT ON files BEGIN SELECT 1; END")
|
||||
|
||||
_, err := OpenReadOnly(context.Background(), path)
|
||||
if !errors.Is(err, errUntrustedSnapshotSchema) {
|
||||
t.Fatalf("expected a trigger to be refused, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestOpenReadOnlyRejectsMissingTable(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// Only one of the expected tables is present.
|
||||
path := forgedDB(t, "CREATE TABLE files (path TEXT)")
|
||||
|
||||
_, err := OpenReadOnly(context.Background(), path)
|
||||
if !errors.Is(err, errUntrustedSnapshotSchema) {
|
||||
t.Fatalf("expected a missing expected table to be refused, got %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestIsVirtualTableSQL(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
cases := []struct {
|
||||
sql string
|
||||
want bool
|
||||
}{
|
||||
{"CREATE VIRTUAL TABLE t USING fts5(x)", true},
|
||||
{" create virtual table t using fts5(x)", true},
|
||||
{"CREATE TABLE t (x)", false},
|
||||
{"CREATE VIEW t AS SELECT 1", false},
|
||||
{"", false},
|
||||
}
|
||||
|
||||
for _, c := range cases {
|
||||
if got := isVirtualTableSQL(c.sql); got != c.want {
|
||||
t.Errorf("isVirtualTableSQL(%q) = %v, want %v", c.sql, got, c.want)
|
||||
}
|
||||
}
|
||||
}
|
||||
+15
-1
@@ -14,8 +14,18 @@ var Module = fx.Module("log",
|
||||
)
|
||||
|
||||
// New creates a new logger configuration from provided options.
|
||||
//
|
||||
// JSON is intentionally not carried into Config: a command emitting a
|
||||
// JSON document on stdout must keep its stderr log level under
|
||||
// --verbose/--debug, so --json must not lower it (issue #112). JSON
|
||||
// silences the stdout UI in setupGlobals instead.
|
||||
func New(opts Options) Config {
|
||||
return Config(opts)
|
||||
return Config{
|
||||
Verbose: opts.Verbose,
|
||||
Debug: opts.Debug,
|
||||
Cron: opts.Cron,
|
||||
Quiet: opts.Quiet,
|
||||
}
|
||||
}
|
||||
|
||||
// Options are provided by the CLI.
|
||||
@@ -24,4 +34,8 @@ type Options struct {
|
||||
Debug bool
|
||||
Cron bool
|
||||
Quiet bool
|
||||
// JSON marks a command whose stdout carries a machine-readable
|
||||
// document. It silences the human UI on stdout (see setupGlobals),
|
||||
// but unlike Quiet it leaves the stderr log level alone.
|
||||
JSON bool
|
||||
}
|
||||
|
||||
@@ -1,67 +0,0 @@
|
||||
// Package models defines shared value types describing files, chunks,
|
||||
// blobs, and snapshots as they move through the backup pipeline.
|
||||
package models
|
||||
|
||||
import (
|
||||
"time"
|
||||
)
|
||||
|
||||
// FileInfo represents a file in the backup system
|
||||
type FileInfo struct {
|
||||
Path string
|
||||
MTime time.Time
|
||||
Size int64
|
||||
}
|
||||
|
||||
// ChunkInfo represents a content-addressed chunk
|
||||
type ChunkInfo struct {
|
||||
Hash string // SHA256 hash
|
||||
Size int64
|
||||
Offset int64 // Offset within source file
|
||||
}
|
||||
|
||||
// ChunkRef represents a reference to a chunk in a blob or file
|
||||
type ChunkRef struct {
|
||||
ChunkHash string
|
||||
Offset int64
|
||||
Length int64
|
||||
}
|
||||
|
||||
// BlobInfo represents an encrypted blob containing multiple chunks
|
||||
type BlobInfo struct {
|
||||
Hash string // SHA256 hash of the blob content (content-addressable)
|
||||
CreatedAt time.Time
|
||||
Size int64
|
||||
ChunkCount int
|
||||
}
|
||||
|
||||
// Snapshot represents a backup snapshot
|
||||
type Snapshot struct {
|
||||
ID string // ISO8601 timestamp
|
||||
Hostname string
|
||||
Version string
|
||||
CreatedAt time.Time
|
||||
FileCount int64
|
||||
ChunkCount int64
|
||||
BlobCount int64
|
||||
TotalSize int64
|
||||
MetadataSize int64
|
||||
}
|
||||
|
||||
// SnapshotMetadata contains the full metadata for a snapshot
|
||||
type SnapshotMetadata struct {
|
||||
Snapshot *Snapshot
|
||||
Files map[string]*FileInfo
|
||||
Chunks map[string]*ChunkInfo
|
||||
Blobs map[string]*BlobInfo
|
||||
FileChunks map[string][]*ChunkRef // path -> chunks
|
||||
BlobChunks map[string][]*ChunkRef // blob hash -> chunks
|
||||
}
|
||||
|
||||
// Chunk represents a data chunk for processing
|
||||
type Chunk struct {
|
||||
Data []byte
|
||||
Hash string
|
||||
Offset int64
|
||||
Length int64
|
||||
}
|
||||
@@ -1,58 +0,0 @@
|
||||
package models_test
|
||||
|
||||
import (
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/models"
|
||||
)
|
||||
|
||||
// TestModelsCompilation ensures all model types can be instantiated
|
||||
func TestModelsCompilation(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// This test primarily serves as a compilation test
|
||||
// to ensure all types are properly defined
|
||||
|
||||
// Test FileInfo
|
||||
fi := &models.FileInfo{
|
||||
Path: "/test/file.txt",
|
||||
MTime: time.Now(),
|
||||
Size: 1024,
|
||||
}
|
||||
if fi.Path != "/test/file.txt" {
|
||||
t.Errorf("FileInfo.Path not set correctly")
|
||||
}
|
||||
|
||||
// Test ChunkInfo
|
||||
ci := &models.ChunkInfo{
|
||||
Hash: "abc123",
|
||||
Size: 512,
|
||||
Offset: 0,
|
||||
}
|
||||
if ci.Hash != "abc123" {
|
||||
t.Errorf("ChunkInfo.Hash not set correctly")
|
||||
}
|
||||
|
||||
// Test BlobInfo
|
||||
bi := &models.BlobInfo{
|
||||
Hash: "blob123",
|
||||
CreatedAt: time.Now(),
|
||||
Size: 1024,
|
||||
ChunkCount: 2,
|
||||
}
|
||||
if bi.Hash != "blob123" {
|
||||
t.Errorf("BlobInfo.Hash not set correctly")
|
||||
}
|
||||
|
||||
// Test Snapshot
|
||||
s := &models.Snapshot{
|
||||
ID: "2024-01-01T00:00:00Z",
|
||||
Hostname: "test-host",
|
||||
Version: "1.0.0",
|
||||
CreatedAt: time.Now(),
|
||||
}
|
||||
if s.ID != "2024-01-01T00:00:00Z" {
|
||||
t.Errorf("Snapshot.ID not set correctly")
|
||||
}
|
||||
}
|
||||
+13
-5
@@ -219,11 +219,7 @@ func (c *Client) HeadObject(ctx context.Context, key string) (bool, error) {
|
||||
Key: aws.String(fullKey),
|
||||
})
|
||||
if err != nil {
|
||||
var (
|
||||
notFound *s3types.NotFound
|
||||
noSuchKey *s3types.NoSuchKey
|
||||
)
|
||||
if errors.As(err, ¬Found) || errors.As(err, &noSuchKey) {
|
||||
if IsNotFound(err) {
|
||||
return false, nil
|
||||
}
|
||||
|
||||
@@ -233,6 +229,18 @@ func (c *Client) HeadObject(ctx context.Context, key string) (bool, error) {
|
||||
return true, nil
|
||||
}
|
||||
|
||||
// IsNotFound reports whether err indicates that an object does not exist.
|
||||
// Head and Get requests surface a missing object as different SDK types,
|
||||
// so both are checked here.
|
||||
func IsNotFound(err error) bool {
|
||||
var (
|
||||
notFound *s3types.NotFound
|
||||
noSuchKey *s3types.NoSuchKey
|
||||
)
|
||||
|
||||
return errors.As(err, ¬Found) || errors.As(err, &noSuchKey)
|
||||
}
|
||||
|
||||
// ObjectInfo contains information about an S3 object.
|
||||
// It is used by ListObjectsStream to return object metadata
|
||||
// along with any errors encountered during listing.
|
||||
|
||||
@@ -0,0 +1,64 @@
|
||||
//nolint:testpackage // exercises the unexported generateBlobManifest
|
||||
package snapshot
|
||||
|
||||
import (
|
||||
"context"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/config"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/types"
|
||||
)
|
||||
|
||||
// TestGenerateBlobManifest_MissingBlobFails is the regression guard for
|
||||
// issue #157: a blob the snapshot references but that is absent from the
|
||||
// blobs table used to be logged and skipped, yielding a manifest with
|
||||
// fewer blobs than the snapshot needs. Since prune trusts the manifest
|
||||
// alone, that omitted blob would be deleted at the next prune. Manifest
|
||||
// generation must fail instead.
|
||||
func TestGenerateBlobManifest_MissingBlobFails(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
dbPath := filepath.Join(t.TempDir(), "snapshot.db")
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
// A real blob row satisfies the snapshot_blobs foreign key on
|
||||
// blob_id; the snapshot then references a different, absent hash.
|
||||
presentBlob := &database.Blob{
|
||||
ID: types.NewBlobID(),
|
||||
Hash: types.BlobHash("present-blob-hash"),
|
||||
CreatedTS: time.Now().Truncate(time.Second),
|
||||
}
|
||||
require.NoError(t, repos.Blobs.Create(ctx, nil, presentBlob))
|
||||
|
||||
snap := &database.Snapshot{
|
||||
ID: "testhost_home_2026-05-01T00:00:00Z",
|
||||
Hostname: "testhost",
|
||||
}
|
||||
require.NoError(t, repos.Snapshots.Create(ctx, nil, snap))
|
||||
require.NoError(t, repos.Snapshots.AddBlob(ctx, nil,
|
||||
snap.ID.String(), presentBlob.ID, types.BlobHash("absent-blob-hash")))
|
||||
|
||||
require.NoError(t, db.Close())
|
||||
|
||||
sm := &SnapshotManager{
|
||||
config: &config.Config{CompressionLevel: 3},
|
||||
fs: afero.NewOsFs(),
|
||||
}
|
||||
|
||||
_, err = sm.generateBlobManifest(ctx, dbPath, snap.ID.String())
|
||||
require.Error(t, err, "manifest generation must fail on a missing blob")
|
||||
assert.Contains(t, err.Error(), "absent-blob-hash")
|
||||
}
|
||||
@@ -22,8 +22,9 @@ const remoteKeyPrefix = "vaultik|"
|
||||
//
|
||||
// - the "metadata/<remote-key>/..." subdirectory on the storage
|
||||
// backend so a directory listing of the bucket / file:// dest
|
||||
// doesn't reveal hostnames, configured snapshot names, or backup
|
||||
// timestamps;
|
||||
// doesn't reveal hostnames or configured snapshot names. (The
|
||||
// backup time is not hidden: the manifest.json.zst inside that
|
||||
// directory carries a plaintext RFC3339 timestamp.)
|
||||
// - the `snapshot_id` field of the unencrypted manifest.json.zst
|
||||
// for the same reason;
|
||||
// - any code path that needs to translate a known local snapshot ID
|
||||
|
||||
@@ -63,7 +63,9 @@ type Scanner struct {
|
||||
exclude []string // Glob patterns for files/directories to exclude
|
||||
compiledExclude []compiledPattern // Compiled glob patterns
|
||||
progress *ProgressReporter
|
||||
skipErrors bool // Skip file read errors (log loudly but continue)
|
||||
// skipErrors skips files that cannot be opened or read (logged loudly);
|
||||
// packer, database, encryption, and upload errors still abort the run.
|
||||
skipErrors bool
|
||||
// ui is the user-facing output; never nil (defaults to a discarding writer).
|
||||
ui *ui.Writer
|
||||
|
||||
@@ -121,7 +123,9 @@ type ScannerConfig struct {
|
||||
EnableProgress bool // Enable the live progress reporter (ETAs, throughput)
|
||||
UI *ui.Writer // Where user-facing scanner messages go; nil = discard
|
||||
Exclude []string // Glob patterns for files/directories to exclude
|
||||
SkipErrors bool // Skip file read errors (log loudly but continue)
|
||||
// SkipErrors skips files that cannot be opened or read (log loudly but
|
||||
// continue); packer, database, encryption, and upload errors still abort.
|
||||
SkipErrors bool
|
||||
}
|
||||
|
||||
// ScanResult contains the results of a scan operation
|
||||
@@ -220,7 +224,14 @@ func (s *Scanner) Scan(
|
||||
defer s.progress.Stop()
|
||||
}
|
||||
|
||||
// Phase 0: Load known files and chunks from database into memory for fast lookup
|
||||
// Phase 0: Repair any state left by an interrupted previous run, then
|
||||
// load known files and chunks from the database into memory for fast
|
||||
// lookup.
|
||||
err := s.repairInterruptedBlobs(ctx)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
knownFiles, err := s.loadDatabaseState(ctx, path)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
@@ -317,6 +328,38 @@ func (s *Scanner) loadDatabaseState(
|
||||
return knownFiles, nil
|
||||
}
|
||||
|
||||
// repairInterruptedBlobs discards blob rows left by a previous run whose
|
||||
// upload never completed. Such a blob has its chunks, blob_chunks, and
|
||||
// blobs rows committed to the local index before the upload is attempted,
|
||||
// so a crash or dropped connection mid-upload leaves them behind while the
|
||||
// data never reaches remote storage. Deduplicating against those chunks on
|
||||
// a later run would produce a snapshot that reports success but cannot be
|
||||
// restored. Dropping the un-uploaded blobs (their blob_chunks cascade) and
|
||||
// then any chunks left unreferenced forces the affected data to be
|
||||
// re-chunked and re-uploaded this run. A blob is attached to a snapshot
|
||||
// only once its upload is recorded, so this never touches a completed
|
||||
// snapshot's data.
|
||||
func (s *Scanner) repairInterruptedBlobs(ctx context.Context) error {
|
||||
removed, err := s.repos.Blobs.DeleteUnuploaded(ctx)
|
||||
if err != nil {
|
||||
return fmt.Errorf("removing un-uploaded blob records: %w", err)
|
||||
}
|
||||
|
||||
if removed == 0 {
|
||||
return nil
|
||||
}
|
||||
|
||||
log.Warn("Discarded blob records from an interrupted previous run; "+
|
||||
"their data will be re-uploaded", "blobs", removed)
|
||||
|
||||
err = s.repos.Chunks.DeleteOrphaned(ctx)
|
||||
if err != nil {
|
||||
return fmt.Errorf("removing orphaned chunks: %w", err)
|
||||
}
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// summarizeScanPhase calculates total size to process, updates progress tracking,
|
||||
// and prints the scan phase summary with file counts and sizes
|
||||
func (s *Scanner) summarizeScanPhase(
|
||||
@@ -392,11 +435,14 @@ func (s *Scanner) loadKnownFiles(
|
||||
return result, nil
|
||||
}
|
||||
|
||||
// loadKnownChunks loads all known chunk hashes from the database into a
|
||||
// map for fast lookup. This avoids per-chunk database queries during file
|
||||
// processing.
|
||||
// loadKnownChunks loads the chunk hashes safe to deduplicate against into
|
||||
// an in-memory map for fast lookup, avoiding per-chunk database queries
|
||||
// during file processing. Only chunks held by a blob whose upload
|
||||
// completed are loaded: a chunk left behind by an interrupted upload
|
||||
// refers to data that never reached remote storage, and deduplicating
|
||||
// against it would silently produce an unrestorable snapshot.
|
||||
func (s *Scanner) loadKnownChunks(ctx context.Context) error {
|
||||
chunks, err := s.repos.Chunks.List(ctx)
|
||||
chunks, err := s.repos.Chunks.ListInUploadedBlobs(ctx)
|
||||
if err != nil {
|
||||
return fmt.Errorf("listing chunks: %w", err)
|
||||
}
|
||||
@@ -1294,6 +1340,15 @@ func (s *Scanner) processFileWithErrorHandling(
|
||||
) (bool, error) {
|
||||
err := s.processFileStreaming(ctx, fileToProcess, result)
|
||||
if err != nil {
|
||||
// A packer/database/encryption/upload failure means the chunk's data
|
||||
// may not have been stored. Skipping the file would let the snapshot
|
||||
// record a file whose chunk is in no blob and cannot be restored, so
|
||||
// abort the run even under --skip-errors. Only open and read errors
|
||||
// are skipped below.
|
||||
var pErr *packerError
|
||||
if errors.As(err, &pErr) {
|
||||
return false, fmt.Errorf("processing file %s: %w", fileToProcess.Path, err)
|
||||
}
|
||||
// Handle files that were deleted between scan and process phases
|
||||
if errors.Is(err, os.ErrNotExist) {
|
||||
log.Warn("File was deleted during backup, skipping",
|
||||
@@ -1303,7 +1358,7 @@ func (s *Scanner) processFileWithErrorHandling(
|
||||
|
||||
return true, nil
|
||||
}
|
||||
// Skip file read errors if --skip-errors is enabled
|
||||
// Skip open/read errors if --skip-errors is enabled
|
||||
if s.skipErrors {
|
||||
log.Error("Failed to process file (skipping due to --skip-errors)",
|
||||
"path", fileToProcess.Path, "error", err)
|
||||
@@ -1401,7 +1456,17 @@ func (s *Scanner) finalizeProcessPhase(ctx context.Context, result *ScanResult)
|
||||
return fmt.Errorf("parsing blob ID: %w", err)
|
||||
}
|
||||
|
||||
// With no remote backend the blob's lifecycle ends here, so
|
||||
// mark it uploaded in the same transaction that attaches it to
|
||||
// the snapshot. This keeps the invariant that any blob a
|
||||
// snapshot references has uploaded_ts set, so deduplication and
|
||||
// interrupted-run repair treat these blobs as trustworthy.
|
||||
err = s.repos.WithTx(ctx, func(ctx context.Context, tx *sql.Tx) error {
|
||||
err := s.repos.Blobs.UpdateUploaded(ctx, tx, b.ID)
|
||||
if err != nil {
|
||||
return fmt.Errorf("marking blob uploaded: %w", err)
|
||||
}
|
||||
|
||||
return s.repos.Snapshots.AddBlob(ctx, tx, s.snapshotID, blobID,
|
||||
types.BlobHash(b.Hash))
|
||||
})
|
||||
@@ -1660,6 +1725,20 @@ type streamingChunkInfo struct {
|
||||
size int64
|
||||
}
|
||||
|
||||
// packerError marks an error that came from adding a chunk to the packer
|
||||
// (packing, database, encryption, or upload). Such an error means the chunk's
|
||||
// data may not have been stored, so the run must abort even under --skip-errors:
|
||||
// skipping the file would leave the chunk recorded as backed up while it lives
|
||||
// in no blob, and a later snapshot could record a file that cannot be restored.
|
||||
// Only open and read errors are safe to skip.
|
||||
type packerError struct {
|
||||
err error
|
||||
}
|
||||
|
||||
func (e *packerError) Error() string { return e.err.Error() }
|
||||
|
||||
func (e *packerError) Unwrap() error { return e.err }
|
||||
|
||||
// processFileStreaming processes a file by streaming chunks directly to the packer
|
||||
func (s *Scanner) processFileStreaming(
|
||||
ctx context.Context, fileToProcess *FileToProcess, result *ScanResult,
|
||||
@@ -1710,7 +1789,11 @@ func (s *Scanner) processFileStreaming(
|
||||
if !chunkExists {
|
||||
err := s.addChunkToPacker(ctx, chunk)
|
||||
if err != nil {
|
||||
return err
|
||||
// Mark as a packer error so --skip-errors cannot swallow it:
|
||||
// the chunk was registered as pending before packing, so a
|
||||
// skipped file here would be recorded as backed up while its
|
||||
// data was never stored.
|
||||
return &packerError{err: err}
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,216 @@
|
||||
package snapshot_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
)
|
||||
|
||||
// errSimTempFail is the one-time temp-file creation failure blobTempFailFs
|
||||
// injects, mirroring a full temp filesystem.
|
||||
var errSimTempFail = errors.New("simulated temp-file creation failure")
|
||||
|
||||
// errSimRead is the read failure readFailFile injects for a file that opens
|
||||
// but cannot be read.
|
||||
var errSimRead = errors.New("simulated read failure")
|
||||
|
||||
// blobTempFailFs fails the first temp-file creation for a packer blob, then
|
||||
// behaves normally, simulating a one-time failure to start a new blob.
|
||||
type blobTempFailFs struct {
|
||||
afero.Fs
|
||||
|
||||
mu sync.Mutex
|
||||
failed bool
|
||||
}
|
||||
|
||||
//nolint:ireturn // afero.Fs.OpenFile is defined to return the interface.
|
||||
func (f *blobTempFailFs) OpenFile(
|
||||
name string, flag int, perm os.FileMode,
|
||||
) (afero.File, error) {
|
||||
if strings.Contains(name, "vaultik-blob-") {
|
||||
f.mu.Lock()
|
||||
firstTime := !f.failed
|
||||
f.failed = true
|
||||
f.mu.Unlock()
|
||||
|
||||
if firstTime {
|
||||
return nil, errSimTempFail
|
||||
}
|
||||
}
|
||||
|
||||
return f.Fs.OpenFile(name, flag, perm)
|
||||
}
|
||||
|
||||
// readFailFile wraps an afero.File whose Read always fails.
|
||||
type readFailFile struct {
|
||||
afero.File
|
||||
}
|
||||
|
||||
func (readFailFile) Read([]byte) (int, error) {
|
||||
return 0, errSimRead
|
||||
}
|
||||
|
||||
// readFailFs fails reads of one target path after a successful open.
|
||||
type readFailFs struct {
|
||||
afero.Fs
|
||||
|
||||
target string
|
||||
}
|
||||
|
||||
//nolint:ireturn // afero.Fs.Open is defined to return the interface.
|
||||
func (f *readFailFs) Open(name string) (afero.File, error) {
|
||||
file, err := f.Fs.Open(name)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
if name == f.target {
|
||||
return readFailFile{File: file}, nil
|
||||
}
|
||||
|
||||
return file, nil
|
||||
}
|
||||
|
||||
// writeSkipErrorTestFile writes one file into fs with a fixed mtime.
|
||||
func writeSkipErrorTestFile(t *testing.T, fs afero.Fs, path, content string) {
|
||||
t.Helper()
|
||||
|
||||
err := fs.MkdirAll(filepath.Dir(path), 0755)
|
||||
if err != nil {
|
||||
t.Fatalf("mkdir: %v", err)
|
||||
}
|
||||
|
||||
err = afero.WriteFile(fs, path, []byte(content), 0644)
|
||||
if err != nil {
|
||||
t.Fatalf("write %s: %v", path, err)
|
||||
}
|
||||
|
||||
when := time.Date(2024, 1, 1, 12, 0, 0, 0, time.UTC)
|
||||
|
||||
err = fs.Chtimes(path, when, when)
|
||||
if err != nil {
|
||||
t.Fatalf("chtimes %s: %v", path, err)
|
||||
}
|
||||
}
|
||||
|
||||
// runSkipErrorScan scans /source on fs with the given skip-errors setting and
|
||||
// returns the repositories (for inspection) and the scan error.
|
||||
func runSkipErrorScan(
|
||||
t *testing.T, fs afero.Fs, skipErrors bool,
|
||||
) (*database.Repositories, error) {
|
||||
t.Helper()
|
||||
|
||||
db, err := database.NewTestDB()
|
||||
if err != nil {
|
||||
t.Fatalf("create test db: %v", err)
|
||||
}
|
||||
|
||||
t.Cleanup(func() {
|
||||
cerr := db.Close()
|
||||
if cerr != nil {
|
||||
t.Errorf("close db: %v", cerr)
|
||||
}
|
||||
})
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
scanner := snapshot.NewScanner(snapshot.ScannerConfig{
|
||||
FS: fs,
|
||||
ChunkSize: int64(1024 * 16),
|
||||
Repositories: repos,
|
||||
MaxBlobSize: int64(1024 * 1024),
|
||||
CompressionLevel: 3,
|
||||
AgeRecipients: []string{testAgePublicKey},
|
||||
SkipErrors: skipErrors,
|
||||
})
|
||||
|
||||
ctx := context.Background()
|
||||
snapshotID := "test-snapshot-skip-errors"
|
||||
createTestSnapshotRecord(ctx, t, repos, snapshotID)
|
||||
|
||||
_, err = scanner.Scan(ctx, "/source", snapshotID)
|
||||
|
||||
return repos, err
|
||||
}
|
||||
|
||||
// TestScannerPackingFailureAbortsUnderSkipErrors checks that a failure to start
|
||||
// a new blob aborts the run even with --skip-errors. Otherwise the file would
|
||||
// be skipped while its chunk had already been registered as pending, letting a
|
||||
// later blob record that chunk in the chunks table with no blob to back it —
|
||||
// a snapshot that completes with a file that cannot be restored.
|
||||
func TestScannerPackingFailureAbortsUnderSkipErrors(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// Two files with distinct content so each yields a distinct chunk: the
|
||||
// first fails to start a blob, and without the fix the second's blob would
|
||||
// commit the first's orphaned chunk row.
|
||||
fs := &blobTempFailFs{Fs: afero.NewMemMapFs()}
|
||||
writeSkipErrorTestFile(t, fs, "/source/file1.txt", "first file content")
|
||||
writeSkipErrorTestFile(t, fs, "/source/file2.txt", "second file content")
|
||||
|
||||
repos, err := runSkipErrorScan(t, fs, true)
|
||||
if err == nil {
|
||||
t.Fatal("expected scan to abort on the packer error, got nil")
|
||||
}
|
||||
|
||||
// ListUnpacked returns chunks recorded with no blob_chunks row: exactly the
|
||||
// unrestorable state this fix prevents.
|
||||
unpacked, err := repos.Chunks.ListUnpacked(context.Background(), 10)
|
||||
if err != nil {
|
||||
t.Fatalf("listing unpacked chunks: %v", err)
|
||||
}
|
||||
|
||||
if len(unpacked) != 0 {
|
||||
t.Fatalf("expected no chunk recorded without a blob, got %d", len(unpacked))
|
||||
}
|
||||
}
|
||||
|
||||
// TestScannerReadErrorAbortsWithoutSkipErrors checks that a file read error
|
||||
// aborts the run when --skip-errors is not set.
|
||||
func TestScannerReadErrorAbortsWithoutSkipErrors(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const target = "/source/unreadable.txt"
|
||||
|
||||
fs := &readFailFs{Fs: afero.NewMemMapFs(), target: target}
|
||||
writeSkipErrorTestFile(t, fs, target, "content that cannot be read")
|
||||
|
||||
_, err := runSkipErrorScan(t, fs, false)
|
||||
if err == nil {
|
||||
t.Fatal("expected scan to fail on the read error, got nil")
|
||||
}
|
||||
}
|
||||
|
||||
// TestScannerReadErrorSkippedWithSkipErrors checks that a file read error is
|
||||
// skipped and the run completes when --skip-errors is set.
|
||||
func TestScannerReadErrorSkippedWithSkipErrors(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const target = "/source/unreadable.txt"
|
||||
|
||||
fs := &readFailFs{Fs: afero.NewMemMapFs(), target: target}
|
||||
writeSkipErrorTestFile(t, fs, target, "content that cannot be read")
|
||||
|
||||
repos, err := runSkipErrorScan(t, fs, true)
|
||||
if err != nil {
|
||||
t.Fatalf("expected scan to complete with --skip-errors, got %v", err)
|
||||
}
|
||||
|
||||
chunks, err := repos.FileChunks.GetByFile(context.Background(), target)
|
||||
if err != nil {
|
||||
t.Fatalf("getting file chunks: %v", err)
|
||||
}
|
||||
|
||||
if len(chunks) != 0 {
|
||||
t.Fatalf("expected unreadable file skipped, got %d chunks", len(chunks))
|
||||
}
|
||||
}
|
||||
@@ -44,7 +44,6 @@ import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
@@ -669,14 +668,31 @@ func (sm *SnapshotManager) collectCleanupStats(
|
||||
|
||||
// vacuumDatabase runs VACUUM on the database to remove deleted data and compact
|
||||
// This is critical for security - ensures no stale/deleted data pages are uploaded
|
||||
//
|
||||
// VACUUM runs through the modernc.org/sqlite driver, on a freshly opened
|
||||
// connection with no transaction in flight (VACUUM cannot run inside one).
|
||||
// The database opens in WAL mode, so VACUUM's rewrite lands in the WAL; the
|
||||
// checkpoint on Close flushes it into the main file, which is the file we
|
||||
// then compress and upload.
|
||||
func (sm *SnapshotManager) vacuumDatabase(ctx context.Context, dbPath string) error {
|
||||
log.Debug("Running VACUUM on database", "path", dbPath)
|
||||
//nolint:gosec // G204: fixed argv; dbPath is our own temp file path
|
||||
cmd := exec.CommandContext(ctx, "sqlite3", dbPath, "VACUUM;")
|
||||
|
||||
output, err := cmd.CombinedOutput()
|
||||
db, err := database.New(ctx, dbPath)
|
||||
if err != nil {
|
||||
return fmt.Errorf("running VACUUM: %w (output: %s)", err, string(output))
|
||||
return fmt.Errorf("opening database for VACUUM: %w", err)
|
||||
}
|
||||
|
||||
defer func() {
|
||||
cerr := db.Close()
|
||||
if cerr != nil {
|
||||
log.Debug("Failed to close database after VACUUM",
|
||||
"path", dbPath, "error", cerr)
|
||||
}
|
||||
}()
|
||||
|
||||
_, err = db.ExecWithLog(ctx, "VACUUM")
|
||||
if err != nil {
|
||||
return fmt.Errorf("running VACUUM: %w", err)
|
||||
}
|
||||
|
||||
return nil
|
||||
@@ -793,6 +809,11 @@ func (sm *SnapshotManager) copyFile(src, dst string) error {
|
||||
return nil
|
||||
}
|
||||
|
||||
// errBlobMissingFromDatabase means a snapshot references a blob that is
|
||||
// absent from the blobs table, so a complete manifest cannot be built.
|
||||
var errBlobMissingFromDatabase = errors.New(
|
||||
"blob referenced by snapshot is not in the database")
|
||||
|
||||
// generateBlobManifest creates a compressed JSON list of all blobs in the snapshot
|
||||
func (sm *SnapshotManager) generateBlobManifest(
|
||||
ctx context.Context, dbPath string, snapshotID string,
|
||||
@@ -823,25 +844,33 @@ func (sm *SnapshotManager) generateBlobManifest(
|
||||
totalCompressedSize := int64(0)
|
||||
|
||||
for _, hash := range blobHashes {
|
||||
// Every blob the snapshot references must appear in the manifest.
|
||||
// Prune consults only the manifest to decide what is still in use,
|
||||
// so silently dropping a blob here would let a later prune delete
|
||||
// it while this snapshot still needs it. A lookup failure or a
|
||||
// missing blob row therefore fails manifest generation.
|
||||
blob, err := repos.Blobs.GetByHash(ctx, hash)
|
||||
if err != nil {
|
||||
log.Warn("Failed to get blob details", "hash", hash, "error", err)
|
||||
return nil, fmt.Errorf("getting blob details for %s: %w", hash, err)
|
||||
}
|
||||
|
||||
continue
|
||||
if blob == nil {
|
||||
return nil, fmt.Errorf("%w: blob %s, snapshot %s",
|
||||
errBlobMissingFromDatabase, hash, snapshotID)
|
||||
}
|
||||
|
||||
if blob != nil {
|
||||
blobs = append(blobs, BlobInfo{
|
||||
Hash: hash,
|
||||
CompressedSize: blob.CompressedSize,
|
||||
})
|
||||
totalCompressedSize += blob.CompressedSize
|
||||
}
|
||||
}
|
||||
|
||||
// Create manifest. SnapshotID in the unencrypted manifest is the
|
||||
// double-SHA256 remote key, not the human ID, so the public bytes
|
||||
// don't reveal hostname/snapshot-name/timestamp metadata.
|
||||
// double-SHA256 remote key (see RemoteSnapshotKey), not the human ID,
|
||||
// so neither this field nor the directory name reveals the hostname or
|
||||
// snapshot name. Timestamp below is written in the clear, so the backup
|
||||
// time is observable to anyone who can read the manifest.
|
||||
manifest := &Manifest{
|
||||
SnapshotID: RemoteSnapshotKey(snapshotID),
|
||||
Timestamp: time.Now().UTC().Format(time.RFC3339),
|
||||
|
||||
@@ -2,6 +2,7 @@
|
||||
package snapshot
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"database/sql"
|
||||
"io"
|
||||
@@ -96,6 +97,97 @@ func verifyCleanedDB(
|
||||
}
|
||||
}
|
||||
|
||||
// TestVacuumDatabaseRemovesDeletedData proves the export path uploads a
|
||||
// compacted database: after rows carrying a recognizable marker are deleted
|
||||
// and vacuumDatabase runs, no page holding that marker survives in the file
|
||||
// on disk (the file compressFile later reads for upload).
|
||||
func TestVacuumDatabaseRemovesDeletedData(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
fs := afero.NewOsFs()
|
||||
|
||||
tempDir := t.TempDir()
|
||||
dbPath := filepath.Join(tempDir, "snapshot.db")
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
if err != nil {
|
||||
t.Fatalf("failed to create database: %v", err)
|
||||
}
|
||||
|
||||
// A marker distinctive enough that its presence in the raw file can only
|
||||
// come from the rows inserted below.
|
||||
marker := []byte("VACUUM_PROBE_DEADBEEF_DELETED_ROW")
|
||||
payload := bytes.Repeat(marker, 128) // ~4 KiB per row
|
||||
|
||||
_, err = db.Conn().ExecContext(ctx,
|
||||
"CREATE TABLE vacuum_probe (id INTEGER PRIMARY KEY, payload BLOB)")
|
||||
if err != nil {
|
||||
t.Fatalf("failed to create probe table: %v", err)
|
||||
}
|
||||
|
||||
for range 512 {
|
||||
_, err = db.Conn().ExecContext(ctx,
|
||||
"INSERT INTO vacuum_probe (payload) VALUES (?)", payload)
|
||||
if err != nil {
|
||||
t.Fatalf("failed to insert probe row: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
_, err = db.Conn().ExecContext(ctx, "DELETE FROM vacuum_probe")
|
||||
if err != nil {
|
||||
t.Fatalf("failed to delete probe rows: %v", err)
|
||||
}
|
||||
|
||||
// Close so the deletes reach the main file, mirroring the state
|
||||
// prepareExportDB hands to vacuumDatabase.
|
||||
err = db.Close()
|
||||
if err != nil {
|
||||
t.Fatalf("failed to close database: %v", err)
|
||||
}
|
||||
|
||||
beforeInfo, err := fs.Stat(dbPath)
|
||||
if err != nil {
|
||||
t.Fatalf("failed to stat database before vacuum: %v", err)
|
||||
}
|
||||
|
||||
beforeBytes, err := afero.ReadFile(fs, dbPath)
|
||||
if err != nil {
|
||||
t.Fatalf("failed to read database before vacuum: %v", err)
|
||||
}
|
||||
|
||||
if !bytes.Contains(beforeBytes, marker) {
|
||||
t.Fatalf("expected deleted-row data to linger before vacuum")
|
||||
}
|
||||
|
||||
sm := &SnapshotManager{fs: fs}
|
||||
|
||||
err = sm.vacuumDatabase(ctx, dbPath)
|
||||
if err != nil {
|
||||
t.Fatalf("vacuumDatabase failed: %v", err)
|
||||
}
|
||||
|
||||
afterBytes, err := afero.ReadFile(fs, dbPath)
|
||||
if err != nil {
|
||||
t.Fatalf("failed to read database after vacuum: %v", err)
|
||||
}
|
||||
|
||||
if bytes.Contains(afterBytes, marker) {
|
||||
t.Fatalf("deleted-row data survived vacuum in the uploaded file")
|
||||
}
|
||||
|
||||
afterInfo, err := fs.Stat(dbPath)
|
||||
if err != nil {
|
||||
t.Fatalf("failed to stat database after vacuum: %v", err)
|
||||
}
|
||||
|
||||
if afterInfo.Size() >= beforeInfo.Size() {
|
||||
t.Fatalf("expected vacuum to shrink the file: before=%d after=%d",
|
||||
beforeInfo.Size(), afterInfo.Size())
|
||||
}
|
||||
}
|
||||
|
||||
func TestCleanSnapshotDBEmptySnapshot(t *testing.T) {
|
||||
// Initialize logger
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
@@ -0,0 +1,198 @@
|
||||
package storage_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"errors"
|
||||
"io"
|
||||
"reflect"
|
||||
"sort"
|
||||
"testing"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
)
|
||||
|
||||
// runStorerConformance is the shared Storer contract. Every backend that
|
||||
// can run in-process is expected to pass it: TestFileStorer runs it against
|
||||
// file://, TestS3Storer against s3://. A new backend inherits this coverage
|
||||
// by passing its own constructor, so the contract is defined once.
|
||||
//
|
||||
// It exercises the public Storer interface: round-trip, stat, list with
|
||||
// prefix filtering, overwrite, delete, delete-of-missing, and not-found on
|
||||
// Get and Stat. Each section takes its own fresh backend instance, so the
|
||||
// order of sections never matters and no section sees another's objects.
|
||||
func runStorerConformance(t *testing.T, newStorer func(*testing.T) storage.Storer) {
|
||||
t.Helper()
|
||||
|
||||
conformanceRoundTrip(t, newStorer(t))
|
||||
conformanceOverwrite(t, newStorer(t))
|
||||
conformanceList(t, newStorer(t))
|
||||
conformanceDelete(t, newStorer(t))
|
||||
conformanceNotFound(t, newStorer(t))
|
||||
}
|
||||
|
||||
// conformanceRoundTrip stores a nested key, then reads it back and stats it.
|
||||
func conformanceRoundTrip(t *testing.T, s storage.Storer) {
|
||||
t.Helper()
|
||||
|
||||
ctx := context.Background()
|
||||
key := "blobs/aa/bb/object.bin"
|
||||
want := []byte("round-trip payload")
|
||||
|
||||
err := s.Put(ctx, key, bytes.NewReader(want))
|
||||
if err != nil {
|
||||
t.Fatalf("Put: %v", err)
|
||||
}
|
||||
|
||||
got := getBytes(t, s, key)
|
||||
if !bytes.Equal(got, want) {
|
||||
t.Errorf("Get returned %q, want %q", got, want)
|
||||
}
|
||||
|
||||
info, err := s.Stat(ctx, key)
|
||||
if err != nil {
|
||||
t.Fatalf("Stat: %v", err)
|
||||
}
|
||||
|
||||
if info.Key != key {
|
||||
t.Errorf("Stat key = %q, want %q", info.Key, key)
|
||||
}
|
||||
|
||||
if info.Size != int64(len(want)) {
|
||||
t.Errorf("Stat size = %d, want %d", info.Size, len(want))
|
||||
}
|
||||
}
|
||||
|
||||
// conformanceOverwrite checks that a second Put replaces the first.
|
||||
func conformanceOverwrite(t *testing.T, s storage.Storer) {
|
||||
t.Helper()
|
||||
|
||||
ctx := context.Background()
|
||||
key := "meta/snapshot.json"
|
||||
|
||||
err := s.Put(ctx, key, bytes.NewReader([]byte("first")))
|
||||
if err != nil {
|
||||
t.Fatalf("first Put: %v", err)
|
||||
}
|
||||
|
||||
want := []byte("second and longer payload")
|
||||
|
||||
err = s.Put(ctx, key, bytes.NewReader(want))
|
||||
if err != nil {
|
||||
t.Fatalf("second Put: %v", err)
|
||||
}
|
||||
|
||||
got := getBytes(t, s, key)
|
||||
if !bytes.Equal(got, want) {
|
||||
t.Errorf("after overwrite Get returned %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
// conformanceList checks prefix filtering and the empty result for a
|
||||
// prefix that matches nothing.
|
||||
func conformanceList(t *testing.T, s storage.Storer) {
|
||||
t.Helper()
|
||||
|
||||
ctx := context.Background()
|
||||
keys := []string{"blobs/aa/one", "blobs/bb/two", "meta/three"}
|
||||
|
||||
for _, k := range keys {
|
||||
err := s.Put(ctx, k, bytes.NewReader([]byte("data")))
|
||||
if err != nil {
|
||||
t.Fatalf("Put %q: %v", k, err)
|
||||
}
|
||||
}
|
||||
|
||||
if got := listSorted(t, s, ""); !reflect.DeepEqual(got, keys) {
|
||||
t.Errorf("List(\"\") = %v, want %v", got, keys)
|
||||
}
|
||||
|
||||
wantBlobs := []string{"blobs/aa/one", "blobs/bb/two"}
|
||||
if got := listSorted(t, s, "blobs/"); !reflect.DeepEqual(got, wantBlobs) {
|
||||
t.Errorf("List(\"blobs/\") = %v, want %v", got, wantBlobs)
|
||||
}
|
||||
|
||||
if got := listSorted(t, s, "absent/"); len(got) != 0 {
|
||||
t.Errorf("List(\"absent/\") = %v, want empty", got)
|
||||
}
|
||||
}
|
||||
|
||||
// conformanceDelete checks that Delete removes an object and that deleting
|
||||
// a missing key is not an error.
|
||||
func conformanceDelete(t *testing.T, s storage.Storer) {
|
||||
t.Helper()
|
||||
|
||||
ctx := context.Background()
|
||||
key := "blobs/cc/gone.bin"
|
||||
|
||||
err := s.Put(ctx, key, bytes.NewReader([]byte("temporary")))
|
||||
if err != nil {
|
||||
t.Fatalf("Put: %v", err)
|
||||
}
|
||||
|
||||
err = s.Delete(ctx, key)
|
||||
if err != nil {
|
||||
t.Fatalf("Delete: %v", err)
|
||||
}
|
||||
|
||||
_, err = s.Get(ctx, key)
|
||||
if !errors.Is(err, storage.ErrNotFound) {
|
||||
t.Errorf("Get after Delete error = %v, want ErrNotFound", err)
|
||||
}
|
||||
|
||||
err = s.Delete(ctx, key)
|
||||
if err != nil {
|
||||
t.Errorf("Delete of missing key = %v, want nil", err)
|
||||
}
|
||||
}
|
||||
|
||||
// conformanceNotFound checks Get and Stat on an absent key.
|
||||
func conformanceNotFound(t *testing.T, s storage.Storer) {
|
||||
t.Helper()
|
||||
|
||||
ctx := context.Background()
|
||||
key := "never/written"
|
||||
|
||||
_, err := s.Get(ctx, key)
|
||||
if !errors.Is(err, storage.ErrNotFound) {
|
||||
t.Errorf("Get error = %v, want ErrNotFound", err)
|
||||
}
|
||||
|
||||
_, err = s.Stat(ctx, key)
|
||||
if !errors.Is(err, storage.ErrNotFound) {
|
||||
t.Errorf("Stat error = %v, want ErrNotFound", err)
|
||||
}
|
||||
}
|
||||
|
||||
// getBytes reads a key fully and closes the reader.
|
||||
func getBytes(t *testing.T, s storage.Storer, key string) []byte {
|
||||
t.Helper()
|
||||
|
||||
rc, err := s.Get(context.Background(), key)
|
||||
if err != nil {
|
||||
t.Fatalf("Get %q: %v", key, err)
|
||||
}
|
||||
|
||||
defer func() { _ = rc.Close() }()
|
||||
|
||||
data, err := io.ReadAll(rc)
|
||||
if err != nil {
|
||||
t.Fatalf("read %q: %v", key, err)
|
||||
}
|
||||
|
||||
return data
|
||||
}
|
||||
|
||||
// listSorted returns the keys under a prefix in a stable order.
|
||||
func listSorted(t *testing.T, s storage.Storer, prefix string) []string {
|
||||
t.Helper()
|
||||
|
||||
keys, err := s.List(context.Background(), prefix)
|
||||
if err != nil {
|
||||
t.Fatalf("List %q: %v", prefix, err)
|
||||
}
|
||||
|
||||
sort.Strings(keys)
|
||||
|
||||
return keys
|
||||
}
|
||||
@@ -0,0 +1,209 @@
|
||||
// Package faultstore provides a storage.Storer wrapper that injects
|
||||
// faults on demand, so tests can reproduce the failure modes a real
|
||||
// backend exhibits: an upload that fails partway, a backend that reports
|
||||
// success while storing nothing, and reads that return corrupt or
|
||||
// truncated bytes. It is the seam called for by the fault-injection
|
||||
// tests (sneak/vaultik issue 72) and is meant to be reused by future
|
||||
// tests rather than re-implemented per case.
|
||||
//
|
||||
// The wrapper delegates every method to the inner Storer. Two hooks
|
||||
// change that: OnPut decides the fate of each write, and OnGet decides
|
||||
// how each read's bytes are returned. Both are keyed by the object key,
|
||||
// so a test can fault only blobs, only metadata, or a single object.
|
||||
package faultstore
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"errors"
|
||||
"fmt"
|
||||
"io"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
)
|
||||
|
||||
// ErrInjectedUpload is returned by a Put the OnPut hook chose to fail.
|
||||
var ErrInjectedUpload = errors.New("faultstore: injected upload failure")
|
||||
|
||||
// PutAction is the disposition OnPut assigns to a write.
|
||||
type PutAction int
|
||||
|
||||
const (
|
||||
// PutNormal writes through to the inner Storer.
|
||||
PutNormal PutAction = iota
|
||||
// PutFail reads part of the stream, then fails without storing the
|
||||
// object — a network upload that dies partway through.
|
||||
PutFail
|
||||
// PutSwallow reports success but stores nothing — a backend that
|
||||
// lies about durability.
|
||||
PutSwallow
|
||||
)
|
||||
|
||||
// GetFault is how OnGet chooses to damage a read.
|
||||
type GetFault int
|
||||
|
||||
const (
|
||||
// GetNormal returns the stored bytes unchanged.
|
||||
GetNormal GetFault = iota
|
||||
// GetCorrupt flips a byte so the returned object no longer matches
|
||||
// what was stored.
|
||||
GetCorrupt
|
||||
// GetTruncate returns a short read: the object's bytes cut off
|
||||
// before the end.
|
||||
GetTruncate
|
||||
)
|
||||
|
||||
// Storer wraps an inner storage.Storer with fault-injection hooks. A
|
||||
// zero-valued hook means "no fault": construct with New and set only the
|
||||
// hook a test needs.
|
||||
type Storer struct {
|
||||
inner storage.Storer
|
||||
|
||||
// OnPut, when set, is consulted before every Put and
|
||||
// PutWithProgress with the object key.
|
||||
OnPut func(key string) PutAction
|
||||
|
||||
// OnGet, when set, is consulted for every Get with the object key
|
||||
// and damages the returned bytes accordingly.
|
||||
OnGet func(key string) GetFault
|
||||
}
|
||||
|
||||
// New wraps inner. inner must be non-nil.
|
||||
func New(inner storage.Storer) *Storer {
|
||||
return &Storer{inner: inner}
|
||||
}
|
||||
|
||||
// midStreamBytes is how far a PutFail reads before failing, enough to be
|
||||
// past the start of any real blob without depending on the blob's size.
|
||||
const midStreamBytes = 512
|
||||
|
||||
// Put stores data unless OnPut faults the write.
|
||||
func (f *Storer) Put(ctx context.Context, key string, data io.Reader) error {
|
||||
handled, err := f.injectPut(key, data)
|
||||
if handled {
|
||||
return err
|
||||
}
|
||||
|
||||
return f.inner.Put(ctx, key, data)
|
||||
}
|
||||
|
||||
// PutWithProgress stores data unless OnPut faults the write.
|
||||
func (f *Storer) PutWithProgress(
|
||||
ctx context.Context, key string, data io.Reader,
|
||||
size int64, progress storage.ProgressCallback,
|
||||
) error {
|
||||
handled, err := f.injectPut(key, data)
|
||||
if handled {
|
||||
return err
|
||||
}
|
||||
|
||||
return f.inner.PutWithProgress(ctx, key, data, size, progress)
|
||||
}
|
||||
|
||||
// Get retrieves data, damaging it if OnGet faults the read.
|
||||
func (f *Storer) Get(ctx context.Context, key string) (io.ReadCloser, error) {
|
||||
rc, err := f.inner.Get(ctx, key)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
fault := GetNormal
|
||||
if f.OnGet != nil {
|
||||
fault = f.OnGet(key)
|
||||
}
|
||||
|
||||
if fault == GetNormal {
|
||||
return rc, nil
|
||||
}
|
||||
|
||||
data, err := io.ReadAll(rc)
|
||||
_ = rc.Close()
|
||||
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
return io.NopCloser(bytes.NewReader(damage(fault, data))), nil
|
||||
}
|
||||
|
||||
// damage returns a faulted copy of the stored bytes. GetCorrupt flips a
|
||||
// byte in the middle so decryption authentication fails; GetTruncate
|
||||
// drops the final byte so the read ends short. Both are no-ops on empty
|
||||
// input, which cannot be damaged into something distinguishable.
|
||||
func damage(fault GetFault, data []byte) []byte {
|
||||
out := make([]byte, len(data))
|
||||
copy(out, data)
|
||||
|
||||
if len(out) == 0 {
|
||||
return out
|
||||
}
|
||||
|
||||
switch fault {
|
||||
case GetCorrupt:
|
||||
out[len(out)/2] ^= 0xff
|
||||
case GetTruncate:
|
||||
out = out[:len(out)-1]
|
||||
case GetNormal:
|
||||
}
|
||||
|
||||
return out
|
||||
}
|
||||
|
||||
// Stat delegates unchanged.
|
||||
func (f *Storer) Stat(ctx context.Context, key string) (*storage.ObjectInfo, error) {
|
||||
return f.inner.Stat(ctx, key)
|
||||
}
|
||||
|
||||
// Delete delegates unchanged.
|
||||
func (f *Storer) Delete(ctx context.Context, key string) error {
|
||||
return f.inner.Delete(ctx, key)
|
||||
}
|
||||
|
||||
// List delegates unchanged.
|
||||
func (f *Storer) List(ctx context.Context, prefix string) ([]string, error) {
|
||||
return f.inner.List(ctx, prefix)
|
||||
}
|
||||
|
||||
// ListStream delegates unchanged.
|
||||
func (f *Storer) ListStream(
|
||||
ctx context.Context, prefix string,
|
||||
) <-chan storage.ObjectInfo {
|
||||
return f.inner.ListStream(ctx, prefix)
|
||||
}
|
||||
|
||||
// Info delegates unchanged.
|
||||
func (f *Storer) Info() storage.Info {
|
||||
return f.inner.Info()
|
||||
}
|
||||
|
||||
func (f *Storer) putAction(key string) PutAction {
|
||||
if f.OnPut == nil {
|
||||
return PutNormal
|
||||
}
|
||||
|
||||
return f.OnPut(key)
|
||||
}
|
||||
|
||||
// injectPut handles the non-normal write dispositions. It reports
|
||||
// whether it handled the write and, if so, with what error.
|
||||
func (f *Storer) injectPut(key string, data io.Reader) (bool, error) {
|
||||
switch f.putAction(key) {
|
||||
case PutFail:
|
||||
// Consume part of the stream so the failure lands mid-transfer,
|
||||
// the way a dropped connection would, then error without
|
||||
// storing anything.
|
||||
_, _ = io.CopyN(io.Discard, data, midStreamBytes)
|
||||
|
||||
return true, fmt.Errorf("%w for %q", ErrInjectedUpload, key)
|
||||
case PutSwallow:
|
||||
// A lying backend still drains the request body, then keeps
|
||||
// nothing.
|
||||
_, _ = io.Copy(io.Discard, data)
|
||||
|
||||
return true, nil
|
||||
case PutNormal:
|
||||
return false, nil
|
||||
default:
|
||||
return false, nil
|
||||
}
|
||||
}
|
||||
+79
-54
@@ -46,31 +46,18 @@ func (f *FileStorer) SetFilesystem(fs afero.Fs) {
|
||||
// storage base path.
|
||||
const storageDirPerm = 0o755
|
||||
|
||||
// tempSuffix marks a partially written object. writeAtomic streams into a
|
||||
// temp file carrying this suffix and only renames it onto the real key once
|
||||
// the whole object is on disk, so an interrupted write can never leave a
|
||||
// truncated object at the key a later run would Stat and trust as a complete
|
||||
// blob. List and ListStream skip these files, so a leftover from an
|
||||
// interrupted write is never listed or trusted as a blob; it is otherwise
|
||||
// harmless and is overwritten when the same key is written again.
|
||||
const tempSuffix = ".partial"
|
||||
|
||||
// Put stores data at the specified key.
|
||||
func (f *FileStorer) Put(_ context.Context, key string, data io.Reader) error {
|
||||
path := f.fullPath(key)
|
||||
|
||||
// Create parent directories
|
||||
dir := filepath.Dir(path)
|
||||
|
||||
err := f.fs.MkdirAll(dir, storageDirPerm)
|
||||
if err != nil {
|
||||
return fmt.Errorf("creating directories: %w", err)
|
||||
}
|
||||
|
||||
file, err := f.fs.Create(path)
|
||||
if err != nil {
|
||||
return fmt.Errorf("creating file: %w", err)
|
||||
}
|
||||
|
||||
defer func() { _ = file.Close() }()
|
||||
|
||||
_, err = io.Copy(file, data)
|
||||
if err != nil {
|
||||
return fmt.Errorf("writing file: %w", err)
|
||||
}
|
||||
|
||||
return nil
|
||||
return f.writeAtomic(key, data, nil)
|
||||
}
|
||||
|
||||
// PutWithProgress stores data with progress reporting.
|
||||
@@ -78,35 +65,7 @@ func (f *FileStorer) PutWithProgress(
|
||||
_ context.Context, key string, data io.Reader,
|
||||
_ int64, progress ProgressCallback,
|
||||
) error {
|
||||
path := f.fullPath(key)
|
||||
|
||||
// Create parent directories
|
||||
dir := filepath.Dir(path)
|
||||
|
||||
err := f.fs.MkdirAll(dir, storageDirPerm)
|
||||
if err != nil {
|
||||
return fmt.Errorf("creating directories: %w", err)
|
||||
}
|
||||
|
||||
file, err := f.fs.Create(path)
|
||||
if err != nil {
|
||||
return fmt.Errorf("creating file: %w", err)
|
||||
}
|
||||
|
||||
defer func() { _ = file.Close() }()
|
||||
|
||||
// Wrap with progress tracking
|
||||
pw := &progressWriter{
|
||||
writer: file,
|
||||
callback: progress,
|
||||
}
|
||||
|
||||
_, err = io.Copy(pw, data)
|
||||
if err != nil {
|
||||
return fmt.Errorf("writing file: %w", err)
|
||||
}
|
||||
|
||||
return nil
|
||||
return f.writeAtomic(key, data, progress)
|
||||
}
|
||||
|
||||
// Get retrieves data from the specified key.
|
||||
@@ -188,7 +147,7 @@ func (f *FileStorer) List(ctx context.Context, prefix string) ([]string, error)
|
||||
default:
|
||||
}
|
||||
|
||||
if !info.IsDir() {
|
||||
if !info.IsDir() && !strings.HasSuffix(info.Name(), tempSuffix) {
|
||||
// Convert back to key (relative path from basePath)
|
||||
relPath, err := filepath.Rel(f.basePath, path)
|
||||
if err != nil {
|
||||
@@ -245,7 +204,7 @@ func (f *FileStorer) ListStream(ctx context.Context, prefix string) <-chan Objec
|
||||
return nil //nolint:nilerr // continue walking despite errors
|
||||
}
|
||||
|
||||
if !info.IsDir() {
|
||||
if !info.IsDir() && !strings.HasSuffix(info.Name(), tempSuffix) {
|
||||
relPath, err := filepath.Rel(f.basePath, path)
|
||||
if err != nil {
|
||||
ch <- ObjectInfo{Err: fmt.Errorf("computing relative path: %w", err)}
|
||||
@@ -275,6 +234,72 @@ func (f *FileStorer) Info() Info {
|
||||
}
|
||||
}
|
||||
|
||||
// writeAtomic streams data into a temp file in the destination directory,
|
||||
// fsyncs it, and renames it onto the final key. The key therefore appears
|
||||
// only once the whole object has been durably written; a failure part-way
|
||||
// leaves a temp file (removed here on the failing path) rather than a
|
||||
// truncated object at the key.
|
||||
func (f *FileStorer) writeAtomic(
|
||||
key string, data io.Reader, progress ProgressCallback,
|
||||
) error {
|
||||
path := f.fullPath(key)
|
||||
dir := filepath.Dir(path)
|
||||
|
||||
err := f.fs.MkdirAll(dir, storageDirPerm)
|
||||
if err != nil {
|
||||
return fmt.Errorf("creating directories: %w", err)
|
||||
}
|
||||
|
||||
tmp, err := afero.TempFile(f.fs, dir, filepath.Base(path)+"-*"+tempSuffix)
|
||||
if err != nil {
|
||||
return fmt.Errorf("creating temp file: %w", err)
|
||||
}
|
||||
|
||||
tmpPath := tmp.Name()
|
||||
|
||||
// Remove the temp file unless the rename below claims it. On the success
|
||||
// path renamed is true, so the deferred Close and Remove are harmless
|
||||
// no-ops on a name that no longer exists.
|
||||
renamed := false
|
||||
|
||||
defer func() {
|
||||
_ = tmp.Close()
|
||||
|
||||
if !renamed {
|
||||
_ = f.fs.Remove(tmpPath)
|
||||
}
|
||||
}()
|
||||
|
||||
var w io.Writer = tmp
|
||||
if progress != nil {
|
||||
w = &progressWriter{writer: tmp, callback: progress}
|
||||
}
|
||||
|
||||
_, err = io.Copy(w, data)
|
||||
if err != nil {
|
||||
return fmt.Errorf("writing file: %w", err)
|
||||
}
|
||||
|
||||
err = tmp.Sync()
|
||||
if err != nil {
|
||||
return fmt.Errorf("syncing temp file: %w", err)
|
||||
}
|
||||
|
||||
err = tmp.Close()
|
||||
if err != nil {
|
||||
return fmt.Errorf("closing temp file: %w", err)
|
||||
}
|
||||
|
||||
err = f.fs.Rename(tmpPath, path)
|
||||
if err != nil {
|
||||
return fmt.Errorf("renaming temp file: %w", err)
|
||||
}
|
||||
|
||||
renamed = true
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// fullPath returns the full filesystem path for a key.
|
||||
func (f *FileStorer) fullPath(key string) string {
|
||||
return filepath.Join(f.basePath, key)
|
||||
|
||||
@@ -0,0 +1,119 @@
|
||||
package storage_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
)
|
||||
|
||||
// errStreamInterrupted stands in for an upload cut off mid-stream.
|
||||
var errStreamInterrupted = errors.New("connection reset mid-upload")
|
||||
|
||||
// failingReader yields its data once, then fails.
|
||||
type failingReader struct {
|
||||
data []byte
|
||||
done bool
|
||||
}
|
||||
|
||||
func (r *failingReader) Read(p []byte) (int, error) {
|
||||
if r.done {
|
||||
return 0, errStreamInterrupted
|
||||
}
|
||||
|
||||
n := copy(p, r.data)
|
||||
r.done = true
|
||||
|
||||
return n, nil
|
||||
}
|
||||
|
||||
// TestFileStorer_InterruptedWriteLeavesNoTrustedObject checks that a write
|
||||
// cut off mid-stream leaves nothing at the destination key, so a later run
|
||||
// cannot Stat a truncated object and trust it as a complete blob.
|
||||
func TestFileStorer_InterruptedWriteLeavesNoTrustedObject(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
f, err := storage.NewFileStorer(t.TempDir())
|
||||
if err != nil {
|
||||
t.Fatalf("NewFileStorer: %v", err)
|
||||
}
|
||||
|
||||
ctx := context.Background()
|
||||
key := "blobs/aa/bb/aabbccddeeff"
|
||||
|
||||
err = f.PutWithProgress(ctx, key, &failingReader{data: []byte("partial")}, 4096, nil)
|
||||
if err == nil {
|
||||
t.Fatal("expected the interrupted write to fail, got nil")
|
||||
}
|
||||
|
||||
_, err = f.Stat(ctx, key)
|
||||
if !errors.Is(err, storage.ErrNotFound) {
|
||||
t.Fatalf("expected key absent after interrupted write, got Stat err %v", err)
|
||||
}
|
||||
|
||||
keys, err := f.List(ctx, "blobs/")
|
||||
if err != nil {
|
||||
t.Fatalf("List: %v", err)
|
||||
}
|
||||
|
||||
if len(keys) != 0 {
|
||||
t.Fatalf("expected no keys listed after interrupted write, got %v", keys)
|
||||
}
|
||||
}
|
||||
|
||||
// TestFileStorer_ListSkipsPartialFiles checks that a leftover temp file (the
|
||||
// storage layer names them with a ".partial" suffix) is never surfaced as a
|
||||
// key by List or ListStream.
|
||||
func TestFileStorer_ListSkipsPartialFiles(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
base := t.TempDir()
|
||||
|
||||
f, err := storage.NewFileStorer(base)
|
||||
if err != nil {
|
||||
t.Fatalf("NewFileStorer: %v", err)
|
||||
}
|
||||
|
||||
ctx := context.Background()
|
||||
realKey := "blobs/aa/bb/aabbccddeeff"
|
||||
|
||||
err = f.Put(ctx, realKey, strings.NewReader("blob-bytes"))
|
||||
if err != nil {
|
||||
t.Fatalf("Put: %v", err)
|
||||
}
|
||||
|
||||
// A stray temp file, as an interrupted write would leave behind.
|
||||
leftover := filepath.Join(base, "blobs/aa/bb/aabbccddeeff-123456.partial")
|
||||
|
||||
err = os.WriteFile(leftover, []byte("half"), 0o600)
|
||||
if err != nil {
|
||||
t.Fatalf("writing leftover temp file: %v", err)
|
||||
}
|
||||
|
||||
keys, err := f.List(ctx, "blobs/")
|
||||
if err != nil {
|
||||
t.Fatalf("List: %v", err)
|
||||
}
|
||||
|
||||
if len(keys) != 1 || keys[0] != realKey {
|
||||
t.Fatalf("List should return only the real key, got %v", keys)
|
||||
}
|
||||
|
||||
var streamed []string
|
||||
|
||||
for obj := range f.ListStream(ctx, "blobs/") {
|
||||
if obj.Err != nil {
|
||||
t.Fatalf("ListStream: %v", obj.Err)
|
||||
}
|
||||
|
||||
streamed = append(streamed, obj.Key)
|
||||
}
|
||||
|
||||
if len(streamed) != 1 || streamed[0] != realKey {
|
||||
t.Fatalf("ListStream should return only the real key, got %v", streamed)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,27 @@
|
||||
package storage_test
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
)
|
||||
|
||||
// newFileStorer builds a file:// backend rooted at a fresh temp directory.
|
||||
//
|
||||
//nolint:ireturn // conformance runs against the Storer interface by design
|
||||
func newFileStorer(t *testing.T) storage.Storer {
|
||||
t.Helper()
|
||||
|
||||
s, err := storage.NewFileStorer(t.TempDir())
|
||||
if err != nil {
|
||||
t.Fatalf("NewFileStorer: %v", err)
|
||||
}
|
||||
|
||||
return s
|
||||
}
|
||||
|
||||
// TestFileStorer runs the shared Storer contract against the file:// backend.
|
||||
func TestFileStorer(t *testing.T) {
|
||||
t.Parallel()
|
||||
runStorerConformance(t, newFileStorer)
|
||||
}
|
||||
@@ -111,10 +111,11 @@ func storerFromParsedS3URL(parsed *URL, cfg *config.Config) (Storer, error) {
|
||||
func storerFromLegacyS3Config(cfg *config.Config) (Storer, error) {
|
||||
endpoint := cfg.S3.Endpoint
|
||||
|
||||
// Ensure protocol is present
|
||||
// Ensure protocol is present. Absent an explicit use_ssl, default to TLS;
|
||||
// plain HTTP only when use_ssl is written as false.
|
||||
if !strings.HasPrefix(endpoint, "http://") &&
|
||||
!strings.HasPrefix(endpoint, "https://") {
|
||||
if cfg.S3.UseSSL {
|
||||
if cfg.S3.UseSSL == nil || *cfg.S3.UseSSL {
|
||||
endpoint = "https://" + endpoint
|
||||
} else {
|
||||
endpoint = "http://" + endpoint
|
||||
|
||||
@@ -0,0 +1,61 @@
|
||||
package storage_test
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/config"
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
)
|
||||
|
||||
// legacyS3Config returns a minimal s3.* (no storage_url) configuration with a
|
||||
// scheme-less endpoint. useSSL mirrors the config file: nil means the key is
|
||||
// omitted, a pointer means it was written explicitly.
|
||||
func legacyS3Config(useSSL *bool) *config.Config {
|
||||
return &config.Config{
|
||||
S3: config.S3Config{
|
||||
Endpoint: "s3.example.com",
|
||||
Bucket: "bucket",
|
||||
AccessKeyID: "key",
|
||||
SecretAccessKey: "secret",
|
||||
Region: "us-east-1",
|
||||
UseSSL: useSSL,
|
||||
},
|
||||
}
|
||||
}
|
||||
|
||||
// endpointScheme builds the storer from cfg and returns the scheme its
|
||||
// resolved endpoint carries (Info().Location is "endpoint/bucket").
|
||||
func endpointScheme(t *testing.T, cfg *config.Config) string {
|
||||
t.Helper()
|
||||
|
||||
storer, err := storage.NewStorer(cfg)
|
||||
if err != nil {
|
||||
t.Fatalf("NewStorer: %v", err)
|
||||
}
|
||||
|
||||
location := storer.Info().Location
|
||||
switch {
|
||||
case strings.HasPrefix(location, "https://"):
|
||||
return "https"
|
||||
case strings.HasPrefix(location, "http://"):
|
||||
return "http"
|
||||
default:
|
||||
t.Fatalf("endpoint has no http(s) scheme: %q", location)
|
||||
|
||||
return ""
|
||||
}
|
||||
}
|
||||
|
||||
func TestLegacyS3SchemelessEndpointDefaultsToTLS(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
if got := endpointScheme(t, legacyS3Config(nil)); got != "https" {
|
||||
t.Errorf("use_ssl omitted: got %q scheme, want https", got)
|
||||
}
|
||||
|
||||
no := false
|
||||
if got := endpointScheme(t, legacyS3Config(&no)); got != "http" {
|
||||
t.Errorf("use_ssl: false: got %q scheme, want http", got)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,58 @@
|
||||
package storage_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"testing"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
)
|
||||
|
||||
// The rclone backend is a thin adapter over the rclone library: it turns a
|
||||
// (remote, path) pair into rclone's "remote:path" string, hands it to
|
||||
// rclone, and maps rclone's own results back to the Storer interface. What
|
||||
// can be tested in-process, without a configured remote or network, is that
|
||||
// adapter layer — how the arguments are shaped and how construction errors
|
||||
// are reported. The data-plane operations (Put/Get/List/Delete) are rclone's
|
||||
// own, exercised against a real provider (drive, s3-via-rclone, ...), which
|
||||
// needs a configured remote with credentials and network access and so is
|
||||
// out of reach of a unit test. The shared Storer conformance suite therefore
|
||||
// runs against the in-process file and s3 backends; the rclone backend
|
||||
// inherits that contract once a remote is configured.
|
||||
//
|
||||
// These tests use rclone's ":local:" on-the-fly backend, which addresses the
|
||||
// local filesystem directly without any configured remote, so construction
|
||||
// runs entirely in-process.
|
||||
|
||||
// TestNewRcloneStorerConstruction checks that a valid remote constructs a
|
||||
// backend and that Info() reports the shaped "remote:path" location.
|
||||
//
|
||||
//nolint:paralleltest // NewRcloneStorer installs the process-global rclone config
|
||||
func TestNewRcloneStorerConstruction(t *testing.T) {
|
||||
dir := t.TempDir()
|
||||
|
||||
s, err := storage.NewRcloneStorer(context.Background(), ":local", dir)
|
||||
if err != nil {
|
||||
t.Fatalf("NewRcloneStorer: %v", err)
|
||||
}
|
||||
|
||||
// Info().Location is the "remote:path" string the adapter builds from
|
||||
// its two arguments, so asserting it confirms the argument shaping.
|
||||
want := ":local:" + dir
|
||||
if got := s.Info().Location; got != want {
|
||||
t.Errorf("Info().Location = %q, want %q", got, want)
|
||||
}
|
||||
}
|
||||
|
||||
// TestNewRcloneStorerUnknownRemote checks that a remote that is not in the
|
||||
// rclone config fails construction with the ErrRemoteNotFound sentinel,
|
||||
// rather than silently returning a backend pointed nowhere.
|
||||
//
|
||||
//nolint:paralleltest // NewRcloneStorer installs the process-global rclone config
|
||||
func TestNewRcloneStorerUnknownRemote(t *testing.T) {
|
||||
_, err := storage.NewRcloneStorer(
|
||||
context.Background(), "vaultik-no-such-remote", "path")
|
||||
if !errors.Is(err, storage.ErrRemoteNotFound) {
|
||||
t.Errorf("NewRcloneStorer error = %v, want ErrRemoteNotFound", err)
|
||||
}
|
||||
}
|
||||
+16
-1
@@ -38,14 +38,29 @@ func (s *S3Storer) PutWithProgress(
|
||||
}
|
||||
|
||||
// Get retrieves data from the specified key.
|
||||
// Returns ErrNotFound if the object does not exist.
|
||||
func (s *S3Storer) Get(ctx context.Context, key string) (io.ReadCloser, error) {
|
||||
return s.client.GetObject(ctx, key)
|
||||
rc, err := s.client.GetObject(ctx, key)
|
||||
if err != nil {
|
||||
if s3.IsNotFound(err) {
|
||||
return nil, fmt.Errorf("get %q: %w", key, ErrNotFound)
|
||||
}
|
||||
|
||||
return nil, err
|
||||
}
|
||||
|
||||
return rc, nil
|
||||
}
|
||||
|
||||
// Stat returns metadata about an object without retrieving its contents.
|
||||
// Returns ErrNotFound if the object does not exist.
|
||||
func (s *S3Storer) Stat(ctx context.Context, key string) (*ObjectInfo, error) {
|
||||
info, err := s.client.StatObject(ctx, key)
|
||||
if err != nil {
|
||||
if s3.IsNotFound(err) {
|
||||
return nil, fmt.Errorf("stat %q: %w", key, ErrNotFound)
|
||||
}
|
||||
|
||||
return nil, err
|
||||
}
|
||||
|
||||
|
||||
@@ -0,0 +1,81 @@
|
||||
package storage_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"net/http/httptest"
|
||||
"testing"
|
||||
|
||||
"github.com/johannesboyne/gofakes3"
|
||||
"github.com/johannesboyne/gofakes3/backend/s3mem"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/s3"
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
)
|
||||
|
||||
// s3TestBucket is the bucket created for each in-process S3 server.
|
||||
const s3TestBucket = "test-bucket"
|
||||
|
||||
// newS3Storer builds an s3:// backend backed by a fresh in-process
|
||||
// S3 server. It reuses the same in-memory S3 harness (gofakes3 + s3mem
|
||||
// over httptest) that internal/s3 and the not-found regression test use,
|
||||
// so no new mock or dependency is introduced. Each call gets its own
|
||||
// server, bucket, and client, so the conformance suite's per-section
|
||||
// instances stay isolated.
|
||||
//
|
||||
//nolint:ireturn // conformance runs against the Storer interface by design
|
||||
func newS3Storer(t *testing.T) storage.Storer {
|
||||
t.Helper()
|
||||
|
||||
backend := s3mem.New()
|
||||
|
||||
err := backend.CreateBucket(s3TestBucket)
|
||||
if err != nil {
|
||||
t.Fatalf("create bucket: %v", err)
|
||||
}
|
||||
|
||||
srv := httptest.NewServer(gofakes3.New(backend).Server())
|
||||
t.Cleanup(srv.Close)
|
||||
|
||||
client, err := s3.NewClient(context.Background(), s3.Config{
|
||||
Endpoint: srv.URL,
|
||||
Bucket: s3TestBucket,
|
||||
AccessKeyID: "test",
|
||||
SecretAccessKey: "test",
|
||||
Region: "us-east-1",
|
||||
})
|
||||
if err != nil {
|
||||
t.Fatalf("new client: %v", err)
|
||||
}
|
||||
|
||||
return storage.NewS3Storer(client)
|
||||
}
|
||||
|
||||
// TestS3Storer runs the shared Storer contract against the s3:// backend,
|
||||
// so it is held to the same round-trip, list, delete, and not-found
|
||||
// behaviour as the file:// backend.
|
||||
func TestS3Storer(t *testing.T) {
|
||||
t.Parallel()
|
||||
runStorerConformance(t, newS3Storer)
|
||||
}
|
||||
|
||||
// TestS3StorerMissingKeyMapsToErrNotFound pins the specific contract that a
|
||||
// missing object surfaces as storage.ErrNotFound rather than the raw AWS SDK
|
||||
// error. Without the mapping, errors.Is(err, storage.ErrNotFound) is false on
|
||||
// s3 and callers would branch differently per backend.
|
||||
func TestS3StorerMissingKeyMapsToErrNotFound(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
storer := newS3Storer(t)
|
||||
ctx := context.Background()
|
||||
|
||||
_, err := storer.Get(ctx, "does-not-exist")
|
||||
if !errors.Is(err, storage.ErrNotFound) {
|
||||
t.Errorf("Get on missing key: got %v, want ErrNotFound", err)
|
||||
}
|
||||
|
||||
_, err = storer.Stat(ctx, "does-not-exist")
|
||||
if !errors.Is(err, storage.ErrNotFound) {
|
||||
t.Errorf("Stat on missing key: got %v, want ErrNotFound", err)
|
||||
}
|
||||
}
|
||||
+71
-16
@@ -4,6 +4,7 @@ import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"net/url"
|
||||
"slices"
|
||||
"strings"
|
||||
)
|
||||
|
||||
@@ -23,6 +24,10 @@ var (
|
||||
ErrUnsupportedScheme = errors.New(
|
||||
"unsupported URL scheme: must start with s3://, file://, or rclone://")
|
||||
ErrUnsupportedStorage = errors.New("unsupported storage scheme")
|
||||
ErrURLCredentials = errors.New(
|
||||
"storage URL must not carry credentials; " +
|
||||
"set s3.access_key_id and s3.secret_access_key in the config instead")
|
||||
ErrURLUnknownParam = errors.New("unknown query parameter in storage URL")
|
||||
)
|
||||
|
||||
// URL represents a parsed storage URL.
|
||||
@@ -59,11 +64,28 @@ func ParseStorageURL(rawURL string) (*URL, error) {
|
||||
}, nil
|
||||
}
|
||||
|
||||
// Handle s3:// URLs
|
||||
if strings.HasPrefix(rawURL, "s3://") {
|
||||
return parseS3URL(rawURL)
|
||||
}
|
||||
|
||||
if strings.HasPrefix(rawURL, "rclone://") {
|
||||
return parseRcloneURL(rawURL)
|
||||
}
|
||||
|
||||
return nil, ErrUnsupportedScheme
|
||||
}
|
||||
|
||||
// parseS3URL parses an s3://bucket/prefix URL. It rejects credentials in
|
||||
// the userinfo and any query parameter other than endpoint, region and
|
||||
// ssl, so a credential-bearing URL is never stored or echoed.
|
||||
func parseS3URL(rawURL string) (*URL, error) {
|
||||
u, err := url.Parse(rawURL)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("invalid URL: %w", err)
|
||||
return nil, wrapParseError(err)
|
||||
}
|
||||
|
||||
if u.User != nil {
|
||||
return nil, ErrURLCredentials
|
||||
}
|
||||
|
||||
bucket := u.Host
|
||||
@@ -71,30 +93,34 @@ func ParseStorageURL(rawURL string) (*URL, error) {
|
||||
return nil, ErrMissingBucket
|
||||
}
|
||||
|
||||
prefix := strings.TrimPrefix(u.Path, "/")
|
||||
|
||||
query := u.Query()
|
||||
|
||||
useSSL := true
|
||||
if query.Get("ssl") == "false" {
|
||||
useSSL = false
|
||||
err = rejectUnknownParams(query, "endpoint", "region", "ssl")
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
return &URL{
|
||||
Scheme: schemeS3,
|
||||
Bucket: bucket,
|
||||
Prefix: prefix,
|
||||
Prefix: strings.TrimPrefix(u.Path, "/"),
|
||||
Endpoint: query.Get("endpoint"),
|
||||
Region: query.Get("region"),
|
||||
UseSSL: useSSL,
|
||||
UseSSL: query.Get("ssl") != "false",
|
||||
}, nil
|
||||
}
|
||||
}
|
||||
|
||||
// Handle rclone:// URLs
|
||||
if strings.HasPrefix(rawURL, "rclone://") {
|
||||
// parseRcloneURL parses an rclone://remote/path URL. rclone:// takes no
|
||||
// query parameters, so credentials in the userinfo and any parameter at
|
||||
// all are rejected rather than silently ignored.
|
||||
func parseRcloneURL(rawURL string) (*URL, error) {
|
||||
u, err := url.Parse(rawURL)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("invalid URL: %w", err)
|
||||
return nil, wrapParseError(err)
|
||||
}
|
||||
|
||||
if u.User != nil {
|
||||
return nil, ErrURLCredentials
|
||||
}
|
||||
|
||||
remote := u.Host
|
||||
@@ -102,16 +128,45 @@ func ParseStorageURL(rawURL string) (*URL, error) {
|
||||
return nil, ErrMissingRemote
|
||||
}
|
||||
|
||||
path := strings.TrimPrefix(u.Path, "/")
|
||||
err = rejectUnknownParams(u.Query())
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
return &URL{
|
||||
Scheme: schemeRclone,
|
||||
Prefix: path,
|
||||
Prefix: strings.TrimPrefix(u.Path, "/"),
|
||||
RcloneRemote: remote,
|
||||
}, nil
|
||||
}
|
||||
|
||||
// rejectUnknownParams returns an error naming the first query parameter
|
||||
// not in allowed. The parameter's name is included (so a misspelt
|
||||
// endpoint= is caught), but never its value, which could be a secret,
|
||||
// and never the whole URL.
|
||||
func rejectUnknownParams(query url.Values, allowed ...string) error {
|
||||
for name := range query {
|
||||
if !slices.Contains(allowed, name) {
|
||||
return fmt.Errorf(
|
||||
"%w: %q; put credentials in s3.access_key_id and "+
|
||||
"s3.secret_access_key, not the URL",
|
||||
ErrURLUnknownParam, name)
|
||||
}
|
||||
}
|
||||
|
||||
return nil, ErrUnsupportedScheme
|
||||
return nil
|
||||
}
|
||||
|
||||
// wrapParseError wraps only the inner cause of a url.Parse failure. The
|
||||
// *url.Error that url.Parse returns embeds the raw URL in its message, so
|
||||
// wrapping it directly would echo a credential-bearing URL into logs.
|
||||
func wrapParseError(err error) error {
|
||||
var uerr *url.Error
|
||||
if errors.As(err, &uerr) {
|
||||
return fmt.Errorf("invalid URL: %w", uerr.Err)
|
||||
}
|
||||
|
||||
return fmt.Errorf("invalid URL: %w", err)
|
||||
}
|
||||
|
||||
// String returns a human-readable representation of the storage URL.
|
||||
|
||||
@@ -0,0 +1,208 @@
|
||||
package storage_test
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"reflect"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
)
|
||||
|
||||
// TestParseStorageURLValid checks that each supported scheme parses into
|
||||
// the expected fields, since those fields decide which backend is built.
|
||||
func TestParseStorageURLValid(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const bucket = "mybucket"
|
||||
|
||||
cases := []struct {
|
||||
name string
|
||||
raw string
|
||||
want *storage.URL
|
||||
}{
|
||||
{
|
||||
name: "file absolute path",
|
||||
raw: "file:///var/backups/vaultik",
|
||||
want: &storage.URL{Scheme: "file", Prefix: "/var/backups/vaultik"},
|
||||
},
|
||||
{
|
||||
name: "s3 bucket and prefix, ssl defaults on",
|
||||
raw: "s3://mybucket/backups/host",
|
||||
want: &storage.URL{
|
||||
Scheme: "s3", Bucket: bucket,
|
||||
Prefix: "backups/host", UseSSL: true,
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "s3 bucket only",
|
||||
raw: "s3://mybucket",
|
||||
want: &storage.URL{Scheme: "s3", Bucket: bucket, UseSSL: true},
|
||||
},
|
||||
{
|
||||
name: "s3 with endpoint, region, ssl off",
|
||||
raw: "s3://mybucket?endpoint=minio.example.com®ion=us-west-2&ssl=false",
|
||||
want: &storage.URL{
|
||||
Scheme: "s3", Bucket: bucket,
|
||||
Endpoint: "minio.example.com", Region: "us-west-2", UseSSL: false,
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "rclone remote and path",
|
||||
raw: "rclone://gdrive/backups/host",
|
||||
want: &storage.URL{
|
||||
Scheme: "rclone", RcloneRemote: "gdrive", Prefix: "backups/host",
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "rclone remote only",
|
||||
raw: "rclone://gdrive",
|
||||
want: &storage.URL{Scheme: "rclone", RcloneRemote: "gdrive"},
|
||||
},
|
||||
}
|
||||
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
got, err := storage.ParseStorageURL(tc.raw)
|
||||
if err != nil {
|
||||
t.Fatalf("ParseStorageURL(%q) returned error: %v", tc.raw, err)
|
||||
}
|
||||
|
||||
if !reflect.DeepEqual(got, tc.want) {
|
||||
t.Errorf("ParseStorageURL(%q) = %+v, want %+v", tc.raw, got, tc.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestParseStorageURLErrors checks that empty, missing, and unknown-scheme
|
||||
// inputs fail with the documented sentinel errors instead of parsing to a
|
||||
// wrong destination.
|
||||
func TestParseStorageURLErrors(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
cases := []struct {
|
||||
name string
|
||||
raw string
|
||||
wantErr error
|
||||
}{
|
||||
{"empty url", "", storage.ErrEmptyStorageURL},
|
||||
{"file empty path", "file://", storage.ErrEmptyFilePath},
|
||||
{"s3 missing bucket", "s3://", storage.ErrMissingBucket},
|
||||
{"s3 missing bucket with path", "s3:///justprefix", storage.ErrMissingBucket},
|
||||
{"rclone missing remote", "rclone://", storage.ErrMissingRemote},
|
||||
{"unknown scheme", "gs://bucket/x", storage.ErrUnsupportedScheme},
|
||||
{"no scheme", "/local/path", storage.ErrUnsupportedScheme},
|
||||
}
|
||||
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
_, err := storage.ParseStorageURL(tc.raw)
|
||||
if !errors.Is(err, tc.wantErr) {
|
||||
t.Errorf("ParseStorageURL(%q) error = %v, want %v",
|
||||
tc.raw, err, tc.wantErr)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestParseStorageURLRejectsCredentials checks that a URL carrying
|
||||
// credentials in its userinfo or in an unknown query parameter is
|
||||
// rejected, and that the error never echoes the secret-bearing URL back
|
||||
// into logs or output.
|
||||
func TestParseStorageURLRejectsCredentials(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
// Split so the literals never form a "user:pass@" URL pattern that
|
||||
// tooling would flag as a real hardcoded credential.
|
||||
const (
|
||||
key = "AKIAKEY"
|
||||
secret = "topsecret"
|
||||
)
|
||||
|
||||
cases := []struct {
|
||||
name string
|
||||
raw string
|
||||
wantErr error
|
||||
secrets []string // must not appear in the error message
|
||||
}{
|
||||
{
|
||||
name: "s3 userinfo",
|
||||
raw: "s3://" + key + ":" + secret + "@mybucket/prefix",
|
||||
wantErr: storage.ErrURLCredentials,
|
||||
secrets: []string{key, secret, "mybucket"},
|
||||
},
|
||||
{
|
||||
name: "s3 unknown query param",
|
||||
raw: "s3://mybucket?access_key=" + key + "&secret=" + secret,
|
||||
wantErr: storage.ErrURLUnknownParam,
|
||||
secrets: []string{key, secret},
|
||||
},
|
||||
{
|
||||
name: "s3 misspelt endpoint",
|
||||
raw: "s3://mybucket?endpiont=minio.example.com",
|
||||
wantErr: storage.ErrURLUnknownParam,
|
||||
secrets: nil,
|
||||
},
|
||||
{
|
||||
name: "rclone userinfo",
|
||||
raw: "rclone://user:" + secret + "@gdrive/backups",
|
||||
wantErr: storage.ErrURLCredentials,
|
||||
secrets: []string{secret},
|
||||
},
|
||||
{
|
||||
name: "rclone query param",
|
||||
raw: "rclone://gdrive/backups?token=" + secret,
|
||||
wantErr: storage.ErrURLUnknownParam,
|
||||
secrets: []string{secret},
|
||||
},
|
||||
}
|
||||
|
||||
for _, tc := range cases {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
_, err := storage.ParseStorageURL(tc.raw)
|
||||
if !errors.Is(err, tc.wantErr) {
|
||||
t.Fatalf("ParseStorageURL(%q) error = %v, want %v",
|
||||
tc.raw, err, tc.wantErr)
|
||||
}
|
||||
|
||||
// The rejection must name the proper config keys so the
|
||||
// operator knows where credentials belong.
|
||||
for _, key := range []string{"s3.access_key_id", "s3.secret_access_key"} {
|
||||
if !strings.Contains(err.Error(), key) {
|
||||
t.Errorf("error %q does not name %q", err.Error(), key)
|
||||
}
|
||||
}
|
||||
|
||||
for _, secret := range tc.secrets {
|
||||
if strings.Contains(err.Error(), secret) {
|
||||
t.Errorf("error message leaked %q: %v", secret, err.Error())
|
||||
}
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestParseStorageURLParseFailureHidesURL checks that when url.Parse
|
||||
// itself fails, the wrapped error carries only the inner cause, not the
|
||||
// *url.Error whose text embeds the raw (possibly credential-bearing) URL.
|
||||
func TestParseStorageURLParseFailureHidesURL(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const raw = "s3://mybucket/%zz"
|
||||
|
||||
_, err := storage.ParseStorageURL(raw)
|
||||
if err == nil {
|
||||
t.Fatalf("ParseStorageURL(%q) returned no error", raw)
|
||||
}
|
||||
|
||||
if strings.Contains(err.Error(), "mybucket") {
|
||||
t.Errorf("error message echoed the raw URL: %v", err.Error())
|
||||
}
|
||||
}
|
||||
@@ -18,6 +18,13 @@ import (
|
||||
// not match the expected double-SHA-256 hash.
|
||||
var errBlobHashMismatch = errors.New("blob hash mismatch")
|
||||
|
||||
// errBlobNotFullyRead is returned when the verifying reader is closed
|
||||
// before its plaintext reached EOF. The hash can only be checked once
|
||||
// the whole stream has been read, so an early or short-read close must
|
||||
// fail rather than silently skip verification.
|
||||
var errBlobNotFullyRead = errors.New(
|
||||
"blob closed before fully read; hash not verified")
|
||||
|
||||
// hashVerifyReader wraps a blobgen.Reader and verifies the double-SHA-256 hash
|
||||
// of decrypted plaintext when Close is called. It reuses the hash that
|
||||
// blobgen.Reader already computes internally via its TeeReader, avoiding
|
||||
@@ -38,12 +45,18 @@ func (h *hashVerifyReader) Read(p []byte) (int, error) {
|
||||
return n, err
|
||||
}
|
||||
|
||||
// Close verifies the hash (if the stream was fully read) and closes underlying readers.
|
||||
// Close closes the underlying readers and verifies the blob hash. The
|
||||
// hash check cannot be skipped: closing before the plaintext reached
|
||||
// EOF (a short read or an early close) is an error, so a caller can
|
||||
// never obtain unverified blob bytes.
|
||||
func (h *hashVerifyReader) Close() error {
|
||||
readerErr := h.reader.Close()
|
||||
fetcherErr := h.fetcher.Close()
|
||||
|
||||
if h.done {
|
||||
if !h.done {
|
||||
return errBlobNotFullyRead
|
||||
}
|
||||
|
||||
firstHash := h.reader.Sum256()
|
||||
secondHasher := sha256.New()
|
||||
secondHasher.Write(firstHash)
|
||||
@@ -53,7 +66,6 @@ func (h *hashVerifyReader) Close() error {
|
||||
return fmt.Errorf("%w: expected %s, got %s",
|
||||
errBlobHashMismatch, h.blobHash[:16], actualHashHex[:16])
|
||||
}
|
||||
}
|
||||
|
||||
if readerErr != nil {
|
||||
return readerErr
|
||||
|
||||
@@ -133,3 +133,51 @@ func TestFetchAndDecryptBlobVerifiesHash(t *testing.T) {
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
// TestFetchAndDecryptBlobCloseBeforeEOFFails verifies the hash check
|
||||
// cannot be skipped: a caller that reads only part of the blob and then
|
||||
// closes gets an error rather than silently unverified bytes.
|
||||
func TestFetchAndDecryptBlobCloseBeforeEOFFails(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
identity, err := age.GenerateX25519Identity()
|
||||
if err != nil {
|
||||
t.Fatalf("generating identity: %v", err)
|
||||
}
|
||||
|
||||
plaintext := []byte("hello world test data for blob hash verification")
|
||||
encryptedData, correctHash := buildHashTestBlob(t, identity, plaintext)
|
||||
|
||||
mockStorage := NewMockStorer()
|
||||
blobPath := "blobs/" + correctHash[:2] + "/" +
|
||||
correctHash[2:4] + "/" + correctHash
|
||||
|
||||
mockStorage.mu.Lock()
|
||||
mockStorage.data[blobPath] = encryptedData
|
||||
mockStorage.mu.Unlock()
|
||||
|
||||
tv := vaultik.NewForTesting(mockStorage)
|
||||
|
||||
rc, err := tv.FetchAndDecryptBlob(
|
||||
context.Background(), correctHash, int64(len(encryptedData)), identity)
|
||||
if err != nil {
|
||||
t.Fatalf("unexpected error opening stream: %v", err)
|
||||
}
|
||||
|
||||
// Read one byte, far short of the plaintext length, then close.
|
||||
buf := make([]byte, 1)
|
||||
|
||||
_, err = rc.Read(buf)
|
||||
if err != nil {
|
||||
t.Fatalf("reading first byte: %v", err)
|
||||
}
|
||||
|
||||
err = rc.Close()
|
||||
if err == nil {
|
||||
t.Fatal("expected error closing before EOF, got nil")
|
||||
}
|
||||
|
||||
if !strings.Contains(err.Error(), "hash not verified") {
|
||||
t.Fatalf("expected not-verified error, got: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
@@ -0,0 +1,108 @@
|
||||
package vaultik_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/ui"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
|
||||
// TestDeepVerifyAcceptsHealthyAndRejectsCorruptBlob backs up a real
|
||||
// snapshot with the on-disk storage backend, runs deep verification on
|
||||
// it, then flips a byte inside one stored blob and runs deep
|
||||
// verification again. A healthy snapshot must pass; a corrupted blob
|
||||
// must fail. The healthy case is the regression guard: deep
|
||||
// verification used to hash the encrypted blob bytes and compare them
|
||||
// to the blob's ID (the double SHA256 of the plaintext), so it reported
|
||||
// every healthy blob as corrupt.
|
||||
func TestDeepVerifyAcceptsHealthyAndRejectsCorruptBlob(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
|
||||
dataDir := filepath.Join(tempDir, "source")
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
|
||||
chunkSize := int64(64 * 1024)
|
||||
maxBlobSize := int64(512 * 1024)
|
||||
|
||||
// One file large enough to span several chunks within a single blob.
|
||||
require.NoError(t, fs.MkdirAll(dataDir, 0o755))
|
||||
require.NoError(t, afero.WriteFile(fs,
|
||||
filepath.Join(dataDir, "data.bin"),
|
||||
bytesPattern("deep-", int(chunkSize*3)), 0o644))
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
// runFileStorageBackup writes a real snapshot to storeDir and closes
|
||||
// the source index, so verification runs from remote bytes only.
|
||||
cfg, storer, snapshotID := runFileStorageBackup(
|
||||
ctx, t, fs, dataDir, storeDir, dbPath, chunkSize, maxBlobSize)
|
||||
|
||||
newVerifier := func() *vaultik.Vaultik {
|
||||
v := &vaultik.Vaultik{
|
||||
Config: cfg,
|
||||
Storage: storer,
|
||||
Fs: fs,
|
||||
Stdout: io.Discard,
|
||||
Stderr: io.Discard,
|
||||
UI: ui.NewWithColor(io.Discard, false),
|
||||
}
|
||||
v.SetContext(ctx)
|
||||
|
||||
return v
|
||||
}
|
||||
|
||||
require.NoError(t,
|
||||
newVerifier().RunDeepVerify(snapshotID, &vaultik.VerifyOptions{Deep: true}),
|
||||
"deep verify should pass on a healthy snapshot")
|
||||
|
||||
// Flip a byte inside one blob without changing its length, so the
|
||||
// blob-existence and size checks still pass and verification reaches
|
||||
// the blob-content stage.
|
||||
corruptOneBlob(t, fs, filepath.Join(storeDir, "blobs"))
|
||||
|
||||
require.Error(t,
|
||||
newVerifier().RunDeepVerify(snapshotID, &vaultik.VerifyOptions{Deep: true}),
|
||||
"deep verify should fail on a corrupted blob")
|
||||
}
|
||||
|
||||
// corruptOneBlob flips a middle byte of the first blob file found under
|
||||
// blobsDir, leaving the file length unchanged.
|
||||
func corruptOneBlob(t *testing.T, fs afero.Fs, blobsDir string) {
|
||||
t.Helper()
|
||||
|
||||
var blobPath string
|
||||
|
||||
err := afero.Walk(fs, blobsDir,
|
||||
func(path string, info os.FileInfo, err error) error {
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
if blobPath == "" && !info.IsDir() {
|
||||
blobPath = path
|
||||
}
|
||||
|
||||
return nil
|
||||
})
|
||||
require.NoError(t, err)
|
||||
require.NotEmpty(t, blobPath, "expected at least one blob on disk")
|
||||
|
||||
data, err := afero.ReadFile(fs, blobPath)
|
||||
require.NoError(t, err)
|
||||
require.NotEmpty(t, data)
|
||||
|
||||
data[len(data)/2] ^= 0xff
|
||||
require.NoError(t, afero.WriteFile(fs, blobPath, data, 0o644))
|
||||
}
|
||||
@@ -0,0 +1,647 @@
|
||||
package vaultik_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"errors"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/config"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
"sneak.berlin/go/vaultik/internal/storage/faultstore"
|
||||
"sneak.berlin/go/vaultik/internal/ui"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
|
||||
// These tests cover the failure modes a backup tool must survive:
|
||||
// interrupted uploads, an interrupted metadata export, corrupt and
|
||||
// truncated reads, a full restore disk, and a backend that reports
|
||||
// success while storing nothing. Faults are injected through the
|
||||
// storage.Storer seam (internal/storage/faultstore), never by patching
|
||||
// production code. Each test asserts on the observable end state — what
|
||||
// is in the index, what is at the destination, what the user is told —
|
||||
// not merely that an error was returned. See
|
||||
// https://git.eeqj.de/sneak/vaultik/issues/72.
|
||||
//
|
||||
// Object-level write atomicity (no partial blob object left behind) is
|
||||
// covered by the file:// backend's atomic-write work
|
||||
// (https://git.eeqj.de/sneak/vaultik/issues/130) and is not re-tested
|
||||
// here; these tests target the layers above the backend.
|
||||
//
|
||||
// The tests run serially, not with t.Parallel: each calls
|
||||
// log.Initialize, which replaces the package-global logger, and a
|
||||
// backup or restore running concurrently reads that same logger. Under
|
||||
// -race the two collide. Running one at a time is the same choice
|
||||
// prune_count_test.go already makes for the same reason.
|
||||
|
||||
const (
|
||||
faultChunkSize = int64(64 * 1024)
|
||||
faultMaxBlobSize = int64(256 * 1024)
|
||||
)
|
||||
|
||||
// faultTestConfig returns the config shared by the fault-injection
|
||||
// tests: a real recipient/secret keypair so blobs are genuinely
|
||||
// encrypted, and a blob size limit the restore sweeper can divide.
|
||||
func faultTestConfig() *config.Config {
|
||||
return &config.Config{
|
||||
AgeRecipients: []string{testAgePublicKey},
|
||||
AgeSecretKey: testAgeSecretKey,
|
||||
CompressionLevel: 3,
|
||||
Hostname: testHostname,
|
||||
BlobSizeLimit: config.Size(faultMaxBlobSize),
|
||||
}
|
||||
}
|
||||
|
||||
// writeFaultSourceTree writes a spread of file sizes that forces several
|
||||
// chunks across more than one blob, so a fault landing on a single blob
|
||||
// still leaves other data intact. Returns the expected content by path.
|
||||
func writeFaultSourceTree(
|
||||
t *testing.T, fs afero.Fs, dataDir string,
|
||||
) map[string][]byte {
|
||||
t.Helper()
|
||||
|
||||
files := map[string][]byte{
|
||||
filepath.Join(dataDir, "small.txt"): []byte("hello vaultik"),
|
||||
filepath.Join(dataDir, "a.bin"): bytesPattern("a-", int(faultChunkSize*3)),
|
||||
filepath.Join(dataDir, "sub", "b.bin"): bytesPattern("b-", int(faultChunkSize*3)),
|
||||
filepath.Join(dataDir, "sub", "c.bin"): bytesPattern("c-", int(faultChunkSize*2)),
|
||||
}
|
||||
|
||||
for path, content := range files {
|
||||
require.NoError(t, fs.MkdirAll(filepath.Dir(path), 0o755))
|
||||
require.NoError(t, afero.WriteFile(fs, path, content, 0o644))
|
||||
}
|
||||
|
||||
return files
|
||||
}
|
||||
|
||||
// newFaultScanner builds a scanner writing through the given storer.
|
||||
func newFaultScanner(
|
||||
fs afero.Fs, storer storage.Storer,
|
||||
cfg *config.Config, repos *database.Repositories,
|
||||
) *snapshot.Scanner {
|
||||
return snapshot.NewScanner(snapshot.ScannerConfig{
|
||||
FS: fs,
|
||||
Storage: storer,
|
||||
ChunkSize: faultChunkSize,
|
||||
MaxBlobSize: faultMaxBlobSize,
|
||||
CompressionLevel: cfg.CompressionLevel,
|
||||
AgeRecipients: cfg.AgeRecipients,
|
||||
Repositories: repos,
|
||||
})
|
||||
}
|
||||
|
||||
// newFaultSnapshotManager builds a snapshot manager writing through the
|
||||
// given storer.
|
||||
func newFaultSnapshotManager(
|
||||
fs afero.Fs, storer storage.Storer,
|
||||
cfg *config.Config, repos *database.Repositories,
|
||||
) *snapshot.SnapshotManager {
|
||||
sm := snapshot.NewSnapshotManager(snapshot.SnapshotManagerParams{
|
||||
Repos: repos,
|
||||
Storage: storer,
|
||||
Config: cfg,
|
||||
})
|
||||
sm.SetFilesystem(fs)
|
||||
|
||||
return sm
|
||||
}
|
||||
|
||||
// fullFaultBackup runs a complete backup (create, scan, complete,
|
||||
// export) through storer and returns the snapshot ID.
|
||||
func fullFaultBackup(
|
||||
ctx context.Context, t *testing.T, fs afero.Fs, storer storage.Storer,
|
||||
cfg *config.Config, repos *database.Repositories,
|
||||
dataDir, dbPath, name string,
|
||||
) string {
|
||||
t.Helper()
|
||||
|
||||
sm := newFaultSnapshotManager(fs, storer, cfg, repos)
|
||||
scanner := newFaultScanner(fs, storer, cfg, repos)
|
||||
|
||||
id, err := sm.CreateSnapshotWithName(ctx, cfg.Hostname, name, "v", "g")
|
||||
require.NoError(t, err)
|
||||
|
||||
_, err = scanner.Scan(ctx, dataDir, id)
|
||||
require.NoError(t, err)
|
||||
|
||||
require.NoError(t, sm.CompleteSnapshot(ctx, id))
|
||||
require.NoError(t, sm.ExportSnapshotMetadata(ctx, dbPath, id))
|
||||
|
||||
return id
|
||||
}
|
||||
|
||||
// newReaderVaultik builds a Vaultik that reads (restore/verify) through
|
||||
// storer, with the given repositories (nil is fine for restore/verify,
|
||||
// which read metadata from storage).
|
||||
func newReaderVaultik(
|
||||
ctx context.Context, cfg *config.Config, storer storage.Storer,
|
||||
repos *database.Repositories, fs afero.Fs,
|
||||
) *vaultik.Vaultik {
|
||||
v := &vaultik.Vaultik{
|
||||
Config: cfg,
|
||||
Storage: storer,
|
||||
Repositories: repos,
|
||||
Fs: fs,
|
||||
Stdout: io.Discard,
|
||||
Stderr: io.Discard,
|
||||
UI: ui.NewWithColor(io.Discard, false),
|
||||
}
|
||||
v.SetContext(ctx)
|
||||
|
||||
return v
|
||||
}
|
||||
|
||||
// Scenario 3: a stored blob's bytes are flipped before restore reads
|
||||
// them. Restore must fail loudly, and no file must be left on the
|
||||
// restore target holding corrupt content.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestRestoreRejectsCorruptBlob(t *testing.T) {
|
||||
assertRestoreRejectsDamagedBlob(t, faultstore.GetCorrupt, "corrupt")
|
||||
}
|
||||
|
||||
// Scenario 4: a stored blob is truncated before restore reads it. Same
|
||||
// contract as the corrupt case.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestRestoreRejectsTruncatedBlob(t *testing.T) {
|
||||
assertRestoreRejectsDamagedBlob(t, faultstore.GetTruncate, "truncated")
|
||||
}
|
||||
|
||||
// assertRestoreRejectsDamagedBlob backs up the source tree, then restores
|
||||
// through a store that damages every blob read with the given fault, and
|
||||
// asserts restore fails naming a blob and leaves no file on the target
|
||||
// holding wrong bytes. Metadata reads are returned intact so the failure
|
||||
// is isolated to the blob.
|
||||
func assertRestoreRejectsDamagedBlob(
|
||||
t *testing.T, fault faultstore.GetFault, name string,
|
||||
) {
|
||||
t.Helper()
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
dataDir := filepath.Join(tempDir, "src")
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
restoreDir := filepath.Join(tempDir, "restored")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
|
||||
ctx := context.Background()
|
||||
cfg := faultTestConfig()
|
||||
testFiles := writeFaultSourceTree(t, fs, dataDir)
|
||||
|
||||
inner, err := storage.NewFileStorer(storeDir)
|
||||
require.NoError(t, err)
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
id := fullFaultBackup(ctx, t, fs, inner, cfg, repos, dataDir, dbPath, name)
|
||||
require.NoError(t, db.Close())
|
||||
|
||||
faultStore := faultstore.New(inner)
|
||||
faultStore.OnGet = func(key string) faultstore.GetFault {
|
||||
if strings.HasPrefix(key, "blobs/") {
|
||||
return fault
|
||||
}
|
||||
|
||||
return faultstore.GetNormal
|
||||
}
|
||||
|
||||
v := newReaderVaultik(ctx, cfg, faultStore, nil, fs)
|
||||
err = v.Restore(&vaultik.RestoreOptions{SnapshotID: id, TargetDir: restoreDir})
|
||||
|
||||
require.Error(t, err, "restore must fail on a damaged blob")
|
||||
assert.Contains(t, err.Error(), "blob",
|
||||
"error should name the blob that failed")
|
||||
assertNoCorruptFiles(t, fs, restoreDir, testFiles)
|
||||
}
|
||||
|
||||
// Scenario 6: the backend accepts blob uploads and reports success but
|
||||
// stores nothing. verify --deep must catch it.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestDeepVerifyCatchesLyingBackend(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
dataDir := filepath.Join(tempDir, "src")
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
|
||||
ctx := context.Background()
|
||||
cfg := faultTestConfig()
|
||||
|
||||
writeFaultSourceTree(t, fs, dataDir)
|
||||
|
||||
inner, err := storage.NewFileStorer(storeDir)
|
||||
require.NoError(t, err)
|
||||
|
||||
// Blob uploads are swallowed; metadata uploads land, so verify can
|
||||
// download the manifest and database and then discover the blobs are
|
||||
// absent.
|
||||
lying := faultstore.New(inner)
|
||||
lying.OnPut = func(key string) faultstore.PutAction {
|
||||
if strings.HasPrefix(key, "blobs/") {
|
||||
return faultstore.PutSwallow
|
||||
}
|
||||
|
||||
return faultstore.PutNormal
|
||||
}
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
id := fullFaultBackup(ctx, t, fs, lying, cfg, repos, dataDir, dbPath, "lying")
|
||||
require.NoError(t, db.Close())
|
||||
|
||||
// No blob objects were actually written.
|
||||
blobKeys, err := inner.List(ctx, "blobs/")
|
||||
require.NoError(t, err)
|
||||
assert.Empty(t, blobKeys, "lying backend should have stored no blobs")
|
||||
|
||||
// Read back through the honest underlying store.
|
||||
v := newReaderVaultik(ctx, cfg, inner, nil, fs)
|
||||
err = v.VerifySnapshotWithOptions(id, &vaultik.VerifyOptions{Deep: true})
|
||||
require.Error(t, err, "deep verify must catch a backend that stored nothing")
|
||||
}
|
||||
|
||||
// Scenario 1a: a blob upload fails partway through. The interrupted run
|
||||
// must not record the blob as uploaded, must not reference it from the
|
||||
// snapshot, and must leave no blob object at the destination.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestInterruptedBlobUploadRecordsNoUploadedBlob(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
dataDir := filepath.Join(tempDir, "src")
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
|
||||
ctx := context.Background()
|
||||
cfg := faultTestConfig()
|
||||
|
||||
writeFaultSourceTree(t, fs, dataDir)
|
||||
|
||||
inner, err := storage.NewFileStorer(storeDir)
|
||||
require.NoError(t, err)
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
|
||||
defer func() { _ = db.Close() }()
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
// Every blob upload fails partway through. The scan must surface it.
|
||||
fault := faultstore.New(inner)
|
||||
fault.OnPut = func(key string) faultstore.PutAction {
|
||||
if strings.HasPrefix(key, "blobs/") {
|
||||
return faultstore.PutFail
|
||||
}
|
||||
|
||||
return faultstore.PutNormal
|
||||
}
|
||||
|
||||
sm := newFaultSnapshotManager(fs, fault, cfg, repos)
|
||||
scanner := newFaultScanner(fs, fault, cfg, repos)
|
||||
|
||||
id, err := sm.CreateSnapshotWithName(ctx, cfg.Hostname, "interrupted", "v", "g")
|
||||
require.NoError(t, err)
|
||||
|
||||
_, err = scanner.Scan(ctx, dataDir, id)
|
||||
require.Error(t, err, "scan must fail when a blob upload fails")
|
||||
|
||||
// No blob may claim to be uploaded.
|
||||
blobs, err := repos.Blobs.GetAll(ctx)
|
||||
require.NoError(t, err)
|
||||
|
||||
for _, b := range blobs {
|
||||
assert.Nilf(t, b.UploadedTS,
|
||||
"blob %s marked uploaded after a failed upload", b.Hash)
|
||||
}
|
||||
|
||||
// The snapshot may reference no blobs, and the destination holds none.
|
||||
hashes, err := repos.Snapshots.GetBlobHashes(ctx, id)
|
||||
require.NoError(t, err)
|
||||
assert.Empty(t, hashes, "interrupted snapshot must reference no blobs")
|
||||
|
||||
blobKeys, err := inner.List(ctx, "blobs/")
|
||||
require.NoError(t, err)
|
||||
assert.Empty(t, blobKeys, "no blob object may survive at the destination")
|
||||
}
|
||||
|
||||
// Scenario 1b: after an interrupted upload, a retry on the same local
|
||||
// index must produce a restorable snapshot. The interrupted run leaves
|
||||
// the blob's chunk rows in the index; the fix for
|
||||
// https://git.eeqj.de/sneak/vaultik/issues/148 discards those un-uploaded
|
||||
// blob rows at the start of the next scan and deduplicates only against
|
||||
// chunks in a blob that was actually uploaded, so the retry re-chunks and
|
||||
// re-uploads the affected data instead of silently referencing data that
|
||||
// never reached storage.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestBackupRetryAfterInterruptedUploadIsRestorable(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
dataDir := filepath.Join(tempDir, "src")
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
restoreDir := filepath.Join(tempDir, "restored")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
|
||||
ctx := context.Background()
|
||||
cfg := faultTestConfig()
|
||||
testFiles := writeFaultSourceTree(t, fs, dataDir)
|
||||
|
||||
inner, err := storage.NewFileStorer(storeDir)
|
||||
require.NoError(t, err)
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
// Attempt 1: every blob upload fails.
|
||||
fault := faultstore.New(inner)
|
||||
fault.OnPut = func(key string) faultstore.PutAction {
|
||||
if strings.HasPrefix(key, "blobs/") {
|
||||
return faultstore.PutFail
|
||||
}
|
||||
|
||||
return faultstore.PutNormal
|
||||
}
|
||||
|
||||
sm := newFaultSnapshotManager(fs, fault, cfg, repos)
|
||||
scanner := newFaultScanner(fs, fault, cfg, repos)
|
||||
|
||||
id1, err := sm.CreateSnapshotWithName(ctx, cfg.Hostname, "interrupted", "v", "g")
|
||||
require.NoError(t, err)
|
||||
|
||||
_, err = scanner.Scan(ctx, dataDir, id1)
|
||||
require.Error(t, err)
|
||||
|
||||
// Retry on the same local index with a working backend.
|
||||
id2 := fullFaultBackup(ctx, t, fs, inner, cfg, repos, dataDir, dbPath, "retry")
|
||||
require.NoError(t, db.Close())
|
||||
|
||||
v := newReaderVaultik(ctx, cfg, inner, nil, fs)
|
||||
require.NoError(t, v.Restore(&vaultik.RestoreOptions{
|
||||
SnapshotID: id2,
|
||||
TargetDir: restoreDir,
|
||||
Verify: true,
|
||||
}), "retry after an interrupted upload must produce a restorable snapshot")
|
||||
|
||||
assertRestoredTree(t, fs, restoreDir, testFiles)
|
||||
}
|
||||
|
||||
// Scenario 2: the process dies during the metadata export, after the
|
||||
// database is uploaded but before the manifest. The destination is left
|
||||
// with blobs and a database but no manifest. verify and snapshot list
|
||||
// must report the damage honestly rather than crashing or passing.
|
||||
// Automatic detection and repair of this partial state on the next run
|
||||
// is tracked in https://git.eeqj.de/sneak/vaultik/issues/177 and is not
|
||||
// asserted here.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestBackupSurvivesMetadataExportInterruption(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
dataDir := filepath.Join(tempDir, "src")
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
|
||||
ctx := context.Background()
|
||||
cfg := faultTestConfig()
|
||||
|
||||
writeFaultSourceTree(t, fs, dataDir)
|
||||
|
||||
inner, err := storage.NewFileStorer(storeDir)
|
||||
require.NoError(t, err)
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
// Back up and complete with a working backend.
|
||||
sm := newFaultSnapshotManager(fs, inner, cfg, repos)
|
||||
scanner := newFaultScanner(fs, inner, cfg, repos)
|
||||
|
||||
id, err := sm.CreateSnapshotWithName(ctx, cfg.Hostname, "export", "v", "g")
|
||||
require.NoError(t, err)
|
||||
|
||||
_, err = scanner.Scan(ctx, dataDir, id)
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, sm.CompleteSnapshot(ctx, id))
|
||||
|
||||
// Export through a backend that fails only the manifest upload. The
|
||||
// database uploads first and lands; the manifest does not.
|
||||
fault := faultstore.New(inner)
|
||||
fault.OnPut = func(key string) faultstore.PutAction {
|
||||
if strings.HasSuffix(key, "manifest.json.zst") {
|
||||
return faultstore.PutFail
|
||||
}
|
||||
|
||||
return faultstore.PutNormal
|
||||
}
|
||||
|
||||
smFault := newFaultSnapshotManager(fs, fault, cfg, repos)
|
||||
|
||||
err = smFault.ExportSnapshotMetadata(ctx, dbPath, id)
|
||||
require.Error(t, err, "export must fail when the manifest upload fails")
|
||||
|
||||
// The destination is in the partial state the scenario describes.
|
||||
key := snapshot.RemoteSnapshotKey(id)
|
||||
|
||||
_, err = inner.Stat(ctx, "metadata/"+key+"/db.zst.age")
|
||||
require.NoError(t, err, "database should have been uploaded before the manifest")
|
||||
|
||||
_, err = inner.Stat(ctx, "metadata/"+key+"/manifest.json.zst")
|
||||
require.ErrorIs(t, err, storage.ErrNotFound, "manifest upload should not have landed")
|
||||
|
||||
// verify must fail loudly for this snapshot, in both modes.
|
||||
reader := newReaderVaultik(ctx, cfg, inner, repos, fs)
|
||||
|
||||
deepOpts := &vaultik.VerifyOptions{Deep: true}
|
||||
require.Error(t, reader.VerifySnapshotWithOptions(id, deepOpts),
|
||||
"deep verify must report the missing manifest")
|
||||
|
||||
shallowOpts := &vaultik.VerifyOptions{Deep: false}
|
||||
require.Error(t, reader.VerifySnapshotWithOptions(id, shallowOpts),
|
||||
"shallow verify must report the missing manifest")
|
||||
|
||||
// snapshot list must not crash on the partial snapshot.
|
||||
require.NoError(t, reader.ListSnapshots(false),
|
||||
"snapshot list must tolerate a partially-exported snapshot")
|
||||
}
|
||||
|
||||
// Scenario 5: the restore target runs out of space mid-file. Restore
|
||||
// must fail with an out-of-space error, and must not leave a truncated
|
||||
// file at the target path presenting as a complete restore. Restore
|
||||
// today writes each file straight to its final path and does not remove
|
||||
// it when a write fails, so the truncated file survives; deleting it is
|
||||
// tracked by https://git.eeqj.de/sneak/vaultik/issues/163. Skipped until
|
||||
// that lands, so the destination assertion below is recorded rather than
|
||||
// dropped.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestRestoreReportsDiskFull(t *testing.T) {
|
||||
t.Skip("blocked on https://git.eeqj.de/sneak/vaultik/issues/163: " +
|
||||
"a disk-full write leaves a truncated file at the target path " +
|
||||
"instead of removing it")
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
osFS := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
dataDir := filepath.Join(tempDir, "src")
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
restoreDir := filepath.Join(tempDir, "restored")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
|
||||
ctx := context.Background()
|
||||
cfg := faultTestConfig()
|
||||
|
||||
testFiles := writeFaultSourceTree(t, osFS, dataDir)
|
||||
|
||||
inner, err := storage.NewFileStorer(storeDir)
|
||||
require.NoError(t, err)
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
id := fullFaultBackup(ctx, t, osFS, inner, cfg, repos, dataDir, dbPath, "diskfull")
|
||||
require.NoError(t, db.Close())
|
||||
|
||||
// Restore onto a filesystem that allows only a few bytes of file
|
||||
// content: enough to create files, far too little to hold them.
|
||||
budget := int64(8)
|
||||
quota := "aFS{Fs: osFS, remaining: &budget}
|
||||
|
||||
v := newReaderVaultik(ctx, cfg, inner, nil, quota)
|
||||
err = v.Restore(&vaultik.RestoreOptions{SnapshotID: id, TargetDir: restoreDir})
|
||||
|
||||
require.Error(t, err, "restore must fail when the target disk is full")
|
||||
assert.Contains(t, err.Error(), errNoSpace.Error(),
|
||||
"restore error should surface the out-of-space cause")
|
||||
|
||||
// The failure must not leave a truncated file behind presenting as a
|
||||
// complete restore: any file at the target must hold the original
|
||||
// bytes, or be absent.
|
||||
assertNoCorruptFiles(t, osFS, restoreDir, testFiles)
|
||||
}
|
||||
|
||||
// assertRestoredTree byte-compares every restored file against the
|
||||
// original.
|
||||
func assertRestoredTree(
|
||||
t *testing.T, fs afero.Fs, restoreDir string, testFiles map[string][]byte,
|
||||
) {
|
||||
t.Helper()
|
||||
|
||||
for origPath, expected := range testFiles {
|
||||
restoredPath := filepath.Join(restoreDir, origPath)
|
||||
got, err := afero.ReadFile(fs, restoredPath)
|
||||
require.NoErrorf(t, err, "restored file missing: %s", origPath)
|
||||
require.Equalf(t, expected, got, "restored content mismatch for %s", origPath)
|
||||
}
|
||||
}
|
||||
|
||||
// errNoSpace is the out-of-space error quotaFS returns once its byte
|
||||
// budget is exhausted, mirroring a real ENOSPC.
|
||||
var errNoSpace = errors.New("no space left on device")
|
||||
|
||||
// quotaFS is an afero.Fs whose files may write only a fixed total number
|
||||
// of content bytes before failing, simulating a full restore target. It
|
||||
// wraps the interface so every method except Create delegates to the
|
||||
// real filesystem; only file writes are capped.
|
||||
type quotaFS struct {
|
||||
afero.Fs
|
||||
|
||||
remaining *int64
|
||||
}
|
||||
|
||||
//nolint:ireturn // afero.Fs.Create's signature requires returning afero.File.
|
||||
func (q *quotaFS) Create(name string) (afero.File, error) {
|
||||
f, err := q.Fs.Create(name)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
return "aFile{File: f, remaining: q.remaining}, nil
|
||||
}
|
||||
|
||||
// quotaFile fails writes once the shared byte budget is exhausted.
|
||||
type quotaFile struct {
|
||||
afero.File
|
||||
|
||||
remaining *int64
|
||||
}
|
||||
|
||||
func (q *quotaFile) Write(p []byte) (int, error) {
|
||||
if *q.remaining <= 0 {
|
||||
return 0, errNoSpace
|
||||
}
|
||||
|
||||
allowed := min(int64(len(p)), *q.remaining)
|
||||
|
||||
n, err := q.File.Write(p[:allowed])
|
||||
*q.remaining -= int64(n)
|
||||
|
||||
if err != nil {
|
||||
return n, err
|
||||
}
|
||||
|
||||
if int64(n) < int64(len(p)) {
|
||||
return n, errNoSpace
|
||||
}
|
||||
|
||||
return n, nil
|
||||
}
|
||||
|
||||
// assertNoCorruptFiles fails if any file that made it to the restore
|
||||
// target holds content that differs from the original: a failed restore
|
||||
// may leave a file absent, but must never leave wrong bytes presenting
|
||||
// as the real file.
|
||||
func assertNoCorruptFiles(
|
||||
t *testing.T, fs afero.Fs, restoreDir string, testFiles map[string][]byte,
|
||||
) {
|
||||
t.Helper()
|
||||
|
||||
for origPath, expected := range testFiles {
|
||||
restoredPath := filepath.Join(restoreDir, origPath)
|
||||
|
||||
got, err := afero.ReadFile(fs, restoredPath)
|
||||
if err != nil {
|
||||
if os.IsNotExist(err) {
|
||||
continue
|
||||
}
|
||||
|
||||
require.NoError(t, err)
|
||||
}
|
||||
|
||||
assert.Equalf(t, expected, got,
|
||||
"restored file %s holds corrupt content", origPath)
|
||||
}
|
||||
}
|
||||
@@ -35,6 +35,7 @@ var (
|
||||
"invalid snapshot ID format: expected hostname_snapshotname_timestamp")
|
||||
errInvalidDuration = errors.New("invalid duration")
|
||||
errUnknownTimeUnit = errors.New("unknown time unit")
|
||||
errNegativeDuration = errors.New("negative durations are not supported")
|
||||
)
|
||||
|
||||
// Time-unit lengths used by parseDuration.
|
||||
@@ -138,8 +139,13 @@ func parseSnapshotName(snapshotID string) string {
|
||||
|
||||
// parseDuration parses a duration string with support for human-friendly units:
|
||||
// d/day/days, w/week/weeks, mo/month/months, y/year/years, plus standard Go
|
||||
// duration units (h, m, s).
|
||||
// duration units. Following Go, m is minutes and mo is months. A bare number,
|
||||
// an unknown unit, and a negative value are all rejected.
|
||||
func parseDuration(s string) (time.Duration, error) {
|
||||
if strings.HasPrefix(strings.TrimSpace(s), "-") {
|
||||
return 0, errNegativeDuration
|
||||
}
|
||||
|
||||
d, err := time.ParseDuration(s)
|
||||
if err == nil {
|
||||
return d, nil
|
||||
|
||||
@@ -51,13 +51,32 @@ func TestParseDuration(t *testing.T) {
|
||||
want time.Duration
|
||||
err bool
|
||||
}{
|
||||
{"30d", 30 * 24 * time.Hour, false},
|
||||
{"4w", 4 * 7 * 24 * time.Hour, false},
|
||||
{"6mo", 6 * 30 * 24 * time.Hour, false},
|
||||
{"1y", 365 * 24 * time.Hour, false},
|
||||
{"2w3d", 2*7*24*time.Hour + 3*24*time.Hour, false},
|
||||
{"1h", time.Hour, false},
|
||||
// Go units, including the m-is-minutes / mo-is-months distinction
|
||||
// that this parser exists to keep straight.
|
||||
{"10ns", 10 * time.Nanosecond, false},
|
||||
{"10us", 10 * time.Microsecond, false},
|
||||
{"500ms", 500 * time.Millisecond, false},
|
||||
{"30s", 30 * time.Second, false},
|
||||
{"6m", 6 * time.Minute, false},
|
||||
{"1h", time.Hour, false},
|
||||
// Extended calendar units.
|
||||
{"30d", 30 * 24 * time.Hour, false},
|
||||
{"3days", 3 * 24 * time.Hour, false},
|
||||
{"4w", 4 * 7 * 24 * time.Hour, false},
|
||||
{"2weeks", 2 * 7 * 24 * time.Hour, false},
|
||||
{"6mo", 180 * 24 * time.Hour, false},
|
||||
{"1month", 30 * 24 * time.Hour, false},
|
||||
{"1y", 365 * 24 * time.Hour, false},
|
||||
{"2years", 2 * 365 * 24 * time.Hour, false},
|
||||
// Combined units.
|
||||
{"2w3d", 2*7*24*time.Hour + 3*24*time.Hour, false},
|
||||
{"1y6mo", 365*24*time.Hour + 180*24*time.Hour, false},
|
||||
// Rejected inputs.
|
||||
{"6", 0, true}, // bare number, no unit
|
||||
{"5x", 0, true}, // unknown unit
|
||||
{"-5d", 0, true}, // negative, extended unit
|
||||
{"-5h", 0, true}, // negative, Go unit
|
||||
{"", 0, true}, // empty
|
||||
{"garbage", 0, true},
|
||||
}
|
||||
|
||||
|
||||
@@ -167,6 +167,12 @@ func (v *Vaultik) PruneBlobs(opts *PruneOptions) error {
|
||||
|
||||
// collectReferencedBlobs downloads all manifests and returns the set of
|
||||
// referenced blob hashes.
|
||||
//
|
||||
// Every manifest must be read successfully. A manifest that cannot be
|
||||
// downloaded or decoded means its snapshot's blobs are unknown, so
|
||||
// treating them as unreferenced would let prune delete data a snapshot
|
||||
// still needs. Rather than risk that silent loss, any failure returns an
|
||||
// error naming the remote key and prune deletes nothing.
|
||||
func (v *Vaultik) collectReferencedBlobs() (map[string]bool, error) {
|
||||
log.Info("Listing remote snapshots")
|
||||
// IDs returned by listUniqueSnapshotIDs are remote keys (hashed
|
||||
@@ -179,27 +185,22 @@ func (v *Vaultik) collectReferencedBlobs() (map[string]bool, error) {
|
||||
log.Info("Found manifests in remote storage", "count", len(remoteKeys))
|
||||
|
||||
allBlobsReferenced := make(map[string]bool)
|
||||
manifestCount := 0
|
||||
|
||||
for _, remoteKey := range remoteKeys {
|
||||
log.Debug("Processing manifest", "remote_key", remoteKey)
|
||||
|
||||
manifest, err := v.downloadManifestByKey(remoteKey)
|
||||
if err != nil {
|
||||
log.Error("Failed to download manifest", "remote_key", remoteKey, "error", err)
|
||||
|
||||
continue
|
||||
return nil, fmt.Errorf("reading manifest %s: %w", remoteKey, err)
|
||||
}
|
||||
|
||||
for _, blob := range manifest.Blobs {
|
||||
allBlobsReferenced[blob.Hash] = true
|
||||
}
|
||||
|
||||
manifestCount++
|
||||
}
|
||||
|
||||
log.Info("Processed manifests",
|
||||
"count", manifestCount, "unique_blobs_referenced", len(allBlobsReferenced))
|
||||
"count", len(remoteKeys), "unique_blobs_referenced", len(allBlobsReferenced))
|
||||
|
||||
return allBlobsReferenced, nil
|
||||
}
|
||||
|
||||
@@ -0,0 +1,79 @@
|
||||
package vaultik //nolint:testpackage // exercises unexported count helpers
|
||||
|
||||
import (
|
||||
"context"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
)
|
||||
|
||||
// TestTableCountForReportSurfacesReadFailure is the regression guard for
|
||||
// the discarded-error bug: getTableCount for a table its query cannot
|
||||
// resolve must not silently become 0. A count that could not be read is
|
||||
// reported as unknown, which a reader can tell apart from an empty table.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestTableCountForReportSurfacesReadFailure(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
db, err := database.New(ctx, ":memory:")
|
||||
require.NoError(t, err)
|
||||
t.Cleanup(func() { _ = db.Close() })
|
||||
|
||||
v := &Vaultik{DB: db}
|
||||
v.SetContext(ctx)
|
||||
|
||||
// A table present in the schema reads as a real count.
|
||||
blobs := v.tableCountForReport("blobs")
|
||||
require.NotNil(t, blobs, "an existing table must read as a real count")
|
||||
assert.Equal(t, int64(0), *blobs)
|
||||
|
||||
// A syntactically valid name the sanitizer accepts but whose table
|
||||
// the query cannot resolve is the exact shape #96 describes: a
|
||||
// would-be loud failure that used to be discarded into a 0.
|
||||
_, err = v.getTableCount("snapshots_missing")
|
||||
require.Error(t, err, "a query against a nonexistent table must fail")
|
||||
|
||||
missing := v.tableCountForReport("snapshots_missing")
|
||||
assert.Nil(t, missing, "a failed read is unknown, not a count")
|
||||
|
||||
// The rendered count for a failed read must say unknown, never 0.
|
||||
assert.Equal(t, countUnknown, countText(missing))
|
||||
assert.NotEqual(t, "0", countText(missing))
|
||||
}
|
||||
|
||||
// TestCountTextDistinguishesEmptyFromUnknown pins the distinction the
|
||||
// output has to preserve: 0 means the table was empty, "unknown" means
|
||||
// the count could not be read.
|
||||
func TestCountTextDistinguishesEmptyFromUnknown(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
zero := int64(0)
|
||||
seven := int64(7)
|
||||
|
||||
assert.Equal(t, "0", countText(&zero))
|
||||
assert.Equal(t, "7", countText(&seven))
|
||||
assert.Equal(t, countUnknown, countText(nil))
|
||||
}
|
||||
|
||||
// TestCountDiffUnknownWhenEitherSideUnknown checks that a delta computed
|
||||
// from an unreadable count is itself unknown rather than a plausible
|
||||
// number.
|
||||
func TestCountDiffUnknownWhenEitherSideUnknown(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
before := int64(10)
|
||||
after := int64(3)
|
||||
|
||||
require.NotNil(t, countDiff(&before, &after))
|
||||
assert.Equal(t, int64(7), *countDiff(&before, &after))
|
||||
|
||||
assert.Nil(t, countDiff(nil, &after), "unknown before yields unknown delta")
|
||||
assert.Nil(t, countDiff(&before, nil), "unknown after yields unknown delta")
|
||||
assert.Nil(t, countDiff(nil, nil))
|
||||
}
|
||||
@@ -0,0 +1,47 @@
|
||||
package vaultik_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
|
||||
// TestPruneBlobs_UnreadableManifestDeletesNothing is the regression guard
|
||||
// for issue #157: prune identifies referenced blobs by reading every
|
||||
// snapshot's manifest, and a manifest it cannot decode used to be logged
|
||||
// and skipped. Blobs referenced only by that snapshot then looked
|
||||
// unreferenced and were deleted, with a zero exit — silent backup loss,
|
||||
// made worse by `snapshot create --prune` running unattended with force.
|
||||
//
|
||||
// The single blob here is referenced only by the snapshot whose manifest
|
||||
// is corrupt, so the old behaviour would delete it and succeed. Prune
|
||||
// must instead delete nothing and return an error.
|
||||
func TestPruneBlobs_UnreadableManifestDeletesNothing(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
env := newListEnv(t)
|
||||
ctx := context.Background()
|
||||
|
||||
blobKey := "blobs/" + testBlobHashA[:2] + "/" + testBlobHashA[2:4] +
|
||||
"/" + testBlobHashA
|
||||
require.NoError(t, env.store.Put(ctx, blobKey,
|
||||
bytes.NewReader([]byte("blob-bytes"))))
|
||||
|
||||
// A manifest at the path prune reads, but with contents it cannot
|
||||
// decode.
|
||||
require.NoError(t, env.store.Put(ctx,
|
||||
"metadata/corruptkey/manifest.json.zst",
|
||||
bytes.NewReader([]byte("not a valid manifest"))))
|
||||
|
||||
err := env.v.PruneBlobs(&vaultik.PruneOptions{Force: true})
|
||||
|
||||
require.Error(t, err, "prune must fail when a manifest cannot be read")
|
||||
assert.True(t, env.store.hasKey(blobKey),
|
||||
"no blob may be deleted when a manifest is unreadable")
|
||||
}
|
||||
@@ -0,0 +1,146 @@
|
||||
package vaultik_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"database/sql"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
"sneak.berlin/go/vaultik/internal/types"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
|
||||
// setupConsistencyTest builds a Vaultik whose local database and mock
|
||||
// remote both hold the given snapshots. Remote metadata is stored under
|
||||
// the production layout, metadata/<RemoteSnapshotKey(id)>/manifest.json.zst.
|
||||
// It returns the instance and the mock so a test can inspect the remote.
|
||||
func setupConsistencyTest(
|
||||
t *testing.T, snapshotIDs []string,
|
||||
) (*vaultik.Vaultik, *MockStorer) {
|
||||
t.Helper()
|
||||
|
||||
ctx := context.Background()
|
||||
db, err := database.New(ctx, ":memory:")
|
||||
require.NoError(t, err)
|
||||
t.Cleanup(func() { _ = db.Close() })
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
mockStorage := NewMockStorer()
|
||||
|
||||
for _, id := range snapshotIDs {
|
||||
parts := strings.Split(id, "_")
|
||||
startedAt, err := time.Parse(time.RFC3339, parts[len(parts)-1])
|
||||
require.NoError(t, err, "parsing timestamp from snapshot ID %q", id)
|
||||
|
||||
completedAt := startedAt.Add(5 * time.Minute)
|
||||
snap := &database.Snapshot{
|
||||
ID: types.SnapshotID(id),
|
||||
Hostname: testHostname,
|
||||
VaultikVersion: testLabel,
|
||||
StartedAt: startedAt,
|
||||
CompletedAt: &completedAt,
|
||||
}
|
||||
err = repos.WithTx(ctx, func(ctx context.Context, tx *sql.Tx) error {
|
||||
return repos.Snapshots.Create(ctx, tx, snap)
|
||||
})
|
||||
require.NoError(t, err, "creating snapshot %s", id)
|
||||
|
||||
metadataKey := "metadata/" + snapshot.RemoteSnapshotKey(id) +
|
||||
"/manifest.json.zst"
|
||||
err = mockStorage.Put(ctx, metadataKey, strings.NewReader("stub"))
|
||||
require.NoError(t, err)
|
||||
}
|
||||
|
||||
v := &vaultik.Vaultik{
|
||||
Storage: mockStorage,
|
||||
Repositories: repos,
|
||||
DB: db,
|
||||
Stdout: &bytes.Buffer{},
|
||||
Stderr: &bytes.Buffer{},
|
||||
Stdin: &bytes.Buffer{},
|
||||
}
|
||||
v.SetContext(ctx)
|
||||
|
||||
return v, mockStorage
|
||||
}
|
||||
|
||||
func remoteHasSnapshot(t *testing.T, m *MockStorer, id string) bool {
|
||||
t.Helper()
|
||||
|
||||
prefix := "metadata/" + snapshot.RemoteSnapshotKey(id) + "/"
|
||||
keys, err := m.List(context.Background(), prefix)
|
||||
require.NoError(t, err)
|
||||
|
||||
return len(keys) > 0
|
||||
}
|
||||
|
||||
// TestPurgeKeepsRemotelyBackedLocalRows guards against issue #160
|
||||
// (https://git.eeqj.de/sneak/vaultik/issues/160): purge reconciles local
|
||||
// rows against the remote first, and that step compared human snapshot IDs
|
||||
// against the hashed remote directory names, which never match — so it
|
||||
// deleted every local record and the purge itself then removed nothing.
|
||||
//
|
||||
// With every snapshot still present remotely and nothing old enough to
|
||||
// purge, all local rows must survive the reconcile untouched.
|
||||
func TestPurgeKeepsRemotelyBackedLocalRows(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
ids := []string{snapHomeT0, snapHomeT1, snapSystemT0}
|
||||
|
||||
v, _ := setupConsistencyTest(t, ids)
|
||||
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
// 100 years: nothing is old enough to delete, so the reconcile
|
||||
// is the only thing that touches the rows.
|
||||
OlderThan: "36500d",
|
||||
Force: true,
|
||||
})
|
||||
require.NoError(t, err)
|
||||
|
||||
remaining := listRemainingSnapshots(t, v)
|
||||
assert.Len(t, remaining, len(ids),
|
||||
"remotely-backed local rows must survive the reconcile")
|
||||
assert.Contains(t, remaining, snapHomeT0)
|
||||
assert.Contains(t, remaining, snapHomeT1)
|
||||
assert.Contains(t, remaining, snapSystemT0)
|
||||
}
|
||||
|
||||
// TestPurgeRemovesLocalAndRemoteTogether proves the two halves stay
|
||||
// consistent: a purged snapshot is gone both locally and remotely, while a
|
||||
// retained one keeps both. Before the fix, the reconcile dropped every
|
||||
// local row yet the remote metadata was left in place.
|
||||
func TestPurgeRemovesLocalAndRemoteTogether(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
ids := []string{snapHomeT0, snapHomeT1, snapSystemT0}
|
||||
|
||||
v, mock := setupConsistencyTest(t, ids)
|
||||
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
KeepLatest: true,
|
||||
Force: true,
|
||||
})
|
||||
require.NoError(t, err)
|
||||
|
||||
// Keep latest per name: newest home and the lone system are kept.
|
||||
remaining := listRemainingSnapshots(t, v)
|
||||
assert.ElementsMatch(t, []string{snapHomeT1, snapSystemT0}, remaining)
|
||||
|
||||
// Local and remote agree: the older home snapshot is gone from both,
|
||||
// the retained ones are present in both.
|
||||
assert.False(t, remoteHasSnapshot(t, mock, snapHomeT0),
|
||||
"purged snapshot must also be removed remotely")
|
||||
assert.True(t, remoteHasSnapshot(t, mock, snapHomeT1),
|
||||
"retained snapshot must remain remotely")
|
||||
assert.True(t, remoteHasSnapshot(t, mock, snapSystemT0),
|
||||
"retained snapshot must remain remotely")
|
||||
}
|
||||
@@ -12,6 +12,7 @@ import (
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
"sneak.berlin/go/vaultik/internal/types"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
@@ -60,8 +61,11 @@ func setupPurgeTest(t *testing.T, snapshotIDs []string) *vaultik.Vaultik {
|
||||
})
|
||||
require.NoError(t, err, "creating snapshot %s", id)
|
||||
|
||||
// Create remote metadata stub so syncWithRemote keeps it
|
||||
metadataKey := "metadata/" + id + "/manifest.json.zst"
|
||||
// Create the remote metadata stub under the production layout so
|
||||
// syncWithRemote keeps the local row. Production stores metadata
|
||||
// under the hashed remote key, not the human snapshot ID.
|
||||
metadataKey := "metadata/" + snapshot.RemoteSnapshotKey(id) +
|
||||
"/manifest.json.zst"
|
||||
err = mockStorage.Put(ctx, metadataKey, strings.NewReader("stub"))
|
||||
require.NoError(t, err)
|
||||
}
|
||||
|
||||
+238
-58
@@ -11,6 +11,7 @@ import (
|
||||
"math"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
"filippo.io/age"
|
||||
@@ -18,7 +19,6 @@ import (
|
||||
"sneak.berlin/go/vaultik/internal/blobgen"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
"sneak.berlin/go/vaultik/internal/types"
|
||||
)
|
||||
|
||||
@@ -35,13 +35,30 @@ var (
|
||||
errChunkNotInAnyBlob = errors.New("chunk not found in any blob")
|
||||
errBlobIDNotInHashIndex = errors.New("blob id missing from hash index")
|
||||
errShortChunkRead = errors.New("short read")
|
||||
errRestorePathEscapesTarget = errors.New(
|
||||
"refusing to restore path outside the target directory")
|
||||
errTrailingRestoreData = errors.New(
|
||||
"restored file has trailing data after its last chunk")
|
||||
errRestoreIncomplete = errors.New(
|
||||
"restore loop ended with files still pending")
|
||||
)
|
||||
|
||||
// snapshotDBFilename is the name the decrypted snapshot database is
|
||||
// written under inside its private temp directory.
|
||||
const snapshotDBFilename = "snapshot.db"
|
||||
|
||||
// restoreDirMode is the permission mode for directories created while
|
||||
// restoring (parent directories and the target root; restored
|
||||
// directories themselves get their stored mode).
|
||||
const restoreDirMode = 0o755
|
||||
|
||||
// restoreFileMode is the restrictive mode a regular file is created with
|
||||
// during restore. Content is written while the file holds this mode; the
|
||||
// stored mode is applied only after the file is fully written and closed,
|
||||
// so a file whose stored mode is restrictive is never briefly readable by
|
||||
// other local users while its content is being written.
|
||||
const restoreFileMode = 0o600
|
||||
|
||||
// sweepIntervalDivisor sets the sweeper threshold to one N-th of the
|
||||
// configured blob size limit.
|
||||
const sweepIntervalDivisor = 100
|
||||
@@ -91,7 +108,7 @@ func (v *Vaultik) Restore(opts *RestoreOptions) error {
|
||||
// Step 1: Download and decrypt the snapshot metadata database
|
||||
log.Info("Downloading snapshot metadata...")
|
||||
|
||||
tempDB, err := v.downloadSnapshotDB(opts.SnapshotID, identity)
|
||||
tempDB, tempDir, err := v.downloadSnapshotDB(opts.SnapshotID, identity)
|
||||
if err != nil {
|
||||
return fmt.Errorf("downloading snapshot database: %w", err)
|
||||
}
|
||||
@@ -101,10 +118,11 @@ func (v *Vaultik) Restore(opts *RestoreOptions) error {
|
||||
if err != nil {
|
||||
log.Debug("Failed to close temp database", "error", err)
|
||||
}
|
||||
// Clean up temp file
|
||||
err = v.Fs.Remove(tempDB.Path())
|
||||
// Remove the whole private directory, so the decrypted database
|
||||
// and any SQLite side files it produced are gone on every path.
|
||||
err = v.Fs.RemoveAll(tempDir)
|
||||
if err != nil {
|
||||
log.Debug("Failed to remove temp database", "error", err)
|
||||
log.Debug("Failed to remove temp database directory", "error", err)
|
||||
}
|
||||
}()
|
||||
|
||||
@@ -357,6 +375,13 @@ func (v *Vaultik) runRestoreLoop(
|
||||
totalBytesExpected, startTime, &lastStatusTime)
|
||||
}
|
||||
|
||||
// The loop above stops as soon as nothing is ready and nothing more
|
||||
// can be downloaded. If files still remain, they were abandoned
|
||||
// rather than restored; fail loudly instead of reporting success.
|
||||
if plan.hasPending() {
|
||||
return errRestoreIncomplete
|
||||
}
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
@@ -371,8 +396,8 @@ func (v *Vaultik) runRestoreLoop(
|
||||
func (s *restoreSession) downloadNextBlobSet(plan *restorePlan) (bool, error) {
|
||||
s.sweeper.sweep()
|
||||
|
||||
next := plan.pickNextDownload()
|
||||
if next.IsZero() {
|
||||
next, ok := plan.pickNextDownload()
|
||||
if !ok {
|
||||
return false, nil
|
||||
}
|
||||
|
||||
@@ -577,18 +602,24 @@ func (v *Vaultik) handleRestoreVerification(
|
||||
}
|
||||
|
||||
// downloadSnapshotDB downloads and decrypts the snapshot metadata
|
||||
// database. The snapshotID is the human ID; we hash it to the remote
|
||||
// key for the storage path.
|
||||
// database. The identifier is resolved to the snapshot's remote key: a
|
||||
// human ID is hashed, and a remote key (or its abbreviation, as printed
|
||||
// for a remote-only snapshot) is used as-is, so a host with no local
|
||||
// index can restore the snapshots it can only see on the store.
|
||||
func (v *Vaultik) downloadSnapshotDB(
|
||||
snapshotID string, identity age.Identity,
|
||||
) (*database.DB, error) {
|
||||
) (*database.DB, string, error) {
|
||||
remoteKey, err := v.resolveSnapshotRemoteKey(snapshotID)
|
||||
if err != nil {
|
||||
return nil, "", err
|
||||
}
|
||||
|
||||
// Download encrypted database from storage
|
||||
dbKey := fmt.Sprintf("metadata/%s/db.zst.age",
|
||||
snapshot.RemoteSnapshotKey(snapshotID))
|
||||
dbKey := fmt.Sprintf("metadata/%s/db.zst.age", remoteKey)
|
||||
|
||||
reader, err := v.Storage.Get(v.ctx, dbKey)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("downloading %s: %w", dbKey, err)
|
||||
return nil, "", fmt.Errorf("downloading %s: %w", dbKey, err)
|
||||
}
|
||||
|
||||
defer func() { _ = reader.Close() }()
|
||||
@@ -596,7 +627,7 @@ func (v *Vaultik) downloadSnapshotDB(
|
||||
// Read all data
|
||||
encryptedData, err := io.ReadAll(reader)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("reading encrypted data: %w", err)
|
||||
return nil, "", fmt.Errorf("reading encrypted data: %w", err)
|
||||
}
|
||||
|
||||
log.Debug("Downloaded encrypted database",
|
||||
@@ -605,7 +636,7 @@ func (v *Vaultik) downloadSnapshotDB(
|
||||
// Decrypt and decompress using blobgen.Reader
|
||||
blobReader, err := blobgen.NewReader(bytes.NewReader(encryptedData), identity)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("creating decryption reader: %w", err)
|
||||
return nil, "", fmt.Errorf("creating decryption reader: %w", err)
|
||||
}
|
||||
|
||||
defer func() { _ = blobReader.Close() }()
|
||||
@@ -613,44 +644,52 @@ func (v *Vaultik) downloadSnapshotDB(
|
||||
// Read the binary SQLite database
|
||||
dbData, err := io.ReadAll(blobReader)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("decrypting and decompressing: %w", err)
|
||||
return nil, "", fmt.Errorf("decrypting and decompressing: %w", err)
|
||||
}
|
||||
|
||||
log.Debug("Decrypted database", "size", ubytes(int64(len(dbData))))
|
||||
|
||||
// Create a temporary database file and write the binary SQLite data directly
|
||||
tempFile, err := afero.TempFile(v.Fs, "", "vaultik-restore-*.db")
|
||||
return v.materializeSnapshotDB(dbData)
|
||||
}
|
||||
|
||||
// materializeSnapshotDB writes the decrypted snapshot database bytes into
|
||||
// a fresh private (0700) temp directory and opens the file read-only. On
|
||||
// any failure it removes the directory before returning, so no decrypted
|
||||
// metadata is left on disk when the open is interrupted or the payload is
|
||||
// damaged. On success the returned directory is the caller's to remove.
|
||||
func (v *Vaultik) materializeSnapshotDB(
|
||||
dbData []byte,
|
||||
) (*database.DB, string, error) {
|
||||
tempDir, err := afero.TempDir(v.Fs, "", "vaultik-restore-")
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("creating temp file: %w", err)
|
||||
return nil, "", fmt.Errorf("creating temp directory: %w", err)
|
||||
}
|
||||
|
||||
tempPath := tempFile.Name()
|
||||
success := false
|
||||
|
||||
// Write the binary SQLite database directly
|
||||
_, err = tempFile.Write(dbData)
|
||||
defer func() {
|
||||
if !success {
|
||||
_ = v.Fs.RemoveAll(tempDir)
|
||||
}
|
||||
}()
|
||||
|
||||
dbPath := filepath.Join(tempDir, snapshotDBFilename)
|
||||
|
||||
err = afero.WriteFile(v.Fs, dbPath, dbData, restoreFileMode)
|
||||
if err != nil {
|
||||
_ = tempFile.Close()
|
||||
_ = v.Fs.Remove(tempPath)
|
||||
|
||||
return nil, fmt.Errorf("writing database file: %w", err)
|
||||
return nil, "", fmt.Errorf("writing database file: %w", err)
|
||||
}
|
||||
|
||||
err = tempFile.Close()
|
||||
if err != nil {
|
||||
_ = v.Fs.Remove(tempPath)
|
||||
log.Debug("Created restore database", "path", dbPath)
|
||||
|
||||
return nil, fmt.Errorf("closing temp file: %w", err)
|
||||
db, err := database.OpenReadOnly(v.ctx, dbPath)
|
||||
if err != nil {
|
||||
return nil, "", fmt.Errorf("opening restore database: %w", err)
|
||||
}
|
||||
|
||||
log.Debug("Created restore database", "path", tempPath)
|
||||
success = true
|
||||
|
||||
// Open the database
|
||||
db, err := database.New(v.ctx, tempPath)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("opening restore database: %w", err)
|
||||
}
|
||||
|
||||
return db, nil
|
||||
return db, tempDir, nil
|
||||
}
|
||||
|
||||
// getFilesToRestore returns the list of files to restore based on path filters
|
||||
@@ -755,13 +794,85 @@ type restoreSession struct {
|
||||
runningAsRoot bool
|
||||
}
|
||||
|
||||
// containedRestorePath resolves rel — a path read from the snapshot
|
||||
// database — to its location under targetDir and confirms the write will
|
||||
// stay inside the target.
|
||||
//
|
||||
// age decryption proves a snapshot is readable, not that it is honest, so
|
||||
// every stored path is treated as hostile. rel is rejected unless
|
||||
// filepath.IsLocal accepts it once the leading separator is stripped:
|
||||
// stored paths are absolute and the join to targetDir drops that
|
||||
// separator, so "/etc/passwd" is judged as the relative "etc/passwd" it
|
||||
// becomes on disk. This bars "..", absolute, and empty paths.
|
||||
//
|
||||
// A stored symlink whose target points outside the tree is still honest
|
||||
// (and restored verbatim), but a later entry must not be written through
|
||||
// it. Each existing ancestor directory below the target is therefore
|
||||
// Lstat'ed and a symlink among them is refused. The leaf itself is not
|
||||
// traversed: honest snapshots restore symlinks at leaf positions, and the
|
||||
// unique-path constraint keeps a leaf from being both a symlink and a
|
||||
// regular file. The target directory itself may be a symlink; only
|
||||
// components below it are checked.
|
||||
func containedRestorePath(fs afero.Fs, targetDir, rel string) (string, error) {
|
||||
local := strings.TrimPrefix(rel, string(filepath.Separator))
|
||||
if !filepath.IsLocal(local) {
|
||||
return "", fmt.Errorf("%w: %s", errRestorePathEscapesTarget, rel)
|
||||
}
|
||||
|
||||
local = filepath.Clean(local)
|
||||
targetPath := filepath.Join(targetDir, local)
|
||||
|
||||
relDir := filepath.Dir(local)
|
||||
if relDir == "." {
|
||||
return targetPath, nil
|
||||
}
|
||||
|
||||
current := targetDir
|
||||
for component := range strings.SplitSeq(relDir, string(filepath.Separator)) {
|
||||
current = filepath.Join(current, component)
|
||||
|
||||
info, err := lstatIfPossible(fs, current)
|
||||
if err != nil {
|
||||
if os.IsNotExist(err) {
|
||||
continue
|
||||
}
|
||||
|
||||
return "", fmt.Errorf("checking restore path %s: %w", current, err)
|
||||
}
|
||||
|
||||
if info.Mode()&os.ModeSymlink != 0 {
|
||||
return "", fmt.Errorf("%w: %s descends through symlink %s",
|
||||
errRestorePathEscapesTarget, rel, current)
|
||||
}
|
||||
}
|
||||
|
||||
return targetPath, nil
|
||||
}
|
||||
|
||||
// lstatIfPossible performs a symlink-aware stat when the filesystem
|
||||
// supports it. afero.OsFs does; MemMapFs, which has no symlinks, reports
|
||||
// that Lstat was not used and its result never carries ModeSymlink.
|
||||
func lstatIfPossible(fs afero.Fs, name string) (os.FileInfo, error) {
|
||||
if lstater, ok := fs.(afero.Lstater); ok {
|
||||
info, _, err := lstater.LstatIfPossible(name)
|
||||
|
||||
return info, err
|
||||
}
|
||||
|
||||
return fs.Stat(name)
|
||||
}
|
||||
|
||||
// restoreFile dispatches to the right per-kind restorer.
|
||||
func (s *restoreSession) restoreFile(file *database.File) error {
|
||||
targetPath := filepath.Join(s.opts.TargetDir, file.Path.String())
|
||||
targetPath, err := containedRestorePath(
|
||||
s.v.Fs, s.opts.TargetDir, file.Path.String())
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
parentDir := filepath.Dir(targetPath)
|
||||
|
||||
err := s.v.Fs.MkdirAll(parentDir, restoreDirMode)
|
||||
err = s.v.Fs.MkdirAll(parentDir, restoreDirMode)
|
||||
if err != nil {
|
||||
return fmt.Errorf("creating parent directory: %w", err)
|
||||
}
|
||||
@@ -809,6 +920,13 @@ func (s *restoreSession) restoreDirectory(
|
||||
return fmt.Errorf("creating directory: %w", err)
|
||||
}
|
||||
|
||||
// MkdirAll applies the process umask, so chmod to the exact stored
|
||||
// mode. A failure here is non-fatal.
|
||||
err = s.v.Fs.Chmod(targetPath, os.FileMode(file.Mode))
|
||||
if err != nil {
|
||||
log.Debug("Failed to set permissions", "path", targetPath, "error", err)
|
||||
}
|
||||
|
||||
s.applyFileMetadata(file, targetPath)
|
||||
|
||||
s.result.FilesRestored++
|
||||
@@ -816,25 +934,22 @@ func (s *restoreSession) restoreDirectory(
|
||||
return nil
|
||||
}
|
||||
|
||||
// applyFileMetadata applies stored permissions, ownership (when running
|
||||
// as root on a real filesystem), and mtime to a restored path. Failures
|
||||
// are logged at debug level and do not abort the restore.
|
||||
// applyFileMetadata applies ownership (when running as root on a real
|
||||
// filesystem) and mtime to a restored path. Permission mode is applied
|
||||
// separately by each caller, with different failure handling, so it is
|
||||
// not touched here. Failures are logged at debug level and do not abort
|
||||
// the restore.
|
||||
func (s *restoreSession) applyFileMetadata(file *database.File, targetPath string) {
|
||||
err := s.v.Fs.Chmod(targetPath, os.FileMode(file.Mode))
|
||||
if err != nil {
|
||||
log.Debug("Failed to set permissions", "path", targetPath, "error", err)
|
||||
}
|
||||
|
||||
if s.runningAsRoot {
|
||||
if _, ok := s.v.Fs.(*afero.OsFs); ok {
|
||||
err = os.Chown(targetPath, int(file.UID), int(file.GID))
|
||||
err := os.Chown(targetPath, int(file.UID), int(file.GID))
|
||||
if err != nil {
|
||||
log.Debug("Failed to set ownership", "path", targetPath, "error", err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
err = s.v.Fs.Chtimes(targetPath, file.MTime, file.MTime)
|
||||
err := s.v.Fs.Chtimes(targetPath, file.MTime, file.MTime)
|
||||
if err != nil {
|
||||
log.Debug("Failed to set mtime", "path", targetPath, "error", err)
|
||||
}
|
||||
@@ -868,17 +983,30 @@ func (s *restoreSession) restoreRegularFile(
|
||||
|
||||
t0 = time.Now()
|
||||
|
||||
outFile, err := s.v.Fs.Create(targetPath)
|
||||
// Remove any existing entry, then create the file with a restrictive
|
||||
// mode via O_EXCL. The stored mode is applied only after the content
|
||||
// is written and the file closed, so a file whose stored mode is
|
||||
// restrictive is never briefly readable by other local users while
|
||||
// its content is written. Removing first (rather than failing on a
|
||||
// leftover file) matches the documented behaviour that re-running
|
||||
// restore overwrites partial output.
|
||||
_ = s.v.Fs.Remove(targetPath)
|
||||
|
||||
outFile, err := s.v.Fs.OpenFile(
|
||||
targetPath, os.O_CREATE|os.O_EXCL|os.O_WRONLY, restoreFileMode)
|
||||
createDur := time.Since(t0)
|
||||
|
||||
if err != nil {
|
||||
return fmt.Errorf("creating output file: %w", err)
|
||||
}
|
||||
|
||||
defer func() { _ = outFile.Close() }()
|
||||
|
||||
bytesWritten, timings, err := s.writeFileChunks(outFile, fileChunks)
|
||||
if err != nil {
|
||||
// Do not leave a partial file behind.
|
||||
_ = outFile.Close()
|
||||
|
||||
s.removePartialRestore(targetPath)
|
||||
|
||||
return err
|
||||
}
|
||||
|
||||
@@ -896,9 +1024,12 @@ func (s *restoreSession) restoreRegularFile(
|
||||
|
||||
err = outFile.Close()
|
||||
if err != nil {
|
||||
s.removePartialRestore(targetPath)
|
||||
|
||||
return fmt.Errorf("closing output file: %w", err)
|
||||
}
|
||||
|
||||
s.applyRestoredFileMode(file, targetPath)
|
||||
s.applyFileMetadata(file, targetPath)
|
||||
|
||||
s.result.FilesRestored++
|
||||
@@ -909,6 +1040,31 @@ func (s *restoreSession) restoreRegularFile(
|
||||
return nil
|
||||
}
|
||||
|
||||
// applyRestoredFileMode applies the stored permission bits to a
|
||||
// just-written regular file (created with restoreFileMode). A failure is
|
||||
// a user-visible warning, not a fatal error: the file's content is
|
||||
// intact and it remains at the restrictive create-time mode, so the
|
||||
// restore is not aborted or discarded over it.
|
||||
func (s *restoreSession) applyRestoredFileMode(
|
||||
file *database.File, targetPath string,
|
||||
) {
|
||||
err := s.v.Fs.Chmod(targetPath, os.FileMode(file.Mode))
|
||||
if err != nil {
|
||||
s.v.UI.Warningf("Failed to set mode %s on %s: %v",
|
||||
os.FileMode(file.Mode).Perm(), s.v.UI.Path(targetPath), err)
|
||||
}
|
||||
}
|
||||
|
||||
// removePartialRestore deletes a restore output file whose write did not
|
||||
// complete, so a failed restore never leaves a partial file behind.
|
||||
func (s *restoreSession) removePartialRestore(targetPath string) {
|
||||
err := s.v.Fs.Remove(targetPath)
|
||||
if err != nil {
|
||||
log.Debug("Failed to remove partial restore file",
|
||||
"path", targetPath, "error", err)
|
||||
}
|
||||
}
|
||||
|
||||
// writeFileChunks streams each of the file's chunks from the blob disk
|
||||
// cache into outFile, crediting restored bytes to the sweeper as it
|
||||
// goes. Returns the bytes written plus per-phase timing accumulators.
|
||||
@@ -990,11 +1146,19 @@ func (s *restoreSession) downloadBlobToCache(
|
||||
streamDur := time.Since(t0)
|
||||
closeErr := rc.Close()
|
||||
|
||||
// closeErr carries the blob's hash-verification result (a mismatch,
|
||||
// or the stream not being fully read). On any failure, drop the
|
||||
// cache entry so a blob that failed verification is never read back
|
||||
// as if it were valid.
|
||||
if copyErr != nil {
|
||||
s.blobCache.Delete(blobHash)
|
||||
|
||||
return copyErr
|
||||
}
|
||||
|
||||
if closeErr != nil {
|
||||
s.blobCache.Delete(blobHash)
|
||||
|
||||
return closeErr
|
||||
}
|
||||
|
||||
@@ -1057,17 +1221,22 @@ func (v *Vaultik) verifyRestoredFiles(
|
||||
return ctx.Err()
|
||||
}
|
||||
|
||||
targetPath := filepath.Join(targetDir, file.Path.String())
|
||||
targetPath, err := containedRestorePath(v.Fs, targetDir, file.Path.String())
|
||||
if err == nil {
|
||||
var bytesVerified int64
|
||||
|
||||
bytesVerified, err = v.verifyFile(ctx, repos, file, targetPath)
|
||||
if err == nil {
|
||||
result.FilesVerified++
|
||||
result.BytesVerified += bytesVerified
|
||||
}
|
||||
}
|
||||
|
||||
bytesVerified, err := v.verifyFile(ctx, repos, file, targetPath)
|
||||
if err != nil {
|
||||
log.Error("File verification failed", "path", file.Path, "error", err)
|
||||
|
||||
result.FilesFailed++
|
||||
result.FailedFiles = append(result.FailedFiles, file.Path.String())
|
||||
} else {
|
||||
result.FilesVerified++
|
||||
result.BytesVerified += bytesVerified
|
||||
}
|
||||
|
||||
bytesProcessed += file.Size
|
||||
@@ -1157,6 +1326,17 @@ func (v *Vaultik) verifyFile(
|
||||
bytesVerified += int64(n)
|
||||
}
|
||||
|
||||
// The stored chunks account for the whole file, so the reader must
|
||||
// be at EOF now. Trailing bytes past the last chunk are corruption
|
||||
// the per-chunk loop cannot see.
|
||||
extra := make([]byte, 1)
|
||||
|
||||
n, err := f.Read(extra)
|
||||
if n != 0 || !errors.Is(err, io.EOF) {
|
||||
return bytesVerified, fmt.Errorf("%w: file longer than its %d chunk(s)",
|
||||
errTrailingRestoreData, len(fileChunks))
|
||||
}
|
||||
|
||||
log.Debug("File verified",
|
||||
"path", file.Path, "bytes", bytesVerified, "chunks", len(fileChunks))
|
||||
|
||||
|
||||
@@ -0,0 +1,167 @@
|
||||
package vaultik_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"io"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/config"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
"sneak.berlin/go/vaultik/internal/ui"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
|
||||
// TestRestoreOnAnotherMachine proves the disaster-recovery path: a host
|
||||
// that has only the vaultik binary, the age secret key, and the storage
|
||||
// credentials — no local index, a different hostname, and no
|
||||
// age_recipients configured — can list, restore, and verify a snapshot
|
||||
// straight from the destination store.
|
||||
//
|
||||
// The backup half writes a snapshot with one index and hostname. The
|
||||
// restore half throws that index away entirely: a fresh, empty index and
|
||||
// a config that shares nothing with the original but the storage location
|
||||
// and the secret key. If restore or verify needed the original local
|
||||
// index — or the human snapshot ID that only that index holds — this test
|
||||
// could not run, because the recovery host can know neither.
|
||||
func TestRestoreOnAnotherMachine(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
|
||||
dataDir := filepath.Join(tempDir, "source")
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
restoreDir := filepath.Join(tempDir, "restored")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
|
||||
chunkSize := int64(64 * 1024)
|
||||
maxBlobSize := int64(512 * 1024)
|
||||
|
||||
sourceFiles := writeRecoverySourceTree(t, fs, dataDir, chunkSize)
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
// Backup host: one index, hostname test-host, age_recipients set.
|
||||
// runFileStorageBackup closes the index before returning, so nothing
|
||||
// below can lean on it.
|
||||
_, storer, originalID := runFileStorageBackup(
|
||||
ctx, t, fs, dataDir, storeDir, dbPath, chunkSize, maxBlobSize)
|
||||
|
||||
// Recovery host: a fresh empty index, a different hostname, and no
|
||||
// age_recipients — only the secret key and the same storage location.
|
||||
recovery, stdout := newRecoveryHost(ctx, t, fs, storer)
|
||||
|
||||
// The recovery index really is empty. This is the assertion that makes
|
||||
// the test a guard against restore quietly depending on the original
|
||||
// index: if it did, an empty index would make restore fail.
|
||||
localSnaps, err := recovery.Repositories.Snapshots.ListRecent(ctx, 100)
|
||||
require.NoError(t, err)
|
||||
require.Empty(t, localSnaps, "recovery host must start with no local index")
|
||||
|
||||
// List: the snapshot shows up as remote-only, identified by its remote
|
||||
// key, with no recoverable human ID.
|
||||
require.NoError(t, recovery.ListSnapshots(true))
|
||||
|
||||
rows := decodeListJSON(t, stdout.String())
|
||||
require.Len(t, rows, 1)
|
||||
|
||||
remote := rows[0]
|
||||
assert.False(t, remote.LocallyTracked, "snapshot must be remote-only here")
|
||||
assert.Empty(t, remote.ID, "the human ID is unknown to the recovery host")
|
||||
require.Len(t, remote.RemoteKey, 64)
|
||||
assert.Equal(t, snapshot.RemoteSnapshotKey(originalID), remote.RemoteKey,
|
||||
"the listed key is the hashed snapshot ID")
|
||||
|
||||
// Restore driven by the abbreviated identifier the table prints (the
|
||||
// first 12 hex of the remote key), then deep-verify from the store
|
||||
// keyed by the full remote key. Both are what a recovery host can know.
|
||||
require.NoError(t, recovery.Restore(&vaultik.RestoreOptions{
|
||||
SnapshotID: remote.RemoteKey[:12],
|
||||
TargetDir: restoreDir,
|
||||
Verify: true,
|
||||
}))
|
||||
require.NoError(t, recovery.RunDeepVerify(
|
||||
remote.RemoteKey, &vaultik.VerifyOptions{Deep: true}))
|
||||
|
||||
assertRestoredTreeMatches(t, fs, restoreDir, sourceFiles)
|
||||
}
|
||||
|
||||
// writeRecoverySourceTree writes a small source tree spanning several
|
||||
// chunks (so restore reassembles real multi-chunk files) and returns the
|
||||
// content keyed by absolute path.
|
||||
func writeRecoverySourceTree(
|
||||
t *testing.T, fs afero.Fs, dataDir string, chunkSize int64,
|
||||
) map[string][]byte {
|
||||
t.Helper()
|
||||
|
||||
sourceFiles := map[string][]byte{
|
||||
filepath.Join(dataDir, "notes.txt"): []byte("recover me"),
|
||||
filepath.Join(dataDir, "sub", "big.bin"): bytesPattern("big-", int(chunkSize*3)),
|
||||
filepath.Join(dataDir, "sub", "small.bin"): bytesPattern("small-", 128),
|
||||
}
|
||||
|
||||
for path, content := range sourceFiles {
|
||||
require.NoError(t, fs.MkdirAll(filepath.Dir(path), 0o755))
|
||||
require.NoError(t, afero.WriteFile(fs, path, content, 0o644))
|
||||
}
|
||||
|
||||
return sourceFiles
|
||||
}
|
||||
|
||||
// newRecoveryHost builds the Vaultik a replacement machine would run: an
|
||||
// empty in-memory index, a hostname different from the backup host, no
|
||||
// age_recipients, and only the secret key plus the shared storer. It
|
||||
// returns the instance and the buffer its stdout is wired to.
|
||||
func newRecoveryHost(
|
||||
ctx context.Context, t *testing.T, fs afero.Fs, storer storage.Storer,
|
||||
) (*vaultik.Vaultik, *bytes.Buffer) {
|
||||
t.Helper()
|
||||
|
||||
recoveryDB, err := database.New(ctx, ":memory:")
|
||||
require.NoError(t, err)
|
||||
t.Cleanup(func() { _ = recoveryDB.Close() })
|
||||
|
||||
stdout := &bytes.Buffer{}
|
||||
|
||||
recovery := &vaultik.Vaultik{
|
||||
Config: &config.Config{
|
||||
AgeSecretKey: testAgeSecretKey,
|
||||
Hostname: "recovery-host",
|
||||
},
|
||||
Storage: storer,
|
||||
Fs: fs,
|
||||
Repositories: database.NewRepositories(recoveryDB),
|
||||
DB: recoveryDB,
|
||||
Stdout: stdout,
|
||||
Stderr: io.Discard,
|
||||
UI: ui.NewWithColor(io.Discard, false),
|
||||
}
|
||||
recovery.SetContext(ctx)
|
||||
|
||||
return recovery, stdout
|
||||
}
|
||||
|
||||
// assertRestoredTreeMatches byte-compares every restored file against its
|
||||
// source content.
|
||||
func assertRestoredTreeMatches(
|
||||
t *testing.T, fs afero.Fs, restoreDir string, sourceFiles map[string][]byte,
|
||||
) {
|
||||
t.Helper()
|
||||
|
||||
for origPath, expected := range sourceFiles {
|
||||
restored := filepath.Join(restoreDir, origPath)
|
||||
got, err := afero.ReadFile(fs, restored)
|
||||
require.NoErrorf(t, err, "restored file missing: %s", restored)
|
||||
require.Truef(t, bytes.Equal(got, expected),
|
||||
"byte mismatch for %s", origPath)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,175 @@
|
||||
package vaultik //nolint:testpackage // drives unexported restore internals
|
||||
|
||||
import (
|
||||
"context"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/config"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/types"
|
||||
"sneak.berlin/go/vaultik/internal/ui"
|
||||
)
|
||||
|
||||
// These tests exercise the path-containment guard that keeps restore from
|
||||
// writing outside its target directory. age decryption proves only that a
|
||||
// snapshot is readable, not that its recorded paths are honest, so restore
|
||||
// treats every stored path as hostile: a compromised backed-up host could
|
||||
// forge a snapshot that decrypts cleanly, and restore usually runs as root.
|
||||
//
|
||||
// They drive restoreAllFiles directly (rather than the full Restore, which
|
||||
// downloads and decrypts the metadata database from storage) so a snapshot
|
||||
// database with adversarial rows can be handed to the restore loop without
|
||||
// the surrounding blob/storage machinery. Directory and symlink entries
|
||||
// carry no chunks, so no blobs are needed.
|
||||
|
||||
// containmentDirMode marks a File row as a directory for the restore loop.
|
||||
const containmentDirMode = uint32(os.ModeDir | 0o755)
|
||||
|
||||
// newContainmentVaultik builds the minimal Vaultik needed to run
|
||||
// restoreAllFiles against fs.
|
||||
func newContainmentVaultik(ctx context.Context, fs afero.Fs) *Vaultik {
|
||||
v := &Vaultik{
|
||||
Config: &config.Config{
|
||||
BlobSizeLimit: config.Size(10 * 1024 * 1024),
|
||||
},
|
||||
Fs: fs,
|
||||
Stdout: io.Discard,
|
||||
Stderr: io.Discard,
|
||||
UI: ui.NewWithColor(io.Discard, false),
|
||||
}
|
||||
v.SetContext(ctx)
|
||||
|
||||
return v
|
||||
}
|
||||
|
||||
// makeFiles inserts the given rows into a fresh in-memory snapshot database
|
||||
// and returns them (with IDs assigned) plus the repositories.
|
||||
func makeFiles(
|
||||
ctx context.Context, t *testing.T, rows []*database.File,
|
||||
) ([]*database.File, *database.Repositories) {
|
||||
t.Helper()
|
||||
|
||||
db, err := database.New(ctx, filepath.Join(t.TempDir(), "index.sqlite"))
|
||||
require.NoError(t, err)
|
||||
t.Cleanup(func() { _ = db.Close() })
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
for _, f := range rows {
|
||||
require.NoError(t, repos.Files.Create(ctx, nil, f))
|
||||
}
|
||||
|
||||
return rows, repos
|
||||
}
|
||||
|
||||
func TestRestoreRejectsPathTraversal(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
tests := []struct {
|
||||
name string
|
||||
// rows are inserted in order; the escape entry is restored after
|
||||
// any entry it depends on (the symlink case needs its link first).
|
||||
rows func(outsideDir string) []*database.File
|
||||
// escaped is the path, outside the target, that must not appear.
|
||||
escaped func(tempDir, outsideDir string) string
|
||||
}{
|
||||
{
|
||||
name: "relative dotdot",
|
||||
rows: func(_ string) []*database.File {
|
||||
return []*database.File{{
|
||||
Path: "../escaped-relative",
|
||||
Mode: containmentDirMode,
|
||||
}}
|
||||
},
|
||||
escaped: func(tempDir, _ string) string {
|
||||
return filepath.Join(tempDir, "escaped-relative")
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "absolute with dotdot",
|
||||
rows: func(_ string) []*database.File {
|
||||
return []*database.File{{
|
||||
Path: "/a/../../escaped-absolute",
|
||||
Mode: containmentDirMode,
|
||||
}}
|
||||
},
|
||||
escaped: func(tempDir, _ string) string {
|
||||
return filepath.Join(tempDir, "escaped-absolute")
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "child through symlink",
|
||||
rows: func(outsideDir string) []*database.File {
|
||||
return []*database.File{
|
||||
// Restored first: an in-target symlink pointing out.
|
||||
{Path: "linkdir", LinkTarget: types.FilePath(outsideDir)},
|
||||
// Restored second: a child written through that link.
|
||||
{Path: "linkdir/child", Mode: containmentDirMode},
|
||||
}
|
||||
},
|
||||
escaped: func(_, outsideDir string) string {
|
||||
return filepath.Join(outsideDir, "child")
|
||||
},
|
||||
},
|
||||
}
|
||||
|
||||
for _, tc := range tests {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
targetDir := filepath.Join(tempDir, "target")
|
||||
outsideDir := filepath.Join(tempDir, "outside")
|
||||
require.NoError(t, fs.MkdirAll(outsideDir, 0o755))
|
||||
|
||||
rows, repos := makeFiles(ctx, t, tc.rows(outsideDir))
|
||||
v := newContainmentVaultik(ctx, fs)
|
||||
|
||||
_, err := v.restoreAllFiles(rows, repos,
|
||||
&RestoreOptions{TargetDir: targetDir}, nil, nil)
|
||||
|
||||
require.ErrorIs(t, err, errRestorePathEscapesTarget)
|
||||
|
||||
escaped := tc.escaped(tempDir, outsideDir)
|
||||
_, statErr := os.Lstat(escaped)
|
||||
require.Truef(t, os.IsNotExist(statErr),
|
||||
"restore wrote outside the target at %s", escaped)
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestRestoreAllowsSymlinkPointingOutsideTree confirms the guard does not
|
||||
// over-block: an honest snapshot may contain a symlink whose target lies
|
||||
// outside the restored tree, and it must still be restored verbatim.
|
||||
func TestRestoreAllowsSymlinkPointingOutsideTree(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
targetDir := filepath.Join(tempDir, "target")
|
||||
linkTarget := filepath.Join(tempDir, "outside", "data")
|
||||
|
||||
rows, repos := makeFiles(ctx, t, []*database.File{
|
||||
{Path: "goodlink", LinkTarget: types.FilePath(linkTarget), MTime: time.Unix(0, 0)},
|
||||
})
|
||||
v := newContainmentVaultik(ctx, fs)
|
||||
|
||||
_, err := v.restoreAllFiles(rows, repos,
|
||||
&RestoreOptions{TargetDir: targetDir}, nil, nil)
|
||||
require.NoError(t, err)
|
||||
|
||||
got, err := os.Readlink(filepath.Join(targetDir, "goodlink"))
|
||||
require.NoError(t, err)
|
||||
require.Equal(t, linkTarget, got)
|
||||
}
|
||||
@@ -0,0 +1,306 @@
|
||||
package vaultik //nolint:testpackage // drives restore through unexported session
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"errors"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"sync"
|
||||
"syscall"
|
||||
"testing"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/config"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
"sneak.berlin/go/vaultik/internal/ui"
|
||||
)
|
||||
|
||||
// errSpyWrite is the injected write failure used to exercise the
|
||||
// partial-file cleanup path.
|
||||
var errSpyWrite = errors.New("injected write failure")
|
||||
|
||||
// modeSpyFs wraps a real filesystem so restore tests can observe and
|
||||
// perturb the single output file whose path contains watch. It records
|
||||
// the on-disk permission bits seen at the moment content is first
|
||||
// written (the window during which another local user could read it),
|
||||
// and can inject a write failure or append trailing bytes on close.
|
||||
type modeSpyFs struct {
|
||||
afero.Fs
|
||||
|
||||
watch string
|
||||
|
||||
mu sync.Mutex
|
||||
writeModes []os.FileMode
|
||||
failWrite bool
|
||||
trailing int
|
||||
}
|
||||
|
||||
//nolint:ireturn // afero.Fs.OpenFile is defined to return the interface
|
||||
func (m *modeSpyFs) OpenFile(
|
||||
name string, flag int, perm os.FileMode,
|
||||
) (afero.File, error) {
|
||||
f, err := m.Fs.OpenFile(name, flag, perm)
|
||||
if err != nil || !strings.Contains(name, m.watch) {
|
||||
return f, err
|
||||
}
|
||||
|
||||
return &modeSpyFile{File: f, fs: m, path: name}, nil
|
||||
}
|
||||
|
||||
type modeSpyFile struct {
|
||||
afero.File
|
||||
|
||||
fs *modeSpyFs
|
||||
path string
|
||||
written bool
|
||||
}
|
||||
|
||||
func (f *modeSpyFile) Write(p []byte) (int, error) {
|
||||
if !f.written {
|
||||
f.written = true
|
||||
|
||||
info, err := f.fs.Stat(f.path)
|
||||
if err == nil {
|
||||
f.fs.mu.Lock()
|
||||
f.fs.writeModes = append(f.fs.writeModes, info.Mode().Perm())
|
||||
f.fs.mu.Unlock()
|
||||
}
|
||||
}
|
||||
|
||||
if f.fs.failWrite {
|
||||
return 0, errSpyWrite
|
||||
}
|
||||
|
||||
return f.File.Write(p)
|
||||
}
|
||||
|
||||
func (f *modeSpyFile) Close() error {
|
||||
if f.fs.trailing > 0 {
|
||||
_, _ = f.File.Write(bytes.Repeat([]byte{'x'}, f.fs.trailing))
|
||||
}
|
||||
|
||||
return f.File.Close()
|
||||
}
|
||||
|
||||
// backupOneFile writes a single source file with the given mode and
|
||||
// backs it up into a fresh file storer, returning everything a restore
|
||||
// needs. The index database is closed before returning so the restore
|
||||
// half runs from the exported metadata and remote bytes only.
|
||||
func backupOneFile(
|
||||
ctx context.Context, t *testing.T, fs afero.Fs, tempDir, name string,
|
||||
content []byte, mode os.FileMode,
|
||||
) (*config.Config, *storage.FileStorer, string, string) {
|
||||
t.Helper()
|
||||
|
||||
dataDir := filepath.Join(tempDir, "src")
|
||||
require.NoError(t, fs.MkdirAll(dataDir, 0o755))
|
||||
|
||||
srcPath := filepath.Join(dataDir, name)
|
||||
require.NoError(t, afero.WriteFile(fs, srcPath, content, mode))
|
||||
require.NoError(t, fs.Chmod(srcPath, mode))
|
||||
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
|
||||
storer, err := storage.NewFileStorer(storeDir)
|
||||
require.NoError(t, err)
|
||||
|
||||
cfg := &config.Config{
|
||||
AgeRecipients: []string{
|
||||
"age1ezrjmfpwsc95svdg0y54mums3zevgzu0x0ecq2f7tp8a05gl0sjq9q9wjg",
|
||||
},
|
||||
AgeSecretKey: "AGE-SECRET-KEY-19CR5YSFW59HM4TLD6GXVEDMZFTVVF7PPHKU" +
|
||||
"T68TXSFPK7APHXA2QS2NJA5",
|
||||
CompressionLevel: 3,
|
||||
Hostname: "test-host",
|
||||
BlobSizeLimit: config.Size(5 * 1024 * 1024),
|
||||
}
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
sm := snapshot.NewSnapshotManager(snapshot.SnapshotManagerParams{
|
||||
Repos: repos,
|
||||
Storage: storer,
|
||||
Config: cfg,
|
||||
})
|
||||
sm.SetFilesystem(fs)
|
||||
|
||||
scanner := snapshot.NewScanner(snapshot.ScannerConfig{
|
||||
FS: fs,
|
||||
Storage: storer,
|
||||
ChunkSize: 4 * 1024 * 1024,
|
||||
MaxBlobSize: 5 * 1024 * 1024,
|
||||
CompressionLevel: cfg.CompressionLevel,
|
||||
AgeRecipients: cfg.AgeRecipients,
|
||||
Repositories: repos,
|
||||
})
|
||||
|
||||
snapshotID, err := sm.CreateSnapshotWithName(
|
||||
ctx, cfg.Hostname, "perms", "test-version", "test-git")
|
||||
require.NoError(t, err)
|
||||
|
||||
_, err = scanner.Scan(ctx, dataDir, snapshotID)
|
||||
require.NoError(t, err)
|
||||
|
||||
require.NoError(t, sm.CompleteSnapshot(ctx, snapshotID))
|
||||
require.NoError(t, sm.ExportSnapshotMetadata(ctx, dbPath, snapshotID))
|
||||
require.NoError(t, db.Close())
|
||||
|
||||
return cfg, storer, snapshotID, srcPath
|
||||
}
|
||||
|
||||
// restoredPathFor returns where backupOneFile's source lands under a
|
||||
// restore target: restore recreates each file at its original absolute
|
||||
// path beneath TargetDir.
|
||||
func restoredPathFor(restoreDir, srcPath string) string {
|
||||
return filepath.Join(restoreDir, srcPath)
|
||||
}
|
||||
|
||||
// withUmask022 forces the process umask to 022 for the duration of a
|
||||
// test, so the difference between a 0600 create and a default create is
|
||||
// observable. Restored serially (no t.Parallel) so it does not race
|
||||
// other tests.
|
||||
func withUmask022(t *testing.T) {
|
||||
t.Helper()
|
||||
|
||||
old := syscall.Umask(0o022)
|
||||
|
||||
t.Cleanup(func() { syscall.Umask(old) })
|
||||
}
|
||||
|
||||
// TestRestoreCreatesFileNeverWiderThanStoredMode checks that a file with
|
||||
// a restrictive stored mode (0600) is never observable with a wider mode
|
||||
// while its content is being written, and ends at its stored mode.
|
||||
//
|
||||
//nolint:paralleltest // sets the process umask; must run serially
|
||||
func TestRestoreCreatesFileNeverWiderThanStoredMode(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
withUmask022(t)
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
ctx := context.Background()
|
||||
|
||||
content := randomBytes(t, 4096)
|
||||
|
||||
cfg, storer, snapshotID, srcPath := backupOneFile(
|
||||
ctx, t, fs, tempDir, "secret.bin", content, 0o600)
|
||||
|
||||
restoreDir := filepath.Join(tempDir, "restored")
|
||||
spy := &modeSpyFs{Fs: fs, watch: "secret.bin"}
|
||||
|
||||
v := newRestoreVaultik(ctx, cfg, storer, spy)
|
||||
require.NoError(t, v.Restore(&RestoreOptions{
|
||||
SnapshotID: snapshotID,
|
||||
TargetDir: restoreDir,
|
||||
}))
|
||||
|
||||
spy.mu.Lock()
|
||||
observed := append([]os.FileMode(nil), spy.writeModes...)
|
||||
spy.mu.Unlock()
|
||||
|
||||
require.NotEmpty(t, observed,
|
||||
"spy never saw the output file being written")
|
||||
|
||||
for _, m := range observed {
|
||||
assert.Equalf(t, os.FileMode(0o600), m,
|
||||
"file was observable at mode %o during write; must be 0600", m)
|
||||
}
|
||||
|
||||
// The stored mode is applied after the content is written.
|
||||
info, err := fs.Stat(restoredPathFor(restoreDir, srcPath))
|
||||
require.NoError(t, err)
|
||||
assert.Equal(t, os.FileMode(0o600), info.Mode().Perm())
|
||||
|
||||
got, err := afero.ReadFile(fs, restoredPathFor(restoreDir, srcPath))
|
||||
require.NoError(t, err)
|
||||
require.True(t, bytes.Equal(got, content))
|
||||
}
|
||||
|
||||
// TestRestoreRemovesPartialFileOnWriteFailure checks that a file whose
|
||||
// content write fails is not left behind.
|
||||
//
|
||||
//nolint:paralleltest // sets the process umask; must run serially
|
||||
func TestRestoreRemovesPartialFileOnWriteFailure(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
withUmask022(t)
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
ctx := context.Background()
|
||||
|
||||
cfg, storer, snapshotID, srcPath := backupOneFile(
|
||||
ctx, t, fs, tempDir, "doomed.bin", randomBytes(t, 4096), 0o600)
|
||||
|
||||
restoreDir := filepath.Join(tempDir, "restored")
|
||||
spy := &modeSpyFs{Fs: fs, watch: "doomed.bin", failWrite: true}
|
||||
|
||||
v := newRestoreVaultik(ctx, cfg, storer, spy)
|
||||
err := v.Restore(&RestoreOptions{
|
||||
SnapshotID: snapshotID,
|
||||
TargetDir: restoreDir,
|
||||
})
|
||||
require.Error(t, err, "restore should fail when the write fails")
|
||||
|
||||
exists, err := afero.Exists(fs, restoredPathFor(restoreDir, srcPath))
|
||||
require.NoError(t, err)
|
||||
assert.False(t, exists, "partial file must be removed after a failed write")
|
||||
}
|
||||
|
||||
// TestVerifyRejectsTrailingBytes checks that --verify fails a restored
|
||||
// file that has bytes past its last chunk.
|
||||
//
|
||||
//nolint:paralleltest // sets the process umask; must run serially
|
||||
func TestVerifyRejectsTrailingBytes(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
withUmask022(t)
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
ctx := context.Background()
|
||||
|
||||
cfg, storer, snapshotID, _ := backupOneFile(
|
||||
ctx, t, fs, tempDir, "padded.bin", randomBytes(t, 4096), 0o600)
|
||||
|
||||
restoreDir := filepath.Join(tempDir, "restored")
|
||||
// Append one byte to the file as it is written, so its content still
|
||||
// matches the stored chunks but it is one byte too long.
|
||||
spy := &modeSpyFs{Fs: fs, watch: "padded.bin", trailing: 1}
|
||||
|
||||
v := newRestoreVaultik(ctx, cfg, storer, spy)
|
||||
err := v.Restore(&RestoreOptions{
|
||||
SnapshotID: snapshotID,
|
||||
TargetDir: restoreDir,
|
||||
Verify: true,
|
||||
})
|
||||
require.Error(t, err, "verify should fail on a file with trailing bytes")
|
||||
assert.ErrorIs(t, err, errFilesFailedVerify)
|
||||
}
|
||||
|
||||
// newRestoreVaultik builds a Vaultik wired for a restore-only test.
|
||||
func newRestoreVaultik(
|
||||
ctx context.Context, cfg *config.Config, storer storage.Storer, fs afero.Fs,
|
||||
) *Vaultik {
|
||||
v := &Vaultik{
|
||||
Config: cfg,
|
||||
Storage: storer,
|
||||
Fs: fs,
|
||||
Stdout: io.Discard,
|
||||
Stderr: io.Discard,
|
||||
UI: ui.NewWithColor(io.Discard, false),
|
||||
}
|
||||
v.SetContext(ctx)
|
||||
|
||||
return v
|
||||
}
|
||||
@@ -171,10 +171,13 @@ func (p *restorePlan) finishFile(fileID types.FileID) {
|
||||
// downloaded next, after which it — together with any other pending
|
||||
// files whose blob sets become empty — moves to the ready queue.
|
||||
//
|
||||
// The zero FileID return means nothing is pending.
|
||||
func (p *restorePlan) pickNextDownload() types.FileID {
|
||||
// The second return value is false when no file needs a download, so a
|
||||
// genuine file carrying the nil UUID is picked rather than mistaken for
|
||||
// "nothing left".
|
||||
func (p *restorePlan) pickNextDownload() (types.FileID, bool) {
|
||||
var best types.FileID
|
||||
|
||||
found := false
|
||||
bestCount := math.MaxInt
|
||||
|
||||
var bestID string
|
||||
@@ -188,14 +191,15 @@ func (p *restorePlan) pickNextDownload() types.FileID {
|
||||
}
|
||||
|
||||
idStr := id.String()
|
||||
if n < bestCount || (n == bestCount && (best.IsZero() || idStr < bestID)) {
|
||||
if !found || n < bestCount || (n == bestCount && idStr < bestID) {
|
||||
best = id
|
||||
found = true
|
||||
bestCount = n
|
||||
bestID = idStr
|
||||
}
|
||||
}
|
||||
|
||||
return best
|
||||
return best, found
|
||||
}
|
||||
|
||||
// blobsNeeded returns the uncached blob hashes for fileID in any order.
|
||||
|
||||
@@ -0,0 +1,88 @@
|
||||
package vaultik //nolint:testpackage // inspects unexported restore plan internals
|
||||
|
||||
import (
|
||||
"context"
|
||||
"math"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/types"
|
||||
)
|
||||
|
||||
// TestPickNextDownloadReturnsNilUUIDFile proves a genuine pending file
|
||||
// carrying the nil UUID is picked for download rather than mistaken for
|
||||
// "nothing left" — the bug that could abandon every remaining file.
|
||||
func TestPickNextDownloadReturnsNilUUIDFile(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
var nilID types.FileID // zero value is the nil UUID
|
||||
|
||||
plan := &restorePlan{
|
||||
fileBlobs: map[types.FileID]map[string]struct{}{
|
||||
nilID: {"blobhash": {}},
|
||||
},
|
||||
}
|
||||
|
||||
id, ok := plan.pickNextDownload()
|
||||
require.True(t, ok,
|
||||
"pickNextDownload treated a pending nil-UUID file as nothing to do")
|
||||
require.True(t, id.IsZero(), "expected the nil-UUID file to be picked")
|
||||
}
|
||||
|
||||
// TestPickNextDownloadEmptyPlan confirms the second return value is false
|
||||
// only when no file needs a download.
|
||||
func TestPickNextDownloadEmptyPlan(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
plan := &restorePlan{
|
||||
fileBlobs: map[types.FileID]map[string]struct{}{},
|
||||
}
|
||||
|
||||
_, ok := plan.pickNextDownload()
|
||||
require.False(t, ok, "pickNextDownload reported work on an empty plan")
|
||||
}
|
||||
|
||||
// TestRunRestoreLoopFailsOnAbandonedFiles proves the loop returns an
|
||||
// error rather than silent success when files remain pending after it
|
||||
// can make no further progress.
|
||||
func TestRunRestoreLoopFailsOnAbandonedFiles(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
db, err := database.NewTestDB()
|
||||
require.NoError(t, err)
|
||||
|
||||
t.Cleanup(func() { _ = db.Close() })
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
cache, err := newBlobDiskCache(math.MaxInt64)
|
||||
require.NoError(t, err)
|
||||
|
||||
t.Cleanup(func() { _ = cache.Close() })
|
||||
|
||||
v := &Vaultik{ctx: ctx}
|
||||
session := &restoreSession{
|
||||
v: v,
|
||||
ctx: ctx,
|
||||
repos: repos,
|
||||
sweeper: newRestoreSweeper(ctx, repos, cache, 1),
|
||||
result: &RestoreResult{},
|
||||
}
|
||||
|
||||
// A file that is still pending but whose uncached-blob set is empty
|
||||
// and which was never queued as ready: the loop can neither restore
|
||||
// nor download it. This is the abandonment the guard must catch.
|
||||
var stuck types.FileID
|
||||
|
||||
plan := &restorePlan{
|
||||
fileBlobs: map[types.FileID]map[string]struct{}{stuck: {}},
|
||||
blobFiles: map[string]map[types.FileID]struct{}{},
|
||||
cached: map[string]struct{}{},
|
||||
}
|
||||
|
||||
err = v.runRestoreLoop(session, plan, map[types.FileID]*database.File{}, 0)
|
||||
require.ErrorIs(t, err, errRestoreIncomplete)
|
||||
}
|
||||
@@ -0,0 +1,73 @@
|
||||
package vaultik //nolint:testpackage // inspects unexported snapshot-db materialization
|
||||
|
||||
import (
|
||||
"context"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
)
|
||||
|
||||
// genuineSnapshotDBBytes returns the on-disk bytes of a real snapshot
|
||||
// database (the full schema applied).
|
||||
func genuineSnapshotDBBytes(t *testing.T) []byte {
|
||||
t.Helper()
|
||||
|
||||
path := filepath.Join(t.TempDir(), "snapshot.db")
|
||||
|
||||
db, err := database.New(context.Background(), path)
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, db.Close())
|
||||
|
||||
data, err := os.ReadFile(path) //nolint:gosec // G304: test-controlled temp path
|
||||
require.NoError(t, err)
|
||||
|
||||
return data
|
||||
}
|
||||
|
||||
// TestMaterializeSnapshotDBPrivateDir proves the decrypted database lands
|
||||
// in a private (0700) directory and opens read-only.
|
||||
func TestMaterializeSnapshotDBPrivateDir(t *testing.T) {
|
||||
dbData := genuineSnapshotDBBytes(t)
|
||||
|
||||
t.Setenv("TMPDIR", t.TempDir())
|
||||
|
||||
v := &Vaultik{ctx: context.Background(), Fs: afero.NewOsFs()}
|
||||
|
||||
db, dir, err := v.materializeSnapshotDB(dbData)
|
||||
require.NoError(t, err)
|
||||
|
||||
t.Cleanup(func() {
|
||||
_ = db.Close()
|
||||
_ = os.RemoveAll(dir)
|
||||
})
|
||||
|
||||
info, err := os.Stat(dir)
|
||||
require.NoError(t, err)
|
||||
require.Equal(t, os.FileMode(0o700), info.Mode().Perm(),
|
||||
"snapshot database directory must not be world-readable")
|
||||
|
||||
_, err = db.Conn().ExecContext(context.Background(),
|
||||
"CREATE TABLE probe_readonly (x)")
|
||||
require.Error(t, err, "materialized snapshot database must be read-only")
|
||||
}
|
||||
|
||||
// TestMaterializeSnapshotDBRemovesDirOnOpenFailure proves a failed open
|
||||
// leaves no temp directory behind.
|
||||
func TestMaterializeSnapshotDBRemovesDirOnOpenFailure(t *testing.T) {
|
||||
base := t.TempDir()
|
||||
|
||||
t.Setenv("TMPDIR", base)
|
||||
|
||||
v := &Vaultik{ctx: context.Background(), Fs: afero.NewOsFs()}
|
||||
|
||||
_, _, err := v.materializeSnapshotDB([]byte("this is not a sqlite database"))
|
||||
require.Error(t, err)
|
||||
|
||||
entries, rerr := os.ReadDir(base)
|
||||
require.NoError(t, rerr)
|
||||
require.Empty(t, entries, "temp directory left behind after open failure")
|
||||
}
|
||||
@@ -8,6 +8,7 @@ import (
|
||||
"path/filepath"
|
||||
"regexp"
|
||||
"sort"
|
||||
"strconv"
|
||||
"strings"
|
||||
"time"
|
||||
|
||||
@@ -669,9 +670,11 @@ func (v *Vaultik) VerifySnapshotWithOptions(
|
||||
|
||||
v.printVerifyHeader(snapshotID, opts)
|
||||
|
||||
// Download and parse manifest. The caller supplies a human
|
||||
// snapshot ID; we hash it to address remote storage.
|
||||
manifest, err := v.downloadManifestByKey(snapshot.RemoteSnapshotKey(snapshotID))
|
||||
// Resolve the identifier to the snapshot's remote key and download the
|
||||
// manifest. A human ID is hashed; a remote key (or its abbreviation,
|
||||
// as printed for a remote-only snapshot) is used as-is, so a host with
|
||||
// no local index can verify a snapshot it can only see on the store.
|
||||
manifest, err := v.resolveAndDownloadManifest(snapshotID)
|
||||
if err != nil {
|
||||
if opts.JSON {
|
||||
result.Status = verifyStatusFailed
|
||||
@@ -932,29 +935,23 @@ func (v *Vaultik) downloadManifestByKey(remoteKey string) (*snapshot.Manifest, e
|
||||
func (v *Vaultik) syncWithRemote() error {
|
||||
log.Info("Syncing with remote snapshots")
|
||||
|
||||
// Get all remote snapshot IDs
|
||||
remoteSnapshots := make(map[string]bool)
|
||||
objectCh := v.Storage.ListStream(v.ctx, "metadata/")
|
||||
|
||||
for object := range objectCh {
|
||||
if object.Err != nil {
|
||||
return fmt.Errorf("listing remote snapshots: %w", object.Err)
|
||||
// Remote metadata lives under metadata/<remote-key>/, where the
|
||||
// directory name is snapshot.RemoteSnapshotKey(id), not the human
|
||||
// snapshot ID. Compare each local row's hashed key against that set
|
||||
// so a row still backed by remote metadata is kept. Comparing human
|
||||
// IDs against the hashed directory names matches nothing and deletes
|
||||
// every local snapshot record (issue #160).
|
||||
remoteKeys, err := v.listAllRemoteSnapshotKeys()
|
||||
if err != nil {
|
||||
return fmt.Errorf("listing remote snapshots: %w", err)
|
||||
}
|
||||
|
||||
// Extract snapshot ID from paths like metadata/hostname-20240115-143052Z/
|
||||
parts := strings.Split(object.Key, "/")
|
||||
if len(parts) >= minSnapshotIDParts &&
|
||||
parts[0] == metadataDirName && parts[1] != "" {
|
||||
// Skip macOS resource fork files (._*) and other hidden files
|
||||
if strings.HasPrefix(parts[1], ".") {
|
||||
continue
|
||||
remoteKeySet := make(map[string]bool, len(remoteKeys))
|
||||
for _, k := range remoteKeys {
|
||||
remoteKeySet[k] = true
|
||||
}
|
||||
|
||||
remoteSnapshots[parts[1]] = true
|
||||
}
|
||||
}
|
||||
|
||||
log.Debug("Found remote snapshots", "count", len(remoteSnapshots))
|
||||
log.Debug("Found remote snapshots", "count", len(remoteKeySet))
|
||||
|
||||
// Get all local snapshots (use a high limit to get all)
|
||||
localSnapshots, err := v.Repositories.Snapshots.ListRecent(v.ctx, listRecentLimit)
|
||||
@@ -962,12 +959,12 @@ func (v *Vaultik) syncWithRemote() error {
|
||||
return fmt.Errorf("listing local snapshots: %w", err)
|
||||
}
|
||||
|
||||
// Remove local snapshots that don't exist remotely
|
||||
// Remove local snapshots whose metadata is absent from the remote.
|
||||
removedCount := 0
|
||||
|
||||
for _, snap := range localSnapshots {
|
||||
snapshotIDStr := snap.ID.String()
|
||||
if !remoteSnapshots[snapshotIDStr] {
|
||||
if !remoteKeySet[snapshot.RemoteSnapshotKey(snapshotIDStr)] {
|
||||
log.Info("Removing local snapshot not found in remote",
|
||||
"snapshot_id", snap.ID)
|
||||
|
||||
@@ -1540,12 +1537,17 @@ func (v *Vaultik) outputRemoveJSON(result *RemoveResult) error {
|
||||
return encoder.Encode(result)
|
||||
}
|
||||
|
||||
// PruneResult contains statistics about the prune operation
|
||||
// PruneResult contains statistics about the prune operation.
|
||||
// SnapshotsDeleted counts snapshots actually deleted. FilesDeleted,
|
||||
// ChunksDeleted, and BlobsDeleted are derived from before/after row
|
||||
// counts of the local index; each is nil when a count could not be read,
|
||||
// so an unreadable count is reported as unknown rather than silently
|
||||
// as 0.
|
||||
type PruneResult struct {
|
||||
SnapshotsDeleted int64
|
||||
FilesDeleted int64
|
||||
ChunksDeleted int64
|
||||
BlobsDeleted int64
|
||||
FilesDeleted *int64
|
||||
ChunksDeleted *int64
|
||||
BlobsDeleted *int64
|
||||
}
|
||||
|
||||
// PruneDatabase removes incomplete snapshots and orphaned files, chunks,
|
||||
@@ -1560,7 +1562,7 @@ func (v *Vaultik) PruneDatabase() (*PruneResult, error) {
|
||||
result := &PruneResult{}
|
||||
|
||||
// Snapshot counts before deletion of incompletes.
|
||||
snapshotCountBefore, _ := v.getTableCount("snapshots")
|
||||
snapshotCountBefore := v.tableCountForReport("snapshots")
|
||||
|
||||
// First, delete any incomplete snapshots
|
||||
incompleteSnapshots, err := v.Repositories.Snapshots.GetIncompleteSnapshots(v.ctx)
|
||||
@@ -1575,9 +1577,9 @@ func (v *Vaultik) PruneDatabase() (*PruneResult, error) {
|
||||
}
|
||||
|
||||
// Get counts before cleanup for reporting
|
||||
fileCountBefore, _ := v.getTableCount("files")
|
||||
chunkCountBefore, _ := v.getTableCount("chunks")
|
||||
blobCountBefore, _ := v.getTableCount("blobs")
|
||||
fileCountBefore := v.tableCountForReport("files")
|
||||
chunkCountBefore := v.tableCountForReport("chunks")
|
||||
blobCountBefore := v.tableCountForReport("blobs")
|
||||
|
||||
// Run the cleanup
|
||||
err = v.SnapshotManager.CleanupOrphanedData(v.ctx)
|
||||
@@ -1586,36 +1588,83 @@ func (v *Vaultik) PruneDatabase() (*PruneResult, error) {
|
||||
}
|
||||
|
||||
// Get counts after cleanup
|
||||
fileCountAfter, _ := v.getTableCount("files")
|
||||
chunkCountAfter, _ := v.getTableCount("chunks")
|
||||
blobCountAfter, _ := v.getTableCount("blobs")
|
||||
fileCountAfter := v.tableCountForReport("files")
|
||||
chunkCountAfter := v.tableCountForReport("chunks")
|
||||
blobCountAfter := v.tableCountForReport("blobs")
|
||||
|
||||
result.FilesDeleted = fileCountBefore - fileCountAfter
|
||||
result.ChunksDeleted = chunkCountBefore - chunkCountAfter
|
||||
result.BlobsDeleted = blobCountBefore - blobCountAfter
|
||||
result.FilesDeleted = countDiff(fileCountBefore, fileCountAfter)
|
||||
result.ChunksDeleted = countDiff(chunkCountBefore, chunkCountAfter)
|
||||
result.BlobsDeleted = countDiff(blobCountBefore, blobCountAfter)
|
||||
|
||||
log.Info("Local database prune complete",
|
||||
"incomplete_snapshots", result.SnapshotsDeleted,
|
||||
"orphaned_files", result.FilesDeleted,
|
||||
"orphaned_chunks", result.ChunksDeleted,
|
||||
"orphaned_blobs", result.BlobsDeleted,
|
||||
"orphaned_files", countText(result.FilesDeleted),
|
||||
"orphaned_chunks", countText(result.ChunksDeleted),
|
||||
"orphaned_blobs", countText(result.BlobsDeleted),
|
||||
)
|
||||
|
||||
snapshotCountAfter := snapshotCountBefore - result.SnapshotsDeleted
|
||||
// Snapshots remaining after removing the incomplete ones; unknown if
|
||||
// the pre-prune snapshot count could not be read.
|
||||
snapshotsRemain := countDiff(snapshotCountBefore, &result.SnapshotsDeleted)
|
||||
|
||||
v.UI.Completef("Pruned local index database.")
|
||||
v.UI.Detailf("Incomplete snapshots: %d removed (%d remain).",
|
||||
result.SnapshotsDeleted, snapshotCountAfter)
|
||||
v.UI.Detailf("Orphaned files: %d removed (%d remain).",
|
||||
result.FilesDeleted, fileCountAfter)
|
||||
v.UI.Detailf("Orphaned chunks: %d removed (%d remain).",
|
||||
result.ChunksDeleted, chunkCountAfter)
|
||||
v.UI.Detailf("Orphaned blobs: %d removed (%d remain).",
|
||||
result.BlobsDeleted, blobCountAfter)
|
||||
v.UI.Detailf("Incomplete snapshots: %s removed (%s remain).",
|
||||
countText(&result.SnapshotsDeleted), countText(snapshotsRemain))
|
||||
v.UI.Detailf("Orphaned files: %s removed (%s remain).",
|
||||
countText(result.FilesDeleted), countText(fileCountAfter))
|
||||
v.UI.Detailf("Orphaned chunks: %s removed (%s remain).",
|
||||
countText(result.ChunksDeleted), countText(chunkCountAfter))
|
||||
v.UI.Detailf("Orphaned blobs: %s removed (%s remain).",
|
||||
countText(result.BlobsDeleted), countText(blobCountAfter))
|
||||
|
||||
return result, nil
|
||||
}
|
||||
|
||||
// countUnknown is what a count reads as when its query could not be run,
|
||||
// distinct from "0", which means the table really was empty.
|
||||
const countUnknown = "unknown"
|
||||
|
||||
// tableCountForReport returns the row count of a table for the prune
|
||||
// summary, or nil if the count could not be read. A read failure is
|
||||
// logged at warn — visible even under --json, which routes warnings to
|
||||
// stderr — and then rendered as unknown rather than silently becoming 0,
|
||||
// so a broken query is a visible failure instead of a plausible wrong
|
||||
// number.
|
||||
func (v *Vaultik) tableCountForReport(tableName string) *int64 {
|
||||
count, err := v.getTableCount(tableName)
|
||||
if err != nil {
|
||||
log.Warn("could not read table row count for prune summary",
|
||||
"table", tableName, "error", err)
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
return &count
|
||||
}
|
||||
|
||||
// countDiff returns before-after, or nil if either count is unknown so
|
||||
// that an unreadable count does not collapse into a plausible delta.
|
||||
func countDiff(before, after *int64) *int64 {
|
||||
if before == nil || after == nil {
|
||||
return nil
|
||||
}
|
||||
|
||||
diff := *before - *after
|
||||
|
||||
return &diff
|
||||
}
|
||||
|
||||
// countText renders a count that may be unknown: nil (the read failed)
|
||||
// becomes "unknown", never "0", so a reader can tell an empty table from
|
||||
// one that could not be queried.
|
||||
func countText(count *int64) string {
|
||||
if count == nil {
|
||||
return countUnknown
|
||||
}
|
||||
|
||||
return strconv.FormatInt(*count, 10)
|
||||
}
|
||||
|
||||
// validTableNameRe matches table names containing only lowercase
|
||||
// alphanumeric characters and underscores.
|
||||
var validTableNameRe = regexp.MustCompile(`^[a-z0-9_]+$`)
|
||||
|
||||
@@ -0,0 +1,101 @@
|
||||
package vaultik
|
||||
|
||||
import (
|
||||
"errors"
|
||||
"fmt"
|
||||
"strings"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
)
|
||||
|
||||
// remoteKeyHexLen is the length of a full remote snapshot key: a SHA256
|
||||
// digest rendered as lowercase hex.
|
||||
const remoteKeyHexLen = 64
|
||||
|
||||
// Sentinel errors for resolving a snapshot identifier against the store.
|
||||
var (
|
||||
errSnapshotKeyNotFound = errors.New(
|
||||
"no snapshot on the destination store matches this identifier")
|
||||
errSnapshotKeyAmbiguous = errors.New(
|
||||
"identifier matches more than one snapshot on the destination store")
|
||||
)
|
||||
|
||||
// resolveSnapshotRemoteKey turns a snapshot identifier supplied on the
|
||||
// command line into the remote key that names the snapshot's metadata
|
||||
// directory on the destination store. Every remote path a restore or
|
||||
// verify reads is built from that key.
|
||||
//
|
||||
// Two forms are accepted, matching the two things a host can know:
|
||||
//
|
||||
// - A human snapshot ID (hostname_name_timestamp), which a host holding
|
||||
// the local index has. It is hashed to its remote key; the store is
|
||||
// not consulted.
|
||||
// - A remote key, or the leading part of one, which is all a host with
|
||||
// no local index can know — it is exactly what `snapshot list` prints
|
||||
// for a remote-only snapshot (see formatRemoteOnlyID). It is resolved
|
||||
// against the destination store's metadata listing; an identifier that
|
||||
// matches no snapshot, or more than one, is an error.
|
||||
//
|
||||
// The two are told apart by shape: a remote key is lowercase hex, and a
|
||||
// human snapshot ID never is (it carries a hostname, underscores, and an
|
||||
// RFC3339 timestamp).
|
||||
func (v *Vaultik) resolveSnapshotRemoteKey(identifier string) (string, error) {
|
||||
if !isRemoteKeyOrPrefix(identifier) {
|
||||
return snapshot.RemoteSnapshotKey(identifier), nil
|
||||
}
|
||||
|
||||
keys, err := v.listAllRemoteSnapshotKeys()
|
||||
if err != nil {
|
||||
return "", fmt.Errorf(
|
||||
"listing destination store to resolve %q: %w", identifier, err)
|
||||
}
|
||||
|
||||
var matches []string
|
||||
|
||||
for _, key := range keys {
|
||||
if strings.HasPrefix(key, identifier) {
|
||||
matches = append(matches, key)
|
||||
}
|
||||
}
|
||||
|
||||
switch len(matches) {
|
||||
case 1:
|
||||
return matches[0], nil
|
||||
case 0:
|
||||
return "", fmt.Errorf("%w: %s", errSnapshotKeyNotFound, identifier)
|
||||
default:
|
||||
return "", fmt.Errorf("%w: %s (%d matches)",
|
||||
errSnapshotKeyAmbiguous, identifier, len(matches))
|
||||
}
|
||||
}
|
||||
|
||||
// resolveAndDownloadManifest resolves a snapshot identifier to its remote
|
||||
// key (see resolveSnapshotRemoteKey) and downloads that snapshot's
|
||||
// manifest.
|
||||
func (v *Vaultik) resolveAndDownloadManifest(
|
||||
identifier string,
|
||||
) (*snapshot.Manifest, error) {
|
||||
remoteKey, err := v.resolveSnapshotRemoteKey(identifier)
|
||||
if err != nil {
|
||||
return nil, err
|
||||
}
|
||||
|
||||
return v.downloadManifestByKey(remoteKey)
|
||||
}
|
||||
|
||||
// isRemoteKeyOrPrefix reports whether s is a full remote key or the
|
||||
// leading part of one: 1 to 64 lowercase hex characters. A human snapshot
|
||||
// ID is never all hex, so this shape test is enough to tell the two apart.
|
||||
func isRemoteKeyOrPrefix(s string) bool {
|
||||
if s == "" || len(s) > remoteKeyHexLen {
|
||||
return false
|
||||
}
|
||||
|
||||
for _, r := range s {
|
||||
if (r < '0' || r > '9') && (r < 'a' || r > 'f') {
|
||||
return false
|
||||
}
|
||||
}
|
||||
|
||||
return true
|
||||
}
|
||||
+90
-56
@@ -9,12 +9,12 @@ import (
|
||||
"hash"
|
||||
"io"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"time"
|
||||
|
||||
"github.com/klauspost/compress/zstd"
|
||||
|
||||
// Blank import registers the pure-Go sqlite driver for database/sql.
|
||||
_ "modernc.org/sqlite"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
)
|
||||
@@ -29,6 +29,8 @@ var (
|
||||
errTrailingBlobData = errors.New(
|
||||
"blob has unexpected trailing bytes not covered by chunk list")
|
||||
errManifestExtraBlob = errors.New("manifest contains blob not in database")
|
||||
errManifestMissingBlob = errors.New(
|
||||
"manifest omits blob present in database")
|
||||
errBlobSizeMismatch = errors.New("blob size mismatch")
|
||||
)
|
||||
|
||||
@@ -138,8 +140,15 @@ func (v *Vaultik) RunDeepVerify(snapshotID string, opts *VerifyOptions) error {
|
||||
func (v *Vaultik) loadVerificationData(
|
||||
snapshotID string, opts *VerifyOptions, result *VerifyResult,
|
||||
) (*snapshot.Manifest, *tempDB, []snapshot.BlobInfo, error) {
|
||||
// All remote paths use the hashed key derived from the human ID.
|
||||
remoteKey := snapshot.RemoteSnapshotKey(snapshotID)
|
||||
// Resolve the identifier to the snapshot's remote key. A human ID is
|
||||
// hashed; a remote key (or its abbreviation, as printed for a
|
||||
// remote-only snapshot) is used as-is, so a host with no local index
|
||||
// can verify a snapshot it can only see on the store.
|
||||
remoteKey, err := v.resolveSnapshotRemoteKey(snapshotID)
|
||||
if err != nil {
|
||||
return nil, nil, nil, v.deepVerifyFailure(result, opts,
|
||||
fmt.Sprintf("resolving snapshot identifier: %v", err), err)
|
||||
}
|
||||
|
||||
// Download manifest. downloadManifestByKey is the single reader for
|
||||
// remote manifests; see its doc comment.
|
||||
@@ -186,7 +195,7 @@ func (v *Vaultik) loadVerificationData(
|
||||
fmt.Errorf("failed to decrypt database: %w", err))
|
||||
}
|
||||
|
||||
dbBlobs, err := v.getBlobsFromDatabase(snapshotID, tdb.DB)
|
||||
dbBlobs, err := v.getBlobsFromDatabase(tdb.db.Conn())
|
||||
if err != nil {
|
||||
_ = tdb.Close()
|
||||
|
||||
@@ -247,7 +256,7 @@ func (v *Vaultik) runVerificationSteps(
|
||||
len(dbBlobs), ubytes(totalSize))
|
||||
}
|
||||
|
||||
err = v.performDeepVerificationFromDB(dbBlobs, tdb.DB, opts)
|
||||
err = v.performDeepVerificationFromDB(dbBlobs, tdb.db.Conn(), opts)
|
||||
if err != nil {
|
||||
return v.deepVerifyFailure(result, opts, err.Error(), err)
|
||||
}
|
||||
@@ -255,16 +264,18 @@ func (v *Vaultik) runVerificationSteps(
|
||||
return nil
|
||||
}
|
||||
|
||||
// tempDB wraps sql.DB with cleanup
|
||||
// tempDB is the downloaded snapshot database opened read-only for deep
|
||||
// verify, held in a private temp directory removed in full on Close.
|
||||
type tempDB struct {
|
||||
*sql.DB
|
||||
|
||||
tempPath string
|
||||
db *database.DB
|
||||
tempDir string
|
||||
}
|
||||
|
||||
func (t *tempDB) Close() error {
|
||||
err := t.DB.Close()
|
||||
_ = os.Remove(t.tempPath)
|
||||
err := t.db.Close()
|
||||
// Remove the whole private directory so the decrypted database and
|
||||
// any SQLite side files are gone on every path.
|
||||
_ = os.RemoveAll(t.tempDir)
|
||||
|
||||
return err
|
||||
}
|
||||
@@ -291,41 +302,56 @@ func (v *Vaultik) decryptAndLoadDatabase(reader io.ReadCloser) (*tempDB, error)
|
||||
}
|
||||
defer decompressor.Close()
|
||||
|
||||
// Create temporary file for the database
|
||||
tempFile, err := os.CreateTemp("", "vaultik-verify-*.db")
|
||||
// Materialize the decrypted database inside a private (0700) temp
|
||||
// directory so it is never world-readable, and remove the whole
|
||||
// directory on any failure below.
|
||||
tempDir, err := os.MkdirTemp("", "vaultik-verify-")
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("failed to create temp directory: %w", err)
|
||||
}
|
||||
|
||||
success := false
|
||||
|
||||
defer func() {
|
||||
if !success {
|
||||
_ = os.RemoveAll(tempDir)
|
||||
}
|
||||
}()
|
||||
|
||||
dbPath := filepath.Join(tempDir, snapshotDBFilename)
|
||||
|
||||
//nolint:gosec // G304: dbPath is our MkdirTemp dir plus a constant filename
|
||||
tempFile, err := os.OpenFile(
|
||||
dbPath, os.O_CREATE|os.O_EXCL|os.O_WRONLY, restoreFileMode)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("failed to create temp file: %w", err)
|
||||
}
|
||||
|
||||
tempPath := tempFile.Name()
|
||||
|
||||
// Stream decompress directly to file
|
||||
log.Info("Decompressing database...")
|
||||
|
||||
written, err := io.Copy(tempFile, decompressor)
|
||||
if err != nil {
|
||||
_ = tempFile.Close()
|
||||
_ = os.Remove(tempPath)
|
||||
|
||||
return nil, fmt.Errorf("failed to decompress database: %w", err)
|
||||
}
|
||||
|
||||
_ = tempFile.Close()
|
||||
err = tempFile.Close()
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("failed to close temp database file: %w", err)
|
||||
}
|
||||
|
||||
log.Info("Database decompressed", "size", ubytes(written))
|
||||
|
||||
// Open the database
|
||||
db, err := sql.Open("sqlite", tempPath)
|
||||
db, err := database.OpenReadOnly(v.ctx, dbPath)
|
||||
if err != nil {
|
||||
_ = os.Remove(tempPath)
|
||||
|
||||
return nil, fmt.Errorf("failed to open database: %w", err)
|
||||
}
|
||||
|
||||
return &tempDB{
|
||||
DB: db,
|
||||
tempPath: tempPath,
|
||||
}, nil
|
||||
success = true
|
||||
|
||||
return &tempDB{db: db, tempDir: tempDir}, nil
|
||||
}
|
||||
|
||||
// verifyBlob downloads and verifies a single blob
|
||||
@@ -344,12 +370,8 @@ func (v *Vaultik) verifyBlob(blobInfo snapshot.BlobInfo, db *sql.DB) error {
|
||||
return fmt.Errorf("failed to get decryptor: %w", err)
|
||||
}
|
||||
|
||||
// Hash the encrypted blob data as it streams through to decryption
|
||||
blobHasher := sha256.New()
|
||||
teeReader := io.TeeReader(reader, blobHasher)
|
||||
|
||||
// Decrypt blob (reading through teeReader to hash encrypted data)
|
||||
decryptedReader, err := decryptor.DecryptStream(teeReader)
|
||||
// Decrypt blob
|
||||
decryptedReader, err := decryptor.DecryptStream(reader)
|
||||
if err != nil {
|
||||
return fmt.Errorf("failed to decrypt: %w", err)
|
||||
}
|
||||
@@ -361,12 +383,19 @@ func (v *Vaultik) verifyBlob(blobInfo snapshot.BlobInfo, db *sql.DB) error {
|
||||
}
|
||||
defer decompressor.Close()
|
||||
|
||||
chunkCount, err := v.verifyBlobChunks(db, blobInfo.Hash, decompressor)
|
||||
// A blob's hash — its remote name — is the double SHA256 of its
|
||||
// decompressed plaintext (see blobgen.Writer.Sum256), not of the
|
||||
// encrypted bytes. Hash the plaintext as chunk verification streams
|
||||
// it, then compare on completion.
|
||||
plaintextHasher := sha256.New()
|
||||
hashedStream := io.TeeReader(decompressor, plaintextHasher)
|
||||
|
||||
chunkCount, err := v.verifyBlobChunks(db, blobInfo.Hash, hashedStream)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
err = v.verifyBlobFinalIntegrity(decompressor, blobHasher, blobInfo.Hash)
|
||||
err = v.verifyBlobFinalIntegrity(hashedStream, plaintextHasher, blobInfo.Hash)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
@@ -470,14 +499,13 @@ func (v *Vaultik) verifyBlobChunks(
|
||||
}
|
||||
|
||||
// verifyBlobFinalIntegrity checks that no trailing data exists in the
|
||||
// decompressed stream and that the encrypted blob hash matches the
|
||||
// expected value.
|
||||
// decompressed stream and that the blob hash matches the expected value.
|
||||
func (v *Vaultik) verifyBlobFinalIntegrity(
|
||||
decompressor io.Reader, blobHasher hash.Hash, expectedHash string,
|
||||
plaintext io.Reader, plaintextHasher hash.Hash, expectedHash string,
|
||||
) error {
|
||||
// Verify no remaining data in blob - if the chunk list is accurate,
|
||||
// the blob should be fully consumed.
|
||||
remaining, err := io.Copy(io.Discard, decompressor)
|
||||
remaining, err := io.Copy(io.Discard, plaintext)
|
||||
if err != nil {
|
||||
return fmt.Errorf("failed to check for remaining blob data: %w", err)
|
||||
}
|
||||
@@ -486,8 +514,11 @@ func (v *Vaultik) verifyBlobFinalIntegrity(
|
||||
return fmt.Errorf("%w: %d bytes", errTrailingBlobData, remaining)
|
||||
}
|
||||
|
||||
// Verify blob hash matches the encrypted data we downloaded
|
||||
calculatedBlobHash := hex.EncodeToString(blobHasher.Sum(nil))
|
||||
// The blob hash is the double SHA256 of its plaintext content.
|
||||
firstHash := plaintextHasher.Sum(nil)
|
||||
secondHash := sha256.Sum256(firstHash)
|
||||
calculatedBlobHash := hex.EncodeToString(secondHash[:])
|
||||
|
||||
if calculatedBlobHash != expectedHash {
|
||||
return fmt.Errorf("%w: calculated %s, expected %s",
|
||||
errBlobHashMismatch, calculatedBlobHash, expectedHash)
|
||||
@@ -496,19 +527,21 @@ func (v *Vaultik) verifyBlobFinalIntegrity(
|
||||
return nil
|
||||
}
|
||||
|
||||
// getBlobsFromDatabase gets all blobs for the snapshot from the database
|
||||
func (v *Vaultik) getBlobsFromDatabase(
|
||||
snapshotID string, db *sql.DB,
|
||||
) ([]snapshot.BlobInfo, error) {
|
||||
// getBlobsFromDatabase gets all blobs for the snapshot from the database.
|
||||
//
|
||||
// The exported per-snapshot database holds exactly one snapshot's data
|
||||
// (see cleanSnapshotDB), so every row in snapshot_blobs belongs to it.
|
||||
// We select them directly rather than filtering by the human snapshot ID,
|
||||
// which a host restoring from the store alone does not have.
|
||||
func (v *Vaultik) getBlobsFromDatabase(db *sql.DB) ([]snapshot.BlobInfo, error) {
|
||||
query := `
|
||||
SELECT b.blob_hash, b.compressed_size
|
||||
FROM snapshot_blobs sb
|
||||
JOIN blobs b ON sb.blob_hash = b.blob_hash
|
||||
WHERE sb.snapshot_id = ?
|
||||
ORDER BY b.blob_hash
|
||||
`
|
||||
|
||||
rows, err := db.QueryContext(v.ctx, query, snapshotID)
|
||||
rows, err := db.QueryContext(v.ctx, query)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("failed to query snapshot blobs: %w", err)
|
||||
}
|
||||
@@ -561,16 +594,11 @@ func (v *Vaultik) verifyManifestAgainstDatabase(
|
||||
manifestBlobMap[blob.Hash] = blob.CompressedSize
|
||||
}
|
||||
|
||||
// Check counts match
|
||||
if len(dbBlobMap) != len(manifestBlobMap) {
|
||||
log.Warn("Manifest blob count mismatch",
|
||||
"database_blobs", len(dbBlobMap),
|
||||
"manifest_blobs", len(manifestBlobMap),
|
||||
)
|
||||
// This is a warning, not an error - database is authoritative
|
||||
}
|
||||
|
||||
// Check each manifest blob exists in database with correct size
|
||||
// The manifest is the only blob list prune consults, so it must match
|
||||
// the database exactly. A blob in the manifest but not the database
|
||||
// points at a corrupt manifest; a blob in the database but omitted
|
||||
// from the manifest would be pruned away while this snapshot still
|
||||
// needs it. Either divergence fails verification.
|
||||
for hash, manifestSize := range manifestBlobMap {
|
||||
dbSize, exists := dbBlobMap[hash]
|
||||
if !exists {
|
||||
@@ -584,6 +612,12 @@ func (v *Vaultik) verifyManifestAgainstDatabase(
|
||||
}
|
||||
}
|
||||
|
||||
for hash := range dbBlobMap {
|
||||
if _, exists := manifestBlobMap[hash]; !exists {
|
||||
return fmt.Errorf("%w: %s", errManifestMissingBlob, hash)
|
||||
}
|
||||
}
|
||||
|
||||
log.Info("✓ Manifest verified against database",
|
||||
"manifest_blobs", len(manifestBlobMap),
|
||||
"database_blobs", len(dbBlobMap),
|
||||
|
||||
@@ -0,0 +1,61 @@
|
||||
package vaultik //nolint:testpackage // calls unexported verifyManifestAgainstDatabase
|
||||
|
||||
import (
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
)
|
||||
|
||||
// Blob hashes shared by the manifest-verification tests below.
|
||||
const (
|
||||
manifestTestBlobA = "blob-a"
|
||||
manifestTestBlobB = "blob-b"
|
||||
)
|
||||
|
||||
// TestVerifyManifestAgainstDatabase_MissingBlobFails is the regression
|
||||
// guard for issue #157: deep verify must fail when the manifest omits a
|
||||
// blob the database records. The divergence used to be logged as a
|
||||
// warning while verification still returned ok, so an incomplete
|
||||
// manifest — the exact defect that lets prune later delete a needed blob
|
||||
// — passed unnoticed.
|
||||
func TestVerifyManifestAgainstDatabase_MissingBlobFails(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
v := &Vaultik{}
|
||||
|
||||
dbBlobs := []snapshot.BlobInfo{
|
||||
{Hash: manifestTestBlobA, CompressedSize: 10},
|
||||
{Hash: manifestTestBlobB, CompressedSize: 20},
|
||||
}
|
||||
manifest := &snapshot.Manifest{
|
||||
Blobs: []snapshot.BlobInfo{
|
||||
{Hash: manifestTestBlobA, CompressedSize: 10},
|
||||
},
|
||||
}
|
||||
|
||||
err := v.verifyManifestAgainstDatabase(manifest, dbBlobs)
|
||||
require.Error(t, err, "verify must fail when the manifest omits a database blob")
|
||||
assert.Contains(t, err.Error(), manifestTestBlobB)
|
||||
}
|
||||
|
||||
// TestVerifyManifestAgainstDatabase_MatchingSetsPass keeps the other half
|
||||
// honest: identical blob sets still verify, so the check above cannot be
|
||||
// satisfied by failing everything.
|
||||
func TestVerifyManifestAgainstDatabase_MatchingSetsPass(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
v := &Vaultik{}
|
||||
|
||||
blobs := []snapshot.BlobInfo{
|
||||
{Hash: manifestTestBlobA, CompressedSize: 10},
|
||||
{Hash: manifestTestBlobB, CompressedSize: 20},
|
||||
}
|
||||
manifest := &snapshot.Manifest{Blobs: blobs}
|
||||
|
||||
require.NoError(t, v.verifyManifestAgainstDatabase(manifest, blobs))
|
||||
}
|
||||
+18
-18
@@ -48,11 +48,12 @@ missing() {
|
||||
! command -v "$1" >/dev/null 2>&1
|
||||
}
|
||||
|
||||
# Docker is a hard requirement, not a nice-to-have: script/lint runs the
|
||||
# digest-pinned golangci-lint image from the Dockerfile's lint stage, and
|
||||
# script/check and script/precommit both run script/lint. A bootstrap
|
||||
# that prints "bootstrap complete" on a machine where `make check` cannot
|
||||
# run is a false success, so this fails instead.
|
||||
# Docker is a hard requirement, not a nice-to-have: script/lint lints by
|
||||
# building Dockerfile.lint, whose digest-pinned golangci-lint image is
|
||||
# the only place the linter runs, and script/check and script/precommit
|
||||
# both run script/lint. A bootstrap that prints "bootstrap complete" on a
|
||||
# machine where `make check` cannot run is a false success, so this fails
|
||||
# instead.
|
||||
#
|
||||
# Installing docker from here was considered and rejected: it needs root,
|
||||
# a running daemon, and on macOS a GUI cask, so an attempt would itself
|
||||
@@ -79,13 +80,15 @@ bootstrap: FAILED - $reason.
|
||||
|
||||
Docker is required to develop this repo. Without it these do not work:
|
||||
|
||||
script/lint runs the digest-pinned golangci-lint image declared
|
||||
by the Dockerfile's lint stage, which is the single
|
||||
source of truth for the linter version
|
||||
script/lint builds Dockerfile.lint, which runs the linter as a
|
||||
build step in a digest-pinned golangci-lint image.
|
||||
That FROM line is the single source of truth for the
|
||||
linter version
|
||||
script/check runs script/lint
|
||||
script/precommit runs script/check, so commits are blocked by the
|
||||
pre-commit hook installed by script/setup
|
||||
script/cibuild builds the Dockerfile, which is what CI runs
|
||||
script/cibuild builds Dockerfile.lint and Dockerfile, which is what
|
||||
CI runs
|
||||
|
||||
Install docker (and start the daemon, checking DOCKER_HOST and your
|
||||
group membership), then re-run script/bootstrap. golangci-lint on PATH
|
||||
@@ -104,15 +107,12 @@ main() {
|
||||
# Go toolchain
|
||||
if missing go; then pkg_install go golang go go; fi
|
||||
|
||||
# golangci-lint is deliberately NOT installed: script/lint runs the
|
||||
# digest-pinned golangci-lint image from the Dockerfile's lint stage,
|
||||
# so whatever a package manager happens to ship would only be a
|
||||
# shadow of the pinned version that could drift from CI. script/lint
|
||||
# will not use a PATH binary on a host at any version, so installing
|
||||
# one here would buy nothing.
|
||||
|
||||
# sqlite3 CLI: the test suite shells out to it (VACUUM).
|
||||
if missing sqlite3; then pkg_install sqlite sqlite3 sqlite sqlite; fi
|
||||
# golangci-lint is deliberately NOT installed: script/lint lints by
|
||||
# building Dockerfile.lint, whose digest-pinned image is the only
|
||||
# place the linter runs, so whatever a package manager happens to
|
||||
# ship would only be a shadow of the pinned version that could drift
|
||||
# from CI. Nothing on the host is ever used as a linter, at any
|
||||
# version, so installing one here would buy nothing.
|
||||
|
||||
# goreleaser, at the version pinned by script/install-goreleaser and
|
||||
# verified against a hardcoded sha256. Package managers are not used
|
||||
|
||||
+48
-12
@@ -1,22 +1,31 @@
|
||||
#!/bin/sh
|
||||
# script/cibuild: run the CI build. The Dockerfile does not run
|
||||
# script/check; it runs `make fmt-check` and `make lint` in its lint
|
||||
# stage and `make test` in its builder stage. A successful build
|
||||
# implies those three passed, provided they actually ran -- which is
|
||||
# what the CHECK_EPOCH below is for.
|
||||
# Generic: needs no adaptation. The Gitea workflow runs this on push.
|
||||
# script/cibuild: run the CI build. This is the full gate, and it is two
|
||||
# builds, in this order:
|
||||
#
|
||||
# Dockerfile.lint the linter, as a build step (a clean build IS a
|
||||
# clean lint)
|
||||
# Dockerfile `make fmt-check` and `make test` in the builder
|
||||
# stage, then the product image
|
||||
#
|
||||
# Either one failing fails this script. Note what follows from the
|
||||
# split: script/docker builds only the product image and so no longer
|
||||
# lints -- this script and script/check (which runs script/lint) are the
|
||||
# things that decide whether the tree is clean.
|
||||
#
|
||||
# Generic apart from the two Dockerfiles: the Gitea workflow runs this
|
||||
# on push.
|
||||
set -eu
|
||||
|
||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||
|
||||
main() {
|
||||
cd "$ROOT"
|
||||
# The Dockerfile's check layers are keyed on CHECK_EPOCH, so a
|
||||
# fresh value here is what forces them to re-run: without it an
|
||||
# Both Dockerfiles key their check layers on CHECK_EPOCH, so a fresh
|
||||
# value is what forces those layers to re-run: without it an
|
||||
# unchanged tree replays them from cache, the checks never execute,
|
||||
# and the build still exits 0. The ARG sits immediately above the
|
||||
# check RUNs, so dependency and module layers still cache. The
|
||||
# Dockerfile also refuses to build at all when CHECK_EPOCH is empty,
|
||||
# and the build still exits 0. Each ARG sits immediately above the
|
||||
# check RUNs, so dependency and module layers still cache. Both
|
||||
# Dockerfiles also refuse to build at all when CHECK_EPOCH is empty,
|
||||
# so a missing value fails loudly here rather than passing quietly.
|
||||
#
|
||||
# The value must be unique per invocation, not per second. `date +%s`
|
||||
@@ -35,8 +44,35 @@ main() {
|
||||
# script exists to prevent -- so the guard would disarm itself and
|
||||
# still exit 0. As a bare assignment, `set -e` catches a failing
|
||||
# `date` and no build starts.
|
||||
#
|
||||
# A separate value per build, because they are separate builds: one
|
||||
# `date` shared between them would still be fresh, but reusing it
|
||||
# invites the two to be collapsed into a single value that is
|
||||
# computed somewhere else and passed in.
|
||||
epoch="$(date +%s%N)$$"
|
||||
docker build --build-arg CHECK_EPOCH="$epoch" .
|
||||
# cacheonly for the lint build: its verdict is the exit status and
|
||||
# the image is never run, so exporting it is pure cost. See
|
||||
# script/lint.
|
||||
docker build --output=type=cacheonly \
|
||||
--build-arg CHECK_EPOCH="$epoch" -f Dockerfile.lint .
|
||||
|
||||
# Version, commit and build date are computed here on the host, the
|
||||
# same way script/docker does, and passed into the product build so
|
||||
# the CI-built image reports its real source. The build context
|
||||
# excludes .git (see .dockerignore), so the build cannot derive them
|
||||
# itself; without these it would stamp the Dockerfile's dev/unknown
|
||||
# fallbacks. VERSION comes from script/version, the source of truth
|
||||
# shared with the Makefile.
|
||||
version="$("$ROOT/script/version")"
|
||||
commit="$(git rev-parse HEAD 2>/dev/null || echo unknown)"
|
||||
commit_date="$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)"
|
||||
|
||||
epoch="$(date +%s%N)$$"
|
||||
docker build --build-arg CHECK_EPOCH="$epoch" \
|
||||
--build-arg VERSION="$version" \
|
||||
--build-arg COMMIT="$commit" \
|
||||
--build-arg COMMIT_DATE="$commit_date" \
|
||||
.
|
||||
}
|
||||
|
||||
main "$@"
|
||||
|
||||
@@ -2,6 +2,13 @@
|
||||
# script/docker: build the Docker image tagged with the project name.
|
||||
# Identical in all repos; the tag comes from script/projectname.
|
||||
# Generic: needs no adaptation.
|
||||
#
|
||||
# This builds the PRODUCT image only, and the product Dockerfile has no
|
||||
# lint stage: linting lives in Dockerfile.lint and is run by
|
||||
# script/lint. So a green here means `make fmt-check` and `make test`
|
||||
# passed and the image built -- it says nothing about lint. The gates
|
||||
# are script/check (which runs script/lint) and script/cibuild (which
|
||||
# builds both files).
|
||||
set -eu
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
|
||||
@@ -17,7 +24,24 @@ main() {
|
||||
# whether the tree is clean. The Dockerfile now refuses to build
|
||||
# without a non-empty value, so this is required, not optional.
|
||||
epoch="$(date +%s%N)$$"
|
||||
|
||||
# Version, commit and build date are computed here on the host,
|
||||
# where .git exists, and passed into the build. The build context
|
||||
# excludes .git (see .dockerignore), so the container cannot derive
|
||||
# them itself -- it used to try and always got "unknown", giving
|
||||
# every image a "commit: unknown" it could not be traced from.
|
||||
# VERSION comes from script/version, the source of truth shared with
|
||||
# the Makefile, so a Docker build reports the same string (tag,
|
||||
# dev-<sha>, or a -dirty variant) that a local build of the same
|
||||
# tree would.
|
||||
version="$("$SCRIPT_DIR/version")"
|
||||
commit="$(git rev-parse HEAD 2>/dev/null || echo unknown)"
|
||||
commit_date="$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)"
|
||||
|
||||
docker build --build-arg CHECK_EPOCH="$epoch" \
|
||||
--build-arg VERSION="$version" \
|
||||
--build-arg COMMIT="$commit" \
|
||||
--build-arg COMMIT_DATE="$commit_date" \
|
||||
-t "$("$SCRIPT_DIR/projectname")" .
|
||||
}
|
||||
|
||||
|
||||
Executable
+161
@@ -0,0 +1,161 @@
|
||||
#!/bin/sh
|
||||
# script/install-go: install the Go toolchain pinned by go.mod into the
|
||||
# repo-local tool directory, verified against a committed sha256. Our
|
||||
# own extension to scripts-to-rule-them-all. Idempotent: exits at once
|
||||
# when the pinned toolchain is already installed.
|
||||
#
|
||||
# Only .gitea/workflows/release.yml calls this. goreleaser is not a
|
||||
# compiler: it shells out to `go` for the `before:` hook and for every
|
||||
# one of the four cross-compiles, so the release runner needs a Go
|
||||
# toolchain on PATH. check.yml never does -- it builds inside the
|
||||
# digest-pinned Dockerfile images -- so this is the release path's only
|
||||
# host Go, and per REPO_POLICIES.md it must be pinned by hash.
|
||||
# actions/setup-go exposes no checksum input, so Go is installed the way
|
||||
# script/install-goreleaser installs goreleaser: download the exact
|
||||
# archive from go.dev and refuse it unless its sha256 matches the value
|
||||
# committed below.
|
||||
#
|
||||
# The version is go.mod's `go` directive, the single source of truth for
|
||||
# the toolchain. GO_VERSION below MUST equal it, and this script fails
|
||||
# when they disagree -- so bumping Go is one reviewed change touching
|
||||
# go.mod, the checksum here, and the Dockerfile golang digest together.
|
||||
#
|
||||
# Linux only, because that is what the release runner is. A darwin dev
|
||||
# building a snapshot uses their own Go; supporting an OS means adding
|
||||
# its checksums.
|
||||
set -eu
|
||||
|
||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||
|
||||
# Go 1.26.1, 2026-09-21. Checksums are the sha256 values go.dev publishes
|
||||
# for each archive at https://go.dev/dl/ (also in its ?mode=json
|
||||
# manifest).
|
||||
GO_VERSION="1.26.1"
|
||||
SHA256_LINUX_AMD64="031f088e5d955bab8657ede27ad4e3bc5b7c1ba281f05f245bcc304f327c987a"
|
||||
SHA256_LINUX_ARM64="a290581cfe4fe28ddd737dde3095f3dbeb7f2e4065cab4eae44dfc53b760c2f7"
|
||||
|
||||
GOROOT_DIR="$ROOT/.tool/go"
|
||||
GOCMD="$GOROOT_DIR/bin/go"
|
||||
|
||||
# The `go` directive in go.mod, e.g. "1.26.1" from `go 1.26.1`.
|
||||
gomod_go_version() {
|
||||
sed -n 's/^go \([0-9][0-9.]*\).*/\1/p' "$ROOT/go.mod" | head -n 1
|
||||
}
|
||||
|
||||
# Print the version of the go at $1 as "1.26.1", or nothing if it is not
|
||||
# usable. `go version` prints "go version go1.26.1 linux/amd64".
|
||||
go_version() {
|
||||
[ -x "$1" ] || return 0
|
||||
"$1" version 2>/dev/null |
|
||||
sed -n 's/^go version go\([0-9][0-9.]*\) .*/\1/p' |
|
||||
head -n 1
|
||||
}
|
||||
|
||||
verify_sha256() {
|
||||
file="$1"
|
||||
want="$2"
|
||||
if command -v sha256sum >/dev/null 2>&1; then
|
||||
got="$(sha256sum "$file" | cut -d' ' -f1)"
|
||||
elif command -v shasum >/dev/null 2>&1; then
|
||||
got="$(shasum -a 256 "$file" | cut -d' ' -f1)"
|
||||
else
|
||||
echo "install-go: no sha256sum or shasum available" >&2
|
||||
return 1
|
||||
fi
|
||||
if [ "$got" != "$want" ]; then
|
||||
echo "install-go: checksum mismatch for $file" >&2
|
||||
echo " expected: $want" >&2
|
||||
echo " actual: $got" >&2
|
||||
return 1
|
||||
fi
|
||||
}
|
||||
|
||||
# On a Gitea/GitHub Actions runner, put the toolchain on PATH for the
|
||||
# steps that follow by appending to the file named by $GITHUB_PATH. A
|
||||
# no-op off CI, where the caller manages its own PATH.
|
||||
export_ci_path() {
|
||||
[ -n "${GITHUB_PATH:-}" ] || return 0
|
||||
echo "$GOROOT_DIR/bin" >>"$GITHUB_PATH"
|
||||
}
|
||||
|
||||
main() {
|
||||
cd "$ROOT"
|
||||
|
||||
want="$(gomod_go_version)"
|
||||
if [ "$want" != "$GO_VERSION" ]; then
|
||||
echo "install-go: go.mod says go $want but this script pins" \
|
||||
"$GO_VERSION." >&2
|
||||
echo " Update GO_VERSION and the checksums in this script to" \
|
||||
"match go.mod." >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Already installed from a previous run? Then just fix PATH and stop.
|
||||
if [ "$(go_version "$GOCMD")" = "$GO_VERSION" ]; then
|
||||
echo "go $GO_VERSION already installed in .tool/go"
|
||||
export_ci_path
|
||||
return 0
|
||||
fi
|
||||
|
||||
os="$(uname -s)"
|
||||
arch="$(uname -m)"
|
||||
case "$os" in
|
||||
Linux) os="linux" ;;
|
||||
*)
|
||||
echo "install-go: unsupported OS $os (release runner is Linux)" >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
case "$arch" in
|
||||
x86_64 | amd64)
|
||||
arch="amd64"
|
||||
sum="$SHA256_LINUX_AMD64"
|
||||
;;
|
||||
arm64 | aarch64)
|
||||
arch="arm64"
|
||||
sum="$SHA256_LINUX_ARM64"
|
||||
;;
|
||||
*)
|
||||
echo "install-go: no pinned checksum for architecture $arch" >&2
|
||||
exit 1
|
||||
;;
|
||||
esac
|
||||
|
||||
archive="go${GO_VERSION}.${os}-${arch}.tar.gz"
|
||||
url="https://go.dev/dl/${archive}"
|
||||
|
||||
if ! command -v curl >/dev/null 2>&1; then
|
||||
echo "install-go: curl is required" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
dl="$(mktemp -d)"
|
||||
mkdir -p "$ROOT/.tool"
|
||||
stage="$(mktemp -d "$ROOT/.tool/.go-install.XXXXXX")"
|
||||
# shellcheck disable=SC2064 # expand the paths now, not at trap time
|
||||
trap "rm -rf '$dl' '$stage'" EXIT INT TERM
|
||||
|
||||
echo "installing go $GO_VERSION for ${os}-${arch}"
|
||||
curl -fsSL --retry 3 -o "$dl/$archive" "$url"
|
||||
verify_sha256 "$dl/$archive" "$sum"
|
||||
|
||||
# The archive unpacks to a top-level `go/` directory. Extract it into
|
||||
# a staging directory on the same filesystem as the destination, then
|
||||
# rename it into place so a concurrent run never observes a
|
||||
# half-written toolchain.
|
||||
tar -xzf "$dl/$archive" -C "$stage"
|
||||
rm -rf "$GOROOT_DIR"
|
||||
mv "$stage/go" "$GOROOT_DIR"
|
||||
|
||||
installed="$(go_version "$GOCMD")"
|
||||
if [ "$installed" != "$GO_VERSION" ]; then
|
||||
echo "install-go: installed toolchain reports '$installed'," \
|
||||
"expected '$GO_VERSION'" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
echo "go $GO_VERSION installed to .tool/go"
|
||||
export_ci_path
|
||||
}
|
||||
|
||||
main "$@"
|
||||
+64
-291
@@ -1,110 +1,42 @@
|
||||
#!/bin/sh
|
||||
# script/lint: run the linter.
|
||||
#
|
||||
# The linter always runs at the version pinned by the Dockerfile's lint
|
||||
# stage, so a local run and a CI run of the same tree cannot disagree.
|
||||
# That FROM line (image tag plus digest) is the single source of truth
|
||||
# for the linter version in this repo: bump it there and nothing else
|
||||
# needs editing.
|
||||
# The linter runs inside the image built by Dockerfile.lint, and it runs
|
||||
# there as a BUILD STEP: a successful build of that file IS a clean
|
||||
# lint. Nothing lints on the host, at any version, ever. That FROM line
|
||||
# is the single source of truth for the linter version in this repo, so
|
||||
# a local run and a CI run of the same tree cannot disagree.
|
||||
#
|
||||
# Normally that means running the pinned image with docker. The one
|
||||
# exception is running INSIDE that image: the Dockerfile's lint stage
|
||||
# runs `make lint`, and there is no docker daemon in there. That stage
|
||||
# sets VAULTIK_LINT_IN_CONTAINER=1, and only when that variable is set
|
||||
# is a golangci-lint on PATH used directly - and then only if its
|
||||
# version is exactly the pin. Version equality alone is deliberately NOT
|
||||
# enough: it also matches a developer's locally installed copy of the
|
||||
# same version, which is a different build with a different Go
|
||||
# toolchain, reached by a different code path, and it would bypass the
|
||||
# digest pin this script exists to enforce. /.dockerenv was considered
|
||||
# as the context signal and rejected: dockerd creates it for `docker
|
||||
# run`, but it is not reliably present during a BuildKit `docker build`,
|
||||
# which is exactly the case the exception exists for.
|
||||
# One container per run means one lint cache and one golangci-lint lock
|
||||
# per run, both private to that run and thrown away with it. That is
|
||||
# what makes concurrent runs on a shared host safe, and it is why this
|
||||
# script no longer carries per-worktree cache directories, a lock-retry
|
||||
# loop, or an output audit: there is no shared state left for them to
|
||||
# defend (issue https://git.eeqj.de/sneak/vaultik/issues/113).
|
||||
#
|
||||
# The linter's output is checked before it is believed: every run is
|
||||
# audited by script/lint-audit for findings that cannot belong to this
|
||||
# tree, and a run refused by golangci-lint's cross-process lock is
|
||||
# retried rather than reported as a verdict. See the lock-retry loop in
|
||||
# main and the header of script/lint-audit.
|
||||
# To watch the linter execute, set BUILDKIT_PROGRESS=plain, which docker
|
||||
# honours directly:
|
||||
#
|
||||
# Extra arguments are passed through to `golangci-lint run`, before
|
||||
# `./...` (see script/lint-fix).
|
||||
# BUILDKIT_PROGRESS=plain script/lint
|
||||
#
|
||||
# The check layers -- `golangci-lint config verify` and then
|
||||
# `golangci-lint run` -- must appear as executing rather than CACHED on
|
||||
# every run; see the CHECK_EPOCH comment in Dockerfile.lint.
|
||||
set -eu
|
||||
|
||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||
DOCKERFILE="$ROOT/Dockerfile"
|
||||
|
||||
# golangci-lint takes a cross-process lock and refuses to start while
|
||||
# another instance holds it. That refusal is not a lint result, and
|
||||
# exiting non-zero on it is indistinguishable to a caller from real
|
||||
# findings - so it is retried rather than reported. Bounded, because a
|
||||
# lock that is never released must fail rather than hang.
|
||||
LOCK_MESSAGE="parallel golangci-lint is running"
|
||||
LOCK_ATTEMPTS=6
|
||||
LOCK_SLEEP=15
|
||||
|
||||
# The image reference of the Dockerfile's lint stage, tag and digest
|
||||
# included, e.g.
|
||||
# golangci/golangci-lint:v2.12.2-alpine@sha256:91b2...
|
||||
lint_image() {
|
||||
awk '$1 == "FROM" && $3 == "AS" && $4 == "lint" { print $2; exit }' \
|
||||
"$DOCKERFILE"
|
||||
}
|
||||
|
||||
# The bare version that image reference pins, e.g. 2.12.2
|
||||
pinned_version() {
|
||||
lint_image | sed -e 's/@.*//' -e 's/.*://' -e 's/^v//' -e 's/-.*//'
|
||||
}
|
||||
|
||||
# The version of the golangci-lint on PATH, if any, e.g. 2.12.2
|
||||
#
|
||||
# `version --short` prints the bare version and is the interface meant
|
||||
# for this (checked against 2.10.1 and 2.12.2). The banner scrape below
|
||||
# it is a fallback for a release where --short is absent or silent; the
|
||||
# banner's exact wording is not a stable interface, which is why it is
|
||||
# no longer the primary parse.
|
||||
installed_version() {
|
||||
command -v golangci-lint >/dev/null 2>&1 || return 0
|
||||
|
||||
short="$(golangci-lint version --short 2>/dev/null |
|
||||
tr -d '[:space:]' | sed -e 's/^v//')"
|
||||
case "$short" in
|
||||
*[0-9].[0-9]*.[0-9]*)
|
||||
echo "$short"
|
||||
return 0
|
||||
;;
|
||||
esac
|
||||
|
||||
golangci-lint version 2>/dev/null | awk '
|
||||
{
|
||||
for (i = 1; i <= NF; i++) {
|
||||
if ($i ~ /^[0-9]+\.[0-9]+\.[0-9]+$/) {
|
||||
print $i
|
||||
exit
|
||||
}
|
||||
}
|
||||
}'
|
||||
}
|
||||
|
||||
# True inside the Dockerfile's lint stage, which sets this. Nothing else
|
||||
# sets it: setting it by hand on a host is an explicit, visible decision
|
||||
# to lint with an unpinned binary, not something reached by accident.
|
||||
in_lint_container() {
|
||||
[ "${VAULTIK_LINT_IN_CONTAINER:-}" = "1" ]
|
||||
}
|
||||
DOCKERFILE="$ROOT/Dockerfile.lint"
|
||||
|
||||
require_docker() {
|
||||
image="$1"
|
||||
if ! command -v docker >/dev/null 2>&1; then
|
||||
cat >&2 <<EOF
|
||||
lint: docker is required to run the pinned linter.
|
||||
|
||||
pinned image: $image
|
||||
lint image declared by: $DOCKERFILE
|
||||
|
||||
Install docker. Linting with any other golangci-lint is not supported:
|
||||
it is what lets a local run pass while CI fails. An installed
|
||||
golangci-lint on PATH is not used, whatever its version; only the lint
|
||||
stage of the Dockerfile itself runs the linter natively.
|
||||
it is what lets a local run pass while CI fails. A golangci-lint on
|
||||
PATH is never used, whatever its version.
|
||||
EOF
|
||||
exit 1
|
||||
fi
|
||||
@@ -113,7 +45,7 @@ EOF
|
||||
lint: the docker daemon is not reachable, so the pinned linter cannot
|
||||
run.
|
||||
|
||||
pinned image: $image
|
||||
lint image declared by: $DOCKERFILE
|
||||
|
||||
Start the daemon (and check DOCKER_HOST / your group membership). This
|
||||
script will not fall back to a different linter version or to an
|
||||
@@ -123,213 +55,54 @@ EOF
|
||||
fi
|
||||
}
|
||||
|
||||
# Where the per-worktree caches live.
|
||||
cache_home() {
|
||||
echo "${XDG_CACHE_HOME:-${HOME:-/tmp}/.cache}/vaultik-lint"
|
||||
}
|
||||
usage() {
|
||||
cat >&2 <<EOF
|
||||
usage: $(basename "$0")
|
||||
|
||||
# A short, stable digest of this worktree's path.
|
||||
path_digest() {
|
||||
if command -v sha256sum >/dev/null 2>&1; then
|
||||
printf '%s' "$ROOT" | sha256sum | cut -c1-12
|
||||
elif command -v shasum >/dev/null 2>&1; then
|
||||
printf '%s' "$ROOT" | shasum -a 256 | cut -c1-12
|
||||
else
|
||||
printf '%s' "$ROOT" | cksum | tr -cd '0-9' | cut -c1-12
|
||||
fi
|
||||
}
|
||||
|
||||
# Caches for the containerized linter, private to THIS worktree.
|
||||
#
|
||||
# Keeping them out of the repo and persisting them between runs is what
|
||||
# keeps the inner loop fast: a warm run costs about the same as a native
|
||||
# one plus container startup. Keeping them keyed on the worktree path is
|
||||
# what keeps them correct. One shared cache for the whole repo was the
|
||||
# defect in issue #99: two worktrees of this repo have identical file
|
||||
# contents, so their cache keys collide, and golangci-lint replays the
|
||||
# stored results - including the file paths recorded when they were
|
||||
# produced. That silently reports one worktree's findings, or one
|
||||
# worktree's clean bill of health, for another.
|
||||
cache_dir() {
|
||||
slug="$(printf '%s' "$(basename "$ROOT")" | tr -c 'A-Za-z0-9._-' '-')"
|
||||
echo "$(cache_home)/$slug-$(path_digest)"
|
||||
}
|
||||
|
||||
# One cache per worktree means throwaway worktrees would otherwise leave
|
||||
# caches behind forever. Each cache records the worktree it belongs to,
|
||||
# and any cache whose worktree no longer exists is collected here, so
|
||||
# growth is bounded by the number of worktrees that actually exist. The
|
||||
# whole tree also sits under XDG_CACHE_HOME (~/.cache by default), so it
|
||||
# is disposable by definition: `rm -rf "${XDG_CACHE_HOME:-~/.cache}/vaultik-lint"`
|
||||
# costs nothing but the next run's cold cache.
|
||||
prune_dead_caches() {
|
||||
home="$(cache_home)"
|
||||
if [ ! -d "$home" ]; then
|
||||
return 0
|
||||
fi
|
||||
for dir in "$home"/*; do
|
||||
if [ ! -f "$dir/worktree" ]; then
|
||||
continue
|
||||
fi
|
||||
owner="$(cat "$dir/worktree")"
|
||||
if [ -z "$owner" ]; then
|
||||
continue
|
||||
fi
|
||||
if [ ! -d "$owner" ]; then
|
||||
# The Go module cache inside is deliberately read-only, and
|
||||
# rm(1) cannot unlink a file out of a directory it may not
|
||||
# write, so the tree has to be made writable first. And
|
||||
# failing to tidy up is a housekeeping problem, never a
|
||||
# reason to fail a lint: without the fallback below, `set
|
||||
# -e` turns a stale cache that will not delete into a lint
|
||||
# error, which is a gate failing for a reason that has
|
||||
# nothing to do with the code. (Observed, not theorised.)
|
||||
chmod -R u+w "$dir" 2>/dev/null || true
|
||||
if ! rm -rf "$dir" 2>/dev/null; then
|
||||
# Restore the marker on a partial removal: an
|
||||
# unmarked leftover would be skipped by every future
|
||||
# run and never collected.
|
||||
mkdir -p "$dir" 2>/dev/null || true
|
||||
echo "$owner" >"$dir/worktree" 2>/dev/null || true
|
||||
echo "lint: could not remove stale cache $dir" >&2
|
||||
fi
|
||||
fi
|
||||
done
|
||||
}
|
||||
|
||||
prepare_cache() {
|
||||
cache="$1"
|
||||
mkdir -p "$cache/go-build" "$cache/go-mod" "$cache/golangci-lint"
|
||||
echo "$ROOT" >"$cache/worktree"
|
||||
}
|
||||
|
||||
# Run the linter, wherever it is that this script is allowed to run it.
|
||||
run_linter() {
|
||||
if in_lint_container; then
|
||||
golangci-lint run "$@" ./...
|
||||
return $?
|
||||
fi
|
||||
|
||||
docker run --rm \
|
||||
--user "$(id -u):$(id -g)" \
|
||||
--env HOME=/tmp \
|
||||
--env GOFLAGS=-buildvcs=false \
|
||||
--env GOCACHE=/cache/go-build \
|
||||
--env GOMODCACHE=/cache/go-mod \
|
||||
--env GOLANGCI_LINT_CACHE=/cache/golangci-lint \
|
||||
--volume "$ROOT:/src" \
|
||||
--volume "$CACHE:/cache" \
|
||||
--workdir /src \
|
||||
"$IMAGE" \
|
||||
golangci-lint run "$@" ./...
|
||||
}
|
||||
|
||||
# Run the linter, streaming its combined output while also capturing it,
|
||||
# and hand back its exit status. The output has to be inspected before
|
||||
# it is believed, which is why this script no longer just execs the
|
||||
# linter. `tee` would swallow the status, so it is smuggled out through
|
||||
# a file: there is no pipefail in POSIX sh.
|
||||
run_capture() {
|
||||
capture="$1"
|
||||
shift
|
||||
rm -f "$capture.status"
|
||||
{
|
||||
rc=0
|
||||
# `set -e` is in force inside this subshell too, so the status
|
||||
# has to be caught here: an unguarded non-zero exit (which is
|
||||
# what "the linter found something" looks like) would abort the
|
||||
# subshell before the status was ever written.
|
||||
run_linter "$@" 2>&1 || rc=$?
|
||||
echo "$rc" >"$capture.status"
|
||||
} | tee "$capture"
|
||||
|
||||
if [ ! -s "$capture.status" ]; then
|
||||
echo "lint: the linter did not report an exit status" >&2
|
||||
exit 1
|
||||
fi
|
||||
read -r captured_status <"$capture.status"
|
||||
rm -f "$capture.status"
|
||||
return "$captured_status"
|
||||
}
|
||||
|
||||
# Reject output that cannot describe this tree. See script/lint-audit
|
||||
# for what that means and why: in short, a finding citing a file that is
|
||||
# not here means the result being reported was produced somewhere else,
|
||||
# and a PASS built out of another checkout's analysis is silent (issue
|
||||
# #99). The audit therefore runs on clean output as well.
|
||||
audit_output() {
|
||||
capture="$1"
|
||||
if ! "$ROOT/script/lint-audit" "$capture"; then
|
||||
if [ -n "$CACHE" ]; then
|
||||
echo " this tree's lint cache: $CACHE" >&2
|
||||
fi
|
||||
exit 1
|
||||
fi
|
||||
script/lint takes no arguments. The linter runs as a build step, so
|
||||
there is no command line to pass flags to; anything accepted here would
|
||||
have to be silently dropped. To apply autofixes, use script/lint-fix,
|
||||
which runs the same pinned image as a container for exactly this
|
||||
reason.
|
||||
EOF
|
||||
exit 2
|
||||
}
|
||||
|
||||
main() {
|
||||
[ "$#" -eq 0 ] || usage
|
||||
|
||||
cd "$ROOT"
|
||||
require_docker
|
||||
|
||||
IMAGE="$(lint_image)"
|
||||
if [ -z "$IMAGE" ]; then
|
||||
echo "lint: no lint stage found in $DOCKERFILE" >&2
|
||||
exit 1
|
||||
fi
|
||||
# A fresh epoch per invocation is what forces the check layers to
|
||||
# execute; the layers above the ARG in Dockerfile.lint still cache,
|
||||
# so a run is not cold. The value must be unique per invocation, not
|
||||
# per second: `date +%s` is second-granular, so two concurrent
|
||||
# invocations in the same second would get identical epochs and the
|
||||
# later one could be served from cache -- the false green in
|
||||
# miniature. `%N` alone does not fix it either, because busybox
|
||||
# silently drops %N, exits 0, and hands back second granularity with
|
||||
# no warning. `$$` is what makes this correct regardless, since
|
||||
# concurrent invocations have different pids.
|
||||
#
|
||||
# Assign it on its own line rather than inline in the argument.
|
||||
# Under `set -eu` a command substitution that fails inside an
|
||||
# argument does NOT abort the script: CHECK_EPOCH would become an
|
||||
# empty string, an empty string is a constant, and a constant epoch
|
||||
# is exactly the cached-lint false green this guards against. As a
|
||||
# bare assignment, `set -e` catches a failing `date` and no build
|
||||
# starts.
|
||||
epoch="$(date +%s%N)$$"
|
||||
|
||||
CACHE=""
|
||||
if in_lint_container; then
|
||||
# No docker daemon in here, so there is no fallback: a mismatch
|
||||
# is a hard error rather than a quiet substitution.
|
||||
installed="$(installed_version)"
|
||||
pinned="$(pinned_version)"
|
||||
if [ -z "$installed" ] || [ "$installed" != "$pinned" ]; then
|
||||
cat >&2 <<EOF
|
||||
lint: VAULTIK_LINT_IN_CONTAINER is set, so this is expected to be
|
||||
running inside the Dockerfile's pinned lint image, but the golangci-lint
|
||||
on PATH does not match the pin.
|
||||
|
||||
pinned: $pinned ($IMAGE)
|
||||
installed: ${installed:-<none>}
|
||||
EOF
|
||||
exit 1
|
||||
fi
|
||||
else
|
||||
require_docker "$IMAGE"
|
||||
prune_dead_caches
|
||||
CACHE="$(cache_dir)"
|
||||
prepare_cache "$CACHE"
|
||||
fi
|
||||
|
||||
capture="$(mktemp "${TMPDIR:-/tmp}/vaultik-lint.XXXXXX")"
|
||||
trap 'rm -f "$capture" "$capture.status"' EXIT HUP INT TERM
|
||||
|
||||
attempt=1
|
||||
while :; do
|
||||
status=0
|
||||
run_capture "$capture" "$@" || status=$?
|
||||
|
||||
if grep -Fq "$LOCK_MESSAGE" "$capture"; then
|
||||
if [ "$attempt" -lt "$LOCK_ATTEMPTS" ]; then
|
||||
echo "lint: another golangci-lint holds the lock;" \
|
||||
"retrying in ${LOCK_SLEEP}s" \
|
||||
"(attempt $attempt of $LOCK_ATTEMPTS)" >&2
|
||||
sleep "$LOCK_SLEEP"
|
||||
attempt=$((attempt + 1))
|
||||
continue
|
||||
fi
|
||||
cat >&2 <<EOF
|
||||
|
||||
lint: gave up after $LOCK_ATTEMPTS attempts, each blocked by another
|
||||
golangci-lint holding the cross-process lock. This is NOT a lint
|
||||
verdict: the tree was never analysed. Re-run when the other run has
|
||||
finished.
|
||||
EOF
|
||||
exit 1
|
||||
fi
|
||||
|
||||
audit_output "$capture"
|
||||
exit "$status"
|
||||
done
|
||||
# cacheonly: the lint verdict is the build's exit status, and the
|
||||
# image it would otherwise produce is never run. Exporting it costs
|
||||
# most of the wall time of a warm run and leaves a dangling image
|
||||
# behind on every invocation, on a host that may be running many.
|
||||
docker build \
|
||||
--output=type=cacheonly \
|
||||
--build-arg CHECK_EPOCH="$epoch" \
|
||||
-f "$DOCKERFILE" \
|
||||
"$ROOT"
|
||||
}
|
||||
|
||||
main "$@"
|
||||
|
||||
@@ -1,113 +0,0 @@
|
||||
#!/bin/sh
|
||||
# script/lint-audit: audit a captured golangci-lint run for output that
|
||||
# cannot describe this tree. Called by script/lint on every run; usable
|
||||
# on its own against any saved lint output.
|
||||
#
|
||||
# script/lint-audit <capture-file>
|
||||
#
|
||||
# Exits 0 when every finding cites a file in this tree, 1 when any does
|
||||
# not. It NEVER certifies that a lint run passed - it has no idea
|
||||
# whether the run found issues, and does not look. It only rejects
|
||||
# output that is impossible for this tree, which is a different and much
|
||||
# weaker claim. Do not use it as a gate; use script/lint.
|
||||
#
|
||||
# Why this exists (issue #99): golangci-lint caches analysis results,
|
||||
# and a cache shared between two checkouts of this repo can serve one
|
||||
# checkout's stored findings for another, file paths included. The
|
||||
# failure is symmetric and only one direction is loud - a clean tree
|
||||
# failed by a dirty sibling gets investigated, while a dirty tree passed
|
||||
# by a clean sibling is silent. This turns the silent direction into a
|
||||
# hard error, which is why it runs on clean output too.
|
||||
#
|
||||
# The primary fix is that script/lint now keys its cache on the worktree
|
||||
# path so the collision cannot happen. This is the backstop, because a
|
||||
# backstop that only runs when we already believe things are fine is
|
||||
# worth more than one more assumption.
|
||||
set -eu
|
||||
|
||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||
|
||||
# Where script/lint bind-mounts the tree inside the pinned image. A
|
||||
# containerized run that prints absolute paths (`--path-mode abs`)
|
||||
# prints them under this, so they are this tree's files under another
|
||||
# name. Note the consequence, and why the cache key rather than this
|
||||
# check is the real fix: two containerized runs of different checkouts
|
||||
# both call themselves /src, so contamination between two container
|
||||
# runs is not distinguishable by path alone.
|
||||
CONTAINER_ROOT="/src"
|
||||
|
||||
usage() {
|
||||
echo "usage: $(basename "$0") <capture-file>" >&2
|
||||
exit 2
|
||||
}
|
||||
|
||||
# Every path cited by a finding that is not a file in this tree.
|
||||
#
|
||||
# The linter runs with the tree root as its working directory, so a
|
||||
# legitimate finding cites either a relative path that resolves inside
|
||||
# the tree or an absolute path under the root. A path that escapes
|
||||
# (absolute and elsewhere, or with a `..` component) or that names a
|
||||
# file which is not here describes something this run did not analyse.
|
||||
foreign_paths() {
|
||||
capture="$1"
|
||||
awk -F: '$1 ~ /\.go$/ && $2 ~ /^[0-9]+$/ { print $1 }' "$capture" |
|
||||
sort -u |
|
||||
while IFS= read -r path; do
|
||||
case "$path" in
|
||||
"$ROOT"/*)
|
||||
path="${path#"$ROOT"/}"
|
||||
;;
|
||||
"$CONTAINER_ROOT"/*)
|
||||
path="${path#"$CONTAINER_ROOT"/}"
|
||||
;;
|
||||
/*)
|
||||
printf '%s\n' "$path"
|
||||
continue
|
||||
;;
|
||||
../* | */../*)
|
||||
printf '%s\n' "$path"
|
||||
continue
|
||||
;;
|
||||
esac
|
||||
if [ ! -e "$ROOT/$path" ]; then
|
||||
printf '%s\n' "$path"
|
||||
fi
|
||||
done
|
||||
}
|
||||
|
||||
main() {
|
||||
[ "$#" -eq 1 ] || usage
|
||||
capture="$1"
|
||||
if [ ! -f "$capture" ]; then
|
||||
echo "lint-audit: no such capture file: $capture" >&2
|
||||
exit 2
|
||||
fi
|
||||
|
||||
foreign="$(foreign_paths "$capture")"
|
||||
if [ -z "$foreign" ]; then
|
||||
exit 0
|
||||
fi
|
||||
|
||||
cat >&2 <<EOF
|
||||
|
||||
lint: REJECTED - the linter reported findings for files that are not in
|
||||
this tree, so its output does not describe the tree that was linted.
|
||||
This result is void, whichever way it went: a pass here would be a pass
|
||||
earned by analysing someone else's code.
|
||||
|
||||
tree: $ROOT
|
||||
|
||||
Paths reported that are not in this tree:
|
||||
EOF
|
||||
printf '%s\n' "$foreign" | sed -e 's/^/ /' >&2
|
||||
cat >&2 <<EOF
|
||||
|
||||
This is the signature of analysis replayed from a cache belonging to
|
||||
another checkout (issue #99). Clear this tree's lint cache and re-run:
|
||||
|
||||
rm -rf "\${XDG_CACHE_HOME:-\$HOME/.cache}/vaultik-lint"
|
||||
EOF
|
||||
exit 1
|
||||
}
|
||||
|
||||
main "$@"
|
||||
+43
-7
@@ -1,18 +1,54 @@
|
||||
#!/bin/sh
|
||||
# script/lint-fix: run the linter's autofixer. Rewrites files in place
|
||||
# for every finding the enabled linters know how to fix; findings
|
||||
# without an autofix are reported but left alone (exit status is
|
||||
# nonzero while any remain).
|
||||
# without an autofix are reported but left alone.
|
||||
#
|
||||
# Delegates to script/lint so the autofixer is the same pinned linter
|
||||
# version that script/lint and CI use - fixes written by a different
|
||||
# version are not necessarily fixes for the version that gates.
|
||||
# THIS IS A DEVELOPER CONVENIENCE AND NEVER A GATE. Nothing in
|
||||
# script/check, script/precommit or script/cibuild calls it, and no gate
|
||||
# reads its exit status. The gate is script/lint, which builds
|
||||
# Dockerfile.lint; run that afterwards to find out whether the tree is
|
||||
# actually clean.
|
||||
#
|
||||
# Unlike script/lint this cannot be a build step: a build step writes
|
||||
# into an image, and fixes have to land in the worktree. So it runs the
|
||||
# same pinned image as a container with the tree bind-mounted, which
|
||||
# means it needs a LOCAL docker daemon -- a remote daemon has no access
|
||||
# to these files, and this script will appear to do nothing there. The
|
||||
# image reference is parsed out of Dockerfile.lint's FROM line, so the
|
||||
# autofixer is always the same version as the linter that gates; fixes
|
||||
# written by a different version are not necessarily fixes for the
|
||||
# version that decides.
|
||||
set -eu
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
|
||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||
DOCKERFILE="$ROOT/Dockerfile.lint"
|
||||
|
||||
# The image reference from Dockerfile.lint, tag and digest included.
|
||||
lint_image() {
|
||||
awk '$1 == "FROM" { print $2; exit }' "$DOCKERFILE"
|
||||
}
|
||||
|
||||
main() {
|
||||
exec "$SCRIPT_DIR/lint" --fix "$@"
|
||||
cd "$ROOT"
|
||||
|
||||
image="$(lint_image)"
|
||||
if [ -z "$image" ]; then
|
||||
echo "lint-fix: no FROM line found in $DOCKERFILE" >&2
|
||||
exit 1
|
||||
fi
|
||||
|
||||
# Run as the invoking user so the rewritten files stay owned by
|
||||
# them. HOME is set because the Go and golangci-lint caches default
|
||||
# under it and that user has no home inside the container; those
|
||||
# caches are per-container and discarded with it.
|
||||
docker run --rm \
|
||||
--user "$(id -u):$(id -g)" \
|
||||
--env HOME=/tmp \
|
||||
--env GOFLAGS=-buildvcs=false \
|
||||
--volume "$ROOT:/src" \
|
||||
--workdir /src \
|
||||
"$image" \
|
||||
golangci-lint run --config .golangci.yml --fix "$@" ./...
|
||||
}
|
||||
|
||||
main "$@"
|
||||
|
||||
@@ -1,27 +0,0 @@
|
||||
# Vaultik test configuration
|
||||
hostname: test-host
|
||||
index_path: /tmp/vaultik-test/index.db
|
||||
source_dirs:
|
||||
- /tmp/vaultik-test/source
|
||||
|
||||
# S3 configuration
|
||||
s3:
|
||||
endpoint: http://localhost:19000 # gofakes3 test endpoint
|
||||
bucket: test-bucket
|
||||
prefix: test-
|
||||
access_key_id: test-key
|
||||
secret_access_key: test-secret
|
||||
region: us-east-1
|
||||
|
||||
# Chunking configuration
|
||||
chunk_size: 65536 # 64KB average chunk size
|
||||
min_chunk_size: 32768 # 32KB minimum
|
||||
max_chunk_size: 131072 # 128KB maximum
|
||||
blob_size: 1048576 # 1MB blobs for testing
|
||||
|
||||
# Compression
|
||||
compression_level: 3
|
||||
|
||||
# Encryption
|
||||
# age_recipients:
|
||||
# - age1qyqszqgpqyqszqgpqyqszqgpqyqszqgpqyqszqgpqyqszqgpqyqs3mw88h
|
||||
@@ -1,24 +0,0 @@
|
||||
age_recipients:
|
||||
- age1278m9q7dp3chsh2dcy82qk27v047zywyvtxwnj4cvt0z65jw6a7q5dqhfj # sneak's long term age key
|
||||
- age1ezrjmfpwsc95svdg0y54mums3zevgzu0x0ecq2f7tp8a05gl0sjq9q9wjg # insecure integration test key
|
||||
source_dirs:
|
||||
- /tmp/vaultik-test-source
|
||||
exclude:
|
||||
- '*.log'
|
||||
- '*.tmp'
|
||||
- '.git'
|
||||
- 'node_modules'
|
||||
s3:
|
||||
endpoint: http://ber1app1.local:3900/
|
||||
bucket: vaultik-integration-test
|
||||
prefix: test-host/
|
||||
access_key_id: GKbc8e6d35fdf50847f155aca5
|
||||
secret_access_key: 217046bee47c050301e3cc13e3cba1a8a943cf5f37f8c7979c349c5254441d18
|
||||
region: us-east-1
|
||||
use_ssl: false
|
||||
part_size: 5242880 # 5MB
|
||||
index_path: /tmp/vaultik-integration-test.sqlite
|
||||
chunk_size: 10MB
|
||||
blob_size_limit: 10GB
|
||||
compression_level: 3
|
||||
hostname: test-host
|
||||
Reference in New Issue
Block a user