23 Commits
Author SHA1 Message Date
clawbot d53202eb86 Fix two misleading messages (closes #240)
check / check (push) Waiting to run
The warning for a config file that others can read always said the
file contained S3 credentials, so a file:// config with none got a
false claim. When s3.access_key_id or s3.secret_access_key is set it
now says the file may contain them, because Load sees the values only
after smartconfig has replaced any ${...} reference, so a set
credential need not be in the file. Otherwise it says the file is
readable by others.

snapshot purge wrapped the listing error, which already starts with
"listing remote snapshots:", in that prefix a second time.
syncWithRemote now returns it unwrapped, as CleanupLocalSnapshots
does.

Model: opus-5-5
2026-10-07 15:29:07 +02:00
clawbot 8b22ae8d42 Pass s3.part_size to the multipart uploader (closes #232)
check / check (push) Waiting to run
s3.part_size was loaded and defaulted but never reached the S3 client,
whose uploader used a fixed 10MiB part. The client now takes the part
size from the config, for storage_url and for the s3.* fields, and
config load rejects a value below 5MiB or above 5GiB, an explicit 0
included. A blob too large for S3's limit of 10,000 parts at that size
is uploaded in larger parts, since the uploader cannot learn the size of
the reader it is given. The docs gave the default as 5MB, which the
config file reads as 5,000,000 bytes, below the minimum; they now say
5MiB.

Judgement call: the 5GiB maximum is enforced along with the 5MiB
minimum the issue names.

Model: opus-5-5
2026-10-07 14:29:10 +02:00
clawbot 3fc8a8f2f4 Read the snapshot name using the stored hostname (closes #230)
check / check (push) Waiting to run
A snapshot ID is hostname_name_timestamp, and purge took the name to be
everything between the first and the last underscore. With a hostname
such as my_host the name home came out as host_home, so
`snapshot purge --keep-latest --snapshot home` found nothing to delete
and `snapshot create --prune` purged nothing without a message. The name
is now read by removing the hostname stored with the snapshot, in the
short form the ID uses, so both may contain underscores. This was chosen
over rejecting underscores in `hostname` when the config loads, which
would also stop restores on such a host.

The purge consistency test stored a hostname that did not match its
snapshot IDs; it now matches, as it always does in production.

Model: opus-5-5
2026-10-07 12:12:08 +02:00
clawbot 7696f83258 Leave remote info orphan figures unknown when a manifest is unreadable (closes #228)
check / check (push) Waiting to run
When a manifest could not be read, remote info skipped it, counted that
snapshot's blobs as orphaned and advised running prune. The orphan
figures are now unknown in that case, with no prune advice; --json gives
them as null and lists the unreadable remote keys in
unreadable_manifests. Only a listed manifest.json.zst is read, so a
directory without one, as an interrupted backup leaves, keeps the
figures known.

Names under metadata/ were used unchecked and printed raw. A name that
is not a remote key is now skipped with a warning. A manifest under it
is not read either, so it also leaves the figures unknown; --json counts
such manifests in skipped_manifest_count.

Model: opus-5-5
2026-10-07 10:59:26 +02:00
clawbot 5d1118d143 Quote a string setting that YAML would read as a number (closes #229)
check / check (push) Waiting to run
config set wrote every value as an unquoted YAML scalar, and config.Load
reads the file through untyped YAML, so an access key 00112233 loaded as
38043 and a hostname 007 as 7.

config set now looks the key up in config.Config by the fields' yaml
tags. A string setting is tagged !!str, which the encoder quotes
wherever YAML would read a number or a boolean. Other settings stay
unquoted, so compression_level 9 is still a number. A value that is not
valid UTF-8 stays untagged and is written as !!binary, which loads back
unchanged.

Judgement call: the type comes from reflection over config.Config.

Model: opus-5-5
2026-10-07 09:29:19 +02:00
clawbot b57ce2277d Store and compare file mtimes to the nanosecond (closes #226)
check / check (push) Waiting to run
The files table held mtime in whole seconds and the scanner compared
whole seconds. A file rewritten with its size unchanged and a new mtime
in the same second as the indexed one was treated as unchanged, and
every later snapshot restored the old content. A new mtime_nsec column
now holds the nanoseconds within the second that mtime holds, and the
scanner compares the full mtime.

A local index created before this change lacks the column and is
rebuilt with `vaultik database delete` and a full backup. A snapshot
made before it cannot be restored by this version.

Model: opus-5-5
2026-10-07 07:12:12 +02:00
clawbot 8496404d8b Make the process-wide lock atomic with flock (closes #227)
check / check (push) Waiting to run
Acquire read vaultik.pid, checked whether that PID was alive, then
wrote its own, so two writers started together could both pass the
check and both run. The lock is now an flock on vaultik.pid, held
while the file stays open; the kernel drops it when the process exits,
so the stale-PID check is gone.

Release empties the file instead of deleting it. Deleting it would let
a process that opened the old file a moment earlier lock it while
another creates and locks a new one.

The new concurrent test fails against the old code only when the race
is hit, not on every run; against the fix it cannot admit two callers.

Model: opus-5-5
2026-10-07 06:12:10 +02:00
clawbot cdc60c4dfa Write only the document to stdout from snapshot remove --json (closes #251)
check / check (push) Waiting to run
When the destination store cannot be reached, `snapshot remove` still
removes the snapshot from the local index and warns. The warning went
through the UI, which writes to stdout, so under `--json` it landed
ahead of the document and `| jq` failed on a command that exited 0.
Under `--json` the UI warning is now skipped; the logger's warning on
stderr carries the follow-up, with the snapshot ID as a field.

The follow-up, in the warning, the README and the command's help, said
`vaultik prune` would finish the cleanup, but `prune` never removes
snapshot metadata. They now say to run `vaultik snapshot remove` for the
snapshot again once the destination store is reachable.

Model: opus-5-5
2026-10-07 04:46:13 +02:00
clawbot 49eed7a3e5 Count each file, byte and upload once in backup statistics (closes #225)
check / check (push) Waiting to run
The scanner added a changed file's bytes again for each new chunk and
counted a file as unchanged for each chunk already stored. It now
counts files and bytes once, in the scan phase, and counts its own
uploads, so a --cron run, which has no progress reporter, records
them. The blob count no longer adds earlier paths' blobs again.

The snapshots row now stores the size of all files in total_size and
the referenced blobs' sizes in blob_size, blob_uncompressed_size and
compression_ratio, as docs/DATAMODEL.md says. Those sizes come from
one query, and a failed query fails the snapshot. DATAMODEL.md now
says chunk_count and blob_count count what the run added.

Removed UpdateSnapshotStats and GetCountBySnapshot, which nothing
calls any more.

Model: opus-5-5
2026-10-07 02:59:27 +02:00
clawbot 85d4ef118d Keep command output to the README's stdout and stderr rules (closes #224)
check / check (push) Successful in 16m49s
The startup banner moves from stdout to stderr, so a `completion`
script, a `config get` value and the hidden `__complete` command print
only their own output. `--quiet`, `--cron` and `--json` still suppress
it.

A failing `remote info`, `prune` or `snapshot remove` under `--json`
now reports its error on stderr. Their reporters returned early under
`--json`, so the failure reached neither stream.

`snapshot verify --quiet` writes no report. A failure is still
returned and printed on stderr, with the same exit status.

Judgement call: the banner's stream, posted on the issue for the owner.

Model: opus-5-5
2026-10-07 01:12:15 +02:00
clawbot b06f992152 Start the progress reporter once per snapshot, not once per path (closes #253)
check / check (push) Successful in 19m54s
A backup without --cron of a snapshot with two or more paths panicked
with "close of closed channel". Scan runs once per path, and it started
the progress reporter and deferred its Stop each time; Stop closes the
reporter's signal channel, so the second path's Stop panicked.
scanAllDirectories now starts the reporter before the first path and
stops it after the last, and Scan no longer starts or stops it. A new
test backs up a two-path snapshot with the reporter on and restores
both paths.

Model: opus-5-5
2026-10-07 00:12:10 +02:00
clawbot 5d685f03ce List the files under a path by path, not by string prefix (closes #223)
check / check (push) Successful in 12m16s
FileRepository.ListByPrefix matched with SQL LIKE: a plain string
prefix that ignores ASCII case and treats _ and % as wildcards.
Restoring /home/u/doc also restored doc2, DOC and doc.txt.bak, and a
backup counted the files of a longer sibling path as deleted. It is
now ListUnderPath, which returns the file at the path and every file
whose path starts with the path plus a slash, compared exactly. A
trailing slash is ignored, so "/" still lists every file.

Three tests relied on string-prefix matching and now name a full path
or call ListAll.

ListIDsWithChunksNotInUploadedBlobs keeps its LIKE: it only adds file
IDs the scan never looks up.

Model: opus-5-5
2026-10-06 22:12:16 +02:00
clawbot f59086c0e5 Return an error, not a panic, on a malformed snapshot database (closes #231)
check / check (push) Successful in 13m39s
Restore cut chunk hashes from the snapshot database to 16 characters
for its error messages, so a shorter hash panicked. Those messages now
use shortHash. Under --verify, a file_chunks row with no chunks row was
dereferenced, and the chunk size from the database was allocated in
one piece, so a negative or huge size panicked. A missing row is now an
error, a negative size is rejected, and each chunk is hashed by
streaming it from the restored file.

A restored file shorter than its chunks now fails verify as a short
read instead of an unexpected EOF.

Model: opus-5-5
2026-10-06 21:16:28 +02:00
clawbot d276d891da Join the S3 prefix to every key with one slash (closes #222)
check / check (push) Successful in 10m59s
The S3 client built each key as prefix + key, and the URL parser keeps
the prefix as written, so s3://bucket/p stored p + "blobs/..." with no
slash while s3://bucket/p/ stored p/blobs/.... A recovery host that wrote
the URL the other way found no snapshots.

NewClient now strips trailing slashes from the prefix and adds one back
when anything is left, giving the README layout for both URL forms; an
empty prefix stays at the bucket root. The s3.prefix config setting
goes through the same client and gets the same join.

A new test writes through each URL shape against an in-process S3
server, checks the key in the bucket, and lists through both List and
ListStream, which every snapshot listing uses.

Model: opus-5-5
2026-10-06 19:46:09 +02:00
clawbot 5aa5ba5544 Apply restored owners, modes and times in an order that keeps them (closes #219)
check / check (push) Successful in 11m53s
A directory got its stored mode and mtime before its contents were
written, so a read-only directory came back without its files and a
non-empty one carried the time of the restore. Directories are now
created owner-only (0700) and get their stored owner, mode and mtime
after the restore loop, each before its parent, skipping any whose
place a symlink has since taken. A file's mode is now applied after its
chown, which on Linux clears setuid and setgid. A symlink gets its
stored owner (as root) and mtime on the link itself, through
golang.org/x/sys/unix, now a direct dependency.

An interrupted restore leaves its directories at 0700.

Model: opus-5-5
2026-10-06 18:29:09 +02:00
clawbot 315b6483b8 Mark github.com/spf13/pflag as a direct dependency in go.mod (closes #246)
check / check (push) Successful in 12m34s
internal/cli/snapshot_restore_test.go imports github.com/spf13/pflag
directly, but go.mod still marked it // indirect. script/precommit runs
go mod tidy and fails when that changes go.mod, so the pre-commit hook
stopped every commit. This is the go mod tidy output: pflag moves to the
direct require block, and go.sum does not change. script/cibuild does
not run the tidy, which is why the gate stayed green.

Model: opus-5-5
2026-10-06 16:59:26 +02:00
clawbot 14fc4c9893 Load a config that has no age recipient (closes #221)
check / check (push) Successful in 13m47s
The README's steps for restoring on another machine failed at the first
command: `config init` wrote a placeholder recipient, and `config.Load`
rejects any recipient that does not parse. `config init` now writes an
empty `age_recipients` list, `config.Load` accepts an empty list, and
`snapshot create` refuses to start without a recipient. A malformed
recipient is still rejected at load.

The recovery-host test now builds its config with `config init` and
`config set` and reads it through `config.Load`, so it imports
`internal/cli`. On a fresh file, `config set age_recipients.0` writes the
list in flow style (`[age1...]`).

Model: opus-5-5
2026-10-06 15:46:13 +02:00
clawbot 66c80a70e3 Skip files of an unreadable blob under restore --skip-errors (closes #218)
check / check (push) Successful in 12m34s
A blob that failed to download ended `snapshot restore` even with
--skip-errors, after restoring whichever files came first. The download
error now goes through the same per-file handling as any other restore
error, once for every pending file that references the blob. With
--skip-errors those files are reported as failed, the rest are
restored, and the command still exits non-zero. Without the flag the
restore still aborts; the error now also names one affected file and
suggests --skip-errors. A cancelled restore still ends at once. The
--skip-errors help text and README now limit the packing and storage
caveat to snapshot creation.

Model: opus-5-5
2026-10-06 14:46:05 +02:00
clawbot 81f83b29f6 Re-vendor the canonical files from sneak/prompts at dd4027b (closes #213)
check / check (push) Successful in 12m47s
Linting and testing become the lint and test phases of the Dockerfile,
and the build stage depends on both. Dockerfile.lint, CHECK_EPOCH and
the tests that checked them are removed. Every docker build in script/
passes --no-cache, and script/cibuild runs script/bootstrap first. A
host without Go gets the go.mod version from script/install-go in
.tool/go, which bootstrap, the Makefile, fmt, fmt-check, precommit and
release add to PATH; fmt-check skips .tool. The image takes its version
from the VERSION build arg or git describe, dev without .git. This
repo's own entries follow the canonical content in .gitignore and
.editorconfig. The golangci-lint v2.14.0 findings are fixed. The rules
in CLAUDE.md move into AGENTS.md. IsDevVersion counts "unknown".

Model: opus-5-5
2026-10-06 12:46:13 +02:00
clawbot c4adb72d80 Run the local index in WAL mode with a busy timeout (closes #217)
check / check (push) Successful in 5m10s
check / check (pull_request) Successful in 5m38s
The connection settings were passed as `_journal_mode=`-style
parameters, which the SQLite driver drops without an error, so the
index ran in rollback-journal mode with no busy timeout. `snapshot
list` or `info` reading during a backup could make the backup's next
write fail with "database is locked". Both open paths now pass
`_pragma=` parameters; foreign keys moved there too.

With WAL on, rows committed to the open index can still be in the
-wal file, which a copy of the main file misses. The metadata export
now copies the index with VACUUM INTO, into an empty 0600 file.

The retry after a failed open no longer claims a TRUNCATE recovery; it
retries with the same settings.

Model: opus-5-5
2026-10-06 11:29:18 +02:00
clawbot 4a167e153a Record the real uid and gid of backed-up files (closes #216)
check / check (push) Successful in 4m58s
check / check (pull_request) Successful in 6m21s
The scanner read uid and gid by asserting the stat result to an
interface with Uid() and Gid() methods. *syscall.Stat_t has Uid and Gid
fields, not methods, so the assertion never matched and every file,
directory and symlink was stored as 0:0; a restore as root then gave
everything to root. The scanner now reads the fields of
*syscall.Stat_t.

The first backup after this change re-reads every file not owned by
root, because its stored uid and gid no longer match the disk.

When the tests run as root, as in the Docker build, the new test
compares 0 with 0 and cannot catch the defect; a non-root run does.

Model: opus-5-5
2026-10-06 09:46:16 +02:00
clawbot ea72697992 List a missing file:// destination directory as an error (closes #220)
check / check (push) Successful in 5m53s
check / check (pull_request) Successful in 4m53s
The file backend listed a destination directory that does not exist as
an empty store. With the volume unplugged, snapshot list reported every
local snapshot as missing from the store, snapshot remove said it had
removed metadata it never reached, and prune dropped every local
snapshot record. List and ListStream now fail when the destination
directory is missing, so those commands take their existing path for a
store that cannot be listed. A missing prefix under an existing
directory is still an empty listing, and a first backup still creates
the directory.

Three tests listed a file:// destination nothing had created; they now
create it.

Model: opus-5-5
2026-10-06 08:46:17 +02:00
clawbot 713be502bd Reject a duration with characters outside its parts (closes #215)
check / check (push) Successful in 7m28s
check / check (pull_request) Successful in 5m39s
parseDuration fell back to an unanchored search for number-and-unit
pieces when time.ParseDuration failed, and skipped everything in
between. 1.0y became 0, so `snapshot create --prune --keep-newer-than
1.0y` deleted every snapshot of the backed-up names, the new one
included. 2.1w became one week and 1,5y five years. The fallback now
requires the whole input to be whole-number-and-unit parts with nothing
between them. A bare number is rejected before time.ParseDuration sees
it, since Go reads 0 and +0 as zero with no unit.

Judgement call: a space between number and unit (`30 days`) was
accepted and is now an error, matching Go's own units.

Model: opus-5-5
2026-10-06 06:12:07 +02:00
82 changed files with 4558 additions and 769 deletions
+4 -4
View File
@@ -18,10 +18,10 @@ jobs:
# `before:` hook and for every one of the four cross-compiles.
# Without this step the release either fails at the before-hook or,
# worse, ships binaries built by whatever Go the runner happens to
# carry. check.yml's script/cibuild also puts Go on its own runner,
# through script/bootstrap from the package manager at whatever
# version that ships, but uses it only for `go mod download` and
# gofmt; it compiles inside the digest-pinned Dockerfile images.
# carry. check.yml's runner gets the same Go through
# script/bootstrap, which calls script/install-go, and uses it only
# for `go mod download` and gofmt; it compiles inside the
# digest-pinned Dockerfile images.
#
# actions/setup-go would pin the action by commit sha, but the Go
# tarball it downloads at runtime is verified against no value in
+2 -2
View File
@@ -335,10 +335,10 @@ CreateSnapshot(opts)
│ │
│ └─► Accumulate statistics
│
├─► SnapshotManager.UpdateSnapshotStatsExtended()
│
├─► SnapshotManager.PopulateSnapshotBlobs() // record referenced blobs
│
├─► SnapshotManager.UpdateSnapshotStatsExtended()
│
├─► SnapshotManager.ExportSnapshotMetadata()
│ │
│ ├─► Copy database to temp file
+14 -7
View File
@@ -40,21 +40,28 @@ RUN go mod download
COPY . .
# The VERSION build arg when one is given, otherwise
# `git describe --tags --always` on the .git in the build context. With
# .git present, a version that is still empty, dev or unknown fails the
# build: git is missing or could not read the checkout. The commit and
# its date always come from that .git. A context without .git, such as
# a source export, stamps "dev" and an "unknown" commit and date.
# `git describe --tags --always` on the .git in the build context. The
# commit and its date always come from that .git. With .git present, a
# version that is still empty, dev or unknown, or a commit or date that
# is unknown, fails the build: git is missing or could not read the
# checkout, as when .git is a file pointing outside the context. A
# context without .git, such as a source export, stamps "dev" and an
# "unknown" commit and date.
ARG VERSION
RUN VERSION="${VERSION:-$(git describe --tags --always || echo dev)}"; \
commit="$(git rev-parse HEAD || echo unknown)"; \
commit_date="$(git show -s --format=%cs HEAD || echo unknown)"; \
if [ -e .git ]; then \
case "$VERSION" in ""|dev|unknown) \
echo "version is '$VERSION' although .git is present" >&2; \
exit 1 ;; \
esac; \
if [ "$commit" = unknown ] || [ "$commit_date" = unknown ]; then \
echo "commit is '$commit' and its date '$commit_date'" \
"although .git is present" >&2; \
exit 1; \
fi; \
fi; \
commit="$(git rev-parse HEAD || echo unknown)"; \
commit_date="$(git show -s --format=%cs HEAD || echo unknown)"; \
globals=sneak.berlin/go/vaultik/internal/globals; \
CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X ${globals}.Version=${VERSION} \
+3
View File
@@ -1,5 +1,8 @@
.PHONY: all bootstrap setup check test lint lint-fix fmt fmt-check build clean deps test-coverage local install release release-snapshot docker hooks
# Where script/bootstrap installs Go when the host has none.
export PATH := $(PATH):$(CURDIR)/.tool/go/bin
# Version number, derived from git by script/version (`git describe
# --tags --always --dirty`). This used to be a hardcoded
# constant, which meant every local build claimed to be a release that
+50 -32
View File
@@ -178,7 +178,7 @@ vaultik version
* `--verbose`, `-v`: Enable verbose output (on stderr — see below)
* `--debug`: Enable debug output (on stderr — see below)
* `--quiet`, `-q`: Suppress non-error output (also suppresses startup banner)
* `--skip-errors`: Skip files that cannot be read when creating a snapshot, or that cannot be restored when restoring, instead of aborting. Packing and storage errors (which would leave a chunk recorded but not stored) still abort the run.
* `--skip-errors`: Skip files that cannot be read when creating a snapshot, or that cannot be restored when restoring, instead of aborting. Packing and storage errors while creating a snapshot (which would leave a chunk recorded but not stored) still abort the run.
### locking
@@ -200,8 +200,10 @@ local index or the destination store. `config`, `database delete`,
### stdout and stderr
Log output — everything from `--verbose` and `--debug`, and every
warning and error the logger emits — goes to **stderr**. stdout carries
the output you asked for: tables, and the documents produced by `--json`.
warning and error the logger emits — goes to **stderr**, and so does the
startup banner. stdout carries the output you asked for: tables, the
documents produced by `--json`, `config get` values, and completion
scripts.
This means `vaultik snapshot list --verbose > out.txt` captures the
listing and leaves the diagnostics on your terminal. To capture both,
@@ -350,8 +352,9 @@ may hold snapshots this host doesn't know about), which is what
prune` invocation to run as a follow-up. Local row cleanup (files,
chunks, blobs the snapshot was the last referrer for) runs
automatically. If the destination store is unreachable, the local-DB
removal still completes and a warning is emitted; rerun `vaultik prune`
once the store is reachable to finish remote cleanup. To wipe everything
removal still completes and a warning is emitted; run `vaultik snapshot
remove <snapshot-id>` again once the store is reachable to remove the
snapshot's metadata from it (`vaultik prune` does not). To wipe everything
on the destination in one go, use `vaultik remote nuke --force`.
* `--local-only`: Skip remote cleanup; only touch the local index
* `--dry-run`: Show what would be deleted without deleting
@@ -387,7 +390,13 @@ recipients, and local database statistics.
**`remote info`**: Show storage backend type and location plus detailed
remote storage inventory: per-snapshot metadata sizes, blob counts, and
orphaned blob detection.
orphaned blob detection. A name under `metadata/` that is not a remote
key is skipped with a warning and is not printed. If a listed
`manifest.json.zst` cannot be read, or sits under a skipped name, the
orphaned blob figures are reported as unknown; `--json` gives them as
`null`, lists the remote key of each unreadable manifest in
`unreadable_manifests` and counts the manifests under skipped names in
`skipped_manifest_count`.
* `--json`: Output as JSON
**`remote nuke`**: Delete every snapshot's metadata and every blob from the
@@ -530,7 +539,7 @@ complete annotated example also lives in
| Field | Default | Description |
|-------|---------|-------------|
| `age_recipients` | (required) | Age public keys for encryption |
| `age_recipients` | (required by `snapshot create`) | Age public keys for encryption. Other commands run without one, so a machine that only restores can leave it empty |
| `age_secret_key` | (unset) | Age private key for decryption (`snapshot restore`, `snapshot verify --deep`). Setting it in the config file places the private key on the backed-up host, defeating the public-key-only design (see "why" above). Prefer the `VAULTIK_AGE_SECRET_KEY` environment variable, supplied only on the machine you restore from. |
| `snapshots` | (required) | Named snapshot definitions with paths and excludes |
| `storage_url` | | Storage backend URL (`s3://`, `file://`, `rclone://`) |
@@ -645,8 +654,8 @@ Work planned after 1.0. Loosely ordered by priority.
## output style
Every command's user-facing output is governed by `internal/ui`, in one
of two ways. Color is enabled when stdout is a TTY and the `NO_COLOR`
environment variable is unset (https://no-color.org/).
of two ways. Color is enabled when the stream written to is a TTY and
the `NO_COLOR` environment variable is unset (https://no-color.org/).
* **Status, progress, warnings, and errors** go through the `internal/ui`
message methods below: marker-prefixed, colored on a TTY, and — except
@@ -656,24 +665,26 @@ environment variable is unset (https://no-color.org/).
`config init`, `config set`, and `database delete`.
* **The data a command exists to produce** is written plain, with no
marker and no color, because a marker would corrupt a table or a
parsed document. This covers the `version`, `info`, and `remote info`
reports, the `snapshot list` table, `config get` values, and every
`--json` document. `--quiet` silences the human reports and tables
(`version`, `info`, `remote info`, `snapshot list`) but never the
machine-consumed `config get` value or the `--json` documents, which a
script depends on. The `database delete` confirmation prompt is also
written this way and always shown: it is an interactive exchange the
operator must see.
parsed document. This covers the `version`, `info`, `remote info` and
`snapshot verify` reports, the `snapshot list` table, `config get`
values, and every `--json` document. `--quiet` silences the human
reports and tables (`version`, `info`, `remote info`, `snapshot
verify`, `snapshot list`) but never the machine-consumed `config get`
value or the `--json` documents, which a script depends on. The
`database delete` confirmation prompt is also written this way and
always shown: it is an interactive exchange the operator must see.
`internal/ui` writes to stdout; it is the output the user asked for.
Structured log records are a different thing and go through
`internal/log`, which writes to stderr (see "stdout and stderr" above).
`internal/ui` writes to stdout; it is the output the user asked for. The
exceptions are the startup banner and the error a failed command ends
with, which go to stderr. Structured log records are a different thing
and go through `internal/log`, which writes to stderr (see "stdout and
stderr" above).
Message classes:
| Class | Marker | Alignment | Use for |
|-------|--------|-----------|---------|
| Banner | none | column 0 | The startup line printed once per invocation |
| Banner | none | column 0 | The startup line printed once per invocation, on stderr |
| Begin | `》` (white) | column 0 | An operation is about to start (present-continuous verb) |
| Complete | `》` (green) | column 0 | An operation just finished (past-tense verb) |
| Info | `》` (white) | column 0 | Neutral status update |
@@ -763,9 +774,13 @@ standard: normalized scripts in `script/` are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call
them. We provide:
* `script/bootstrap` — install all development dependencies (go, Go
module download). It deliberately does not install `golangci-lint`;
see `script/lint` below.
* `script/bootstrap` — install all development dependencies (Go, Go
module download). A host without Go gets the `go.mod` version through
`script/install-go`, in `.tool/go`, which `script/bootstrap` itself,
the `Makefile`, `script/fmt`, `script/fmt-check`, `script/precommit`
and `script/release` add to their `PATH`. It
deliberately does not install `golangci-lint`; see `script/lint`
below.
* `script/setup` — make a fresh clone ready for development: runs
`script/bootstrap`, then `script/install-precommit`
* `script/projectname` — print the project name (used for the Docker
@@ -780,12 +795,14 @@ them. We provide:
`script/bootstrap` insists on.
* `script/install-go` — install the Go toolchain named by `go.mod`'s
`go` directive into `.tool/go` from a sha256-verified `go.dev`
archive, and put it on `PATH`. Idempotent. Called only by the release
workflow, which needs a host Go for `goreleaser` to shell out to;
nothing else on the release runner does. `actions/setup-go` is not
used because it verifies the downloaded toolchain against no value in
this repo. Bumping Go edits `go.mod`, the checksum in this script, and
the two `golang` digests in the `Dockerfile` together.
archive, for Linux or macOS on amd64 or arm64. Idempotent. On a CI
runner it also puts `.tool/go/bin` on `PATH` for the steps that
follow. Called by the release workflow, which needs a host Go for
`goreleaser` to shell out to, and by `script/bootstrap` on a host
without Go. `actions/setup-go` is not used because it verifies the
downloaded toolchain against no value in this repo. Bumping Go edits
`go.mod`, the checksums in this script, and the two `golang` digests
in the `Dockerfile` together.
* `script/release` — cross-compile and publish the release artifacts
with the pinned `goreleaser`. Refuses a `goreleaser` on `PATH` whose
version is not the pinned one, on the same reasoning as `script/lint`.
@@ -813,7 +830,7 @@ them. We provide:
find out whether the tree is clean.
* `script/fmt` — format all code (writes)
* `script/fmt-check` — check formatting (read-only). It runs `gofmt` on
the host.
the host, over every Go file outside `.tool`.
* `script/check` — run `script/test`, `script/lint`, and
`script/fmt-check`.
* `script/docker` — build the Docker image tagged via
@@ -854,7 +871,8 @@ file. It is `git describe --tags --always --dirty`, which
A `docker build .` of a clone, with no build arguments, runs the same
`git describe` (without `--dirty`) on the `.git` in its build context,
so it stamps the same value for a clean commit; the build fails if the
context carries `.git` and no version comes out. A binary built without
context carries `.git` and no version, commit or commit date comes out.
A binary built without
git metadata reports `dev`, or `unknown` when `script/docker` or
`script/cibuild` built it outside a git checkout.
+216
View File
@@ -22,6 +22,186 @@ the tag exists and is exercised; what is left is merging `next` to
# Completed Steps
- 2026-10-07: Made two messages say only what is true
([issue #240](https://git.eeqj.de/sneak/vaultik/issues/240)). A config
file that others can read was warned about as containing S3
credentials even when it set none, as a `file://` config does. The
warning now says the file may contain S3 credentials only when
`s3.access_key_id` or `s3.secret_access_key` is set, since either may
come from a `${...}` reference rather than the file, and otherwise
says the file is readable by others. `snapshot purge` against a
destination store it could not list gave an error with
`listing remote snapshots:` in it twice; the prefix now appears once.
- 2026-10-07: Made `s3.part_size` set the multipart upload part size
([issue #232](https://git.eeqj.de/sneak/vaultik/issues/232)). It was
loaded and defaulted but never passed to the S3 client, whose uploader
used a fixed 10MiB part. It now reaches the uploader for `storage_url`
and for the `s3.*` fields, and a part size S3 refuses, below 5MiB or
above 5GiB, `0` included, fails at config load. A blob too large for
S3's limit of 10,000 parts at the configured size is uploaded in larger
parts. The docs gave the default as `5MB`, which the config file reads
as 5,000,000 bytes, below the minimum; they now say `5MiB`.
- 2026-10-07: Made per-name retention work when the hostname contains `_`
([issue #230](https://git.eeqj.de/sneak/vaultik/issues/230)). A
snapshot ID is `hostname_name_timestamp`, and the name was read as
everything between the first and the last `_`, so with
`hostname: my_host` the name `home` came out as `host_home`.
`snapshot purge --keep-latest --snapshot home` then printed "No
snapshots to delete", and `snapshot create --prune` purged nothing
without a message. The name is now read using the hostname the
`snapshots` table stores with each snapshot, cut at its first `.` as it
is in the ID.
- 2026-10-07: Made `remote info` stop reporting a snapshot's blobs as
orphaned when its manifest cannot be read, and stop printing raw
names from under `metadata/`
([issue #228](https://git.eeqj.de/sneak/vaultik/issues/228)). A
manifest it failed to read was skipped, so that snapshot's blobs were
counted as orphaned and the report advised running `vaultik prune`.
The orphan figures are now unknown in that case, with no prune
advice, and `--json` gives them as `null` with the unreadable remote
keys in `unreadable_manifests`. A name under `metadata/` that is not
64 lowercase hex characters is now skipped with a warning instead of
being printed, control characters included. A manifest under a
skipped name is then not read either, so it also leaves the orphan
figures unknown, and `--json` counts such manifests in
`skipped_manifest_count`. A directory with no manifest in it, as left
by an interrupted backup, leaves the figures known.
- 2026-10-07: Made `config set` keep a string that looks like a number
([issue #229](https://git.eeqj.de/sneak/vaultik/issues/229)). It wrote
every value unquoted, and `config.Load` reads the file through untyped
YAML, so an access key `00112233` loaded as `38043` and a hostname `007`
as `7`. A value for a string setting in `config.Config` is now tagged as
a YAML string, which the file quotes wherever YAML would read a number or
a boolean; other settings are still written unquoted.
- 2026-10-07: Made a backup notice a file rewritten with its size
unchanged and a new mtime in the same second as the one in the index
([issue #226](https://git.eeqj.de/sneak/vaultik/issues/226)). The
`files` table held mtime in whole seconds and the scanner compared
whole seconds, so every later snapshot kept the old content. A new
`mtime_nsec` column holds the nanoseconds within the second that
`mtime` holds, and the scanner compares the full mtime. A local index
created before the change lacks the column and is rebuilt with
`vaultik database delete` and a full backup.
- 2026-10-07: Made taking the process-wide lock atomic
([issue #227](https://git.eeqj.de/sneak/vaultik/issues/227)). The lock
read `vaultik.pid`, checked whether that PID was alive and then wrote
its own, so two writers started together could both pass the check and
both run. It is now an `flock` on `vaultik.pid`, held until the run
ends; the kernel drops it when the process exits, so a crash leaves no
lock behind. A clean exit now empties the file instead of deleting it,
because deleting it would let two later runs each lock a different
file.
- 2026-10-06: Made `snapshot remove --json` write only its document to
stdout when the destination store cannot be reached
([issue #251](https://git.eeqj.de/sneak/vaultik/issues/251)). Its
warning that the snapshot's metadata was left on the destination store
went to stdout ahead of the document, breaking `| jq` on a command that
exited 0. Under `--json` the warning now reaches stderr only, through
the logger. The warning, the README and the command's help said
`vaultik prune` would finish the cleanup, but `prune` never removes
snapshot metadata; they now say to run `vaultik snapshot remove` for the
snapshot again once the destination store is reachable.
- 2026-10-06: Made the backup summary and the `snapshots` row count each
file, byte and upload once
([issue #225](https://git.eeqj.de/sneak/vaultik/issues/225)). The
scanner added a file's bytes again for each new chunk and counted a
file as unchanged for each chunk already stored, so a first backup
reported twice its size and "backed up" could go negative. Upload
figures came from the progress reporter, which `--cron` turns off, and
`blob_count` counted earlier paths' blobs again for each later path.
The scanner now counts uploads itself; `blob_size`,
`blob_uncompressed_size` and `compression_ratio` describe the blobs
the snapshot references, and `docs/DATAMODEL.md` now says
`chunk_count` and `blob_count` count what the run added.
- 2026-10-06: Made command output follow the README's stdout and stderr
rules ([issue #224](https://git.eeqj.de/sneak/vaultik/issues/224)). The
startup banner went to stdout, so a `completion` script or a
`config get` value started with it; the banner now goes to stderr. A
failing `remote info`, `prune` or `snapshot remove` under `--json`
printed nothing on either stream, and now reports its error on stderr.
`snapshot verify --quiet` printed its whole report; it now prints
none, and a failure still reaches stderr with the same exit status.
- 2026-10-06: Made a backup without `--cron` of a snapshot with two or
more `paths` complete instead of panicking with `close of closed
channel` ([issue #253](https://git.eeqj.de/sneak/vaultik/issues/253)).
`Scan` runs once per path and started and stopped the progress
reporter each time, and a second stop panics. The reporter is now
started and stopped once per snapshot, around the scans of all its
paths.
- 2026-10-06: Made a restore path argument select only that path and
what is beneath it
([issue #223](https://git.eeqj.de/sneak/vaultik/issues/223)). The
lookup matched with SQL `LIKE`, so `/home/u/doc` also restored
`doc2`, `DOC` and `doc.txt.bak`, and a `_` or `%` in the path acted
as a wildcard. A backup used the same lookup to load the known files
of each configured path, so files of a longer sibling path were
counted as deleted. `FileRepository.ListUnderPath`, which replaces
`ListByPrefix`, returns the file at the path and every file whose
path starts with the path plus `/`, compared case-sensitively.
- 2026-10-06: Made restore return an error instead of panicking on a
malformed snapshot database
([issue #231](https://git.eeqj.de/sneak/vaultik/issues/231)). A chunk
hash shorter than 16 characters crashed the error message naming it,
and `--verify` dereferenced a missing `chunks` row and allocated
whatever chunk size the database gave. Those messages now go through
`shortHash`, a missing row is an error, and `--verify` rejects a
negative size and hashes each chunk as a stream.
- 2026-10-06: Made `s3://bucket/prefix` and `s3://bucket/prefix/` the same
destination ([issue #222](https://git.eeqj.de/sneak/vaultik/issues/222)).
The S3 client put the prefix directly in front of each key, so a prefix
without a trailing slash stored `prefixblobs/...`. A non-empty prefix is
now joined to every key with one `/`, giving the README's
`<bucket>/<prefix>/blobs/...` layout. The `s3.prefix` config setting goes
through the same client and gets the same join.
- 2026-10-06: Made restore apply owners, modes and times in an order
that keeps them
([issue #219](https://git.eeqj.de/sneak/vaultik/issues/219)). A
directory got its stored mode and mtime before its contents were
written, so a read-only directory came back without its files and a
non-empty one carried the time of the restore. Directories are now
created owner-only and get their stored owner, mode and mtime once the
restore loop is done, each before its parent. A file's mode is applied
after its chown, which on Linux clears setuid and setgid, and a
symlink gets its stored owner (as root) and mtime on the link itself.
- 2026-10-06: Made `go.mod` what `go mod tidy` writes, so the
pre-commit hook no longer stops every commit
([issue #246](https://git.eeqj.de/sneak/vaultik/issues/246)). A test
in `internal/cli` imports `github.com/spf13/pflag` directly, but
`go.mod` still marked it `// indirect`, and `script/precommit` fails
whenever the tidy changes `go.mod`. It is now in the direct `require`
block.
- 2026-10-06: Made the README's steps for restoring on another machine
work ([issue #221](https://git.eeqj.de/sneak/vaultik/issues/221)).
`config init` wrote a placeholder recipient that `config.Load`
rejects, so every command on the new machine failed before reaching
the store. The file now has an empty `age_recipients` list,
`config.Load` accepts an empty list, and `snapshot create` refuses to
run without a recipient.
- 2026-10-06: Made `snapshot restore --skip-errors` skip the files that
need a blob it cannot download
([issue #218](https://git.eeqj.de/sneak/vaultik/issues/218)). A missing
or damaged blob ended the restore even with the flag, after restoring
whichever files happened to come first. Every file that needs such a
blob is now reported as failed, the rest are restored, and the command
still exits non-zero. Without the flag the blob error still aborts.
- 2026-10-06: Re-vendored the canonical files from `sneak/prompts` at
`dd4027b` ([issue #213](https://git.eeqj.de/sneak/vaultik/issues/213)).
Linting and testing are now the `lint` and `test` phases of the
@@ -31,6 +211,42 @@ the tag exists and is exercised; what is left is merging `next` to
is v2.14.0, with its new findings fixed in the code, and the rules in
`CLAUDE.md` now live in `AGENTS.md`.
- 2026-10-06: Made the local index actually run in WAL mode with a busy
timeout ([issue #217](https://git.eeqj.de/sneak/vaultik/issues/217)).
The connection settings were written in a form the SQLite driver
ignores, so the index ran without either and `snapshot list` or `info`
during a backup could make the backup's next write fail with
`database is locked`. They are now `_pragma=` parameters, and the
metadata export copies the open index with `VACUUM INTO`, because a
copy of the file alone misses rows still in the `-wal` file.
- 2026-10-06: Made a backup record the real uid and gid of files,
directories and symlinks
([issue #216](https://git.eeqj.de/sneak/vaultik/issues/216)). The
scanner asked the stat result for `Uid()` and `Gid()` methods, which
`*syscall.Stat_t` does not have, so every entry was stored as `0:0`
and a restore as root gave every file to root. It now reads the
`Uid` and `Gid` fields of `*syscall.Stat_t`.
- 2026-10-06: Made a `file://` destination whose directory is missing
count as one that cannot be listed
([issue #220](https://git.eeqj.de/sneak/vaultik/issues/220)). The file
backend listed a missing directory as an empty store, so with the
quickstart's USB stick unplugged `snapshot list` reported every local
snapshot as missing from the store, `snapshot remove` claimed to have
removed metadata it never reached, and `prune` dropped every local
snapshot record. Listing a missing destination directory is now an
error; a first backup still creates the directory.
- 2026-10-06: Made `--older-than` and `--keep-newer-than` reject a
duration with characters outside its number-and-unit parts
([issue #215](https://git.eeqj.de/sneak/vaultik/issues/215)). The
parser picked out the parts it recognised and skipped the rest, so
`1.0y` became zero and `--prune --keep-newer-than 1.0y` deleted every
snapshot of the backed-up names, the new one included. `1.0y`,
`1,5y`, `30 days`, `x7d` and a bare `0` are now errors; decimals still
work in Go units such as `1.5h`.
- 2026-10-06: Made a backup re-chunk a known file that lists a chunk no
uploaded blob holds
([issue #214](https://git.eeqj.de/sneak/vaultik/issues/214)). File
+7 -1
View File
@@ -47,7 +47,9 @@ func TestProductDockerfileTakesVersionAsBuildArg(t *testing.T) {
// no VERSION, such as a plain `docker build .` of a clone, takes it from
// `git describe` of the .git in its context, stamps "dev" when the
// context has no .git, and fails rather than stamp "dev" when that .git
// yields no version. The commit and its date come from the same .git.
// yields no version. The commit and its date come from the same .git,
// and the build fails rather than stamp them "unknown" when it is
// present.
func TestProductDockerfileDerivesVersionFromGit(t *testing.T) {
t.Parallel()
@@ -66,6 +68,10 @@ func TestProductDockerfileDerivesVersionFromGit(t *testing.T) {
"%s must stamp the commit from git", productDockerfile)
assert.Contains(t, found[buildAt], "git show -s --format=%cs HEAD",
"%s must stamp the commit date from git", productDockerfile)
assert.Contains(t, found[buildAt],
`[ "$commit" = unknown ] || [ "$commit_date" = unknown ]`,
"%s must fail when the context carries .git but yields no commit"+
" or date", productDockerfile)
}
// TestDockerScriptComputesVersionOnTheHost fails unless script/docker
+7 -5
View File
@@ -3,7 +3,8 @@
# Copy this file and uncomment/modify the values you need
# Age recipient public keys for encryption
# This is REQUIRED - backups are encrypted to these public keys
# Backups are encrypted to these public keys. snapshot create needs at least
# one; listing, verifying and restoring do not
# Generate with: age-keygen | grep "public key"
age_recipients:
- age1cj2k2addawy294f6k2gr2mf9gps9r3syplryxca3nvxj3daqm96qfp84tz
@@ -286,10 +287,11 @@ storage_url: "rclone://myremote/path/to/backups"
# #use_ssl: true
#
# # Part size for multipart uploads
# # Minimum 5MB, affects memory usage during upload
# # Supports: 5MB, 10M, 100MiB, etc.
# # Default: 5MB
# #part_size: 5MB
# # Minimum 5MiB, maximum 5GiB; affects memory usage during upload
# # A blob too large for 10,000 parts of this size gets larger parts
# # Supports: 10MB, 16MiB, 100MiB, etc. (5MB is below the minimum)
# # Default: 5MiB
# #part_size: 5MiB
# Path to local SQLite index database
# This database tracks file state for incremental backups
+6 -5
View File
@@ -36,7 +36,8 @@ Stores metadata about files in the filesystem being backed up.
**Columns:**
- `id` (TEXT PRIMARY KEY) - UUID for the file record
- `path` (TEXT NOT NULL UNIQUE) - Absolute file path
- `mtime` (INTEGER NOT NULL) - Modification time as Unix timestamp
- `mtime` (INTEGER NOT NULL) - Modification time, whole seconds since the Unix epoch
- `mtime_nsec` (INTEGER NOT NULL) - Nanoseconds within that second, 0 to 999999999
- `size` (INTEGER NOT NULL) - File size in bytes
- `mode` (INTEGER NOT NULL) - Unix file permissions and type
- `uid` (INTEGER NOT NULL) - User ID of file owner
@@ -117,10 +118,10 @@ Tracks backup snapshots.
- `started_at` (INTEGER) - Start timestamp
- `completed_at` (INTEGER) - Completion timestamp (NULL if in progress)
- `file_count` (INTEGER) - Number of files in snapshot
- `chunk_count` (INTEGER) - Number of unique chunks
- `blob_count` (INTEGER) - Number of blobs referenced
- `chunk_count` (INTEGER) - Number of chunks this snapshot stored that were not stored before
- `blob_count` (INTEGER) - Number of blobs this snapshot created
- `total_size` (INTEGER) - Total size of all files
- `blob_size` (INTEGER) - Total size of all blobs (compressed)
- `blob_size` (INTEGER) - Total compressed size of all referenced blobs
- `blob_uncompressed_size` (INTEGER) - Total uncompressed size of all referenced blobs
- `compression_ratio` (REAL) - Compression ratio achieved
- `compression_level` (INTEGER) - Compression level used for this snapshot
@@ -280,7 +281,7 @@ This ensures consistency, especially important for operations like:
3. **Batch Operations**: Where possible, operations are batched within transactions
4. **Write-Ahead Logging**: SQLite WAL mode is enabled for better concurrency
4. **Write-Ahead Logging**: The local index runs in SQLite WAL mode with a 10-second busy timeout, so a read-only command such as `snapshot list` can read it while a backup writes to it. Committed rows can sit in the `-wal` file beside the index until a checkpoint, so the metadata export copies the index through SQLite (`VACUUM INTO`), not as a file
## Data Integrity
+2 -2
View File
@@ -20,9 +20,11 @@ require (
github.com/rclone/rclone v1.72.1
github.com/spf13/afero v1.15.0
github.com/spf13/cobra v1.10.1
github.com/spf13/pflag v1.0.10
github.com/stretchr/testify v1.11.1
go.uber.org/fx v1.24.0
golang.org/x/sync v0.18.0
golang.org/x/sys v0.38.0
golang.org/x/term v0.37.0
gopkg.in/yaml.v3 v3.0.1
modernc.org/sqlite v1.38.0
@@ -225,7 +227,6 @@ require (
github.com/smarty/assertions v1.16.0 // indirect
github.com/sony/gobreaker v1.0.0 // indirect
github.com/spacemonkeygo/monkit/v3 v3.0.25-0.20251022131615-eb24eb109368 // indirect
github.com/spf13/pflag v1.0.10 // indirect
github.com/t3rm1n4l/go-mega v0.0.0-20251031123324-a804aaa87491 // indirect
github.com/tidwall/gjson v1.18.0 // indirect
github.com/tidwall/match v1.1.1 // indirect
@@ -263,7 +264,6 @@ require (
golang.org/x/exp v0.0.0-20251023183803-a4bb9ffd2546 // indirect
golang.org/x/net v0.47.0 // indirect
golang.org/x/oauth2 v0.33.0 // indirect
golang.org/x/sys v0.38.0 // indirect
golang.org/x/text v0.31.0 // indirect
golang.org/x/time v0.14.0 // indirect
golang.org/x/tools v0.38.0 // indirect
+13 -18
View File
@@ -200,10 +200,10 @@ func RunApp(ctx context.Context, app *fx.App) error {
}
// errReported marks a failure the operation has already shown the user
// (and deliberately withheld under --json). Entry turns it into a
// non-zero exit status without printing anything further, so the error
// line is not doubled. It flows up from RunOperation through cobra to
// Entry.
// (or, under `snapshot verify --json`, put in its document). Entry
// turns it into a non-zero exit status without printing anything
// further, so the error line is not doubled. It flows up from
// RunOperation through cobra to Entry.
var errReported = errors.New("operation failed")
// RunOperation runs op against the Vaultik instance inside the fx app
@@ -220,10 +220,10 @@ var errReported = errors.New("operation failed")
// interrupt OnStop cancels op and waits for the goroutine to return, so
// op's cleanup (removing decrypted scratch files) runs before the
// process exits; the wait is bounded by shutdownTimeout. report is
// called with a non-canceled failure so the caller can log it (and
// suppress it under --json) before it becomes errReported. A context
// cancellation is the interrupt path, not a failure: it is neither
// reported nor counted as one.
// called with a non-canceled failure so the caller can show it to the
// user before it becomes errReported. A context cancellation is the
// interrupt path, not a failure: it is neither reported nor counted as
// one.
func RunOperation(
ctx context.Context, opts AppOptions,
op func(v *vaultik.Vaultik) error, report func(err error),
@@ -293,13 +293,12 @@ func RunOperation(
// runVaultikApp runs the standard single-operation command lifecycle
// shared by the snapshot list/purge/remove and remote nuke subcommands:
// resolve the config, then run op against the Vaultik instance through
// RunOperation, reporting a failure prefixed with failMsg (suppressed
// while suppressErrors is true, e.g. under --json). mode says whether the
// command takes the PID lock. jsonOutput marks a command whose stdout is a
// JSON document: it quiets the UI but, unlike Quiet, leaves the stderr log
// level alone.
// RunOperation, reporting a failure prefixed with failMsg on stderr. mode
// says whether the command takes the PID lock. jsonOutput marks a command
// whose stdout is a JSON document: it quiets the UI but, unlike Quiet,
// leaves the stderr log level alone.
func runVaultikApp(
cmd *cobra.Command, mode lockMode, jsonOutput, suppressErrors bool,
cmd *cobra.Command, mode lockMode, jsonOutput bool,
failMsg string, op func(v *vaultik.Vaultik) error,
) error {
configPath, err := ResolveConfigPath()
@@ -319,10 +318,6 @@ func runVaultikApp(
},
Mode: mode,
}, op, func(err error) {
if suppressErrors {
return
}
log.Error(failMsg, "error", err)
ReportErrorf("%s: %v", failMsg, err)
})
+63 -5
View File
@@ -7,11 +7,14 @@ import (
"os"
"os/exec"
"path/filepath"
"reflect"
"strconv"
"strings"
"unicode/utf8"
"github.com/spf13/cobra"
"gopkg.in/yaml.v3"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/ui"
)
@@ -31,6 +34,9 @@ const configDirMode = 0o755
// yaml.Marshal's 4-space default.
const configYAMLIndent = 2
// yamlStringTag is YAML's tag for a string scalar.
const yamlStringTag = "!!str"
var (
errConfigExists = errors.New("config file already exists")
errEmptyConfig = errors.New("empty config file")
@@ -45,16 +51,19 @@ const defaultConfigTemplate = `# vaultik configuration
# ─── REQUIRED ────────────────────────────────────────────────────────────────
# Age recipient public keys for encryption.
# Age recipient public keys for encryption. snapshot create needs at least
# one; listing, verifying and restoring do not, so a machine that only
# restores can leave this empty.
# Backups are encrypted to ALL listed recipients; any one of the corresponding
# private keys can decrypt. Adding a recipient later does not re-encrypt data
# already stored: deduplicated chunks and existing blobs stay encrypted to the
# earlier recipients, so a newly added key cannot restore them on its own (see
# docs/REPOSTRUCTURE.md, Accepted Risks). Generate a keypair with:
# docs/REPOSTRUCTURE.md, Accepted Risks). Generate a keypair and add its
# public key with:
# age-keygen -o vaultik_backup_private_key.txt
# grep 'public key' vaultik_backup_private_key.txt
age_recipients:
- age1REPLACE_WITH_YOUR_PUBLIC_KEY
# vaultik config set age_recipients.0 age1...
age_recipients: []
# Named snapshots. Each snapshot backs up one or more paths and can have its
# own exclude patterns in addition to the global excludes below.
@@ -196,7 +205,7 @@ storage_url: ""
# access_key_id: YOUR_ACCESS_KEY
# secret_access_key: YOUR_SECRET_KEY
# # region: us-east-1 # Default: us-east-1
# # part_size: 5MB # Multipart upload part size. Default: 5MB
# # part_size: 5MiB # Upload part size, 5MiB to 5GiB. Default: 5MiB
# # For the s3:// form, disable TLS with ?ssl=false in the URL, not use_ssl.
# ─── OPTIONAL ────────────────────────────────────────────────────────────────
@@ -580,9 +589,58 @@ func yamlPathSet(root *yaml.Node, keys []string, value string) error {
}
}
// config.Load reads the file through untyped YAML, which turns an
// unquoted 00112233 into the number 38043 and 1e5 into 100000. Tagging
// a string setting as a string makes the encoder quote such a value.
// Other settings stay unquoted, so compression_level 9 is a number.
// The encoder refuses to write a value that is not valid UTF-8 as a
// string. Left untagged, such a value is written as base64 !!binary and
// loads back unchanged.
if configKeyIsString(keys) && utf8.ValidString(value) {
node.Tag = yamlStringTag
}
return nil
}
// configKeyIsString reports whether the dotted key names a string in
// config.Config, following the fields' yaml tags, as s3.access_key_id and
// snapshots.home.exclude.0 do.
func configKeyIsString(keys []string) bool {
typ := reflect.TypeFor[config.Config]()
for _, key := range keys {
switch {
case typ.Kind() == reflect.Map || typ.Kind() == reflect.Slice:
// The key is a snapshot name or a list index.
typ = typ.Elem()
case typ.Kind() == reflect.Struct:
field, ok := yamlField(typ, key)
if !ok {
return false
}
typ = field.Type
default:
return false
}
}
return typ.Kind() == reflect.String
}
// yamlField returns the field of struct type typ whose yaml tag names key.
func yamlField(typ reflect.Type, key string) (reflect.StructField, bool) {
for field := range typ.Fields() {
name, _, _ := strings.Cut(field.Tag.Get("yaml"), ",")
if name == key {
return field, true
}
}
return reflect.StructField{}, false
}
// yamlSetInMapping resolves (creating if needed) the value node for key
// within a mapping node, setting it to value when it is the final path
// element, and returns the node to descend into.
+146 -2
View File
@@ -4,6 +4,7 @@ import (
"bytes"
"os"
"path/filepath"
"strconv"
"strings"
"testing"
@@ -24,8 +25,10 @@ func TestDefaultConfigTemplateParses(t *testing.T) {
t.Fatalf("default config template is not valid YAML: %v", err)
}
if len(cfg.AgeRecipients) != 1 {
t.Errorf("expected 1 placeholder age recipient, got %d", len(cfg.AgeRecipients))
// A placeholder recipient would fail config.Load, so the template
// leaves the list empty.
if len(cfg.AgeRecipients) != 0 {
t.Errorf("expected no age recipients, got %d", len(cfg.AgeRecipients))
}
home, ok := cfg.Snapshots["home"]
@@ -55,6 +58,147 @@ func TestDefaultConfigTemplateParses(t *testing.T) {
}
}
// TestConfigSetRecipientOnFreshConfig follows the README quickstart: on the
// file `config init` writes, `config set age_recipients.0` and
// `config set storage_url` give a config that loads with that recipient.
func TestConfigSetRecipientOnFreshConfig(t *testing.T) {
t.Parallel()
const recipient = "age1278m9q7dp3chsh2dcy82qk27v047zywyvtxwnj4cvt0z65jw6a7q5dqhfj"
path := filepath.Join(t.TempDir(), "config.yml")
err := os.WriteFile(path, []byte(defaultConfigTemplate), configFileMode)
if err != nil {
t.Fatalf("write config: %v", err)
}
out := ui.NewWithColor(&bytes.Buffer{}, false)
err = writeConfigSet(out, path, "age_recipients.0", recipient)
if err != nil {
t.Fatalf("config set age_recipients.0: %v", err)
}
err = writeConfigSet(out, path, "storage_url", "file:///mnt/backups")
if err != nil {
t.Fatalf("config set storage_url: %v", err)
}
cfg, err := config.Load(path)
if err != nil {
t.Fatalf("config.Load: %v", err)
}
if len(cfg.AgeRecipients) != 1 || cfg.AgeRecipients[0] != recipient {
t.Errorf("age_recipients = %v, want [%s]", cfg.AgeRecipients, recipient)
}
}
// TestConfigSetStringLooksLikeNumber sets string settings to values that
// YAML reads as numbers or booleans when they are unquoted, and checks that
// config.Load returns each one unchanged.
func TestConfigSetStringLooksLikeNumber(t *testing.T) {
t.Parallel()
tests := []struct {
key string
value string
field func(cfg *config.Config) string
}{
{"s3.access_key_id", "00112233",
func(cfg *config.Config) string { return cfg.S3.AccessKeyID }},
{"s3.secret_access_key", "12345678901234567890123456789012",
func(cfg *config.Config) string { return cfg.S3.SecretAccessKey }},
{"hostname", "007",
func(cfg *config.Config) string { return cfg.Hostname }},
{"s3.prefix", "1e5",
func(cfg *config.Config) string { return cfg.S3.Prefix }},
{"s3.bucket", "true",
func(cfg *config.Config) string { return cfg.S3.Bucket }},
{"s3.region", "FALSE",
func(cfg *config.Config) string { return cfg.S3.Region }},
{"snapshots.home.exclude.0", "1.10",
func(cfg *config.Config) string { return cfg.Snapshots["home"].Exclude[0] }},
}
for _, tt := range tests {
t.Run(tt.key+"="+tt.value, func(t *testing.T) {
t.Parallel()
cfg := loadAfterConfigSet(t, tt.key, tt.value)
got := tt.field(cfg)
if got != tt.value {
t.Errorf("%s = %q after config set %q", tt.key, got, tt.value)
}
})
}
}
// TestConfigSetNonUTF8Path checks that config set still accepts a value that
// is not valid UTF-8, such as a path with a Latin-1 file name, and that
// config.Load returns it unchanged.
func TestConfigSetNonUTF8Path(t *testing.T) {
t.Parallel()
const dir = "/srv/caf\xe9"
cfg := loadAfterConfigSet(t, "snapshots.home.paths.0", dir)
got := cfg.Snapshots["home"].Paths[0]
if got != dir {
t.Errorf("snapshots.home.paths.0 = %q, want %q", got, dir)
}
}
// TestConfigSetNumberStaysNumber checks that a number set for an integer
// setting is still read as a number, not as a quoted string.
func TestConfigSetNumberStaysNumber(t *testing.T) {
t.Parallel()
const level = 9
cfg := loadAfterConfigSet(t, "compression_level", strconv.Itoa(level))
if cfg.CompressionLevel != level {
t.Errorf("compression_level = %d, want %d", cfg.CompressionLevel, level)
}
}
// loadAfterConfigSet writes the file `config init` writes, sets storage_url
// to a local directory so that the file passes validation, applies
// `config set key value` and returns what config.Load reads back.
func loadAfterConfigSet(t *testing.T, key, value string) *config.Config {
t.Helper()
path := filepath.Join(t.TempDir(), "config.yml")
err := os.WriteFile(path, []byte(defaultConfigTemplate), configFileMode)
if err != nil {
t.Fatalf("write config: %v", err)
}
out := ui.NewWithColor(&bytes.Buffer{}, false)
err = writeConfigSet(out, path, "storage_url", "file:///mnt/backups")
if err != nil {
t.Fatalf("config set storage_url: %v", err)
}
err = writeConfigSet(out, path, key, value)
if err != nil {
t.Fatalf("config set %s: %v", key, err)
}
cfg, err := config.Load(path)
if err != nil {
t.Fatalf("config.Load: %v", err)
}
return cfg
}
const testYAML = `# top comment
compression_level: 3
age_recipients:
+10 -9
View File
@@ -16,16 +16,18 @@ import (
const shortCommitLen = 12
// Entry is the main entry point for the CLI application.
// It prints the startup banner to stdout (unless a banner-suppressing
// It prints the startup banner to stderr (unless a banner-suppressing
// flag is present in os.Args — see bannerSuppressedInArgs), executes the
// root cobra command, and routes any returned error through the
// ui.Writer so the user sees a properly formatted "🛑 ERROR:" line.
// The banner goes to stderr because stdout carries only the output the
// user asked for, such as a completion script or a `config get` value.
//
// It returns the process exit code (0 on success, 1 on error) rather
// than calling os.Exit, so that main's deferred profile writers run
// before the process ends. See run in cmd/vaultik/main.go.
func Entry() int {
emitStartupBanner(os.Args[1:], os.Stdout)
emitStartupBanner(os.Args[1:], os.Stderr)
rootCmd := NewRootCommand()
rootCmd.SilenceErrors = true
@@ -33,8 +35,9 @@ func Entry() int {
err := rootCmd.Execute()
if err != nil {
// An operation that ran inside the fx app has already reported
// its own failure (and suppressed it under --json); errReported
// says so. Printing it again here would double the error line.
// its own failure (`snapshot verify --json` puts it in the
// document instead); errReported says so. Printing it again
// here would double the error line.
// Every other error — bad arguments, a config that would not
// load — reaches Entry unreported, so it is shown here.
if !errors.Is(err, errReported) {
@@ -49,9 +52,8 @@ func Entry() int {
// emitStartupBanner writes the startup banner to w unless args (the
// argument vector with the program name already stripped) contains a
// flag that suppresses it. Split out of Entry so that the decision — the
// only thing standing between a --json invocation and a parseable
// stdout — is reachable from a test without running the whole CLI.
// flag that suppresses it. Split out of Entry so that the decision is
// reachable from a test without running the whole CLI.
func emitStartupBanner(args []string, w io.Writer) {
if bannerSuppressedInArgs(args) {
return
@@ -86,8 +88,7 @@ func ReportErrorf(format string, args ...any) {
// --json is a subcommand flag rather than a persistent one, but so is
// --cron (it exists only on `snapshot create`), so this adds no new
// class of imprecision. The only cost of a false positive is a missing
// decorative banner; the cost of a false negative is a corrupt document
// on stdout, so the scan errs deliberately in that direction.
// decorative banner.
func bannerSuppressedInArgs(args []string) bool {
for _, a := range args {
if a == "--" {
+25 -49
View File
@@ -36,23 +36,14 @@ const (
// strips it before scanning, so it has to be present.
programName = "vaultik"
// someSnapshotID is any snapshot identifier: these tests never run
// the command, so it only has to occupy the positional argument.
// someSnapshotID only fills the positional argument; no test needs
// the snapshot to exist.
someSnapshotID = "host_2026-01-01T00:00:00Z"
)
// placeholderJSONDocument stands in for whatever document a --json
// command writes to stdout. `snapshot list --json` with no snapshots
// prints exactly this; the other --json commands print an object rather
// than an array, but this test is not about their shape. It is about
// what is on stdout *before* them, which is the same for all of them
// because Entry prints the banner before cobra has parsed anything and
// therefore before it can know which command is running.
const placeholderJSONDocument = "[]\n"
// jsonArgumentVectors are the argument vectors of every --json
// invocation the CLI accepts, with the program name stripped exactly as
// Entry strips it. Each one must leave stdout untouched by the banner.
// Entry strips it. Each one must suppress the banner.
//
//nolint:gochecknoglobals // read-only test fixture shared by two tests
var jsonArgumentVectors = map[string][]string{
@@ -74,39 +65,23 @@ var jsonArgumentVectors = map[string][]string{
},
}
// TestJSONInvocationStdoutIsExactlyOneDocument is the CLI-layer
// regression guard for issue #106: `vaultik snapshot list --json | jq`
// must work with no other flags.
//
// internal/vaultik's TestListSnapshots_JSONStdoutIsOnlyTheDocument
// guards the same contract one layer down, but it calls the library
// function directly and so cannot see Entry, which is where the
// contamination was: the startup banner is written to stdout before
// cobra parses anything, and the suppression scan did not know about
// --json. The two banner lines and the blank line landed ahead of the
// document and `jq` refused the result.
//
// The document is a constant here because this test is about the
// argument vectors, one per --json command; the one that runs a real
// command end to end is TestEntryJSONStdoutIsExactlyOneDocument below.
func TestJSONInvocationStdoutIsExactlyOneDocument(t *testing.T) {
// TestJSONInvocationSuppressesBanner checks that every --json
// invocation suppresses the startup banner, as the README says --json
// does along with --quiet and --cron. The scan is over the raw argument
// vector, so each position and spelling of --json is listed.
func TestJSONInvocationSuppressesBanner(t *testing.T) {
t.Parallel()
for name, argv := range jsonArgumentVectors {
t.Run(name, func(t *testing.T) {
t.Parallel()
var stdout bytes.Buffer
var banner bytes.Buffer
emitStartupBanner(argv, &stdout)
emitStartupBanner(argv, &banner)
require.Empty(t, stdout.String(),
"nothing may reach stdout ahead of a --json document")
_, err := stdout.WriteString(placeholderJSONDocument)
require.NoError(t, err)
requireExactlyOneJSONDocument(t, stdout.String())
assert.Empty(t, banner.String(),
"--json suppresses the banner")
})
}
}
@@ -127,11 +102,11 @@ func TestBannerStillPrintedWithoutSuppressingFlag(t *testing.T) {
t.Run(name, func(t *testing.T) {
t.Parallel()
var stdout bytes.Buffer
var banner bytes.Buffer
emitStartupBanner(argv, &stdout)
emitStartupBanner(argv, &banner)
assert.Contains(t, stdout.String(), "starting up at",
assert.Contains(t, banner.String(), "starting up at",
"the banner belongs on invocations that did not opt out")
})
}
@@ -172,9 +147,9 @@ func TestBannerSuppressedInArgs(t *testing.T) {
// hermeticConfig is a complete, valid config that needs no network and
// no credentials: file:// storage is exempt from the S3 credential
// checks, and FileStorer over a directory that does not exist lists
// zero objects without erroring. Chunk, blob and compression settings
// are filled in by config.Load.
// checks. A test that lists the destination must create its directory
// first, because listing a directory that does not exist is an error.
// Chunk, blob and compression settings are filled in by config.Load.
const hermeticConfig = `age_recipients:
- age1278m9q7dp3chsh2dcy82qk27v047zywyvtxwnj4cvt0z65jw6a7q5dqhfj
snapshots:
@@ -196,22 +171,25 @@ hostname: test-host
// `snapshot list` is the command chosen because it is the only --json
// command that reaches its document without a populated destination
// store: it reads the local index, streams `metadata/` (empty here),
// and treats a barren destination as an empty list rather than a
// failure.
// and treats an empty destination directory as an empty list rather
// than a failure.
//
// Not parallel: it replaces os.Args, os.Stdout and the xdg globals.
func TestEntryJSONStdoutIsExactlyOneDocument(t *testing.T) {
dir := t.TempDir()
configPath := filepath.Join(dir, "config.yml")
storeDir := filepath.Join(dir, "store")
contents := fmt.Sprintf(hermeticConfig,
filepath.Join(dir, "source"),
filepath.Join(dir, "store"),
storeDir,
filepath.Join(dir, "index.sqlite"))
require.NoError(t,
os.WriteFile(configPath, []byte(contents), configFileMode))
require.NoError(t, os.Mkdir(storeDir, 0o750))
// The PID lock lives under xdg.DataHome, which xdg resolves at
// package init; point it at the temp dir so the test neither
// touches nor collides with the real one.
@@ -244,9 +222,7 @@ func TestEntryJSONStdoutIsExactlyOneDocument(t *testing.T) {
// captureProcessStdout redirects the process's own stdout to a pipe for
// the duration of fn and returns what was written to it. The redirection
// has to be at the file-descriptor level rather than through an injected
// writer, because the banner and the JSON encoder reach os.Stdout
// independently and the point of the test is that both land in the same
// place.
// writer, because the commands Entry runs reach os.Stdout directly.
//
// Not parallel-safe: os.Stdout is process-global.
func captureProcessStdout(t *testing.T, fn func()) string {
+10 -5
View File
@@ -98,25 +98,30 @@ func TestEntryPruneJSONStdoutIsExactlyOneDocument(t *testing.T) {
}
}
// writeHermeticPruneConfig builds a config over a temp directory and, if
// seedStale is set, creates the index database up front with one
// snapshot record that has no counterpart on the destination store.
// Returns the config path.
// writeHermeticPruneConfig builds a config over a temp directory with an
// empty destination directory and, if seedStale is set, creates the
// index database up front with one snapshot record that has no
// counterpart on the destination store. Returns the config path.
func writeHermeticPruneConfig(t *testing.T, seedStale bool) string {
t.Helper()
dir := t.TempDir()
configPath := filepath.Join(dir, "config.yml")
indexPath := filepath.Join(dir, "index.sqlite")
storeDir := filepath.Join(dir, "store")
contents := fmt.Sprintf(hermeticConfig,
filepath.Join(dir, "source"),
filepath.Join(dir, "store"),
storeDir,
indexPath)
require.NoError(t,
os.WriteFile(configPath, []byte(contents), configFileMode))
// prune fails on a destination directory that does not exist, so the
// empty store is created here rather than left to a first backup.
require.NoError(t, os.Mkdir(storeDir, 0o750))
// The PID lock lives under xdg.DataHome, which xdg resolves at
// package init; point it at the temp dir so the test neither
// touches nor collides with the real one.
+3 -3
View File
@@ -15,8 +15,8 @@ import (
// run, so a failing command must come back with a non-zero code rather
// than ending the process here.
//
// Stdout is captured only to keep the banner and command output off the
// test log; the assertion is on the returned code.
// Stdout and stderr are captured only to keep the banner and command
// output off the test log; the assertion is on the returned code.
//
//nolint:paralleltest // replaces os.Args and rootFlags
func TestEntryReturnsStatusCode(t *testing.T) {
@@ -50,7 +50,7 @@ func TestEntryReturnsStatusCode(t *testing.T) {
var code int
_ = captureProcessStdout(t, func() { code = Entry() })
_, _ = captureProcessStdoutAndStderr(t, func() { code = Entry() })
assert.Equal(t, testCase.want, code)
})
+200
View File
@@ -0,0 +1,200 @@
package cli //nolint:testpackage // shares hermeticConfig and the capture helpers
import (
"context"
"encoding/json"
"fmt"
"log/slog"
"os"
"path/filepath"
"strings"
"testing"
"github.com/adrg/xdg"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/database"
)
// TestEntryCompletionStdoutIsTheScript runs `vaultik completion bash`,
// whose stdout the README tells the user to source. The script has to
// start on the first line.
//
//nolint:paralleltest // replaces os.Args, os.Stdout and os.Stderr
func TestEntryCompletionStdoutIsTheScript(t *testing.T) {
code, stdout, _ := runEntry(t, "completion", "bash")
require.Equal(t, 0, code)
firstLine, _, _ := strings.Cut(stdout, "\n")
assert.True(t, strings.HasPrefix(firstLine, "# bash completion"),
"the first line of stdout must be the script's, got %q", firstLine)
}
// TestEntryConfigGetStdoutIsTheValue runs `vaultik config get`, whose
// stdout a script reads as the value and nothing else.
//
//nolint:paralleltest // replaces os.Args, os.Stdout and os.Stderr
func TestEntryConfigGetStdoutIsTheValue(t *testing.T) {
configPath := filepath.Join(t.TempDir(), "config.yml")
require.NoError(t, os.WriteFile(configPath,
[]byte("hostname: test-host\n"), configFileMode))
code, stdout, _ := runEntry(t,
flagConfig, configPath, "config", "get", "hostname")
require.Equal(t, 0, code)
assert.Equal(t, "test-host\n", stdout)
}
// TestEntryJSONFailureIsReportedOnStderr runs each --json command that
// writes no document when it fails, against a destination it cannot
// use. The error must reach stderr, and stdout must stay empty.
//
//nolint:paralleltest // replaces os.Args, os.Stdout, os.Stderr and the xdg globals
func TestEntryJSONFailureIsReportedOnStderr(t *testing.T) {
for _, testCase := range []struct {
name string
args []string
wantOnStderr string
}{
{
name: "remote info",
args: []string{cmdRemote, cmdInfo, flagJSON},
wantOnStderr: "Failed to get remote info",
},
{
name: "prune",
args: []string{cmdPrune, flagJSON},
wantOnStderr: "Prune failed",
},
{
name: "snapshot remove",
args: []string{cmdSnapshot, cmdRemove, someSnapshotID, flagJSON},
wantOnStderr: "Failed to remove snapshot",
},
} {
t.Run(testCase.name, func(t *testing.T) {
configPath := writeUnusableDestinationConfig(t)
code, stdout, stderr := runEntry(t,
append([]string{flagConfig, configPath}, testCase.args...)...)
assert.Equal(t, 1, code)
assert.Empty(t, stdout,
"a failed --json command has no document to write")
assert.Contains(t, stderr, testCase.wantOnStderr,
"the failure must be reported on stderr")
})
}
}
// TestEntrySnapshotRemoveJSONWarningIsOnStderr runs `snapshot remove
// --json` on a snapshot in the local index, against a destination
// directory that does not exist. The command removes the snapshot from
// the local index and still exits 0. Its stdout must hold the document
// alone, with the warning about the destination store on stderr: the
// command to run again once it is reachable, and the snapshot's ID in
// the record's snapshot_id field.
//
//nolint:paralleltest // replaces os.Args, os.Stdout, os.Stderr and the xdg globals
func TestEntrySnapshotRemoveJSONWarningIsOnStderr(t *testing.T) {
configPath, indexPath := writeMissingDestinationConfig(t)
seedStaleSnapshotRecord(t, indexPath)
code, stdout, stderr := runEntry(t, flagConfig, configPath,
cmdSnapshot, cmdRemove, stalePruneSnapshotID, flagJSON)
require.Equal(t, 0, code)
requireExactlyOneJSONDocument(t, stdout)
// stderr is a pipe here, so the logger writes one JSON record a line.
var warning map[string]any
for line := range strings.Lines(stderr) {
if strings.Contains(line,
"Could not remove snapshot metadata from remote storage") {
require.NoError(t, json.Unmarshal([]byte(line), &warning))
}
}
require.NotNil(t, warning, "the warning must reach stderr")
assert.Contains(t, warning[slog.MessageKey],
"run 'vaultik snapshot remove' with the snapshot's ID again")
assert.Equal(t, stalePruneSnapshotID, warning["snapshot_id"])
}
// writeMissingDestinationConfig builds a config whose destination
// directory does not exist. Returns the config path and the path of
// its local index, which is not created here.
func writeMissingDestinationConfig(t *testing.T) (string, string) {
t.Helper()
dir := t.TempDir()
configPath := filepath.Join(dir, "config.yml")
indexPath := filepath.Join(dir, "index.sqlite")
contents := fmt.Sprintf(hermeticConfig,
filepath.Join(dir, "source"),
filepath.Join(dir, "missing-store"),
indexPath)
require.NoError(t,
os.WriteFile(configPath, []byte(contents), configFileMode))
// The PID lock lives under xdg.DataHome, which xdg resolves at
// package init; point it at the temp dir so the test neither
// touches nor collides with the real one.
t.Setenv("XDG_DATA_HOME", filepath.Join(dir, "data"))
xdg.Reload()
t.Cleanup(xdg.Reload)
return configPath, indexPath
}
// writeUnusableDestinationConfig builds a config whose destination
// directory does not exist, which fails `remote info`, and whose local
// index is bound to another destination, which fails `prune` and
// `snapshot remove` (a missing destination alone only makes `snapshot
// remove` warn). Returns the config path.
func writeUnusableDestinationConfig(t *testing.T) string {
t.Helper()
configPath, indexPath := writeMissingDestinationConfig(t)
ctx := context.Background()
db, err := database.New(ctx, indexPath)
require.NoError(t, err)
defer func() { require.NoError(t, db.Close()) }()
require.NoError(t, database.NewRepositories(db).LocalMeta.Set(ctx,
database.LocalMetaKeyStorageURL, "file://"+t.TempDir()))
return configPath
}
// runEntry runs Entry with args after the program name and returns its
// exit code and what it wrote to stdout and stderr.
//
// Not parallel-safe: it replaces os.Args, os.Stdout and os.Stderr.
func runEntry(t *testing.T, args ...string) (int, string, string) {
t.Helper()
previousArgs := os.Args
t.Cleanup(func() {
os.Args = previousArgs
rootFlags = RootFlags{}
})
os.Args = append([]string{programName}, args...)
var code int
stdout, stderr := captureProcessStdoutAndStderr(t,
func() { code = Entry() })
return code, stdout, stderr
}
-4
View File
@@ -48,10 +48,6 @@ work (e.g. after a crashed backup or to reclaim storage).`,
}, func(v *vaultik.Vaultik) error {
return v.Prune(opts)
}, func(err error) {
if opts.JSON {
return
}
log.Error("Prune operation failed", "error", err)
ReportErrorf("Prune failed: %v", err)
})
+1 -5
View File
@@ -45,7 +45,7 @@ This is destructive and irreversible. Requires --force.`,
return errNukeNeedsForce
}
return runVaultikApp(cmd, mutating, false, false, "Remote nuke failed",
return runVaultikApp(cmd, mutating, false, "Remote nuke failed",
func(v *vaultik.Vaultik) error {
return v.NukeRemote(true)
})
@@ -92,10 +92,6 @@ func newRemoteInfoCommand() *cobra.Command {
}, func(v *vaultik.Vaultik) error {
return v.RemoteInfo(jsonOutput)
}, func(err error) {
if jsonOutput {
return
}
log.Error("Failed to get remote info", "error", err)
ReportErrorf("Failed to get remote info: %v", err)
})
+2 -1
View File
@@ -60,7 +60,8 @@ on the source system.`,
cmd.PersistentFlags().BoolVar(&rootFlags.SkipErrors, "skip-errors", false,
"Skip files that cannot be read when creating a snapshot, or "+
"that cannot be restored when restoring, instead of aborting "+
"(packing and storage errors still abort)")
"(packing and storage errors while creating a snapshot still "+
"abort)")
// Add subcommands
cmd.AddCommand(
+6 -5
View File
@@ -126,7 +126,7 @@ func newSnapshotListCommand() *cobra.Command {
Long: "Lists all snapshots with their ID, timestamp, and compressed size",
Args: cobra.NoArgs,
RunE: func(cmd *cobra.Command, _ []string) error {
return runVaultikApp(cmd, readOnly, false, false,
return runVaultikApp(cmd, readOnly, false,
"Failed to list snapshots",
func(v *vaultik.Vaultik) error {
return v.ListSnapshots(jsonOutput)
@@ -162,7 +162,7 @@ restrict the operation to specific snapshot names.`,
return errPurgeCriteriaBoth
}
return runVaultikApp(cmd, mutating, false, false,
return runVaultikApp(cmd, mutating, false,
"Failed to purge snapshots",
func(v *vaultik.Vaultik) error {
return v.PurgeSnapshotsWithOptions(opts)
@@ -258,14 +258,15 @@ Use --local-only to skip the remote half (e.g. when you want to forget a
snapshot locally without touching the destination store).
If the remote is unreachable, the local-database removal still completes
and a warning is emitted; rerun 'vaultik prune' once the destination store
is reachable to finish remote cleanup.
and a warning is emitted; run 'vaultik snapshot remove <snapshot-id>' again
once the destination store is reachable to remove the snapshot's metadata
from it ('vaultik prune' does not).
To wipe the entire destination store and start over, use 'vaultik remote
nuke --force' — it is the single supported entry point for that.`,
Args: requireSnapshotIDArg,
RunE: func(cmd *cobra.Command, args []string) error {
return runVaultikApp(cmd, mutating, opts.JSON, opts.JSON,
return runVaultikApp(cmd, mutating, opts.JSON,
"Failed to remove snapshot",
func(v *vaultik.Vaultik) error {
_, err := v.RemoveSnapshot(args[0], opts)
+30 -16
View File
@@ -33,18 +33,19 @@ const secretKeyPrefix = "AGE-SECRET-KEY-"
const (
defaultBlobSizeLimit = Size(10 * 1024 * 1024 * 1024) // 10GB
defaultChunkSize = Size(10 * 1024 * 1024) // 10MB
defaultS3PartSize = Size(5 * 1024 * 1024) // 5MB
defaultS3PartSize = Size(5 * 1024 * 1024) // 5MiB
defaultCompressionLevel = 3
minChunkSize = 1024 * 1024 // 1MB
minCompressionLevel = 1
maxCompressionLevel = 19
// S3 accepts a multipart upload part from 5MiB to 5GiB.
minS3PartSize = 5 * 1024 * 1024
maxS3PartSize = 5 * 1024 * 1024 * 1024
)
// Sentinel validation errors.
var (
errNoConfigPath = errors.New("config path not provided")
errNoAgeRecipients = errors.New(
"at least one age_recipient is required (generate with: age-keygen)")
errNoConfigPath = errors.New("config path not provided")
errRecipientIsSecretKey = errors.New(
"an age secret key was given where a public key (age1...) belongs")
errRecipientNotX25519 = errors.New(
@@ -57,6 +58,7 @@ var (
"blob_size_limit must be at least the largest chunk the chunker can " +
"emit (chunk_size times the FastCDC size spread)")
errBadCompression = errors.New("compression_level must be between 1 and 19")
errBadS3PartSize = errors.New("s3.part_size must be between 5MiB and 5GiB")
errBadStorageScheme = errors.New(
"storage_url must start with s3://, file://, or rclone://")
errStorageNotConfigured = errors.New(
@@ -246,6 +248,7 @@ func Load(path string) (*Config, error) {
ChunkSize: defaultChunkSize,
IndexPath: filepath.Join(xdg.DataHome, appName, "index.sqlite"),
CompressionLevel: defaultCompressionLevel,
S3: S3Config{PartSize: defaultS3PartSize},
}
// Convert smartconfig data to YAML then unmarshal
@@ -296,17 +299,13 @@ func Load(path string) (*Config, error) {
cfg.S3.Region = "us-east-1"
}
if cfg.S3.PartSize == 0 {
cfg.S3.PartSize = defaultS3PartSize
}
// Check config file permissions (warn if world or group readable)
//nolint:gosec // G703: config path is operator-supplied by design
info, statErr := os.Stat(path)
if statErr == nil {
mode := info.Mode().Perm()
if mode&0044 != 0 { // group or world readable
log.Warn("Config file has insecure permissions (contains S3 credentials)",
log.Warn(cfg.readableByOthersWarning(),
"path", path,
"mode", fmt.Sprintf("%04o", mode),
"recommendation", "chmod 600 "+path)
@@ -323,9 +322,10 @@ func Load(path string) (*Config, error) {
// Validate checks if the configuration is valid and complete.
// It ensures all required fields are present and have valid values:
// - At least one age recipient must be specified, and every recipient must
// parse as an X25519 age1... public key (so a bad entry fails at load, not
// mid-backup); errors name the position, never the value
// - Every age recipient must parse as an X25519 age1... public key (so a
// bad entry fails at load, not mid-backup); errors name the position,
// never the value. An empty list is accepted, because only snapshot
// create needs a recipient and it checks for one itself
// - At least one snapshot must be configured with at least one path
// - Storage must be configured (either storage_url or s3.* fields)
// - Chunk size must be at least 1MB
@@ -333,13 +333,10 @@ func Load(path string) (*Config, error) {
// (chunk_size times chunker.ChunkSizeSpread), so a single-chunk blob never
// exceeds the configured limit
// - Compression level must be between 1 and 19
// - S3 part size must be between 5MiB and 5GiB, the part sizes S3 accepts
//
// Returns an error describing the first validation failure encountered.
func (c *Config) Validate() error {
if len(c.AgeRecipients) == 0 {
return errNoAgeRecipients
}
for i, recipient := range c.AgeRecipients {
err := validateAgeRecipient(recipient)
if err != nil {
@@ -381,6 +378,11 @@ func (c *Config) Validate() error {
return errBadCompression
}
if c.S3.PartSize.Int64() < minS3PartSize ||
c.S3.PartSize.Int64() > maxS3PartSize {
return errBadS3PartSize
}
return nil
}
@@ -416,6 +418,18 @@ func (c *Config) setAgeSecretKey() {
}
}
// readableByOthersWarning is the warning Load logs when others can read
// the config file. It says "may contain" because the S3 credentials are
// seen only after smartconfig has replaced any ${...} reference in the
// file with its value, so a set credential need not be in the file.
func (c *Config) readableByOthersWarning() string {
if c.S3.AccessKeyID != "" || c.S3.SecretAccessKey != "" {
return "Config file is readable by others and may contain S3 credentials"
}
return "Config file is readable by others"
}
// validateStorage validates storage configuration.
// If StorageURL is set, it takes precedence. S3 URLs require credentials.
// File URLs don't require any S3 configuration.
+247 -2
View File
@@ -8,6 +8,7 @@ import (
"testing"
"sneak.berlin/go/vaultik/internal/chunker"
"sneak.berlin/go/vaultik/internal/log"
)
const (
@@ -166,6 +167,7 @@ func TestValidateBlobSizeLimit(t *testing.T) {
ChunkSize: chunkSize,
BlobSizeLimit: blobLimit,
CompressionLevel: 3,
S3: S3Config{PartSize: defaultS3PartSize},
}
}
@@ -221,9 +223,124 @@ func TestValidateBlobSizeLimit(t *testing.T) {
}
}
// TestValidateS3PartSize checks that s3.part_size is held to the part sizes
// S3 accepts, 5MiB to 5GiB, by changing only the part size of the test
// config. "5MB" in the config file is 5,000,000 bytes, below the minimum.
func TestValidateS3PartSize(t *testing.T) {
t.Parallel()
base, err := Load(os.Getenv("VAULTIK_CONFIG"))
if err != nil {
t.Fatalf("Failed to load config: %v", err)
}
tests := []struct {
name string
partSize Size
wantErr bool
}{
{
name: "5MB is rejected",
partSize: 5_000_000,
wantErr: true,
},
{
name: "one byte below 5MiB is rejected",
partSize: minS3PartSize - 1,
wantErr: true,
},
{
name: "5MiB is accepted",
partSize: minS3PartSize,
wantErr: false,
},
{
name: "5GiB is accepted",
partSize: maxS3PartSize,
wantErr: false,
},
{
name: "one byte above 5GiB is rejected",
partSize: maxS3PartSize + 1,
wantErr: true,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
cfg := *base
cfg.S3.PartSize = tt.partSize
err := cfg.Validate()
if tt.wantErr {
if !errors.Is(err, errBadS3PartSize) {
t.Fatalf("Validate() error = %v, want errBadS3PartSize", err)
}
return
}
if err != nil {
t.Fatalf("Validate() unexpected error: %v", err)
}
})
}
}
// TestLoadS3PartSize checks that a config file without s3.part_size loads
// with the 5MiB default, and that an explicit 0 fails at load like any other
// part size S3 refuses.
func TestLoadS3PartSize(t *testing.T) {
t.Parallel()
const withoutPartSize = "snapshots:\n" +
" test:\n" +
" paths: [/tmp/vaultik-test-source]\n" +
"storage_url: file:///tmp/vaultik-test-storage\n"
writeConfig := func(t *testing.T, text string) string {
t.Helper()
path := filepath.Join(t.TempDir(), "config.yml")
err := os.WriteFile(path, []byte(text), 0o600)
if err != nil {
t.Fatalf("write config: %v", err)
}
return path
}
t.Run("absent loads as 5MiB", func(t *testing.T) {
t.Parallel()
cfg, err := Load(writeConfig(t, withoutPartSize))
if err != nil {
t.Fatalf("Load() unexpected error: %v", err)
}
if cfg.S3.PartSize != defaultS3PartSize {
t.Errorf("s3.part_size = %d, want %d",
cfg.S3.PartSize, defaultS3PartSize)
}
})
t.Run("0 is rejected", func(t *testing.T) {
t.Parallel()
_, err := Load(writeConfig(t, withoutPartSize+"s3:\n part_size: 0\n"))
if !errors.Is(err, errBadS3PartSize) {
t.Fatalf("Load() error = %v, want errBadS3PartSize", err)
}
})
}
// TestValidateAgeRecipients checks that recipients are parsed at config load
// (a bad entry fails immediately, not mid-backup) and that no invalid entry —
// least of all a pasted secret key — is echoed in the error.
// least of all a pasted secret key — is echoed in the error. An empty list
// loads, because only snapshot create needs a recipient.
func TestValidateAgeRecipients(t *testing.T) {
t.Parallel()
@@ -235,6 +352,7 @@ func TestValidateAgeRecipients(t *testing.T) {
ChunkSize: Size(10 * 1024 * 1024),
BlobSizeLimit: Size(10 * 1024 * 1024 * 1024),
CompressionLevel: 3,
S3: S3Config{PartSize: defaultS3PartSize},
}
}
@@ -244,7 +362,12 @@ func TestValidateAgeRecipients(t *testing.T) {
wantErr bool
}{
{
name: "config init placeholder is rejected",
name: "no recipients is accepted",
recipients: nil,
wantErr: false,
},
{
name: "placeholder recipient is rejected",
recipients: []string{"age1REPLACE_WITH_YOUR_PUBLIC_KEY"},
wantErr: true,
},
@@ -337,3 +460,125 @@ func TestAgeSecretKeySourceName(t *testing.T) {
})
}
}
// loadReadableConfig writes configYAML to a file that others can read,
// loads it, and returns what the logger wrote to stderr meanwhile. The
// logger writes to the os.Stderr it finds when it is initialized, so
// os.Stderr is pointed at a file first. Not parallel-safe: os.Stderr and
// the logger are process-global.
func loadReadableConfig(t *testing.T, configYAML string) string {
t.Helper()
dir := t.TempDir()
configPath := filepath.Join(dir, "config.yml")
stderrPath := filepath.Join(dir, "stderr")
err := os.WriteFile(configPath, []byte(configYAML), 0o600)
if err != nil {
t.Fatalf("writing config: %v", err)
}
//nolint:gosec // G302: the test needs a config file others can read
err = os.Chmod(configPath, 0o644)
if err != nil {
t.Fatalf("chmod config: %v", err)
}
stderrFile, err := os.Create(stderrPath) //nolint:gosec // G304: test temp path
if err != nil {
t.Fatalf("creating stderr file: %v", err)
}
previous := os.Stderr
os.Stderr = stderrFile
log.Initialize(log.Config{})
_, loadErr := Load(configPath)
os.Stderr = previous
log.Initialize(log.Config{})
_ = stderrFile.Close()
if loadErr != nil {
t.Fatalf("Load() error = %v", loadErr)
}
captured, err := os.ReadFile(stderrPath) //nolint:gosec // G304: test temp path
if err != nil {
t.Fatalf("reading stderr file: %v", err)
}
return string(captured)
}
// TestLoadWarnsReadableConfigWithoutS3Credentials checks that a config
// file others can read, holding no S3 credentials, is warned about
// without a claim that it holds them.
//
//nolint:paralleltest // loadReadableConfig replaces os.Stderr
func TestLoadWarnsReadableConfigWithoutS3Credentials(t *testing.T) {
stderr := loadReadableConfig(t, `
storage_url: file:///var/backups/vaultik
snapshots:
home:
paths:
- /home
`)
if !strings.Contains(stderr, "Config file is readable by others") {
t.Errorf("expected a warning that the file is readable by others, got %q",
stderr)
}
if strings.Contains(stderr, "S3 credentials") {
t.Errorf("warning names S3 credentials the file does not set: %q", stderr)
}
}
// TestLoadWarnsReadableConfigWithS3Credentials checks that a config file
// others can read and that sets S3 credentials, as values or as ${ENV:...}
// references, is warned about as one that may contain them.
//
//nolint:paralleltest // loadReadableConfig replaces os.Stderr
func TestLoadWarnsReadableConfigWithS3Credentials(t *testing.T) {
t.Setenv("VAULTIK_TEST_ACCESS_KEY_ID", "test-access-key")
t.Setenv("VAULTIK_TEST_SECRET_ACCESS_KEY", "test-secret-key")
configs := map[string]string{
"values": `
storage_url: s3://bucket/prefix?endpoint=s3.example.com
s3:
access_key_id: test-access-key
secret_access_key: test-secret-key
snapshots:
home:
paths:
- /home
`,
"references": `
storage_url: s3://bucket/prefix?endpoint=s3.example.com
s3:
access_key_id: ${ENV:VAULTIK_TEST_ACCESS_KEY_ID}
secret_access_key: ${ENV:VAULTIK_TEST_SECRET_ACCESS_KEY}
snapshots:
home:
paths:
- /home
`,
}
for name, configYAML := range configs {
t.Run(name, func(t *testing.T) {
stderr := loadReadableConfig(t, configYAML)
if !strings.Contains(stderr,
"Config file is readable by others and may contain S3 credentials") {
t.Errorf("expected a warning naming the S3 credentials, got %q",
stderr)
}
})
}
}
+39 -44
View File
@@ -42,6 +42,10 @@ var schemaFS embed.FS
// table itself. It is applied before the normal migration loop.
const bootstrapVersion = 0
// busyTimeoutMs is how long a connection to the index waits for another
// connection's lock before failing with "database is locked".
const busyTimeoutMs = 10000
// DB represents the Vaultik local index database connection.
// It uses SQLite to track file metadata, content-defined chunks, and blob associations.
// The database enables incremental backups by detecting changed files and
@@ -94,10 +98,27 @@ func ParseMigrationVersion(filename string) (int, error) {
return version, nil
}
// indexDSN returns the driver DSN that opens the database at path. The
// driver runs each _pragma parameter on every connection it opens and drops
// any parameter it does not know without an error, so a setting written in
// another form silently does nothing. In WAL mode one connection can write
// while others read; the busy timeout makes a connection wait for a lock
// instead of failing at once.
func indexDSN(path string) string {
return fmt.Sprintf(
"%s?_pragma=busy_timeout(%d)&_pragma=journal_mode(WAL)"+
"&_pragma=synchronous(NORMAL)&_pragma=foreign_keys(1)",
path, busyTimeoutMs)
}
// New creates a new database connection at the specified path.
// It creates the schema if needed and configures SQLite with WAL mode for
// better concurrency. SQLite handles crash recovery automatically when
// opening a database with journal/WAL files present.
// It creates the schema if needed. Every connection runs in WAL mode with
// a busy timeout and foreign keys on (see indexDSN), so a read-only command
// can read the index while a backup writes to it. Committed rows can sit in
// the -wal file beside the database until a checkpoint, so a copy of the
// database file alone may miss them.
// SQLite handles crash recovery automatically when opening a database with
// journal/WAL files present.
// The path parameter can be a file path for persistent storage or ":memory:"
// for an in-memory database (useful for testing).
func New(ctx context.Context, path string) (*DB, error) {
@@ -110,11 +131,7 @@ func New(ctx context.Context, path string) (*DB, error) {
// First attempt with standard WAL mode
log.Debug("Attempting to open database with WAL mode", "path", path)
conn, err := sql.Open(
"sqlite",
path+"?_journal_mode=WAL&_synchronous=NORMAL&_busy_timeout=10000"+
"&_locking_mode=NORMAL&_foreign_keys=ON",
)
conn, err := sql.Open("sqlite", indexDSN(path))
if err == nil {
configureConnPool(conn)
@@ -134,8 +151,8 @@ func New(ctx context.Context, path string) (*DB, error) {
_ = conn.Close()
}
// If first attempt failed, try with TRUNCATE mode to clear any locks
return openWithRecovery(ctx, path)
// If the first attempt failed, try once more
return retryOpen(ctx, path)
}
// configureConnPool serializes all database access through one connection.
@@ -147,18 +164,12 @@ func configureConnPool(conn *sql.DB) {
conn.SetMaxIdleConns(1)
}
// finishOpen enables foreign keys, wraps the connection, and applies any
// pending migrations. On migration failure the connection is closed.
// finishOpen wraps the connection and applies any pending migrations. On
// migration failure the connection is closed.
func finishOpen(ctx context.Context, conn *sql.DB, path string) (*DB, error) {
// Enable foreign keys explicitly
_, err := conn.ExecContext(ctx, "PRAGMA foreign_keys = ON")
if err != nil {
log.Warn("Failed to enable foreign keys", "path", path, "error", err)
}
db := &DB{conn: conn, path: path}
err = applyMigrations(ctx, conn)
err := applyMigrations(ctx, conn)
if err != nil {
_ = conn.Close()
@@ -168,21 +179,15 @@ func finishOpen(ctx context.Context, conn *sql.DB, path string) (*DB, error) {
return db, nil
}
// openWithRecovery retries opening the database in TRUNCATE journal mode to
// clear stale locks, then switches back to WAL mode.
func openWithRecovery(ctx context.Context, path string) (*DB, error) {
log.Info(
"Database appears locked, attempting recovery with TRUNCATE mode",
"path", path,
)
// retryOpen makes a second attempt to open the database, with the same
// settings, after the first attempt failed, for example because another
// process held a lock for longer than the busy timeout.
func retryOpen(ctx context.Context, path string) (*DB, error) {
log.Info("Database appears locked, retrying open", "path", path)
conn, err := sql.Open(
"sqlite",
path+"?_journal_mode=TRUNCATE&_synchronous=NORMAL&_busy_timeout=10000"+
"&_foreign_keys=ON",
)
conn, err := sql.Open("sqlite", indexDSN(path))
if err != nil {
return nil, fmt.Errorf("opening database in recovery mode: %w", err)
return nil, fmt.Errorf("opening database on retry: %w", err)
}
configureConnPool(conn)
@@ -190,28 +195,18 @@ func openWithRecovery(ctx context.Context, path string) (*DB, error) {
err = conn.PingContext(ctx)
if err != nil {
log.Debug(
"Failed to ping database in recovery mode, closing",
"Failed to ping database on retry, closing",
"path", path, "error", err,
)
_ = conn.Close()
return nil, fmt.Errorf(
"database still locked after recovery attempt: %w",
"database still locked on retry: %w",
err,
)
}
log.Debug("Database opened in TRUNCATE mode", "path", path)
// Switch back to WAL mode
log.Debug("Switching database back to WAL mode", "path", path)
_, err = conn.ExecContext(ctx, "PRAGMA journal_mode=WAL")
if err != nil {
log.Warn("Failed to switch back to WAL mode", "path", path, "error", err)
}
db, err := finishOpen(ctx, conn, path)
if err != nil {
return nil, err
+83
View File
@@ -120,6 +120,89 @@ func TestDatabaseConcurrentAccess(t *testing.T) {
}
}
// TestNewSetsJournalModeAndBusyTimeout checks that the connection settings
// New passes reach SQLite. The driver drops a setting it does not recognise
// without an error, so only reading the value back shows it took effect.
func TestNewSetsJournalModeAndBusyTimeout(t *testing.T) {
t.Parallel()
ctx := context.Background()
db, err := New(ctx, filepath.Join(t.TempDir(), "index.db"))
if err != nil {
t.Fatalf("failed to create database: %v", err)
}
defer func() { _ = db.Close() }()
var journalMode string
err = db.conn.QueryRowContext(ctx, "PRAGMA journal_mode").Scan(&journalMode)
if err != nil {
t.Fatalf("reading journal_mode: %v", err)
}
if journalMode != "wal" {
t.Errorf("journal_mode = %q, want %q", journalMode, "wal")
}
var busyTimeout int
err = db.conn.QueryRowContext(ctx, "PRAGMA busy_timeout").Scan(&busyTimeout)
if err != nil {
t.Fatalf("reading busy_timeout: %v", err)
}
if busyTimeout != busyTimeoutMs {
t.Errorf("busy_timeout = %d, want %d", busyTimeout, busyTimeoutMs)
}
}
// TestNewWriteSucceedsWhileAnotherHandleReads opens the same index twice,
// as a read-only command does while a backup runs, and checks that a write
// on one handle commits while the other is in the middle of a read.
func TestNewWriteSucceedsWhileAnotherHandleReads(t *testing.T) {
t.Parallel()
ctx := context.Background()
dbPath := filepath.Join(t.TempDir(), "index.db")
reader, err := New(ctx, dbPath)
if err != nil {
t.Fatalf("failed to open reading handle: %v", err)
}
defer func() { _ = reader.Close() }()
writer, err := New(ctx, dbPath)
if err != nil {
t.Fatalf("failed to open writing handle: %v", err)
}
defer func() { _ = writer.Close() }()
// The read lock taken by the SELECT is held until the transaction ends.
readTx, err := reader.BeginTx(ctx, nil)
if err != nil {
t.Fatalf("beginning read transaction: %v", err)
}
defer func() { _ = readTx.Rollback() }()
var count int
err = readTx.QueryRowContext(ctx, "SELECT COUNT(*) FROM chunks").Scan(&count)
if err != nil {
t.Fatalf("reading chunks: %v", err)
}
_, err = writer.ExecWithLog(ctx,
"INSERT INTO chunks (chunk_hash, size) VALUES (?, ?)", "hash", 1024)
if err != nil {
t.Fatalf("write while another handle reads: %v", err)
}
}
func TestParseMigrationVersion(t *testing.T) {
t.Parallel()
+39 -25
View File
@@ -33,11 +33,13 @@ func (r *FileRepository) Create(ctx context.Context, tx *sql.Tx, file *File) err
}
query := `
INSERT INTO files (id, path, source_path, mtime, size, mode, uid, gid, link_target)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)
INSERT INTO files
(id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target)
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
ON CONFLICT(path) DO UPDATE SET
source_path = excluded.source_path,
mtime = excluded.mtime,
mtime_nsec = excluded.mtime_nsec,
size = excluded.size,
mode = excluded.mode,
uid = excluded.uid,
@@ -54,16 +56,19 @@ func (r *FileRepository) Create(ctx context.Context, tx *sql.Tx, file *File) err
if tx != nil {
LogSQL("Execute", query,
file.ID.String(), file.Path.String(), file.SourcePath.String(),
file.MTime.Unix(), file.Size, file.Mode, file.UID, file.GID,
file.MTime.Unix(), file.MTime.Nanosecond(),
file.Size, file.Mode, file.UID, file.GID,
file.LinkTarget.String())
err = tx.QueryRowContext(ctx, query,
file.ID.String(), file.Path.String(), file.SourcePath.String(),
file.MTime.Unix(), file.Size, file.Mode, file.UID, file.GID,
file.MTime.Unix(), file.MTime.Nanosecond(),
file.Size, file.Mode, file.UID, file.GID,
file.LinkTarget.String()).Scan(&idStr)
} else {
err = r.db.QueryRowWithLog(ctx, query,
file.ID.String(), file.Path.String(), file.SourcePath.String(),
file.MTime.Unix(), file.Size, file.Mode, file.UID, file.GID,
file.MTime.Unix(), file.MTime.Nanosecond(),
file.Size, file.Mode, file.UID, file.GID,
file.LinkTarget.String()).Scan(&idStr)
}
@@ -84,7 +89,7 @@ func (r *FileRepository) Create(ctx context.Context, tx *sql.Tx, file *File) err
// in the index.
func (r *FileRepository) GetByPath(ctx context.Context, path string) (*File, error) {
query := `
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
FROM files
WHERE path = ?
`
@@ -104,7 +109,7 @@ func (r *FileRepository) GetByPath(ctx context.Context, path string) (*File, err
// GetByID retrieves a file by its UUID
func (r *FileRepository) GetByID(ctx context.Context, id types.FileID) (*File, error) {
query := `
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
FROM files
WHERE id = ?
`
@@ -127,7 +132,7 @@ func (r *FileRepository) GetByPathTx(
ctx context.Context, tx *sql.Tx, path string,
) (*File, error) {
query := `
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
FROM files
WHERE path = ?
`
@@ -158,13 +163,14 @@ func (r *FileRepository) ListModifiedSince(
ctx context.Context, since time.Time,
) ([]*File, error) {
query := `
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
FROM files
WHERE mtime >= ?
WHERE (mtime, mtime_nsec) >= (?, ?)
ORDER BY path
`
rows, err := r.db.conn.QueryContext(ctx, query, since.Unix())
rows, err := r.db.conn.QueryContext(ctx, query,
since.Unix(), since.Nanosecond())
if err != nil {
return nil, fmt.Errorf("querying files: %w", err)
}
@@ -228,19 +234,24 @@ func (r *FileRepository) DeleteByID(
return nil
}
// ListByPrefix returns all files whose path starts with prefix, ordered by
// path.
func (r *FileRepository) ListByPrefix(
ctx context.Context, prefix string,
// ListUnderPath returns the file at path and every file beneath it,
// ordered by path. Paths are compared case-sensitively, and a trailing
// slash on path is ignored, so "/" lists every file.
func (r *FileRepository) ListUnderPath(
ctx context.Context, path string,
) ([]*File, error) {
path = strings.TrimRight(path, "/")
dirPrefix := path + "/"
// LIKE would ignore ASCII case and treat _ and % in path as wildcards.
query := `
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
FROM files
WHERE path LIKE ? || '%'
WHERE path = ? OR substr(path, 1, length(?)) = ?
ORDER BY path
`
rows, err := r.db.conn.QueryContext(ctx, query, prefix)
rows, err := r.db.conn.QueryContext(ctx, query, path, dirPrefix, dirPrefix)
if err != nil {
return nil, fmt.Errorf("querying files: %w", err)
}
@@ -319,7 +330,7 @@ func (r *FileRepository) ListIDsWithChunksNotInUploadedBlobs(
// ListAll returns all files in the database
func (r *FileRepository) ListAll(ctx context.Context) ([]*File, error) {
query := `
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
FROM files
ORDER BY path
`
@@ -360,7 +371,7 @@ func (r *FileRepository) CreateBatch(
}
// Each files row binds this many SQL variables.
const fileCols = 9
const fileCols = 10
// Batch at 100 rows to be safe with SQLite's variable limit.
const batchSize = 100
@@ -371,7 +382,7 @@ func (r *FileRepository) CreateBatch(
batch := files[i:end]
query := `INSERT INTO files
(id, path, source_path, mtime, size, mode, uid, gid, link_target)
(id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target)
VALUES `
args := make([]any, 0, len(batch)*fileCols)
@@ -383,11 +394,12 @@ func (r *FileRepository) CreateBatch(
querySb325.WriteString(", ")
}
querySb325.WriteString("(?, ?, ?, ?, ?, ?, ?, ?, ?)")
querySb325.WriteString("(?, ?, ?, ?, ?, ?, ?, ?, ?, ?)")
args = append(args,
f.ID.String(), f.Path.String(), f.SourcePath.String(),
f.MTime.Unix(), f.Size, f.Mode, f.UID, f.GID,
f.MTime.Unix(), f.MTime.Nanosecond(),
f.Size, f.Mode, f.UID, f.GID,
f.LinkTarget.String())
}
@@ -396,6 +408,7 @@ func (r *FileRepository) CreateBatch(
query += ` ON CONFLICT(path) DO UPDATE SET
source_path = excluded.source_path,
mtime = excluded.mtime,
mtime_nsec = excluded.mtime_nsec,
size = excluded.size,
mode = excluded.mode,
uid = excluded.uid,
@@ -455,7 +468,7 @@ func (r *FileRepository) scanFileFrom(row fileRowScanner) (*File, error) {
var (
file File
idStr, pathStr, sourcePathStr string
mtimeUnix int64
mtimeUnix, mtimeNsec int64
linkTarget sql.NullString
)
@@ -464,6 +477,7 @@ func (r *FileRepository) scanFileFrom(row fileRowScanner) (*File, error) {
&pathStr,
&sourcePathStr,
&mtimeUnix,
&mtimeNsec,
&file.Size,
&file.Mode,
&file.UID,
@@ -482,7 +496,7 @@ func (r *FileRepository) scanFileFrom(row fileRowScanner) (*File, error) {
file.Path = types.FilePath(pathStr)
file.SourcePath = types.SourcePath(sourcePathStr)
file.MTime = time.Unix(mtimeUnix, 0).UTC()
file.MTime = time.Unix(mtimeUnix, mtimeNsec).UTC()
if linkTarget.Valid {
file.LinkTarget = types.FilePath(linkTarget.String)
}
+187
View File
@@ -5,10 +5,12 @@ import (
"database/sql"
"errors"
"os"
"slices"
"testing"
"time"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/types"
)
// errTestRollback is the sentinel returned from transaction bodies to
@@ -134,6 +136,82 @@ func TestFileRepositoryListDelete(t *testing.T) {
}
}
func TestFileRepositoryListUnderPath(t *testing.T) {
t.Parallel()
db, cleanup := setupTestDB(t)
defer cleanup()
ctx := context.Background()
repo := database.NewFileRepository(db)
const (
docDir = "/home/u/doc"
docFile = "/home/u/doc/a.txt"
)
// In path order, so the root case can expect all of them as listed.
paths := []string{
"/home/u/50%/x.txt",
"/home/u/50percent/y.txt",
"/home/u/DOC/c.txt",
"/home/u/a_b/x.txt",
"/home/u/axb/y.txt",
docDir,
"/home/u/doc.txt.bak",
docFile,
"/home/u/doc/sub/b.txt",
"/home/u/doc2/b.txt",
}
for _, path := range paths {
err := repo.Create(ctx, nil, &database.File{
Path: types.FilePath(path),
MTime: time.Now().Truncate(time.Second),
Mode: 0644,
})
if err != nil {
t.Fatalf("failed to create %s: %v", path, err)
}
}
docTree := []string{docDir, docFile, "/home/u/doc/sub/b.txt"}
tests := []struct {
name string
path string
want []string
}{
{"directory", docDir, docTree},
{"directory with trailing slash", docDir + "/", docTree},
{"directory differing only in case", "/home/u/DOC",
[]string{"/home/u/DOC/c.txt"}},
{"file", docFile, []string{docFile}},
{"underscore is literal", "/home/u/a_b",
[]string{"/home/u/a_b/x.txt"}},
{"percent is literal", "/home/u/50%",
[]string{"/home/u/50%/x.txt"}},
{"root", "/", paths},
}
for _, tt := range tests {
files, err := repo.ListUnderPath(ctx, tt.path)
if err != nil {
t.Fatalf("%s: failed to list files: %v", tt.name, err)
}
got := make([]string, 0, len(files))
for _, f := range files {
got = append(got, f.Path.String())
}
if !slices.Equal(got, tt.want) {
t.Errorf("%s: listing %q got %q, want %q",
tt.name, tt.path, got, tt.want)
}
}
}
func TestFileRepositorySymlink(t *testing.T) {
t.Parallel()
@@ -174,6 +252,115 @@ func TestFileRepositorySymlink(t *testing.T) {
}
}
// An mtime after 2262 or before 1678 does not fit in int64 nanoseconds
// since the epoch, and must still come back from the database unchanged.
func TestFileRepositoryMTimeOutsideInt64NanosecondRange(t *testing.T) {
t.Parallel()
db, cleanup := setupTestDB(t)
defer cleanup()
ctx := context.Background()
repo := database.NewFileRepository(db)
mtimes := []time.Time{
time.Date(2300, time.January, 1, 0, 0, 0, 123456789, time.UTC),
time.Date(1601, time.January, 1, 0, 0, 0, 987654321, time.UTC),
}
for _, mtime := range mtimes {
created := &database.File{
Path: types.FilePath("/created-" + mtime.Format(time.RFC3339Nano)),
MTime: mtime,
}
err := repo.Create(ctx, nil, created)
if err != nil {
t.Fatalf("failed to create file: %v", err)
}
batched := &database.File{
ID: types.NewFileID(),
Path: types.FilePath("/batched-" + mtime.Format(time.RFC3339Nano)),
MTime: mtime,
}
err = repo.CreateBatch(ctx, nil, []*database.File{batched})
if err != nil {
t.Fatalf("failed to batch create file: %v", err)
}
for _, path := range []types.FilePath{created.Path, batched.Path} {
retrieved, err := repo.GetByPath(ctx, path.String())
if err != nil {
t.Fatalf("failed to get file: %v", err)
}
if !retrieved.MTime.Equal(mtime) {
t.Errorf("%s: mtime got %v, want %v",
path, retrieved.MTime, mtime)
}
}
}
}
// A file already in the index and rewritten within the same second must get
// its new nanoseconds stored, through both Create and CreateBatch.
func TestFileRepositoryUpsertMTimeInSameSecond(t *testing.T) {
t.Parallel()
db, cleanup := setupTestDB(t)
defer cleanup()
ctx := context.Background()
repo := database.NewFileRepository(db)
indexed := time.Date(2026, time.October, 7, 12, 0, 0, 100000000, time.UTC)
rewritten := indexed.Add(800 * time.Millisecond)
tests := []struct {
name string
upsert func(file *database.File) error
}{
{"Create", func(file *database.File) error {
return repo.Create(ctx, nil, file)
}},
{"CreateBatch", func(file *database.File) error {
return repo.CreateBatch(ctx, nil, []*database.File{file})
}},
}
for _, tt := range tests {
file := &database.File{
ID: types.NewFileID(),
Path: types.FilePath("/" + tt.name),
MTime: indexed,
}
err := tt.upsert(file)
if err != nil {
t.Fatalf("%s: failed to create file: %v", tt.name, err)
}
file.MTime = rewritten
err = tt.upsert(file)
if err != nil {
t.Fatalf("%s: failed to update file: %v", tt.name, err)
}
retrieved, err := repo.GetByPath(ctx, file.Path.String())
if err != nil {
t.Fatalf("%s: failed to get file: %v", tt.name, err)
}
if !retrieved.MTime.Equal(rewritten) {
t.Errorf("%s: mtime got %v, want %v",
tt.name, retrieved.MTime, rewritten)
}
}
}
func TestFileRepositoryTransaction(t *testing.T) {
t.Parallel()
+2 -2
View File
@@ -99,8 +99,8 @@ type Snapshot struct {
StartedAt time.Time
CompletedAt *time.Time // nil if still in progress
FileCount int64
ChunkCount int64
BlobCount int64
ChunkCount int64 // Chunks this snapshot stored that were not stored before
BlobCount int64 // Blobs this snapshot created
TotalSize int64 // Total size of all referenced files
// BlobSize is the total size of all referenced blobs (compressed and
@@ -824,7 +824,7 @@ func TestTransactionIsolation(t *testing.T) {
}
// Verify the file was not created (transaction rolled back)
files, err := repos.Files.ListByPrefix(ctx, "/tx-test")
files, err := repos.Files.ListUnderPath(ctx, "/tx-test.txt")
if err != nil {
t.Fatal(err)
}
@@ -916,7 +916,7 @@ func TestConcurrentOrphanedCleanup(t *testing.T) {
}
// Verify correct files were deleted
files, err := repos.Files.ListByPrefix(ctx, "/concurrent-")
files, err := repos.Files.ListAll(ctx)
if err != nil {
t.Fatal(err)
}
+1 -1
View File
@@ -147,7 +147,7 @@ func TestOrphanedFileCleanupDebug(t *testing.T) {
t.Logf("Files count after cleanup: %d", count)
// List remaining files
files, err := repos.Files.ListByPrefix(ctx, "/")
files, err := repos.Files.ListUnderPath(ctx, "/")
if err != nil {
t.Fatal(err)
}
@@ -442,12 +442,12 @@ func TestLargeDatasets(t *testing.T) {
createLargeDatasetFiles(t, repos, snapshot.ID.String(), fileCount)
})
// Test ListByPrefix performance
// Test ListUnderPath performance
//nolint:paralleltest // phases share one database and are order-dependent
t.Run("list by prefix performance", func(t *testing.T) {
t.Run("list under path performance", func(t *testing.T) {
start := time.Now()
files, err := repos.Files.ListByPrefix(ctx, "/large/")
files, err := repos.Files.ListUnderPath(ctx, "/large/")
if err != nil {
t.Fatal(err)
}
@@ -472,7 +472,7 @@ func TestLargeDatasets(t *testing.T) {
t.Logf("Cleaned up orphaned files in %v", time.Since(start))
// Verify correct number remain
files, err := repos.Files.ListByPrefix(ctx, "/large/")
files, err := repos.Files.ListUnderPath(ctx, "/large/")
if err != nil {
t.Fatal(err)
}
@@ -606,8 +606,7 @@ func TestTimezoneHandling(t *testing.T) {
t.Skip("timezone not available")
}
// Use Truncate to remove sub-second precision since we store as Unix timestamps
nyTime := time.Now().In(loc).Truncate(time.Second)
nyTime := time.Now().In(loc)
file := &File{
Path: "/timezone-test.txt",
MTime: nyTime,
+2 -1
View File
@@ -6,7 +6,8 @@ CREATE TABLE IF NOT EXISTS files (
id TEXT PRIMARY KEY, -- UUID
path TEXT NOT NULL UNIQUE,
source_path TEXT NOT NULL DEFAULT '', -- The source directory this file came from (for restore path stripping)
mtime INTEGER NOT NULL,
mtime INTEGER NOT NULL, -- whole seconds since the Unix epoch
mtime_nsec INTEGER NOT NULL, -- nanoseconds within that second, 0 to 999999999
size INTEGER NOT NULL,
mode INTEGER NOT NULL,
uid INTEGER NOT NULL,
+28 -3
View File
@@ -127,6 +127,7 @@ func (r *SnapshotRepository) UpdateExtendedStats(
snapshotID string,
blobUncompressedSize int64,
compressionLevel int,
uploadBytes int64,
uploadDurationMs int64,
) error {
compressionRatio, err := r.extendedCompressionRatio(
@@ -141,7 +142,7 @@ func (r *SnapshotRepository) UpdateExtendedStats(
SET blob_uncompressed_size = ?,
compression_ratio = ?,
compression_level = ?,
upload_bytes = blob_size,
upload_bytes = ?,
upload_duration_ms = ?
WHERE id = ?
`
@@ -149,11 +150,11 @@ func (r *SnapshotRepository) UpdateExtendedStats(
if tx != nil {
_, err = tx.ExecContext(ctx, query,
blobUncompressedSize, compressionRatio, compressionLevel,
uploadDurationMs, snapshotID)
uploadBytes, uploadDurationMs, snapshotID)
} else {
_, err = r.db.ExecWithLog(ctx, query,
blobUncompressedSize, compressionRatio, compressionLevel,
uploadDurationMs, snapshotID)
uploadBytes, uploadDurationMs, snapshotID)
}
if err != nil {
@@ -543,6 +544,30 @@ func (r *SnapshotRepository) GetSnapshotTotalCompressedSize(
return totalSize, nil
}
// GetSnapshotBlobSizes returns the total compressed and uncompressed sizes
// of all blobs referenced by a snapshot.
func (r *SnapshotRepository) GetSnapshotBlobSizes(
ctx context.Context, snapshotID string,
) (int64, int64, error) {
query := `
SELECT COALESCE(SUM(b.compressed_size), 0),
COALESCE(SUM(b.uncompressed_size), 0)
FROM snapshot_blobs sb
JOIN blobs b ON sb.blob_hash = b.blob_hash
WHERE sb.snapshot_id = ?
`
var compressed, uncompressed int64
err := r.db.conn.QueryRowContext(ctx, query, snapshotID).Scan(
&compressed, &uncompressed)
if err != nil {
return 0, 0, fmt.Errorf("querying snapshot blob sizes: %w", err)
}
return compressed, uncompressed, nil
}
// GetSnapshotUncompressedChunkSize returns the sum of plaintext sizes of all unique
// chunks referenced by a snapshot (via snapshot_files → file_chunks → chunks).
func (r *SnapshotRepository) GetSnapshotUncompressedChunkSize(
+59
View File
@@ -145,6 +145,65 @@ func TestSnapshotRepositoryUpdateCounts(t *testing.T) {
}
}
// GetSnapshotBlobSizes totals the blobs the snapshot references, and only
// those.
func TestSnapshotRepositoryGetSnapshotBlobSizes(t *testing.T) {
t.Parallel()
db, cleanup := setupTestDB(t)
defer cleanup()
ctx := context.Background()
repos := database.NewRepositories(db)
snapshot := &database.Snapshot{
ID: "2024-01-03T12:00:00Z",
Hostname: testHostname,
VaultikVersion: testVersion,
StartedAt: time.Now().Truncate(time.Second),
}
err := repos.Snapshots.Create(ctx, nil, snapshot)
if err != nil {
t.Fatalf("failed to create snapshot: %v", err)
}
blobs := []*database.Blob{
{Hash: "referenced-1", CompressedSize: 10, UncompressedSize: 100},
{Hash: "referenced-2", CompressedSize: 20, UncompressedSize: 200},
{Hash: "unreferenced", CompressedSize: 40, UncompressedSize: 400},
}
for _, blob := range blobs {
blob.ID = types.NewBlobID()
blob.CreatedTS = time.Now().Truncate(time.Second)
err = repos.Blobs.Create(ctx, nil, blob)
if err != nil {
t.Fatalf("failed to create blob %s: %v", blob.Hash, err)
}
}
for _, blob := range blobs[:2] {
err = repos.Snapshots.AddBlob(ctx, nil, snapshot.ID.String(),
blob.ID, blob.Hash)
if err != nil {
t.Fatalf("failed to add blob %s to snapshot: %v", blob.Hash, err)
}
}
compressed, uncompressed, err := repos.Snapshots.GetSnapshotBlobSizes(
ctx, snapshot.ID.String())
if err != nil {
t.Fatalf("failed to get snapshot blob sizes: %v", err)
}
if compressed != 30 || uncompressed != 300 {
t.Errorf("blob sizes: got %d and %d, want 30 and 300",
compressed, uncompressed)
}
}
func TestSnapshotRepositoryListRecent(t *testing.T) {
t.Parallel()
-16
View File
@@ -158,19 +158,3 @@ type UploadStats struct {
MinDurationMs int64
MaxDurationMs int64
}
// GetCountBySnapshot returns the count of uploads for a specific snapshot
func (r *UploadRepository) GetCountBySnapshot(
ctx context.Context, snapshotID string,
) (int64, error) {
query := `SELECT COUNT(*) FROM uploads WHERE snapshot_id = ?`
var count int64
err := r.conn.QueryRowContext(ctx, query, snapshotID).Scan(&count)
if err != nil {
return 0, err
}
return count, nil
}
+10 -6
View File
@@ -18,14 +18,19 @@ var Appname = "vaultik" //nolint:gochecknoglobals // set via -ldflags at build t
// deliberately not a number.
const DevVersion = "dev"
// Unknown is what Commit and CommitDate hold when the build did not
// stamp them, and the version script/docker and script/cibuild stamp
// when the host has no git checkout.
const Unknown = "unknown"
// Version is the application version, populated from main().
var Version = DevVersion //nolint:gochecknoglobals // set via -ldflags at build time
// Commit is the git commit hash, populated from main().
var Commit = "unknown" //nolint:gochecknoglobals // set via -ldflags at build time
var Commit = Unknown //nolint:gochecknoglobals // set via -ldflags at build time
// CommitDate is the ISO-8601 date of the commit, populated from main().
var CommitDate = "unknown" //nolint:gochecknoglobals // set via -ldflags at build time
var CommitDate = Unknown //nolint:gochecknoglobals // set via -ldflags at build time
// Author identifies the upstream author of vaultik.
const Author = "Jeffrey Paul <sneak@sneak.berlin>"
@@ -72,11 +77,10 @@ func New() (*Globals, error) {
// safe reading of "we could not establish that this is a release" is
// that it is not one. The Makefile refuses to build at all in that
// case; this is the second line of defence, for a binary linked by
// something other than the Makefile. "unknown" counts for the same
// reason: script/docker and script/cibuild stamp it when the host has
// no git checkout.
// something other than the Makefile. Unknown counts for the same
// reason.
func IsDevVersion(v string) bool {
if v == "" || v == "unknown" || v == DevVersion ||
if v == "" || v == Unknown || v == DevVersion ||
strings.HasPrefix(v, DevVersion+"-") ||
strings.HasSuffix(v, "-dirty") {
return true
+67 -52
View File
@@ -10,15 +10,18 @@ import (
"path/filepath"
"strconv"
"strings"
"syscall"
"golang.org/x/sys/unix"
)
// ErrAlreadyRunning indicates another vaultik instance is running.
var ErrAlreadyRunning = errors.New("another vaultik instance is already running")
// Lock represents an acquired PID lock.
// Lock represents an acquired PID lock: an flock(2) on the PID file,
// held while the file stays open. The kernel drops it when the process
// exits, however it exits, so a crashed run never leaves the lock held.
type Lock struct {
path string
file *os.File
}
const (
@@ -29,10 +32,9 @@ const (
)
// Acquire attempts to acquire a PID lock in the specified directory.
// If the lock file exists and the process is still running, it returns
// ErrAlreadyRunning with details about the existing process.
// On success, it writes the current PID to the lock file and returns
// a Lock that must be released with Release().
// If another process holds the lock, it returns ErrAlreadyRunning with
// that process's PID. On success, it writes the current PID to the lock
// file and returns a Lock that must be released with Release().
func Acquire(lockDir string) (*Lock, error) {
// Ensure lock directory exists
err := os.MkdirAll(lockDir, lockDirPerm)
@@ -42,56 +44,82 @@ func Acquire(lockDir string) (*Lock, error) {
lockPath := filepath.Join(lockDir, "vaultik.pid")
// Check for existing lock
existingPID, err := readPIDFile(lockPath)
if err == nil {
// Lock file exists, check if process is running
if isProcessRunning(existingPID) {
return nil, fmt.Errorf("%w (PID %d)", ErrAlreadyRunning, existingPID)
}
// Process is not running, stale lock file - we can take over
}
// Write our PID
pid := os.Getpid()
err = os.WriteFile(lockPath, []byte(strconv.Itoa(pid)), pidFilePerm)
// No O_TRUNC: the file may hold the PID of the process that has the
// lock, which the error below reports.
file, err := os.OpenFile( //nolint:gosec // G304: path is our own lock file
lockPath, os.O_RDWR|os.O_CREATE, pidFilePerm)
if err != nil {
return nil, fmt.Errorf("writing PID file: %w", err)
return nil, fmt.Errorf("opening PID file: %w", err)
}
return &Lock{path: lockPath}, nil
err = unix.Flock(int(file.Fd()), unix.LOCK_EX|unix.LOCK_NB)
if err != nil {
_ = file.Close()
if errors.Is(err, unix.EWOULDBLOCK) {
return nil, alreadyRunningError(lockPath)
}
return nil, fmt.Errorf("locking PID file: %w", err)
}
err = writePID(file)
if err != nil {
_ = file.Close()
return nil, err
}
return &Lock{file: file}, nil
}
// Release removes the PID lock file.
// Release empties the PID file and closes it, which drops the lock.
// It is safe to call Release multiple times.
func (l *Lock) Release() error {
if l == nil || l.path == "" {
if l == nil || l.file == nil {
return nil
}
// Verify we still own the lock (our PID is in the file)
existingPID, err := readPIDFile(l.path)
file := l.file
l.file = nil
// Do not remove the file here. A process that opened it a moment
// earlier could then lock the removed file while another creates and
// locks a new one, and both would run.
truncateErr := file.Truncate(0)
closeErr := file.Close()
return errors.Join(truncateErr, closeErr)
}
// writePID replaces the contents of the locked PID file with the current
// PID.
func writePID(file *os.File) error {
err := file.Truncate(0)
if err != nil {
// File already gone or unreadable - that's fine
return nil //nolint:nilerr // unreadable lock file means nothing to release
return fmt.Errorf("truncating PID file: %w", err)
}
if existingPID != os.Getpid() {
// Someone else wrote to our lock file - don't remove it
return nil
_, err = file.WriteAt([]byte(strconv.Itoa(os.Getpid())), 0)
if err != nil {
return fmt.Errorf("writing PID file: %w", err)
}
err = os.Remove(l.path)
if err != nil && !os.IsNotExist(err) {
return fmt.Errorf("removing PID file: %w", err)
}
l.path = "" // Prevent double-release
return nil
}
// alreadyRunningError reports that another process holds the lock,
// naming its PID when the file holds one. The holder writes its PID just
// after it locks, so the file can briefly be empty.
func alreadyRunningError(lockPath string) error {
pid, err := readPIDFile(lockPath)
if err != nil {
return ErrAlreadyRunning
}
return fmt.Errorf("%w (PID %d)", ErrAlreadyRunning, pid)
}
// readPIDFile reads and parses the PID from a lock file.
func readPIDFile(path string) (int, error) {
data, err := os.ReadFile(path) //nolint:gosec // G304: path is our own lock file
@@ -106,16 +134,3 @@ func readPIDFile(path string) (int, error) {
return pid, nil
}
// isProcessRunning checks if a process with the given PID is running.
func isProcessRunning(pid int) bool {
process, err := os.FindProcess(pid)
if err != nil {
return false
}
// On Unix, FindProcess always succeeds. We need to send signal 0 to check.
err = process.Signal(syscall.Signal(0))
return err == nil
}
+63 -3
View File
@@ -4,6 +4,7 @@ import (
"os"
"path/filepath"
"strconv"
"sync"
"testing"
"github.com/stretchr/testify/assert"
@@ -33,9 +34,10 @@ func TestAcquireAndRelease(t *testing.T) {
err = lock.Release()
require.NoError(t, err)
// Verify PID file is gone
_, err = os.Stat(pidPath)
assert.True(t, os.IsNotExist(err))
// Verify PID file is empty
data, err = os.ReadFile(pidPath) //nolint:gosec // G304: test's own temp file
require.NoError(t, err)
assert.Empty(t, data)
}
func TestAcquireBlocksSecondInstance(t *testing.T) {
@@ -55,6 +57,64 @@ func TestAcquireBlocksSecondInstance(t *testing.T) {
lock2, err := pidlock.Acquire(tmpDir)
require.ErrorIs(t, err, pidlock.ErrAlreadyRunning)
assert.Nil(t, lock2)
// Once the first lock is released, the next Acquire succeeds
require.NoError(t, lock1.Release())
lock3, err := pidlock.Acquire(tmpDir)
require.NoError(t, err)
require.NoError(t, lock3.Release())
}
// TestConcurrentAcquireAdmitsOne starts many Acquire calls at the same
// moment, as two cron entries firing together would, and checks that
// exactly one of them gets the lock.
func TestConcurrentAcquireAdmitsOne(t *testing.T) {
t.Parallel()
const callers = 50
tmpDir := t.TempDir()
start := make(chan struct{})
var (
mu sync.Mutex
acquired []*pidlock.Lock
failures []error
wg sync.WaitGroup
)
for range callers {
wg.Go(func() {
<-start
lock, err := pidlock.Acquire(tmpDir)
mu.Lock()
defer mu.Unlock()
if err != nil {
failures = append(failures, err)
return
}
acquired = append(acquired, lock)
})
}
close(start)
wg.Wait()
for _, lock := range acquired {
require.NoError(t, lock.Release())
}
assert.Len(t, acquired, 1, "exactly one caller should hold the lock")
for _, err := range failures {
require.ErrorIs(t, err, pidlock.ErrAlreadyRunning)
}
}
func TestAcquireWithStaleLock(t *testing.T) {
+34 -6
View File
@@ -6,6 +6,7 @@ import (
"context"
"errors"
"io"
"strings"
"sync/atomic"
"github.com/aws/aws-sdk-go-v2/aws"
@@ -26,10 +27,14 @@ type Client struct {
bucket string
prefix string
endpoint string
partSize int64
}
// Config contains S3 client configuration.
// All fields are required except Prefix, which defaults to an empty string.
// All fields are required except Prefix, which defaults to an empty string,
// and PartSize, where zero means the SDK default of 5 MiB.
// A non-empty Prefix is joined to every key with one "/", whether or not
// it ends with one.
// The Endpoint field should include the protocol (http:// or https://).
type Config struct {
Endpoint string
@@ -38,6 +43,9 @@ type Config struct {
AccessKeyID string
SecretAccessKey string
Region string
// PartSize is the size in bytes of each part of a multipart upload.
// An upload too large for S3's limit of 10,000 parts gets larger parts.
PartSize int64
}
// nopLogger is a logger that discards all output.
@@ -75,11 +83,19 @@ func NewClient(ctx context.Context, cfg Config) (*Client, error) {
s3Client := s3.NewFromConfig(awsCfg, s3Opts)
// Every method below builds a key as prefix + key, so the prefix
// must carry its own trailing "/".
prefix := strings.TrimRight(cfg.Prefix, "/")
if prefix != "" {
prefix += "/"
}
return &Client{
s3Client: s3Client,
bucket: cfg.Bucket,
prefix: cfg.Prefix,
prefix: prefix,
endpoint: cfg.Endpoint,
partSize: cfg.PartSize,
}, nil
}
@@ -113,12 +129,9 @@ func (c *Client) PutObjectWithProgress(
) error {
fullKey := c.prefix + key
// uploadPartSize is 10MB for better progress granularity.
const uploadPartSize = 10 * 1024 * 1024
// Create an uploader with the S3 client
uploader := manager.NewUploader(c.s3Client, func(u *manager.Uploader) {
u.PartSize = uploadPartSize
u.PartSize = uploadPartSize(c.partSize, size)
})
// Create a progress reader that tracks upload progress
@@ -139,6 +152,21 @@ func (c *Client) PutObjectWithProgress(
return err
}
// uploadPartSize returns the part size for an upload of size bytes: the
// configured part size (the SDK default when zero), raised where needed so
// the upload fits in S3's limit of 10,000 parts. The uploader cannot raise
// it itself, because it cannot seek the progress reader to learn its size.
func uploadPartSize(configured, size int64) int64 {
if configured == 0 {
configured = manager.DefaultUploadPartSize
}
maxParts := int64(manager.MaxUploadParts)
smallestThatFits := (size + maxParts - 1) / maxParts // rounded up
return max(configured, smallestThatFits)
}
// GetObject downloads an object from S3 with the specified key.
// The key is automatically prefixed with the configured prefix.
// Returns a ReadCloser containing the object data. The caller must
+55
View File
@@ -0,0 +1,55 @@
package s3
import "testing"
// TestUploadPartSize checks that an upload too large for 10,000 parts of the
// configured size gets parts just large enough to fit in 10,000.
func TestUploadPartSize(t *testing.T) {
t.Parallel()
const mib = 1024 * 1024
tests := []struct {
name string
configured int64
size int64
want int64
}{
{
name: "an upload that fits keeps the configured size",
configured: 5 * mib,
size: 10 * 1024 * mib,
want: 5 * mib,
},
{
name: "exactly 10,000 parts keeps the configured size",
configured: 6 * mib,
size: 10_000 * 6 * mib,
want: 6 * mib,
},
{
name: "one byte more than 10,000 parts adds a byte to each",
configured: 6 * mib,
size: 10_000*6*mib + 1,
want: 6*mib + 1,
},
{
name: "zero means the SDK default of 5MiB",
configured: 0,
size: 1,
want: 5 * mib,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
got := uploadPartSize(tt.configured, tt.size)
if got != tt.want {
t.Errorf("uploadPartSize(%d, %d) = %d, want %d",
tt.configured, tt.size, got, tt.want)
}
})
}
}
+1
View File
@@ -28,6 +28,7 @@ func provideClient(lc fx.Lifecycle, cfg *config.Config) (*Client, error) {
AccessKeyID: cfg.S3.AccessKeyID,
SecretAccessKey: cfg.S3.SecretAccessKey,
Region: cfg.S3.Region,
PartSize: cfg.S3.PartSize.Int64(),
})
if err != nil {
return nil, err
+1 -1
View File
@@ -71,7 +71,7 @@ func verifyBackupFiles(
) {
t.Helper()
files, err := repos.Files.ListByPrefix(ctx, "")
files, err := repos.Files.ListAll(ctx)
if err != nil {
t.Fatalf("Failed to list files: %v", err)
}
+16 -8
View File
@@ -1,40 +1,48 @@
//nolint:testpackage // exercises the unexported copyFile helper
//nolint:testpackage // exercises the unexported copyDatabase helper
package snapshot
import (
"context"
"os"
"path/filepath"
"syscall"
"testing"
"github.com/spf13/afero"
"sneak.berlin/go/vaultik/internal/database"
)
// TestCopyFileExportCopyMode verifies that the exported snapshot database
// copy is created owner-only (0600), even under a lenient 022 umask that
// would otherwise leave a fresh file world-readable.
// TestCopyDatabaseExportCopyMode verifies that the exported snapshot
// database copy is created owner-only (0600), even under a lenient 022
// umask that would otherwise leave a fresh file world-readable.
//
//nolint:paralleltest // syscall.Umask is process-global; parallel tests would clash
func TestCopyFileExportCopyMode(t *testing.T) {
func TestCopyDatabaseExportCopyMode(t *testing.T) {
restore := syscall.Umask(0o022)
defer syscall.Umask(restore)
ctx := context.Background()
dir := t.TempDir()
src := filepath.Join(dir, "index.sqlite")
err := os.WriteFile(src, []byte("index data"), 0o600)
db, err := database.New(ctx, src)
if err != nil {
t.Fatalf("creating source index: %v", err)
}
err = db.Close()
if err != nil {
t.Fatalf("closing source index: %v", err)
}
dst := filepath.Join(dir, "snapshot.db")
sm := &SnapshotManager{fs: afero.NewOsFs()}
err = sm.copyFile(src, dst)
err = sm.copyDatabase(ctx, src, dst)
if err != nil {
t.Fatalf("copyFile: %v", err)
t.Fatalf("copyDatabase: %v", err)
}
info, err := os.Stat(dst)
-4
View File
@@ -66,7 +66,6 @@ type ProgressStats struct {
BlobsCreated atomic.Int64
BlobsUploaded atomic.Int64
BytesUploaded atomic.Int64
UploadDurationMs atomic.Int64 // Total milliseconds spent uploading
CurrentFile atomic.Value // stores string
TotalSize atomic.Int64 // Total size to process (set after scan phase)
TotalFiles atomic.Int64 // Total files to process in phase 2
@@ -231,9 +230,6 @@ func (pr *ProgressReporter) ReportUploadComplete(
// Clear current upload
pr.stats.CurrentUpload.Store((*UploadInfo)(nil))
// Add to total upload duration
pr.stats.UploadDurationMs.Add(duration.Milliseconds())
// Calculate speed
if duration < time.Millisecond {
duration = time.Millisecond
+63 -100
View File
@@ -11,6 +11,7 @@ import (
"runtime"
"strings"
"sync"
"syscall"
"time"
"github.com/dustin/go-humanize"
@@ -91,9 +92,6 @@ type Scanner struct {
// Mutex for coordinating blob creation
packerMu sync.Mutex // Blocks chunk production during blob creation
// Context for cancellation
scanCtx context.Context //nolint:containedctx // set per-Scan for packer callbacks
}
// Periodic status output intervals and thresholds for the scan and
@@ -133,18 +131,23 @@ type ScannerConfig struct {
SkipErrors bool
}
// ScanResult contains the results of a scan operation
// ScanResult contains the results of a scan operation. Files and bytes
// are counted per file: BytesScanned is the size of the new and changed
// files, BytesSkipped that of the unchanged ones.
type ScanResult struct {
FilesScanned int
FilesSkipped int
FilesDeleted int
BytesScanned int64
BytesSkipped int64
BytesDeleted int64
ChunksCreated int
BlobsCreated int
StartTime time.Time
EndTime time.Time
FilesScanned int
FilesSkipped int
FilesDeleted int
BytesScanned int64
BytesSkipped int64
BytesDeleted int64
ChunksCreated int
BlobsCreated int
BlobsUploaded int
BytesUploaded int64
UploadDuration time.Duration
StartTime time.Time
EndTime time.Time
}
// NewScanner creates a new scanner instance
@@ -210,7 +213,6 @@ func (s *Scanner) Scan(
s.snapshotID = snapshotID
// Store source path for file records (used during restore)
s.currentSourcePath = path
s.scanCtx = ctx
result := &ScanResult{
StartTime: time.Now().UTC(),
}
@@ -218,17 +220,13 @@ func (s *Scanner) Scan(
// Set blob handler for concurrent upload
if s.storage != nil {
log.Debug("Setting blob handler for storage uploads")
s.packer.SetBlobHandler(s.handleBlobReady)
s.packer.SetBlobHandler(func(blobWithReader *blob.WithReader) error {
return s.handleBlobReady(ctx, blobWithReader, result)
})
} else {
log.Debug("No storage configured, blobs will not be uploaded")
}
// Start progress reporting if enabled
if s.progress != nil {
s.progress.Start()
defer s.progress.Stop()
}
// Phase 0: Repair any state left by an interrupted previous run, then
// load known files and chunks from the database into memory for fast
// lookup.
@@ -293,13 +291,14 @@ func (s *Scanner) Scan(
log.Info("Phase 2/3: Skipping (no files need processing, metadata-only snapshot)")
}
// Finalize result with blob statistics
s.finalizeScanResult(ctx, result)
result.EndTime = time.Now().UTC()
return result, nil
}
// GetProgress returns the progress reporter for this scanner
// GetProgress returns the progress reporter for this scanner, or nil when
// progress is off. Scan neither starts nor stops it: the caller does,
// once for all the paths it scans, because a second Stop panics.
func (s *Scanner) GetProgress() *ProgressReporter {
return s.progress
}
@@ -434,35 +433,16 @@ func (s *Scanner) summarizeScanPhase(
s.ui.Completef("%s.", msg)
}
// finalizeScanResult populates final blob statistics in the scan result
// by querying the packer and database for blob/upload counts
func (s *Scanner) finalizeScanResult(ctx context.Context, result *ScanResult) {
blobs := s.packer.GetFinishedBlobs()
result.BlobsCreated += len(blobs)
// Query database for actual blob count created during this snapshot
// The database is authoritative, especially for concurrent blob uploads
// We count uploads rather than all snapshot_blobs to get only NEW blobs
if s.snapshotID != "" {
uploadCount, err := s.repos.Uploads.GetCountBySnapshot(ctx, s.snapshotID)
if err != nil {
log.Warn("Failed to query upload count from database", "error", err)
} else {
result.BlobsCreated = int(uploadCount)
}
}
result.EndTime = time.Now().UTC()
}
// loadKnownFiles loads all known files from the database into a map for fast lookup
// This avoids per-file database queries during the scan phase
// loadKnownFiles loads the known files at and beneath path from the
// database into a map for fast lookup. Every loaded file the scan does
// not find is counted as deleted. This avoids per-file database queries
// during the scan phase.
func (s *Scanner) loadKnownFiles(
ctx context.Context, path string,
) (map[string]*database.File, error) {
files, err := s.repos.Files.ListByPrefix(ctx, path)
files, err := s.repos.Files.ListUnderPath(ctx, path)
if err != nil {
return nil, fmt.Errorf("listing files by prefix: %w", err)
return nil, fmt.Errorf("listing files under %s: %w", path, err)
}
result := make(map[string]*database.File, len(files))
@@ -1142,12 +1122,9 @@ func (s *Scanner) buildSymlinkEntry(path string, info os.FileInfo) *database.Fil
}
var uid, gid uint32
if stat, ok := info.Sys().(interface {
Uid() uint32
Gid() uint32
}); ok {
uid = stat.Uid()
gid = stat.Gid()
if stat, ok := info.Sys().(*syscall.Stat_t); ok {
uid = stat.Uid
gid = stat.Gid
}
return &database.File{
@@ -1166,12 +1143,9 @@ func (s *Scanner) buildSymlinkEntry(path string, info os.FileInfo) *database.Fil
// buildDirectoryEntry creates a File record for a directory.
func (s *Scanner) buildDirectoryEntry(path string, info os.FileInfo) *database.File {
var uid, gid uint32
if stat, ok := info.Sys().(interface {
Uid() uint32
Gid() uint32
}); ok {
uid = stat.Uid()
gid = stat.Gid()
if stat, ok := info.Sys().(*syscall.Stat_t); ok {
uid = stat.Uid
gid = stat.Gid
}
return &database.File{
@@ -1204,16 +1178,10 @@ func (s *Scanner) recordNonRegularFile(ctx context.Context, ftp *FileToProcess)
func (s *Scanner) checkFileInMemory(
path string, info os.FileInfo, knownFiles map[string]*database.File,
) (*database.File, bool) {
// Get file stats
stat, ok := info.Sys().(interface {
Uid() uint32
Gid() uint32
})
var uid, gid uint32
if ok {
uid = stat.Uid()
gid = stat.Gid()
if stat, ok := info.Sys().(*syscall.Stat_t); ok {
uid = stat.Uid
gid = stat.Gid
}
// Check against in-memory map first to get existing ID if available
@@ -1253,7 +1221,7 @@ func (s *Scanner) checkFileInMemory(
// Check if file has changed
if existingFile.Size != file.Size ||
existingFile.MTime.Unix() != file.MTime.Unix() ||
!existingFile.MTime.Equal(file.MTime) ||
existingFile.Mode != file.Mode ||
existingFile.UID != file.UID ||
existingFile.GID != file.GID {
@@ -1524,24 +1492,24 @@ func (s *Scanner) finalizeProcessPhase(ctx context.Context, result *ScanResult)
}
// handleBlobReady is called by the packer when a blob is finalized
func (s *Scanner) handleBlobReady(blobWithReader *blob.WithReader) error {
func (s *Scanner) handleBlobReady(
ctx context.Context, blobWithReader *blob.WithReader, result *ScanResult,
) error {
startTime := time.Now().UTC()
finishedBlob := blobWithReader.FinishedBlob
result.BlobsCreated++
if s.progress != nil {
s.progress.ReportUploadStart(finishedBlob.Hash, finishedBlob.Compressed)
s.progress.GetStats().BlobsCreated.Add(1)
}
ctx := s.scanCtx
if ctx == nil {
ctx = context.Background()
}
blobPath := fmt.Sprintf("blobs/%s/%s/%s",
finishedBlob.Hash[:2], finishedBlob.Hash[2:4], finishedBlob.Hash)
blobExists, err := s.uploadBlobIfNeeded(ctx, blobPath, blobWithReader, startTime)
blobExists, err := s.uploadBlobIfNeeded(
ctx, blobPath, blobWithReader, startTime, result)
if err != nil {
s.cleanupBlobTempFile(blobWithReader)
@@ -1576,6 +1544,7 @@ func (s *Scanner) uploadBlobIfNeeded(
blobPath string,
blobWithReader *blob.WithReader,
startTime time.Time,
result *ScanResult,
) (bool, error) {
finishedBlob := blobWithReader.FinishedBlob
@@ -1611,6 +1580,10 @@ func (s *Scanner) uploadBlobIfNeeded(
uploadDuration := time.Since(startTime)
uploadSpeedBps := float64(finishedBlob.Compressed) / uploadDuration.Seconds()
result.BlobsUploaded++
result.BytesUploaded += finishedBlob.Compressed
result.UploadDuration += uploadDuration
s.ui.Completef("Uploaded blob %s (%s) in %s at %s.",
s.ui.Hex(finishedBlob.Hash),
s.ui.Size(finishedBlob.Compressed),
@@ -1826,9 +1799,9 @@ func (s *Scanner) processFileStreaming(
size: chunk.Size,
})
s.updateChunkStats(chunkExists, chunk.Size, result)
if !chunkExists {
s.updateChunkStats(chunk.Size, result)
err := s.addChunkToPacker(ctx, chunk)
if err != nil {
// Mark as a packer error so --skip-errors cannot swallow it:
@@ -1856,26 +1829,16 @@ func (s *Scanner) processFileStreaming(
return nil
}
// updateChunkStats updates scan result and progress stats for a processed chunk
func (s *Scanner) updateChunkStats(
chunkExists bool, chunkSize int64, result *ScanResult,
) {
if chunkExists {
result.FilesSkipped++
// updateChunkStats counts a chunk that was not already stored. The scan
// result's file counts, BytesScanned and BytesSkipped are not touched
// here: the scan phase counts each file once.
func (s *Scanner) updateChunkStats(chunkSize int64, result *ScanResult) {
result.ChunksCreated++
result.BytesSkipped += chunkSize
if s.progress != nil {
s.progress.GetStats().BytesSkipped.Add(chunkSize)
}
} else {
result.ChunksCreated++
result.BytesScanned += chunkSize
if s.progress != nil {
s.progress.GetStats().ChunksCreated.Add(1)
s.progress.GetStats().BytesProcessed.Add(chunkSize)
s.progress.UpdateChunkingActivity()
}
if s.progress != nil {
s.progress.GetStats().ChunksCreated.Add(1)
s.progress.GetStats().BytesProcessed.Add(chunkSize)
s.progress.UpdateChunkingActivity()
}
}
+78 -1
View File
@@ -71,7 +71,7 @@ func verifySimpleScanDatabase(
t.Helper()
// Verify files in database - includes regular files and directories
files, err := repos.Files.ListByPrefix(ctx, "/source")
files, err := repos.Files.ListUnderPath(ctx, "/source")
if err != nil {
t.Fatalf("failed to list files: %v", err)
}
@@ -307,3 +307,80 @@ func TestScannerLargeFile(t *testing.T) {
}
}
}
// TestScannerRecordsOwnership backs up real files on disk and checks that
// the uid and gid of a file, a directory and a symlink are recorded.
// When the tests run as root both sides are 0, so only a run as another
// user can catch ownership recorded as 0.
func TestScannerRecordsOwnership(t *testing.T) {
t.Parallel()
sourceDir := t.TempDir()
filePath := filepath.Join(sourceDir, "file.txt")
dirPath := filepath.Join(sourceDir, "subdir")
linkPath := filepath.Join(sourceDir, "link")
err := os.WriteFile(filePath, []byte("owned"), 0o600)
if err != nil {
t.Fatal(err)
}
err = os.Mkdir(dirPath, 0o700)
if err != nil {
t.Fatal(err)
}
err = os.Symlink("file.txt", linkPath)
if err != nil {
t.Fatal(err)
}
db, err := database.NewTestDB()
if err != nil {
t.Fatalf("failed to create test database: %v", err)
}
defer func() {
err := db.Close()
if err != nil {
t.Errorf("failed to close database: %v", err)
}
}()
repos := database.NewRepositories(db)
scanner := snapshot.NewScanner(snapshot.ScannerConfig{
FS: afero.NewOsFs(),
ChunkSize: int64(1024 * 16),
Repositories: repos,
MaxBlobSize: int64(1024 * 1024),
CompressionLevel: 3,
AgeRecipients: []string{testAgePublicKey},
})
ctx := context.Background()
snapshotID := "test-snapshot-ownership"
createTestSnapshotRecord(ctx, t, repos, snapshotID)
_, err = scanner.Scan(ctx, sourceDir, snapshotID)
if err != nil {
t.Fatalf("scan failed: %v", err)
}
for _, path := range []string{filePath, dirPath, linkPath} {
file, err := repos.Files.GetByPath(ctx, path)
if err != nil {
t.Fatalf("failed to get %s: %v", path, err)
}
if file == nil {
t.Fatalf("%s was not recorded", path)
}
if int(file.UID) != os.Getuid() || int(file.GID) != os.Getgid() {
t.Errorf("%s recorded as uid %d gid %d, want uid %d gid %d",
path, file.UID, file.GID, os.Getuid(), os.Getgid())
}
}
}
+47 -66
View File
@@ -105,17 +105,21 @@ func (sm *SnapshotManager) CreateSnapshot(
return sm.CreateSnapshotWithName(ctx, hostname, "", version, gitRevision)
}
// ShortHostname returns hostname up to its first dot. A snapshot ID starts
// with this form, while the snapshots table stores the full hostname.
func ShortHostname(hostname string) string {
short, _, _ := strings.Cut(hostname, ".")
return short
}
// CreateSnapshotWithName creates a new snapshot record with an optional
// snapshot name. The snapshot ID format is: hostname_name_timestamp or
// hostname_timestamp if name is empty.
func (sm *SnapshotManager) CreateSnapshotWithName(
ctx context.Context, hostname, name, version, gitRevision string,
) (string, error) {
// Use short hostname (strip domain if present)
shortHostname := hostname
if before, _, ok := strings.Cut(hostname, "."); ok {
shortHostname = before
}
shortHostname := ShortHostname(hostname)
// Build snapshot ID with optional name
timestamp := time.Now().UTC().Format("2006-01-02T15:04:05Z")
@@ -154,26 +158,6 @@ func (sm *SnapshotManager) CreateSnapshotWithName(
return snapshotID, nil
}
// UpdateSnapshotStats updates the statistics for a snapshot during backup
func (sm *SnapshotManager) UpdateSnapshotStats(
ctx context.Context, snapshotID string, stats BackupStats,
) error {
err := sm.repos.WithTx(ctx, func(ctx context.Context, tx *sql.Tx) error {
return sm.repos.Snapshots.UpdateCounts(ctx, tx, snapshotID,
int64(stats.FilesScanned),
int64(stats.ChunksCreated),
int64(stats.BlobsCreated),
stats.BytesScanned,
stats.BytesUploaded,
)
})
if err != nil {
return fmt.Errorf("updating snapshot stats: %w", err)
}
return nil
}
// UpdateSnapshotStatsExtended updates snapshot statistics with extended metrics.
// This includes compression level, uncompressed blob size, and upload duration.
func (sm *SnapshotManager) UpdateSnapshotStatsExtended(
@@ -185,8 +169,8 @@ func (sm *SnapshotManager) UpdateSnapshotStatsExtended(
int64(stats.FilesScanned),
int64(stats.ChunksCreated),
int64(stats.BlobsCreated),
stats.BytesScanned,
stats.BytesUploaded,
stats.TotalSize,
stats.BlobSize,
)
if err != nil {
return err
@@ -196,6 +180,7 @@ func (sm *SnapshotManager) UpdateSnapshotStatsExtended(
return sm.repos.Snapshots.UpdateExtendedStats(ctx, tx, snapshotID,
stats.BlobUncompressedSize,
stats.CompressionLevel,
stats.BytesUploaded,
stats.UploadDurationMs,
)
})
@@ -386,12 +371,12 @@ func (sm *SnapshotManager) prepareExportDB(
ctx context.Context, dbPath, snapshotID, tempDir string,
) ([]byte, string, error) {
// Step 1: Copy database to temp file
// The main database should be closed at this point
// The main database is still open here, so it is copied through SQLite
tempDBPath := filepath.Join(tempDir, "snapshot.db")
log.Debug("Copying database to temporary location",
"source", dbPath, "destination", tempDBPath)
err := sm.copyFile(dbPath, tempDBPath)
err := sm.copyDatabase(ctx, dbPath, tempDBPath)
if err != nil {
return nil, "", fmt.Errorf("copying database: %w", err)
}
@@ -648,9 +633,10 @@ func (sm *SnapshotManager) collectCleanupStats(
//
// VACUUM runs through the modernc.org/sqlite driver, on a freshly opened
// connection with no transaction in flight (VACUUM cannot run inside one).
// The database opens in WAL mode, so VACUUM's rewrite lands in the WAL; the
// checkpoint on Close flushes it into the main file, which is the file we
// then compress and upload.
// database.New opens the file in WAL mode, so VACUUM's rewrite lands in the
// -wal file. This is the only connection to the file, so closing it
// checkpoints the rewrite into the main file and removes the -wal file; the
// main file is the one compressFile then compresses and uploads.
func (sm *SnapshotManager) vacuumDatabase(ctx context.Context, dbPath string) error {
log.Debug("Running VACUUM on database", "path", dbPath)
@@ -744,26 +730,15 @@ func (sm *SnapshotManager) compressFile(inputPath, outputPath string) error {
// user; it holds the same private index data as the local index file.
const exportCopyPerm = 0o600
// copyFile copies a file from src to dst. The destination is the exported
// snapshot database, so it is created owner-only rather than with the
// umask-dependent default.
func (sm *SnapshotManager) copyFile(src, dst string) error {
log.Debug("Opening source file for copy", "path", src)
sourceFile, err := sm.fs.Open(src)
if err != nil {
return err
}
defer func() {
log.Debug("Closing source file", "path", src)
err := sourceFile.Close()
if err != nil {
log.Debug("Failed to close source file", "path", src, "error", err)
}
}()
// copyDatabase copies the database at src to dst with VACUUM INTO. It reads
// through SQLite, so the copy holds rows committed to src that are still in
// its -wal file, which a copy of the file alone would miss. The destination
// is the exported snapshot database, so it is created empty and owner-only
// first rather than with the umask-dependent default: VACUUM INTO writes
// into an existing empty file and keeps its mode.
func (sm *SnapshotManager) copyDatabase(
ctx context.Context, src, dst string,
) error {
log.Debug("Creating destination file", "path", dst)
destFile, err := sm.fs.OpenFile(
@@ -773,23 +748,28 @@ func (sm *SnapshotManager) copyFile(src, dst string) error {
return err
}
defer func() {
log.Debug("Closing destination file", "path", dst)
err := destFile.Close()
if err != nil {
log.Debug("Failed to close destination file", "path", dst, "error", err)
}
}()
log.Debug("Copying file data")
n, err := io.Copy(destFile, sourceFile)
err = destFile.Close()
if err != nil {
return err
}
log.Debug("File copy complete", "bytes_copied", n)
db, err := database.New(ctx, src)
if err != nil {
return fmt.Errorf("opening database to copy: %w", err)
}
defer func() {
cerr := db.Close()
if cerr != nil {
log.Debug("Failed to close database after copy",
"path", src, "error", cerr)
}
}()
_, err = db.ExecWithLog(ctx, "VACUUM INTO ?", dst)
if err != nil {
return fmt.Errorf("running VACUUM INTO: %w", err)
}
return nil
}
@@ -895,7 +875,7 @@ func (sm *SnapshotManager) getFileSize(path string) int64 {
// BackupStats contains statistics from a backup operation
type BackupStats struct {
FilesScanned int
BytesScanned int64
TotalSize int64 // Total size of all files examined
ChunksCreated int
BlobsCreated int
BytesUploaded int64
@@ -905,6 +885,7 @@ type BackupStats struct {
type ExtendedBackupStats struct {
BackupStats
BlobSize int64 // Total compressed size of all referenced blobs
BlobUncompressedSize int64 // Total uncompressed size of all referenced blobs
CompressionLevel int // Compression level used for this snapshot
UploadDurationMs int64 // Total milliseconds spent uploading to S3
+71
View File
@@ -188,6 +188,77 @@ func TestVacuumDatabaseRemovesDeletedData(t *testing.T) {
}
}
// TestPrepareExportDBKeepsRowsCommittedToOpenIndex exports from an index
// that is still open, as a backup does. A row committed there can still be
// in the index's -wal file, and the export must hold it all the same.
func TestPrepareExportDBKeepsRowsCommittedToOpenIndex(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
ctx := context.Background()
fs := afero.NewOsFs()
dbPath := filepath.Join(t.TempDir(), "index.sqlite")
db, err := database.New(ctx, dbPath)
if err != nil {
t.Fatalf("failed to create database: %v", err)
}
defer func() { _ = db.Close() }()
repos := database.NewRepositories(db)
snapshot := &database.Snapshot{ID: "open-index-snapshot", Hostname: "test-host"}
err = repos.WithTx(ctx, func(ctx context.Context, tx *sql.Tx) error {
return repos.Snapshots.Create(ctx, tx, snapshot)
})
if err != nil {
t.Fatalf("failed to create snapshot: %v", err)
}
sm := &SnapshotManager{
config: &config.Config{
CompressionLevel: 3,
AgeRecipients: []string{testAgeRecipient},
},
fs: fs,
}
_, tempDBPath, err := sm.prepareExportDB(
ctx, dbPath, snapshot.ID.String(), t.TempDir())
if err != nil {
t.Fatalf("prepareExportDB failed: %v", err)
}
// Only the main database file is compressed and uploaded, so open a
// copy of that file alone.
uploadedPath := filepath.Join(t.TempDir(), "uploaded.db")
err = copyFile(fs, tempDBPath, uploadedPath)
if err != nil {
t.Fatalf("failed to copy exported database: %v", err)
}
exported, err := database.OpenReadOnly(ctx, uploadedPath)
if err != nil {
t.Fatalf("failed to open exported database: %v", err)
}
defer func() { _ = exported.Close() }()
got, err := database.NewRepositories(exported).Snapshots.GetByID(
ctx, snapshot.ID.String())
if err != nil {
t.Fatalf("failed to read snapshot from export: %v", err)
}
if got == nil {
t.Fatal("exported database is missing the snapshot row")
}
}
func TestCleanSnapshotDBEmptySnapshot(t *testing.T) {
// Initialize logger
log.Initialize(log.Config{})
+23 -8
View File
@@ -22,11 +22,11 @@ type FileStorer struct {
//
// Construction is intentionally cheap and does not touch the filesystem.
// The basePath is recorded; the directory is created lazily on first
// write. Reads (Get/Stat/List) tolerate a missing basePath — a missing
// or unmounted destination during `snapshot list` should NOT block the
// command, it should degrade to "no remote snapshots reachable" with a
// warning. Write operations (Put/PutWithProgress) call MkdirAll for the
// write. Write operations (Put/PutWithProgress) call MkdirAll for the
// per-blob parent directory, which also covers basePath on first use.
// Get and Stat report a key under a missing basePath as ErrNotFound.
// List and ListStream fail on a missing basePath, because listing it as
// an empty store would make `prune` drop every local snapshot record.
//
// Uses the real OS filesystem by default; call SetFilesystem to
// override for testing.
@@ -119,13 +119,20 @@ func (f *FileStorer) Delete(_ context.Context, key string) error {
return nil
}
// List returns all keys with the given prefix.
// List returns all keys with the given prefix. It fails when the
// destination directory is missing; a missing prefix under it is an
// empty listing.
func (f *FileStorer) List(ctx context.Context, prefix string) ([]string, error) {
var keys []string
_, err := f.fs.Stat(f.basePath)
if err != nil {
return nil, fmt.Errorf("checking destination directory: %w", err)
}
basePath := f.fullPath(prefix)
// Check if base path exists
// Check if the prefix exists
exists, err := afero.Exists(f.fs, basePath)
if err != nil {
return nil, fmt.Errorf("checking path: %w", err)
@@ -167,16 +174,24 @@ func (f *FileStorer) List(ctx context.Context, prefix string) ([]string, error)
return keys, nil
}
// ListStream returns a channel of ObjectInfo for large result sets.
// ListStream returns a channel of ObjectInfo for large result sets. Like
// List, it sends an error when the destination directory is missing.
func (f *FileStorer) ListStream(ctx context.Context, prefix string) <-chan ObjectInfo {
ch := make(chan ObjectInfo)
go func() {
defer close(ch)
_, err := f.fs.Stat(f.basePath)
if err != nil {
ch <- ObjectInfo{Err: fmt.Errorf("checking destination directory: %w", err)}
return
}
basePath := f.fullPath(prefix)
// Check if base path exists
// Check if the prefix exists
exists, err := afero.Exists(f.fs, basePath)
if err != nil {
ch <- ObjectInfo{Err: fmt.Errorf("checking path: %w", err)}
+35
View File
@@ -1,6 +1,10 @@
package storage_test
import (
"context"
"errors"
"io/fs"
"path/filepath"
"testing"
"sneak.berlin/go/vaultik/internal/storage"
@@ -25,3 +29,34 @@ func TestFileStorer(t *testing.T) {
t.Parallel()
runStorerConformance(t, newFileStorer)
}
// TestFileStorerListMissingDestination checks that List and ListStream
// fail when the destination directory does not exist, instead of
// reporting an empty store.
func TestFileStorerListMissingDestination(t *testing.T) {
t.Parallel()
ctx := context.Background()
s, err := storage.NewFileStorer(filepath.Join(t.TempDir(), "unmounted"))
if err != nil {
t.Fatalf("NewFileStorer: %v", err)
}
keys, err := s.List(ctx, "metadata/")
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("List = %v, %v; want a not-exist error", keys, err)
}
var streamErr error
for object := range s.ListStream(ctx, "metadata/") {
if object.Err != nil {
streamErr = object.Err
}
}
if !errors.Is(streamErr, fs.ErrNotExist) {
t.Errorf("ListStream error = %v, want a not-exist error", streamErr)
}
}
+2
View File
@@ -99,6 +99,7 @@ func storerFromParsedS3URL(parsed *URL, cfg *config.Config) (Storer, error) {
AccessKeyID: cfg.S3.AccessKeyID,
SecretAccessKey: cfg.S3.SecretAccessKey,
Region: region,
PartSize: cfg.S3.PartSize.Int64(),
})
if err != nil {
return nil, fmt.Errorf("creating S3 client: %w", err)
@@ -134,6 +135,7 @@ func storerFromLegacyS3Config(cfg *config.Config) (Storer, error) {
AccessKeyID: cfg.S3.AccessKeyID,
SecretAccessKey: cfg.S3.SecretAccessKey,
Region: region,
PartSize: cfg.S3.PartSize.Int64(),
})
if err != nil {
return nil, fmt.Errorf("creating S3 client: %w", err)
+203
View File
@@ -1,14 +1,20 @@
package storage_test
import (
"bytes"
"context"
"errors"
"net/http"
"net/http/httptest"
"slices"
"strings"
"sync/atomic"
"testing"
"github.com/johannesboyne/gofakes3"
"github.com/johannesboyne/gofakes3/backend/s3mem"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/s3"
"sneak.berlin/go/vaultik/internal/storage"
)
@@ -16,6 +22,13 @@ import (
// s3TestBucket is the bucket created for each in-process S3 server.
const s3TestBucket = "test-bucket"
// Credentials for the tests that build a storer from a config.Config. The
// in-process S3 server accepts any.
const (
s3TestAccessKeyID = "key"
s3TestSecretAccessKey = "secret"
)
// newS3Storer builds an s3:// backend backed by a fresh in-process
// S3 server. It reuses the same in-memory S3 harness (gofakes3 + s3mem
// over httptest) that internal/s3 and the not-found regression test use,
@@ -79,3 +92,193 @@ func TestS3StorerMissingKeyMapsToErrNotFound(t *testing.T) {
t.Errorf("Stat on missing key: got %v, want ErrNotFound", err)
}
}
// TestS3URLPrefixKeyLayout pins the bucket keys an s3:// URL reads and
// writes: the README's remote storage layout, with the prefix joined to
// each key by one "/". s3://b/p and s3://b/p/ must be the same
// destination, or a host that writes the URL the other way finds no
// snapshots. The listed object is put straight into the bucket, as
// another host would have written it. Both List and ListStream are
// checked: ListStream is what every snapshot listing goes through.
func TestS3URLPrefixKeyLayout(t *testing.T) {
t.Parallel()
const (
blobKey = "blobs/aa/bb/aabbccdd"
listPrefix = "metadata/"
manifestKey = listPrefix + "snap/manifest.json.zst"
manifestBody = "manifest"
)
cases := []struct {
urlPath string // URL path after the bucket name
keyPrefix string // what every key in the bucket must start with
}{
{urlPath: "/p", keyPrefix: "p/"},
{urlPath: "/p/", keyPrefix: "p/"},
{urlPath: "", keyPrefix: ""},
}
for _, tc := range cases {
storageURL := "s3://" + s3TestBucket + tc.urlPath
t.Run(storageURL, func(t *testing.T) {
t.Parallel()
backend := s3mem.New()
err := backend.CreateBucket(s3TestBucket)
if err != nil {
t.Fatalf("create bucket: %v", err)
}
srv := httptest.NewServer(gofakes3.New(backend).Server())
t.Cleanup(srv.Close)
storer, err := storage.NewStorer(&config.Config{
StorageURL: storageURL + "?endpoint=" + srv.URL,
S3: config.S3Config{
AccessKeyID: s3TestAccessKeyID,
SecretAccessKey: s3TestSecretAccessKey,
},
})
if err != nil {
t.Fatalf("NewStorer: %v", err)
}
ctx := context.Background()
err = storer.Put(ctx, blobKey, strings.NewReader("blob"))
if err != nil {
t.Fatalf("Put: %v", err)
}
_, err = backend.HeadObject(s3TestBucket, tc.keyPrefix+blobKey)
if err != nil {
t.Errorf("blob not stored at %q: %v", tc.keyPrefix+blobKey, err)
}
_, err = backend.PutObject(s3TestBucket, tc.keyPrefix+manifestKey,
nil, strings.NewReader(manifestBody), int64(len(manifestBody)))
if err != nil {
t.Fatalf("seed manifest: %v", err)
}
keys, err := storer.List(ctx, listPrefix)
if err != nil {
t.Fatalf("List: %v", err)
}
if !slices.Equal(keys, []string{manifestKey}) {
t.Errorf("List(%q) = %q, want [%q]", listPrefix, keys, manifestKey)
}
streamed := listStreamKeys(t, storer, listPrefix)
if !slices.Equal(streamed, []string{manifestKey}) {
t.Errorf("ListStream(%q) = %q, want [%q]", listPrefix, streamed, manifestKey)
}
})
}
}
// TestS3UploadUsesConfiguredPartSize checks that s3.part_size reaches the
// multipart uploader, through storage_url and through the s3.* fields. An
// object three parts long must arrive as three parts; at the SDK's default
// of 5 MiB it would arrive as four.
func TestS3UploadUsesConfiguredPartSize(t *testing.T) {
t.Parallel()
const (
partSize = 6 * 1024 * 1024
wantParts = 3
)
backend := s3mem.New()
err := backend.CreateBucket(s3TestBucket)
if err != nil {
t.Fatalf("create bucket: %v", err)
}
// Every part of a multipart upload is one request with a partNumber.
var parts atomic.Int32
fake := gofakes3.New(backend).Server()
srv := httptest.NewServer(http.HandlerFunc(
func(w http.ResponseWriter, r *http.Request) {
if r.URL.Query().Has("partNumber") {
parts.Add(1)
}
fake.ServeHTTP(w, r)
}))
t.Cleanup(srv.Close)
cases := []struct {
name string
cfg *config.Config
}{
{
name: "storage_url",
cfg: &config.Config{
StorageURL: "s3://" + s3TestBucket + "?endpoint=" + srv.URL,
S3: config.S3Config{
AccessKeyID: s3TestAccessKeyID,
SecretAccessKey: s3TestSecretAccessKey,
PartSize: partSize,
},
},
},
{
name: "s3.endpoint",
cfg: &config.Config{
S3: config.S3Config{
Endpoint: srv.URL,
Bucket: s3TestBucket,
AccessKeyID: s3TestAccessKeyID,
SecretAccessKey: s3TestSecretAccessKey,
PartSize: partSize,
},
},
},
}
for _, tc := range cases {
parts.Store(0)
storer, err := storage.NewStorer(tc.cfg)
if err != nil {
t.Fatalf("%s: NewStorer: %v", tc.name, err)
}
data := bytes.NewReader(make([]byte, wantParts*partSize))
err = storer.PutWithProgress(
context.Background(), "blob", data, data.Size(), nil)
if err != nil {
t.Fatalf("%s: PutWithProgress: %v", tc.name, err)
}
if got := parts.Load(); got != wantParts {
t.Errorf("%s: uploaded in %d parts, want %d", tc.name, got, wantParts)
}
}
}
// listStreamKeys returns the keys ListStream yields under a prefix, and
// fails the test on a listing error.
func listStreamKeys(t *testing.T, s storage.Storer, prefix string) []string {
t.Helper()
var keys []string
for obj := range s.ListStream(context.Background(), prefix) {
if obj.Err != nil {
t.Fatalf("ListStream %q: %v", prefix, obj.Err)
}
keys = append(keys, obj.Key)
}
return keys
}
+4
View File
@@ -301,6 +301,10 @@ func TestInterruptedBlobUploadRecordsNoUploadedBlob(t *testing.T) {
writeFaultSourceTree(t, fs, dataDir)
// No upload succeeds, so nothing creates the destination directory.
// It must exist for the listing below to show that no blob survived.
require.NoError(t, os.Mkdir(storeDir, 0o750))
inner, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
+38 -23
View File
@@ -9,6 +9,7 @@ import (
"time"
"github.com/dustin/go-humanize"
"sneak.berlin/go/vaultik/internal/snapshot"
"sneak.berlin/go/vaultik/internal/types"
)
@@ -46,12 +47,9 @@ const (
year = 365 * day
)
// Snapshot IDs split on "_" into hostname, optional name parts, and a
// trailing timestamp.
const (
minSnapshotIDParts = 2
minSnapshotIDNameParts = 3
)
// A snapshot ID split on "_" has at least a hostname and a trailing
// timestamp.
const minSnapshotIDParts = 2
// SnapshotInfo contains information about a snapshot.
//
@@ -121,43 +119,60 @@ func parseSnapshotTimestamp(snapshotID string) (time.Time, error) {
return timestamp.UTC(), nil
}
// parseSnapshotName extracts the snapshot name from a snapshot ID.
// Format: hostname_snapshotname_timestamp — the middle part(s) between hostname
// and the RFC3339 timestamp are the snapshot name (may contain underscores).
// Returns the snapshot name, or empty string if the ID is malformed.
func parseSnapshotName(snapshotID string) string {
parts := strings.Split(snapshotID, "_")
if len(parts) < minSnapshotIDNameParts {
// Format: hostname_timestamp — no snapshot name
// parseSnapshotName extracts the snapshot name from a snapshot ID of the
// form hostname_name_timestamp, given the hostname stored with that
// snapshot. The hostname and the name may both contain underscores, so the
// name is what is left after removing the short hostname and its "_" from
// the front and the last "_" and the timestamp from the end. Returns "" for
// an ID with no name (hostname_timestamp), and for an ID that does not start
// with that hostname, which CreateSnapshotWithName never writes.
func parseSnapshotName(snapshotID, hostname string) string {
prefix := snapshot.ShortHostname(hostname) + "_"
rest, ok := strings.CutPrefix(snapshotID, prefix)
if !ok {
return ""
}
// Format: hostname_name_timestamp — middle parts are the name.
// The last part is the RFC3339 timestamp, the first part is the hostname,
// everything in between is the snapshot name (which may itself contain underscores).
return strings.Join(parts[1:len(parts)-1], "_")
end := strings.LastIndex(rest, "_")
if end < 0 {
return ""
}
return rest[:end]
}
// parseDuration parses a duration string with support for human-friendly units:
// d/day/days, w/week/weeks, mo/month/months, y/year/years, plus standard Go
// duration units. Following Go, m is minutes and mo is months. A bare number,
// an unknown unit, and a negative value are all rejected.
// an unknown unit, and a negative value are all rejected. Outside the Go
// units the input must be whole numbers each followed directly by a unit
// (2w3d), so 1.0y and 30 days are errors; decimals work only in Go units
// (1.5h).
func parseDuration(s string) (time.Duration, error) {
if strings.HasPrefix(strings.TrimSpace(s), "-") {
return 0, errNegativeDuration
}
// A bare number has no unit, but time.ParseDuration accepts 0, which
// would put the cutoff at now and select every snapshot.
_, err := strconv.ParseFloat(s, 64)
if err == nil {
return 0, fmt.Errorf("%w: %q", errInvalidDuration, s)
}
d, err := time.ParseDuration(s)
if err == nil {
return d, nil
}
re := regexp.MustCompile(`(\d+)\s*([a-zA-Z]+)`)
matches := re.FindAllStringSubmatch(s, -1)
if len(matches) == 0 {
if !regexp.MustCompile(`^(\d+[a-zA-Z]+)+$`).MatchString(s) {
return 0, fmt.Errorf("%w: %q", errInvalidDuration, s)
}
re := regexp.MustCompile(`(\d+)([a-zA-Z]+)`)
matches := re.FindAllStringSubmatch(s, -1)
var total time.Duration
for _, match := range matches {
+36 -3
View File
@@ -11,33 +11,55 @@ func TestParseSnapshotName(t *testing.T) {
tests := []struct {
name string
snapshotID string
hostname string
want string
}{
{
name: "standard format with name",
snapshotID: "myhost_home_2026-01-12T14:41:15Z",
hostname: "myhost",
want: "home",
},
{
name: "standard format with different name",
snapshotID: "server1_system_2026-02-15T09:30:00Z",
hostname: "server1",
want: "system",
},
{
name: "name with underscores",
snapshotID: "myhost_my_special_backup_2026-03-01T00:00:00Z",
hostname: "myhost",
want: "my_special_backup",
},
{
name: "hostname with underscores",
snapshotID: "my_host_docs_2026-03-01T00:00:00Z",
hostname: "my_host",
want: "docs",
},
{
name: "stored hostname with domain",
snapshotID: "my_host_mail_2026-03-01T00:00:00Z",
hostname: "my_host.example.com",
want: "mail",
},
{
name: "no name",
snapshotID: "my_host_2026-03-01T00:00:00Z",
hostname: "my_host",
want: "",
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
got := parseSnapshotName(tt.snapshotID)
got := parseSnapshotName(tt.snapshotID, tt.hostname)
if got != tt.want {
t.Errorf("parseSnapshotName(%q) = %q, want %q",
tt.snapshotID, got, tt.want)
t.Errorf("parseSnapshotName(%q, %q) = %q, want %q",
tt.snapshotID, tt.hostname, got, tt.want)
}
})
}
@@ -59,6 +81,7 @@ func TestParseDuration(t *testing.T) {
{"30s", 30 * time.Second, false},
{"6m", 6 * time.Minute, false},
{"1h", time.Hour, false},
{"1.5h", 90 * time.Minute, false},
// Extended calendar units.
{"30d", 30 * 24 * time.Hour, false},
{"3days", 3 * 24 * time.Hour, false},
@@ -73,11 +96,21 @@ func TestParseDuration(t *testing.T) {
{"1y6mo", 365*24*time.Hour + 180*24*time.Hour, false},
// Rejected inputs.
{"6", 0, true}, // bare number, no unit
{"0", 0, true}, // bare number; time.ParseDuration accepts it
{"+0", 0, true}, // bare number; time.ParseDuration accepts it
{"5x", 0, true}, // unknown unit
{"-5d", 0, true}, // negative, extended unit
{"-5h", 0, true}, // negative, Go unit
{"", 0, true}, // empty
{"garbage", 0, true},
// Characters outside the number-and-unit parts must not be
// skipped: 1.0y read as 0y would select every snapshot.
{"1.0y", 0, true},
{"2.1w", 0, true},
{"1,5y", 0, true},
{"30 days", 0, true},
{"30 days ago", 0, true},
{"x7d", 0, true},
}
for _, tt := range tests {
+109 -25
View File
@@ -183,6 +183,11 @@ type SnapshotMetadataInfo struct {
TotalSize int64 `json:"total_size"`
BlobCount int `json:"blob_count"`
BlobsSize int64 `json:"blobs_size"`
// Set when the listing holds this snapshot's manifest.json.zst. A
// backup interrupted before its manifest upload leaves a directory
// without one, which prune does not treat as a snapshot.
hasManifest bool
}
// RemoteInfoResult contains all remote storage information
@@ -206,9 +211,20 @@ type RemoteInfoResult struct {
ReferencedBlobCount int `json:"referenced_blob_count"`
ReferencedBlobSize int64 `json:"referenced_blob_size"`
// Orphaned blobs
OrphanedBlobCount int `json:"orphaned_blob_count"`
OrphanedBlobSize int64 `json:"orphaned_blob_size"`
// Orphaned blobs. Both stay nil (null in the JSON) when a manifest
// was listed but not read, since that snapshot's blobs would be
// counted as orphaned.
OrphanedBlobCount *int `json:"orphaned_blob_count"`
OrphanedBlobSize *int64 `json:"orphaned_blob_size"`
// Remote key of each snapshot whose manifest could not be read
UnreadableManifests []string `json:"unreadable_manifests,omitempty"`
// Number of manifests not read because the name above them under
// metadata/ is not a remote key. The names themselves are not
// reported: they come from the destination store and may hold
// control characters.
SkippedManifestCount int `json:"skipped_manifest_count,omitempty"`
}
// RemoteInfo displays information about remote storage
@@ -234,16 +250,28 @@ func (v *Vaultik) RemoteInfo(jsonOutput bool) error {
v.stdoutf("Scanning snapshot metadata...\n")
}
snapshotMetadata, snapshotIDs, err := v.collectSnapshotMetadata()
snapshotMetadata, snapshotIDs, skippedManifestCount, err := v.collectSnapshotMetadata()
if err != nil {
return err
}
result.SkippedManifestCount = skippedManifestCount
if showText {
v.stdoutf("Downloading %d manifest(s)...\n", len(snapshotIDs))
manifestCount := 0
for _, info := range snapshotMetadata {
if info.hasManifest {
manifestCount++
}
}
v.stdoutf("Downloading %d manifest(s)...\n", manifestCount)
}
referencedBlobs := v.collectReferencedBlobsFromManifests(snapshotIDs, snapshotMetadata)
referencedBlobs, unreadableManifests := v.collectReferencedBlobsFromManifests(
snapshotIDs, snapshotMetadata)
result.UnreadableManifests = unreadableManifests
v.populateRemoteInfoResult(result, snapshotMetadata, snapshotIDs, referencedBlobs)
@@ -256,7 +284,7 @@ func (v *Vaultik) RemoteInfo(jsonOutput bool) error {
"snapshots", result.TotalMetadataCount,
"total_blobs", result.TotalBlobCount,
"referenced_blobs", result.ReferencedBlobCount,
"orphaned_blobs", result.OrphanedBlobCount)
"unreadable_manifests", len(result.UnreadableManifests))
if jsonOutput {
enc := json.NewEncoder(v.Stdout)
@@ -273,16 +301,18 @@ func (v *Vaultik) RemoteInfo(jsonOutput bool) error {
}
// collectSnapshotMetadata scans remote metadata and returns
// per-snapshot info and sorted IDs.
// per-snapshot info, sorted IDs and the number of manifests it skipped
// because the name above them is not a remote key.
func (v *Vaultik) collectSnapshotMetadata() (
map[string]*SnapshotMetadataInfo, []string, error,
map[string]*SnapshotMetadataInfo, []string, int, error,
) {
snapshotMetadata := make(map[string]*SnapshotMetadataInfo)
skippedManifestCount := 0
metadataCh := v.Storage.ListStream(v.ctx, "metadata/")
for obj := range metadataCh {
if obj.Err != nil {
return nil, nil, fmt.Errorf("listing metadata: %w", obj.Err)
return nil, nil, 0, fmt.Errorf("listing metadata: %w", obj.Err)
}
parts := strings.Split(obj.Key, "/")
@@ -291,6 +321,22 @@ func (v *Vaultik) collectSnapshotMetadata() (
}
snapshotID := parts[1]
filename := parts[2]
isManifest := filename == "manifest.json.zst"
// The name comes from the destination store, which is not
// trusted, and is printed in the report. Accept it only in the
// form of a remote key.
if !isBlobHash(snapshotID) {
log.Warn("Skipping non-conforming key under metadata/",
"key", obj.Key)
if isManifest {
skippedManifestCount++
}
continue
}
if _, exists := snapshotMetadata[snapshotID]; !exists {
snapshotMetadata[snapshotID] = &SnapshotMetadataInfo{SnapshotID: snapshotID}
@@ -298,7 +344,10 @@ func (v *Vaultik) collectSnapshotMetadata() (
info := snapshotMetadata[snapshotID]
filename := parts[2]
if isManifest {
info.hasManifest = true
}
if strings.HasPrefix(filename, "manifest") {
info.ManifestSize = obj.Size
} else if strings.HasPrefix(filename, "db") {
@@ -315,17 +364,25 @@ func (v *Vaultik) collectSnapshotMetadata() (
sort.Strings(snapshotIDs)
return snapshotMetadata, snapshotIDs, nil
return snapshotMetadata, snapshotIDs, skippedManifestCount, nil
}
// collectReferencedBlobsFromManifests downloads manifests and returns
// referenced blob hashes with sizes.
// collectReferencedBlobsFromManifests downloads the listed manifests
// and returns referenced blob hashes with sizes, and the remote keys
// of the manifests it could not read.
func (v *Vaultik) collectReferencedBlobsFromManifests(
snapshotIDs []string, snapshotMetadata map[string]*SnapshotMetadataInfo,
) map[string]int64 {
) (map[string]int64, []string) {
referencedBlobs := make(map[string]int64)
var unreadable []string
for _, snapshotID := range snapshotIDs {
info := snapshotMetadata[snapshotID]
if !info.hasManifest {
continue
}
// snapshotIDs here are remote keys, taken straight from the
// metadata/ listing. downloadManifestByKey is the single reader
// for remote manifests; see its doc comment.
@@ -333,10 +390,11 @@ func (v *Vaultik) collectReferencedBlobsFromManifests(
if err != nil {
log.Warn("Failed to read manifest", "snapshot", snapshotID, "error", err)
unreadable = append(unreadable, snapshotID)
continue
}
info := snapshotMetadata[snapshotID]
info.BlobCount = manifest.BlobCount
var blobsSize int64
@@ -349,7 +407,7 @@ func (v *Vaultik) collectReferencedBlobsFromManifests(
info.BlobsSize = blobsSize
}
return referencedBlobs
return referencedBlobs, unreadable
}
// populateRemoteInfoResult fills in the result's snapshot and
@@ -378,8 +436,9 @@ func (v *Vaultik) populateRemoteInfoResult(
}
// scanRemoteBlobStorage lists all blobs on remote and computes orphan
// stats. showText is true only when the human report is being printed
// (not --json, not --quiet), gating the progress line.
// stats when every listed manifest was read. showText is true only
// when the human report is being printed (not --json, not --quiet),
// gating the progress line.
func (v *Vaultik) scanRemoteBlobStorage(
result *RemoteInfoResult, referencedBlobs map[string]int64, showText bool,
) error {
@@ -406,13 +465,28 @@ func (v *Vaultik) scanRemoteBlobStorage(
result.TotalBlobSize += obj.Size
}
// A blob named only by a manifest that could not be read, or by one
// under a skipped name, would be counted as orphaned, so the orphan
// figures stay unknown.
if len(result.UnreadableManifests) > 0 || result.SkippedManifestCount > 0 {
return nil
}
var (
orphanedCount int
orphanedSize int64
)
for hash, size := range allBlobs {
if _, referenced := referencedBlobs[hash]; !referenced {
result.OrphanedBlobCount++
result.OrphanedBlobSize += size
orphanedCount++
orphanedSize += size
}
}
result.OrphanedBlobCount = &orphanedCount
result.OrphanedBlobSize = &orphanedSize
return nil
}
@@ -465,11 +539,21 @@ func (v *Vaultik) printRemoteInfoTable(result *RemoteInfoResult) {
v.stdoutf("Referenced by snapshots: %s (%s)\n",
humanize.Comma(int64(result.ReferencedBlobCount)),
ubytes(result.ReferencedBlobSize))
v.stdoutf("Orphaned (unreferenced): %s (%s)\n",
humanize.Comma(int64(result.OrphanedBlobCount)),
ubytes(result.OrphanedBlobSize))
if result.OrphanedBlobCount > 0 {
if result.OrphanedBlobCount == nil {
v.stdoutf("Orphaned (unreferenced): unknown "+
"(%d manifest(s) could not be read, "+
"%d manifest(s) under a non-conforming name skipped)\n",
len(result.UnreadableManifests), result.SkippedManifestCount)
return
}
v.stdoutf("Orphaned (unreferenced): %s (%s)\n",
humanize.Comma(int64(*result.OrphanedBlobCount)),
ubytes(*result.OrphanedBlobSize))
if *result.OrphanedBlobCount > 0 {
v.stdoutf("\nRun 'vaultik prune' to remove orphaned blobs.\n")
}
}
+1 -1
View File
@@ -249,7 +249,7 @@ func verifyEndToEndBackupState(
assert.Positive(t, blobUploads, "Should upload at least one blob")
// Verify files in database
files, err := repos.Files.ListByPrefix(ctx, "/home/user")
files, err := repos.Files.ListUnderPath(ctx, "/home/user")
require.NoError(t, err)
// Count only regular files (not directories)
regularFiles := 0
@@ -0,0 +1,176 @@
package vaultik_test
import (
"bytes"
"context"
"io/fs"
"os"
"path/filepath"
"strings"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/ui"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// These tests cover https://git.eeqj.de/sneak/vaultik/issues/220: a
// file:// destination whose directory is missing, such as a USB stick
// that is not plugged in, cannot be listed. It is not an empty store, so
// no command may conclude from it that the local snapshots are gone.
// backUpToFileDestination backs up the snapshot named "first" to a
// file:// destination at storeDir, which need not exist yet. Everything
// the returned Vaultik prints after the backup goes to the returned
// buffer.
func backUpToFileDestination(
ctx context.Context, t *testing.T, storeDir string,
) (*vaultik.Vaultik, *database.Repositories, *bytes.Buffer) {
t.Helper()
osFs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
dbPath := filepath.Join(tempDir, "index.sqlite")
writeFaultSourceTree(t, osFs, dataDir)
store, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
t.Cleanup(func() { _ = db.Close() })
repos := database.NewRepositories(db)
cfg := changedFileConfig(dataDir, dbPath)
v := newBackupVaultik(ctx, cfg, store, repos, db, osFs)
require.NoError(t, backUp(v, "first"))
out := &bytes.Buffer{}
v.Stdout = out
v.UI = ui.NewWithColor(out, false)
return v, repos, out
}
// backUpThenUnplug backs up to a file:// destination, then moves the
// destination directory away, as unplugging the volume it lives on would.
func backUpThenUnplug(
ctx context.Context, t *testing.T,
) (*vaultik.Vaultik, *database.Repositories, *bytes.Buffer) {
t.Helper()
storeDir := filepath.Join(t.TempDir(), "usbstick")
v, repos, out := backUpToFileDestination(ctx, t, storeDir)
require.NoError(t, os.Rename(storeDir, storeDir+"-unplugged"))
return v, repos, out
}
// TestFirstBackupCreatesDestinationDirectory checks that a first backup
// to a destination directory that does not exist yet creates it, and
// that the destination can be listed afterwards.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestFirstBackupCreatesDestinationDirectory(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
storeDir := filepath.Join(t.TempDir(), "volume", "backup")
v, _, out := backUpToFileDestination(ctx, t, storeDir)
require.NoError(t, v.ListSnapshots(false))
assert.NotContains(t, out.String(), "Could not list backup destination store")
assert.NotContains(t, out.String(), "not found in backup destination store")
}
// TestListSnapshotsWarnsWhenDestinationMissing checks that snapshot list
// warns and shows the local index alone, without reporting the local
// snapshot as missing from the destination.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestListSnapshotsWarnsWhenDestinationMissing(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
v, repos, out := backUpThenUnplug(ctx, t)
id := localSnapshotID(ctx, t, repos, "first")
require.NoError(t, v.ListSnapshots(false))
assert.Contains(t, out.String(), "Could not list backup destination store")
assert.Contains(t, out.String(), "Showing snapshots from the local index only.")
assert.Contains(t, out.String(), id)
assert.NotContains(t, out.String(), "not found in backup destination store")
}
// TestRemoveSnapshotWarnsWhenDestinationMissing checks that snapshot
// remove warns that the metadata could not be removed from the
// destination, instead of reporting that it was.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestRemoveSnapshotWarnsWhenDestinationMissing(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
v, repos, out := backUpThenUnplug(ctx, t)
result, err := v.RemoveSnapshot(localSnapshotID(ctx, t, repos, "first"),
&vaultik.RemoveOptions{Force: true})
require.NoError(t, err)
assert.False(t, result.RemoteRemoved)
assert.Contains(t, out.String(),
"Could not remove snapshot metadata from remote")
assert.NotContains(t, out.String(),
"Removed snapshot metadata from remote storage")
}
// TestPruneKeepsLocalRecordsWhenDestinationMissing checks that prune
// fails on a destination it cannot list and deletes no local snapshot
// record.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestPruneKeepsLocalRecordsWhenDestinationMissing(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
v, repos, _ := backUpThenUnplug(ctx, t)
err := v.Prune(&vaultik.PruneOptions{Force: true})
require.ErrorIs(t, err, fs.ErrNotExist)
require.ErrorContains(t, err, "listing remote snapshots")
snapshots, err := repos.Snapshots.ListRecent(ctx, listRecentTestLimit)
require.NoError(t, err)
assert.Len(t, snapshots, 1, "prune must delete no local snapshot record")
}
// TestPurgeSaysListingFailedOnceWhenDestinationMissing checks that
// snapshot purge fails on a destination it cannot list, with an error
// that says "listing remote snapshots" once.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestPurgeSaysListingFailedOnceWhenDestinationMissing(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
v, _, _ := backUpThenUnplug(ctx, t)
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
KeepLatest: true,
Force: true,
})
require.ErrorIs(t, err, fs.ErrNotExist)
assert.Equal(t, 1, strings.Count(err.Error(), "listing remote snapshots"),
err.Error())
}
@@ -42,7 +42,7 @@ func setupConsistencyTest(
completedAt := startedAt.Add(5 * time.Minute)
snap := &database.Snapshot{
ID: types.SnapshotID(id),
Hostname: testHostname,
Hostname: snapHostname,
VaultikVersion: testLabel,
StartedAt: startedAt,
CompletedAt: &completedAt,
+42 -12
View File
@@ -17,8 +17,10 @@ import (
"sneak.berlin/go/vaultik/internal/vaultik"
)
// Snapshot IDs reused across the purge tests.
// Snapshot IDs reused across the purge tests, and the hostname they were
// taken on.
const (
snapHostname = "testhost"
snapSystemT0 = "testhost_system_2026-01-01T00:00:00Z"
snapHomeT0 = "testhost_home_2026-01-01T00:00:00Z"
snapHomeT1 = "testhost_home_2026-01-01T01:00:00Z"
@@ -26,9 +28,12 @@ const (
)
// setupPurgeTest creates a Vaultik instance with an in-memory database and mock
// storage pre-populated with the given snapshot IDs. Each snapshot is marked as
// completed. Remote metadata stubs are created so syncWithRemote keeps them.
func setupPurgeTest(t *testing.T, snapshotIDs []string) *vaultik.Vaultik {
// storage pre-populated with the given snapshot IDs, all taken on hostname.
// Each snapshot is marked as completed. Remote metadata stubs are created so
// syncWithRemote keeps them.
func setupPurgeTest(
t *testing.T, hostname string, snapshotIDs []string,
) *vaultik.Vaultik {
t.Helper()
ctx := context.Background()
@@ -51,7 +56,7 @@ func setupPurgeTest(t *testing.T, snapshotIDs []string) *vaultik.Vaultik {
completedAt := startedAt.Add(5 * time.Minute)
snap := &database.Snapshot{
ID: types.SnapshotID(id),
Hostname: "testhost",
Hostname: types.Hostname(hostname),
VaultikVersion: testLabel,
StartedAt: startedAt,
CompletedAt: &completedAt,
@@ -120,7 +125,7 @@ func TestPurgeKeepLatest_PerName(t *testing.T) {
"testhost_system_2026-01-01T04:00:00Z",
}
v := setupPurgeTest(t, snapshotIDs)
v := setupPurgeTest(t, snapHostname, snapshotIDs)
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
KeepLatest: true,
@@ -148,7 +153,7 @@ func TestPurgeKeepLatest_SingleName(t *testing.T) {
"testhost_home_2026-01-01T02:00:00Z",
}
v := setupPurgeTest(t, snapshotIDs)
v := setupPurgeTest(t, snapHostname, snapshotIDs)
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
KeepLatest: true,
@@ -176,7 +181,7 @@ func TestPurgeKeepLatest_WithNameFilter(t *testing.T) {
"testhost_home_2026-01-01T04:00:00Z",
}
v := setupPurgeTest(t, snapshotIDs)
v := setupPurgeTest(t, snapHostname, snapshotIDs)
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
KeepLatest: true,
@@ -198,7 +203,7 @@ func TestPurgeKeepLatest_NoSnapshots(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
v := setupPurgeTest(t, nil)
v := setupPurgeTest(t, snapHostname, nil)
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
KeepLatest: true,
@@ -216,7 +221,7 @@ func TestPurgeKeepLatest_NameFilterNoMatch(t *testing.T) {
"testhost_system_2026-01-01T01:00:00Z",
}
v := setupPurgeTest(t, snapshotIDs)
v := setupPurgeTest(t, snapHostname, snapshotIDs)
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
KeepLatest: true,
@@ -243,7 +248,7 @@ func TestPurgeOlderThan_WithNameFilter(t *testing.T) {
snapHomeT0,
}
v := setupPurgeTest(t, snapshotIDs)
v := setupPurgeTest(t, snapHostname, snapshotIDs)
// Purge only "home" snapshots older than 365 days
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
@@ -277,7 +282,7 @@ func TestPurgeKeepLatest_ThreeNames(t *testing.T) {
"testhost_home_2026-01-01T06:00:00Z",
}
v := setupPurgeTest(t, snapshotIDs)
v := setupPurgeTest(t, snapHostname, snapshotIDs)
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
KeepLatest: true,
@@ -291,3 +296,28 @@ func TestPurgeKeepLatest_ThreeNames(t *testing.T) {
assert.Contains(t, remaining, "testhost_system_2026-01-01T04:00:00Z")
assert.Contains(t, remaining, "testhost_media_2026-01-01T05:00:00Z")
}
// A hostname may contain underscores, so the snapshot name cannot be found
// by splitting the ID at them. A purge by name must still select "docs".
func TestPurgeKeepLatest_HostnameWithUnderscore(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
const (
system = "my_host_system_2026-01-01T00:00:00Z"
docsT1 = "my_host_docs_2026-01-01T01:00:00Z"
docsT2 = "my_host_docs_2026-01-01T02:00:00Z"
)
v := setupPurgeTest(t, "my_host", []string{system, docsT1, docsT2})
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
KeepLatest: true,
Force: true,
Names: []string{"docs"},
})
require.NoError(t, err)
assert.ElementsMatch(t, []string{system, docsT2},
listRemainingSnapshots(t, v))
}
+159
View File
@@ -0,0 +1,159 @@
package vaultik_test
import (
"bytes"
"context"
"encoding/json"
"testing"
"time"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/snapshot"
)
// testBlobHashB is a blob that the manifest written by addRemote does
// not reference.
const testBlobHashB = "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb" +
"bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb"
// TestRemoteInfo_UnreadableManifestLeavesOrphansUnknown checks that a
// manifest remote info cannot read makes the orphan figures unknown. A
// blob referenced only by that snapshot would otherwise be counted as
// orphaned, and the report would advise running prune.
func TestRemoteInfo_UnreadableManifestLeavesOrphansUnknown(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
env := newListEnv(t)
// The readable manifest references blob A only.
env.addRemote(t, listRemoteID, time.Date(2026, 3, 2, 0, 0, 0, 0, time.UTC))
addBlob(t, env.store.testStorer, testBlobHashA)
addBlob(t, env.store.testStorer, testBlobHashB)
// With every manifest readable, blob B is orphaned.
require.NoError(t, env.v.RemoteInfo(true))
var doc map[string]any
require.NoError(t, json.Unmarshal(env.stdout.Bytes(), &doc))
assert.InDelta(t, 1, doc["orphaned_blob_count"], 0)
// A second snapshot whose manifest cannot be decoded. Blob B may be
// one of its blobs.
unreadableKey := snapshot.RemoteSnapshotKey(listLocalID)
require.NoError(t, env.store.Put(context.Background(),
"metadata/"+unreadableKey+"/manifest.json.zst",
bytes.NewReader([]byte("not a valid manifest"))))
env.stdout.Reset()
require.NoError(t, env.v.RemoteInfo(false))
text := env.stdout.String()
assert.Contains(t, text, "Orphaned (unreferenced): unknown "+
"(1 manifest(s) could not be read, "+
"0 manifest(s) under a non-conforming name skipped)")
assert.NotContains(t, text, "vaultik prune")
env.stdout.Reset()
require.NoError(t, env.v.RemoteInfo(true))
doc = nil
require.NoError(t, json.Unmarshal(env.stdout.Bytes(), &doc))
assert.Contains(t, doc, "orphaned_blob_count")
assert.Nil(t, doc["orphaned_blob_count"])
assert.Contains(t, doc, "orphaned_blob_size")
assert.Nil(t, doc["orphaned_blob_size"])
assert.Equal(t, []any{unreadableKey}, doc["unreadable_manifests"])
}
// TestRemoteInfo_SkipsNonConformingMetadataName checks that a directory
// under metadata/ whose name is not a remote key is left out of the
// report, and that the orphan figures are unknown when it holds a
// manifest. The name comes from the destination store; printed raw, its
// control characters would reach the terminal. Its manifest is not
// read, so a blob only it references would otherwise be counted as
// orphaned.
func TestRemoteInfo_SkipsNonConformingMetadataName(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
env := newListEnv(t)
env.addRemote(t, listRemoteID, time.Date(2026, 3, 2, 0, 0, 0, 0, time.UTC))
addBlob(t, env.store.testStorer, testBlobHashA)
addBlob(t, env.store.testStorer, testBlobHashB)
require.NoError(t, env.store.Put(context.Background(),
"metadata/\x1b[31mred/manifest.json.zst",
bytes.NewReader([]byte("not a valid manifest"))))
require.NoError(t, env.v.RemoteInfo(false))
text := env.stdout.String()
assert.NotContains(t, text, "\x1b")
assert.NotContains(t, text, "31mred")
assert.Contains(t, text, "Total (1 snapshots)")
assert.Contains(t, text, "Orphaned (unreferenced): unknown "+
"(0 manifest(s) could not be read, "+
"1 manifest(s) under a non-conforming name skipped)")
assert.NotContains(t, text, "vaultik prune")
env.stdout.Reset()
require.NoError(t, env.v.RemoteInfo(true))
out := env.stdout.String()
assert.NotContains(t, out, "31mred")
var doc map[string]any
require.NoError(t, json.Unmarshal([]byte(out), &doc))
assert.Contains(t, doc, "orphaned_blob_count")
assert.Nil(t, doc["orphaned_blob_count"])
assert.Contains(t, doc, "orphaned_blob_size")
assert.Nil(t, doc["orphaned_blob_size"])
assert.InDelta(t, 1, doc["skipped_manifest_count"], 0)
assert.NotContains(t, doc, "unreadable_manifests")
}
// TestRemoteInfo_DirectoryWithoutManifestLeavesOrphansKnown checks that
// a directory under metadata/ holding no manifest.json.zst, such as one
// left by a backup interrupted before its manifest upload, leaves the
// orphan figures known. prune does not treat such a directory as a
// snapshot and deletes the blobs the report lists as orphaned.
func TestRemoteInfo_DirectoryWithoutManifestLeavesOrphansKnown(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
env := newListEnv(t)
env.addRemote(t, listRemoteID, time.Date(2026, 3, 2, 0, 0, 0, 0, time.UTC))
addBlob(t, env.store.testStorer, testBlobHashA)
addBlob(t, env.store.testStorer, testBlobHashB)
// One directory under a remote key and one under a non-conforming
// name, each holding only a database.
names := []string{snapshot.RemoteSnapshotKey(listLocalID), "\x1b[31mred"}
for _, name := range names {
require.NoError(t, env.store.Put(context.Background(),
"metadata/"+name+"/db.zst.age",
bytes.NewReader([]byte("not a valid database"))))
}
require.NoError(t, env.v.RemoteInfo(false))
text := env.stdout.String()
assert.NotContains(t, text, "\x1b")
assert.Contains(t, text, "Downloading 1 manifest(s)...")
assert.Contains(t, text, "Orphaned (unreferenced): 1 (")
assert.Contains(t, text, "Run 'vaultik prune' to remove orphaned blobs.")
env.stdout.Reset()
require.NoError(t, env.v.RemoteInfo(true))
var doc map[string]any
require.NoError(t, json.Unmarshal(env.stdout.Bytes(), &doc))
assert.InDelta(t, 1, doc["orphaned_blob_count"], 0)
assert.NotContains(t, doc, "unreadable_manifests")
assert.NotContains(t, doc, "skipped_manifest_count")
}
+142 -37
View File
@@ -10,11 +10,13 @@ import (
"math"
"os"
"path/filepath"
"slices"
"strings"
"time"
"filippo.io/age"
"github.com/spf13/afero"
"golang.org/x/sys/unix"
"sneak.berlin/go/vaultik/internal/blobgen"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
@@ -40,6 +42,7 @@ var (
errChunkNotInAnyBlob = errors.New("chunk not found in any blob")
errBlobIDNotInHashIndex = errors.New("blob id missing from hash index")
errShortChunkRead = errors.New("short read")
errChunkRowMissing = errors.New("chunk has no row in the chunks table")
errRestorePathEscapesTarget = errors.New(
"refusing to restore path outside the target directory")
errTrailingRestoreData = errors.New(
@@ -63,6 +66,12 @@ const snapshotDBFilename = "snapshot.db"
// directories themselves get their stored mode).
const restoreDirMode = 0o755
// restoreDirCreateMode is the owner-only mode a directory from the
// snapshot is created with during restore, so its contents can be written
// whatever its stored mode. The stored mode is applied after the restore
// loop, by applyDirectoryMetadata.
const restoreDirCreateMode = 0o700
// restoreFileMode is the restrictive mode a regular file is created with
// during restore. Content is written while the file holds this mode; the
// stored mode is applied only after the file is fully written and closed,
@@ -332,6 +341,8 @@ func (v *Vaultik) restoreAllFiles(
return nil, err
}
session.applyDirectoryMetadata()
return result, nil
}
@@ -355,7 +366,7 @@ func (v *Vaultik) runRestoreLoop(
fileID, ready := plan.popReady()
if !ready {
downloaded, err := session.downloadNextBlobSet(plan)
downloaded, err := session.downloadNextBlobSet(plan, filesByID)
if err != nil {
return err
}
@@ -411,7 +422,13 @@ func (v *Vaultik) runRestoreLoop(
// blob set and downloads its blobs; after each blob lands, the plan
// moves any pending file whose set just emptied onto the ready queue.
// Returns false when nothing is pending download (the caller stops).
func (s *restoreSession) downloadNextBlobSet(plan *restorePlan) (bool, error) {
//
// A blob that cannot be downloaded is reported through
// handleRestoreFileError for every pending file that references it, so
// it aborts the restore unless --skip-errors is set.
func (s *restoreSession) downloadNextBlobSet(
plan *restorePlan, filesByID map[types.FileID]*database.File,
) (bool, error) {
s.sweeper.sweep()
next, ok := plan.pickNextDownload()
@@ -433,7 +450,25 @@ func (s *restoreSession) downloadNextBlobSet(plan *restorePlan) (bool, error) {
err := s.downloadBlobToCache(hash, blob.CompressedSize, blob.UncompressedSize)
if err != nil {
return false, fmt.Errorf("downloading blob %s: %w", shortHash(hash), err)
err = fmt.Errorf("downloading blob %s: %w", shortHash(hash), err)
// On cancel the error says nothing about the blob, so it ends
// the restore instead of failing the files that need it.
if s.ctx.Err() != nil {
return false, err
}
for _, fileID := range plan.filesReferencingBlob(hash) {
fileErr := s.v.handleRestoreFileError(
plan, s.opts, s.result, filesByID[fileID], fileID, err)
if fileErr != nil {
return false, fileErr
}
}
// next is among the failed files, so the rest of its blob set
// is left for any other file that still needs it.
return true, nil
}
s.result.BlobsDownloaded++
@@ -795,10 +830,9 @@ func (v *Vaultik) getFilesToRestore(
// Normalize the filter path
filter = filepath.Clean(filter)
// Get files with this prefix
files, err := repos.Files.ListByPrefix(ctx, filter)
files, err := repos.Files.ListUnderPath(ctx, filter)
if err != nil {
return nil, fmt.Errorf("listing files with prefix %s: %w", filter, err)
return nil, fmt.Errorf("listing files under %s: %w", filter, err)
}
for _, file := range files {
@@ -877,6 +911,9 @@ type restoreSession struct {
// the call entirely as non-root and emit one warning at the end
// of the restore explaining that ownership was not preserved.
runningAsRoot bool
// directories holds every restored directory, for
// applyDirectoryMetadata to finish after the restore loop.
directories []*database.File
}
// containedRestorePath resolves rel — a path read from the snapshot
@@ -983,6 +1020,8 @@ func (s *restoreSession) restoreSymlink(file *database.File, targetPath string)
if err != nil {
return fmt.Errorf("creating symlink: %w", err)
}
s.applySymlinkMetadata(file, targetPath)
} else {
log.Debug("Symlink creation not supported on this filesystem",
"path", file.Path, "target", file.LinkTarget)
@@ -995,35 +1034,90 @@ func (s *restoreSession) restoreSymlink(file *database.File, targetPath string)
return nil
}
// restoreDirectory restores a directory with its permissions, mtime,
// and (on real filesystems, with sufficient privileges) ownership.
// restoreDirectory creates a directory with restoreDirCreateMode. Its
// stored mode, owner and mtime are applied after the restore loop by
// applyDirectoryMetadata: a read-only stored mode would block writing
// its contents, and writing them changes its mtime.
func (s *restoreSession) restoreDirectory(
file *database.File, targetPath string,
) error {
err := s.v.Fs.MkdirAll(targetPath, os.FileMode(file.Mode))
err := s.v.Fs.MkdirAll(targetPath, restoreDirCreateMode)
if err != nil {
return fmt.Errorf("creating directory: %w", err)
}
// MkdirAll applies the process umask, so chmod to the exact stored
// mode. A failure here is non-fatal.
err = s.v.Fs.Chmod(targetPath, os.FileMode(file.Mode))
if err != nil {
log.Debug("Failed to set permissions", "path", targetPath, "error", err)
}
s.applyFileMetadata(file, targetPath)
s.directories = append(s.directories, file)
s.result.FilesRestored++
return nil
}
// applyDirectoryMetadata applies the stored owner, mtime and mode to every
// restored directory, each before its parent, so a parent whose stored
// mode denies search does not block its children. Failures are logged at
// debug level and do not abort the restore.
func (s *restoreSession) applyDirectoryMetadata() {
// A path sorts after its parent's, so reverse order puts every
// directory before its parent.
slices.SortFunc(s.directories, func(a, b *database.File) int {
return strings.Compare(b.Path.String(), a.Path.String())
})
for _, dir := range s.directories {
targetPath, err := containedRestorePath(
s.v.Fs, s.opts.TargetDir, dir.Path.String())
if err != nil {
log.Debug("Failed to set directory metadata",
"path", dir.Path, "error", err)
continue
}
// A later entry can have put a symlink in the directory's place,
// for example one stored as "/d/" next to the directory "/d". The
// calls below follow symlinks, so they would change its target.
info, err := lstatIfPossible(s.v.Fs, targetPath)
if err != nil || !info.IsDir() {
log.Debug("Not setting directory metadata: no longer a directory",
"path", targetPath, "error", err)
continue
}
s.applyFileMetadata(dir, targetPath)
err = s.v.Fs.Chmod(targetPath, os.FileMode(dir.Mode))
if err != nil {
log.Debug("Failed to set permissions", "path", targetPath, "error", err)
}
}
}
// applySymlinkMetadata applies ownership (when running as root) and mtime
// to a restored symlink itself; os.Chown and Chtimes would follow it.
// Failures are logged at debug level and do not abort the restore.
func (s *restoreSession) applySymlinkMetadata(file *database.File, targetPath string) {
if s.runningAsRoot {
err := os.Lchown(targetPath, int(file.UID), int(file.GID))
if err != nil {
log.Debug("Failed to set ownership", "path", targetPath, "error", err)
}
}
mtime := unix.NsecToTimeval(file.MTime.UnixNano())
err := unix.Lutimes(targetPath, []unix.Timeval{mtime, mtime})
if err != nil {
log.Debug("Failed to set mtime", "path", targetPath, "error", err)
}
}
// applyFileMetadata applies ownership (when running as root on a real
// filesystem) and mtime to a restored path. Permission mode is applied
// separately by each caller, with different failure handling, so it is
// not touched here. Failures are logged at debug level and do not abort
// the restore.
// filesystem) and mtime to a restored path. The caller applies the mode
// afterwards: on Linux a chown clears the setuid and setgid bits of a
// regular file. Failures are logged at debug level and do not abort the
// restore.
func (s *restoreSession) applyFileMetadata(file *database.File, targetPath string) {
if s.runningAsRoot {
if _, ok := s.v.Fs.(*afero.OsFs); ok {
@@ -1114,8 +1208,8 @@ func (s *restoreSession) restoreRegularFile(
return fmt.Errorf("closing output file: %w", err)
}
s.applyRestoredFileMode(file, targetPath)
s.applyFileMetadata(file, targetPath)
s.applyRestoredFileMode(file, targetPath)
s.result.FilesRestored++
s.result.BytesRestored += bytesWritten
@@ -1173,7 +1267,7 @@ func (s *restoreSession) writeFileChunks(
blobChunk, ok := s.chunkToBlobMap[chunkHashStr]
if !ok {
return bytesWritten, timings, fmt.Errorf(
"%w: %s", errChunkNotInAnyBlob, chunkHashStr[:16])
"%w: %s", errChunkNotInAnyBlob, shortHash(chunkHashStr))
}
blobHash, ok := s.blobIDToHash[blobChunk.BlobID.String()]
@@ -1190,7 +1284,7 @@ func (s *restoreSession) writeFileChunks(
if err != nil {
return bytesWritten, timings, fmt.Errorf(
"reading chunk %s from cached blob %s: %w",
fc.ChunkHash[:16], blobHash[:16], err)
shortHash(chunkHashStr), shortHash(blobHash), err)
}
t0 = time.Now()
@@ -1388,33 +1482,44 @@ func (v *Vaultik) verifyFile(
chunk, err := repos.Chunks.GetByHash(ctx, fc.ChunkHash.String())
if err != nil {
return bytesVerified, fmt.Errorf("getting chunk %s: %w",
fc.ChunkHash.String()[:16], err)
shortHash(fc.ChunkHash.String()), err)
}
// Read chunk data from file
chunkData := make([]byte, chunk.Size)
n, err := io.ReadFull(f, chunkData)
if err != nil {
return bytesVerified, fmt.Errorf("reading chunk data: %w", err)
if chunk == nil {
return bytesVerified, fmt.Errorf("%w: %s",
errChunkRowMissing, shortHash(fc.ChunkHash.String()))
}
if int64(n) != chunk.Size {
// chunk.Size comes from the snapshot database, which is not
// trusted: reject a negative size, and hash the chunk by
// streaming it rather than allocating that many bytes.
if chunk.Size < 0 {
return bytesVerified, fmt.Errorf("%w: chunk %d size %d",
errNegativeChunkLength, fc.Idx, chunk.Size)
}
hasher := sha256.New()
n, err := io.CopyN(hasher, f, chunk.Size)
if errors.Is(err, io.EOF) {
return bytesVerified, fmt.Errorf("%w: expected %d bytes, got %d",
errShortChunkRead, chunk.Size, n)
}
// Calculate hash and compare
hash := sha256.Sum256(chunkData)
actualHash := hex.EncodeToString(hash[:])
if err != nil {
return bytesVerified, fmt.Errorf("reading chunk data: %w", err)
}
actualHash := hex.EncodeToString(hasher.Sum(nil))
expectedHash := fc.ChunkHash.String()
if actualHash != expectedHash {
return bytesVerified, fmt.Errorf("%w: chunk %d: expected %s, got %s",
errChunkHashMismatch, fc.Idx, expectedHash[:16], actualHash[:16])
errChunkHashMismatch, fc.Idx,
shortHash(expectedHash), shortHash(actualHash))
}
bytesVerified += int64(n)
bytesVerified += n
}
// The stored chunks account for the whole file, so the reader must
@@ -10,6 +10,7 @@ import (
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/cli"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
@@ -27,10 +28,12 @@ import (
//
// The backup half writes a snapshot with one index and hostname. The
// restore half throws that index away entirely: a fresh, empty index and
// a config that shares nothing with the original but the storage location
// and the secret key. If restore or verify needed the original local
// index — or the human snapshot ID that only that index holds — this test
// could not run, because the recovery host can know neither.
// a config written by `config init` and `config set storage_url`, as in
// the README's steps for restoring on another machine, which shares
// nothing with the original but the storage location and the secret key.
// If restore or verify needed the original local index — or
// the human snapshot ID that only that index holds — this test could not
// run, because the recovery host can know neither.
func TestRestoreOnAnotherMachine(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
@@ -58,7 +61,12 @@ func TestRestoreOnAnotherMachine(t *testing.T) {
// Recovery host: a fresh empty index, a different hostname, and no
// age_recipients — only the secret key and the same storage location.
recovery, stdout := newRecoveryHost(ctx, t, fs, storer)
configPath := filepath.Join(tempDir, "config.yml")
runVaultikCommand(t, "--config", configPath, "config", "init")
runVaultikCommand(t, "--config", configPath,
"config", "set", "storage_url", "file://"+storeDir)
recovery, stdout := newRecoveryHost(ctx, t, fs, storer, configPath)
// The recovery index really is empty. This is the assertion that makes
// the test a guard against restore quietly depending on the original
@@ -93,6 +101,10 @@ func TestRestoreOnAnotherMachine(t *testing.T) {
remote.RemoteKey, &vaultik.VerifyOptions{Deep: true}))
assertRestoredTreeMatches(t, fs, restoreDir, sourceFiles)
// With no public key configured, a backup must refuse to start.
err = recovery.CreateSnapshot(&vaultik.SnapshotCreateOptions{Cron: true})
require.ErrorContains(t, err, "age_recipients")
}
// writeRecoverySourceTree writes a small source tree spanning several
@@ -118,14 +130,24 @@ func writeRecoverySourceTree(
}
// newRecoveryHost builds the Vaultik a replacement machine would run: an
// empty in-memory index, a hostname different from the backup host, no
// age_recipients, and only the secret key plus the shared storer. It
// returns the instance and the buffer its stdout is wired to.
// empty in-memory index, a hostname different from the backup host, and
// only the secret key plus the shared storer. Its config is read from
// configPath by config.Load, as every command reads it. It returns the
// instance and the buffer its stdout is wired to.
func newRecoveryHost(
ctx context.Context, t *testing.T, fs afero.Fs, storer storage.Storer,
configPath string,
) (*vaultik.Vaultik, *bytes.Buffer) {
t.Helper()
cfg, err := config.Load(configPath)
require.NoError(t, err)
// Set directly rather than through VAULTIK_AGE_SECRET_KEY, which a
// parallel test cannot change.
cfg.AgeSecretKey = testAgeSecretKey
cfg.Hostname = "recovery-host"
recoveryDB, err := database.New(ctx, ":memory:")
require.NoError(t, err)
t.Cleanup(func() { _ = recoveryDB.Close() })
@@ -133,10 +155,7 @@ func newRecoveryHost(
stdout := &bytes.Buffer{}
recovery := &vaultik.Vaultik{
Config: &config.Config{
AgeSecretKey: testAgeSecretKey,
Hostname: "recovery-host",
},
Config: cfg,
Storage: storer,
Fs: fs,
Repositories: database.NewRepositories(recoveryDB),
@@ -150,6 +169,20 @@ func newRecoveryHost(
return recovery, stdout
}
// runVaultikCommand runs one vaultik command line in-process and fails the
// test if it returns an error. The command writes the cli package's global
// flag variables, so it must not be called from two tests that run at the
// same time.
func runVaultikCommand(t *testing.T, args ...string) {
t.Helper()
cmd := cli.NewRootCommand()
cmd.SetArgs(args)
cmd.SetOut(io.Discard)
cmd.SetErr(io.Discard)
require.NoError(t, cmd.Execute())
}
// assertRestoredTreeMatches byte-compares every restored file against its
// source content.
func assertRestoredTreeMatches(
@@ -1,6 +1,7 @@
package vaultik //nolint:testpackage // sets ctx/cancel and inspects scratch files
import (
"bytes"
"context"
"io"
"path/filepath"
@@ -11,6 +12,7 @@ import (
"time"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
@@ -157,3 +159,51 @@ func scratchEntries(t *testing.T, dir string) []string {
return matches
}
// TestRestoreSkipErrorsCancelDuringBlobDownload cancels a SkipErrors
// restore while a blob download is in progress. The download fails only
// because of the cancel, so Restore must return context.Canceled without
// reporting the file that needs the blob as failed.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestRestoreSkipErrorsCancelDuringBlobDownload(t *testing.T) {
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
cfg, storer, snapshotID, srcPath := backupOneFile(context.Background(),
t, fs, tempDir, "a.txt", []byte("hello vaultik"), 0o644)
gate := newBlockingBlobStorer(storer)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
var out bytes.Buffer
v := newRestoreVaultik(ctx, cfg, gate, fs)
v.UI = ui.NewWithColor(&out, false)
restoreErr := make(chan error, 1)
go func() {
restoreErr <- v.Restore(&RestoreOptions{
SnapshotID: snapshotID,
TargetDir: filepath.Join(tempDir, "restored"),
SkipErrors: true,
})
}()
select {
case <-gate.entered:
case <-time.After(30 * time.Second):
t.Fatal("restore never reached the blob-download phase")
}
cancel()
require.ErrorIs(t, <-restoreErr, context.Canceled)
assert.NotContains(t, out.String(), srcPath,
"the cancel was reported as a failed file")
}
@@ -0,0 +1,221 @@
package vaultik //nolint:testpackage // drives unexported restore and verify steps
import (
"context"
"math"
"path/filepath"
"strings"
"testing"
"time"
"github.com/spf13/afero"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/types"
)
// These tests feed restore and --verify a snapshot database written by
// hand, as a damaged or hostile store could serve one. Each malformed row
// must end in an error, not a panic.
// shortChunkHash is shorter than the hash prefix that error messages print.
const shortChunkHash = "abc"
// restoredFileContent is the content of the restored file under verify.
const restoredFileContent = "xyz"
// craftedSnapshotDB opens an empty snapshot database in a temp directory.
func craftedSnapshotDB(t *testing.T) (*database.DB, *database.Repositories) {
t.Helper()
db, err := database.New(context.Background(),
filepath.Join(t.TempDir(), "snapshot.db"))
require.NoError(t, err)
t.Cleanup(func() { _ = db.Close() })
return db, database.NewRepositories(db)
}
// craftedFile adds a regular file whose only chunk has the given hash.
// Adding the chunks row, if any, is left to the caller.
func craftedFile(
t *testing.T, repos *database.Repositories, chunkHash string,
) *database.File {
t.Helper()
ctx := context.Background()
file := &database.File{
Path: "/src/f",
MTime: time.Now().UTC(),
Size: int64(len(restoredFileContent)),
Mode: 0o644,
}
require.NoError(t, repos.Files.Create(ctx, nil, file))
require.NoError(t, repos.FileChunks.Create(ctx, nil, &database.FileChunk{
FileID: file.ID,
ChunkHash: types.ChunkHash(chunkHash),
}))
return file
}
// TestRestoreShortChunkHashInNoBlob proves a file whose short chunk hash
// has no blob_chunks row fails restore planning and the chunk write with
// an error.
func TestRestoreShortChunkHashInNoBlob(t *testing.T) {
t.Parallel()
ctx := context.Background()
_, repos := craftedSnapshotDB(t)
require.NoError(t, repos.Chunks.Create(ctx, nil,
&database.Chunk{ChunkHash: shortChunkHash, Size: 3}))
file := craftedFile(t, repos, shortChunkHash)
v := NewForTesting(nil)
chunkToBlobMap, err := v.buildChunkToBlobMap(ctx, repos)
require.NoError(t, err)
_, err = newRestorePlan(ctx, repos, []*database.File{file},
chunkToBlobMap, map[string]string{})
require.ErrorIs(t, err, errPlanChunkMissing)
fileChunks, err := repos.FileChunks.GetByFileID(ctx, file.ID)
require.NoError(t, err)
out, err := afero.NewMemMapFs().Create("out")
require.NoError(t, err)
session := &restoreSession{
v: v.Vaultik, ctx: ctx, chunkToBlobMap: chunkToBlobMap,
}
_, _, err = session.writeFileChunks(out, fileChunks)
require.ErrorIs(t, err, errChunkNotInAnyBlob)
}
// TestRestoreShortChunkHashReadPastBlobEnd proves a short chunk hash
// whose blob_chunks row reads past the end of its blob fails the chunk
// write with an error.
func TestRestoreShortChunkHashReadPastBlobEnd(t *testing.T) {
t.Parallel()
ctx := context.Background()
_, repos := craftedSnapshotDB(t)
blobHash := strings.Repeat("b", blobHashHexLen)
blob := &database.Blob{
ID: types.NewBlobID(),
Hash: types.BlobHash(blobHash),
CreatedTS: time.Now().UTC(),
}
require.NoError(t, repos.Blobs.Create(ctx, nil, blob))
require.NoError(t, repos.Chunks.Create(ctx, nil,
&database.Chunk{ChunkHash: shortChunkHash, Size: 3}))
require.NoError(t, repos.BlobChunks.Create(ctx, nil, &database.BlobChunk{
BlobID: blob.ID,
ChunkHash: shortChunkHash,
Length: 100,
}))
file := craftedFile(t, repos, shortChunkHash)
cache, err := newBlobDiskCache(1 << 20)
require.NoError(t, err)
t.Cleanup(func() { _ = cache.Close() })
require.NoError(t, cache.Put(blobHash, []byte("abc")))
v := NewForTesting(nil)
chunkToBlobMap, err := v.buildChunkToBlobMap(ctx, repos)
require.NoError(t, err)
_, blobIDToHash, err := v.buildBlobIndexes(repos)
require.NoError(t, err)
fileChunks, err := repos.FileChunks.GetByFileID(ctx, file.ID)
require.NoError(t, err)
out, err := afero.NewMemMapFs().Create("out")
require.NoError(t, err)
session := &restoreSession{
v: v.Vaultik,
ctx: ctx,
chunkToBlobMap: chunkToBlobMap,
blobIDToHash: blobIDToHash,
blobCache: cache,
}
_, _, err = session.writeFileChunks(out, fileChunks)
require.ErrorIs(t, err, errCacheReadBeyondBlob)
}
// TestVerifyFileMalformedChunkRow proves --verify returns an error for a
// chunk with no chunks row, a short chunk hash, and a chunk size from the
// database that is negative or larger than the restored file.
func TestVerifyFileMalformedChunkRow(t *testing.T) {
t.Parallel()
fullHash := types.ChunkHash(strings.Repeat("c", blobHashHexLen))
tests := []struct {
name string
hash types.ChunkHash
chunk *database.Chunk // nil adds no chunks row
want error
}{
{
name: "missing chunk row",
hash: fullHash,
want: errChunkRowMissing,
},
{
name: "short hash",
hash: shortChunkHash,
chunk: &database.Chunk{ChunkHash: shortChunkHash, Size: 3},
want: errChunkHashMismatch,
},
{
name: "size larger than the file",
hash: fullHash,
chunk: &database.Chunk{ChunkHash: fullHash, Size: math.MaxInt64},
want: errShortChunkRead,
},
{
name: "negative size",
hash: fullHash,
chunk: &database.Chunk{ChunkHash: fullHash, Size: -1},
want: errNegativeChunkLength,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
ctx := context.Background()
db, repos := craftedSnapshotDB(t)
if tt.chunk == nil {
// A crafted database need not satisfy its foreign keys.
_, err := db.Conn().ExecContext(ctx, "PRAGMA foreign_keys = OFF")
require.NoError(t, err)
} else {
require.NoError(t, repos.Chunks.Create(ctx, nil, tt.chunk))
}
file := craftedFile(t, repos, tt.hash.String())
v := NewForTesting(nil)
v.Fs = afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(v.Fs, "/restore/f",
[]byte(restoredFileContent), 0o600))
_, err := v.verifyFile(ctx, repos, file, "/restore/f")
require.ErrorIs(t, err, tt.want)
})
}
}
+349
View File
@@ -0,0 +1,349 @@
package vaultik //nolint:testpackage // drives unexported restore internals
import (
"context"
"os"
"path/filepath"
"syscall"
"testing"
"time"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/types"
)
// These tests check that restore applies each entry's owner, mode and
// mtime in an order that keeps them.
const (
readOnlyDirMode = uint32(os.ModeDir | 0o555)
unsearchableMode = uint32(os.ModeDir | 0o600)
plainDirMode = uint32(os.ModeDir | 0o755)
plainFileMode = uint32(0o644)
setuidFileMode = uint32(os.ModeSetuid | 0o755)
otherOwnerID = uint32(4321)
writableTestMode = 0o755
symlinkTargetPath = "/nonexistent/target"
memTargetDir = "/restore"
// Owner bits a normal user needs on a directory to create an entry
// in it, and to change an entry in it.
ownerWriteAndSearch = os.FileMode(0o300)
ownerSearch = os.FileMode(0o100)
)
// normalUserFs refuses what the kernel refuses a normal user. make test
// runs as root, which a read-only or unsearchable directory does not
// stop, so without it the tests below could not fail there. Creating an
// entry needs owner write and search on the directory holding it;
// changing an entry's mode or times needs owner search. Only that one
// directory is checked, not every ancestor.
type normalUserFs struct {
afero.Fs
}
//nolint:ireturn // afero.Fs.OpenFile is defined to return the interface
func (fs normalUserFs) OpenFile(
name string, flag int, perm os.FileMode,
) (afero.File, error) {
if flag&os.O_CREATE != 0 {
err := fs.checkParent(name, ownerWriteAndSearch)
if err != nil {
return nil, err
}
}
return fs.Fs.OpenFile(name, flag, perm)
}
func (fs normalUserFs) MkdirAll(path string, perm os.FileMode) error {
_, err := fs.Stat(path)
if err != nil {
err = fs.checkParent(path, ownerWriteAndSearch)
if err != nil {
return err
}
}
return fs.Fs.MkdirAll(path, perm)
}
func (fs normalUserFs) Chmod(name string, mode os.FileMode) error {
err := fs.checkParent(name, ownerSearch)
if err != nil {
return err
}
return fs.Fs.Chmod(name, mode)
}
func (fs normalUserFs) Chtimes(name string, atime, mtime time.Time) error {
err := fs.checkParent(name, ownerSearch)
if err != nil {
return err
}
return fs.Fs.Chtimes(name, atime, mtime)
}
// checkParent returns a permission error when the directory holding name
// exists and its owner bits lack any of need.
func (fs normalUserFs) checkParent(name string, need os.FileMode) error {
info, err := fs.Stat(filepath.Dir(name))
if err == nil && info.Mode().Perm()&need != need {
return &os.PathError{Op: "access", Path: name, Err: os.ErrPermission}
}
return nil
}
// TestRestoreFillsReadOnlyDirectory checks that a read-only directory
// still receives the entries inside it, and ends with its stored mode.
func TestRestoreFillsReadOnlyDirectory(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
ctx := context.Background()
rows, repos := makeFiles(ctx, t, []*database.File{
{Path: "/ro", Mode: readOnlyDirMode},
{Path: "/ro/file", Mode: plainFileMode},
{Path: "/ro/sub", Mode: readOnlyDirMode},
{Path: "/ro/sub/file", Mode: plainFileMode},
})
fs := afero.NewMemMapFs()
v := newContainmentVaultik(ctx, normalUserFs{Fs: fs})
_, err := v.restoreAllFiles(rows, repos,
&RestoreOptions{TargetDir: memTargetDir}, nil, nil)
require.NoError(t, err)
for _, path := range []string{"/ro/file", "/ro/sub/file"} {
_, err := fs.Stat(filepath.Join(memTargetDir, path))
require.NoErrorf(t, err, "file inside a read-only directory: %s", path)
}
for _, dir := range []string{"/ro", "/ro/sub"} {
info, err := fs.Stat(filepath.Join(memTargetDir, dir))
require.NoError(t, err)
assert.Equalf(t, os.FileMode(0o555), info.Mode().Perm(), "mode of %s", dir)
}
}
// TestRestoreKeepsNonEmptyDirectoryMTime checks that a directory keeps
// its stored mtime although entries were written into it. It runs on the
// real filesystem, where writing an entry changes its directory's mtime.
func TestRestoreKeepsNonEmptyDirectoryMTime(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
ctx := context.Background()
targetDir := t.TempDir()
mtime := time.Date(2001, time.February, 3, 4, 5, 6, 0, time.UTC)
rows, repos := makeFiles(ctx, t, []*database.File{
{Path: "/dir", Mode: plainDirMode, MTime: mtime},
{Path: "/dir/file", Mode: plainFileMode, MTime: mtime},
})
v := newContainmentVaultik(ctx, afero.NewOsFs())
_, err := v.restoreAllFiles(rows, repos,
&RestoreOptions{TargetDir: targetDir}, nil, nil)
require.NoError(t, err)
info, err := os.Stat(filepath.Join(targetDir, "dir"))
require.NoError(t, err)
assert.Truef(t, info.ModTime().Equal(mtime),
"directory mtime is %s, stored %s", info.ModTime(), mtime)
}
// TestRestoreFinishesChildBeforeUnsearchableParent checks that a
// directory inside one whose stored mode denies search still gets its
// own stored mode and mtime.
func TestRestoreFinishesChildBeforeUnsearchableParent(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
ctx := context.Background()
mtime := time.Date(2001, time.February, 3, 4, 5, 6, 0, time.UTC)
rows, repos := makeFiles(ctx, t, []*database.File{
{Path: "/locked", Mode: unsearchableMode, MTime: mtime},
{Path: "/locked/sub", Mode: plainDirMode, MTime: mtime},
})
fs := afero.NewMemMapFs()
v := newContainmentVaultik(ctx, normalUserFs{Fs: fs})
_, err := v.restoreAllFiles(rows, repos,
&RestoreOptions{TargetDir: memTargetDir}, nil, nil)
require.NoError(t, err)
info, err := fs.Stat(filepath.Join(memTargetDir, "locked"))
require.NoError(t, err)
assert.Equal(t, os.FileMode(0o600), info.Mode().Perm())
info, err = fs.Stat(filepath.Join(memTargetDir, "locked", "sub"))
require.NoError(t, err)
assert.Equal(t, os.FileMode(0o755), info.Mode().Perm())
assert.Truef(t, info.ModTime().Equal(mtime),
"mtime of sub is %s, stored %s", info.ModTime(), mtime)
}
// TestRestoreLeavesSymlinkedDirectoryTargetAlone checks that a directory
// whose place a later entry takes with a symlink does not hand its stored
// owner, mode and mtime to whatever the symlink points at. "/d" and "/d/"
// are different stored paths for the same place on disk.
func TestRestoreLeavesSymlinkedDirectoryTargetAlone(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
ctx := context.Background()
tempDir := t.TempDir()
targetDir := filepath.Join(tempDir, "target")
outsideDir := filepath.Join(tempDir, "outside")
require.NoError(t, os.Mkdir(outsideDir, writableTestMode))
before, err := os.Stat(outsideDir)
require.NoError(t, err)
mtime := time.Date(2001, time.February, 3, 4, 5, 6, 0, time.UTC)
rows, repos := makeFiles(ctx, t, []*database.File{
{
Path: "/d",
Mode: readOnlyDirMode,
UID: otherOwnerID,
GID: otherOwnerID,
MTime: mtime,
},
{Path: "/d/", LinkTarget: types.FilePath(outsideDir), MTime: mtime},
})
v := newContainmentVaultik(ctx, afero.NewOsFs())
_, err = v.restoreAllFiles(rows, repos,
&RestoreOptions{TargetDir: targetDir}, nil, nil)
require.NoError(t, err)
after, err := os.Stat(outsideDir)
require.NoError(t, err)
assert.Equal(t, before.Mode(), after.Mode())
assert.Truef(t, after.ModTime().Equal(before.ModTime()),
"mtime changed from %s to %s", before.ModTime(), after.ModTime())
beforeOwner, ok := before.Sys().(*syscall.Stat_t)
require.True(t, ok)
afterOwner, ok := after.Sys().(*syscall.Stat_t)
require.True(t, ok)
assert.Equal(t, beforeOwner.Uid, afterOwner.Uid)
assert.Equal(t, beforeOwner.Gid, afterOwner.Gid)
}
// TestRestoreKeepsSetuidThroughChown checks that a setuid file keeps the
// bit when restore changes its owner. Linux clears setuid on any chown of
// a regular file, even one to its current owner, so the file is recorded
// with the current user as owner and the session is told it runs as
// root: the chown then needs no privilege.
func TestRestoreKeepsSetuidThroughChown(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
ctx := context.Background()
targetDir := t.TempDir()
info, err := os.Stat(targetDir)
require.NoError(t, err)
owner, ok := info.Sys().(*syscall.Stat_t)
require.True(t, ok)
rows, repos := makeFiles(ctx, t, []*database.File{{
Path: "/suid",
Mode: setuidFileMode,
UID: owner.Uid,
GID: owner.Gid,
MTime: time.Date(2001, time.February, 3, 4, 5, 6, 0, time.UTC),
}})
session := &restoreSession{
v: newContainmentVaultik(ctx, afero.NewOsFs()),
ctx: ctx,
repos: repos,
opts: &RestoreOptions{TargetDir: targetDir},
result: &RestoreResult{},
runningAsRoot: true,
}
require.NoError(t, session.restoreFile(rows[0]))
info, err = os.Stat(filepath.Join(targetDir, "suid"))
require.NoError(t, err)
assert.Equal(t, os.FileMode(setuidFileMode),
info.Mode()&(os.ModeSetuid|os.ModePerm))
}
// TestRestoreSetsSymlinkMTime checks that a restored symlink gets its
// stored mtime on the link itself. The link dangles, so a call that
// follows it would fail.
func TestRestoreSetsSymlinkMTime(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
ctx := context.Background()
targetDir := t.TempDir()
mtime := time.Date(2001, time.February, 3, 4, 5, 6, 0, time.UTC)
rows, repos := makeFiles(ctx, t, []*database.File{{
Path: "/link",
LinkTarget: types.FilePath(symlinkTargetPath),
MTime: mtime,
}})
v := newContainmentVaultik(ctx, afero.NewOsFs())
_, err := v.restoreAllFiles(rows, repos,
&RestoreOptions{TargetDir: targetDir}, nil, nil)
require.NoError(t, err)
info, err := os.Lstat(filepath.Join(targetDir, "link"))
require.NoError(t, err)
assert.Truef(t, info.ModTime().Equal(mtime),
"symlink mtime is %s, stored %s", info.ModTime(), mtime)
}
// TestRestoreSetsSymlinkOwnerAsRoot checks that a symlink restored as
// root gets its stored owner on the link itself.
func TestRestoreSetsSymlinkOwnerAsRoot(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
if os.Geteuid() != 0 {
t.Skip("giving a file to another user needs root")
}
ctx := context.Background()
targetDir := t.TempDir()
rows, repos := makeFiles(ctx, t, []*database.File{{
Path: "/link",
LinkTarget: types.FilePath(symlinkTargetPath),
UID: otherOwnerID,
GID: otherOwnerID,
MTime: time.Date(2001, time.February, 3, 4, 5, 6, 0, time.UTC),
}})
v := newContainmentVaultik(ctx, afero.NewOsFs())
_, err := v.restoreAllFiles(rows, repos,
&RestoreOptions{TargetDir: targetDir}, nil, nil)
require.NoError(t, err)
info, err := os.Lstat(filepath.Join(targetDir, "link"))
require.NoError(t, err)
owner, ok := info.Sys().(*syscall.Stat_t)
require.True(t, ok)
assert.Equal(t, otherOwnerID, owner.Uid)
assert.Equal(t, otherOwnerID, owner.Gid)
}
+15 -1
View File
@@ -75,7 +75,7 @@ func newRestorePlan(
bc, ok := chunkToBlobMap[fc.ChunkHash.String()]
if !ok {
return nil, fmt.Errorf("planning %s: %w: %s",
f.Path, errPlanChunkMissing, fc.ChunkHash.String()[:16])
f.Path, errPlanChunkMissing, shortHash(fc.ChunkHash.String()))
}
hash, ok := blobIDToHash[bc.BlobID.String()]
@@ -214,6 +214,20 @@ func (p *restorePlan) blobsNeeded(fileID types.FileID) []string {
return out
}
// filesReferencingBlob returns the pending files that reference the
// named blob, in any order. The result is a copy, so the caller may
// finishFile each of them while ranging over it.
func (p *restorePlan) filesReferencingBlob(blobHash string) []types.FileID {
files := p.blobFiles[blobHash]
out := make([]types.FileID, 0, len(files))
for id := range files {
out = append(out, id)
}
return out
}
// hasPending reports whether any unfinished files remain.
func (p *restorePlan) hasPending() bool {
return len(p.fileBlobs) > 0
@@ -0,0 +1,194 @@
package vaultik_test
import (
"bytes"
"context"
"fmt"
"path/filepath"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/ui"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// A file no larger than the chunker's minimum chunk size (a quarter of
// faultChunkSize) is stored as one chunk, so it lives in exactly one blob.
// More of them than fit in faultMaxBlobSize make the snapshot span two
// blobs.
const (
missingBlobFileBytes = int(faultChunkSize / 4)
missingBlobFileCount = 20
)
// missingBlobBackup is a snapshot of single-chunk files spread over two
// blobs, with one of those blobs deleted from the store.
type missingBlobBackup struct {
fs afero.Fs
cfg *config.Config
storer storage.Storer
snapshotID string
restoreDir string
// files holds the original content by source path.
files map[string][]byte
// lost holds the source paths whose chunk is in the deleted blob.
lost map[string]bool
}
// TestRestoreSkipErrorsSkipsFilesOfMissingBlob restores with SkipErrors
// after one blob of a two-blob snapshot was deleted. Every file stored in
// that blob must be reported as failed and left absent, every other file
// must be restored intact, and Restore must still return an error.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestRestoreSkipErrorsSkipsFilesOfMissingBlob(t *testing.T) {
ctx := context.Background()
backup := backupThenDeleteOneBlob(ctx, t)
var out bytes.Buffer
v := newReaderVaultik(ctx, backup.cfg, backup.storer, nil, backup.fs)
v.UI = ui.NewWithColor(&out, false)
err := v.Restore(&vaultik.RestoreOptions{
SnapshotID: backup.snapshotID,
TargetDir: backup.restoreDir,
SkipErrors: true,
})
require.Error(t, err, "restore must fail when files were skipped")
assert.Contains(t, err.Error(),
fmt.Sprintf("%d file(s) failed to restore", len(backup.lost)))
for path, content := range backup.files {
restored := filepath.Join(backup.restoreDir, path)
if backup.lost[path] {
assert.Containsf(t, out.String(), path,
"%s needs the deleted blob and must be reported", path)
assert.NoFileExists(t, restored)
continue
}
got, err := afero.ReadFile(backup.fs, restored)
require.NoErrorf(t, err, "%s does not need the deleted blob", path)
assert.Equalf(t, content, got, "%s restored with wrong content", path)
}
}
// TestRestoreMissingBlobAbortsWithoutSkipErrors checks that a deleted blob
// still ends the restore with an error when SkipErrors is not set.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestRestoreMissingBlobAbortsWithoutSkipErrors(t *testing.T) {
ctx := context.Background()
backup := backupThenDeleteOneBlob(ctx, t)
v := newReaderVaultik(ctx, backup.cfg, backup.storer, nil, backup.fs)
err := v.Restore(&vaultik.RestoreOptions{
SnapshotID: backup.snapshotID,
TargetDir: backup.restoreDir,
})
require.ErrorIs(t, err, storage.ErrNotFound)
}
// backupThenDeleteOneBlob backs up missingBlobFileCount single-chunk
// files, then deletes from the store the blob holding the first of them.
func backupThenDeleteOneBlob(
ctx context.Context, t *testing.T,
) *missingBlobBackup {
t.Helper()
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
dbPath := filepath.Join(tempDir, "index.sqlite")
cfg := faultTestConfig()
require.NoError(t, fs.MkdirAll(dataDir, 0o755))
files := make(map[string][]byte, missingBlobFileCount)
for i := range missingBlobFileCount {
path := filepath.Join(dataDir, fmt.Sprintf("file-%02d.bin", i))
files[path] = bytesPattern(
fmt.Sprintf("file-%02d-", i), missingBlobFileBytes)
require.NoError(t, afero.WriteFile(fs, path, files[path], 0o644))
}
storer, err := storage.NewFileStorer(filepath.Join(tempDir, "remote"))
require.NoError(t, err)
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
repos := database.NewRepositories(db)
id := fullFaultBackup(
ctx, t, fs, storer, cfg, repos, dataDir, dbPath, "missingblob")
blobOfFile := make(map[string]string, len(files))
for path := range files {
blobOfFile[path] = blobHashOfSingleChunkFile(ctx, t, repos, path)
}
require.NoError(t, db.Close())
deleted := blobOfFile[filepath.Join(dataDir, "file-00.bin")]
require.NoError(t, storer.Delete(ctx, fmt.Sprintf(
"blobs/%s/%s/%s", deleted[:2], deleted[2:4], deleted)))
lost := make(map[string]bool)
for path, hash := range blobOfFile {
if hash == deleted {
lost[path] = true
}
}
require.Less(t, len(lost), len(files),
"the snapshot must span more than one blob")
return &missingBlobBackup{
fs: fs,
cfg: cfg,
storer: storer,
snapshotID: id,
restoreDir: filepath.Join(tempDir, "restored"),
files: files,
lost: lost,
}
}
// blobHashOfSingleChunkFile returns the hash of the blob holding the one
// chunk of the file at path, as recorded in the local index.
func blobHashOfSingleChunkFile(
ctx context.Context, t *testing.T,
repos *database.Repositories, path string,
) string {
t.Helper()
chunks, err := repos.FileChunks.GetByPath(ctx, path)
require.NoError(t, err)
require.Lenf(t, chunks, 1, "%s must be a single chunk", path)
blobChunk, err := repos.BlobChunks.GetByChunkHash(
ctx, chunks[0].ChunkHash.String())
require.NoError(t, err)
require.NotNilf(t, blobChunk, "chunk of %s is in no blob", path)
blob, err := repos.Blobs.GetByID(ctx, blobChunk.BlobID.String())
require.NoError(t, err)
return blob.Hash.String()
}
@@ -0,0 +1,73 @@
package vaultik_test
import (
"bytes"
"context"
"path/filepath"
"testing"
"time"
"github.com/spf13/afero"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// A file rewritten with its size unchanged and a new mtime in the same
// second as the mtime the index holds must still be backed up. See
// https://git.eeqj.de/sneak/vaultik/issues/226.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestBackupOfSameSecondRewriteRestoresNewContent(t *testing.T) {
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
storeDir := filepath.Join(tempDir, "remote")
restoreDir := filepath.Join(tempDir, "restored")
dbPath := filepath.Join(tempDir, "index.sqlite")
rewrittenPath := filepath.Join(dataDir, "small.txt")
ctx := context.Background()
files := writeFaultSourceTree(t, fs, dataDir)
cfg := changedFileConfig(dataDir, dbPath)
firstMTime := time.Date(2026, time.January, 2, 3, 4, 5, 0, time.UTC).
Add(100 * time.Millisecond)
secondMTime := firstMTime.Add(800 * time.Millisecond)
require.NoError(t, fs.Chtimes(rewrittenPath, firstMTime, firstMTime))
store, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
repos := database.NewRepositories(db)
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
require.NoError(t, backUp(v, "first"))
// Upper-casing ASCII text keeps its size.
files[rewrittenPath] = bytes.ToUpper(files[rewrittenPath])
require.NoError(t, afero.WriteFile(fs, rewrittenPath, files[rewrittenPath], 0o644))
require.NoError(t, fs.Chtimes(rewrittenPath, secondMTime, secondMTime))
require.NoError(t, backUp(v, "second"))
id := localSnapshotID(ctx, t, repos, "second")
require.NoError(t, db.Close())
reader := newReaderVaultik(ctx, cfg, store, nil, fs)
require.NoError(t, reader.Restore(&vaultik.RestoreOptions{
SnapshotID: id,
TargetDir: restoreDir,
Verify: true,
}))
assertRestoredTree(t, fs, restoreDir, files)
}
+101 -75
View File
@@ -14,6 +14,7 @@ import (
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/snapshot"
"sneak.berlin/go/vaultik/internal/types"
)
// Sentinel errors for snapshot management.
@@ -23,6 +24,9 @@ var (
errSnapshotVerifyFailed = errors.New("verification failed")
errRemoveAllNeedsForce = errors.New("--all requires --force")
errInvalidTableName = errors.New("invalid table name")
errNoAgeRecipients = errors.New(
"creating a snapshot needs at least one public key in " +
"age_recipients (generate a keypair with: age-keygen)")
)
// listRecentLimit caps how many snapshot rows are fetched from the
@@ -42,6 +46,12 @@ type SnapshotCreateOptions struct {
// CreateSnapshot executes the snapshot creation operation
func (v *Vaultik) CreateSnapshot(opts *SnapshotCreateOptions) error {
// config.Load accepts an empty list, since listing, verifying and
// restoring need no public key.
if len(v.Config.AgeRecipients) == 0 {
return errNoAgeRecipients
}
overallStartTime := time.Now()
log.Info("Starting snapshot creation",
@@ -180,6 +190,11 @@ type snapshotStats struct {
totalBytesUploaded int64
totalBlobsUploaded int
uploadDuration time.Duration
// The sizes of all blobs the snapshot references, set by
// finalizeSnapshotMetadata once snapshot_blobs is populated.
blobSize int64
blobUncompressedSize int64
}
// createNamedSnapshot creates a single named snapshot
@@ -219,8 +234,6 @@ func (v *Vaultik) createNamedSnapshot(
return err
}
v.collectUploadStats(scanner, stats)
err = v.finalizeSnapshotMetadata(snapshotID, stats)
if err != nil {
return err
@@ -272,6 +285,11 @@ func (v *Vaultik) resolveSnapshotPaths(snapName string) ([]string, error) {
func (v *Vaultik) scanAllDirectories(
scanner *snapshot.Scanner, resolvedDirs []string, snapshotID string,
) (*snapshotStats, error) {
if progress := scanner.GetProgress(); progress != nil {
progress.Start()
defer progress.Stop()
}
stats := &snapshotStats{}
for i, dir := range resolvedDirs {
@@ -300,6 +318,9 @@ func (v *Vaultik) scanAllDirectories(
stats.totalBytesSkipped += result.BytesSkipped
stats.totalFilesDeleted += result.FilesDeleted
stats.totalBytesDeleted += result.BytesDeleted
stats.totalBlobsUploaded += result.BlobsUploaded
stats.totalBytesUploaded += result.BytesUploaded
stats.uploadDuration += result.UploadDuration
log.Info("Directory scan complete",
"path", dir,
@@ -315,18 +336,6 @@ func (v *Vaultik) scanAllDirectories(
return stats, nil
}
// collectUploadStats gathers upload statistics from the scanner's
// progress reporter.
func (v *Vaultik) collectUploadStats(scanner *snapshot.Scanner, stats *snapshotStats) {
if s := scanner.GetProgress(); s != nil {
progressStats := s.GetStats()
stats.totalBytesUploaded = progressStats.BytesUploaded.Load()
stats.totalBlobsUploaded = int(progressStats.BlobsUploaded.Load())
stats.uploadDuration = time.Duration(
progressStats.UploadDurationMs.Load()) * time.Millisecond
}
}
// finalizeSnapshotMetadata updates stats, exports metadata, and only then
// marks the snapshot complete. Recording completion last is deliberate: an
// export interrupted by a crash leaves the snapshot incomplete rather than
@@ -336,31 +345,39 @@ func (v *Vaultik) collectUploadStats(scanner *snapshot.Scanner, stats *snapshotS
func (v *Vaultik) finalizeSnapshotMetadata(
snapshotID string, stats *snapshotStats,
) error {
// snapshot_blobs must be populated before the blob sizes below, which
// total the snapshot's blobs, and before the export, which builds the
// manifest and the trimmed metadata database from it.
err := v.SnapshotManager.PopulateSnapshotBlobs(v.ctx, snapshotID)
if err != nil {
return fmt.Errorf("populating snapshot blobs: %w", err)
}
stats.blobSize, stats.blobUncompressedSize, err =
v.Repositories.Snapshots.GetSnapshotBlobSizes(v.ctx, snapshotID)
if err != nil {
return fmt.Errorf("getting snapshot blob sizes: %w", err)
}
extStats := snapshot.ExtendedBackupStats{
BackupStats: snapshot.BackupStats{
FilesScanned: stats.totalFiles,
BytesScanned: stats.totalBytes,
TotalSize: stats.totalBytes + stats.totalBytesSkipped,
ChunksCreated: stats.totalChunks,
BlobsCreated: stats.totalBlobs,
BytesUploaded: stats.totalBytesUploaded,
},
BlobUncompressedSize: 0,
BlobSize: stats.blobSize,
BlobUncompressedSize: stats.blobUncompressedSize,
CompressionLevel: v.Config.CompressionLevel,
UploadDurationMs: stats.uploadDuration.Milliseconds(),
}
err := v.SnapshotManager.UpdateSnapshotStatsExtended(v.ctx, snapshotID, extStats)
err = v.SnapshotManager.UpdateSnapshotStatsExtended(v.ctx, snapshotID, extStats)
if err != nil {
return fmt.Errorf("updating snapshot stats: %w", err)
}
// snapshot_blobs must be populated before the export, which builds the
// manifest and the trimmed metadata database from it.
err = v.SnapshotManager.PopulateSnapshotBlobs(v.ctx, snapshotID)
if err != nil {
return fmt.Errorf("populating snapshot blobs: %w", err)
}
err = v.SnapshotManager.ExportSnapshotMetadata(
v.ctx, v.Config.IndexPath, snapshotID)
if err != nil {
@@ -395,12 +412,10 @@ func (v *Vaultik) printSnapshotSummary(
totalFilesChanged := stats.totalFiles - stats.totalFilesSkipped
totalBytesAll := stats.totalBytes + stats.totalBytesSkipped
// Get total blob sizes from database
compressedSize, uncompressedSize := v.getSnapshotBlobSizes(snapshotID)
var compressionRatio float64
if uncompressedSize > 0 {
compressionRatio = float64(compressedSize) / float64(uncompressedSize)
if stats.blobUncompressedSize > 0 {
compressionRatio = float64(stats.blobSize) /
float64(stats.blobUncompressedSize)
} else {
compressionRatio = 1.0
}
@@ -428,8 +443,8 @@ func (v *Vaultik) printSnapshotSummary(
if stats.totalBlobsUploaded > 0 {
v.UI.Detailf("Storage: %s compressed from %s (%.2fx ratio).",
v.UI.Size(compressedSize),
v.UI.Size(uncompressedSize),
v.UI.Size(stats.blobSize),
v.UI.Size(stats.blobUncompressedSize),
compressionRatio)
v.UI.Detailf("Upload: %d blobs, %s in %s (%s).",
stats.totalBlobsUploaded,
@@ -441,27 +456,6 @@ func (v *Vaultik) printSnapshotSummary(
v.UI.Detailf("Snapshot create duration: %s.", v.UI.Duration(snapshotDuration))
}
// getSnapshotBlobSizes returns total compressed and uncompressed blob
// sizes for a snapshot.
func (v *Vaultik) getSnapshotBlobSizes(snapshotID string) (int64, int64) {
var compressed, uncompressed int64
blobHashes, err := v.Repositories.Snapshots.GetBlobHashes(v.ctx, snapshotID)
if err != nil {
return 0, 0
}
for _, hash := range blobHashes {
blob, err := v.Repositories.Blobs.GetByHash(v.ctx, hash)
if err == nil && blob != nil {
compressed += blob.CompressedSize
uncompressed += blob.UncompressedSize
}
}
return compressed, uncompressed
}
// SnapshotPurgeOptions contains options for the snapshot purge command.
type SnapshotPurgeOptions struct {
KeepLatest bool // Keep only the most recent snapshot per name
@@ -502,19 +496,23 @@ func (v *Vaultik) PurgeSnapshotsWithOptions(opts *SnapshotPurgeOptions) error {
nameFilter[n] = struct{}{}
}
// Collect completed snapshots, applying the name filter.
// Collect completed snapshots and their names, applying the name filter.
snapshots := make([]SnapshotInfo, 0, len(dbSnapshots))
names := make(map[types.SnapshotID]string, len(dbSnapshots))
for _, s := range dbSnapshots {
if s.CompletedAt == nil {
continue
}
name := parseSnapshotName(s.ID.String(), s.Hostname.String())
if len(nameFilter) > 0 {
if _, ok := nameFilter[parseSnapshotName(s.ID.String())]; !ok {
if _, ok := nameFilter[name]; !ok {
continue
}
}
names[s.ID] = name
snapshots = append(snapshots, SnapshotInfo{
ID: s.ID,
Timestamp: s.StartedAt,
@@ -527,7 +525,7 @@ func (v *Vaultik) PurgeSnapshotsWithOptions(opts *SnapshotPurgeOptions) error {
return snapshots[i].Timestamp.After(snapshots[j].Timestamp)
})
toDelete, err := selectSnapshotsToPurge(snapshots, opts)
toDelete, err := selectSnapshotsToPurge(snapshots, names, opts)
if err != nil {
return err
}
@@ -545,9 +543,11 @@ func (v *Vaultik) PurgeSnapshotsWithOptions(opts *SnapshotPurgeOptions) error {
// selectSnapshotsToPurge applies the purge retention criteria to the
// newest-first sorted snapshot list and returns the deletion
// candidates.
// candidates. names maps each snapshot's ID to its snapshot name.
func selectSnapshotsToPurge(
snapshots []SnapshotInfo, opts *SnapshotPurgeOptions,
snapshots []SnapshotInfo,
names map[types.SnapshotID]string,
opts *SnapshotPurgeOptions,
) ([]SnapshotInfo, error) {
var toDelete []SnapshotInfo
@@ -558,7 +558,7 @@ func selectSnapshotsToPurge(
seen := make(map[string]bool)
for _, snap := range snapshots {
name := parseSnapshotName(snap.ID.String())
name := names[snap.ID]
if seen[name] {
toDelete = append(toDelete, snap)
@@ -718,7 +718,7 @@ func (v *Vaultik) VerifySnapshotWithOptions(
result.BlobCount = manifest.BlobCount
result.TotalSize = manifest.TotalCompressedSize
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf("Snapshot information:\n")
v.stdoutf(" Blob count: %d\n", manifest.BlobCount)
v.stdoutf(" Total size: %s\n", ubytes(manifest.TotalCompressedSize))
@@ -773,7 +773,7 @@ func (v *Vaultik) printVerifyHeader(snapshotID string, opts *VerifyOptions) {
snapshotTime = t
}
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf("Verifying snapshot %s\n", snapshotID)
if !snapshotTime.IsZero() {
@@ -813,7 +813,7 @@ func (v *Vaultik) verifyManifestBlobs(
stat, err := v.Storage.Stat(v.ctx, blobPath)
switch {
case err != nil:
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf(" Missing: %s (%s)\n",
blob.Hash, ubytes(blob.CompressedSize))
}
@@ -821,7 +821,7 @@ func (v *Vaultik) verifyManifestBlobs(
missing++
missingSize += blob.CompressedSize
case stat.Size != blob.CompressedSize:
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf(" Wrong size: %s (store has %s, manifest lists %s)\n",
blob.Hash, ubytes(stat.Size), ubytes(blob.CompressedSize))
}
@@ -853,6 +853,22 @@ func (v *Vaultik) formatVerifyResult(
return v.outputVerifyJSON(result)
}
// Under --quiet a failure is still returned, and the cli layer
// prints it on stderr.
if !v.UI.Quiet() {
v.printVerifySummary(result, failure)
}
if failure != "" {
return fmt.Errorf("%w: %s", errSnapshotVerifyFailed, failure)
}
return nil
}
// printVerifySummary prints the counts and the status line that end the
// human-readable shallow verify report. failure is empty when it passed.
func (v *Vaultik) printVerifySummary(result *VerifyResult, failure string) {
v.stdoutf("\nVerification complete:\n")
v.stdoutf(" Present with listed size: %d blobs\n", result.Verified)
@@ -874,14 +890,12 @@ func (v *Vaultik) formatVerifyResult(
if failure != "" {
v.stdoutf("FAILED - %s\n", failure)
return fmt.Errorf("%w: %s", errSnapshotVerifyFailed, failure)
return
}
// Report only what was actually checked: presence and size, not contents.
v.stdoutf("OK - all %d blobs listed in the manifest are present with the "+
"listed size; contents not checked (use --deep)\n", result.Verified)
return nil
}
// shallowVerifyFailure returns a human-readable description of everything
@@ -1038,7 +1052,7 @@ func (v *Vaultik) syncWithRemote() error {
// every local snapshot record (issue #160).
remoteKeys, err := v.listAllRemoteSnapshotKeys()
if err != nil {
return fmt.Errorf("listing remote snapshots: %w", err)
return err
}
remoteKeySet := make(map[string]bool, len(remoteKeys))
@@ -1103,6 +1117,12 @@ type RemoveResult struct {
// just-removed snapshot left behind on the destination store.
const pruneCommandHint = "vaultik prune"
// snapshotRemoveCommandHint is the command suggested, with the
// snapshot's ID, when a remove could not reach the destination store:
// running it again removes the snapshot's metadata there, which
// `vaultik prune` never does.
const snapshotRemoveCommandHint = "vaultik snapshot remove"
// RemoveSnapshot removes a snapshot from the local index database and,
// unless LocalOnly is set, also strips the snapshot's metadata from the
// destination store. Blobs are NOT touched: removing a snapshot's
@@ -1139,7 +1159,7 @@ func (v *Vaultik) RemoveSnapshot(
}
if !opts.LocalOnly {
result.RemoteRemoved = v.removeSnapshotRemote(snapshotID)
result.RemoteRemoved = v.removeSnapshotRemote(snapshotID, opts)
}
if v.SnapshotManager != nil {
@@ -1222,9 +1242,11 @@ func (v *Vaultik) confirmRemoveSnapshot(snapshotID string, opts *RemoveOptions)
// removeSnapshotRemote strips the snapshot's metadata from the
// destination store, warning and proceeding on failure: the local-DB
// removal has already happened, so the user is told the remote half
// didn't finish and can retry with `vaultik prune` once the destination
// store is reachable. Returns true when the remote removal succeeded.
func (v *Vaultik) removeSnapshotRemote(snapshotID string) bool {
// didn't finish and to run `vaultik snapshot remove` for the snapshot
// again once the destination store is reachable (`vaultik prune` never
// removes snapshot metadata). Returns true when the remote removal
// succeeded.
func (v *Vaultik) removeSnapshotRemote(snapshotID string, opts *RemoveOptions) bool {
log.Info("Removing snapshot metadata from remote storage",
"snapshot_id", snapshotID)
@@ -1232,13 +1254,17 @@ func (v *Vaultik) removeSnapshotRemote(snapshotID string) bool {
err := v.deleteRemoteSnapshotByKey(remoteKey)
if err != nil {
log.Warn("Could not remove snapshot metadata from remote storage",
"error", err)
log.Warn("Could not remove snapshot metadata from remote storage; "+
"run '"+snapshotRemoveCommandHint+"' with the snapshot's ID "+
"again once the remote is reachable",
"snapshot_id", snapshotID, "error", err)
if v.UI != nil {
// The UI writes to stdout, which under --json holds only the
// document; the log record above is the warning on stderr.
if v.UI != nil && !opts.JSON {
v.UI.Warningf("Could not remove snapshot metadata from remote: "+
"%v. Run '%s' once the remote is reachable to finish cleanup.",
err, pruneCommandHint)
"%v. Run '%s %s' again once the remote is reachable.",
err, snapshotRemoveCommandHint, snapshotID)
}
return false
+285
View File
@@ -0,0 +1,285 @@
package vaultik_test
import (
"bytes"
"context"
"fmt"
"os"
"path/filepath"
"strings"
"testing"
"time"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/storage/faultstore"
"sneak.berlin/go/vaultik/internal/ui"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// These tests cover https://git.eeqj.de/sneak/vaultik/issues/225: the
// summary printed after a backup, and the statistics stored in the
// snapshots table, count each file, byte and upload once, and a --cron
// run records its uploads.
// summaryUploadDelay slows every blob upload, so a run's upload time is
// at least this long per blob even on a local store.
const summaryUploadDelay = 20 * time.Millisecond
// summaryEnv is a backup setup whose user-facing output is kept in out.
type summaryEnv struct {
v *vaultik.Vaultik
db *database.DB
repos *database.Repositories
out *bytes.Buffer
// aPath is a.bin, whose content copy.bin repeats; aSize is its size
// and totalSize the size of all three source files.
aPath string
aSize int64
totalSize int64
}
// newSummaryEnv writes src/one/a.bin, src/one/small.txt and
// src/two/copy.bin, a copy of a.bin. Every chunk of copy.bin is therefore
// already stored by the time the backup reaches it.
//
// The snapshot names "first" and "second" back up src; "split" backs up
// src/one and src/two as two paths.
func newSummaryEnv(t *testing.T) *summaryEnv {
t.Helper()
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
srcDir := filepath.Join(tempDir, "src")
dirOne := filepath.Join(srcDir, "one")
dirTwo := filepath.Join(srcDir, "two")
dbPath := filepath.Join(tempDir, "index.sqlite")
ctx := context.Background()
aContent := bytesPattern("a-", int(3*faultChunkSize))
smallContent := []byte("hello vaultik")
files := map[string][]byte{
filepath.Join(dirOne, "a.bin"): aContent,
filepath.Join(dirOne, "small.txt"): smallContent,
filepath.Join(dirTwo, "copy.bin"): aContent,
}
for path, content := range files {
require.NoError(t, fs.MkdirAll(filepath.Dir(path), 0o755))
require.NoError(t, afero.WriteFile(fs, path, content, 0o644))
}
cfg := faultTestConfig()
cfg.IndexPath = dbPath
cfg.ChunkSize = config.Size(faultChunkSize)
cfg.Snapshots = map[string]config.SnapshotConfig{
"first": {Paths: []string{srcDir}},
"second": {Paths: []string{srcDir}},
"split": {Paths: []string{dirOne, dirTwo}},
}
inner, err := storage.NewFileStorer(filepath.Join(tempDir, "remote"))
require.NoError(t, err)
store := faultstore.New(inner)
store.OnPut = func(key string) faultstore.PutAction {
if strings.HasPrefix(key, "blobs/") {
time.Sleep(summaryUploadDelay)
}
return faultstore.PutNormal
}
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
t.Cleanup(func() { _ = db.Close() })
repos := database.NewRepositories(db)
out := &bytes.Buffer{}
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
v.UI = ui.NewWithColor(out, false)
return &summaryEnv{
v: v,
db: db,
repos: repos,
out: out,
aPath: filepath.Join(dirOne, "a.bin"),
aSize: int64(len(aContent)),
totalSize: int64(2*len(aContent) + len(smallContent)),
}
}
// backUp runs a backup of the named snapshot and returns its output.
func (e *summaryEnv) backUp(t *testing.T, name string, cron bool) string {
t.Helper()
e.out.Reset()
require.NoError(t, e.v.CreateSnapshot(&vaultik.SnapshotCreateOptions{
Cron: cron,
Snapshots: []string{name},
}))
return e.out.String()
}
// snapshot returns the local snapshots row of the snapshot named name.
func (e *summaryEnv) snapshot(t *testing.T, name string) *database.Snapshot {
t.Helper()
ctx := context.Background()
snap, err := e.repos.Snapshots.GetByID(ctx,
localSnapshotID(ctx, t, e.repos, name))
require.NoError(t, err)
require.NotNil(t, snap)
return snap
}
// uploads returns how many blobs the snapshot uploaded and their
// total size, as recorded in the uploads table.
func (e *summaryEnv) uploads(t *testing.T, snapshotID string) (int64, int64) {
t.Helper()
var count, size int64
err := e.db.Conn().QueryRowContext(context.Background(), `
SELECT COUNT(*), COALESCE(SUM(size), 0)
FROM uploads WHERE snapshot_id = ?`, snapshotID).Scan(&count, &size)
require.NoError(t, err)
return count, size
}
// referencedBlobSizes returns the compressed and uncompressed sizes of
// all blobs the snapshot references.
func (e *summaryEnv) referencedBlobSizes(
t *testing.T, snapshotID string,
) (int64, int64) {
t.Helper()
var compressed, uncompressed int64
err := e.db.Conn().QueryRowContext(context.Background(), `
SELECT COALESCE(SUM(b.compressed_size), 0),
COALESCE(SUM(b.uncompressed_size), 0)
FROM snapshot_blobs sb JOIN blobs b ON b.blob_hash = sb.blob_hash
WHERE sb.snapshot_id = ?`, snapshotID).Scan(&compressed, &uncompressed)
require.NoError(t, err)
return compressed, uncompressed
}
// filesLine returns the summary's line of file counts.
func filesLine(examined, backedUp, unchanged int) string {
return fmt.Sprintf("Files: %d examined, %d backed up, %d unchanged.",
examined, backedUp, unchanged)
}
// dataLine returns the summary's line of byte counts.
func (e *summaryEnv) dataLine(total, backedUp int64) string {
return fmt.Sprintf("Data: %s total (%s backed up).",
e.v.UI.Size(total), e.v.UI.Size(backedUp))
}
// A first backup stores copy.bin's chunks while backing up a.bin, so
// copy.bin's chunks are deduplicated within the run. Each file and byte
// is still counted once.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestSnapshotSummaryFirstRun(t *testing.T) {
env := newSummaryEnv(t)
summary := env.backUp(t, "first", false)
assert.Contains(t, summary, filesLine(3, 3, 0))
assert.Contains(t, summary, env.dataLine(env.totalSize, env.totalSize))
snap := env.snapshot(t, "first")
uploadCount, uploadBytes := env.uploads(t, snap.ID.String())
require.Positive(t, uploadCount)
assert.Contains(t, summary, fmt.Sprintf("Upload: %d blobs, %s in ",
uploadCount, env.v.UI.Size(uploadBytes)))
assert.Equal(t, int64(3), snap.FileCount)
assert.Equal(t, env.totalSize, snap.TotalSize)
assert.Equal(t, uploadCount, snap.BlobCount)
assert.Equal(t, uploadBytes, snap.UploadBytes)
}
// An incremental backup where a.bin's mtime changed but its content did
// not: a.bin is backed up again and every one of its chunks is already
// stored.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestSnapshotSummaryIncrementalRunWithDeduplicatedChunks(t *testing.T) {
env := newSummaryEnv(t)
env.backUp(t, "first", false)
later := time.Now().Add(time.Hour)
require.NoError(t, os.Chtimes(env.aPath, later, later))
summary := env.backUp(t, "second", false)
assert.Contains(t, summary, filesLine(3, 1, 2))
assert.Contains(t, summary, env.dataLine(env.totalSize, env.aSize))
assert.NotContains(t, summary, "Upload:")
snap := env.snapshot(t, "second")
compressed, uncompressed := env.referencedBlobSizes(t, snap.ID.String())
require.Positive(t, compressed)
assert.Equal(t, env.totalSize, snap.TotalSize)
assert.Zero(t, snap.ChunkCount)
assert.Zero(t, snap.BlobCount)
assert.Zero(t, snap.UploadBytes)
assert.Equal(t, compressed, snap.BlobSize,
"blob_size must total the blobs the snapshot references")
assert.Equal(t, uncompressed, snap.BlobUncompressedSize)
}
// Under --cron the progress reporter is off; the upload figures must
// still reach the summary and the snapshots row. The snapshot has two
// paths, each backed up by its own scan.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestSnapshotSummaryCronRunRecordsUploads(t *testing.T) {
env := newSummaryEnv(t)
summary := env.backUp(t, "split", true)
snap := env.snapshot(t, "split")
uploadCount, uploadBytes := env.uploads(t, snap.ID.String())
require.Positive(t, uploadCount)
assert.Contains(t, summary, filesLine(3, 3, 0))
assert.Contains(t, summary, env.dataLine(env.totalSize, env.totalSize))
assert.Contains(t, summary, fmt.Sprintf("Upload: %d blobs, %s in ",
uploadCount, env.v.UI.Size(uploadBytes)))
assert.Equal(t, env.totalSize, snap.TotalSize)
assert.Equal(t, uploadCount, snap.BlobCount,
"blob_count must count each blob once, however many paths the "+
"snapshot has")
assert.Equal(t, uploadBytes, snap.UploadBytes)
assert.GreaterOrEqual(t, snap.UploadDurationMs,
uploadCount*summaryUploadDelay.Milliseconds())
compressed, uncompressed := env.referencedBlobSizes(t, snap.ID.String())
require.Positive(t, uncompressed)
assert.Equal(t, compressed, snap.BlobSize)
assert.Equal(t, uncompressed, snap.BlobUncompressedSize)
assert.InDelta(t, float64(compressed)/float64(uncompressed),
snap.CompressionRatio, 1e-9)
}
@@ -0,0 +1,71 @@
package vaultik_test
import (
"context"
"maps"
"path/filepath"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// A backup without --cron runs the progress reporter while one scanner
// scans each path of the snapshot in turn. See
// https://git.eeqj.de/sneak/vaultik/issues/253.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestBackupWithoutCronOfTwoPathSnapshotRestoresBothPaths(t *testing.T) {
log.Initialize(log.Config{})
const snapshotName = "data"
fs := afero.NewOsFs()
tempDir := t.TempDir()
firstDir := filepath.Join(tempDir, "first")
secondDir := filepath.Join(tempDir, "second")
storeDir := filepath.Join(tempDir, "remote")
restoreDir := filepath.Join(tempDir, "restored")
dbPath := filepath.Join(tempDir, "index.sqlite")
ctx := context.Background()
files := writeFaultSourceTree(t, fs, firstDir)
maps.Copy(files, writeFaultSourceTree(t, fs, secondDir))
cfg := faultTestConfig()
cfg.IndexPath = dbPath
cfg.ChunkSize = config.Size(faultChunkSize)
cfg.Snapshots = map[string]config.SnapshotConfig{
snapshotName: {Paths: []string{firstDir, secondDir}},
}
store, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
repos := database.NewRepositories(db)
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
require.NoError(t, v.CreateSnapshot(&vaultik.SnapshotCreateOptions{
Snapshots: []string{snapshotName},
}))
id := localSnapshotID(ctx, t, repos, snapshotName)
require.NoError(t, db.Close())
reader := newReaderVaultik(ctx, cfg, store, nil, fs)
require.NoError(t, reader.Restore(&vaultik.RestoreOptions{
SnapshotID: id,
TargetDir: restoreDir,
Verify: true,
}))
assertRestoredTree(t, fs, restoreDir, files)
}
+15 -12
View File
@@ -103,7 +103,7 @@ func (v *Vaultik) RunDeepVerify(snapshotID string, opts *VerifyOptions) error {
log.Info("Starting snapshot verification", "snapshot_id", snapshotID, "mode", "deep")
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf("Deep verification of snapshot: %s\n\n", snapshotID)
}
@@ -143,10 +143,13 @@ func (v *Vaultik) RunDeepVerify(snapshotID string, opts *VerifyOptions) error {
log.Info("✓ Verification completed successfully",
"snapshot_id", snapshotID, "mode", "deep", "blobs_verified", len(dbBlobs))
v.stdoutf("\n✓ Verification completed successfully\n")
v.stdoutf(" Snapshot: %s\n", snapshotID)
v.stdoutf(" Blobs verified: %d\n", len(dbBlobs))
v.stdoutf(" Total size: %s\n", ubytes(totalSize))
if !v.UI.Quiet() {
v.stdoutf("\n✓ Verification completed successfully\n")
v.stdoutf(" Snapshot: %s\n", snapshotID)
v.stdoutf(" Blobs verified: %d\n", len(dbBlobs))
v.stdoutf(" Total size: %s\n", ubytes(totalSize))
}
return nil
}
@@ -170,7 +173,7 @@ func (v *Vaultik) loadVerificationData(
// remote manifests; see its doc comment.
log.Info("Downloading manifest", "remote_key", remoteKey)
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf("Downloading manifest...\n")
}
@@ -185,7 +188,7 @@ func (v *Vaultik) loadVerificationData(
"manifest_blob_count", manifest.BlobCount,
"manifest_total_size", ubytes(manifest.TotalCompressedSize))
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf("Manifest loaded: %d blobs (%s)\n",
manifest.BlobCount, ubytes(manifest.TotalCompressedSize))
v.stdoutf("Downloading and decrypting database...\n")
@@ -215,7 +218,7 @@ func (v *Vaultik) loadVerificationData(
"db_blob_count", len(dbBlobs),
"db_total_size", ubytes(dbTotalSize))
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf("Database loaded: %d blobs (%s)\n",
len(dbBlobs), ubytes(dbTotalSize))
}
@@ -273,7 +276,7 @@ func (v *Vaultik) runVerificationSteps(
totalSize int64,
identities []age.Identity,
) error {
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf("Verifying manifest against database...\n")
}
@@ -282,7 +285,7 @@ func (v *Vaultik) runVerificationSteps(
return v.deepVerifyFailure(result, opts, err.Error(), err)
}
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf("Manifest verified.\n")
v.stdoutf("Checking blob existence in remote storage...\n")
}
@@ -292,7 +295,7 @@ func (v *Vaultik) runVerificationSteps(
return v.deepVerifyFailure(result, opts, err.Error(), err)
}
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf("All blobs exist.\n")
v.stdoutf("Downloading and verifying blob contents (%d blobs, %s)...\n",
len(dbBlobs), ubytes(totalSize))
@@ -748,7 +751,7 @@ func (v *Vaultik) performDeepVerificationFromDB(
"eta", eta.Round(time.Second),
)
if !opts.JSON {
if !opts.JSON && !v.UI.Quiet() {
v.stdoutf(" Verified %d/%d blobs (%d remaining) - %s/%s - elapsed %s, eta %s\n",
i+1, len(blobs), remaining,
ubytes(bytesProcessed),
+102
View File
@@ -0,0 +1,102 @@
package vaultik_test
import (
"bytes"
"context"
"io"
"os"
"path/filepath"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/snapshot"
"sneak.berlin/go/vaultik/internal/ui"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// TestVerify_QuietSuppressesReport is the --quiet contract for
// `snapshot verify`: neither shallow nor deep verify writes its report,
// a failed verify still returns its error (which the cli layer prints
// on stderr), and the --json document still emits.
func TestVerify_QuietSuppressesReport(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "source")
storeDir := filepath.Join(tempDir, "remote")
dbPath := filepath.Join(tempDir, "index.sqlite")
chunkSize := int64(32 * 1024)
maxBlobSize := int64(128 * 1024)
require.NoError(t, fs.MkdirAll(dataDir, 0o755))
require.NoError(t, afero.WriteFile(fs,
filepath.Join(dataDir, "data.bin"),
bytesPattern("quiet-", int(maxBlobSize*2)), 0o644))
ctx := context.Background()
cfg, storer, snapshotID := runFileStorageBackup(
ctx, t, fs, dataDir, storeDir, dbPath, chunkSize, maxBlobSize)
// The UI writes to the same buffer as Stdout, as both write to the
// process's stdout in production.
var stdout bytes.Buffer
newQuietVerifier := func() *vaultik.Vaultik {
v := &vaultik.Vaultik{
Config: cfg,
Storage: storer,
Fs: fs,
Stdout: &stdout,
Stderr: io.Discard,
UI: ui.NewWithColor(&stdout, false),
}
v.SetContext(ctx)
v.UI.SetQuiet(true)
return v
}
require.NoError(t, newQuietVerifier().VerifySnapshotWithOptions(
snapshotID, &vaultik.VerifyOptions{}))
require.Empty(t, stdout.String(),
"shallow verify must write no report under --quiet")
require.NoError(t, newQuietVerifier().VerifySnapshotWithOptions(
snapshotID, &vaultik.VerifyOptions{Deep: true}))
require.Empty(t, stdout.String(),
"deep verify must write no report under --quiet")
require.NoError(t, newQuietVerifier().VerifySnapshotWithOptions(
snapshotID, &vaultik.VerifyOptions{JSON: true}))
require.Equal(t, "ok", decodeVerifyResult(t, stdout.Bytes()).Status,
"the --json document must still emit under --quiet")
// A snapshot without its encrypted database fails shallow verify. A
// failed report also lists each missing blob and each blob of the
// wrong size, so remove one blob and grow another.
require.NoError(t, os.Remove(filepath.Join(storeDir, "metadata",
snapshot.RemoteSnapshotKey(snapshotID), "db.zst.age")))
blobFiles, err := filepath.Glob(
filepath.Join(storeDir, "blobs", "*", "*", "*"))
require.NoError(t, err)
require.GreaterOrEqual(t, len(blobFiles), 2,
"the snapshot must span two blobs, one to remove and one to grow")
require.NoError(t, os.Remove(blobFiles[0]))
growOneBlob(t, fs, filepath.Join(storeDir, "blobs"))
stdout.Reset()
require.Error(t, newQuietVerifier().VerifySnapshotWithOptions(
snapshotID, &vaultik.VerifyOptions{}),
"--quiet must not change the outcome of a failed verify")
require.Empty(t, stdout.String(),
"a failed verify must write no report under --quiet")
}
+9 -8
View File
@@ -38,12 +38,7 @@ pkg_install() {
detect_pkgmgr
case "$PKGMGR" in
nix) nix-env -iA "nixpkgs.$1" ;;
# A fresh image, such as the CI runner's, has no package lists,
# so apt-get install finds nothing until they are fetched.
apt)
$SUDO apt-get update
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2"
;;
apt) $SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2" ;;
brew) brew install "$3" ;;
apk) apk add --no-cache "$4" ;;
esac
@@ -109,8 +104,14 @@ main() {
if missing git; then pkg_install git git git git; fi
if missing make; then pkg_install gnumake make make make; fi
# Go toolchain
if missing go; then pkg_install go golang go go; fi
# Go toolchain: the host's own, or else the version go.mod names,
# hash-verified, in .tool/go. That directory is not on the caller's
# PATH; this script, the Makefile, script/fmt, script/fmt-check,
# script/precommit and script/release add it to theirs.
if missing go; then
"$ROOT/script/install-go"
PATH="$PATH:$ROOT/.tool/go/bin"
fi
# golangci-lint is deliberately NOT installed: script/lint lints by
# building the lint phase of the Dockerfile, whose digest-pinned
+2
View File
@@ -6,6 +6,8 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
# Where script/bootstrap installs Go when the host has none.
PATH="$PATH:$ROOT/.tool/go/bin"
go fmt ./...
}
+8 -1
View File
@@ -7,7 +7,14 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
unformatted="$(gofmt -l .)"
# Where script/bootstrap installs Go when the host has none.
PATH="$PATH:$ROOT/.tool/go/bin"
# .tool/go holds that toolchain's own sources, which are not ours.
if ! unformatted="$(find . -path ./.tool -prune -o -name '*.go' -type f \
-exec gofmt -l {} +)"; then
echo "fmt-check: gofmt failed" >&2
exit 1
fi
if [ -n "$unformatted" ]; then
echo "Files not formatted:" >&2
echo "$unformatted" >&2
+22 -20
View File
@@ -4,14 +4,13 @@
# own extension to scripts-to-rule-them-all. Idempotent: exits at once
# when the pinned toolchain is already installed.
#
# Only .gitea/workflows/release.yml calls this. goreleaser is not a
# .gitea/workflows/release.yml calls this, and so does script/bootstrap
# when the host has no Go, as on the check runner. goreleaser is not a
# compiler: it shells out to `go` for the `before:` hook and for every
# one of the four cross-compiles, so the release runner needs a Go
# toolchain on PATH. check.yml compiles nothing on the host: it uses
# whatever Go script/bootstrap finds or installs only for that script's
# `go mod download` and for gofmt in script/fmt-check. So this is the
# only host Go that builds anything, and per REPO_POLICIES.md it must be
# pinned by hash.
# toolchain on PATH. The check runner compiles nothing on the host; it
# uses this Go only for bootstrap's `go mod download` and for gofmt in
# script/fmt-check. Per REPO_POLICIES.md a host Go is pinned by hash.
# actions/setup-go exposes no checksum input, so Go is installed the way
# script/install-goreleaser installs goreleaser: download the exact
# archive from go.dev and refuse it unless its sha256 matches the value
@@ -20,12 +19,11 @@
# The version is go.mod's `go` directive, the single source of truth for
# the toolchain. GO_VERSION below MUST equal it, and this script fails
# when they disagree -- so bumping Go is one reviewed change touching
# go.mod, the checksum here, and the Dockerfile's two golang digests
# go.mod, the checksums here, and the Dockerfile's two golang digests
# together.
#
# Linux only, because that is what the release runner is. A darwin dev
# building a snapshot uses their own Go; supporting an OS means adding
# its checksums.
# Linux and macOS, each on amd64 and arm64: the four archives whose
# checksums are committed below.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
@@ -36,6 +34,8 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
GO_VERSION="1.26.1"
SHA256_LINUX_AMD64="031f088e5d955bab8657ede27ad4e3bc5b7c1ba281f05f245bcc304f327c987a"
SHA256_LINUX_ARM64="a290581cfe4fe28ddd737dde3095f3dbeb7f2e4065cab4eae44dfc53b760c2f7"
SHA256_DARWIN_AMD64="65773dab2f8cc4cd23d93ba6d0a805de150ca0b78378879292be0b903b8cdd08"
SHA256_DARWIN_ARM64="353df43a7811ce284c8938b5f3c7df40b7bfb6f56cb165b150bc40b5e2dd541f"
GOROOT_DIR="$ROOT/.tool/go"
GOCMD="$GOROOT_DIR/bin/go"
@@ -104,26 +104,28 @@ main() {
arch="$(uname -m)"
case "$os" in
Linux) os="linux" ;;
Darwin) os="darwin" ;;
*)
echo "install-go: unsupported OS $os (release runner is Linux)" >&2
echo "install-go: unsupported OS $os" >&2
exit 1
;;
esac
case "$arch" in
x86_64 | amd64)
arch="amd64"
sum="$SHA256_LINUX_AMD64"
;;
arm64 | aarch64)
arch="arm64"
sum="$SHA256_LINUX_ARM64"
;;
x86_64 | amd64) arch="amd64" ;;
arm64 | aarch64) arch="arm64" ;;
*)
echo "install-go: no pinned checksum for architecture $arch" >&2
echo "install-go: unsupported architecture $arch" >&2
exit 1
;;
esac
case "${os}-${arch}" in
linux-amd64) sum="$SHA256_LINUX_AMD64" ;;
linux-arm64) sum="$SHA256_LINUX_ARM64" ;;
darwin-amd64) sum="$SHA256_DARWIN_AMD64" ;;
darwin-arm64) sum="$SHA256_DARWIN_ARM64" ;;
esac
archive="go${GO_VERSION}.${os}-${arch}.tar.gz"
url="https://go.dev/dl/${archive}"
+2
View File
@@ -9,6 +9,8 @@ ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
# Where script/bootstrap installs Go when the host has none.
PATH="$PATH:$ROOT/.tool/go/bin"
go mod tidy
go fmt ./...
git diff --exit-code -- go.mod go.sum || {
+2
View File
@@ -44,6 +44,8 @@ resolve_goreleaser() {
main() {
cd "$ROOT"
# Where script/bootstrap installs Go when the host has none.
PATH="$PATH:$ROOT/.tool/go/bin"
if ! bin="$(resolve_goreleaser)"; then
cat >&2 <<EOF
+1 -1
View File
@@ -19,7 +19,7 @@ s3:
secret_access_key: test-secret-key
region: us-east-1
use_ssl: true
part_size: 5242880 # 5MB
part_size: 5242880 # 5MiB
index_path: /tmp/vaultik-test.sqlite
chunk_size: 10MB
blob_size_limit: 10GB