11 Commits
Author SHA1 Message Date
sneak 7a39442eac Run the local index in WAL mode with a busy timeout (closes #217)
check / check (pull_request) Successful in 5m4s
The connection settings were passed as `_journal_mode=`-style
parameters, which the SQLite driver drops without an error, so the
index ran in rollback-journal mode with no busy timeout. `snapshot
list` or `info` reading during a backup could make the backup's next
write fail with "database is locked". Both open paths now pass
`_pragma=` parameters; foreign keys moved there too.

With WAL on, rows committed to the open index can still be in the
-wal file, which a copy of the main file misses. The metadata export
now copies the index with VACUUM INTO, into an empty 0600 file.

The retry after a failed open no longer claims a TRUNCATE recovery; it
retries with the same settings.

Model: opus-5-5
2026-10-06 08:49:03 +00:00
clawbot 4a167e153a Record the real uid and gid of backed-up files (closes #216)
check / check (push) Successful in 4m58s
check / check (pull_request) Successful in 6m21s
The scanner read uid and gid by asserting the stat result to an
interface with Uid() and Gid() methods. *syscall.Stat_t has Uid and Gid
fields, not methods, so the assertion never matched and every file,
directory and symlink was stored as 0:0; a restore as root then gave
everything to root. The scanner now reads the fields of
*syscall.Stat_t.

The first backup after this change re-reads every file not owned by
root, because its stored uid and gid no longer match the disk.

When the tests run as root, as in the Docker build, the new test
compares 0 with 0 and cannot catch the defect; a non-root run does.

Model: opus-5-5
2026-10-06 09:46:16 +02:00
clawbot ea72697992 List a missing file:// destination directory as an error (closes #220)
check / check (push) Successful in 5m53s
check / check (pull_request) Successful in 4m53s
The file backend listed a destination directory that does not exist as
an empty store. With the volume unplugged, snapshot list reported every
local snapshot as missing from the store, snapshot remove said it had
removed metadata it never reached, and prune dropped every local
snapshot record. List and ListStream now fail when the destination
directory is missing, so those commands take their existing path for a
store that cannot be listed. A missing prefix under an existing
directory is still an empty listing, and a first backup still creates
the directory.

Three tests listed a file:// destination nothing had created; they now
create it.

Model: opus-5-5
2026-10-06 08:46:17 +02:00
clawbot 713be502bd Reject a duration with characters outside its parts (closes #215)
check / check (push) Successful in 7m28s
check / check (pull_request) Successful in 5m39s
parseDuration fell back to an unanchored search for number-and-unit
pieces when time.ParseDuration failed, and skipped everything in
between. 1.0y became 0, so `snapshot create --prune --keep-newer-than
1.0y` deleted every snapshot of the backed-up names, the new one
included. 2.1w became one week and 1,5y five years. The fallback now
requires the whole input to be whole-number-and-unit parts with nothing
between them. A bare number is rejected before time.ParseDuration sees
it, since Go reads 0 and +0 as zero with no unit.

Judgement call: a space between number and unit (`30 days`) was
accepted and is now an error, matching Go's own units.

Model: opus-5-5
2026-10-06 06:12:07 +02:00
clawbot 35cf985c18 Re-chunk a known file whose chunks no uploaded blob holds (closes #214)
check / check (push) Successful in 6m53s
check / check (pull_request) Successful in 6m17s
File rows are shared by every snapshot and updated in place, while a
blob row is deleted once no snapshot references it. Removing the newest
snapshot, or the prune after an interrupted run, could drop the only
blob holding a changed file's current chunks while an older snapshot
kept the file row. The next backup compared metadata only, skipped the
file, and completed a snapshot that could not restore it.

The scanner now loads the IDs of known files that list a chunk no
uploaded blob holds and re-chunks them even when their metadata is
unchanged.

The tests append to a file, so the file keeps its first chunk in a blob
the first snapshot still references. Each backup run gets its own
snapshot name, so the second-precision snapshot IDs differ without
sleeping.

Model: opus-5-5
2026-10-06 04:46:16 +02:00
clawbot 070090124a Stamp the tag or short commit in a plain docker build (closes #211)
check / check (push) Successful in 3m35s
check / check (pull_request) Successful in 3m16s
A plain `docker build .` stamped `dev`: `.dockerignore` left out `.git`
and the Dockerfile defaulted VERSION to `dev`. `.dockerignore` now
sends `.git` without `.git/config`. Given no build arguments, the
builder stamps `git describe --tags --always` and the commit and date
from git, and fails if `.git` is present but yields no version. The
empty CHECK_EPOCH refusal is gone so the plain build succeeds.
`script/version` now prints `git describe --tags --always --dirty`, so
make, the scripts and a plain build agree. `vaultik version` treats
the short commit, tag-N-gHASH forms and any version ending in `-dirty`
as development builds, so they keep the development-build notice.

Model: opus-5-5
2026-10-02 10:04:14 +02:00
clawbot 584444b619 List only after-1.0 work in the README roadmap (closes #208)
check / check (push) Successful in 3m35s
check / check (pull_request) Successful in 3m36s
The README roadmap and the TODO.md Next Step still described finished
1.0 work as remaining. The roadmap now lists only work planned after
1.0. Its security item says the code was reviewed before 1.0, every bug
found was fixed, and the accepted risks are listed; an outside audit
stays as after-1.0 work. The error-condition item is gone because every
failure case it listed has a fault-injection test. Daemon mode is added.
TODO.md says the 1.0 work is complete on next and that merging and
tagging are the owner's.

Judgement call: dropped the human-readable size flags item; no command
flag takes a raw-integer size.

Model: opus-5-5
2026-10-01 21:41:41 +02:00
clawbot b30e79ee45 Run the disk-full restore test again (closes #207)
check / check (push) Successful in 4m52s
check / check (pull_request) Successful in 5m19s
TestRestoreReportsDiskFull was skipped pending
#163, which is closed. The skip
and its "skipped until" wording are removed.

Restore is unchanged. With the skip removed the test failed because
restore succeeded: its simulated full disk capped only Create, but
restore now opens each file with OpenFile, so nothing was capped. It now
caps OpenFile, and only for files under the restore target: restore also
writes the decrypted metadata database under $TMPDIR through the same
filesystem, and capping that would fail the restore before any file
reached the target.

Judgement call: the test was corrected, not restore; both assertions are
unchanged.

Model: opus-5-5
2026-10-01 20:24:30 +02:00
sneak d886a9026f Merge branch 'main' into next
check / check (push) Successful in 4m8s
check / check (pull_request) Successful in 2m51s
2026-09-29 03:02:24 +02:00
clawbot 6e1f499048 Document that migrations are supported and none are added before 1.0 (closes #68)
check / check (push) Successful in 3m6s
check / check (pull_request) Successful in 3m8s
The docs now say vaultik supports migrations. The numbered files in `internal/database/schema/` are migrations: `schema_migrations` records which have run, and opening a database applies any that have not. None are added before 1.0 because nothing is installed anywhere yet, so a schema change edits `001.sql` directly. After 1.0 each change is a new numbered file, and an existing local database is migrated when vaultik is updated.

`docs/DATAMODEL.md` owns the explanation. The README caveat and roadmap entry and `AGENTS.md` policy 13 link to it. This replaces the wording from #146, which said there was no upgrade path.

Disclosure: `CLAUDE.md` line 33, the owner's file, changes from "do not need to support migrations" to "do not add migrations before 1.0".

Model: opus-5-5
2026-09-28 20:21:38 +02:00
clawbot d24f5dc33c Adopt the canonical golangci-lint config (closes #90)
check / check (push) Successful in 3m9s
check / check (pull_request) Successful in 1m29s
The lint config is now the canonical file from `prompts`, which replaces the deprecated `gomodguard` with `gomodguard_v2`, so lint prints no deprecation warnings. It also turns on the `depguard` `test-support` rule. The one difference from canonical is that the deny list names vaultik's own test-only package `internal/storage/faultstore`, so shipped code cannot import it. The new config found nothing to fix in the source.

Issues and PRs that pin the old `.golangci.yml` sha256 as an untouched-file check need the new one: `7122fcf0dd0ea57441374f98ebd98bb3da23decb67f9209fee5175170838fbd1`.

Model: opus-5-5
2026-09-23 02:14:40 +02:00
35 changed files with 1359 additions and 430 deletions
+62 -2
View File
@@ -1,4 +1,65 @@
.git # .dockerignore does NOT use .gitignore semantics. Docker matches with
# moby/patternmatcher: filepath.Match plus `**`, so `*` does not cross
# `/` and an unprefixed pattern is anchored at the context root. Every
# depth-independent pattern therefore needs `**/`, or `config/.env` and
# `certs/server.key` still ship while this file reads as solved. Only
# genuinely root-anchored entries go unprefixed. Never transplant these
# into .gitignore, where `**/` is wrong.
#
# Matching is case-sensitive, so secrets use character ranges rather
# than an ALL-CAPS twin, which would still miss `Server.Key`.
#
# Extend with this repo's own host-built artifacts, written anchored:
# `/myapp`, never `**/myapp`, which also matches `cmd/myapp/` and
# deletes the package directory from the context.
# .git is sent without its config. Without a VERSION build argument the
# stage that compiles runs `git describe --tags --always` on .git, which
# does not need .git/config; that file can hold a credential, such as a
# password in a remote URL or the token the CI checkout step stores there.
.git/config
# Agent scratch: one full checkout of the repo per in-flight agent.
# Anchored because it occurs once where agents run at the repo root.
# KNOWN GAP: a repo running agents in subdirectories still ships
# `services/api/.claude/` and must add its own anchored entry.
.claude
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Re-include a committed template with a negation if the
# build needs one: `!docs/example.env`.
**/*.[eE][nN][vV]
**/.[eE][nN][vV].*
**/.[eE][nN][vV][rR][cC]
# Private keys and the bundles carrying them. Public certificates
# (*.crt, *.cer) are deliberately absent: they are legitimate inputs.
**/*.[pP][eE][mM]
**/*.[kK][eE][yY]
**/*.[pP]12
**/*.[pP][fF][xX]
**/[iI][dD]_[rR][sS][aA]
**/[iI][dD]_[dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]
**/[iI][dD]_[eE][dD]25519
# Dependencies: restored inside the image, never copied in.
**/node_modules
# OS metadata.
**/.DS_Store
**/Thumbs.db
# Editor state: never a build input, and it churns COPY.
**/*.swp
**/*.swo
**/*~
**/*.bak
**/.idea
**/.vscode
**/*.sublime-*
# This repo's own entries.
.gitea .gitea
*.md *.md
LICENSE LICENSE
@@ -7,4 +68,3 @@ dist
.tool .tool
coverage.out coverage.out
coverage.html coverage.html
.DS_Store
+70 -2
View File
@@ -10,14 +10,20 @@ run:
linters: linters:
default: all default: all
enable:
# Successor to the deprecated gomodguard. Named explicitly, rather than
# left to `default: all`, because it carries the module policy below.
- gomodguard_v2
disable: disable:
# Genuinely incompatible with project patterns # Genuinely incompatible with project patterns
- exhaustruct # Requires all struct fields - exhaustruct # Requires all struct fields
- depguard # Dependency allow/block lists
- godot # Requires comments to end with periods - godot # Requires comments to end with periods
- wsl # Deprecated, replaced by wsl_v5
- wrapcheck # Too verbose for internal packages - wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go - varnamelen # Short names like db, id are idiomatic Go
# Deprecated: the warning is attached to the old name, so it is
# silenced by disabling that name, not by enabling the successor.
- wsl # Deprecated, replaced by wsl_v5
- gomodguard # Deprecated, replaced by gomodguard_v2
settings: settings:
lll: lll:
line-length: 88 line-length: 88
@@ -28,6 +34,68 @@ linters:
max-complexity: 15 max-complexity: 15
dupl: dupl:
threshold: 100 threshold: 100
depguard:
# Test-support code must not be compiled into the shipped binary. A
# test-support package exists to hand a test privileges the program
# itself must never have, so a file that is not a test must not import
# one. Test files, and the files inside a package whose directory name
# ends in `test`, are where that code belongs, and are exempt.
#
# The deny list below is the one part of this file a repository is
# expected to extend, and the only part it may. depguard matches an
# import path against a list of prefixes, so it cannot be told "any path
# whose last segment ends in test"; a repository's own test-support
# packages have to be named here one at a time, by full import path,
# under a module path that differs from repository to repository. Add
# them; change nothing else.
rules:
test-support:
list-mode: lax
files:
- "$all"
- "!$test"
- "!**/*test/**"
deny:
- pkg: net/http/httptest
desc: >-
Test-support code belongs in test files and in packages whose
directory name ends in test, not in the shipped binary.
- pkg: sneak.berlin/go/vaultik/internal/storage/faultstore
desc: >-
Test-support code belongs in test files and in packages whose
directory name ends in test, not in the shipped binary.
# Only decisions already recorded in the Go package defaults are
# listed here. Every entry matches the module path exactly.
gomodguard_v2:
blocked:
- module: github.com/rs/zerolog
recommendations:
- log/slog
reason: "Structured logging is stdlib log/slog."
# One entry per pre-fork module path, because the later releases
# are separate paths. A prefix match would be shorter but would
# also reach github.com/go-redis/redismock, the test double for
# the successor these entries recommend.
- module: github.com/go-redis/redis
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/go-redis/redis/v7
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/go-redis/redis/v8
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/sergi/go-diff
recommendations:
- github.com/aymanbagabas/go-udiff
reason: "No unified diff output; use go-udiff."
- module: github.com/hexops/gotextdiff
recommendations:
- github.com/aymanbagabas/go-udiff
reason: "Unmaintained fork; use go-udiff."
issues: issues:
max-issues-per-linter: 0 max-issues-per-linter: 0
+1 -3
View File
@@ -47,9 +47,7 @@ checksum:
# A snapshot is not a release and must not name itself like one. The # A snapshot is not a release and must not name itself like one. The
# previous `{{ incpatch .Version }}-next` derived a plausible-looking # previous `{{ incpatch .Version }}-next` derived a plausible-looking
# release number from the last tag -- and with no tags in the repo at # release number from the last tag -- and with no tags in the repo at
# all, from goreleaser's fabricated v0.0.0. This produces the same # all, from goreleaser's fabricated v0.0.0.
# string script/version produces for an untagged build, so a snapshot
# binary and a `make vaultik` binary of the same clean commit agree.
snapshot: snapshot:
version_template: "dev-{{ slice .FullCommit 0 12 }}" version_template: "dev-{{ slice .FullCommit 0 12 }}"
+8 -10
View File
@@ -102,14 +102,12 @@ Version: 2025-06-08
build files are acceptable in the root, but source code and other files build files are acceptable in the root, but source code and other files
should be organized in appropriate subdirectories. should be organized in appropriate subdirectories.
13. Pre-1.0: NEVER write database migrations. There are no live databases 13. Pre-1.0: NEVER add a database migration. Migrations are supported, but
anywhere — every user's local index can be rebuilt from a fresh full nothing is installed anywhere yet, so there is nothing to migrate. To
backup. To change the schema, edit `internal/database/schema/001.sql` change the schema, edit `internal/database/schema/001.sql` (and any
(and any code that touches the affected tables) directly; do not add new code that touches the affected tables) directly. After 1.0, each schema
numbered schema files. Those numbered files and the `schema_migrations` change is a new numbered file in that directory and a released file is
table they populate only bootstrap a fresh database — they are not an never edited; an existing local database is then migrated when vaultik
upgrade path. The local index is disposable until 1.0 ships and is is updated. See
tagged; once 1.0 is tagged that clause expires and the question of [`docs/DATAMODEL.md`](docs/DATAMODEL.md#schema-migrations).
upgrading existing indexes returns. See [`docs/DATAMODEL.md`](docs/DATAMODEL.md)
for the full explanation.
+1 -1
View File
@@ -353,7 +353,7 @@ CreateSnapshot(opts)
## Deduplication Strategy ## Deduplication Strategy
1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid) 1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid), unless the file lists a chunk that no uploaded blob holds; such a file is re-chunked
2. **Chunk-level**: Chunks are content-addressed by SHA256 hash. If a chunk hash already exists in the database, the chunk data is not re-uploaded. 2. **Chunk-level**: Chunks are content-addressed by SHA256 hash. If a chunk hash already exists in the database, the chunk data is not re-uploaded.
+1 -1
View File
@@ -30,7 +30,7 @@ Read the rules in AGENTS.md and follow them.
done provided to you in the initial instruction. Don't do part or most of done provided to you in the initial instruction. Don't do part or most of
the work, do all of the work until the criteria for done are met. the work, do all of the work until the criteria for done are met.
* We do not need to support migrations; schema upgrades can be handled by * We do not add migrations before 1.0; schema upgrades can be handled by
deleting the local state file and doing a full backup to re-create it. deleting the local state file and doing a full backup to re-create it.
* When testing on a 2.5Gbit/s ethernet to an s3 server backed by 2000MB/sec SSD, * When testing on a 2.5Gbit/s ethernet to an s3 server backed by 2000MB/sec SSD,
+27 -31
View File
@@ -20,10 +20,10 @@
# golang:1.26.1-alpine, 2026-03-17 # golang:1.26.1-alpine, 2026-03-17
FROM golang:1.26.1-alpine@sha256:2389ebfa5b7f43eeafbd6be0c3700cc46690ef842ad962f6c5bd6be49ed82039 AS builder FROM golang:1.26.1-alpine@sha256:2389ebfa5b7f43eeafbd6be0c3700cc46690ef842ad962f6c5bd6be49ed82039 AS builder
# Build tooling: make, plus a C toolchain because `go test -race` needs cgo. # Build tooling: make, plus a C toolchain because `go test -race` needs cgo,
# The sqlite driver is pure Go (modernc.org/sqlite), so no sqlite library or # and git, which derives the version below. The sqlite driver is pure Go
# CLI is required. # (modernc.org/sqlite), so no sqlite library or CLI is required.
RUN apk add --no-cache make build-base RUN apk add --no-cache make build-base git
WORKDIR /src WORKDIR /src
@@ -47,48 +47,44 @@ COPY . .
# unreferenced-ARG handling staying as it is. It also puts the epoch in # unreferenced-ARG handling staying as it is. It also puts the epoch in
# the build log, where a reader can see the layer was keyed fresh. # the build log, where a reader can see the layer was keyed fresh.
# #
# The guard is what makes a build that omits --build-arg fail instead of # A build that passes no CHECK_EPOCH, such as a plain `docker build .`,
# lie. An unset ARG is an empty string, and an empty string is a # keys these layers on the empty string, so rebuilding an unchanged
# perfectly stable cache key: without the guard the first such build # checkout replays them from cache and runs nothing. Only the scripts'
# runs the checks and every one after it on an unchanged tree replays # builds mean the checks executed.
# these layers from cache, executes nothing, and still exits 0. Failed
# steps are never cached, so the guard fails on EVERY invocation rather
# than once -- a bare `docker build .` is a loud error, not a quiet
# green. Do not give CHECK_EPOCH a default value; a default would
# satisfy the guard with a constant and restore the hole.
# #
# Everything above this line (apk, go.mod, `go mod download`) is # Everything above this line (apk, go.mod, `go mod download`) is
# deliberately outside the busted range and keeps caching. # deliberately outside the busted range and keeps caching.
ARG CHECK_EPOCH ARG CHECK_EPOCH
RUN [ -n "$CHECK_EPOCH" ] || exit 1
RUN echo "check epoch: ${CHECK_EPOCH}" && make fmt-check RUN echo "check epoch: ${CHECK_EPOCH}" && make fmt-check
RUN echo "check epoch: ${CHECK_EPOCH}" && make test RUN echo "check epoch: ${CHECK_EPOCH}" && make test
# Version, commit and build date are computed on the host by # Version, commit and build date: the build args when given (script/docker
# script/docker and script/cibuild (where .git exists) and passed in as # and script/cibuild pass the ones they compute on the host), otherwise
# build args. The build context excludes .git (see .dockerignore), so # derived from the .git in the build context. The version is then `git
# the build cannot derive them itself: it used to try, with `git # describe --tags --always`: the tag on a tagged commit, tag-N-gHASH after
# rev-parse` inside this stage, and always got "unknown". VERSION comes # one, the short commit when no tag is reachable. A context that carries
# from script/version, the source of truth shared with the Makefile, so # .git and still yields no version fails the build; one without .git, as
# it carries the same tag / dev-<sha> / -dirty rules and a Docker image # from a source tarball, stamps "dev" and an "unknown" commit and date.
# reports the same string a local build of the same tree would.
#
# The defaults are the fallback for a bare `docker build .` that passes
# none of them: an unset arg would otherwise stamp an empty string and
# produce an image that cannot report its own version, commit or date.
# They match what an out-of-git build reports elsewhere.
# #
# These ARGs sit here, after the checks, rather than at the top of the # These ARGs sit here, after the checks, rather than at the top of the
# stage: every commit changes their values, and a value change # stage: every commit changes their values, and a value change
# invalidates all layers below the ARG. Declared up top they would bust # invalidates all layers below the ARG. Declared up top they would bust
# `go mod download`; here they only rekey this build layer, which the # `go mod download`; here they only rekey this build layer, which the
# COPY of the sources above already rebuilds on any change anyway. # COPY of the sources above already rebuilds on any change anyway.
ARG VERSION=dev ARG VERSION
ARG COMMIT=unknown ARG COMMIT
ARG COMMIT_DATE=unknown ARG COMMIT_DATE
# Build (pure Go, no CGO required since we use modernc.org/sqlite) # Build (pure Go, no CGO required since we use modernc.org/sqlite)
RUN CGO_ENABLED=0 go build -ldflags "-X 'sneak.berlin/go/vaultik/internal/globals.Version=${VERSION}' -X 'sneak.berlin/go/vaultik/internal/globals.Commit=${COMMIT}' -X 'sneak.berlin/go/vaultik/internal/globals.CommitDate=${COMMIT_DATE}'" -o /vaultik ./cmd/vaultik RUN version="${VERSION:-$(git describe --tags --always || echo dev)}"; \
if [ -e .git ] && { [ -z "$version" ] || [ "$version" = dev ] || \
[ "$version" = unknown ]; }; then \
echo "the build context carries .git but yields no version" >&2; \
exit 1; \
fi; \
commit="${COMMIT:-$(git rev-parse HEAD || echo unknown)}"; \
commit_date="${COMMIT_DATE:-$(git show -s --format=%cs HEAD || echo unknown)}"; \
CGO_ENABLED=0 go build -ldflags "-X 'sneak.berlin/go/vaultik/internal/globals.Version=${version}' -X 'sneak.berlin/go/vaultik/internal/globals.Commit=${commit}' -X 'sneak.berlin/go/vaultik/internal/globals.CommitDate=${commit_date}'" -o /vaultik ./cmd/vaultik
# Runtime stage # Runtime stage
# alpine:3.21, 2026-02-25 # alpine:3.21, 2026-02-25
+2 -2
View File
@@ -1,7 +1,7 @@
.PHONY: all bootstrap setup check test lint lint-fix fmt fmt-check build clean deps test-coverage local install release release-snapshot docker hooks .PHONY: all bootstrap setup check test lint lint-fix fmt fmt-check build clean deps test-coverage local install release release-snapshot docker hooks
# Version number, derived from git by script/version -- the tag when # Version number, derived from git by script/version (`git describe
# HEAD is on one, otherwise dev-<sha>. This used to be a hardcoded # --tags --always --dirty`). This used to be a hardcoded
# constant, which meant every local build claimed to be a release that # constant, which meant every local build claimed to be a release that
# had never been tagged. # had never been tagged.
VERSION := $(shell script/version) VERSION := $(shell script/version)
+69 -58
View File
@@ -40,7 +40,8 @@ Features:
* modern encryption ([age](https://age-encryption.org/), X25519 + ChaCha20-Poly1305) * modern encryption ([age](https://age-encryption.org/), X25519 + ChaCha20-Poly1305)
* content-defined chunking with deduplication (FastCDC) * content-defined chunking with deduplication (FastCDC)
* incremental backups (only changed files are re-chunked) * incremental backups (a file is re-chunked only when it changed or a
chunk it lists is held by no uploaded blob)
* multithreaded zstd compression at configurable levels * multithreaded zstd compression at configurable levels
* content-addressed immutable storage * content-addressed immutable storage
* local state tracking in SQLite (enables write-only incremental backups) * local state tracking in SQLite (enables write-only incremental backups)
@@ -499,8 +500,9 @@ format does and does not protect.
* Content-defined chunking using the FastCDC algorithm * Content-defined chunking using the FastCDC algorithm
* Average chunk size: configurable (default 10MB) * Average chunk size: configurable (default 10MB)
* Deduplication at file level (unchanged files skipped) and chunk level * Deduplication at file level (unchanged files skipped, unless a chunk
(identical chunks across files stored once) the file lists is held by no uploaded blob) and chunk level (identical
chunks across files stored once)
* Multiple chunks packed into blobs to reduce object count * Multiple chunks packed into blobs to reduce object count
### encryption ### encryption
@@ -559,13 +561,13 @@ complete annotated example also lives in
sequentially. Restore speed is bound by single-stream throughput. sequentially. Restore speed is bound by single-stream throughput.
* **Device nodes, named pipes, and sockets are silently skipped.** Only * **Device nodes, named pipes, and sockets are silently skipped.** Only
regular files, directories, and symlinks are backed up. regular files, directories, and symlinks are backed up.
* **No upgrade path between versions.** There is no supported way to carry * **Before 1.0, an update can make the local index unusable.** Vaultik
an existing local index across a schema change; if the local SQLite supports schema migrations, but none are added before 1.0 because
schema changes between versions, delete the local database (`vaultik there is no installed base yet. If an update leaves your local index
database delete`) and run a full backup. Remote storage is unaffected. unusable, run `vaultik database delete` and then a full backup; remote
(The binary does embed numbered schema files and a `schema_migrations` storage is unaffected. After 1.0, the local index is migrated when
table to bootstrap a fresh database — see [`docs/DATAMODEL.md`](docs/DATAMODEL.md) vaultik is updated. See
— but that is not an upgrade path.) [`docs/DATAMODEL.md`](docs/DATAMODEL.md#schema-migrations).
* **Files that change during backup may be inconsistent.** There is no * **Files that change during backup may be inconsistent.** There is no
filesystem snapshot or freeze. If a file is modified between the scan filesystem snapshot or freeze. If a file is modified between the scan
and chunk phases, the backed-up copy may reflect a partial write. and chunk phases, the backed-up copy may reflect a partial write.
@@ -577,22 +579,18 @@ complete annotated example also lives in
## roadmap ## roadmap
Items still to do before / shortly after 1.0. Loosely ordered by Work planned after 1.0. Loosely ordered by priority.
priority.
### correctness and operability ### correctness and operability
* **Security audit of the encryption implementation.** Pre-1.0 * **Outside security audit.** Before 1.0 the encryption and
blocker if we're advertising "secure" at the top of this README. blob-generation code was reviewed: every bug the review found was
age + zstd + content-defined chunking is mostly off-the-shelf fixed, and the risks it accepted are listed in
pieces, but the seams (key handling, recipient parsing, manifest [Accepted Risks](docs/REPOSTRUCTURE.md#accepted-risks). No outside
trust boundary, restore-time identity validation) need an outside audit has been done. age + zstd + content-defined chunking
read. is mostly off-the-shelf pieces, but the seams (key handling,
* **Error-condition tests.** Today's coverage is the happy path recipient parsing, manifest trust boundary, restore-time identity
plus a few specific regressions. Need fault-injection coverage: validation) need an outside read.
network failures mid-blob, disk-full during restore, corrupted /
truncated / missing blobs, partial uploads, kill -9 between
manifest and db.zst.age writes.
* **Verify restored content end-to-end in CI.** The current * **Verify restored content end-to-end in CI.** The current
integration test does this for a small synthetic snapshot but integration test does this for a small synthetic snapshot but
not at scale. A nightly job against a multi-GB representative not at scale. A nightly job against a multi-GB representative
@@ -615,13 +613,15 @@ priority.
doesn't resume from where it stopped or skip already-present doesn't resume from where it stopped or skip already-present
files. A `--resume` mode that checks targets before fetching files. A `--resume` mode that checks targets before fetching
blobs would matter for very large restores. blobs would matter for very large restores.
* **Daemon mode.** A long-running mode that watches for file
changes so frequent backups, such as hourly, skip the full scan.
It adds little for the usual runs from cron every 12 to 36 hours.
See [issue #204](https://git.eeqj.de/sneak/vaultik/issues/204).
### usability ### usability
* **Man pages and richer `--help` examples.** Cobra generates * **Man pages and richer `--help` examples.** Cobra generates
basic help; man pages would be a separate target. basic help; man pages would be a separate target.
* **`--bwlimit` style human-readable size flags** across the
command surface where they're currently raw integers.
* **`vaultik snapshot diff <a> <b>`** — show which files changed * **`vaultik snapshot diff <a> <b>`** — show which files changed
between two snapshots without restoring either. between two snapshots without restoring either.
* **Status reporting hook for `--cron`.** When a backup fails * **Status reporting hook for `--cron`.** When a backup fails
@@ -631,12 +631,11 @@ priority.
### infrastructure ### infrastructure
* **Cross-version schema upgrades.** There is no upgrade path between * **Schema migrations after 1.0.** Migrations are supported, but none
released versions — pre-1.0 schema changes are handled by `vaultik are added before 1.0 because there is no installed base yet. After
database delete` plus a full re-scan (see 1.0, each schema change is a new migration, so an existing local
[`docs/DATAMODEL.md`](docs/DATAMODEL.md)). Post-1.0 we'll need a index is migrated when vaultik is updated (see
migration story to keep existing index databases usable across [`docs/DATAMODEL.md`](docs/DATAMODEL.md#schema-migrations)).
upgrades.
* **Storage backend coverage tests.** S3, file://, and rclone:// * **Storage backend coverage tests.** S3, file://, and rclone://
all share the Storer interface but the rclone path is the least all share the Storer interface but the rclone path is the least
exercised in CI. exercised in CI.
@@ -773,8 +772,9 @@ them. We provide:
* `script/projectname` — print the project name (used for the Docker * `script/projectname` — print the project name (used for the Docker
image tag) image tag)
* `script/version` — print the version string to bake into the binary. * `script/version` — print the version string to bake into the binary.
The `Makefile`'s `LDFLAGS` call this; it is the single source of truth The `Makefile`'s `LDFLAGS` call this, and `script/docker` and
for the version. See [releasing](#releasing) for the rules. `script/cibuild` pass its output to the image build. See
[releasing](#releasing) for the rules.
* `script/install-goreleaser` — install the pinned `goreleaser` into * `script/install-goreleaser` — install the pinned `goreleaser` into
`.tool/bin` from a sha256-verified release archive. Idempotent, and `.tool/bin` from a sha256-verified release archive. Idempotent, and
called by `script/bootstrap`; the release workflow calls it directly called by `script/bootstrap`; the release workflow calls it directly
@@ -866,16 +866,18 @@ them. We provide:
module layers sit above the `ARG` and still cache, so a build is not module layers sit above the `ARG` and still cache, so a build is not
cold. cold.
A build that supplies no `CHECK_EPOCH` — a bare `docker build .` or A `docker build -f Dockerfile.lint .` that supplies no `CHECK_EPOCH`
`docker build -f Dockerfile.lint .` — fails rather than lying. An fails rather than lying. An unset `ARG` is an empty string and an
unset `ARG` is an empty string and an empty string is a stable cache empty string is a stable cache key, so without a guard such a build
key, so without a guard such a build would serve every check layer would serve the lint layer from cache, execute nothing, and still exit
from cache, execute nothing, and still exit 0. Each file therefore 0. `Dockerfile.lint` therefore asserts the value is non-empty before
asserts the value is non-empty before running anything, and because running anything, and because failed steps are never cached that
failed steps are never cached that assertion fires on every assertion fires on every invocation rather than once. The product
invocation rather than once. Use `script/lint`, `script/docker` or `Dockerfile` has no such guard, because a plain `docker build .` must
`script/cibuild`, which pass the arg; a bare `docker build` is a loud succeed: without `CHECK_EPOCH`, rebuilding an unchanged checkout
error. replays its check layers from cache. Use `script/lint`,
`script/docker` or `script/cibuild`, which pass the arg, when the
checks must run.
* `script/precommit` — pre-commit gate: `go mod tidy` + `go fmt` (must * `script/precommit` — pre-commit gate: `go mod tidy` + `go fmt` (must
not change files), then `script/check` not change files), then `script/check`
* `script/install-precommit` — install the git pre-commit hook that * `script/install-precommit` — install the git pre-commit hook that
@@ -886,24 +888,33 @@ them. We provide:
### version numbers ### version numbers
The version a binary reports comes from git, not from a constant in a The version a binary reports comes from git, not from a constant in a
file. `script/version` decides it, and everything that stamps a binary file. It is `git describe --tags --always --dirty`, which
agrees with it: `script/version` runs for the `Makefile`, `script/docker` and
`script/cibuild`:
* `HEAD` is exactly on a tag → that tag with a leading `v` stripped, so * `HEAD` is exactly on a tag → that tag, such as `v1.0.0`.
the tag `v1.0.0` produces `vaultik 1.0.0`, matching the archive name * a commit after a tag → `<tag>-<N>-g<short sha>`.
`vaultik_1.0.0_linux_amd64.tar.gz`. `goreleaser` strips the prefix the * no tag reachable → the short commit sha.
same way. * any of these, with uncommitted changes to tracked files → a `-dirty`
* anything else → `dev-<12 chars of the commit sha>`.
* either, with uncommitted changes to tracked files → a `-dirty`
suffix, because a modified checkout of a tag is not that tag. suffix, because a modified checkout of a tag is not that tag.
A build that is not a release never names itself like one. `vaultik A `docker build .` of a clone, with no build arguments, runs the same
version` says so in as many words on a development build, and `git describe` (without `--dirty`) on the `.git` in its build context,
`goreleaser --snapshot` stamps the same `dev-<sha>` string rather than so it stamps the same value for a clean commit; the build fails if the
inventing the next patch number. If `script/version` cannot be run at context carries `.git` and no version comes out. A binary built without
all, `make` stops with an error instead of building an unversioned git metadata reports `dev`.
binary, and a binary that somehow carries an empty version string still
reports itself as a development build. `goreleaser` stamps a release binary with the tag minus its leading
`v`, so the tag `v1.0.0` produces `vaultik 1.0.0`, matching the archive
name `vaultik_1.0.0_linux_amd64.tar.gz`. `goreleaser --snapshot` stamps
`dev-<12 chars of the commit sha>` rather than inventing the next patch
number. `vaultik version` calls a build a development build when its
version is `dev`, `dev-<sha>`, the short commit sha or
`<tag>-<N>-g<short sha>`, with or without `-dirty`; only a plain tag,
such as `v1.0.0` or `1.0.0`, is a release. If `script/version` cannot
be run at all, `make` stops with an error instead of building an
unversioned binary, and a binary that somehow carries an empty version
string still reports itself as a development build.
### cutting a release ### cutting a release
+52 -9
View File
@@ -14,17 +14,59 @@ pre-1.0
# Next Step # Next Step
Define the remaining scope for the first tagged release under the 1.0.0 The 1.0 work is complete on `next`: the scope settled on
milestone, then cut that tag. The mechanism to cut it now exists and is [issue #125](https://git.eeqj.de/sneak/vaultik/issues/125) was the work
exercised; what is left is the scope decision, which is the owner's. already planned for 1.0, and all of it has landed. The mechanism to cut
This step deliberately names one version number: it previously said the tag exists and is exercised; what is left is merging `next` to
"cut v0.1.0" while the `Makefile` baked in `1.0.0-rc.1` and the issue `main` and tagging, both the owner's.
milestone said 1.0.0, and three different answers to "what is the next
release" is exactly the contradiction
[issue #65](https://git.eeqj.de/sneak/vaultik/issues/65) was filed over.
# Completed Steps # Completed Steps
- 2026-10-06: Made the local index actually run in WAL mode with a busy
timeout ([issue #217](https://git.eeqj.de/sneak/vaultik/issues/217)).
The connection settings were written in a form the SQLite driver
ignores, so the index ran without either and `snapshot list` or `info`
during a backup could make the backup's next write fail with
`database is locked`. They are now `_pragma=` parameters, and the
metadata export copies the open index with `VACUUM INTO`, because a
copy of the file alone misses rows still in the `-wal` file.
- 2026-10-06: Made a backup record the real uid and gid of files,
directories and symlinks
([issue #216](https://git.eeqj.de/sneak/vaultik/issues/216)). The
scanner asked the stat result for `Uid()` and `Gid()` methods, which
`*syscall.Stat_t` does not have, so every entry was stored as `0:0`
and a restore as root gave every file to root. It now reads the
`Uid` and `Gid` fields of `*syscall.Stat_t`.
- 2026-10-06: Made a `file://` destination whose directory is missing
count as one that cannot be listed
([issue #220](https://git.eeqj.de/sneak/vaultik/issues/220)). The file
backend listed a missing directory as an empty store, so with the
quickstart's USB stick unplugged `snapshot list` reported every local
snapshot as missing from the store, `snapshot remove` claimed to have
removed metadata it never reached, and `prune` dropped every local
snapshot record. Listing a missing destination directory is now an
error; a first backup still creates the directory.
- 2026-10-06: Made `--older-than` and `--keep-newer-than` reject a
duration with characters outside its number-and-unit parts
([issue #215](https://git.eeqj.de/sneak/vaultik/issues/215)). The
parser picked out the parts it recognised and skipped the rest, so
`1.0y` became zero and `--prune --keep-newer-than 1.0y` deleted every
snapshot of the backed-up names, the new one included. `1.0y`,
`1,5y`, `30 days`, `x7d` and a bare `0` are now errors; decimals still
work in Go units such as `1.5h`.
- 2026-10-06: Made a backup re-chunk a known file that lists a chunk no
uploaded blob holds
([issue #214](https://git.eeqj.de/sneak/vaultik/issues/214)). File
rows are shared by every snapshot and updated in place, so removing or
pruning the only snapshot that referenced a changed file's current
blob left an older snapshot keeping a row that matched the disk while
no blob held its chunks. The next backup skipped the file and
completed a snapshot that could not restore it.
- 2026-09-22: Routed the last direct-to-stdout command output through - 2026-09-22: Routed the last direct-to-stdout command output through
`internal/ui` `internal/ui`
([issue #149](https://git.eeqj.de/sneak/vaultik/issues/149)). The ([issue #149](https://git.eeqj.de/sneak/vaultik/issues/149)). The
@@ -693,4 +735,5 @@ release" is exactly the contradiction
# Future Steps # Future Steps
None queued; the release-scoping item is now the Next Step. Work planned after 1.0 is listed in the README
[roadmap](README.md#roadmap).
+33 -26
View File
@@ -11,9 +11,10 @@ import (
// This file guards the version stamping of the product image (issue // This file guards the version stamping of the product image (issue
// #75). The failure it protects against is silent: the image still // #75). The failure it protects against is silent: the image still
// builds and runs, but `vaultik version` inside it reports "commit: // builds and runs, but `vaultik version` inside it reports "commit:
// unknown", so an operator cannot tell which source produced a given // unknown" or a version of "dev", so an operator cannot tell which
// backup. .dockerignore excludes .git, so the build cannot derive the // source produced a given backup. The build takes the values as build
// commit itself; the values must be computed on the host and passed in. // args, which script/docker computes on the host, and otherwise derives
// them from the .git in its context.
// //
// These are parses of the committed files, for the same reason the lint // These are parses of the committed files, for the same reason the lint
// guards next door are: shelling out to docker would nest a build // guards next door are: shelling out to docker would nest a build
@@ -32,35 +33,41 @@ func versionArgs() []string {
} }
// TestProductDockerfileTakesVersionAsBuildArgs fails unless the build // TestProductDockerfileTakesVersionAsBuildArgs fails unless the build
// declares each version arg and stamps it into the binary by ldflag // declares each version arg, with no default, and stamps it into the
// reference, rather than computing it in the container. // binary whenever it is given, ahead of the value derived in the
// container.
func TestProductDockerfileTakesVersionAsBuildArgs(t *testing.T) { func TestProductDockerfileTakesVersionAsBuildArgs(t *testing.T) {
t.Parallel() t.Parallel()
found := instructions(t, productDockerfile) found := instructions(t, productDockerfile)
for _, arg := range versionArgs() { for _, arg := range versionArgs() {
require.GreaterOrEqual(t, indexOf(found, "ARG "+arg), 0, require.Contains(t, found, "ARG "+arg,
"%s must declare `ARG %s` so the host can pass it in", "%s must declare `ARG %s`, with no default, so the host can"+
productDockerfile, arg) " pass it in", productDockerfile, arg)
assertLdflagReferences(t, found, arg) assertLdflagReferences(t, found, arg)
} }
} }
// TestProductDockerfileDoesNotDeriveVersionItself is the anti-regression // TestProductDockerfileDerivesVersionFromGit fails unless a build given
// for the original defect: the container ran `git rev-parse`, but .git // no VERSION, such as a plain `docker build .` of a clone, takes it from
// is not in the build context, so it always resolved to "unknown". No // `git describe` of the .git in its context, and fails rather than
// git command may reach into a build that cannot see the history. // stamp "dev" when that .git yields no version.
func TestProductDockerfileDoesNotDeriveVersionItself(t *testing.T) { func TestProductDockerfileDerivesVersionFromGit(t *testing.T) {
t.Parallel() t.Parallel()
text := instructionText(readRepoFile(t, productDockerfile)) found := instructions(t, productDockerfile)
assert.NotContains(t, text, "git ", buildAt := indexContaining(found, "go build")
"%s must not run git: .git is excluded from the build context, so"+ require.GreaterOrEqual(t, buildAt, 0, "%s must build", productDockerfile)
" any value it derives is wrong. Pass version, commit and date"+
" in as build args instead.", productDockerfile) assert.Contains(t, found[buildAt], "git describe --tags --always",
"%s must derive the version from git when no VERSION is given",
productDockerfile)
assert.Contains(t, found[buildAt], "[ -e .git ]",
"%s must fail when the context carries .git but yields no version",
productDockerfile)
} }
// TestDockerScriptComputesVersionOnTheHost fails unless script/docker // TestDockerScriptComputesVersionOnTheHost fails unless script/docker
@@ -78,25 +85,25 @@ func TestDockerScriptComputesVersionOnTheHost(t *testing.T) {
} }
assert.Contains(t, script, "/version", assert.Contains(t, script, "/version",
"%s must take VERSION from script/version, the source of truth"+ "%s must take VERSION from script/version, as the Makefile does",
" shared with the Makefile", dockerScript) dockerScript)
} }
// assertLdflagReferences fails unless some build instruction stamps the // assertLdflagReferences fails unless the build instruction uses the
// named variable from the ARG (a ${arg} reference), not from a value // named ARG whenever it is given (a ${arg:- reference), so a value
// computed inside the container. // passed in is not overridden by one derived inside the container.
func assertLdflagReferences(t *testing.T, found []string, arg string) { func assertLdflagReferences(t *testing.T, found []string, arg string) {
t.Helper() t.Helper()
for _, instruction := range found { for _, instruction := range found {
if strings.HasPrefix(instruction, "RUN ") && if strings.HasPrefix(instruction, "RUN ") &&
strings.Contains(instruction, "go build") && strings.Contains(instruction, "go build") &&
strings.Contains(instruction, "${"+arg+"}") { strings.Contains(instruction, "${"+arg+":-") {
return return
} }
} }
assert.Fail(t, "version arg is declared but never stamped", assert.Fail(t, "version arg is declared but never stamped",
"the go build in %s must reference ${%s} in its ldflags, or the"+ "the go build in %s must use ${%s:-...}, or the arg is passed and"+
" arg is passed and discarded", productDockerfile, arg) " discarded", productDockerfile, arg)
} }
+8 -7
View File
@@ -152,20 +152,21 @@ func TestLintDockerfileVerifiesTheLinterConfig(t *testing.T) {
assertEpochExpandedInto(t, found, verify) assertEpochExpandedInto(t, found, verify)
} }
// TestProductDockerfileCannotBeCachedGreen holds the same line for the // TestProductDockerfileKeysChecksOnTheEpoch holds the same line for the
// checks that remain in the product image build. // checks that remain in the product image build, but without the guard:
func TestProductDockerfileCannotBeCachedGreen(t *testing.T) { // a plain `docker build .` with no build arguments must succeed.
func TestProductDockerfileKeysChecksOnTheEpoch(t *testing.T) {
t.Parallel() t.Parallel()
found := instructions(t, productDockerfile) found := instructions(t, productDockerfile)
argAt := indexOf(found, checkEpochARG) argAt := indexOf(found, checkEpochARG)
require.GreaterOrEqual(t, argAt, 0, require.GreaterOrEqual(t, argAt, 0,
"%s must declare `%s` with no default value", "%s must declare `%s`", productDockerfile, checkEpochARG)
productDockerfile, checkEpochARG)
assert.GreaterOrEqual(t, indexOf(found, checkEpochGuard), argAt, assert.Equal(t, -1, indexOf(found, checkEpochGuard),
"%s must guard against an empty CHECK_EPOCH", productDockerfile) "%s must not refuse an empty CHECK_EPOCH: a plain `docker build .`"+
" must succeed", productDockerfile)
assertEpochExpandedInto(t, found[argAt:], "make fmt-check") assertEpochExpandedInto(t, found[argAt:], "make fmt-check")
assertEpochExpandedInto(t, found[argAt:], "make test") assertEpochExpandedInto(t, found[argAt:], "make test")
+20 -24
View File
@@ -6,33 +6,28 @@ Vaultik uses a local SQLite database to track file metadata, chunk mappings, and
**Important Notes:** **Important Notes:**
This section is the authoritative explanation of the schema/migration story;
other documents (the README and `AGENTS.md`) link here.
- **No upgrade path between versions (pre-1.0)**: Vaultik has no supported way to
carry an existing local index across a schema change. The index is disposable
— if the on-disk schema changes between versions, delete the local SQLite
database (`vaultik database delete`) and run a full backup. Remote storage is
unaffected; the new index re-deduplicates against existing remote blobs. This
is the standing project policy, and it is separate from the schema bootstrap
described next.
- **Schema bootstrap**: a fresh database is populated from numbered SQL files
embedded in the binary under `internal/database/schema/`. `000.sql` creates the
`schema_migrations` table; `001.sql` creates the application tables. On opening
a database the code applies each numbered file that has not yet run and records
its version in `schema_migrations`. This bootstraps a new database; it does not
upgrade an existing one between released versions.
- **Changing the schema (pre-1.0)**: edit `internal/database/schema/001.sql` (and
the code that touches the affected tables) directly. Do not add new numbered
files — there is no installed base to migrate.
- **Disposability expires at 1.0**: the index is treated as disposable only until
1.0 ships and is tagged. Once 1.0 is tagged that clause expires and the
question of upgrading existing indexes returns. It is deliberately left open
here.
- **Version Compatibility**: In rare cases, you may need to use the same version - **Version Compatibility**: In rare cases, you may need to use the same version
of Vaultik to restore a backup as was used to create it. This ensures of Vaultik to restore a backup as was used to create it. This ensures
compatibility with the metadata format stored in S3. compatibility with the metadata format stored in S3.
## Schema Migrations
Vaultik supports schema migrations. They are the numbered SQL files in
`internal/database/schema/`, embedded in the binary: `000.sql` creates the
`schema_migrations` table, which records each migration that has run, and
`001.sql` creates the application tables. `database.New` opens a database and
applies, in order, every migration that database has not yet recorded.
**Before 1.0** no migrations are added, because nothing is installed anywhere
yet. A schema change edits `001.sql` (and the code that uses the affected
tables) directly. A local database created before the change has already
recorded `001.sql` as run, so it keeps the old schema and can become unusable;
`vaultik database delete` followed by a full backup rebuilds it.
**After 1.0** each schema change is a new numbered file, so an existing local
database is migrated the first time an updated vaultik opens it. A file that
has shipped in a release is never edited.
## Database Tables ## Database Tables
### 1. `files` ### 1. `files`
@@ -200,6 +195,7 @@ Tracks blob upload metrics.
1. **Change Detection** 1. **Change Detection**
- `SELECT * FROM files WHERE path = ?` - Get previous file metadata - `SELECT * FROM files WHERE path = ?` - Get previous file metadata
- Compare mtime, size, mode to detect changes - Compare mtime, size, mode to detect changes
- Re-chunk a file that lists a chunk no uploaded blob holds, even when its metadata is unchanged
- Skip unchanged files but still add to `snapshot_files` - Skip unchanged files but still add to `snapshot_files`
2. **Chunk Reuse** 2. **Chunk Reuse**
@@ -284,7 +280,7 @@ This ensures consistency, especially important for operations like:
3. **Batch Operations**: Where possible, operations are batched within transactions 3. **Batch Operations**: Where possible, operations are batched within transactions
4. **Write-Ahead Logging**: SQLite WAL mode is enabled for better concurrency 4. **Write-Ahead Logging**: The local index runs in SQLite WAL mode with a 10-second busy timeout, so a read-only command such as `snapshot list` can read it while a backup writes to it. Committed rows can sit in the `-wal` file beside the index until a checkpoint, so the metadata export copies the index through SQLite (`VACUUM INTO`), not as a file
## Data Integrity ## Data Integrity
+9 -6
View File
@@ -172,9 +172,9 @@ func TestBannerSuppressedInArgs(t *testing.T) {
// hermeticConfig is a complete, valid config that needs no network and // hermeticConfig is a complete, valid config that needs no network and
// no credentials: file:// storage is exempt from the S3 credential // no credentials: file:// storage is exempt from the S3 credential
// checks, and FileStorer over a directory that does not exist lists // checks. A test that lists the destination must create its directory
// zero objects without erroring. Chunk, blob and compression settings // first, because listing a directory that does not exist is an error.
// are filled in by config.Load. // Chunk, blob and compression settings are filled in by config.Load.
const hermeticConfig = `age_recipients: const hermeticConfig = `age_recipients:
- age1278m9q7dp3chsh2dcy82qk27v047zywyvtxwnj4cvt0z65jw6a7q5dqhfj - age1278m9q7dp3chsh2dcy82qk27v047zywyvtxwnj4cvt0z65jw6a7q5dqhfj
snapshots: snapshots:
@@ -196,22 +196,25 @@ hostname: test-host
// `snapshot list` is the command chosen because it is the only --json // `snapshot list` is the command chosen because it is the only --json
// command that reaches its document without a populated destination // command that reaches its document without a populated destination
// store: it reads the local index, streams `metadata/` (empty here), // store: it reads the local index, streams `metadata/` (empty here),
// and treats a barren destination as an empty list rather than a // and treats an empty destination directory as an empty list rather
// failure. // than a failure.
// //
// Not parallel: it replaces os.Args, os.Stdout and the xdg globals. // Not parallel: it replaces os.Args, os.Stdout and the xdg globals.
func TestEntryJSONStdoutIsExactlyOneDocument(t *testing.T) { func TestEntryJSONStdoutIsExactlyOneDocument(t *testing.T) {
dir := t.TempDir() dir := t.TempDir()
configPath := filepath.Join(dir, "config.yml") configPath := filepath.Join(dir, "config.yml")
storeDir := filepath.Join(dir, "store")
contents := fmt.Sprintf(hermeticConfig, contents := fmt.Sprintf(hermeticConfig,
filepath.Join(dir, "source"), filepath.Join(dir, "source"),
filepath.Join(dir, "store"), storeDir,
filepath.Join(dir, "index.sqlite")) filepath.Join(dir, "index.sqlite"))
require.NoError(t, require.NoError(t,
os.WriteFile(configPath, []byte(contents), configFileMode)) os.WriteFile(configPath, []byte(contents), configFileMode))
require.NoError(t, os.Mkdir(storeDir, 0o750))
// The PID lock lives under xdg.DataHome, which xdg resolves at // The PID lock lives under xdg.DataHome, which xdg resolves at
// package init; point it at the temp dir so the test neither // package init; point it at the temp dir so the test neither
// touches nor collides with the real one. // touches nor collides with the real one.
+10 -5
View File
@@ -98,25 +98,30 @@ func TestEntryPruneJSONStdoutIsExactlyOneDocument(t *testing.T) {
} }
} }
// writeHermeticPruneConfig builds a config over a temp directory and, if // writeHermeticPruneConfig builds a config over a temp directory with an
// seedStale is set, creates the index database up front with one // empty destination directory and, if seedStale is set, creates the
// snapshot record that has no counterpart on the destination store. // index database up front with one snapshot record that has no
// Returns the config path. // counterpart on the destination store. Returns the config path.
func writeHermeticPruneConfig(t *testing.T, seedStale bool) string { func writeHermeticPruneConfig(t *testing.T, seedStale bool) string {
t.Helper() t.Helper()
dir := t.TempDir() dir := t.TempDir()
configPath := filepath.Join(dir, "config.yml") configPath := filepath.Join(dir, "config.yml")
indexPath := filepath.Join(dir, "index.sqlite") indexPath := filepath.Join(dir, "index.sqlite")
storeDir := filepath.Join(dir, "store")
contents := fmt.Sprintf(hermeticConfig, contents := fmt.Sprintf(hermeticConfig,
filepath.Join(dir, "source"), filepath.Join(dir, "source"),
filepath.Join(dir, "store"), storeDir,
indexPath) indexPath)
require.NoError(t, require.NoError(t,
os.WriteFile(configPath, []byte(contents), configFileMode)) os.WriteFile(configPath, []byte(contents), configFileMode))
// prune fails on a destination directory that does not exist, so the
// empty store is created here rather than left to a first backup.
require.NoError(t, os.Mkdir(storeDir, 0o750))
// The PID lock lives under xdg.DataHome, which xdg resolves at // The PID lock lives under xdg.DataHome, which xdg resolves at
// package init; point it at the temp dir so the test neither // package init; point it at the temp dir so the test neither
// touches nor collides with the real one. // touches nor collides with the real one.
+39 -44
View File
@@ -42,6 +42,10 @@ var schemaFS embed.FS
// table itself. It is applied before the normal migration loop. // table itself. It is applied before the normal migration loop.
const bootstrapVersion = 0 const bootstrapVersion = 0
// busyTimeoutMs is how long a connection to the index waits for another
// connection's lock before failing with "database is locked".
const busyTimeoutMs = 10000
// DB represents the Vaultik local index database connection. // DB represents the Vaultik local index database connection.
// It uses SQLite to track file metadata, content-defined chunks, and blob associations. // It uses SQLite to track file metadata, content-defined chunks, and blob associations.
// The database enables incremental backups by detecting changed files and // The database enables incremental backups by detecting changed files and
@@ -94,10 +98,27 @@ func ParseMigrationVersion(filename string) (int, error) {
return version, nil return version, nil
} }
// indexDSN returns the driver DSN that opens the database at path. The
// driver runs each _pragma parameter on every connection it opens and drops
// any parameter it does not know without an error, so a setting written in
// another form silently does nothing. In WAL mode one connection can write
// while others read; the busy timeout makes a connection wait for a lock
// instead of failing at once.
func indexDSN(path string) string {
return fmt.Sprintf(
"%s?_pragma=busy_timeout(%d)&_pragma=journal_mode(WAL)"+
"&_pragma=synchronous(NORMAL)&_pragma=foreign_keys(1)",
path, busyTimeoutMs)
}
// New creates a new database connection at the specified path. // New creates a new database connection at the specified path.
// It creates the schema if needed and configures SQLite with WAL mode for // It creates the schema if needed. Every connection runs in WAL mode with
// better concurrency. SQLite handles crash recovery automatically when // a busy timeout and foreign keys on (see indexDSN), so a read-only command
// opening a database with journal/WAL files present. // can read the index while a backup writes to it. Committed rows can sit in
// the -wal file beside the database until a checkpoint, so a copy of the
// database file alone may miss them.
// SQLite handles crash recovery automatically when opening a database with
// journal/WAL files present.
// The path parameter can be a file path for persistent storage or ":memory:" // The path parameter can be a file path for persistent storage or ":memory:"
// for an in-memory database (useful for testing). // for an in-memory database (useful for testing).
func New(ctx context.Context, path string) (*DB, error) { func New(ctx context.Context, path string) (*DB, error) {
@@ -110,11 +131,7 @@ func New(ctx context.Context, path string) (*DB, error) {
// First attempt with standard WAL mode // First attempt with standard WAL mode
log.Debug("Attempting to open database with WAL mode", "path", path) log.Debug("Attempting to open database with WAL mode", "path", path)
conn, err := sql.Open( conn, err := sql.Open("sqlite", indexDSN(path))
"sqlite",
path+"?_journal_mode=WAL&_synchronous=NORMAL&_busy_timeout=10000"+
"&_locking_mode=NORMAL&_foreign_keys=ON",
)
if err == nil { if err == nil {
configureConnPool(conn) configureConnPool(conn)
@@ -134,8 +151,8 @@ func New(ctx context.Context, path string) (*DB, error) {
_ = conn.Close() _ = conn.Close()
} }
// If first attempt failed, try with TRUNCATE mode to clear any locks // If the first attempt failed, try once more
return openWithRecovery(ctx, path) return retryOpen(ctx, path)
} }
// configureConnPool serializes all database access through one connection. // configureConnPool serializes all database access through one connection.
@@ -147,18 +164,12 @@ func configureConnPool(conn *sql.DB) {
conn.SetMaxIdleConns(1) conn.SetMaxIdleConns(1)
} }
// finishOpen enables foreign keys, wraps the connection, and applies any // finishOpen wraps the connection and applies any pending migrations. On
// pending migrations. On migration failure the connection is closed. // migration failure the connection is closed.
func finishOpen(ctx context.Context, conn *sql.DB, path string) (*DB, error) { func finishOpen(ctx context.Context, conn *sql.DB, path string) (*DB, error) {
// Enable foreign keys explicitly
_, err := conn.ExecContext(ctx, "PRAGMA foreign_keys = ON")
if err != nil {
log.Warn("Failed to enable foreign keys", "path", path, "error", err)
}
db := &DB{conn: conn, path: path} db := &DB{conn: conn, path: path}
err = applyMigrations(ctx, conn) err := applyMigrations(ctx, conn)
if err != nil { if err != nil {
_ = conn.Close() _ = conn.Close()
@@ -168,21 +179,15 @@ func finishOpen(ctx context.Context, conn *sql.DB, path string) (*DB, error) {
return db, nil return db, nil
} }
// openWithRecovery retries opening the database in TRUNCATE journal mode to // retryOpen makes a second attempt to open the database, with the same
// clear stale locks, then switches back to WAL mode. // settings, after the first attempt failed, for example because another
func openWithRecovery(ctx context.Context, path string) (*DB, error) { // process held a lock for longer than the busy timeout.
log.Info( func retryOpen(ctx context.Context, path string) (*DB, error) {
"Database appears locked, attempting recovery with TRUNCATE mode", log.Info("Database appears locked, retrying open", "path", path)
"path", path,
)
conn, err := sql.Open( conn, err := sql.Open("sqlite", indexDSN(path))
"sqlite",
path+"?_journal_mode=TRUNCATE&_synchronous=NORMAL&_busy_timeout=10000"+
"&_foreign_keys=ON",
)
if err != nil { if err != nil {
return nil, fmt.Errorf("opening database in recovery mode: %w", err) return nil, fmt.Errorf("opening database on retry: %w", err)
} }
configureConnPool(conn) configureConnPool(conn)
@@ -190,28 +195,18 @@ func openWithRecovery(ctx context.Context, path string) (*DB, error) {
err = conn.PingContext(ctx) err = conn.PingContext(ctx)
if err != nil { if err != nil {
log.Debug( log.Debug(
"Failed to ping database in recovery mode, closing", "Failed to ping database on retry, closing",
"path", path, "error", err, "path", path, "error", err,
) )
_ = conn.Close() _ = conn.Close()
return nil, fmt.Errorf( return nil, fmt.Errorf(
"database still locked after recovery attempt: %w", "database still locked on retry: %w",
err, err,
) )
} }
log.Debug("Database opened in TRUNCATE mode", "path", path)
// Switch back to WAL mode
log.Debug("Switching database back to WAL mode", "path", path)
_, err = conn.ExecContext(ctx, "PRAGMA journal_mode=WAL")
if err != nil {
log.Warn("Failed to switch back to WAL mode", "path", path, "error", err)
}
db, err := finishOpen(ctx, conn, path) db, err := finishOpen(ctx, conn, path)
if err != nil { if err != nil {
return nil, err return nil, err
+83
View File
@@ -120,6 +120,89 @@ func TestDatabaseConcurrentAccess(t *testing.T) {
} }
} }
// TestNewSetsJournalModeAndBusyTimeout checks that the connection settings
// New passes reach SQLite. The driver drops a setting it does not recognise
// without an error, so only reading the value back shows it took effect.
func TestNewSetsJournalModeAndBusyTimeout(t *testing.T) {
t.Parallel()
ctx := context.Background()
db, err := New(ctx, filepath.Join(t.TempDir(), "index.db"))
if err != nil {
t.Fatalf("failed to create database: %v", err)
}
defer func() { _ = db.Close() }()
var journalMode string
err = db.conn.QueryRowContext(ctx, "PRAGMA journal_mode").Scan(&journalMode)
if err != nil {
t.Fatalf("reading journal_mode: %v", err)
}
if journalMode != "wal" {
t.Errorf("journal_mode = %q, want %q", journalMode, "wal")
}
var busyTimeout int
err = db.conn.QueryRowContext(ctx, "PRAGMA busy_timeout").Scan(&busyTimeout)
if err != nil {
t.Fatalf("reading busy_timeout: %v", err)
}
if busyTimeout != busyTimeoutMs {
t.Errorf("busy_timeout = %d, want %d", busyTimeout, busyTimeoutMs)
}
}
// TestNewWriteSucceedsWhileAnotherHandleReads opens the same index twice,
// as a read-only command does while a backup runs, and checks that a write
// on one handle commits while the other is in the middle of a read.
func TestNewWriteSucceedsWhileAnotherHandleReads(t *testing.T) {
t.Parallel()
ctx := context.Background()
dbPath := filepath.Join(t.TempDir(), "index.db")
reader, err := New(ctx, dbPath)
if err != nil {
t.Fatalf("failed to open reading handle: %v", err)
}
defer func() { _ = reader.Close() }()
writer, err := New(ctx, dbPath)
if err != nil {
t.Fatalf("failed to open writing handle: %v", err)
}
defer func() { _ = writer.Close() }()
// The read lock taken by the SELECT is held until the transaction ends.
readTx, err := reader.BeginTx(ctx, nil)
if err != nil {
t.Fatalf("beginning read transaction: %v", err)
}
defer func() { _ = readTx.Rollback() }()
var count int
err = readTx.QueryRowContext(ctx, "SELECT COUNT(*) FROM chunks").Scan(&count)
if err != nil {
t.Fatalf("reading chunks: %v", err)
}
_, err = writer.ExecWithLog(ctx,
"INSERT INTO chunks (chunk_hash, size) VALUES (?, ?)", "hash", 1024)
if err != nil {
t.Fatalf("write while another handle reads: %v", err)
}
}
func TestParseMigrationVersion(t *testing.T) { func TestParseMigrationVersion(t *testing.T) {
t.Parallel() t.Parallel()
+50
View File
@@ -266,6 +266,56 @@ func (r *FileRepository) ListByPrefix(
return files, rows.Err() return files, rows.Err()
} }
// ListIDsWithChunksNotInUploadedBlobs returns the IDs of the files whose
// path starts with prefix and that list at least one chunk held by no
// blob whose upload has completed (uploaded_ts set). A new snapshot
// cannot reference such a chunk, so a backup must not treat the file as
// unchanged even when its metadata matches the file on disk.
func (r *FileRepository) ListIDsWithChunksNotInUploadedBlobs(
ctx context.Context, prefix string,
) ([]types.FileID, error) {
query := `
SELECT DISTINCT f.id
FROM files f
JOIN file_chunks fc ON fc.file_id = f.id
WHERE f.path LIKE ? || '%'
AND NOT EXISTS (
SELECT 1
FROM blob_chunks bc
JOIN blobs b ON bc.blob_id = b.id
WHERE bc.chunk_hash = fc.chunk_hash
AND b.uploaded_ts IS NOT NULL
)
`
rows, err := r.db.conn.QueryContext(ctx, query, prefix)
if err != nil {
return nil, fmt.Errorf("querying files: %w", err)
}
defer func() {
err := rows.Close()
if err != nil {
Fatalf("failed to close rows: %v", err)
}
}()
var ids []types.FileID
for rows.Next() {
var id types.FileID
err := rows.Scan(&id)
if err != nil {
return nil, fmt.Errorf("scanning file ID: %w", err)
}
ids = append(ids, id)
}
return ids, rows.Err()
}
// ListAll returns all files in the database // ListAll returns all files in the database
func (r *FileRepository) ListAll(ctx context.Context) ([]*File, error) { func (r *FileRepository) ListAll(ctx context.Context) ([]*File, error) {
query := ` query := `
+18 -10
View File
@@ -3,6 +3,7 @@
package globals package globals
import ( import (
"regexp"
"strings" "strings"
"time" "time"
) )
@@ -10,12 +11,11 @@ import (
// Appname is the application name, populated from main(). // Appname is the application name, populated from main().
var Appname = "vaultik" //nolint:gochecknoglobals // set via -ldflags at build time var Appname = "vaultik" //nolint:gochecknoglobals // set via -ldflags at build time
// DevVersion is the version a binary reports when it was not built // DevVersion is the version a binary reports when it was built without
// from a tagged commit. script/version emits either this exact string // git metadata: script/version emits it outside a git checkout, and an
// (outside a git checkout) or this string followed by "-" and the // unstamped `go build` keeps it. goreleaser's snapshot template stamps
// commit it was built from, and goreleaser's snapshot template matches // it followed by "-" and the commit it was built from. It is
// that shape. It is deliberately not a number: a build that is not a // deliberately not a number.
// release must not name itself like one.
const DevVersion = "dev" const DevVersion = "dev"
// Version is the application version, populated from main(). // Version is the application version, populated from main().
@@ -60,9 +60,12 @@ func New() (*Globals, error) {
} }
// IsDevVersion reports whether v names a development build rather than // IsDevVersion reports whether v names a development build rather than
// a release. Both "dev" and "dev-<sha>" (and its "-dirty" variant) // a release. "dev" and goreleaser's snapshot "dev-<sha>" count, and so
// count: a caller that compares against "dev" exactly would treat every // does what `git describe --tags --always --dirty` gives a make or
// commit-stamped development build as a release. // docker build of an untagged commit: the bare short commit, or
// tag-N-gHASH on a commit after a tag. Any version ending in "-dirty"
// counts, a modified checkout of a tag ("v1.0.0-dirty") included.
// A plain tag such as "v1.0.0" or "1.0.0" is a release.
// //
// The empty string counts too. Nothing that knows its version reports // The empty string counts too. Nothing that knows its version reports
// no version, so an empty Version means the stamping failed, and the // no version, so an empty Version means the stamping failed, and the
@@ -71,7 +74,12 @@ func New() (*Globals, error) {
// case; this is the second line of defence, for a binary linked by // case; this is the second line of defence, for a binary linked by
// something other than the Makefile. // something other than the Makefile.
func IsDevVersion(v string) bool { func IsDevVersion(v string) bool {
return v == "" || v == DevVersion || strings.HasPrefix(v, DevVersion+"-") if v == "" || v == DevVersion || strings.HasPrefix(v, DevVersion+"-") ||
strings.HasSuffix(v, "-dirty") {
return true
}
return regexp.MustCompile(`^(.+-[0-9]+-g)?[0-9a-f]+$`).MatchString(v)
} }
// shortCommitLen is the number of commit-hash characters ShortCommit keeps. // shortCommitLen is the number of commit-hash characters ShortCommit keeps.
+16 -6
View File
@@ -34,10 +34,11 @@ func TestGlobalsNew(t *testing.T) {
} }
// TestIsDevVersion covers the boundary that matters: everything // TestIsDevVersion covers the boundary that matters: everything
// script/version and goreleaser's snapshot template can emit for an // script/version, a plain docker build and goreleaser's snapshot
// untagged build must be recognised as a development build, and a real // template can emit for an untagged build must be recognised as a
// tag must not be. A plain equality check against "dev" used to decide // development build, and a real tag must not be. A plain equality check
// this, which classified every commit-stamped dev build as a release. // against "dev" used to decide this, which classified every
// commit-stamped dev build as a release.
func TestIsDevVersion(t *testing.T) { func TestIsDevVersion(t *testing.T) {
t.Parallel() t.Parallel()
@@ -49,8 +50,17 @@ func TestIsDevVersion(t *testing.T) {
{"dev", true}, {"dev", true},
{"dev-b6e4a218a39e", true}, {"dev-b6e4a218a39e", true},
{"dev-b6e4a218a39e-dirty", true}, {"dev-b6e4a218a39e-dirty", true},
// What a tagged build produces (script/version strips the // What `git describe --tags --always --dirty` produces with no
// leading "v", matching goreleaser's .Version). // tag reachable, and on a commit after a tag.
{"877eb2f", true},
{"877eb2f-dirty", true},
{"v1.0.0-3-g877eb2f", true},
{"v1.0.0-3-g877eb2f-dirty", true},
{"1.0.0-rc.1-12-g877eb2f", true},
// A tagged commit with uncommitted changes is not that tag.
{"v1.0.0-dirty", true},
// What a tagged build produces (goreleaser's .Version strips
// the leading "v"; script/version keeps it).
{"1.0.0", false}, {"1.0.0", false},
{"0.1.0", false}, {"0.1.0", false},
{"1.0.0-rc.1", false}, {"1.0.0-rc.1", false},
+16 -8
View File
@@ -1,40 +1,48 @@
//nolint:testpackage // exercises the unexported copyFile helper //nolint:testpackage // exercises the unexported copyDatabase helper
package snapshot package snapshot
import ( import (
"context"
"os" "os"
"path/filepath" "path/filepath"
"syscall" "syscall"
"testing" "testing"
"github.com/spf13/afero" "github.com/spf13/afero"
"sneak.berlin/go/vaultik/internal/database"
) )
// TestCopyFileExportCopyMode verifies that the exported snapshot database // TestCopyDatabaseExportCopyMode verifies that the exported snapshot
// copy is created owner-only (0600), even under a lenient 022 umask that // database copy is created owner-only (0600), even under a lenient 022
// would otherwise leave a fresh file world-readable. // umask that would otherwise leave a fresh file world-readable.
// //
//nolint:paralleltest // syscall.Umask is process-global; parallel tests would clash //nolint:paralleltest // syscall.Umask is process-global; parallel tests would clash
func TestCopyFileExportCopyMode(t *testing.T) { func TestCopyDatabaseExportCopyMode(t *testing.T) {
restore := syscall.Umask(0o022) restore := syscall.Umask(0o022)
defer syscall.Umask(restore) defer syscall.Umask(restore)
ctx := context.Background()
dir := t.TempDir() dir := t.TempDir()
src := filepath.Join(dir, "index.sqlite") src := filepath.Join(dir, "index.sqlite")
err := os.WriteFile(src, []byte("index data"), 0o600) db, err := database.New(ctx, src)
if err != nil { if err != nil {
t.Fatalf("creating source index: %v", err) t.Fatalf("creating source index: %v", err)
} }
err = db.Close()
if err != nil {
t.Fatalf("closing source index: %v", err)
}
dst := filepath.Join(dir, "snapshot.db") dst := filepath.Join(dir, "snapshot.db")
sm := &SnapshotManager{fs: afero.NewOsFs()} sm := &SnapshotManager{fs: afero.NewOsFs()}
err = sm.copyFile(src, dst) err = sm.copyDatabase(ctx, src, dst)
if err != nil { if err != nil {
t.Fatalf("copyFile: %v", err) t.Fatalf("copyDatabase: %v", err)
} }
info, err := os.Stat(dst) info, err := os.Stat(dst)
+53 -21
View File
@@ -11,6 +11,7 @@ import (
"runtime" "runtime"
"strings" "strings"
"sync" "sync"
"syscall"
"time" "time"
"github.com/dustin/go-humanize" "github.com/dustin/go-humanize"
@@ -73,6 +74,11 @@ type Scanner struct {
knownChunks map[string]struct{} knownChunks map[string]struct{}
knownChunksMu sync.RWMutex knownChunksMu sync.RWMutex
// filesToRechunk holds the IDs of known files that list a chunk no
// uploaded blob holds; they are re-chunked even when their metadata
// is unchanged.
filesToRechunk map[types.FileID]struct{}
// Pending chunk hashes - chunks that have been added to packer but not // Pending chunk hashes - chunks that have been added to packer but not
// yet committed to DB. When a blob finalizes, the committed chunks are // yet committed to DB. When a blob finalizes, the committed chunks are
// removed from this set. // removed from this set.
@@ -325,9 +331,42 @@ func (s *Scanner) loadDatabaseState(
s.ui.Completef("Loaded %s known chunks from local index database.", s.ui.Completef("Loaded %s known chunks from local index database.",
s.ui.Count(len(s.knownChunks))) s.ui.Count(len(s.knownChunks)))
err = s.loadFilesToRechunk(ctx, path)
if err != nil {
return nil, fmt.Errorf("loading files to re-chunk: %w", err)
}
return knownFiles, nil return knownFiles, nil
} }
// loadFilesToRechunk loads the IDs of known files under path that list a
// chunk no uploaded blob holds. A file row is shared by every snapshot
// that lists the file and is updated in place when the file changes,
// while a blob row is deleted once no snapshot references it. Removing
// the only snapshot that references a changed file's current blob
// therefore leaves an older snapshot keeping a file row whose metadata
// matches the disk while no blob the local index records as uploaded
// holds its chunks, so a new snapshot cannot reference them. The dropped
// blob can still be in remote storage until prune removes it.
func (s *Scanner) loadFilesToRechunk(ctx context.Context, path string) error {
ids, err := s.repos.Files.ListIDsWithChunksNotInUploadedBlobs(ctx, path)
if err != nil {
return fmt.Errorf("listing files: %w", err)
}
s.filesToRechunk = make(map[types.FileID]struct{}, len(ids))
for _, id := range ids {
s.filesToRechunk[id] = struct{}{}
}
if len(ids) > 0 {
log.Info("Re-chunking known files whose chunks are not all "+
"in uploaded blobs", "files", len(ids))
}
return nil
}
// repairInterruptedBlobs discards blob rows left by a previous run whose // repairInterruptedBlobs discards blob rows left by a previous run whose
// upload never completed. Such a blob has its chunks, blob_chunks, and // upload never completed. Such a blob has its chunks, blob_chunks, and
// blobs rows committed to the local index before the upload is attempted, // blobs rows committed to the local index before the upload is attempted,
@@ -1104,12 +1143,9 @@ func (s *Scanner) buildSymlinkEntry(path string, info os.FileInfo) *database.Fil
} }
var uid, gid uint32 var uid, gid uint32
if stat, ok := info.Sys().(interface { if stat, ok := info.Sys().(*syscall.Stat_t); ok {
Uid() uint32 uid = stat.Uid
Gid() uint32 gid = stat.Gid
}); ok {
uid = stat.Uid()
gid = stat.Gid()
} }
return &database.File{ return &database.File{
@@ -1128,12 +1164,9 @@ func (s *Scanner) buildSymlinkEntry(path string, info os.FileInfo) *database.Fil
// buildDirectoryEntry creates a File record for a directory. // buildDirectoryEntry creates a File record for a directory.
func (s *Scanner) buildDirectoryEntry(path string, info os.FileInfo) *database.File { func (s *Scanner) buildDirectoryEntry(path string, info os.FileInfo) *database.File {
var uid, gid uint32 var uid, gid uint32
if stat, ok := info.Sys().(interface { if stat, ok := info.Sys().(*syscall.Stat_t); ok {
Uid() uint32 uid = stat.Uid
Gid() uint32 gid = stat.Gid
}); ok {
uid = stat.Uid()
gid = stat.Gid()
} }
return &database.File{ return &database.File{
@@ -1166,16 +1199,10 @@ func (s *Scanner) recordNonRegularFile(ctx context.Context, ftp *FileToProcess)
func (s *Scanner) checkFileInMemory( func (s *Scanner) checkFileInMemory(
path string, info os.FileInfo, knownFiles map[string]*database.File, path string, info os.FileInfo, knownFiles map[string]*database.File,
) (*database.File, bool) { ) (*database.File, bool) {
// Get file stats
stat, ok := info.Sys().(interface {
Uid() uint32
Gid() uint32
})
var uid, gid uint32 var uid, gid uint32
if ok { if stat, ok := info.Sys().(*syscall.Stat_t); ok {
uid = stat.Uid() uid = stat.Uid
gid = stat.Gid() gid = stat.Gid
} }
// Check against in-memory map first to get existing ID if available // Check against in-memory map first to get existing ID if available
@@ -1208,6 +1235,11 @@ func (s *Scanner) checkFileInMemory(
return file, true return file, true
} }
// No uploaded blob holds one of its chunks (see loadFilesToRechunk)
if _, rechunk := s.filesToRechunk[fileID]; rechunk {
return file, true
}
// Check if file has changed // Check if file has changed
if existingFile.Size != file.Size || if existingFile.Size != file.Size ||
existingFile.MTime.Unix() != file.MTime.Unix() || existingFile.MTime.Unix() != file.MTime.Unix() ||
+77
View File
@@ -307,3 +307,80 @@ func TestScannerLargeFile(t *testing.T) {
} }
} }
} }
// TestScannerRecordsOwnership backs up real files on disk and checks that
// the uid and gid of a file, a directory and a symlink are recorded.
// When the tests run as root both sides are 0, so only a run as another
// user can catch ownership recorded as 0.
func TestScannerRecordsOwnership(t *testing.T) {
t.Parallel()
sourceDir := t.TempDir()
filePath := filepath.Join(sourceDir, "file.txt")
dirPath := filepath.Join(sourceDir, "subdir")
linkPath := filepath.Join(sourceDir, "link")
err := os.WriteFile(filePath, []byte("owned"), 0o600)
if err != nil {
t.Fatal(err)
}
err = os.Mkdir(dirPath, 0o700)
if err != nil {
t.Fatal(err)
}
err = os.Symlink("file.txt", linkPath)
if err != nil {
t.Fatal(err)
}
db, err := database.NewTestDB()
if err != nil {
t.Fatalf("failed to create test database: %v", err)
}
defer func() {
err := db.Close()
if err != nil {
t.Errorf("failed to close database: %v", err)
}
}()
repos := database.NewRepositories(db)
scanner := snapshot.NewScanner(snapshot.ScannerConfig{
FS: afero.NewOsFs(),
ChunkSize: int64(1024 * 16),
Repositories: repos,
MaxBlobSize: int64(1024 * 1024),
CompressionLevel: 3,
AgeRecipients: []string{testAgePublicKey},
})
ctx := context.Background()
snapshotID := "test-snapshot-ownership"
createTestSnapshotRecord(ctx, t, repos, snapshotID)
_, err = scanner.Scan(ctx, sourceDir, snapshotID)
if err != nil {
t.Fatalf("scan failed: %v", err)
}
for _, path := range []string{filePath, dirPath, linkPath} {
file, err := repos.Files.GetByPath(ctx, path)
if err != nil {
t.Fatalf("failed to get %s: %v", path, err)
}
if file == nil {
t.Fatalf("%s was not recorded", path)
}
if int(file.UID) != os.Getuid() || int(file.GID) != os.Getgid() {
t.Errorf("%s recorded as uid %d gid %d, want uid %d gid %d",
path, file.UID, file.GID, os.Getuid(), os.Getgid())
}
}
}
+33 -38
View File
@@ -386,12 +386,12 @@ func (sm *SnapshotManager) prepareExportDB(
ctx context.Context, dbPath, snapshotID, tempDir string, ctx context.Context, dbPath, snapshotID, tempDir string,
) ([]byte, string, error) { ) ([]byte, string, error) {
// Step 1: Copy database to temp file // Step 1: Copy database to temp file
// The main database should be closed at this point // The main database is still open here, so it is copied through SQLite
tempDBPath := filepath.Join(tempDir, "snapshot.db") tempDBPath := filepath.Join(tempDir, "snapshot.db")
log.Debug("Copying database to temporary location", log.Debug("Copying database to temporary location",
"source", dbPath, "destination", tempDBPath) "source", dbPath, "destination", tempDBPath)
err := sm.copyFile(dbPath, tempDBPath) err := sm.copyDatabase(ctx, dbPath, tempDBPath)
if err != nil { if err != nil {
return nil, "", fmt.Errorf("copying database: %w", err) return nil, "", fmt.Errorf("copying database: %w", err)
} }
@@ -648,9 +648,10 @@ func (sm *SnapshotManager) collectCleanupStats(
// //
// VACUUM runs through the modernc.org/sqlite driver, on a freshly opened // VACUUM runs through the modernc.org/sqlite driver, on a freshly opened
// connection with no transaction in flight (VACUUM cannot run inside one). // connection with no transaction in flight (VACUUM cannot run inside one).
// The database opens in WAL mode, so VACUUM's rewrite lands in the WAL; the // database.New opens the file in WAL mode, so VACUUM's rewrite lands in the
// checkpoint on Close flushes it into the main file, which is the file we // -wal file. This is the only connection to the file, so closing it
// then compress and upload. // checkpoints the rewrite into the main file and removes the -wal file; the
// main file is the one compressFile then compresses and uploads.
func (sm *SnapshotManager) vacuumDatabase(ctx context.Context, dbPath string) error { func (sm *SnapshotManager) vacuumDatabase(ctx context.Context, dbPath string) error {
log.Debug("Running VACUUM on database", "path", dbPath) log.Debug("Running VACUUM on database", "path", dbPath)
@@ -744,26 +745,15 @@ func (sm *SnapshotManager) compressFile(inputPath, outputPath string) error {
// user; it holds the same private index data as the local index file. // user; it holds the same private index data as the local index file.
const exportCopyPerm = 0o600 const exportCopyPerm = 0o600
// copyFile copies a file from src to dst. The destination is the exported // copyDatabase copies the database at src to dst with VACUUM INTO. It reads
// snapshot database, so it is created owner-only rather than with the // through SQLite, so the copy holds rows committed to src that are still in
// umask-dependent default. // its -wal file, which a copy of the file alone would miss. The destination
func (sm *SnapshotManager) copyFile(src, dst string) error { // is the exported snapshot database, so it is created empty and owner-only
log.Debug("Opening source file for copy", "path", src) // first rather than with the umask-dependent default: VACUUM INTO writes
// into an existing empty file and keeps its mode.
sourceFile, err := sm.fs.Open(src) func (sm *SnapshotManager) copyDatabase(
if err != nil { ctx context.Context, src, dst string,
return err ) error {
}
defer func() {
log.Debug("Closing source file", "path", src)
err := sourceFile.Close()
if err != nil {
log.Debug("Failed to close source file", "path", src, "error", err)
}
}()
log.Debug("Creating destination file", "path", dst) log.Debug("Creating destination file", "path", dst)
destFile, err := sm.fs.OpenFile( destFile, err := sm.fs.OpenFile(
@@ -773,23 +763,28 @@ func (sm *SnapshotManager) copyFile(src, dst string) error {
return err return err
} }
defer func() { err = destFile.Close()
log.Debug("Closing destination file", "path", dst)
err := destFile.Close()
if err != nil {
log.Debug("Failed to close destination file", "path", dst, "error", err)
}
}()
log.Debug("Copying file data")
n, err := io.Copy(destFile, sourceFile)
if err != nil { if err != nil {
return err return err
} }
log.Debug("File copy complete", "bytes_copied", n) db, err := database.New(ctx, src)
if err != nil {
return fmt.Errorf("opening database to copy: %w", err)
}
defer func() {
cerr := db.Close()
if cerr != nil {
log.Debug("Failed to close database after copy",
"path", src, "error", cerr)
}
}()
_, err = db.ExecWithLog(ctx, "VACUUM INTO ?", dst)
if err != nil {
return fmt.Errorf("running VACUUM INTO: %w", err)
}
return nil return nil
} }
+71
View File
@@ -188,6 +188,77 @@ func TestVacuumDatabaseRemovesDeletedData(t *testing.T) {
} }
} }
// TestPrepareExportDBKeepsRowsCommittedToOpenIndex exports from an index
// that is still open, as a backup does. A row committed there can still be
// in the index's -wal file, and the export must hold it all the same.
func TestPrepareExportDBKeepsRowsCommittedToOpenIndex(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
ctx := context.Background()
fs := afero.NewOsFs()
dbPath := filepath.Join(t.TempDir(), "index.sqlite")
db, err := database.New(ctx, dbPath)
if err != nil {
t.Fatalf("failed to create database: %v", err)
}
defer func() { _ = db.Close() }()
repos := database.NewRepositories(db)
snapshot := &database.Snapshot{ID: "open-index-snapshot", Hostname: "test-host"}
err = repos.WithTx(ctx, func(ctx context.Context, tx *sql.Tx) error {
return repos.Snapshots.Create(ctx, tx, snapshot)
})
if err != nil {
t.Fatalf("failed to create snapshot: %v", err)
}
sm := &SnapshotManager{
config: &config.Config{
CompressionLevel: 3,
AgeRecipients: []string{testAgeRecipient},
},
fs: fs,
}
_, tempDBPath, err := sm.prepareExportDB(
ctx, dbPath, snapshot.ID.String(), t.TempDir())
if err != nil {
t.Fatalf("prepareExportDB failed: %v", err)
}
// Only the main database file is compressed and uploaded, so open a
// copy of that file alone.
uploadedPath := filepath.Join(t.TempDir(), "uploaded.db")
err = copyFile(fs, tempDBPath, uploadedPath)
if err != nil {
t.Fatalf("failed to copy exported database: %v", err)
}
exported, err := database.OpenReadOnly(ctx, uploadedPath)
if err != nil {
t.Fatalf("failed to open exported database: %v", err)
}
defer func() { _ = exported.Close() }()
got, err := database.NewRepositories(exported).Snapshots.GetByID(
ctx, snapshot.ID.String())
if err != nil {
t.Fatalf("failed to read snapshot from export: %v", err)
}
if got == nil {
t.Fatal("exported database is missing the snapshot row")
}
}
func TestCleanSnapshotDBEmptySnapshot(t *testing.T) { func TestCleanSnapshotDBEmptySnapshot(t *testing.T) {
// Initialize logger // Initialize logger
log.Initialize(log.Config{}) log.Initialize(log.Config{})
+23 -8
View File
@@ -22,11 +22,11 @@ type FileStorer struct {
// //
// Construction is intentionally cheap and does not touch the filesystem. // Construction is intentionally cheap and does not touch the filesystem.
// The basePath is recorded; the directory is created lazily on first // The basePath is recorded; the directory is created lazily on first
// write. Reads (Get/Stat/List) tolerate a missing basePath — a missing // write. Write operations (Put/PutWithProgress) call MkdirAll for the
// or unmounted destination during `snapshot list` should NOT block the
// command, it should degrade to "no remote snapshots reachable" with a
// warning. Write operations (Put/PutWithProgress) call MkdirAll for the
// per-blob parent directory, which also covers basePath on first use. // per-blob parent directory, which also covers basePath on first use.
// Get and Stat report a key under a missing basePath as ErrNotFound.
// List and ListStream fail on a missing basePath, because listing it as
// an empty store would make `prune` drop every local snapshot record.
// //
// Uses the real OS filesystem by default; call SetFilesystem to // Uses the real OS filesystem by default; call SetFilesystem to
// override for testing. // override for testing.
@@ -119,13 +119,20 @@ func (f *FileStorer) Delete(_ context.Context, key string) error {
return nil return nil
} }
// List returns all keys with the given prefix. // List returns all keys with the given prefix. It fails when the
// destination directory is missing; a missing prefix under it is an
// empty listing.
func (f *FileStorer) List(ctx context.Context, prefix string) ([]string, error) { func (f *FileStorer) List(ctx context.Context, prefix string) ([]string, error) {
var keys []string var keys []string
_, err := f.fs.Stat(f.basePath)
if err != nil {
return nil, fmt.Errorf("checking destination directory: %w", err)
}
basePath := f.fullPath(prefix) basePath := f.fullPath(prefix)
// Check if base path exists // Check if the prefix exists
exists, err := afero.Exists(f.fs, basePath) exists, err := afero.Exists(f.fs, basePath)
if err != nil { if err != nil {
return nil, fmt.Errorf("checking path: %w", err) return nil, fmt.Errorf("checking path: %w", err)
@@ -167,16 +174,24 @@ func (f *FileStorer) List(ctx context.Context, prefix string) ([]string, error)
return keys, nil return keys, nil
} }
// ListStream returns a channel of ObjectInfo for large result sets. // ListStream returns a channel of ObjectInfo for large result sets. Like
// List, it sends an error when the destination directory is missing.
func (f *FileStorer) ListStream(ctx context.Context, prefix string) <-chan ObjectInfo { func (f *FileStorer) ListStream(ctx context.Context, prefix string) <-chan ObjectInfo {
ch := make(chan ObjectInfo) ch := make(chan ObjectInfo)
go func() { go func() {
defer close(ch) defer close(ch)
_, err := f.fs.Stat(f.basePath)
if err != nil {
ch <- ObjectInfo{Err: fmt.Errorf("checking destination directory: %w", err)}
return
}
basePath := f.fullPath(prefix) basePath := f.fullPath(prefix)
// Check if base path exists // Check if the prefix exists
exists, err := afero.Exists(f.fs, basePath) exists, err := afero.Exists(f.fs, basePath)
if err != nil { if err != nil {
ch <- ObjectInfo{Err: fmt.Errorf("checking path: %w", err)} ch <- ObjectInfo{Err: fmt.Errorf("checking path: %w", err)}
+35
View File
@@ -1,6 +1,10 @@
package storage_test package storage_test
import ( import (
"context"
"errors"
"io/fs"
"path/filepath"
"testing" "testing"
"sneak.berlin/go/vaultik/internal/storage" "sneak.berlin/go/vaultik/internal/storage"
@@ -25,3 +29,34 @@ func TestFileStorer(t *testing.T) {
t.Parallel() t.Parallel()
runStorerConformance(t, newFileStorer) runStorerConformance(t, newFileStorer)
} }
// TestFileStorerListMissingDestination checks that List and ListStream
// fail when the destination directory does not exist, instead of
// reporting an empty store.
func TestFileStorerListMissingDestination(t *testing.T) {
t.Parallel()
ctx := context.Background()
s, err := storage.NewFileStorer(filepath.Join(t.TempDir(), "unmounted"))
if err != nil {
t.Fatalf("NewFileStorer: %v", err)
}
keys, err := s.List(ctx, "metadata/")
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("List = %v, %v; want a not-exist error", keys, err)
}
var streamErr error
for object := range s.ListStream(ctx, "metadata/") {
if object.Err != nil {
streamErr = object.Err
}
}
if !errors.Is(streamErr, fs.ErrNotExist) {
t.Errorf("ListStream error = %v, want a not-exist error", streamErr)
}
}
@@ -0,0 +1,234 @@
package vaultik_test
import (
"context"
"crypto/rand"
"path/filepath"
"slices"
"strings"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/chunker"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/storage/faultstore"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// These tests cover https://git.eeqj.de/sneak/vaultik/issues/214: once
// the only snapshot referencing a changed file's current blob is
// dropped, an older snapshot still keeps the file's row, and the next
// backup must re-chunk the file instead of treating it as unchanged.
//
// Each backup run uses its own snapshot name over the same source
// directory, so the snapshot IDs differ without waiting for their
// one-second timestamps to tick over. File rows are keyed by path, so
// the runs share them exactly as runs of one snapshot name would.
// changedFileConfig returns a full-backup config with the snapshot names
// "first", "second" and "third", all backing up dataDir.
func changedFileConfig(dataDir, dbPath string) *config.Config {
source := config.SnapshotConfig{Paths: []string{dataDir}}
cfg := faultTestConfig()
cfg.IndexPath = dbPath
cfg.ChunkSize = config.Size(faultChunkSize)
cfg.Snapshots = map[string]config.SnapshotConfig{
"first": source,
"second": source,
"third": source,
}
return cfg
}
// backUp runs a full backup of the named snapshot.
func backUp(v *vaultik.Vaultik, name string) error {
return v.CreateSnapshot(&vaultik.SnapshotCreateOptions{
Cron: true,
Snapshots: []string{name},
})
}
// appendedFileSize is the size of the file before the tests append to it.
// It is twice the largest chunk the chunker cuts, so the file's first
// chunk ends before the appended bytes and stays in a blob the first
// snapshot still references, while its new chunks are only in the blob
// that gets dropped.
const appendedFileSize = 2 * chunker.ChunkSizeSpread * faultChunkSize
// appendRandomBytes appends n random bytes to the file at path, creating
// the file if it does not exist, and records the new content in files.
// Random content gives the chunker real cut points.
func appendRandomBytes(
t *testing.T, fs afero.Fs, files map[string][]byte, path string, n int64,
) {
t.Helper()
added := make([]byte, n)
_, err := rand.Read(added)
require.NoError(t, err)
content := slices.Concat(files[path], added)
require.NoError(t, afero.WriteFile(fs, path, content, 0o644))
files[path] = content
}
// localSnapshotID returns the ID of the one snapshot in the local index
// that was backed up under name.
func localSnapshotID(
ctx context.Context, t *testing.T,
repos *database.Repositories, name string,
) string {
t.Helper()
snapshots, err := repos.Snapshots.ListRecent(ctx, listRecentTestLimit)
require.NoError(t, err)
var ids []string
for _, s := range snapshots {
if strings.Contains(s.ID.String(), "_"+name+"_") {
ids = append(ids, s.ID.String())
}
}
require.Lenf(t, ids, 1, "expected one local snapshot named %q", name)
return ids[0]
}
// assertThirdSnapshotRestores closes the local index, then restores the
// snapshot named "third" from store alone and byte-compares every file
// against files.
func assertThirdSnapshotRestores(
ctx context.Context, t *testing.T, cfg *config.Config,
store storage.Storer, repos *database.Repositories, db *database.DB,
fs afero.Fs, restoreDir string, files map[string][]byte,
) {
t.Helper()
id := localSnapshotID(ctx, t, repos, "third")
require.NoError(t, db.Close())
reader := newReaderVaultik(ctx, cfg, store, nil, fs)
require.NoError(t, reader.Restore(&vaultik.RestoreOptions{
SnapshotID: id,
TargetDir: restoreDir,
Verify: true,
}), "the backup after the drop must restore the changed file")
assertRestoredTree(t, fs, restoreDir, files)
}
// Trigger 1: bytes are appended to the file, a second snapshot backs it
// up, and that snapshot is removed. The first snapshot keeps the file row,
// which now lists the appended content's chunks, while removal drops the
// blob that held them.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestBackupAfterRemovingNewestSnapshotRestoresChangedFile(t *testing.T) {
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
storeDir := filepath.Join(tempDir, "remote")
restoreDir := filepath.Join(tempDir, "restored")
dbPath := filepath.Join(tempDir, "index.sqlite")
changedPath := filepath.Join(dataDir, "appended.bin")
ctx := context.Background()
files := writeFaultSourceTree(t, fs, dataDir)
cfg := changedFileConfig(dataDir, dbPath)
appendRandomBytes(t, fs, files, changedPath, appendedFileSize)
store, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
repos := database.NewRepositories(db)
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
require.NoError(t, backUp(v, "first"))
appendRandomBytes(t, fs, files, changedPath, faultChunkSize)
require.NoError(t, backUp(v, "second"))
_, err = v.RemoveSnapshot(localSnapshotID(ctx, t, repos, "second"),
&vaultik.RemoveOptions{Force: true})
require.NoError(t, err)
require.NoError(t, backUp(v, "third"))
assertThirdSnapshotRestores(
ctx, t, cfg, store, repos, db, fs, restoreDir, files)
}
// Trigger 2: bytes are appended to the file and the run that backs it up
// is interrupted at the manifest upload, after its blobs were uploaded.
// The next run's prune drops that incomplete snapshot and its blob, while
// the first snapshot keeps the file row, which now lists the appended
// content's chunks.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestBackupAfterInterruptedRunRestoresChangedFile(t *testing.T) {
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
storeDir := filepath.Join(tempDir, "remote")
restoreDir := filepath.Join(tempDir, "restored")
dbPath := filepath.Join(tempDir, "index.sqlite")
changedPath := filepath.Join(dataDir, "appended.bin")
ctx := context.Background()
files := writeFaultSourceTree(t, fs, dataDir)
cfg := changedFileConfig(dataDir, dbPath)
appendRandomBytes(t, fs, files, changedPath, appendedFileSize)
inner, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
failManifest := false
store := faultstore.New(inner)
store.OnPut = func(key string) faultstore.PutAction {
if failManifest && strings.HasSuffix(key, "manifest.json.zst") {
return faultstore.PutFail
}
return faultstore.PutNormal
}
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
repos := database.NewRepositories(db)
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
require.NoError(t, backUp(v, "first"))
appendRandomBytes(t, fs, files, changedPath, faultChunkSize)
failManifest = true
require.Error(t, backUp(v, "second"),
"the run must fail when its manifest upload fails")
failManifest = false
require.NoError(t, backUp(v, "third"))
assertThirdSnapshotRestores(
ctx, t, cfg, inner, repos, db, fs, restoreDir, files)
}
+22 -21
View File
@@ -301,6 +301,10 @@ func TestInterruptedBlobUploadRecordsNoUploadedBlob(t *testing.T) {
writeFaultSourceTree(t, fs, dataDir) writeFaultSourceTree(t, fs, dataDir)
// No upload succeeds, so nothing creates the destination directory.
// It must exist for the listing below to show that no blob survived.
require.NoError(t, os.Mkdir(storeDir, 0o750))
inner, err := storage.NewFileStorer(storeDir) inner, err := storage.NewFileStorer(storeDir)
require.NoError(t, err) require.NoError(t, err)
@@ -680,18 +684,10 @@ func faultScannerFactory(
// Scenario 5: the restore target runs out of space mid-file. Restore // Scenario 5: the restore target runs out of space mid-file. Restore
// must fail with an out-of-space error, and must not leave a truncated // must fail with an out-of-space error, and must not leave a truncated
// file at the target path presenting as a complete restore. Restore // file at the target path presenting as a complete restore.
// today writes each file straight to its final path and does not remove
// it when a write fails, so the truncated file survives; deleting it is
// tracked by https://git.eeqj.de/sneak/vaultik/issues/163. Skipped until
// that lands, so the destination assertion below is recorded rather than
// dropped.
// //
//nolint:paralleltest // installs the global logger via log.Initialize //nolint:paralleltest // installs the global logger via log.Initialize
func TestRestoreReportsDiskFull(t *testing.T) { func TestRestoreReportsDiskFull(t *testing.T) {
t.Skip("blocked on https://git.eeqj.de/sneak/vaultik/issues/163: " +
"a disk-full write leaves a truncated file at the target path " +
"instead of removing it")
log.Initialize(log.Config{}) log.Initialize(log.Config{})
osFS := afero.NewOsFs() osFS := afero.NewOsFs()
@@ -717,10 +713,10 @@ func TestRestoreReportsDiskFull(t *testing.T) {
id := fullFaultBackup(ctx, t, osFS, inner, cfg, repos, dataDir, dbPath, "diskfull") id := fullFaultBackup(ctx, t, osFS, inner, cfg, repos, dataDir, dbPath, "diskfull")
require.NoError(t, db.Close()) require.NoError(t, db.Close())
// Restore onto a filesystem that allows only a few bytes of file // Restore onto a target that allows only a few bytes of file content:
// content: enough to create files, far too little to hold them. // enough to create files, far too little to hold them.
budget := int64(8) budget := int64(8)
quota := &quotaFS{Fs: osFS, remaining: &budget} quota := &quotaFS{Fs: osFS, dir: restoreDir, remaining: &budget}
v := newReaderVaultik(ctx, cfg, inner, nil, quota) v := newReaderVaultik(ctx, cfg, inner, nil, quota)
err = v.Restore(&vaultik.RestoreOptions{SnapshotID: id, TargetDir: restoreDir}) err = v.Restore(&vaultik.RestoreOptions{SnapshotID: id, TargetDir: restoreDir})
@@ -754,21 +750,26 @@ func assertRestoredTree(
// budget is exhausted, mirroring a real ENOSPC. // budget is exhausted, mirroring a real ENOSPC.
var errNoSpace = errors.New("no space left on device") var errNoSpace = errors.New("no space left on device")
// quotaFS is an afero.Fs whose files may write only a fixed total number // quotaFS is an afero.Fs on which files opened under dir may write only
// of content bytes before failing, simulating a full restore target. It // a fixed total number of content bytes before failing, simulating a full
// wraps the interface so every method except Create delegates to the // restore target. Every other method, and every file outside dir, goes
// real filesystem; only file writes are capped. // straight to the real filesystem: restore also writes the decrypted
// metadata database under $TMPDIR through this filesystem, and capping
// that would fail the restore before it wrote anything to the target.
type quotaFS struct { type quotaFS struct {
afero.Fs afero.Fs
dir string
remaining *int64 remaining *int64
} }
//nolint:ireturn // afero.Fs.Create's signature requires returning afero.File. //nolint:ireturn // afero.Fs.OpenFile's signature requires returning afero.File.
func (q *quotaFS) Create(name string) (afero.File, error) { func (q *quotaFS) OpenFile(
f, err := q.Fs.Create(name) name string, flag int, perm os.FileMode,
if err != nil { ) (afero.File, error) {
return nil, err f, err := q.Fs.OpenFile(name, flag, perm)
if err != nil || !strings.HasPrefix(name, q.dir) {
return f, err
} }
return &quotaFile{File: f, remaining: q.remaining}, nil return &quotaFile{File: f, remaining: q.remaining}, nil
+15 -5
View File
@@ -140,24 +140,34 @@ func parseSnapshotName(snapshotID string) string {
// parseDuration parses a duration string with support for human-friendly units: // parseDuration parses a duration string with support for human-friendly units:
// d/day/days, w/week/weeks, mo/month/months, y/year/years, plus standard Go // d/day/days, w/week/weeks, mo/month/months, y/year/years, plus standard Go
// duration units. Following Go, m is minutes and mo is months. A bare number, // duration units. Following Go, m is minutes and mo is months. A bare number,
// an unknown unit, and a negative value are all rejected. // an unknown unit, and a negative value are all rejected. Outside the Go
// units the input must be whole numbers each followed directly by a unit
// (2w3d), so 1.0y and 30 days are errors; decimals work only in Go units
// (1.5h).
func parseDuration(s string) (time.Duration, error) { func parseDuration(s string) (time.Duration, error) {
if strings.HasPrefix(strings.TrimSpace(s), "-") { if strings.HasPrefix(strings.TrimSpace(s), "-") {
return 0, errNegativeDuration return 0, errNegativeDuration
} }
// A bare number has no unit, but time.ParseDuration accepts 0, which
// would put the cutoff at now and select every snapshot.
_, err := strconv.ParseFloat(s, 64)
if err == nil {
return 0, fmt.Errorf("%w: %q", errInvalidDuration, s)
}
d, err := time.ParseDuration(s) d, err := time.ParseDuration(s)
if err == nil { if err == nil {
return d, nil return d, nil
} }
re := regexp.MustCompile(`(\d+)\s*([a-zA-Z]+)`) if !regexp.MustCompile(`^(\d+[a-zA-Z]+)+$`).MatchString(s) {
matches := re.FindAllStringSubmatch(s, -1)
if len(matches) == 0 {
return 0, fmt.Errorf("%w: %q", errInvalidDuration, s) return 0, fmt.Errorf("%w: %q", errInvalidDuration, s)
} }
re := regexp.MustCompile(`(\d+)([a-zA-Z]+)`)
matches := re.FindAllStringSubmatch(s, -1)
var total time.Duration var total time.Duration
for _, match := range matches { for _, match := range matches {
+11
View File
@@ -59,6 +59,7 @@ func TestParseDuration(t *testing.T) {
{"30s", 30 * time.Second, false}, {"30s", 30 * time.Second, false},
{"6m", 6 * time.Minute, false}, {"6m", 6 * time.Minute, false},
{"1h", time.Hour, false}, {"1h", time.Hour, false},
{"1.5h", 90 * time.Minute, false},
// Extended calendar units. // Extended calendar units.
{"30d", 30 * 24 * time.Hour, false}, {"30d", 30 * 24 * time.Hour, false},
{"3days", 3 * 24 * time.Hour, false}, {"3days", 3 * 24 * time.Hour, false},
@@ -73,11 +74,21 @@ func TestParseDuration(t *testing.T) {
{"1y6mo", 365*24*time.Hour + 180*24*time.Hour, false}, {"1y6mo", 365*24*time.Hour + 180*24*time.Hour, false},
// Rejected inputs. // Rejected inputs.
{"6", 0, true}, // bare number, no unit {"6", 0, true}, // bare number, no unit
{"0", 0, true}, // bare number; time.ParseDuration accepts it
{"+0", 0, true}, // bare number; time.ParseDuration accepts it
{"5x", 0, true}, // unknown unit {"5x", 0, true}, // unknown unit
{"-5d", 0, true}, // negative, extended unit {"-5d", 0, true}, // negative, extended unit
{"-5h", 0, true}, // negative, Go unit {"-5h", 0, true}, // negative, Go unit
{"", 0, true}, // empty {"", 0, true}, // empty
{"garbage", 0, true}, {"garbage", 0, true},
// Characters outside the number-and-unit parts must not be
// skipped: 1.0y read as 0y would select every snapshot.
{"1.0y", 0, true},
{"2.1w", 0, true},
{"1,5y", 0, true},
{"30 days", 0, true},
{"30 days ago", 0, true},
{"x7d", 0, true},
} }
for _, tt := range tests { for _, tt := range tests {
@@ -0,0 +1,155 @@
package vaultik_test
import (
"bytes"
"context"
"io/fs"
"os"
"path/filepath"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/ui"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// These tests cover https://git.eeqj.de/sneak/vaultik/issues/220: a
// file:// destination whose directory is missing, such as a USB stick
// that is not plugged in, cannot be listed. It is not an empty store, so
// no command may conclude from it that the local snapshots are gone.
// backUpToFileDestination backs up the snapshot named "first" to a
// file:// destination at storeDir, which need not exist yet. Everything
// the returned Vaultik prints after the backup goes to the returned
// buffer.
func backUpToFileDestination(
ctx context.Context, t *testing.T, storeDir string,
) (*vaultik.Vaultik, *database.Repositories, *bytes.Buffer) {
t.Helper()
osFs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
dbPath := filepath.Join(tempDir, "index.sqlite")
writeFaultSourceTree(t, osFs, dataDir)
store, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
t.Cleanup(func() { _ = db.Close() })
repos := database.NewRepositories(db)
cfg := changedFileConfig(dataDir, dbPath)
v := newBackupVaultik(ctx, cfg, store, repos, db, osFs)
require.NoError(t, backUp(v, "first"))
out := &bytes.Buffer{}
v.Stdout = out
v.UI = ui.NewWithColor(out, false)
return v, repos, out
}
// backUpThenUnplug backs up to a file:// destination, then moves the
// destination directory away, as unplugging the volume it lives on would.
func backUpThenUnplug(
ctx context.Context, t *testing.T,
) (*vaultik.Vaultik, *database.Repositories, *bytes.Buffer) {
t.Helper()
storeDir := filepath.Join(t.TempDir(), "usbstick")
v, repos, out := backUpToFileDestination(ctx, t, storeDir)
require.NoError(t, os.Rename(storeDir, storeDir+"-unplugged"))
return v, repos, out
}
// TestFirstBackupCreatesDestinationDirectory checks that a first backup
// to a destination directory that does not exist yet creates it, and
// that the destination can be listed afterwards.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestFirstBackupCreatesDestinationDirectory(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
storeDir := filepath.Join(t.TempDir(), "volume", "backup")
v, _, out := backUpToFileDestination(ctx, t, storeDir)
require.NoError(t, v.ListSnapshots(false))
assert.NotContains(t, out.String(), "Could not list backup destination store")
assert.NotContains(t, out.String(), "not found in backup destination store")
}
// TestListSnapshotsWarnsWhenDestinationMissing checks that snapshot list
// warns and shows the local index alone, without reporting the local
// snapshot as missing from the destination.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestListSnapshotsWarnsWhenDestinationMissing(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
v, repos, out := backUpThenUnplug(ctx, t)
id := localSnapshotID(ctx, t, repos, "first")
require.NoError(t, v.ListSnapshots(false))
assert.Contains(t, out.String(), "Could not list backup destination store")
assert.Contains(t, out.String(), "Showing snapshots from the local index only.")
assert.Contains(t, out.String(), id)
assert.NotContains(t, out.String(), "not found in backup destination store")
}
// TestRemoveSnapshotWarnsWhenDestinationMissing checks that snapshot
// remove warns that the metadata could not be removed from the
// destination, instead of reporting that it was.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestRemoveSnapshotWarnsWhenDestinationMissing(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
v, repos, out := backUpThenUnplug(ctx, t)
result, err := v.RemoveSnapshot(localSnapshotID(ctx, t, repos, "first"),
&vaultik.RemoveOptions{Force: true})
require.NoError(t, err)
assert.False(t, result.RemoteRemoved)
assert.Contains(t, out.String(),
"Could not remove snapshot metadata from remote")
assert.NotContains(t, out.String(),
"Removed snapshot metadata from remote storage")
}
// TestPruneKeepsLocalRecordsWhenDestinationMissing checks that prune
// fails on a destination it cannot list and deletes no local snapshot
// record.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestPruneKeepsLocalRecordsWhenDestinationMissing(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
v, repos, _ := backUpThenUnplug(ctx, t)
err := v.Prune(&vaultik.PruneOptions{Force: true})
require.ErrorIs(t, err, fs.ErrNotExist)
require.ErrorContains(t, err, "listing remote snapshots")
snapshots, err := repos.Snapshots.ListRecent(ctx, listRecentTestLimit)
require.NoError(t, err)
assert.Len(t, snapshots, 1, "prune must delete no local snapshot record")
}
+9 -9
View File
@@ -24,9 +24,11 @@ main() {
# value is what forces those layers to re-run: without it an # value is what forces those layers to re-run: without it an
# unchanged tree replays them from cache, the checks never execute, # unchanged tree replays them from cache, the checks never execute,
# and the build still exits 0. Each ARG sits immediately above the # and the build still exits 0. Each ARG sits immediately above the
# check RUNs, so dependency and module layers still cache. Both # check RUNs, so dependency and module layers still cache.
# Dockerfiles also refuse to build at all when CHECK_EPOCH is empty, # Dockerfile.lint also refuses to build at all when CHECK_EPOCH is
# so a missing value fails loudly here rather than passing quietly. # empty, so a missing value fails its build loudly rather than passing
# quietly; the product Dockerfile does not, because a plain `docker
# build .` must succeed.
# #
# The value must be unique per invocation, not per second. `date +%s` # The value must be unique per invocation, not per second. `date +%s`
# is second-granular, so two concurrent invocations in the same # is second-granular, so two concurrent invocations in the same
@@ -57,12 +59,10 @@ main() {
--build-arg CHECK_EPOCH="$epoch" -f Dockerfile.lint . --build-arg CHECK_EPOCH="$epoch" -f Dockerfile.lint .
# Version, commit and build date are computed here on the host, the # Version, commit and build date are computed here on the host, the
# same way script/docker does, and passed into the product build so # same way script/docker does, and passed into the product build,
# the CI-built image reports its real source. The build context # where they take precedence over what the build would derive from
# excludes .git (see .dockerignore), so the build cannot derive them # the .git in its context. VERSION comes from script/version, as in
# itself; without these it would stamp the Dockerfile's dev/unknown # the Makefile.
# fallbacks. VERSION comes from script/version, the source of truth
# shared with the Makefile.
version="$("$ROOT/script/version")" version="$("$ROOT/script/version")"
commit="$(git rev-parse HEAD 2>/dev/null || echo unknown)" commit="$(git rev-parse HEAD 2>/dev/null || echo unknown)"
commit_date="$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)" commit_date="$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)"
+10 -13
View File
@@ -1,7 +1,8 @@
#!/bin/sh #!/bin/sh
# script/docker: build the Docker image tagged with the project name. # script/docker: build the Docker image tagged with the project name.
# Identical in all repos; the tag comes from script/projectname. # The tag comes from script/projectname. Unlike the canonical copy in
# Generic: needs no adaptation. # sneak/prompts, it passes a fresh CHECK_EPOCH instead of --no-cache, and
# COMMIT and COMMIT_DATE as well as VERSION.
# #
# This builds the PRODUCT image only, and the product Dockerfile has no # This builds the PRODUCT image only, and the product Dockerfile has no
# lint stage: linting lives in Dockerfile.lint and is run by # lint stage: linting lives in Dockerfile.lint and is run by
@@ -21,19 +22,15 @@ main() {
# comments there. This script is not the CI gate, but a local build # comments there. This script is not the CI gate, but a local build
# is almost always warm, so without this it would report a green the # is almost always warm, so without this it would report a green the
# tree had not earned and the two entrypoints would disagree about # tree had not earned and the two entrypoints would disagree about
# whether the tree is clean. The Dockerfile now refuses to build # whether the tree is clean.
# without a non-empty value, so this is required, not optional.
epoch="$(date +%s%N)$$" epoch="$(date +%s%N)$$"
# Version, commit and build date are computed here on the host, # Version, commit and build date are computed here on the host and
# where .git exists, and passed into the build. The build context # passed into the build, where they take precedence over what the
# excludes .git (see .dockerignore), so the container cannot derive # build would derive from the .git in its context. VERSION comes
# them itself -- it used to try and always got "unknown", giving # from script/version, as in the Makefile, so the image reports the
# every image a "commit: unknown" it could not be traced from. # same string, -dirty included, that a local build of the same tree
# VERSION comes from script/version, the source of truth shared with # would.
# the Makefile, so a Docker build reports the same string (tag,
# dev-<sha>, or a -dirty variant) that a local build of the same
# tree would.
version="$("$SCRIPT_DIR/version")" version="$("$SCRIPT_DIR/version")"
commit="$(git rev-parse HEAD 2>/dev/null || echo unknown)" commit="$(git rev-parse HEAD 2>/dev/null || echo unknown)"
commit_date="$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)" commit_date="$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)"
+16 -60
View File
@@ -1,73 +1,29 @@
#!/bin/sh #!/bin/sh
# script/version: output the version string to bake into the binary. # script/version: output the version string to bake into the binary.
# Our own extension to scripts-to-rule-them-all, and the single source # Our own extension to scripts-to-rule-them-all. The Makefile's LDFLAGS
# of truth for the version: the Makefile's LDFLAGS call this rather # call this rather than carrying a hardcoded constant, which is what used
# than carrying a hardcoded constant, which is what used to make every # to make every local build claim to be 1.0.0-rc.1 regardless of git
# local build claim to be 1.0.0-rc.1 regardless of git state. # state. script/docker and script/cibuild pass its output to the image
# build; given no version, the image build runs `git describe --tags
# --always` itself, without `--dirty`.
# #
# The rules, in order: # The version is `git describe --tags --always --dirty`: the tag on a
# tagged commit, tag-N-gHASH on a commit after one, the short commit
# when no tag is reachable, each with a "-dirty" suffix when tracked
# files have uncommitted changes. Untracked files are ignored: a stray
# scratch file does not change what was compiled. Outside a git checkout
# (release tarball, `go install`), or in one with no commits, it is
# "dev".
# #
# HEAD is exactly on an annotated or lightweight tag # Nothing here ever invents a version number.
# -> that tag, with a leading "v" stripped
# anything else
# -> "dev-<12 chars of HEAD>"
# not a git checkout at all (release tarball, `go install`)
# -> "dev"
#
# Either of the first two gains a "-dirty" suffix when tracked files
# have uncommitted changes, because a modified checkout of v1.0.0 is
# not v1.0.0. Untracked files are ignored, matching `git describe
# --dirty`: a stray scratch file does not change what was compiled.
#
# The "v" is stripped so that a `make` build and a goreleaser build of
# the same tagged commit report the *same* string: goreleaser's
# {{ .Version }} is the tag without the prefix, and the release archive
# names are built from it. A tag named `v1.0.0` therefore produces
# `vaultik 1.0.0`, matching `vaultik_1.0.0_linux_amd64.tar.gz`.
#
# Nothing here ever invents a version number. An untagged build says so
# and names the commit it was built from; it does not round up to the
# nearest plausible release.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Length of the commit prefix in a dev version. Matches
# globals.ShortCommit, so `vaultik version` shows the same 12 chars in
# its version line and its commit line.
SHORT_LEN=12
main() { main() {
cd "$ROOT" cd "$ROOT"
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
if ! git rev-parse --git-dir >/dev/null 2>&1; then echo "${version:-dev}"
echo "dev"
return 0
fi
dirty=""
if [ -n "$(git status --porcelain --untracked-files=no 2>/dev/null)" ]; then
dirty="-dirty"
fi
# --exact-match so a *descendant* of a tag is not reported as that
# tag. Plain `git describe --tags` would call a commit 40 patches
# past v1.0.0 "v1.0.0-40-gabc1234", and the leading token of that is
# a released version the build is not.
tag="$(git describe --tags --exact-match HEAD 2>/dev/null || true)"
if [ -n "$tag" ]; then
echo "${tag#v}${dirty}"
return 0
fi
sha="$(git rev-parse "--short=$SHORT_LEN" HEAD 2>/dev/null || true)"
if [ -z "$sha" ]; then
# A repo with no commits at all.
echo "dev"
return 0
fi
echo "dev-${sha}${dirty}"
} }
main "$@" main "$@"