Author SHA1 Message Date
sneak f131e59d56 Join the S3 prefix to every key with one slash (closes #222)
check / check (push) Waiting to run
The S3 client built each key as prefix + key, and the URL parser keeps
the prefix as written, so s3://bucket/p stored p + "blobs/..." with no
slash while s3://bucket/p/ stored p/blobs/.... A recovery host that wrote
the URL the other way found no snapshots.

NewClient now strips trailing slashes from the prefix and adds one back
when anything is left, giving the README layout for both URL forms; an
empty prefix stays at the bucket root. The s3.prefix config setting
goes through the same client and gets the same join.

A new test writes and lists through each URL shape against an
in-process S3 server and checks the keys in the bucket.

Model: opus-5-5
2026-10-06 15:35:06 +00:00
clawbot 315b6483b8 Mark github.com/spf13/pflag as a direct dependency in go.mod (closes #246)
check / check (push) Waiting to run
internal/cli/snapshot_restore_test.go imports github.com/spf13/pflag
directly, but go.mod still marked it // indirect. script/precommit runs
go mod tidy and fails when that changes go.mod, so the pre-commit hook
stopped every commit. This is the go mod tidy output: pflag moves to the
direct require block, and go.sum does not change. script/cibuild does
not run the tidy, which is why the gate stayed green.

Model: opus-5-5
2026-10-06 16:59:26 +02:00
clawbot 14fc4c9893 Load a config that has no age recipient (closes #221)
check / check (push) Waiting to run
The README's steps for restoring on another machine failed at the first
command: `config init` wrote a placeholder recipient, and `config.Load`
rejects any recipient that does not parse. `config init` now writes an
empty `age_recipients` list, `config.Load` accepts an empty list, and
`snapshot create` refuses to start without a recipient. A malformed
recipient is still rejected at load.

The recovery-host test now builds its config with `config init` and
`config set` and reads it through `config.Load`, so it imports
`internal/cli`. On a fresh file, `config set age_recipients.0` writes the
list in flow style (`[age1...]`).

Model: opus-5-5
2026-10-06 15:46:13 +02:00
clawbot 66c80a70e3 Skip files of an unreadable blob under restore --skip-errors (closes #218)
check / check (push) Waiting to run
A blob that failed to download ended `snapshot restore` even with
--skip-errors, after restoring whichever files came first. The download
error now goes through the same per-file handling as any other restore
error, once for every pending file that references the blob. With
--skip-errors those files are reported as failed, the rest are
restored, and the command still exits non-zero. Without the flag the
restore still aborts; the error now also names one affected file and
suggests --skip-errors. A cancelled restore still ends at once. The
--skip-errors help text and README now limit the packing and storage
caveat to snapshot creation.

Model: opus-5-5
2026-10-06 14:46:05 +02:00
clawbot 81f83b29f6 Re-vendor the canonical files from sneak/prompts at dd4027b (closes #213)
check / check (push) Waiting to run
Linting and testing become the lint and test phases of the Dockerfile,
and the build stage depends on both. Dockerfile.lint, CHECK_EPOCH and
the tests that checked them are removed. Every docker build in script/
passes --no-cache, and script/cibuild runs script/bootstrap first. A
host without Go gets the go.mod version from script/install-go in
.tool/go, which bootstrap, the Makefile, fmt, fmt-check, precommit and
release add to PATH; fmt-check skips .tool. The image takes its version
from the VERSION build arg or git describe, dev without .git. This
repo's own entries follow the canonical content in .gitignore and
.editorconfig. The golangci-lint v2.14.0 findings are fixed. The rules
in CLAUDE.md move into AGENTS.md. IsDevVersion counts "unknown".

Model: opus-5-5
2026-10-06 12:46:13 +02:00
clawbot c4adb72d80 Run the local index in WAL mode with a busy timeout (closes #217)
check / check (pull_request) Waiting to run
check / check (push) In progress
The connection settings were passed as `_journal_mode=`-style
parameters, which the SQLite driver drops without an error, so the
index ran in rollback-journal mode with no busy timeout. `snapshot
list` or `info` reading during a backup could make the backup's next
write fail with "database is locked". Both open paths now pass
`_pragma=` parameters; foreign keys moved there too.

With WAL on, rows committed to the open index can still be in the
-wal file, which a copy of the main file misses. The metadata export
now copies the index with VACUUM INTO, into an empty 0600 file.

The retry after a failed open no longer claims a TRUNCATE recovery; it
retries with the same settings.

Model: opus-5-5
2026-10-06 11:29:18 +02:00
clawbot 4a167e153a Record the real uid and gid of backed-up files (closes #216)
check / check (push) Successful in 4m58s
check / check (pull_request) Successful in 6m21s
The scanner read uid and gid by asserting the stat result to an
interface with Uid() and Gid() methods. *syscall.Stat_t has Uid and Gid
fields, not methods, so the assertion never matched and every file,
directory and symlink was stored as 0:0; a restore as root then gave
everything to root. The scanner now reads the fields of
*syscall.Stat_t.

The first backup after this change re-reads every file not owned by
root, because its stored uid and gid no longer match the disk.

When the tests run as root, as in the Docker build, the new test
compares 0 with 0 and cannot catch the defect; a non-root run does.

Model: opus-5-5
2026-10-06 09:46:16 +02:00
clawbot ea72697992 List a missing file:// destination directory as an error (closes #220)
check / check (push) Successful in 5m53s
check / check (pull_request) Successful in 4m53s
The file backend listed a destination directory that does not exist as
an empty store. With the volume unplugged, snapshot list reported every
local snapshot as missing from the store, snapshot remove said it had
removed metadata it never reached, and prune dropped every local
snapshot record. List and ListStream now fail when the destination
directory is missing, so those commands take their existing path for a
store that cannot be listed. A missing prefix under an existing
directory is still an empty listing, and a first backup still creates
the directory.

Three tests listed a file:// destination nothing had created; they now
create it.

Model: opus-5-5
2026-10-06 08:46:17 +02:00
clawbot 713be502bd Reject a duration with characters outside its parts (closes #215)
check / check (push) Successful in 7m28s
check / check (pull_request) Successful in 5m39s
parseDuration fell back to an unanchored search for number-and-unit
pieces when time.ParseDuration failed, and skipped everything in
between. 1.0y became 0, so `snapshot create --prune --keep-newer-than
1.0y` deleted every snapshot of the backed-up names, the new one
included. 2.1w became one week and 1,5y five years. The fallback now
requires the whole input to be whole-number-and-unit parts with nothing
between them. A bare number is rejected before time.ParseDuration sees
it, since Go reads 0 and +0 as zero with no unit.

Judgement call: a space between number and unit (`30 days`) was
accepted and is now an error, matching Go's own units.

Model: opus-5-5
2026-10-06 06:12:07 +02:00
clawbot 35cf985c18 Re-chunk a known file whose chunks no uploaded blob holds (closes #214)
check / check (push) Successful in 6m53s
check / check (pull_request) Successful in 6m17s
File rows are shared by every snapshot and updated in place, while a
blob row is deleted once no snapshot references it. Removing the newest
snapshot, or the prune after an interrupted run, could drop the only
blob holding a changed file's current chunks while an older snapshot
kept the file row. The next backup compared metadata only, skipped the
file, and completed a snapshot that could not restore it.

The scanner now loads the IDs of known files that list a chunk no
uploaded blob holds and re-chunks them even when their metadata is
unchanged.

The tests append to a file, so the file keeps its first chunk in a blob
the first snapshot still references. Each backup run gets its own
snapshot name, so the second-precision snapshot IDs differ without
sleeping.

Model: opus-5-5
2026-10-06 04:46:16 +02:00
clawbot 070090124a Stamp the tag or short commit in a plain docker build (closes #211)
check / check (push) Successful in 3m35s
check / check (pull_request) Successful in 3m16s
A plain `docker build .` stamped `dev`: `.dockerignore` left out `.git`
and the Dockerfile defaulted VERSION to `dev`. `.dockerignore` now
sends `.git` without `.git/config`. Given no build arguments, the
builder stamps `git describe --tags --always` and the commit and date
from git, and fails if `.git` is present but yields no version. The
empty CHECK_EPOCH refusal is gone so the plain build succeeds.
`script/version` now prints `git describe --tags --always --dirty`, so
make, the scripts and a plain build agree. `vaultik version` treats
the short commit, tag-N-gHASH forms and any version ending in `-dirty`
as development builds, so they keep the development-build notice.

Model: opus-5-5
2026-10-02 10:04:14 +02:00
clawbot 584444b619 List only after-1.0 work in the README roadmap (closes #208)
check / check (push) Successful in 3m35s
check / check (pull_request) Successful in 3m36s
The README roadmap and the TODO.md Next Step still described finished
1.0 work as remaining. The roadmap now lists only work planned after
1.0. Its security item says the code was reviewed before 1.0, every bug
found was fixed, and the accepted risks are listed; an outside audit
stays as after-1.0 work. The error-condition item is gone because every
failure case it listed has a fault-injection test. Daemon mode is added.
TODO.md says the 1.0 work is complete on next and that merging and
tagging are the owner's.

Judgement call: dropped the human-readable size flags item; no command
flag takes a raw-integer size.

Model: opus-5-5
2026-10-01 21:41:41 +02:00
clawbot b30e79ee45 Run the disk-full restore test again (closes #207)
check / check (push) Successful in 4m52s
check / check (pull_request) Successful in 5m19s
TestRestoreReportsDiskFull was skipped pending
#163, which is closed. The skip
and its "skipped until" wording are removed.

Restore is unchanged. With the skip removed the test failed because
restore succeeded: its simulated full disk capped only Create, but
restore now opens each file with OpenFile, so nothing was capped. It now
caps OpenFile, and only for files under the restore target: restore also
writes the decrypted metadata database under $TMPDIR through the same
filesystem, and capping that would fail the restore before any file
reached the target.

Judgement call: the test was corrected, not restore; both assertions are
unchanged.

Model: opus-5-5
2026-10-01 20:24:30 +02:00
sneak d886a9026f Merge branch 'main' into next
check / check (push) Successful in 4m8s
check / check (pull_request) Successful in 2m51s
2026-09-29 03:02:24 +02:00
clawbot 6e1f499048 Document that migrations are supported and none are added before 1.0 (closes #68)
check / check (push) Successful in 3m6s
check / check (pull_request) Successful in 3m8s
The docs now say vaultik supports migrations. The numbered files in `internal/database/schema/` are migrations: `schema_migrations` records which have run, and opening a database applies any that have not. None are added before 1.0 because nothing is installed anywhere yet, so a schema change edits `001.sql` directly. After 1.0 each change is a new numbered file, and an existing local database is migrated when vaultik is updated.

`docs/DATAMODEL.md` owns the explanation. The README caveat and roadmap entry and `AGENTS.md` policy 13 link to it. This replaces the wording from #146, which said there was no upgrade path.

Disclosure: `CLAUDE.md` line 33, the owner's file, changes from "do not need to support migrations" to "do not add migrations before 1.0".

Model: opus-5-5
2026-09-28 20:21:38 +02:00
clawbot d24f5dc33c Adopt the canonical golangci-lint config (closes #90)
check / check (push) Successful in 3m9s
check / check (pull_request) Successful in 1m29s
The lint config is now the canonical file from `prompts`, which replaces the deprecated `gomodguard` with `gomodguard_v2`, so lint prints no deprecation warnings. It also turns on the `depguard` `test-support` rule. The one difference from canonical is that the deny list names vaultik's own test-only package `internal/storage/faultstore`, so shipped code cannot import it. The new config found nothing to fix in the source.

Issues and PRs that pin the old `.golangci.yml` sha256 as an untouched-file check need the new one: `7122fcf0dd0ea57441374f98ebd98bb3da23decb67f9209fee5175170838fbd1`.

Model: opus-5-5
2026-09-23 02:14:40 +02:00
73 changed files with 2688 additions and 1589 deletions
+79 -10
View File
@@ -1,10 +1,79 @@
.git
.gitea
*.md
LICENSE
vaultik
dist
.tool
coverage.out
coverage.html
.DS_Store
# .dockerignore does NOT use .gitignore semantics. Docker matches with
# moby/patternmatcher: filepath.Match plus `**`, so `*` does not cross
# `/` and an unprefixed pattern is anchored at the context root. Every
# depth-independent pattern therefore needs `**/`, or `config/.env` and
# `certs/server.key` still ship while this file reads as solved. Only
# genuinely root-anchored entries go unprefixed. Never transplant these
# into .gitignore, where `**/` is wrong.
#
# Matching is case-sensitive, so secrets use character ranges rather
# than an ALL-CAPS twin, which would still miss `Server.Key`.
#
# Extend with this repo's own host-built artifacts, written anchored:
# `/myapp`, never `**/myapp`, which also matches `cmd/myapp/` and
# deletes the package directory from the context.
# .git is sent without its config. Without a VERSION build argument the
# stage that compiles runs `git describe --tags --always` on .git, which
# does not need .git/config; that file can hold a credential, such as a
# password in a remote URL or the token the CI checkout step stores there.
# Each submodule keeps a config with the same exposure in its git directory
# under .git/modules/, nested again for a submodule's own submodules, or in
# its own .git directory when it keeps one.
# KNOWN GAP: a submodule whose name has a `config` segment (`config`,
# `deploy/config`, `config/lib`) loses its whole git directory, because
# `**/.git/modules/**/config` also matches that segment's directory
# under .git/modules/. Go's version stamping then fails the build;
# nothing leaks. Name such a submodule without that segment:
# `git submodule add --name`.
**/.git/config
**/.git/modules/**/config
# Agent scratch: one full checkout of the repo per in-flight agent.
# Anchored because it occurs once where agents run at the repo root.
# KNOWN GAP: a repo running agents in subdirectories still ships
# `services/api/.claude/` and must add its own anchored entry.
.claude
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Re-include a committed template with a negation if the
# build needs one: `!docs/example.env`.
**/*.[eE][nN][vV]
**/.[eE][nN][vV].*
**/.[eE][nN][vV][rR][cC]
# Private keys and the bundles carrying them. Public certificates
# (*.crt, *.cer) are deliberately absent: they are legitimate inputs.
**/*.[pP][eE][mM]
**/*.[kK][eE][yY]
**/*.[pP]12
**/*.[pP][fF][xX]
**/[iI][dD]_[rR][sS][aA]
**/[iI][dD]_[dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]_[sS][kK]
**/[iI][dD]_[eE][dD]25519
**/[iI][dD]_[eE][dD]25519_[sS][kK]
# Dependencies: restored inside the image, never copied in.
**/node_modules
# OS metadata.
**/.DS_Store
**/Thumbs.db
# Editor state: never a build input, and it churns COPY.
**/*.swp
**/*.swo
**/*~
**/*.bak
**/.idea
**/.vscode
**/*.sublime-*
# This repo's own host-built artifacts.
/vaultik
/dist
/.tool
/coverage.out
/coverage.html
+3
View File
@@ -10,3 +10,6 @@ insert_final_newline = true
[Makefile]
indent_style = tab
[*.go]
indent_style = tab
+7 -12
View File
@@ -1,14 +1,9 @@
name: check
on:
push:
branches: [main, next]
pull_request:
branches: [main, next]
on: [push]
jobs:
check:
runs-on: ubuntu-latest
steps:
# actions/checkout v4, 2024-09-16
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5
- name: Build and check
run: script/cibuild
check:
runs-on: ubuntu-latest
steps:
# actions/checkout v4.2.2, 2026-02-22
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
- run: script/cibuild
+6 -5
View File
@@ -16,11 +16,12 @@ jobs:
fetch-depth: 0
# goreleaser is not a compiler: it shells out to `go` for the
# `before:` hook and for every one of the four cross-compiles.
# Nothing else in this repo puts a Go toolchain on the runner --
# check.yml runs script/cibuild, which does all of its work inside
# the digest-pinned Dockerfile images -- so without this step the
# release either fails at the before-hook or, worse, ships binaries
# built by whatever Go the runner happens to carry.
# Without this step the release either fails at the before-hook or,
# worse, ships binaries built by whatever Go the runner happens to
# carry. check.yml's runner gets the same Go through
# script/bootstrap, which calls script/install-go, and uses it only
# for `go mod download` and gofmt; it compiles inside the
# digest-pinned Dockerfile images.
#
# actions/setup-go would pin the action by commit sha, but the Go
# tarball it downloads at runtime is verified against no value in
+57 -23
View File
@@ -1,28 +1,62 @@
# Binary
/vaultik
# goreleaser output
/dist/
# Locally installed pinned tools (script/install-goreleaser)
/.tool/
# Test artifacts
*.out
*.test
coverage.html
coverage.out
# IDE
.vscode/
.idea/
*.swp
*.swo
# OS
.DS_Store
Thumbs.db
# Local config for development
# Editors
*.swp
*.swo
*~
*.bak
.idea/
.vscode/
*.sublime-*
# Agent scratch (worktrees of this repo, created and destroyed by
# in-flight tooling). Unanchored: .gitignore patterns already match at
# every depth, so no prefix is wanted here. This is not a .dockerignore
# entry and must not be given a `**/` prefix on the way into one.
.claude/
# Node
node_modules/
# Secrets. Unanchored like every entry above, so each matches at every
# depth. Matching is case-sensitive on Linux, so names use character
# ranges rather than a lowercase form that misses `Server.Key`.
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Only the templates `example.env` and `sample.env` are
# re-included below. A repository that commits any other template adds
# its own negation after these lines, for example `!.env.example`.
*.[eE][nN][vV]
.[eE][nN][vV].*
.[eE][nN][vV][rR][cC]
!example.env
!sample.env
# Private keys and the bundles carrying them.
*.[pP][eE][mM]
*.[kK][eE][yY]
*.[pP]12
*.[pP][fF][xX]
[iI][dD]_[rR][sS][aA]
[iI][dD]_[dD][sS][aA]
[iI][dD]_[eE][cC][dD][sS][aA]
[iI][dD]_[eE][cC][dD][sS][aA]_[sS][kK]
[iI][dD]_[eE][dD]25519
[iI][dD]_[eE][dD]25519_[sS][kK]
# Go build and test output.
*.log
*.out
*.test
coverage.html
# This repo's own host-built artifacts.
/vaultik
/dist/
/.tool/
# Local configs for development; they hold storage credentials.
local-config.yaml
dev-config.yaml
dev-config.yaml
+71 -2
View File
@@ -10,14 +10,21 @@ run:
linters:
default: all
enable:
# Successor to the deprecated gomodguard. Named explicitly, rather than
# left to `default: all`, because it carries the module policy below.
- gomodguard_v2
disable:
# Genuinely incompatible with project patterns
- exhaustruct # Requires all struct fields
- depguard # Dependency allow/block lists
- exhaustruct_v5 # Requires all struct fields (successor to exhaustruct)
- godot # Requires comments to end with periods
- wsl # Deprecated, replaced by wsl_v5
- wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go
# Deprecated: the warning is attached to the old name, so it is
# silenced by disabling that name, not by enabling the successor.
- wsl # Deprecated, replaced by wsl_v5
- gomodguard # Deprecated, replaced by gomodguard_v2
settings:
lll:
line-length: 88
@@ -28,6 +35,68 @@ linters:
max-complexity: 15
dupl:
threshold: 100
depguard:
# Test-support code must not be compiled into the shipped binary. A
# test-support package exists to hand a test privileges the program
# itself must never have, so a file that is not a test must not import
# one. Test files, and the files inside a package whose directory name
# ends in `test`, are where that code belongs, and are exempt.
#
# The deny list below is the one part of this file a repository is
# expected to extend, and the only part it may. depguard matches an
# import path against a list of prefixes, so it cannot be told "any path
# whose last segment ends in test"; a repository's own test-support
# packages have to be named here one at a time, by full import path,
# under a module path that differs from repository to repository. Add
# them; change nothing else.
rules:
test-support:
list-mode: lax
files:
- "$all"
- "!$test"
- "!**/*test/**"
deny:
- pkg: net/http/httptest
desc: >-
Test-support code belongs in test files and in packages whose
directory name ends in test, not in the shipped binary.
- pkg: sneak.berlin/go/vaultik/internal/storage/faultstore
desc: >-
Test-support code belongs in test files and in packages whose
directory name ends in test, not in the shipped binary.
# Only decisions already recorded in the Go package defaults are
# listed here. Every entry matches the module path exactly.
gomodguard_v2:
blocked:
- module: github.com/rs/zerolog
recommendations:
- log/slog
reason: "Structured logging is stdlib log/slog."
# One entry per pre-fork module path, because the later releases
# are separate paths. A prefix match would be shorter but would
# also reach github.com/go-redis/redismock, the test double for
# the successor these entries recommend.
- module: github.com/go-redis/redis
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/go-redis/redis/v7
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/go-redis/redis/v8
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/sergi/go-diff
recommendations:
- github.com/aymanbagabas/go-udiff
reason: "No unified diff output; use go-udiff."
- module: github.com/hexops/gotextdiff
recommendations:
- github.com/aymanbagabas/go-udiff
reason: "Unmaintained fork; use go-udiff."
issues:
max-issues-per-linter: 0
+1 -3
View File
@@ -47,9 +47,7 @@ checksum:
# A snapshot is not a release and must not name itself like one. The
# previous `{{ incpatch .Version }}-next` derived a plausible-looking
# release number from the last tag -- and with no tags in the repo at
# all, from goreleaser's fabricated v0.0.0. This produces the same
# string script/version produces for an untagged build, so a snapshot
# binary and a `make vaultik` binary of the same clean commit agree.
# all, from goreleaser's fabricated v0.0.0.
snapshot:
version_template: "dev-{{ slice .FullCommit 0 12 }}"
+25 -10
View File
@@ -102,14 +102,29 @@ Version: 2025-06-08
build files are acceptable in the root, but source code and other files
should be organized in appropriate subdirectories.
13. Pre-1.0: NEVER write database migrations. There are no live databases
anywhere — every user's local index can be rebuilt from a fresh full
backup. To change the schema, edit `internal/database/schema/001.sql`
(and any code that touches the affected tables) directly; do not add new
numbered schema files. Those numbered files and the `schema_migrations`
table they populate only bootstrap a fresh database — they are not an
upgrade path. The local index is disposable until 1.0 ships and is
tagged; once 1.0 is tagged that clause expires and the question of
upgrading existing indexes returns. See [`docs/DATAMODEL.md`](docs/DATAMODEL.md)
for the full explanation.
13. Pre-1.0: NEVER add a database migration. Migrations are supported, but
nothing is installed anywhere yet, so there is nothing to migrate. To
change the schema, edit `internal/database/schema/001.sql` (and any
code that touches the affected tables) directly. After 1.0, each schema
change is a new numbered file in that directory and a released file is
never edited; an existing local database is then migrated when vaultik
is updated. Before 1.0, a local database left on an older schema is
deleted and re-created by a full backup. See
[`docs/DATAMODEL.md`](docs/DATAMODEL.md#schema-migrations).
14. Never use `git add -A`. Stage only the files you intentionally
changed.
15. Commit messages carry no attribution or advertising trailers for the
tool that helped write the code, or for its vendor: the owner is the
sole author of code written with a tool.
16. Run the whole test suite with `make test` every time, and read its full
output. Never run `go test`, a single test or a single package, and
never grep the output.
17. Do not stop working on a task until the definition of done given in the
initial instruction is met: all of the work, not part or most of it.
18. For estimates: backing up over 2.5Gbit/s ethernet to an S3 server
backed by 2000MB/sec SSD takes about 4 seconds per gigabyte.
+1 -1
View File
@@ -353,7 +353,7 @@ CreateSnapshot(opts)
## Deduplication Strategy
1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid)
1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid), unless the file lists a chunk that no uploaded blob holds; such a file is re-chunked
2. **Chunk-level**: Chunks are content-addressed by SHA256 hash. If a chunk hash already exists in the database, the chunk data is not re-uploaded.
-44
View File
@@ -1,44 +0,0 @@
# Rules
Read the rules in AGENTS.md and follow them.
# Memory
* Claude is an inanimate tool. The spam that Claude attempts to insert into
commit messages (which it erroneously refers to as "attribution") is not
attribution, as I am the sole author of code created using Claude. It is
corporate advertising for Anthropic and is therefore completely
unacceptable in commit messages.
* NEVER use `git add -A`. Always add only the files you intentionally
changed.
* Tests should always be run before committing code. No commits should be
made that do not pass tests.
* Code should always be formatted before committing. Do not commit
unformatted code.
* Code should always be linted before committing. Do not commit
unlinted code.
* The test suite is fast and local. When running tests, don't run
individual parts of the test suite, always run the whole thing by running
"make test".
* Do not stop working on a task until you have reached the definition of
done provided to you in the initial instruction. Don't do part or most of
the work, do all of the work until the criteria for done are met.
* We do not need to support migrations; schema upgrades can be handled by
deleting the local state file and doing a full backup to re-create it.
* When testing on a 2.5Gbit/s ethernet to an s3 server backed by 2000MB/sec SSD,
estimate about 4 seconds per gigabyte of backup time.
* When running tests, don't run individual tests, or grep the output. run
the entire test suite every time and read the full output.
* When running tests, don't run individual tests, or try to grep the output.
never run "go test". only ever run "make test" to run the full test
suite, and examine the full output.
+66 -86
View File
@@ -1,96 +1,76 @@
# This file has no lint stage, deliberately.
#
# Linting lives in Dockerfile.lint, built by script/lint, and
# script/cibuild builds both. A lint stage here would have to either
# shell out to `make lint` -- which is now `docker build`, so
# docker-in-docker inside a BuildKit step with no daemon -- or call
# golangci-lint directly, which would mean a second, independently
# bumpable digest pin for the linter alongside the one in
# Dockerfile.lint. Two pins for one tool is the drift that
# https://git.eeqj.de/sneak/vaultik/issues/78 was filed over. See
# https://git.eeqj.de/sneak/vaultik/issues/113 for the ruling.
#
# Consequence, stated rather than left to be discovered: script/docker
# builds this file only and therefore does not lint. `make fmt-check`
# and `make test` still run here, so what a green build of this file
# means is "formatted, tested, and it compiles" -- the lint verdict
# comes from script/lint or script/cibuild.
# Build stage
# golang:1.26.1-alpine, 2026-03-17
FROM golang:1.26.1-alpine@sha256:2389ebfa5b7f43eeafbd6be0c3700cc46690ef842ad962f6c5bd6be49ed82039 AS builder
# Build tooling: make, plus a C toolchain because `go test -race` needs cgo.
# The sqlite driver is pure Go (modernc.org/sqlite), so no sqlite library or
# CLI is required.
RUN apk add --no-cache make build-base
# Lint phase. The linter is invoked directly rather than through `make
# lint` or `script/lint`, which are themselves a docker build and would
# recurse into a daemon that does not exist in a build step.
# golangci/golangci-lint:v2.14.0, 2026-10-05
FROM golangci/golangci-lint:v2.14.0@sha256:ad862ba6b3798cbe0fd9fd7408d498fd74fbd2623a92406b2fd3898faf0bf98f AS lint
WORKDIR /src
# Copy go mod files first for better layer caching
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# `golangci-lint run` silently ignores an unknown top-level key in
# .golangci.yml, such as a misspelt `linters:`; `config verify` fails on it.
RUN golangci-lint config verify --config .golangci.yml
RUN golangci-lint run --config .golangci.yml ./...
# Copy source code
# Test phase. -race needs cgo and so a C compiler, which the Debian Go
# image ships and the alpine one does not.
# golang:1.26.1 (Debian trixie), 2026-10-05
FROM golang:1.26.1@sha256:cd78d88e00afadbedd272f977d375a6247455f3a4b1178f8ae8bbcb201743a8a AS test
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN go test -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 90s -race -v ./...; exit 1; }
# Build stage. Nothing is wanted from either phase above; the copies
# are what make BuildKit build them first, so this stage cannot run
# unless lint and test passed.
# golang:1.26.1-alpine, 2026-03-17
FROM golang:1.26.1-alpine@sha256:2389ebfa5b7f43eeafbd6be0c3700cc46690ef842ad962f6c5bd6be49ed82039 AS builder
COPY --from=lint /src/go.sum /dev/null
COPY --from=test /src/go.sum /dev/null
RUN apk add --no-cache git
# A tar-stream context keeps the sender's file owners, which git refuses.
RUN git config --system --add safe.directory /src
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Run the format check and the tests.
#
# CHECK_EPOCH must stay immediately above these RUNs. These layers are
# keyed on its value, so they are cache-eligible only for a value
# already built against this same tree. script/cibuild and script/docker
# each pass a fresh value on every invocation, which is what makes their
# green mean the checks really executed.
#
# The value is expanded into each check command rather than left to a
# bare declaration, so the cache miss does not depend on BuildKit's
# unreferenced-ARG handling staying as it is. It also puts the epoch in
# the build log, where a reader can see the layer was keyed fresh.
#
# The guard is what makes a build that omits --build-arg fail instead of
# lie. An unset ARG is an empty string, and an empty string is a
# perfectly stable cache key: without the guard the first such build
# runs the checks and every one after it on an unchanged tree replays
# these layers from cache, executes nothing, and still exits 0. Failed
# steps are never cached, so the guard fails on EVERY invocation rather
# than once -- a bare `docker build .` is a loud error, not a quiet
# green. Do not give CHECK_EPOCH a default value; a default would
# satisfy the guard with a constant and restore the hole.
#
# Everything above this line (apk, go.mod, `go mod download`) is
# deliberately outside the busted range and keeps caching.
ARG CHECK_EPOCH
RUN [ -n "$CHECK_EPOCH" ] || exit 1
RUN echo "check epoch: ${CHECK_EPOCH}" && make fmt-check
RUN echo "check epoch: ${CHECK_EPOCH}" && make test
# The VERSION build arg when one is given, otherwise
# `git describe --tags --always` on the .git in the build context. The
# commit and its date always come from that .git. With .git present, a
# version that is still empty, dev or unknown, or a commit or date that
# is unknown, fails the build: git is missing or could not read the
# checkout, as when .git is a file pointing outside the context. A
# context without .git, such as a source export, stamps "dev" and an
# "unknown" commit and date.
ARG VERSION
RUN VERSION="${VERSION:-$(git describe --tags --always || echo dev)}"; \
commit="$(git rev-parse HEAD || echo unknown)"; \
commit_date="$(git show -s --format=%cs HEAD || echo unknown)"; \
if [ -e .git ]; then \
case "$VERSION" in ""|dev|unknown) \
echo "version is '$VERSION' although .git is present" >&2; \
exit 1 ;; \
esac; \
if [ "$commit" = unknown ] || [ "$commit_date" = unknown ]; then \
echo "commit is '$commit' and its date '$commit_date'" \
"although .git is present" >&2; \
exit 1; \
fi; \
fi; \
globals=sneak.berlin/go/vaultik/internal/globals; \
CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X ${globals}.Version=${VERSION} \
-X ${globals}.Commit=${commit} \
-X ${globals}.CommitDate=${commit_date}" \
-o /vaultik ./cmd/vaultik
# Version, commit and build date are computed on the host by
# script/docker and script/cibuild (where .git exists) and passed in as
# build args. The build context excludes .git (see .dockerignore), so
# the build cannot derive them itself: it used to try, with `git
# rev-parse` inside this stage, and always got "unknown". VERSION comes
# from script/version, the source of truth shared with the Makefile, so
# it carries the same tag / dev-<sha> / -dirty rules and a Docker image
# reports the same string a local build of the same tree would.
#
# The defaults are the fallback for a bare `docker build .` that passes
# none of them: an unset arg would otherwise stamp an empty string and
# produce an image that cannot report its own version, commit or date.
# They match what an out-of-git build reports elsewhere.
#
# These ARGs sit here, after the checks, rather than at the top of the
# stage: every commit changes their values, and a value change
# invalidates all layers below the ARG. Declared up top they would bust
# `go mod download`; here they only rekey this build layer, which the
# COPY of the sources above already rebuilds on any change anyway.
ARG VERSION=dev
ARG COMMIT=unknown
ARG COMMIT_DATE=unknown
# Build (pure Go, no CGO required since we use modernc.org/sqlite)
RUN CGO_ENABLED=0 go build -ldflags "-X 'sneak.berlin/go/vaultik/internal/globals.Version=${VERSION}' -X 'sneak.berlin/go/vaultik/internal/globals.Commit=${COMMIT}' -X 'sneak.berlin/go/vaultik/internal/globals.CommitDate=${COMMIT_DATE}'" -o /vaultik ./cmd/vaultik
# Runtime stage
# Runtime stage, and the last one: a plain `docker build .` builds this
# stage's chain and nothing else.
# alpine:3.21, 2026-02-25
FROM alpine:3.21@sha256:c3f8e73fdb79deaebaa2037150150191b9dcbfba68b4a46d70103204c53f4709
-104
View File
@@ -1,104 +0,0 @@
# Lint image.
#
# Every lint run in this repo happens inside this image, invoked through
# script/lint, and linting is a BUILD STEP rather than a container
# command: a successful build of this file IS a clean lint. That shape
# also works where the docker daemon is remote and bind mounts are
# impossible, which `docker run` against a mounted worktree does not.
#
# This FROM line is the single source of truth for the linter version in
# this repo. Nothing else pins golangci-lint: the product Dockerfile has
# no lint stage, deliberately, so there is no second digest to bump and
# no pair of pins that can drift apart. Bump the tag AND the digest here
# and nowhere else.
#
# Note for readers coming from REPO_POLICIES.md: that document still
# describes the older pattern, a lint stage inside the product
# Dockerfile wired up with `COPY --from=lint /src/go.sum /dev/null`.
# That pattern is superseded here by the owner's ruling recorded in
# https://git.eeqj.de/sneak/vaultik/issues/113 -- lint runs in its own
# image, per run, with its own cache and its own lock, which is what
# makes concurrent runs on one host safe. The policy text is org-wide
# and is being amended separately; this file is what this repo does.
#
# golangci/golangci-lint:v2.12.2, 2026-08-10
FROM golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240
WORKDIR /src
# Copy the dependency manifests first so the module download layer stays
# cached until they change. Everything above the ARG below is cacheable
# on purpose; a cold module download on every lint would make the inner
# loop unusable and buys nothing, because it is not what the gate is
# asserting.
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Force the check layers to execute on every invocation.
#
# CHECK_EPOCH must stay immediately above the RUNs below. Those layers
# are keyed on its value, so they are cache-eligible only for a value
# already built against this same tree; script/lint and script/cibuild
# each pass a fresh value on every invocation, which is what makes their
# green mean the linter really ran. Without it, `docker build -f
# Dockerfile.lint .` on an unchanged tree exits 0 in well under a second
# having linted nothing.
#
# The value is expanded into each check command itself rather than left
# to a bare declaration, so the cache miss does not depend on BuildKit's
# unreferenced-ARG handling staying as it is. It also puts the epoch in
# the build log, where a reader can see the layer was keyed fresh.
#
# The guard is what makes a build that omits --build-arg fail instead of
# lie. An unset ARG is an empty string, and an empty string is a
# perfectly stable cache key: without the guard the first such build
# lints and every one after it on an unchanged tree replays this layer,
# executes nothing, and still exits 0. Failed steps are never cached, so
# the guard fails on EVERY invocation rather than once. Do not give
# CHECK_EPOCH a default value; a default would satisfy the guard with a
# constant and restore the hole.
ARG CHECK_EPOCH
RUN [ -n "$CHECK_EPOCH" ] || exit 1
# Validate .golangci.yml before linting with it.
#
# This is not belt-and-braces; it closes a hole that `golangci-lint run`
# leaves wide open. `run` rejects YAML it cannot PARSE, but it silently
# IGNORES an unknown top-level KEY. Renaming `linters:` to `linterz:` --
# one character -- discards `default: all`, the whole disable list and
# every threshold, leaves only golangci-lint's small default linter set
# running, and exits 0 reporting `0 issues.` on a tree the real config
# fails. Demonstrated on this repo at this pin, recorded on
# https://git.eeqj.de/sneak/vaultik/pulls/114: with a planted
# over-length line, `script/lint` exits 1 naming the `revive` finding
# with `linters:` and exits 0 with `linterz:`. A set-but-ineffective
# config quietly falling back to defaults is precisely the false-green
# class this gate exists to eliminate, so it must not sit in the gate's
# own configuration.
#
# `config verify` catches it, and it does so OFFLINE at this pinned
# version -- verified, not assumed. Under `docker run --network none`
# against the pinned digest it exits 0 on this repo's config and exits 3
# on the `linterz:` variant with `additional properties 'linterz' not
# allowed`. An earlier revision of this file asserted the opposite, that
# the schema is fetched over live HTTPS from an unpinned URL, and used
# that to justify omitting this line. That claim was false at v2.12.2;
# the schema is embedded. If a future bump reintroduces a network fetch
# the failure is loud and this comment is where to record it.
#
# It is keyed on CHECK_EPOCH, like the lint run below, so it executes on
# every invocation. Content-addressing alone would arguably be enough --
# .golangci.yml arrives through `COPY . .`, so a cache hit here implies
# a byte-identical config was validated when the layer really ran. That
# argument is exactly the one that would also excuse caching the lint
# layer, and this repo has ruled it insufficient: a cached check layer
# checks nothing, and the cost of being wrong is silent. Forcing it costs
# milliseconds and puts the epoch in the log, where a reader can see that
# this validation ran rather than being replayed.
RUN echo "check epoch: ${CHECK_EPOCH}" && \
golangci-lint config verify --config .golangci.yml
RUN echo "check epoch: ${CHECK_EPOCH}" && \
golangci-lint run --config .golangci.yml ./...
+17 -13
View File
@@ -1,7 +1,10 @@
.PHONY: all bootstrap setup check test lint lint-fix fmt fmt-check build clean deps test-coverage local install release release-snapshot docker hooks
# Version number, derived from git by script/version -- the tag when
# HEAD is on one, otherwise dev-<sha>. This used to be a hardcoded
# Where script/bootstrap installs Go when the host has none.
export PATH := $(PATH):$(CURDIR)/.tool/go/bin
# Version number, derived from git by script/version (`git describe
# --tags --always --dirty`). This used to be a hardcoded
# constant, which meant every local build claimed to be a release that
# had never been tagged.
VERSION := $(shell script/version)
@@ -40,9 +43,10 @@ setup:
check:
@script/check
# Run tests only. This runs the ENTIRE suite -- there is no separate
# integration target and no build-tagged subset held back. In
# particular internal/vaultik/integration_test.go, which does full
# Run tests only, by building the test phase of the Dockerfile. This
# runs the ENTIRE suite -- there is no separate integration target and
# no build-tagged subset held back. In particular
# internal/vaultik/integration_test.go, which does full
# chunk -> pack -> encrypt -> upload -> restore round-trips, runs here.
# A `test-integration` target used to exist and was removed: no file in
# the repo carried a build tag, so `-tags=integration` selected nothing
@@ -87,17 +91,17 @@ clean:
go clean
# Install dependencies. The linter is deliberately not installed here:
# script/lint lints by building Dockerfile.lint, whose FROM line is the
# single source of truth for the linter version. A second, separately
# pinned copy on PATH could drift from it and make a local `make lint`
# disagree with CI.
# script/lint lints by building the lint phase of the Dockerfile, whose
# FROM line is the single source of truth for the linter version. A
# second, separately pinned copy on PATH could drift from it and make a
# local `make lint` disagree with CI.
deps:
go mod download
# Run tests with coverage. -count=1 for the same reason script/test
# uses it: without it an unchanged package is served from Go's test
# result cache, and a coverage profile assembled from cached results
# describes a run that did not happen.
# Run tests with coverage, on the host. -count=1 because without it an
# unchanged package is served from Go's test result cache, and a
# coverage profile assembled from cached results describes a run that
# did not happen.
test-coverage:
go test -v -count=1 -coverprofile=coverage.out ./...
go tool cover -html=coverage.out -o coverage.html
+115 -143
View File
@@ -40,7 +40,8 @@ Features:
* modern encryption ([age](https://age-encryption.org/), X25519 + ChaCha20-Poly1305)
* content-defined chunking with deduplication (FastCDC)
* incremental backups (only changed files are re-chunked)
* incremental backups (a file is re-chunked only when it changed or a
chunk it lists is held by no uploaded blob)
* multithreaded zstd compression at configurable levels
* content-addressed immutable storage
* local state tracking in SQLite (enables write-only incremental backups)
@@ -177,7 +178,7 @@ vaultik version
* `--verbose`, `-v`: Enable verbose output (on stderr — see below)
* `--debug`: Enable debug output (on stderr — see below)
* `--quiet`, `-q`: Suppress non-error output (also suppresses startup banner)
* `--skip-errors`: Skip files that cannot be read when creating a snapshot, or that cannot be restored when restoring, instead of aborting. Packing and storage errors (which would leave a chunk recorded but not stored) still abort the run.
* `--skip-errors`: Skip files that cannot be read when creating a snapshot, or that cannot be restored when restoring, instead of aborting. Packing and storage errors while creating a snapshot (which would leave a chunk recorded but not stored) still abort the run.
### locking
@@ -499,8 +500,9 @@ format does and does not protect.
* Content-defined chunking using the FastCDC algorithm
* Average chunk size: configurable (default 10MB)
* Deduplication at file level (unchanged files skipped) and chunk level
(identical chunks across files stored once)
* Deduplication at file level (unchanged files skipped, unless a chunk
the file lists is held by no uploaded blob) and chunk level (identical
chunks across files stored once)
* Multiple chunks packed into blobs to reduce object count
### encryption
@@ -528,7 +530,7 @@ complete annotated example also lives in
| Field | Default | Description |
|-------|---------|-------------|
| `age_recipients` | (required) | Age public keys for encryption |
| `age_recipients` | (required by `snapshot create`) | Age public keys for encryption. Other commands run without one, so a machine that only restores can leave it empty |
| `age_secret_key` | (unset) | Age private key for decryption (`snapshot restore`, `snapshot verify --deep`). Setting it in the config file places the private key on the backed-up host, defeating the public-key-only design (see "why" above). Prefer the `VAULTIK_AGE_SECRET_KEY` environment variable, supplied only on the machine you restore from. |
| `snapshots` | (required) | Named snapshot definitions with paths and excludes |
| `storage_url` | | Storage backend URL (`s3://`, `file://`, `rclone://`) |
@@ -559,13 +561,13 @@ complete annotated example also lives in
sequentially. Restore speed is bound by single-stream throughput.
* **Device nodes, named pipes, and sockets are silently skipped.** Only
regular files, directories, and symlinks are backed up.
* **No upgrade path between versions.** There is no supported way to carry
an existing local index across a schema change; if the local SQLite
schema changes between versions, delete the local database (`vaultik
database delete`) and run a full backup. Remote storage is unaffected.
(The binary does embed numbered schema files and a `schema_migrations`
table to bootstrap a fresh database — see [`docs/DATAMODEL.md`](docs/DATAMODEL.md)
— but that is not an upgrade path.)
* **Before 1.0, an update can make the local index unusable.** Vaultik
supports schema migrations, but none are added before 1.0 because
there is no installed base yet. If an update leaves your local index
unusable, run `vaultik database delete` and then a full backup; remote
storage is unaffected. After 1.0, the local index is migrated when
vaultik is updated. See
[`docs/DATAMODEL.md`](docs/DATAMODEL.md#schema-migrations).
* **Files that change during backup may be inconsistent.** There is no
filesystem snapshot or freeze. If a file is modified between the scan
and chunk phases, the backed-up copy may reflect a partial write.
@@ -577,22 +579,18 @@ complete annotated example also lives in
## roadmap
Items still to do before / shortly after 1.0. Loosely ordered by
priority.
Work planned after 1.0. Loosely ordered by priority.
### correctness and operability
* **Security audit of the encryption implementation.** Pre-1.0
blocker if we're advertising "secure" at the top of this README.
age + zstd + content-defined chunking is mostly off-the-shelf
pieces, but the seams (key handling, recipient parsing, manifest
trust boundary, restore-time identity validation) need an outside
read.
* **Error-condition tests.** Today's coverage is the happy path
plus a few specific regressions. Need fault-injection coverage:
network failures mid-blob, disk-full during restore, corrupted /
truncated / missing blobs, partial uploads, kill -9 between
manifest and db.zst.age writes.
* **Outside security audit.** Before 1.0 the encryption and
blob-generation code was reviewed: every bug the review found was
fixed, and the risks it accepted are listed in
[Accepted Risks](docs/REPOSTRUCTURE.md#accepted-risks). No outside
audit has been done. age + zstd + content-defined chunking
is mostly off-the-shelf pieces, but the seams (key handling,
recipient parsing, manifest trust boundary, restore-time identity
validation) need an outside read.
* **Verify restored content end-to-end in CI.** The current
integration test does this for a small synthetic snapshot but
not at scale. A nightly job against a multi-GB representative
@@ -615,13 +613,15 @@ priority.
doesn't resume from where it stopped or skip already-present
files. A `--resume` mode that checks targets before fetching
blobs would matter for very large restores.
* **Daemon mode.** A long-running mode that watches for file
changes so frequent backups, such as hourly, skip the full scan.
It adds little for the usual runs from cron every 12 to 36 hours.
See [issue #204](https://git.eeqj.de/sneak/vaultik/issues/204).
### usability
* **Man pages and richer `--help` examples.** Cobra generates
basic help; man pages would be a separate target.
* **`--bwlimit` style human-readable size flags** across the
command surface where they're currently raw integers.
* **`vaultik snapshot diff <a> <b>`** — show which files changed
between two snapshots without restoring either.
* **Status reporting hook for `--cron`.** When a backup fails
@@ -631,12 +631,11 @@ priority.
### infrastructure
* **Cross-version schema upgrades.** There is no upgrade path between
released versions — pre-1.0 schema changes are handled by `vaultik
database delete` plus a full re-scan (see
[`docs/DATAMODEL.md`](docs/DATAMODEL.md)). Post-1.0 we'll need a
migration story to keep existing index databases usable across
upgrades.
* **Schema migrations after 1.0.** Migrations are supported, but none
are added before 1.0 because there is no installed base yet. After
1.0, each schema change is a new migration, so an existing local
index is migrated when vaultik is updated (see
[`docs/DATAMODEL.md`](docs/DATAMODEL.md#schema-migrations)).
* **Storage backend coverage tests.** S3, file://, and rclone://
all share the Storer interface but the rclone path is the least
exercised in CI.
@@ -729,12 +728,11 @@ regardless of color setting (emoji are not color).
## requirements
* Go 1.26 or later
* Docker, with a reachable daemon, to lint, check, or commit:
`script/lint` lints by building `Dockerfile.lint`, which runs the
digest-pinned `golangci-lint` image as a build step, and `make check`
and the pre-commit hook both run it. A `golangci-lint` installed on
`PATH` is not a substitute and is never used on a host, whatever its
version.
* Docker, with a reachable daemon, to test, lint, check, or commit:
`script/test` and `script/lint` build the `test` and `lint` phases of
the `Dockerfile`, and `make check` and the pre-commit hook both run
them. A `golangci-lint` installed on `PATH` is not a substitute and is
never used on a host, whatever its version.
* S3-compatible object storage (or local filesystem, or rclone remote)
## development workflow
@@ -765,16 +763,20 @@ standard: normalized scripts in `script/` are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call
them. We provide:
* `script/bootstrap` — install all development dependencies (go, Go
module download). It deliberately does not install `golangci-lint`;
see `script/lint` below.
* `script/bootstrap` — install all development dependencies (Go, Go
module download). A host without Go gets the `go.mod` version through
`script/install-go`, in `.tool/go`, which `script/bootstrap` itself,
the `Makefile`, `script/fmt`, `script/fmt-check`, `script/precommit`
and `script/release` add to their `PATH`. It
deliberately does not install `golangci-lint`; see `script/lint`
below.
* `script/setup` — make a fresh clone ready for development: runs
`script/bootstrap`, then `script/install-precommit`
* `script/projectname` — print the project name (used for the Docker
image tag)
* `script/version` — print the version string to bake into the binary.
The `Makefile`'s `LDFLAGS` call this; it is the single source of truth
for the version. See [releasing](#releasing) for the rules.
The `Makefile`'s `LDFLAGS` call this. See [releasing](#releasing) for
the rules.
* `script/install-goreleaser` — install the pinned `goreleaser` into
`.tool/bin` from a sha256-verified release archive. Idempotent, and
called by `script/bootstrap`; the release workflow calls it directly
@@ -782,100 +784,59 @@ them. We provide:
`script/bootstrap` insists on.
* `script/install-go` — install the Go toolchain named by `go.mod`'s
`go` directive into `.tool/go` from a sha256-verified `go.dev`
archive, and put it on `PATH`. Idempotent. Called only by the release
workflow, which needs a host Go for `goreleaser` to shell out to;
nothing else on the release runner does. `actions/setup-go` is not
used because it verifies the downloaded toolchain against no value in
this repo. Bumping Go edits `go.mod`, the checksum in this script, and
the `Dockerfile` `golang` digest together.
archive, for Linux or macOS on amd64 or arm64. Idempotent. On a CI
runner it also puts `.tool/go/bin` on `PATH` for the steps that
follow. Called by the release workflow, which needs a host Go for
`goreleaser` to shell out to, and by `script/bootstrap` on a host
without Go. `actions/setup-go` is not used because it verifies the
downloaded toolchain against no value in this repo. Bumping Go edits
`go.mod`, the checksums in this script, and the two `golang` digests
in the `Dockerfile` together.
* `script/release` — cross-compile and publish the release artifacts
with the pinned `goreleaser`. Refuses a `goreleaser` on `PATH` whose
version is not the pinned one, on the same reasoning as `script/lint`.
* `script/release-snapshot` — the same build with no publishing and no
tagging, into `./dist`
* `script/test` — run the test suite (verbose rerun on failure). This
runs *everything*: there is no separate integration target and no
build-tagged subset held back, so the full round-trip tests in
`internal/vaultik/integration_test.go` run on every invocation. It
passes `-count=1`, which disables Go's test result cache. That is
deliberate and it is not free: on this repo's suite it costs about 11
seconds on every repeat run (measured, back to back: 0.4s cached
versus 11.6s with `-count=1`). That is the price of the run meaning
anything, because without it an unchanged package prints
`ok <pkg> (cached)`, which is indistinguishable from a package that
really ran, so the whole suite can report a full set of `ok` lines in
under half a second having executed nothing. The `-timeout` is a hang
backstop rather than a performance budget — it applies per test binary
to test execution only, not to compilation — and is set well above the
slowest package's measured runtime. Its 120s value deliberately
diverges from the 30s `REPO_POLICIES.md` mandates; the reasoning is in
the comment in the script, and issue #101 proposes amending the policy
text.
* `script/lint` — lint by building `Dockerfile.lint`, which runs
`golangci-lint run --config .golangci.yml ./...` as a build step
inside the digest-pinned `golangci-lint` image, so a successful build
*is* a clean lint. Nothing lints on the host, at any version, ever;
the script requires Docker and fails loudly rather than falling back
to a `golangci-lint` on `PATH`. That `FROM` line is the single source
of truth for the linter version — bump it there and nowhere else.
It takes no arguments, because a build step has no command line to
pass flags to, and it passes a fresh `--build-arg CHECK_EPOCH` on
every invocation so the lint layer cannot be replayed from cache (see
`script/cibuild` below for what that mechanism defends against). To
watch the linter execute, run it as
`BUILDKIT_PROGRESS=plain script/lint` and check that the lint layer
says `RUN … golangci-lint` rather than `CACHED`.
One container per run means one lint cache and one `golangci-lint`
lock per run, both private to it and discarded with it, so concurrent
runs on one host cannot contaminate or block each other.
* `script/test` — run the test suite by building the `test` phase of
the `Dockerfile` (verbose rerun on failure). This runs *everything*:
there is no separate integration target and no build-tagged subset
held back, so the full round-trip tests in
`internal/vaultik/integration_test.go` run on every invocation. The
90s `-timeout` is a hang backstop rather than a performance budget; it
applies per test binary, to test execution only.
* `script/lint` — lint by building the `lint` phase of the
`Dockerfile`, which runs `golangci-lint config verify` and then
`golangci-lint run --config .golangci.yml ./...` as build steps in
the digest-pinned `golangci-lint` image. Nothing lints on the host.
That `FROM` line is the only pin of the linter version, and it changes
in the same commit as a re-vendored `.golangci.yml`.
* `script/lint-fix` — apply the linter's autofixes (rewrites files),
using the same pinned image, parsed out of `Dockerfile.lint`. It
cannot be a build step, because fixes have to land in the worktree, so
it bind-mounts the tree into a `docker run` and therefore needs a
*local* daemon. It is a developer convenience and never a gate: no
gate reads its exit status. Run `make lint` afterwards to find out
whether the tree is clean.
using the same pinned image, parsed out of the `lint` phase's `FROM`
line. It cannot be a build step, because fixes have to land in the
worktree, so it bind-mounts the tree into a `docker run` and therefore
needs a *local* daemon. It is a developer convenience and never a
gate: no gate reads its exit status. Run `make lint` afterwards to
find out whether the tree is clean.
* `script/fmt` — format all code (writes)
* `script/fmt-check` — check formatting (read-only)
* `script/fmt-check` — check formatting (read-only). It runs `gofmt` on
the host, over every Go file outside `.tool`.
* `script/check` — run `script/test`, `script/lint`, and
`script/fmt-check`. This is authoritative *because* `script/lint`
builds `Dockerfile.lint`: a local `make check` and CI cannot disagree
about lint findings.
`script/fmt-check`.
* `script/docker` — build the Docker image tagged via
`script/projectname`. Passes a fresh `--build-arg CHECK_EPOCH` for the
same reason `script/cibuild` does, so a local image build cannot be
green on checks it replayed from cache. It builds the *product* image
only, and the product `Dockerfile` has no lint stage, so it does not
lint: a green here means formatted, tested, and it compiles.
* `script/cibuild` — CI entrypoint, and the full gate. Two builds, in
order: `Dockerfile.lint` (the linter, as a build step) and then
`Dockerfile` (`make fmt-check` and `make test` in its builder stage,
then the product image). Either failing fails the script. It runs the
checks in the same containers CI does, from a clean copy of the tree,
so it also catches anything that depends on host state.
`.gitea/workflows/check.yml` runs it on every push to `main` and
`next` and on every pull request against either.
`script/projectname`, stamped with the version
`git describe --tags --always --dirty` gives on the host, or `unknown`
outside a git checkout. The image's build stage depends on the `lint`
and `test` phases, so this lints and tests too.
* `script/cibuild` — CI entrypoint: runs `script/bootstrap`,
`script/check`, and then the same image build as `script/docker`.
`.gitea/workflows/check.yml` runs it on every push.
It passes a fresh `--build-arg CHECK_EPOCH` to each build, unique per
invocation, which both files declare immediately above their check
`RUN`s and expand into each check command. Those layers are keyed on
that value, so a new value re-runs them even on a byte-identical tree,
and a green from this script means the checks executed. Dependency and
module layers sit above the `ARG` and still cache, so a build is not
cold.
A build that supplies no `CHECK_EPOCH` — a bare `docker build .` or
`docker build -f Dockerfile.lint .` — fails rather than lying. An
unset `ARG` is an empty string and an empty string is a stable cache
key, so without a guard such a build would serve every check layer
from cache, execute nothing, and still exit 0. Each file therefore
asserts the value is non-empty before running anything, and because
failed steps are never cached that assertion fires on every
invocation rather than once. Use `script/lint`, `script/docker` or
`script/cibuild`, which pass the arg; a bare `docker build` is a loud
error.
Every `docker build` in these scripts passes `--no-cache`, because on
an unchanged tree a cached check layer is replayed without running and
the build still exits 0. A plain `docker build .` is therefore no
evidence that the checks ran. The cost is that `script/cibuild` runs
the `lint` and `test` phases twice: once in `script/check` and again
in the image build.
* `script/precommit` — pre-commit gate: `go mod tidy` + `go fmt` (must
not change files), then `script/check`
* `script/install-precommit` — install the git pre-commit hook that
@@ -886,24 +847,35 @@ them. We provide:
### version numbers
The version a binary reports comes from git, not from a constant in a
file. `script/version` decides it, and everything that stamps a binary
agrees with it:
file. It is `git describe --tags --always --dirty`, which
`script/version` runs for the `Makefile`, and `script/docker` and
`script/cibuild` run themselves:
* `HEAD` is exactly on a tag → that tag with a leading `v` stripped, so
the tag `v1.0.0` produces `vaultik 1.0.0`, matching the archive name
`vaultik_1.0.0_linux_amd64.tar.gz`. `goreleaser` strips the prefix the
same way.
* anything else → `dev-<12 chars of the commit sha>`.
* either, with uncommitted changes to tracked files → a `-dirty`
* `HEAD` is exactly on a tag → that tag, such as `v1.0.0`.
* a commit after a tag → `<tag>-<N>-g<short sha>`.
* no tag reachable → the short commit sha.
* any of these, with uncommitted changes to tracked files → a `-dirty`
suffix, because a modified checkout of a tag is not that tag.
A build that is not a release never names itself like one. `vaultik
version` says so in as many words on a development build, and
`goreleaser --snapshot` stamps the same `dev-<sha>` string rather than
inventing the next patch number. If `script/version` cannot be run at
all, `make` stops with an error instead of building an unversioned
binary, and a binary that somehow carries an empty version string still
reports itself as a development build.
A `docker build .` of a clone, with no build arguments, runs the same
`git describe` (without `--dirty`) on the `.git` in its build context,
so it stamps the same value for a clean commit; the build fails if the
context carries `.git` and no version, commit or commit date comes out.
A binary built without
git metadata reports `dev`, or `unknown` when `script/docker` or
`script/cibuild` built it outside a git checkout.
`goreleaser` stamps a release binary with the tag minus its leading
`v`, so the tag `v1.0.0` produces `vaultik 1.0.0`, matching the archive
name `vaultik_1.0.0_linux_amd64.tar.gz`. `goreleaser --snapshot` stamps
`dev-<12 chars of the commit sha>` rather than inventing the next patch
number. `vaultik version` calls a build a development build when its
version is `dev`, `unknown`, `dev-<sha>`, the short commit sha or
`<tag>-<N>-g<short sha>`, with or without `-dirty`; only a plain tag,
such as `v1.0.0` or `1.0.0`, is a release. If `script/version` cannot
be run at all, `make` stops with an error instead of building an
unversioned binary, and a binary that somehow carries an empty version
string still reports itself as a development build.
### cutting a release
+357 -86
View File
@@ -1,6 +1,6 @@
---
title: Repository Policies
last_modified: 2026-07-06
last_modified: 2026-10-04
---
This document covers repository structure, tooling, and workflow standards. Code
@@ -60,17 +60,28 @@ style conventions are in separate documents:
prerequisite since nvm requires bash. yarn is then pinned via
`corepack prepare yarn@<version> --activate`. Never install "latest" or "lts";
always exact versions. `script/cibuild` runs the CI build: it changes to the
repo root and runs `docker build .`; the Gitea workflow calls it. Four further
scripts are our own extensions to the standard: `script/check` runs
`script/test`, `script/lint`, and `script/fmt-check`; `script/precommit` is
what the git pre-commit hook runs, and it calls `script/check`;
`script/install-precommit` installs the git pre-commit hook (the `make hooks`
target shims to it); and `script/projectname` (literally that filename) simply
outputs the project's name. Scripts that need the name call
`script/projectname` — e.g. `script/docker` assembles its image tag from it —
so those scripts stay byte-identical across all repos. Repo-type-specific
pre-commit extras (e.g. `go mod tidy` verification in Go repos) belong in
`script/precommit`, not in the hook itself. Model scripts are at
repo root, runs `script/bootstrap`, runs `script/check`, and builds the image
with the version; the Gitea workflow calls it. **`script/cibuild` runs
`script/bootstrap` first**, because the workflow checks out the repo and runs
nothing else, while `script/fmt-check` runs the formatter on the host: on a
pristine checkout with nothing installed the run dies there, after the
containerised gates have passed. **The bootstrap alone is not enough**:
`script/bootstrap` installs node and yarn under nvm and leaves neither on the
`PATH` of the shell that called it, so a bare `yarn` still exits 127. The host
entrypoints that need yarn — `script/fmt` and `script/fmt-check` — therefore
source nvm for the pinned node version before invoking it, exactly as
`script/bootstrap`'s own install step does. A runner carrying nothing but
docker and git then gets through `script/check`. Four further scripts are our
own extensions to the standard: `script/check` runs `script/test`,
`script/lint` and `script/fmt-check`; `script/precommit` is what the git
pre-commit hook runs, and it calls `script/check`; `script/install-precommit`
installs the git pre-commit hook (the `make hooks` target shims to it); and
`script/projectname` (literally that filename) simply outputs the project's
name. Scripts that need the name call `script/projectname` — e.g.
`script/docker` assembles its image tag from it — so those scripts stay
byte-identical across all repos. Repo-type-specific pre-commit extras (e.g.
`go mod tidy` verification in Go repos) belong in `script/precommit`, not in
the hook itself. Model scripts are at
`https://git.eeqj.de/sneak/prompts/raw/branch/main/script/<name>`. The README
must document the provided scripts in an **Entrypoints** section (see the
README requirements below).
@@ -89,87 +100,198 @@ style conventions are in separate documents:
contributor should be able to understand the entire development workflow by
reading the Makefile.
- Every repo should have a `Dockerfile`. All Dockerfiles must run `make check`
as a build step so the build fails if the branch is not green. For non-server
repos, the Dockerfile should bring up a development environment and run
`make check`. For server repos, `make check` should run as an early build
stage before the final image is assembled. Dockerfiles install development
prerequisites by running `script/bootstrap` rather than duplicating installs
inline; COPY `script/` and the dependency manifests (`package.json` +
`yarn.lock`, `go.mod` + `go.sum`, etc.) before running it so the bootstrap
layer stays cached until dependencies change.
- Every repo should have a `Dockerfile`, and it carries the repo's gates: a
`lint` phase and a `test` phase, with the final stage depending on both so the
image cannot be built unless they pass. For non-server repos the final stage
brings up a development environment; for server repos it is the runtime image.
The gate phases and the build stage start from their pinned base images and
install what those images lack either inline, as the canonical Go `Dockerfile`
below does for `git`, or by running `script/bootstrap`, as the `prompts`
repo's own `Dockerfile` does for its yarn packages. The development
environment stage installs development prerequisites by running
`script/bootstrap` rather than duplicating its installs inline. A stage that
runs `script/bootstrap` COPYs `script/` and the dependency manifests
(`package.json` + `yarn.lock`, `go.mod` + `go.sum`, etc.) before running it.
- **Dockerfiles must use a separate lint stage for fail-fast feedback.** Go
repos use a multistage build where linting runs in an independent stage based
on the `golangci/golangci-lint` image (pinned by hash). This stage runs
`make fmt-check` and `make lint` before the full build begins. The build stage
then declares an explicit dependency on the lint stage via
`COPY --from=lint /src/go.sum /dev/null`, which forces BuildKit to complete
linting before proceeding to compilation and tests. This ensures lint failures
surface in seconds rather than minutes, without blocking on dependency
download or compilation in the build stage.
- **Linting and testing run in Docker, as phases of the `Dockerfile`.** There is
no separate lint file. `script/lint` and `script/test` each build one phase
and nothing else:
The standard pattern for a Go repo Dockerfile is:
```sh
docker build --no-cache --target lint -t "$(script/projectname)-lint" .
docker build --no-cache --target test -t "$(script/projectname)-test" .
```
**A stage that is not the last one in the file is built only when the final
stage's chain depends on it, or when `--target` names it.** That is why the
two gates are always invoked by name here, and why the final stage carries a
`COPY --from=` of a harmless file from each of them: without that edge a
plain `docker build .` builds the last stage alone and exits 0 having linted
and tested nothing.
**Every `docker build` in `script/` is tagged**, here and in
`script/cibuild` and `script/docker`. An untagged build leaves a dangling
image behind on every invocation, on every developer host and every CI
runner; a tagged one replaces the previous image.
Inside a phase the tool is invoked directly — `golangci-lint`, `go test`,
`eslint`, `prettier` — never through `make lint` or `script/test`, which are
themselves a `docker build` and would recurse into a daemon that does not
exist in a build step. Formatting is the exception and stays on the host:
`script/fmt` writes the working tree, and `script/fmt-check` is its
read-only twin.
**No lint verdict may come from a host invocation of the linter.** On a
shared host golangci-lint reads a result cache keyed on file content rather
than location, so a second checkout of the same content is served the first
one's findings, and a host-global lock in `$TMPDIR` makes concurrent runs
exit non-zero with `parallel golangci-lint is running` — a status a caller
cannot tell from real findings. Both have produced wrong verdicts in this
org, in both directions. A container has its own cache, its own `TMPDIR` and
a digest-pinned binary, so neither is reachable.
- **Any build that runs checks is built with `--no-cache`.** Docker invalidates
a `COPY` layer only when the copied content changes, so on an unchanged tree
the check `RUN` is served from cache, nothing executes, and the build still
exits 0. Every `docker build` in `script/` therefore passes `--no-cache`:
`script/lint`, `script/test`, `script/cibuild` and `script/docker` are the
four, and there is no fifth — `script/check` runs the two gate phases and
`script/fmt-check`, and builds no image of its own. A bare `docker build .` is
not evidence that anything ran: a sub-second build reporting success is a
cache hit, not a result. Never invalidate by pruning — `docker builder prune`
and friends destroy a build cache shared with every other build on the host.
When a check is added or changed, prove it works by planting a defect it must
catch and watching the run fail on it, then revert the defect. A green run
alone shows neither that the check ran nor that it covers what it should.
- **The gate phases are separate stages, and the build stage depends on both.**
The lint phase is based on the `golangci/golangci-lint` image (pinned by
hash), so lint failures surface in seconds rather than after a full compile,
and the test phase is based on the Debian Go image. The canonical Go repo
`Dockerfile`:
```dockerfile
# Lint stage — fast feedback on formatting and lint issues
# Lint phase
# golangci/golangci-lint:v2.x.x, YYYY-MM-DD
FROM golangci/golangci-lint@sha256:... AS lint
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make fmt-check
RUN make lint
RUN golangci-lint run --config .golangci.yml ./...
# Build stage
# golang:1.x-alpine, YYYY-MM-DD
FROM golang@sha256:... AS builder
# Test phase. -race needs cgo and so a C compiler, which the Debian Go
# image ships and the alpine one does not.
# golang:1.x, YYYY-MM-DD
FROM golang@sha256:... AS test
WORKDIR /src
# Force BuildKit to run the lint stage before proceeding
COPY --from=lint /src/go.sum /dev/null
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make test
RUN go test -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 90s -race -v ./...; exit 1; }
ARG VERSION=dev
RUN CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X main.Version=${VERSION}" \
-o /app ./cmd/app/
# Build stage. Nothing is wanted from either phase above; the copies
# are what make BuildKit build them first, so this stage cannot run
# unless lint and test passed.
# golang:1.x-alpine, YYYY-MM-DD
FROM golang@sha256:... AS builder
COPY --from=lint /src/go.sum /dev/null
COPY --from=test /src/go.sum /dev/null
RUN apk add --no-cache git
# A tar-stream context keeps the sender's file owners, which git refuses.
RUN git config --system --add safe.directory /src
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Runtime stage
# The VERSION build arg when one is given, otherwise
# `git describe --tags --always` on the .git in the build context. With
# .git present, a version that is still empty, dev or unknown fails the
# build: git is missing or could not read the checkout.
ARG VERSION
RUN VERSION="${VERSION:-$(git describe --tags --always)}"; \
if [ -e .git ]; then \
case "$VERSION" in ""|dev|unknown) \
echo "version is '$VERSION' although .git is present" >&2; \
exit 1 ;; \
esac; \
fi; \
CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X main.Version=${VERSION}" \
-o /app ./cmd/app/
# Runtime stage, and the last one
FROM alpine@sha256:...
COPY --from=builder /app /usr/local/bin/app
ENTRYPOINT ["app"]
```
Key points:
- The lint stage uses the `golangci/golangci-lint` image directly (it
includes both Go and the linter), so there is no need to install the
linter separately.
- `COPY --from=lint /src/go.sum /dev/null` is a no-op file copy that creates
a stage dependency. BuildKit runs stages in parallel by default; without
this line, the build stage would not wait for lint to finish and a lint
failure might not fail the overall build.
- The lint phase uses the `golangci/golangci-lint` image directly (it has
both Go and the linter), so nothing needs installing.
- `COPY --from=<phase> /src/go.sum /dev/null` is a no-op copy whose only
purpose is the ordering edge. BuildKit runs stages in parallel by default,
and a stage nothing depends on is not built at all, so without these two
lines a red gate would not fail the build.
- Keep the runtime stage last, and if you add a stage after it, give it the
same two copies. A plain `docker build .` builds the last stage's chain
and nothing else.
- If the project uses `//go:embed` directives that reference build artifacts
(e.g. a web frontend compiled in a separate stage), the lint stage must
(e.g. a web frontend compiled in a separate stage), the lint phase must
create placeholder files so the embed directives resolve. Example:
`RUN mkdir -p web/dist && touch web/dist/index.html web/dist/style.css`.
The lint stage should not depend on the actual build output — it exists to
fail fast.
- If the project requires CGO or system libraries for linting (e.g.
`vips-dev`), install them in the lint stage with `apk add`.
- The build stage runs `make test` after compilation setup. Tests run in the
build stage, not the lint stage, because they may require compiled
artifacts or heavier dependencies.
- If the project requires CGO or system libraries for linting, install them
in the lint phase. The `golangci/golangci-lint` image is Debian-based and
has no `apk`, so install with `apt-get` under the Debian package name
(`libvips-dev`, where alpine says `vips-dev`), and delete the package
lists in the same `RUN`, so the layer does not keep them:
```dockerfile
RUN apt-get update \
&& apt-get install -y --no-install-recommends libvips-dev \
&& rm -rf /var/lib/apt/lists/*
```
- `.dockerignore` lets `.git` into the build context. It keeps out every git
`config` at any depth (`**/.git/config`, `**/.git/modules/**/config`): the
repository's own, each submodule's under `.git/modules/`, and that of a
submodule keeping its own `.git` directory. `git describe` does not need
them, and each can hold a credential: a password in a remote URL, or the
token the CI checkout step stores there. A submodule whose name has a
`config` segment (`config`, `deploy/config`, `config/lib`) loses its whole
git directory to `**/.git/modules/**/config`, and Go's version stamping
then fails the build: give it a name without that segment
(`git submodule add --name`). The stage that compiles has `git` (the
Debian Go image has it; an alpine one needs `apk add --no-cache git`) and
takes the version from the `VERSION` build argument when one is given,
otherwise from `git describe --tags --always`. That gives the tag on a
tagged commit; on a later commit, the tag, the number of commits since it
and the short commit (`v1.2.3-4-gabc1234`); and the short commit when no
tag is reachable. The stage that compiles also marks its working directory
safe for git (`git config --system --add safe.directory /src`): a context
sent as a tar stream keeps the sender's file owners, and git refuses a
checkout owned by another user, so the version would come out empty.
`ARG VERSION` has no default, and the build fails if the context carries
`.git` and the version still comes out empty, `dev` or `unknown`. A plain
`docker build .` with no build arguments must succeed; a Dockerfile that
refuses an empty build argument drops that refusal and keeps the argument.
- Every repo should have a Gitea Actions workflow (`.gitea/workflows/`) that
runs `script/cibuild` (which runs `docker build .`) on push. Since the
Dockerfile already runs `make check`, a successful build implies all checks
pass.
runs `script/cibuild` on push, and checks out the repo as its only other step.
That script bootstraps, runs the gate phases, and then builds the image, so a
successful run means every check passed; a bare `docker build .` does not
carry the same guarantee, because its gate phases may come from the cache. The
image build is uncached and so runs the gate phases a second time. That is the
price of the rule above, and it is worth paying: the image that ships is built
from a run of its own gates rather than from a cache entry. A separate
workflow limited to `main` by a `branches` list under `on: push` cannot be
checked by review: to try a change to it, add the feature branch to that list
and push, then remove the branch from the list again before merging. Keep any
job in it that publishes behind `if: github.ref_name == 'main'`, so the run
from the feature branch publishes nothing.
- Use platform-standard formatters: `black` for Python, `prettier` for
JS/CSS/Markdown/HTML, `go fmt` for Go. Always use default configuration with
@@ -189,14 +311,21 @@ style conventions are in separate documents:
module under test to verify it compiles/parses. There is no excuse for
`make test` to be a no-op.
- `make test` must complete in under 20 seconds. Add a 30-second timeout in the
Makefile.
- `make test` must complete in under 60 seconds. That is the hard cap, and a
suite that exceeds it fails. Under 20 seconds is the target. A suite between
20 and 60 seconds is still green, but the overage must be filed as an
improvement bug against that repo. Add a 90-second timeout to the test
invocation (`go test -timeout 90s`). The backstop deliberately sits above the
hard cap so that it catches a genuinely hung test rather than a merely slow
one.
- **`make test` should use the conditional verbose rerun pattern.** Run tests
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to
show full output. This keeps CI logs and `docker build` output clean on
success (just package/suite summaries) while providing full diagnostic detail
on failure (every test case, every assertion). The general shell pattern:
- **The test command should use the conditional verbose rerun pattern.** Run
tests without `-v` (verbose) first. If tests fail, automatically rerun with
`-v` to show full output. This keeps CI logs and `docker build` output clean
on success (just package/suite summaries) while providing full diagnostic
detail on failure (every test case, every assertion). The command lives in the
`test` phase of the `Dockerfile`, since `script/test` builds that phase; the
Makefile form below is the same pattern for any repo-local invocation:
```makefile
test:
@@ -209,11 +338,26 @@ style conventions are in separate documents:
```makefile
test:
@go test -timeout 30s -race -cover ./... || \
@go test -count=1 -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 30s -race -v ./...; exit 1; }
go test -count=1 -timeout 90s -race -v ./...; exit 1; }
```
`-count=1` is required on both invocations: it defeats Go's test _result_
cache, so neither run can report a stored pass in place of running the
tests. It leaves the build cache alone, so it costs the runtime of the suite
and no recompilation.
That cache is Go's own, separate from Docker's layer cache. Go stores a
passing result in its cache directory (`GOCACHE`), and when the same tests
run again on unchanged code it prints that result, marked `(cached)`,
without running them. That matters on a developer's machine, where this
target runs and the directory lasts from one run to the next. The `test`
phase of the `Dockerfile` needs no `-count=1`: its base image holds no
result for this repo's tests and nothing before its `go test` step runs a
test, so there is nothing to replay. `--no-cache` (above) is what makes that
step run on an unchanged tree.
Python example:
```makefile
@@ -239,10 +383,84 @@ style conventions are in separate documents:
must be in `.gitignore`. No exceptions.
- `.gitignore` should be comprehensive from the start: OS files (`.DS_Store`),
editor files (`.swp`, `*~`), language build artifacts, and `node_modules/`.
Fetch the standard `.gitignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.gitignore` when setting up
a new repo.
editor files (`.swp`, `*~`), in-repo agent scratch directories (`.claude/`),
language build artifacts, and `node_modules/`. Fetch the standard `.gitignore`
from `https://git.eeqj.de/sneak/prompts/raw/branch/main/.gitignore` when
setting up a new repo. These patterns are written to `.gitignore`'s own
semantics, in which an unanchored pattern already matches at every depth; they
are not a `.dockerignore` and must not be transplanted into one unmodified.
- **`.dockerignore` does not use `.gitignore` semantics, and copying patterns
across unmodified leaves secrets in the build context.** Docker matches with
`moby/patternmatcher`: `filepath.Match` semantics plus a `**` extension, so
`*` does not cross `/` and a pattern without a leading `**/` is anchored at
the build-context root. A `.dockerignore` listing `.env`, `*.pem` and `*.key`
therefore excludes only the copies at the repository root, while `config/.env`
and `certs/server.key` still reach the context and can land in an image layer
— which is more dangerous than a short file with no secret patterns at all,
because it reads as solved and stops anyone looking. Give every
depth-independent pattern the `**/` prefix and leave only genuinely
root-anchored entries unprefixed: `.claude`, and the repo's own host-built
binary, written `/myapp` and never `**/myapp`, which would also match
`cmd/myapp/` and delete the package directory from the context. Matching is
case-sensitive, and an ALL-CAPS twin per pattern still misses `Server.Key`, so
secret names use character ranges — `**/*.[kK][eE][yY]`, `**/*.[pP][eE][mM]`,
and likewise for `.envrc` and the extensionless SSH keys. Where such a pattern
also catches something the build needs, re-include it with a negation
(`!docs/example.env`); deleting the pattern reopens the exposure for every
other file it covers. Fetch the standard `.dockerignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.dockerignore` and extend
it with the repo's own artifacts.
- **In-repo agent scratch belongs in both files, written to each file's own
semantics.** `.claude/` holds one worktree per in-flight agent — an entire
additional checkout of the repo — so under `COPY . .` the build context
inflates by a multiple of the repo and another session's unreviewed work can
be copied into an image layer. In `.gitignore` the entry is `.claude/`,
unanchored. In `.dockerignore` it is `.claude`, anchored and with **no** `**/`
prefix, because the prefixed form would also delete any nested directory of
that name from the build. Anchoring carries a known gap that the canonical
`.dockerignore` states in its own comment, since consuming repos receive the
file and not the tracker: the directory is created in the agent's working
directory, so a repo running agents in subdirectories still ships
`services/api/.claude/` and must add its own anchored entry there.
- **A plain `docker build .` of a clone stamps the version that
`git describe --tags --always` gives**, derived from the `.git` in the build
context as the canonical `Dockerfile` above shows. Without its failure check,
a missing `git` or an unreadable checkout would leave `-X main.Version=` empty
and the build would still exit 0. `script/docker` and `script/cibuild` pass
the version they compute on the host; it takes precedence. They do this
byte-identically across repos:
```sh
# Own line: a failing command substitution inside an argument does not
# trip `set -e`, so the inline form degrades to an empty constant.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$(script/projectname)" .
```
`--always` makes an untagged repo yield an abbreviated commit hash rather
than failing, and the `[ -n "$version" ]` line is the single place the
fallback is applied — a live check that fires on a build from an export with
no `.git` and on a repository with no commits yet. Do not fold it into the
substitution as `|| echo unknown`, which makes the guard unreachable. The
Dockerfile's side is `ARG VERSION` in the stage that compiles, declared
there because `ARG` is stage-scoped; passing `VERSION` to a repo whose
Dockerfile declares no such `ARG` is ignored and costs nothing, which is why
the scripts stay byte-identical. One consequence for CI: the standard
checkout action clones shallow and fetches no tags, so a repo that embeds a
tag-derived version must set `fetch-depth: 0` on its checkout step.
- **Verify `.dockerignore` by enumerating the image, not by reading the
patterns.** Plant files at the root _and_ at least two directories deep, build
a probe image that does `COPY . .`, and list what actually landed
(`docker run --rm --entrypoint find IMAGE /app`). The `transferring context`
size is not a substitute: a nested secret is a few bytes, and BuildKit
transfers only the delta from the previous build.
- **No build artifacts in version control.** Code-derived data (compiled
bundles, minified output, generated assets) must never be committed to the
@@ -258,9 +476,56 @@ style conventions are in separate documents:
- Make all changes on a feature branch. You can do whatever you want on a
feature branch.
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only
manually by the user. Fetch from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`.
- `.golangci.yml` is standardized. The vendored copy in a consuming repo must
_NEVER_ be modified by an agent: fetch it from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml` and keep it
byte-identical, so that no repo can quietly loosen its own linting. Linter
configuration changes are made to the canonical copy in the `prompts` repo and
reach consuming repos by re-vendoring; an agent may open a PR against
canonical, which only the user merges. One list is exempt from byte-identity,
because it cannot be written once for every repo: the `deny` list of the
`test-support` depguard rule, where a repo names its own test-support packages
by full import path. A repo adds entries there and changes nothing else, and a
re-vendor carries its entries forward. The canonical golangci-lint version is
v2.14.0 (released 2026-09-24), pinned as the digest of the lint phase's base
image
(`golangci/golangci-lint@sha256:ad862ba6b3798cbe0fd9fd7408d498fd74fbd2623a92406b2fd3898faf0bf98f`,
which reports `2.14.0 built with go1.27.0 from 114493f9`). A module's `go`
directive must not name a newer Go minor version than the one golangci-lint
was built with, or golangci-lint refuses to lint it: this release lints
`go 1.27.1` but not `go 1.28`. That digest is the only pin, since no repo
installs golangci-lint on the host. A repo sets the lint phase digest to the
one named here and re-vendors `.golangci.yml` in the same commit, whichever of
the two prompted the change: the canonical copy can name linters that an older
golangci-lint rejects, and a newer golangci-lint can add linters that
`default: all` switches on until the canonical copy disables them.
- **`script/bootstrap` installs a pinned tool by comparing versions, never by
testing presence.** An `if ! command -v <tool>; then install; fi` guard tests
`PATH` only, so on an already-provisioned machine the pin is inert and a
version bump is a silent no-op — while the Dockerfile, installing into a clean
image, gets the pinned version, so a local `make check` and `make docker` can
disagree about what the tool even is. The canonical form:
- compares the installed version against the pin over the **whole** version
token; a parser that stops at the first `-` reports `2.12.2` for a host
running `2.12.2-rc1` and skips the install;
- treats absent, non-zero, empty or unrecognised `--version` output as a
mismatch, so the failure direction is a redundant install and never a
skipped one;
- after installing, re-resolves the binary the way callers do — `hash -r`,
then through `PATH`, not through the directory the installer wrote to —
and fails naming the resolved path, since an install that a shadowing
binary hides succeeds while changing nothing any caller sees;
- is actually called, and prints the version on both success paths: a
function defined and never invoked has the same exit status and the same
empty output as one that worked.
Keep it POSIX sh: no arrays, no `[[`, no `grep -P`.
A Go tool a repo needs on the host is installed with `go install` pinned to
a commit hash (`go install <package>@<commit hash>`). It is never tracked as
a `go.mod` tool dependency or through a `tools.go` file, either of which
pulls the tool's own dependencies into the repo's `go.mod` and `go.sum`.
- When pinning images or packages by hash, add a comment above the reference
with the version and date (YYYY-MM-DD).
@@ -374,12 +639,14 @@ style conventions are in separate documents:
settings.
- Avoid putting files in the repo root unless necessary. Root should contain
only project-level config files (`README.md`, `Makefile`, `Dockerfile`,
`LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`, and
language-specific config). Everything else goes in a subdirectory. Canonical
subdirectory names:
only project-level config files (`README.md`, `AGENTS.md`, `Makefile`,
`Dockerfile`, `LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`,
and language-specific config). Everything else goes in a subdirectory.
Canonical subdirectory names:
- `bin/` — executable scripts and tools
- `cmd/` — Go command entrypoints
- `cmd/` — Go command entrypoints; thin only: one `main.go` per binary whose
body is a single call into `internal/` or `pkg/`, no project logic in
`cmd/`
- `configs/` — configuration templates and examples
- `deploy/` — deployment manifests (k8s, compose, terraform)
- `docs/` — documentation and markdown (README.md stays in root)
@@ -406,3 +673,7 @@ style conventions are in separate documents:
- Go: `go.mod`, `go.sum`, `.golangci.yml`
- JS: `package.json`, `yarn.lock`, `.prettierrc`, `.prettierignore`
- Python: `pyproject.toml`
- Guidance for coding agents lives in one `AGENTS.md` at the repository root. It
is never committed under a file or directory named after one agent tool, such
as `CLAUDE.md` or `.claude/`, and never split into separate memory files.
+93 -9
View File
@@ -14,17 +14,100 @@ pre-1.0
# Next Step
Define the remaining scope for the first tagged release under the 1.0.0
milestone, then cut that tag. The mechanism to cut it now exists and is
exercised; what is left is the scope decision, which is the owner's.
This step deliberately names one version number: it previously said
"cut v0.1.0" while the `Makefile` baked in `1.0.0-rc.1` and the issue
milestone said 1.0.0, and three different answers to "what is the next
release" is exactly the contradiction
[issue #65](https://git.eeqj.de/sneak/vaultik/issues/65) was filed over.
The 1.0 work is complete on `next`: the scope settled on
[issue #125](https://git.eeqj.de/sneak/vaultik/issues/125) was the work
already planned for 1.0, and all of it has landed. The mechanism to cut
the tag exists and is exercised; what is left is merging `next` to
`main` and tagging, both the owner's.
# Completed Steps
- 2026-10-06: Made `s3://bucket/prefix` and `s3://bucket/prefix/` the same
destination ([issue #222](https://git.eeqj.de/sneak/vaultik/issues/222)).
The S3 client put the prefix directly in front of each key, so a prefix
without a trailing slash stored `prefixblobs/...`. A non-empty prefix is
now joined to every key with one `/`, giving the README's
`<bucket>/<prefix>/blobs/...` layout. The `s3.prefix` config setting goes
through the same client and gets the same join.
- 2026-10-06: Made `go.mod` what `go mod tidy` writes, so the
pre-commit hook no longer stops every commit
([issue #246](https://git.eeqj.de/sneak/vaultik/issues/246)). A test
in `internal/cli` imports `github.com/spf13/pflag` directly, but
`go.mod` still marked it `// indirect`, and `script/precommit` fails
whenever the tidy changes `go.mod`. It is now in the direct `require`
block.
- 2026-10-06: Made the README's steps for restoring on another machine
work ([issue #221](https://git.eeqj.de/sneak/vaultik/issues/221)).
`config init` wrote a placeholder recipient that `config.Load`
rejects, so every command on the new machine failed before reaching
the store. The file now has an empty `age_recipients` list,
`config.Load` accepts an empty list, and `snapshot create` refuses to
run without a recipient.
- 2026-10-06: Made `snapshot restore --skip-errors` skip the files that
need a blob it cannot download
([issue #218](https://git.eeqj.de/sneak/vaultik/issues/218)). A missing
or damaged blob ended the restore even with the flag, after restoring
whichever files happened to come first. Every file that needs such a
blob is now reported as failed, the rest are restored, and the command
still exits non-zero. Without the flag the blob error still aborts.
- 2026-10-06: Re-vendored the canonical files from `sneak/prompts` at
`dd4027b` ([issue #213](https://git.eeqj.de/sneak/vaultik/issues/213)).
Linting and testing are now the `lint` and `test` phases of the
`Dockerfile`, and the build stage depends on both. `Dockerfile.lint`,
`CHECK_EPOCH` and the tests that checked them are gone; every
`docker build` in `script/` passes `--no-cache` instead. golangci-lint
is v2.14.0, with its new findings fixed in the code, and the rules in
`CLAUDE.md` now live in `AGENTS.md`.
- 2026-10-06: Made the local index actually run in WAL mode with a busy
timeout ([issue #217](https://git.eeqj.de/sneak/vaultik/issues/217)).
The connection settings were written in a form the SQLite driver
ignores, so the index ran without either and `snapshot list` or `info`
during a backup could make the backup's next write fail with
`database is locked`. They are now `_pragma=` parameters, and the
metadata export copies the open index with `VACUUM INTO`, because a
copy of the file alone misses rows still in the `-wal` file.
- 2026-10-06: Made a backup record the real uid and gid of files,
directories and symlinks
([issue #216](https://git.eeqj.de/sneak/vaultik/issues/216)). The
scanner asked the stat result for `Uid()` and `Gid()` methods, which
`*syscall.Stat_t` does not have, so every entry was stored as `0:0`
and a restore as root gave every file to root. It now reads the
`Uid` and `Gid` fields of `*syscall.Stat_t`.
- 2026-10-06: Made a `file://` destination whose directory is missing
count as one that cannot be listed
([issue #220](https://git.eeqj.de/sneak/vaultik/issues/220)). The file
backend listed a missing directory as an empty store, so with the
quickstart's USB stick unplugged `snapshot list` reported every local
snapshot as missing from the store, `snapshot remove` claimed to have
removed metadata it never reached, and `prune` dropped every local
snapshot record. Listing a missing destination directory is now an
error; a first backup still creates the directory.
- 2026-10-06: Made `--older-than` and `--keep-newer-than` reject a
duration with characters outside its number-and-unit parts
([issue #215](https://git.eeqj.de/sneak/vaultik/issues/215)). The
parser picked out the parts it recognised and skipped the rest, so
`1.0y` became zero and `--prune --keep-newer-than 1.0y` deleted every
snapshot of the backed-up names, the new one included. `1.0y`,
`1,5y`, `30 days`, `x7d` and a bare `0` are now errors; decimals still
work in Go units such as `1.5h`.
- 2026-10-06: Made a backup re-chunk a known file that lists a chunk no
uploaded blob holds
([issue #214](https://git.eeqj.de/sneak/vaultik/issues/214)). File
rows are shared by every snapshot and updated in place, so removing or
pruning the only snapshot that referenced a changed file's current
blob left an older snapshot keeping a row that matched the disk while
no blob held its chunks. The next backup skipped the file and
completed a snapshot that could not restore it.
- 2026-09-22: Routed the last direct-to-stdout command output through
`internal/ui`
([issue #149](https://git.eeqj.de/sneak/vaultik/issues/149)). The
@@ -693,4 +776,5 @@ release" is exactly the contradiction
# Future Steps
None queued; the release-scoping item is now the Next Step.
Work planned after 1.0 is listed in the README
[roadmap](README.md#roadmap).
+101 -54
View File
@@ -11,92 +11,139 @@ import (
// This file guards the version stamping of the product image (issue
// #75). The failure it protects against is silent: the image still
// builds and runs, but `vaultik version` inside it reports "commit:
// unknown", so an operator cannot tell which source produced a given
// backup. .dockerignore excludes .git, so the build cannot derive the
// commit itself; the values must be computed on the host and passed in.
// unknown" or a version of "dev", so an operator cannot tell which
// source produced a given backup. The build takes the version as a
// build arg, which script/docker computes on the host, and otherwise
// derives it from the .git in its context; the commit and its date
// always come from that .git.
//
// These are parses of the committed files, for the same reason the lint
// guards next door are: shelling out to docker would nest a build
// inside `make test`. That `vaultik version` in the built image really
// prints the host's version is verified by hand and recorded on the
// pull request.
// These are parses of the committed files, because shelling out to
// docker would nest a build inside `make test`. That `vaultik version`
// in the built image really prints the host's version is verified by
// hand.
// dockerScript is script/docker, relative to the repository root.
const dockerScript = "script/docker"
// The files under guard, relative to the repository root.
const (
productDockerfile = "Dockerfile"
dockerScript = "script/docker"
)
// versionArgs are the ldflag targets the build stamps and, matching
// them, the build args the host must supply. The names line up so the
// same list checks both files.
func versionArgs() []string {
return []string{"VERSION", "COMMIT", "COMMIT_DATE"}
}
// TestProductDockerfileTakesVersionAsBuildArgs fails unless the build
// declares each version arg and stamps it into the binary by ldflag
// reference, rather than computing it in the container.
func TestProductDockerfileTakesVersionAsBuildArgs(t *testing.T) {
// TestProductDockerfileTakesVersionAsBuildArg fails unless the build
// declares ARG VERSION, with no default, and stamps it into the binary
// whenever it is given, ahead of the value derived in the container.
func TestProductDockerfileTakesVersionAsBuildArg(t *testing.T) {
t.Parallel()
found := instructions(t, productDockerfile)
for _, arg := range versionArgs() {
require.GreaterOrEqual(t, indexOf(found, "ARG "+arg), 0,
"%s must declare `ARG %s` so the host can pass it in",
productDockerfile, arg)
require.Contains(t, found, "ARG VERSION",
"%s must declare `ARG VERSION`, with no default, so the host can"+
" pass it in", productDockerfile)
assertLdflagReferences(t, found, arg)
}
assertLdflagReferences(t, found, "VERSION")
}
// TestProductDockerfileDoesNotDeriveVersionItself is the anti-regression
// for the original defect: the container ran `git rev-parse`, but .git
// is not in the build context, so it always resolved to "unknown". No
// git command may reach into a build that cannot see the history.
func TestProductDockerfileDoesNotDeriveVersionItself(t *testing.T) {
// TestProductDockerfileDerivesVersionFromGit fails unless a build given
// no VERSION, such as a plain `docker build .` of a clone, takes it from
// `git describe` of the .git in its context, stamps "dev" when the
// context has no .git, and fails rather than stamp "dev" when that .git
// yields no version. The commit and its date come from the same .git,
// and the build fails rather than stamp them "unknown" when it is
// present.
func TestProductDockerfileDerivesVersionFromGit(t *testing.T) {
t.Parallel()
text := instructionText(readRepoFile(t, productDockerfile))
found := instructions(t, productDockerfile)
assert.NotContains(t, text, "git ",
"%s must not run git: .git is excluded from the build context, so"+
" any value it derives is wrong. Pass version, commit and date"+
" in as build args instead.", productDockerfile)
buildAt := indexContaining(found, "go build")
require.GreaterOrEqual(t, buildAt, 0, "%s must build", productDockerfile)
assert.Contains(t, found[buildAt], "git describe --tags --always || echo dev",
"%s must derive the version from git when no VERSION is given,"+
" and stamp dev when the context has no .git", productDockerfile)
assert.Contains(t, found[buildAt], "[ -e .git ]",
"%s must fail when the context carries .git but yields no version",
productDockerfile)
assert.Contains(t, found[buildAt], "git rev-parse HEAD",
"%s must stamp the commit from git", productDockerfile)
assert.Contains(t, found[buildAt], "git show -s --format=%cs HEAD",
"%s must stamp the commit date from git", productDockerfile)
assert.Contains(t, found[buildAt],
`[ "$commit" = unknown ] || [ "$commit_date" = unknown ]`,
"%s must fail when the context carries .git but yields no commit"+
" or date", productDockerfile)
}
// TestDockerScriptComputesVersionOnTheHost fails unless script/docker
// derives each value where .git exists and passes it as a build arg,
// with VERSION coming from script/version so a Docker build reports the
// same string a local build of the same tree would.
// passes the version it derives where .git exists as a build arg.
func TestDockerScriptComputesVersionOnTheHost(t *testing.T) {
t.Parallel()
script := readRepoFile(t, dockerScript)
for _, arg := range versionArgs() {
assert.Contains(t, script, "--build-arg "+arg+"=",
"%s must pass --build-arg %s to the build", dockerScript, arg)
}
assert.Contains(t, script, "/version",
"%s must take VERSION from script/version, the source of truth"+
" shared with the Makefile", dockerScript)
assert.Contains(t, script, "--build-arg VERSION=",
"%s must pass --build-arg VERSION to the build", dockerScript)
}
// assertLdflagReferences fails unless some build instruction stamps the
// named variable from the ARG (a ${arg} reference), not from a value
// computed inside the container.
// assertLdflagReferences fails unless the build instruction uses the
// named ARG whenever it is given (a ${arg:- reference), so a value
// passed in is not overridden by one derived inside the container.
func assertLdflagReferences(t *testing.T, found []string, arg string) {
t.Helper()
for _, instruction := range found {
if strings.HasPrefix(instruction, "RUN ") &&
strings.Contains(instruction, "go build") &&
strings.Contains(instruction, "${"+arg+"}") {
strings.Contains(instruction, "${"+arg+":-") {
return
}
}
assert.Fail(t, "version arg is declared but never stamped",
"the go build in %s must reference ${%s} in its ldflags, or the"+
" arg is passed and discarded", productDockerfile, arg)
"the go build in %s must use ${%s:-...}, or the arg is passed and"+
" discarded", productDockerfile, arg)
}
// instructions returns the Dockerfile's instructions, one per element,
// with comments and blank lines dropped and continuation lines joined,
// so a multi-line RUN is one string.
func instructions(t *testing.T, name string) []string {
t.Helper()
var (
out []string
continued string
isContinued bool
)
for line := range strings.SplitSeq(readRepoFile(t, name), "\n") {
trimmed := strings.TrimSpace(line)
if !isContinued && (trimmed == "" || strings.HasPrefix(trimmed, "#")) {
continue
}
isContinued = strings.HasSuffix(trimmed, `\`)
continued += strings.TrimSuffix(trimmed, `\`)
if isContinued {
continue
}
out = append(out, strings.Join(strings.Fields(continued), " "))
continued = ""
}
return out
}
// indexContaining returns the position of the first instruction
// containing want; -1 if there is none.
func indexContaining(found []string, want string) int {
for i, instruction := range found {
if strings.Contains(instruction, want) {
return i
}
}
return -1
}
-367
View File
@@ -1,367 +0,0 @@
package main_test
import (
"os"
"path/filepath"
"strings"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// This file guards the shape of the lint gate. Every property asserted
// here is one whose loss is SILENT: the build still exits 0, the gate
// still looks green, and nothing was linted or tested.
//
// The gate is a build step. script/lint builds Dockerfile.lint, which
// runs golangci-lint as a RUN instruction, so a successful build is a
// clean lint. BuildKit will happily replay that RUN from cache on an
// unchanged tree in well under a second, which is why the check layers
// are keyed on a CHECK_EPOCH build arg that the calling script
// regenerates per invocation, and why an empty value is a hard error
// rather than a stable cache key.
//
// These are parses rather than invocations. Shelling out to docker from
// the test suite would nest a build inside `make test`, which itself
// runs inside a build in CI. The one property a parse cannot establish
// -- that a real finding actually fails the build -- is verified by
// hand against a deliberately broken tree, recorded on the pull
// request.
//
// One property is deliberately NOT tested here: that no script runs the
// linter on the host. script/lint is the only lint entry point, and it
// runs golangci-lint only inside the container; keeping it that way is a
// review matter, not something a test in this file establishes.
// The files under guard, relative to the repository root.
const (
lintDockerfile = "Dockerfile.lint"
productDockerfile = "Dockerfile"
lintScript = "script/lint"
cibuildScript = "script/cibuild"
)
// linterBinary is the linter's command name, used to locate the
// config-verify and lint steps in Dockerfile.lint.
const linterBinary = "golangci-lint"
// checkEpochARG is the declaration, with no default value. A default
// would satisfy the non-empty guard with a constant, and a constant is
// a stable cache key: the checks would be replayed from cache forever
// after the first build.
const checkEpochARG = "ARG CHECK_EPOCH"
// checkEpochGuard is what turns a build that omits --build-arg into a
// loud failure instead of a quiet green. Failed steps are never cached,
// so it fires on every such invocation rather than once.
const checkEpochGuard = `RUN [ -n "$CHECK_EPOCH" ] || exit 1`
// freshEpoch is the epoch computation the calling scripts must use, as
// a bare assignment on its own line. Inline in an argument, a failing
// `date` would not abort under `set -eu`; CHECK_EPOCH would become the
// empty string, and the guard above would be the only thing standing
// between that and a permanently cached green. `$$` is required because
// `date +%s` is second-granular and busybox silently drops `%N`, so
// without the pid two concurrent runs in one second can collide.
const freshEpoch = `epoch="$(date +%s%N)$$"`
// TestLintDockerfilePinsTheLinterByDigest fails if the lint image stops
// being pinned. An unpinned tag makes the gate's verdict depend on
// whatever the registry currently serves under that name.
func TestLintDockerfilePinsTheLinterByDigest(t *testing.T) {
t.Parallel()
from := ""
for _, instruction := range instructions(t, lintDockerfile) {
if strings.HasPrefix(instruction, "FROM ") {
from = instruction
break
}
}
require.NotEmpty(t, from, "%s declares no FROM", lintDockerfile)
assert.Contains(t, from, "golangci/golangci-lint",
"the lint image must be the golangci-lint image")
assert.Contains(t, from, "@sha256:",
"the lint image must be pinned by digest, not by tag alone")
}
// TestLintDockerfileCannotBeCachedGreen pins the whole cache-busting
// mechanism in the file that lints: the declaration with no default,
// the non-empty guard, and the value expanded into the lint command
// itself rather than merely declared.
func TestLintDockerfileCannotBeCachedGreen(t *testing.T) {
t.Parallel()
found := instructions(t, lintDockerfile)
argAt := indexOf(found, checkEpochARG)
require.GreaterOrEqual(t, argAt, 0,
"%s must declare `%s` with no default value",
lintDockerfile, checkEpochARG)
assert.GreaterOrEqual(t, indexOf(found, checkEpochGuard), argAt,
"%s must guard against an empty CHECK_EPOCH with `%s`",
lintDockerfile, checkEpochGuard)
assertEpochExpandedInto(t, found[argAt:], "golangci-lint run")
// Dependency layers must stay above the ARG, or every lint run
// re-downloads the module cache and the inner loop becomes
// unusable.
download := indexOf(found, "RUN go mod download")
require.GreaterOrEqual(t, download, 0,
"%s must download modules in their own layer", lintDockerfile)
assert.Less(t, download, argAt,
"`%s` must come after `go mod download` so dependency layers"+
" still cache", checkEpochARG)
}
// TestLintDockerfileVerifiesTheLinterConfig guards the validation of
// .golangci.yml itself. `golangci-lint run` rejects a config it cannot
// parse but silently IGNORES an unknown top-level key, so renaming
// `linters:` to `linterz:` discards `default: all` and every threshold
// and still exits 0 reporting no issues. `config verify` is what turns
// that into a failure, and it has to run BEFORE the lint, or the lint
// spends a minute reporting a verdict from a config already known to be
// wrong.
func TestLintDockerfileVerifiesTheLinterConfig(t *testing.T) {
t.Parallel()
found := instructions(t, lintDockerfile)
verify := linterBinary + " config verify"
verifyAt := indexContaining(found, verify)
require.GreaterOrEqual(t, verifyAt, 0,
"%s must run `%s --config .golangci.yml`: without it a typo'd"+
" top-level key in .golangci.yml is silently ignored and the"+
" gate passes with only the default linter set", lintDockerfile,
verify)
runAt := indexContaining(found, linterBinary+" run")
require.GreaterOrEqual(t, runAt, 0, "%s must lint", lintDockerfile)
assert.Less(t, verifyAt, runAt,
"%s must verify the config before linting with it", lintDockerfile)
// Keyed on the epoch like every other check layer, so it executes
// per invocation rather than being replayed. A cached validation
// validates nothing.
assertEpochExpandedInto(t, found, verify)
}
// TestProductDockerfileCannotBeCachedGreen holds the same line for the
// checks that remain in the product image build.
func TestProductDockerfileCannotBeCachedGreen(t *testing.T) {
t.Parallel()
found := instructions(t, productDockerfile)
argAt := indexOf(found, checkEpochARG)
require.GreaterOrEqual(t, argAt, 0,
"%s must declare `%s` with no default value",
productDockerfile, checkEpochARG)
assert.GreaterOrEqual(t, indexOf(found, checkEpochGuard), argAt,
"%s must guard against an empty CHECK_EPOCH", productDockerfile)
assertEpochExpandedInto(t, found[argAt:], "make fmt-check")
assertEpochExpandedInto(t, found[argAt:], "make test")
}
// TestProductDockerfileDoesNotLint records the split deliberately: the
// linter lives in Dockerfile.lint and nowhere else, so there is exactly
// one digest pinning it. A lint stage reintroduced here would either be
// docker-in-docker (`make lint` is now `docker build`) or a second,
// independently bumpable pin.
func TestProductDockerfileDoesNotLint(t *testing.T) {
t.Parallel()
contents := readRepoFile(t, productDockerfile)
for _, forbidden := range []string{"golangci", "make lint"} {
assert.NotContains(t, instructionText(contents), forbidden,
"%s must not lint: the linter is pinned once, in %s",
productDockerfile, lintDockerfile)
}
}
// TestLintScriptBuildsTheLintDockerfileWithAFreshEpoch is the other
// half of the mechanism. The Dockerfile's guard only rejects an EMPTY
// epoch; a constant non-empty one would satisfy it and still be served
// from cache forever.
func TestLintScriptBuildsTheLintDockerfileWithAFreshEpoch(t *testing.T) {
t.Parallel()
script := readRepoFile(t, lintScript)
assertBareEpochAssignment(t, script, lintScript)
assert.Contains(t, script, `--build-arg CHECK_EPOCH="$epoch"`,
"%s must pass the fresh epoch to the build", lintScript)
assert.Contains(t, script, lintDockerfile,
"%s must build %s", lintScript, lintDockerfile)
}
// TestCibuildBuildsBothDockerfilesWithFreshEpochs guards the CI gate:
// dropping either build silently removes a whole class of check from
// CI while leaving it green.
func TestCibuildBuildsBothDockerfilesWithFreshEpochs(t *testing.T) {
t.Parallel()
script := readRepoFile(t, cibuildScript)
assertBareEpochAssignment(t, script, cibuildScript)
assert.Equal(t, 2, strings.Count(script, freshEpoch),
"%s must compute a fresh epoch for each of its two builds",
cibuildScript)
assert.Equal(t, 2,
strings.Count(script, `--build-arg CHECK_EPOCH="$epoch"`),
"%s must pass a fresh epoch to both builds", cibuildScript)
assert.Contains(t, script, "-f Dockerfile.lint",
"%s must build %s", cibuildScript, lintDockerfile)
}
// assertEpochExpandedInto fails unless some instruction runs the named
// command with the epoch expanded into it. Expansion, not mere
// declaration: an ARG that no instruction references is not guaranteed
// to key the layer, and the expansion also puts the value in the build
// log where a reader can see the layer was keyed fresh.
func assertEpochExpandedInto(t *testing.T, found []string, command string) {
t.Helper()
for _, instruction := range found {
if !strings.HasPrefix(instruction, "RUN ") {
continue
}
if strings.Contains(instruction, command) &&
strings.Contains(instruction, "${CHECK_EPOCH}") {
return
}
}
assert.Fail(t, "no epoch-keyed layer runs the command",
"`%s` must run in a layer that expands ${CHECK_EPOCH}, or it"+
" will be replayed from cache without executing", command)
}
// assertBareEpochAssignment fails unless the script computes the epoch
// as a bare assignment on its own line.
func assertBareEpochAssignment(t *testing.T, script, name string) {
t.Helper()
for line := range strings.SplitSeq(script, "\n") {
if strings.TrimSpace(line) == freshEpoch {
return
}
}
assert.Fail(t, "no bare epoch assignment",
"%s must compute `%s` as a bare assignment on its own line, so"+
" `set -e` catches a failing date instead of quietly"+
" building with an empty epoch", name, freshEpoch)
}
// instructions returns the Dockerfile's instructions, one per element,
// with comments and blank lines dropped and continuation lines joined,
// so a multi-line RUN is one string.
func instructions(t *testing.T, name string) []string {
t.Helper()
return strings.Split(instructionText(readRepoFile(t, name)), "\n")
}
// instructionText is instructions' parse, before splitting: it is also
// what a "must not contain" assertion should look at, so that a word
// appearing only in a comment is not mistaken for behaviour.
func instructionText(contents string) string {
var (
out []string
continued string
isContinued bool
)
for line := range strings.SplitSeq(contents, "\n") {
trimmed := strings.TrimSpace(line)
if !isContinued && (trimmed == "" || strings.HasPrefix(trimmed, "#")) {
continue
}
isContinued = strings.HasSuffix(trimmed, `\`)
continued += strings.TrimSuffix(trimmed, `\`)
if isContinued {
continue
}
out = append(out, strings.Join(strings.Fields(continued), " "))
continued = ""
}
return strings.Join(out, "\n")
}
// indexOf returns the position of the first instruction equal to, or
// beginning with, want; -1 if there is none. An `ARG NAME=default`
// counts as beginning with `ARG NAME`, so a declared arg is found
// whether or not it carries a default.
func indexOf(found []string, want string) int {
for i, instruction := range found {
if instruction == want ||
strings.HasPrefix(instruction, want+" ") ||
strings.HasPrefix(instruction, want+"=") {
return i
}
}
return -1
}
// indexContaining returns the position of the first instruction
// containing want; -1 if there is none.
func indexContaining(found []string, want string) int {
for i, instruction := range found {
if strings.Contains(instruction, want) {
return i
}
}
return -1
}
// readRepoFile reads a file by its path relative to the repository
// root.
func readRepoFile(t *testing.T, name string) string {
t.Helper()
//nolint:gosec // G304: the path is a constant relative to this repo
contents, err := os.ReadFile(filepath.Join(repoRoot(t), name))
require.NoError(t, err)
return string(contents)
}
// repoRoot returns the repository root. The test binary runs with its
// package directory as the working directory, so the root is found by
// walking up until the module file appears.
func repoRoot(t *testing.T) string {
t.Helper()
dir, err := os.Getwd()
require.NoError(t, err)
for {
_, err = os.Stat(filepath.Join(dir, "go.mod"))
if err == nil {
return dir
}
parent := filepath.Dir(dir)
require.NotEqual(t, dir, parent,
"walked to the filesystem root without finding a go.mod")
dir = parent
}
}
+38 -2
View File
@@ -1,6 +1,8 @@
package main_test
import (
"os"
"path/filepath"
"regexp"
"slices"
"strings"
@@ -84,14 +86,48 @@ func TestBuildTargetBuildsTheBinary(t *testing.T) {
"`make build` must depend on the rule that builds the binary")
}
// readMakefile returns the contents of the repository's Makefile. The
// root is located by the shared walk in lintdocker_test.go.
// readMakefile returns the contents of the repository's Makefile.
func readMakefile(t *testing.T) string {
t.Helper()
return readRepoFile(t, "Makefile")
}
// readRepoFile reads a file by its path relative to the repository
// root.
func readRepoFile(t *testing.T, name string) string {
t.Helper()
//nolint:gosec // G304: the path is a constant relative to this repo
contents, err := os.ReadFile(filepath.Join(repoRoot(t), name))
require.NoError(t, err)
return string(contents)
}
// repoRoot returns the repository root. The test binary runs with its
// package directory as the working directory, so the root is found by
// walking up until the module file appears.
func repoRoot(t *testing.T) string {
t.Helper()
dir, err := os.Getwd()
require.NoError(t, err)
for {
_, err = os.Stat(filepath.Join(dir, "go.mod"))
if err == nil {
return dir
}
parent := filepath.Dir(dir)
require.NotEqual(t, dir, parent,
"walked to the filesystem root without finding a go.mod")
dir = parent
}
}
// phonyTargets returns every name declared phony, across all .PHONY
// lines.
func phonyTargets(makefile string) []string {
+2 -1
View File
@@ -3,7 +3,8 @@
# Copy this file and uncomment/modify the values you need
# Age recipient public keys for encryption
# This is REQUIRED - backups are encrypted to these public keys
# Backups are encrypted to these public keys. snapshot create needs at least
# one; listing, verifying and restoring do not
# Generate with: age-keygen | grep "public key"
age_recipients:
- age1cj2k2addawy294f6k2gr2mf9gps9r3syplryxca3nvxj3daqm96qfp84tz
+20 -24
View File
@@ -6,33 +6,28 @@ Vaultik uses a local SQLite database to track file metadata, chunk mappings, and
**Important Notes:**
This section is the authoritative explanation of the schema/migration story;
other documents (the README and `AGENTS.md`) link here.
- **No upgrade path between versions (pre-1.0)**: Vaultik has no supported way to
carry an existing local index across a schema change. The index is disposable
— if the on-disk schema changes between versions, delete the local SQLite
database (`vaultik database delete`) and run a full backup. Remote storage is
unaffected; the new index re-deduplicates against existing remote blobs. This
is the standing project policy, and it is separate from the schema bootstrap
described next.
- **Schema bootstrap**: a fresh database is populated from numbered SQL files
embedded in the binary under `internal/database/schema/`. `000.sql` creates the
`schema_migrations` table; `001.sql` creates the application tables. On opening
a database the code applies each numbered file that has not yet run and records
its version in `schema_migrations`. This bootstraps a new database; it does not
upgrade an existing one between released versions.
- **Changing the schema (pre-1.0)**: edit `internal/database/schema/001.sql` (and
the code that touches the affected tables) directly. Do not add new numbered
files — there is no installed base to migrate.
- **Disposability expires at 1.0**: the index is treated as disposable only until
1.0 ships and is tagged. Once 1.0 is tagged that clause expires and the
question of upgrading existing indexes returns. It is deliberately left open
here.
- **Version Compatibility**: In rare cases, you may need to use the same version
of Vaultik to restore a backup as was used to create it. This ensures
compatibility with the metadata format stored in S3.
## Schema Migrations
Vaultik supports schema migrations. They are the numbered SQL files in
`internal/database/schema/`, embedded in the binary: `000.sql` creates the
`schema_migrations` table, which records each migration that has run, and
`001.sql` creates the application tables. `database.New` opens a database and
applies, in order, every migration that database has not yet recorded.
**Before 1.0** no migrations are added, because nothing is installed anywhere
yet. A schema change edits `001.sql` (and the code that uses the affected
tables) directly. A local database created before the change has already
recorded `001.sql` as run, so it keeps the old schema and can become unusable;
`vaultik database delete` followed by a full backup rebuilds it.
**After 1.0** each schema change is a new numbered file, so an existing local
database is migrated the first time an updated vaultik opens it. A file that
has shipped in a release is never edited.
## Database Tables
### 1. `files`
@@ -200,6 +195,7 @@ Tracks blob upload metrics.
1. **Change Detection**
- `SELECT * FROM files WHERE path = ?` - Get previous file metadata
- Compare mtime, size, mode to detect changes
- Re-chunk a file that lists a chunk no uploaded blob holds, even when its metadata is unchanged
- Skip unchanged files but still add to `snapshot_files`
2. **Chunk Reuse**
@@ -284,7 +280,7 @@ This ensures consistency, especially important for operations like:
3. **Batch Operations**: Where possible, operations are batched within transactions
4. **Write-Ahead Logging**: SQLite WAL mode is enabled for better concurrency
4. **Write-Ahead Logging**: The local index runs in SQLite WAL mode with a 10-second busy timeout, so a read-only command such as `snapshot list` can read it while a backup writes to it. Committed rows can sit in the `-wal` file beside the index until a checkpoint, so the metadata export copies the index through SQLite (`VACUUM INTO`), not as a file
## Data Integrity
+1 -1
View File
@@ -20,6 +20,7 @@ require (
github.com/rclone/rclone v1.72.1
github.com/spf13/afero v1.15.0
github.com/spf13/cobra v1.10.1
github.com/spf13/pflag v1.0.10
github.com/stretchr/testify v1.11.1
go.uber.org/fx v1.24.0
golang.org/x/sync v0.18.0
@@ -225,7 +226,6 @@ require (
github.com/smarty/assertions v1.16.0 // indirect
github.com/sony/gobreaker v1.0.0 // indirect
github.com/spacemonkeygo/monkit/v3 v3.0.25-0.20251022131615-eb24eb109368 // indirect
github.com/spf13/pflag v1.0.10 // indirect
github.com/t3rm1n4l/go-mega v0.0.0-20251031123324-a804aaa87491 // indirect
github.com/tidwall/gjson v1.18.0 // indirect
github.com/tidwall/match v1.1.1 // indirect
+7 -4
View File
@@ -45,16 +45,19 @@ const defaultConfigTemplate = `# vaultik configuration
# ─── REQUIRED ────────────────────────────────────────────────────────────────
# Age recipient public keys for encryption.
# Age recipient public keys for encryption. snapshot create needs at least
# one; listing, verifying and restoring do not, so a machine that only
# restores can leave this empty.
# Backups are encrypted to ALL listed recipients; any one of the corresponding
# private keys can decrypt. Adding a recipient later does not re-encrypt data
# already stored: deduplicated chunks and existing blobs stay encrypted to the
# earlier recipients, so a newly added key cannot restore them on its own (see
# docs/REPOSTRUCTURE.md, Accepted Risks). Generate a keypair with:
# docs/REPOSTRUCTURE.md, Accepted Risks). Generate a keypair and add its
# public key with:
# age-keygen -o vaultik_backup_private_key.txt
# grep 'public key' vaultik_backup_private_key.txt
age_recipients:
- age1REPLACE_WITH_YOUR_PUBLIC_KEY
# vaultik config set age_recipients.0 age1...
age_recipients: []
# Named snapshots. Each snapshot backs up one or more paths and can have its
# own exclude patterns in addition to the global excludes below.
+41 -2
View File
@@ -24,8 +24,10 @@ func TestDefaultConfigTemplateParses(t *testing.T) {
t.Fatalf("default config template is not valid YAML: %v", err)
}
if len(cfg.AgeRecipients) != 1 {
t.Errorf("expected 1 placeholder age recipient, got %d", len(cfg.AgeRecipients))
// A placeholder recipient would fail config.Load, so the template
// leaves the list empty.
if len(cfg.AgeRecipients) != 0 {
t.Errorf("expected no age recipients, got %d", len(cfg.AgeRecipients))
}
home, ok := cfg.Snapshots["home"]
@@ -55,6 +57,43 @@ func TestDefaultConfigTemplateParses(t *testing.T) {
}
}
// TestConfigSetRecipientOnFreshConfig follows the README quickstart: on the
// file `config init` writes, `config set age_recipients.0` and
// `config set storage_url` give a config that loads with that recipient.
func TestConfigSetRecipientOnFreshConfig(t *testing.T) {
t.Parallel()
const recipient = "age1278m9q7dp3chsh2dcy82qk27v047zywyvtxwnj4cvt0z65jw6a7q5dqhfj"
path := filepath.Join(t.TempDir(), "config.yml")
err := os.WriteFile(path, []byte(defaultConfigTemplate), configFileMode)
if err != nil {
t.Fatalf("write config: %v", err)
}
out := ui.NewWithColor(&bytes.Buffer{}, false)
err = writeConfigSet(out, path, "age_recipients.0", recipient)
if err != nil {
t.Fatalf("config set age_recipients.0: %v", err)
}
err = writeConfigSet(out, path, "storage_url", "file:///mnt/backups")
if err != nil {
t.Fatalf("config set storage_url: %v", err)
}
cfg, err := config.Load(path)
if err != nil {
t.Fatalf("config.Load: %v", err)
}
if len(cfg.AgeRecipients) != 1 || cfg.AgeRecipients[0] != recipient {
t.Errorf("age_recipients = %v, want [%s]", cfg.AgeRecipients, recipient)
}
}
const testYAML = `# top comment
compression_level: 3
age_recipients:
+9 -6
View File
@@ -172,9 +172,9 @@ func TestBannerSuppressedInArgs(t *testing.T) {
// hermeticConfig is a complete, valid config that needs no network and
// no credentials: file:// storage is exempt from the S3 credential
// checks, and FileStorer over a directory that does not exist lists
// zero objects without erroring. Chunk, blob and compression settings
// are filled in by config.Load.
// checks. A test that lists the destination must create its directory
// first, because listing a directory that does not exist is an error.
// Chunk, blob and compression settings are filled in by config.Load.
const hermeticConfig = `age_recipients:
- age1278m9q7dp3chsh2dcy82qk27v047zywyvtxwnj4cvt0z65jw6a7q5dqhfj
snapshots:
@@ -196,22 +196,25 @@ hostname: test-host
// `snapshot list` is the command chosen because it is the only --json
// command that reaches its document without a populated destination
// store: it reads the local index, streams `metadata/` (empty here),
// and treats a barren destination as an empty list rather than a
// failure.
// and treats an empty destination directory as an empty list rather
// than a failure.
//
// Not parallel: it replaces os.Args, os.Stdout and the xdg globals.
func TestEntryJSONStdoutIsExactlyOneDocument(t *testing.T) {
dir := t.TempDir()
configPath := filepath.Join(dir, "config.yml")
storeDir := filepath.Join(dir, "store")
contents := fmt.Sprintf(hermeticConfig,
filepath.Join(dir, "source"),
filepath.Join(dir, "store"),
storeDir,
filepath.Join(dir, "index.sqlite"))
require.NoError(t,
os.WriteFile(configPath, []byte(contents), configFileMode))
require.NoError(t, os.Mkdir(storeDir, 0o750))
// The PID lock lives under xdg.DataHome, which xdg resolves at
// package init; point it at the temp dir so the test neither
// touches nor collides with the real one.
+10 -5
View File
@@ -98,25 +98,30 @@ func TestEntryPruneJSONStdoutIsExactlyOneDocument(t *testing.T) {
}
}
// writeHermeticPruneConfig builds a config over a temp directory and, if
// seedStale is set, creates the index database up front with one
// snapshot record that has no counterpart on the destination store.
// Returns the config path.
// writeHermeticPruneConfig builds a config over a temp directory with an
// empty destination directory and, if seedStale is set, creates the
// index database up front with one snapshot record that has no
// counterpart on the destination store. Returns the config path.
func writeHermeticPruneConfig(t *testing.T, seedStale bool) string {
t.Helper()
dir := t.TempDir()
configPath := filepath.Join(dir, "config.yml")
indexPath := filepath.Join(dir, "index.sqlite")
storeDir := filepath.Join(dir, "store")
contents := fmt.Sprintf(hermeticConfig,
filepath.Join(dir, "source"),
filepath.Join(dir, "store"),
storeDir,
indexPath)
require.NoError(t,
os.WriteFile(configPath, []byte(contents), configFileMode))
// prune fails on a destination directory that does not exist, so the
// empty store is created here rather than left to a first backup.
require.NoError(t, os.Mkdir(storeDir, 0o750))
// The PID lock lives under xdg.DataHome, which xdg resolves at
// package init; point it at the temp dir so the test neither
// touches nor collides with the real one.
+2 -1
View File
@@ -60,7 +60,8 @@ on the source system.`,
cmd.PersistentFlags().BoolVar(&rootFlags.SkipErrors, "skip-errors", false,
"Skip files that cannot be read when creating a snapshot, or "+
"that cannot be restored when restoring, instead of aborting "+
"(packing and storage errors still abort)")
"(packing and storage errors while creating a snapshot still "+
"abort)")
// Add subcommands
cmd.AddCommand(
+5 -10
View File
@@ -42,9 +42,7 @@ const (
// Sentinel validation errors.
var (
errNoConfigPath = errors.New("config path not provided")
errNoAgeRecipients = errors.New(
"at least one age_recipient is required (generate with: age-keygen)")
errNoConfigPath = errors.New("config path not provided")
errRecipientIsSecretKey = errors.New(
"an age secret key was given where a public key (age1...) belongs")
errRecipientNotX25519 = errors.New(
@@ -323,9 +321,10 @@ func Load(path string) (*Config, error) {
// Validate checks if the configuration is valid and complete.
// It ensures all required fields are present and have valid values:
// - At least one age recipient must be specified, and every recipient must
// parse as an X25519 age1... public key (so a bad entry fails at load, not
// mid-backup); errors name the position, never the value
// - Every age recipient must parse as an X25519 age1... public key (so a
// bad entry fails at load, not mid-backup); errors name the position,
// never the value. An empty list is accepted, because only snapshot
// create needs a recipient and it checks for one itself
// - At least one snapshot must be configured with at least one path
// - Storage must be configured (either storage_url or s3.* fields)
// - Chunk size must be at least 1MB
@@ -336,10 +335,6 @@ func Load(path string) (*Config, error) {
//
// Returns an error describing the first validation failure encountered.
func (c *Config) Validate() error {
if len(c.AgeRecipients) == 0 {
return errNoAgeRecipients
}
for i, recipient := range c.AgeRecipients {
err := validateAgeRecipient(recipient)
if err != nil {
+8 -2
View File
@@ -223,7 +223,8 @@ func TestValidateBlobSizeLimit(t *testing.T) {
// TestValidateAgeRecipients checks that recipients are parsed at config load
// (a bad entry fails immediately, not mid-backup) and that no invalid entry —
// least of all a pasted secret key — is echoed in the error.
// least of all a pasted secret key — is echoed in the error. An empty list
// loads, because only snapshot create needs a recipient.
func TestValidateAgeRecipients(t *testing.T) {
t.Parallel()
@@ -244,7 +245,12 @@ func TestValidateAgeRecipients(t *testing.T) {
wantErr bool
}{
{
name: "config init placeholder is rejected",
name: "no recipients is accepted",
recipients: nil,
wantErr: false,
},
{
name: "placeholder recipient is rejected",
recipients: []string{"age1REPLACE_WITH_YOUR_PUBLIC_KEY"},
wantErr: true,
},
-2
View File
@@ -16,8 +16,6 @@ var (
// Size represents a byte size that can be specified in configuration files.
// It can unmarshal from both numeric values (interpreted as bytes) and
// human-readable strings like "10MB", "2.5GB", or "1TB".
//
//nolint:recvcheck // UnmarshalYAML requires a pointer; String/Int64 are value reads
type Size int64
// UnmarshalYAML implements yaml.Unmarshaler for Size, allowing it to be
+1 -1
View File
@@ -220,7 +220,7 @@ func (r *ChunkFileRepository) CreateBatch(
cf.ChunkHash.String(), cf.FileID.String(), cf.FileOffset, cf.Length)
}
query += querySb183.String() //nolint:gosec // G202: appends "?" placeholders only
query += querySb183.String()
query += " ON CONFLICT(chunk_hash, file_id) DO NOTHING"
+1 -1
View File
@@ -98,7 +98,7 @@ func (r *ChunkRepository) GetByHashes(
args[i] = hash
}
query += querySb75.String() //nolint:gosec // G202: appends "?" placeholders only
query += querySb75.String()
query += ") ORDER BY chunk_hash"
+39 -44
View File
@@ -42,6 +42,10 @@ var schemaFS embed.FS
// table itself. It is applied before the normal migration loop.
const bootstrapVersion = 0
// busyTimeoutMs is how long a connection to the index waits for another
// connection's lock before failing with "database is locked".
const busyTimeoutMs = 10000
// DB represents the Vaultik local index database connection.
// It uses SQLite to track file metadata, content-defined chunks, and blob associations.
// The database enables incremental backups by detecting changed files and
@@ -94,10 +98,27 @@ func ParseMigrationVersion(filename string) (int, error) {
return version, nil
}
// indexDSN returns the driver DSN that opens the database at path. The
// driver runs each _pragma parameter on every connection it opens and drops
// any parameter it does not know without an error, so a setting written in
// another form silently does nothing. In WAL mode one connection can write
// while others read; the busy timeout makes a connection wait for a lock
// instead of failing at once.
func indexDSN(path string) string {
return fmt.Sprintf(
"%s?_pragma=busy_timeout(%d)&_pragma=journal_mode(WAL)"+
"&_pragma=synchronous(NORMAL)&_pragma=foreign_keys(1)",
path, busyTimeoutMs)
}
// New creates a new database connection at the specified path.
// It creates the schema if needed and configures SQLite with WAL mode for
// better concurrency. SQLite handles crash recovery automatically when
// opening a database with journal/WAL files present.
// It creates the schema if needed. Every connection runs in WAL mode with
// a busy timeout and foreign keys on (see indexDSN), so a read-only command
// can read the index while a backup writes to it. Committed rows can sit in
// the -wal file beside the database until a checkpoint, so a copy of the
// database file alone may miss them.
// SQLite handles crash recovery automatically when opening a database with
// journal/WAL files present.
// The path parameter can be a file path for persistent storage or ":memory:"
// for an in-memory database (useful for testing).
func New(ctx context.Context, path string) (*DB, error) {
@@ -110,11 +131,7 @@ func New(ctx context.Context, path string) (*DB, error) {
// First attempt with standard WAL mode
log.Debug("Attempting to open database with WAL mode", "path", path)
conn, err := sql.Open(
"sqlite",
path+"?_journal_mode=WAL&_synchronous=NORMAL&_busy_timeout=10000"+
"&_locking_mode=NORMAL&_foreign_keys=ON",
)
conn, err := sql.Open("sqlite", indexDSN(path))
if err == nil {
configureConnPool(conn)
@@ -134,8 +151,8 @@ func New(ctx context.Context, path string) (*DB, error) {
_ = conn.Close()
}
// If first attempt failed, try with TRUNCATE mode to clear any locks
return openWithRecovery(ctx, path)
// If the first attempt failed, try once more
return retryOpen(ctx, path)
}
// configureConnPool serializes all database access through one connection.
@@ -147,18 +164,12 @@ func configureConnPool(conn *sql.DB) {
conn.SetMaxIdleConns(1)
}
// finishOpen enables foreign keys, wraps the connection, and applies any
// pending migrations. On migration failure the connection is closed.
// finishOpen wraps the connection and applies any pending migrations. On
// migration failure the connection is closed.
func finishOpen(ctx context.Context, conn *sql.DB, path string) (*DB, error) {
// Enable foreign keys explicitly
_, err := conn.ExecContext(ctx, "PRAGMA foreign_keys = ON")
if err != nil {
log.Warn("Failed to enable foreign keys", "path", path, "error", err)
}
db := &DB{conn: conn, path: path}
err = applyMigrations(ctx, conn)
err := applyMigrations(ctx, conn)
if err != nil {
_ = conn.Close()
@@ -168,21 +179,15 @@ func finishOpen(ctx context.Context, conn *sql.DB, path string) (*DB, error) {
return db, nil
}
// openWithRecovery retries opening the database in TRUNCATE journal mode to
// clear stale locks, then switches back to WAL mode.
func openWithRecovery(ctx context.Context, path string) (*DB, error) {
log.Info(
"Database appears locked, attempting recovery with TRUNCATE mode",
"path", path,
)
// retryOpen makes a second attempt to open the database, with the same
// settings, after the first attempt failed, for example because another
// process held a lock for longer than the busy timeout.
func retryOpen(ctx context.Context, path string) (*DB, error) {
log.Info("Database appears locked, retrying open", "path", path)
conn, err := sql.Open(
"sqlite",
path+"?_journal_mode=TRUNCATE&_synchronous=NORMAL&_busy_timeout=10000"+
"&_foreign_keys=ON",
)
conn, err := sql.Open("sqlite", indexDSN(path))
if err != nil {
return nil, fmt.Errorf("opening database in recovery mode: %w", err)
return nil, fmt.Errorf("opening database on retry: %w", err)
}
configureConnPool(conn)
@@ -190,28 +195,18 @@ func openWithRecovery(ctx context.Context, path string) (*DB, error) {
err = conn.PingContext(ctx)
if err != nil {
log.Debug(
"Failed to ping database in recovery mode, closing",
"Failed to ping database on retry, closing",
"path", path, "error", err,
)
_ = conn.Close()
return nil, fmt.Errorf(
"database still locked after recovery attempt: %w",
"database still locked on retry: %w",
err,
)
}
log.Debug("Database opened in TRUNCATE mode", "path", path)
// Switch back to WAL mode
log.Debug("Switching database back to WAL mode", "path", path)
_, err = conn.ExecContext(ctx, "PRAGMA journal_mode=WAL")
if err != nil {
log.Warn("Failed to switch back to WAL mode", "path", path, "error", err)
}
db, err := finishOpen(ctx, conn, path)
if err != nil {
return nil, err
+83
View File
@@ -120,6 +120,89 @@ func TestDatabaseConcurrentAccess(t *testing.T) {
}
}
// TestNewSetsJournalModeAndBusyTimeout checks that the connection settings
// New passes reach SQLite. The driver drops a setting it does not recognise
// without an error, so only reading the value back shows it took effect.
func TestNewSetsJournalModeAndBusyTimeout(t *testing.T) {
t.Parallel()
ctx := context.Background()
db, err := New(ctx, filepath.Join(t.TempDir(), "index.db"))
if err != nil {
t.Fatalf("failed to create database: %v", err)
}
defer func() { _ = db.Close() }()
var journalMode string
err = db.conn.QueryRowContext(ctx, "PRAGMA journal_mode").Scan(&journalMode)
if err != nil {
t.Fatalf("reading journal_mode: %v", err)
}
if journalMode != "wal" {
t.Errorf("journal_mode = %q, want %q", journalMode, "wal")
}
var busyTimeout int
err = db.conn.QueryRowContext(ctx, "PRAGMA busy_timeout").Scan(&busyTimeout)
if err != nil {
t.Fatalf("reading busy_timeout: %v", err)
}
if busyTimeout != busyTimeoutMs {
t.Errorf("busy_timeout = %d, want %d", busyTimeout, busyTimeoutMs)
}
}
// TestNewWriteSucceedsWhileAnotherHandleReads opens the same index twice,
// as a read-only command does while a backup runs, and checks that a write
// on one handle commits while the other is in the middle of a read.
func TestNewWriteSucceedsWhileAnotherHandleReads(t *testing.T) {
t.Parallel()
ctx := context.Background()
dbPath := filepath.Join(t.TempDir(), "index.db")
reader, err := New(ctx, dbPath)
if err != nil {
t.Fatalf("failed to open reading handle: %v", err)
}
defer func() { _ = reader.Close() }()
writer, err := New(ctx, dbPath)
if err != nil {
t.Fatalf("failed to open writing handle: %v", err)
}
defer func() { _ = writer.Close() }()
// The read lock taken by the SELECT is held until the transaction ends.
readTx, err := reader.BeginTx(ctx, nil)
if err != nil {
t.Fatalf("beginning read transaction: %v", err)
}
defer func() { _ = readTx.Rollback() }()
var count int
err = readTx.QueryRowContext(ctx, "SELECT COUNT(*) FROM chunks").Scan(&count)
if err != nil {
t.Fatalf("reading chunks: %v", err)
}
_, err = writer.ExecWithLog(ctx,
"INSERT INTO chunks (chunk_hash, size) VALUES (?, ?)", "hash", 1024)
if err != nil {
t.Fatalf("write while another handle reads: %v", err)
}
}
func TestParseMigrationVersion(t *testing.T) {
t.Parallel()
+1 -1
View File
@@ -253,7 +253,7 @@ func (r *FileChunkRepository) CreateBatch(
args = append(args, fc.FileID.String(), fc.Idx, fc.ChunkHash.String())
}
query += querySb211.String() //nolint:gosec // G202: appends "?" placeholders only
query += querySb211.String()
query += " ON CONFLICT(file_id, idx) DO NOTHING"
+51 -1
View File
@@ -266,6 +266,56 @@ func (r *FileRepository) ListByPrefix(
return files, rows.Err()
}
// ListIDsWithChunksNotInUploadedBlobs returns the IDs of the files whose
// path starts with prefix and that list at least one chunk held by no
// blob whose upload has completed (uploaded_ts set). A new snapshot
// cannot reference such a chunk, so a backup must not treat the file as
// unchanged even when its metadata matches the file on disk.
func (r *FileRepository) ListIDsWithChunksNotInUploadedBlobs(
ctx context.Context, prefix string,
) ([]types.FileID, error) {
query := `
SELECT DISTINCT f.id
FROM files f
JOIN file_chunks fc ON fc.file_id = f.id
WHERE f.path LIKE ? || '%'
AND NOT EXISTS (
SELECT 1
FROM blob_chunks bc
JOIN blobs b ON bc.blob_id = b.id
WHERE bc.chunk_hash = fc.chunk_hash
AND b.uploaded_ts IS NOT NULL
)
`
rows, err := r.db.conn.QueryContext(ctx, query, prefix)
if err != nil {
return nil, fmt.Errorf("querying files: %w", err)
}
defer func() {
err := rows.Close()
if err != nil {
Fatalf("failed to close rows: %v", err)
}
}()
var ids []types.FileID
for rows.Next() {
var id types.FileID
err := rows.Scan(&id)
if err != nil {
return nil, fmt.Errorf("scanning file ID: %w", err)
}
ids = append(ids, id)
}
return ids, rows.Err()
}
// ListAll returns all files in the database
func (r *FileRepository) ListAll(ctx context.Context) ([]*File, error) {
query := `
@@ -341,7 +391,7 @@ func (r *FileRepository) CreateBatch(
f.LinkTarget.String())
}
query += querySb325.String() //nolint:gosec // G202: appends "?" placeholders only
query += querySb325.String()
query += ` ON CONFLICT(path) DO UPDATE SET
source_path = excluded.source_path,
+1 -1
View File
@@ -395,7 +395,7 @@ func (r *SnapshotRepository) AddFilesByIDBatch(
args = append(args, snapshotID, fileID.String())
}
query += querySb312.String() //nolint:gosec // G202: appends "?" placeholders only
query += querySb312.String()
var err error
if tx != nil {
+28 -13
View File
@@ -3,6 +3,7 @@
package globals
import (
"regexp"
"strings"
"time"
)
@@ -10,22 +11,26 @@ import (
// Appname is the application name, populated from main().
var Appname = "vaultik" //nolint:gochecknoglobals // set via -ldflags at build time
// DevVersion is the version a binary reports when it was not built
// from a tagged commit. script/version emits either this exact string
// (outside a git checkout) or this string followed by "-" and the
// commit it was built from, and goreleaser's snapshot template matches
// that shape. It is deliberately not a number: a build that is not a
// release must not name itself like one.
// DevVersion is the version a binary reports when it was built without
// git metadata: script/version emits it outside a git checkout, and an
// unstamped `go build` keeps it. goreleaser's snapshot template stamps
// it followed by "-" and the commit it was built from. It is
// deliberately not a number.
const DevVersion = "dev"
// Unknown is what Commit and CommitDate hold when the build did not
// stamp them, and the version script/docker and script/cibuild stamp
// when the host has no git checkout.
const Unknown = "unknown"
// Version is the application version, populated from main().
var Version = DevVersion //nolint:gochecknoglobals // set via -ldflags at build time
// Commit is the git commit hash, populated from main().
var Commit = "unknown" //nolint:gochecknoglobals // set via -ldflags at build time
var Commit = Unknown //nolint:gochecknoglobals // set via -ldflags at build time
// CommitDate is the ISO-8601 date of the commit, populated from main().
var CommitDate = "unknown" //nolint:gochecknoglobals // set via -ldflags at build time
var CommitDate = Unknown //nolint:gochecknoglobals // set via -ldflags at build time
// Author identifies the upstream author of vaultik.
const Author = "Jeffrey Paul <sneak@sneak.berlin>"
@@ -60,18 +65,28 @@ func New() (*Globals, error) {
}
// IsDevVersion reports whether v names a development build rather than
// a release. Both "dev" and "dev-<sha>" (and its "-dirty" variant)
// count: a caller that compares against "dev" exactly would treat every
// commit-stamped development build as a release.
// a release. "dev" and goreleaser's snapshot "dev-<sha>" count, and so
// does what `git describe --tags --always --dirty` gives a make or
// docker build of an untagged commit: the bare short commit, or
// tag-N-gHASH on a commit after a tag. Any version ending in "-dirty"
// counts, a modified checkout of a tag ("v1.0.0-dirty") included.
// A plain tag such as "v1.0.0" or "1.0.0" is a release.
//
// The empty string counts too. Nothing that knows its version reports
// no version, so an empty Version means the stamping failed, and the
// safe reading of "we could not establish that this is a release" is
// that it is not one. The Makefile refuses to build at all in that
// case; this is the second line of defence, for a binary linked by
// something other than the Makefile.
// something other than the Makefile. Unknown counts for the same
// reason.
func IsDevVersion(v string) bool {
return v == "" || v == DevVersion || strings.HasPrefix(v, DevVersion+"-")
if v == "" || v == Unknown || v == DevVersion ||
strings.HasPrefix(v, DevVersion+"-") ||
strings.HasSuffix(v, "-dirty") {
return true
}
return regexp.MustCompile(`^(.+-[0-9]+-g)?[0-9a-f]+$`).MatchString(v)
}
// shortCommitLen is the number of commit-hash characters ShortCommit keeps.
+19 -6
View File
@@ -34,10 +34,11 @@ func TestGlobalsNew(t *testing.T) {
}
// TestIsDevVersion covers the boundary that matters: everything
// script/version and goreleaser's snapshot template can emit for an
// untagged build must be recognised as a development build, and a real
// tag must not be. A plain equality check against "dev" used to decide
// this, which classified every commit-stamped dev build as a release.
// script/version, a plain docker build and goreleaser's snapshot
// template can emit for an untagged build must be recognised as a
// development build, and a real tag must not be. A plain equality check
// against "dev" used to decide this, which classified every
// commit-stamped dev build as a release.
func TestIsDevVersion(t *testing.T) {
t.Parallel()
@@ -49,8 +50,17 @@ func TestIsDevVersion(t *testing.T) {
{"dev", true},
{"dev-b6e4a218a39e", true},
{"dev-b6e4a218a39e-dirty", true},
// What a tagged build produces (script/version strips the
// leading "v", matching goreleaser's .Version).
// What `git describe --tags --always --dirty` produces with no
// tag reachable, and on a commit after a tag.
{"877eb2f", true},
{"877eb2f-dirty", true},
{"v1.0.0-3-g877eb2f", true},
{"v1.0.0-3-g877eb2f-dirty", true},
{"1.0.0-rc.1-12-g877eb2f", true},
// A tagged commit with uncommitted changes is not that tag.
{"v1.0.0-dirty", true},
// What a tagged build produces (goreleaser's .Version strips
// the leading "v"; script/version keeps it).
{"1.0.0", false},
{"0.1.0", false},
{"1.0.0-rc.1", false},
@@ -64,6 +74,9 @@ func TestIsDevVersion(t *testing.T) {
// as one. The Makefile refuses to build when script/version
// yields nothing; this covers a binary linked some other way.
{"", true},
// What script/docker and script/cibuild stamp when the host
// has no git checkout.
{"unknown", true},
}
for _, tc := range cases {
+11 -1
View File
@@ -6,6 +6,7 @@ import (
"context"
"errors"
"io"
"strings"
"sync/atomic"
"github.com/aws/aws-sdk-go-v2/aws"
@@ -30,6 +31,8 @@ type Client struct {
// Config contains S3 client configuration.
// All fields are required except Prefix, which defaults to an empty string.
// A non-empty Prefix is joined to every key with one "/", whether or not
// it ends with one.
// The Endpoint field should include the protocol (http:// or https://).
type Config struct {
Endpoint string
@@ -75,10 +78,17 @@ func NewClient(ctx context.Context, cfg Config) (*Client, error) {
s3Client := s3.NewFromConfig(awsCfg, s3Opts)
// Every method below builds a key as prefix + key, so the prefix
// must carry its own trailing "/".
prefix := strings.TrimRight(cfg.Prefix, "/")
if prefix != "" {
prefix += "/"
}
return &Client{
s3Client: s3Client,
bucket: cfg.Bucket,
prefix: cfg.Prefix,
prefix: prefix,
endpoint: cfg.Endpoint,
}, nil
}
+16 -8
View File
@@ -1,40 +1,48 @@
//nolint:testpackage // exercises the unexported copyFile helper
//nolint:testpackage // exercises the unexported copyDatabase helper
package snapshot
import (
"context"
"os"
"path/filepath"
"syscall"
"testing"
"github.com/spf13/afero"
"sneak.berlin/go/vaultik/internal/database"
)
// TestCopyFileExportCopyMode verifies that the exported snapshot database
// copy is created owner-only (0600), even under a lenient 022 umask that
// would otherwise leave a fresh file world-readable.
// TestCopyDatabaseExportCopyMode verifies that the exported snapshot
// database copy is created owner-only (0600), even under a lenient 022
// umask that would otherwise leave a fresh file world-readable.
//
//nolint:paralleltest // syscall.Umask is process-global; parallel tests would clash
func TestCopyFileExportCopyMode(t *testing.T) {
func TestCopyDatabaseExportCopyMode(t *testing.T) {
restore := syscall.Umask(0o022)
defer syscall.Umask(restore)
ctx := context.Background()
dir := t.TempDir()
src := filepath.Join(dir, "index.sqlite")
err := os.WriteFile(src, []byte("index data"), 0o600)
db, err := database.New(ctx, src)
if err != nil {
t.Fatalf("creating source index: %v", err)
}
err = db.Close()
if err != nil {
t.Fatalf("closing source index: %v", err)
}
dst := filepath.Join(dir, "snapshot.db")
sm := &SnapshotManager{fs: afero.NewOsFs()}
err = sm.copyFile(src, dst)
err = sm.copyDatabase(ctx, src, dst)
if err != nil {
t.Fatalf("copyFile: %v", err)
t.Fatalf("copyDatabase: %v", err)
}
info, err := os.Stat(dst)
+54 -23
View File
@@ -11,6 +11,7 @@ import (
"runtime"
"strings"
"sync"
"syscall"
"time"
"github.com/dustin/go-humanize"
@@ -73,6 +74,11 @@ type Scanner struct {
knownChunks map[string]struct{}
knownChunksMu sync.RWMutex
// filesToRechunk holds the IDs of known files that list a chunk no
// uploaded blob holds; they are re-chunked even when their metadata
// is unchanged.
filesToRechunk map[types.FileID]struct{}
// Pending chunk hashes - chunks that have been added to packer but not
// yet committed to DB. When a blob finalizes, the committed chunks are
// removed from this set.
@@ -325,9 +331,42 @@ func (s *Scanner) loadDatabaseState(
s.ui.Completef("Loaded %s known chunks from local index database.",
s.ui.Count(len(s.knownChunks)))
err = s.loadFilesToRechunk(ctx, path)
if err != nil {
return nil, fmt.Errorf("loading files to re-chunk: %w", err)
}
return knownFiles, nil
}
// loadFilesToRechunk loads the IDs of known files under path that list a
// chunk no uploaded blob holds. A file row is shared by every snapshot
// that lists the file and is updated in place when the file changes,
// while a blob row is deleted once no snapshot references it. Removing
// the only snapshot that references a changed file's current blob
// therefore leaves an older snapshot keeping a file row whose metadata
// matches the disk while no blob the local index records as uploaded
// holds its chunks, so a new snapshot cannot reference them. The dropped
// blob can still be in remote storage until prune removes it.
func (s *Scanner) loadFilesToRechunk(ctx context.Context, path string) error {
ids, err := s.repos.Files.ListIDsWithChunksNotInUploadedBlobs(ctx, path)
if err != nil {
return fmt.Errorf("listing files: %w", err)
}
s.filesToRechunk = make(map[types.FileID]struct{}, len(ids))
for _, id := range ids {
s.filesToRechunk[id] = struct{}{}
}
if len(ids) > 0 {
log.Info("Re-chunking known files whose chunks are not all "+
"in uploaded blobs", "files", len(ids))
}
return nil
}
// repairInterruptedBlobs discards blob rows left by a previous run whose
// upload never completed. Such a blob has its chunks, blob_chunks, and
// blobs rows committed to the local index before the upload is attempted,
@@ -1104,12 +1143,9 @@ func (s *Scanner) buildSymlinkEntry(path string, info os.FileInfo) *database.Fil
}
var uid, gid uint32
if stat, ok := info.Sys().(interface {
Uid() uint32
Gid() uint32
}); ok {
uid = stat.Uid()
gid = stat.Gid()
if stat, ok := info.Sys().(*syscall.Stat_t); ok {
uid = stat.Uid
gid = stat.Gid
}
return &database.File{
@@ -1128,12 +1164,9 @@ func (s *Scanner) buildSymlinkEntry(path string, info os.FileInfo) *database.Fil
// buildDirectoryEntry creates a File record for a directory.
func (s *Scanner) buildDirectoryEntry(path string, info os.FileInfo) *database.File {
var uid, gid uint32
if stat, ok := info.Sys().(interface {
Uid() uint32
Gid() uint32
}); ok {
uid = stat.Uid()
gid = stat.Gid()
if stat, ok := info.Sys().(*syscall.Stat_t); ok {
uid = stat.Uid
gid = stat.Gid
}
return &database.File{
@@ -1166,16 +1199,10 @@ func (s *Scanner) recordNonRegularFile(ctx context.Context, ftp *FileToProcess)
func (s *Scanner) checkFileInMemory(
path string, info os.FileInfo, knownFiles map[string]*database.File,
) (*database.File, bool) {
// Get file stats
stat, ok := info.Sys().(interface {
Uid() uint32
Gid() uint32
})
var uid, gid uint32
if ok {
uid = stat.Uid()
gid = stat.Gid()
if stat, ok := info.Sys().(*syscall.Stat_t); ok {
uid = stat.Uid
gid = stat.Gid
}
// Check against in-memory map first to get existing ID if available
@@ -1208,6 +1235,11 @@ func (s *Scanner) checkFileInMemory(
return file, true
}
// No uploaded blob holds one of its chunks (see loadFilesToRechunk)
if _, rechunk := s.filesToRechunk[fileID]; rechunk {
return file, true
}
// Check if file has changed
if existingFile.Size != file.Size ||
existingFile.MTime.Unix() != file.MTime.Unix() ||
@@ -1345,8 +1377,7 @@ func (s *Scanner) processFileWithErrorHandling(
// record a file whose chunk is in no blob and cannot be restored, so
// abort the run even under --skip-errors. Only open and read errors
// are skipped below.
var pErr *packerError
if errors.As(err, &pErr) {
if _, ok := errors.AsType[*packerError](err); ok {
return false, fmt.Errorf("processing file %s: %w", fileToProcess.Path, err)
}
// Handle files that were deleted between scan and process phases
+77
View File
@@ -307,3 +307,80 @@ func TestScannerLargeFile(t *testing.T) {
}
}
}
// TestScannerRecordsOwnership backs up real files on disk and checks that
// the uid and gid of a file, a directory and a symlink are recorded.
// When the tests run as root both sides are 0, so only a run as another
// user can catch ownership recorded as 0.
func TestScannerRecordsOwnership(t *testing.T) {
t.Parallel()
sourceDir := t.TempDir()
filePath := filepath.Join(sourceDir, "file.txt")
dirPath := filepath.Join(sourceDir, "subdir")
linkPath := filepath.Join(sourceDir, "link")
err := os.WriteFile(filePath, []byte("owned"), 0o600)
if err != nil {
t.Fatal(err)
}
err = os.Mkdir(dirPath, 0o700)
if err != nil {
t.Fatal(err)
}
err = os.Symlink("file.txt", linkPath)
if err != nil {
t.Fatal(err)
}
db, err := database.NewTestDB()
if err != nil {
t.Fatalf("failed to create test database: %v", err)
}
defer func() {
err := db.Close()
if err != nil {
t.Errorf("failed to close database: %v", err)
}
}()
repos := database.NewRepositories(db)
scanner := snapshot.NewScanner(snapshot.ScannerConfig{
FS: afero.NewOsFs(),
ChunkSize: int64(1024 * 16),
Repositories: repos,
MaxBlobSize: int64(1024 * 1024),
CompressionLevel: 3,
AgeRecipients: []string{testAgePublicKey},
})
ctx := context.Background()
snapshotID := "test-snapshot-ownership"
createTestSnapshotRecord(ctx, t, repos, snapshotID)
_, err = scanner.Scan(ctx, sourceDir, snapshotID)
if err != nil {
t.Fatalf("scan failed: %v", err)
}
for _, path := range []string{filePath, dirPath, linkPath} {
file, err := repos.Files.GetByPath(ctx, path)
if err != nil {
t.Fatalf("failed to get %s: %v", path, err)
}
if file == nil {
t.Fatalf("%s was not recorded", path)
}
if int(file.UID) != os.Getuid() || int(file.GID) != os.Getgid() {
t.Errorf("%s recorded as uid %d gid %d, want uid %d gid %d",
path, file.UID, file.GID, os.Getuid(), os.Getgid())
}
}
}
+33 -38
View File
@@ -386,12 +386,12 @@ func (sm *SnapshotManager) prepareExportDB(
ctx context.Context, dbPath, snapshotID, tempDir string,
) ([]byte, string, error) {
// Step 1: Copy database to temp file
// The main database should be closed at this point
// The main database is still open here, so it is copied through SQLite
tempDBPath := filepath.Join(tempDir, "snapshot.db")
log.Debug("Copying database to temporary location",
"source", dbPath, "destination", tempDBPath)
err := sm.copyFile(dbPath, tempDBPath)
err := sm.copyDatabase(ctx, dbPath, tempDBPath)
if err != nil {
return nil, "", fmt.Errorf("copying database: %w", err)
}
@@ -648,9 +648,10 @@ func (sm *SnapshotManager) collectCleanupStats(
//
// VACUUM runs through the modernc.org/sqlite driver, on a freshly opened
// connection with no transaction in flight (VACUUM cannot run inside one).
// The database opens in WAL mode, so VACUUM's rewrite lands in the WAL; the
// checkpoint on Close flushes it into the main file, which is the file we
// then compress and upload.
// database.New opens the file in WAL mode, so VACUUM's rewrite lands in the
// -wal file. This is the only connection to the file, so closing it
// checkpoints the rewrite into the main file and removes the -wal file; the
// main file is the one compressFile then compresses and uploads.
func (sm *SnapshotManager) vacuumDatabase(ctx context.Context, dbPath string) error {
log.Debug("Running VACUUM on database", "path", dbPath)
@@ -744,26 +745,15 @@ func (sm *SnapshotManager) compressFile(inputPath, outputPath string) error {
// user; it holds the same private index data as the local index file.
const exportCopyPerm = 0o600
// copyFile copies a file from src to dst. The destination is the exported
// snapshot database, so it is created owner-only rather than with the
// umask-dependent default.
func (sm *SnapshotManager) copyFile(src, dst string) error {
log.Debug("Opening source file for copy", "path", src)
sourceFile, err := sm.fs.Open(src)
if err != nil {
return err
}
defer func() {
log.Debug("Closing source file", "path", src)
err := sourceFile.Close()
if err != nil {
log.Debug("Failed to close source file", "path", src, "error", err)
}
}()
// copyDatabase copies the database at src to dst with VACUUM INTO. It reads
// through SQLite, so the copy holds rows committed to src that are still in
// its -wal file, which a copy of the file alone would miss. The destination
// is the exported snapshot database, so it is created empty and owner-only
// first rather than with the umask-dependent default: VACUUM INTO writes
// into an existing empty file and keeps its mode.
func (sm *SnapshotManager) copyDatabase(
ctx context.Context, src, dst string,
) error {
log.Debug("Creating destination file", "path", dst)
destFile, err := sm.fs.OpenFile(
@@ -773,23 +763,28 @@ func (sm *SnapshotManager) copyFile(src, dst string) error {
return err
}
defer func() {
log.Debug("Closing destination file", "path", dst)
err := destFile.Close()
if err != nil {
log.Debug("Failed to close destination file", "path", dst, "error", err)
}
}()
log.Debug("Copying file data")
n, err := io.Copy(destFile, sourceFile)
err = destFile.Close()
if err != nil {
return err
}
log.Debug("File copy complete", "bytes_copied", n)
db, err := database.New(ctx, src)
if err != nil {
return fmt.Errorf("opening database to copy: %w", err)
}
defer func() {
cerr := db.Close()
if cerr != nil {
log.Debug("Failed to close database after copy",
"path", src, "error", cerr)
}
}()
_, err = db.ExecWithLog(ctx, "VACUUM INTO ?", dst)
if err != nil {
return fmt.Errorf("running VACUUM INTO: %w", err)
}
return nil
}
+71
View File
@@ -188,6 +188,77 @@ func TestVacuumDatabaseRemovesDeletedData(t *testing.T) {
}
}
// TestPrepareExportDBKeepsRowsCommittedToOpenIndex exports from an index
// that is still open, as a backup does. A row committed there can still be
// in the index's -wal file, and the export must hold it all the same.
func TestPrepareExportDBKeepsRowsCommittedToOpenIndex(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
ctx := context.Background()
fs := afero.NewOsFs()
dbPath := filepath.Join(t.TempDir(), "index.sqlite")
db, err := database.New(ctx, dbPath)
if err != nil {
t.Fatalf("failed to create database: %v", err)
}
defer func() { _ = db.Close() }()
repos := database.NewRepositories(db)
snapshot := &database.Snapshot{ID: "open-index-snapshot", Hostname: "test-host"}
err = repos.WithTx(ctx, func(ctx context.Context, tx *sql.Tx) error {
return repos.Snapshots.Create(ctx, tx, snapshot)
})
if err != nil {
t.Fatalf("failed to create snapshot: %v", err)
}
sm := &SnapshotManager{
config: &config.Config{
CompressionLevel: 3,
AgeRecipients: []string{testAgeRecipient},
},
fs: fs,
}
_, tempDBPath, err := sm.prepareExportDB(
ctx, dbPath, snapshot.ID.String(), t.TempDir())
if err != nil {
t.Fatalf("prepareExportDB failed: %v", err)
}
// Only the main database file is compressed and uploaded, so open a
// copy of that file alone.
uploadedPath := filepath.Join(t.TempDir(), "uploaded.db")
err = copyFile(fs, tempDBPath, uploadedPath)
if err != nil {
t.Fatalf("failed to copy exported database: %v", err)
}
exported, err := database.OpenReadOnly(ctx, uploadedPath)
if err != nil {
t.Fatalf("failed to open exported database: %v", err)
}
defer func() { _ = exported.Close() }()
got, err := database.NewRepositories(exported).Snapshots.GetByID(
ctx, snapshot.ID.String())
if err != nil {
t.Fatalf("failed to read snapshot from export: %v", err)
}
if got == nil {
t.Fatal("exported database is missing the snapshot row")
}
}
func TestCleanSnapshotDBEmptySnapshot(t *testing.T) {
// Initialize logger
log.Initialize(log.Config{})
+23 -8
View File
@@ -22,11 +22,11 @@ type FileStorer struct {
//
// Construction is intentionally cheap and does not touch the filesystem.
// The basePath is recorded; the directory is created lazily on first
// write. Reads (Get/Stat/List) tolerate a missing basePath — a missing
// or unmounted destination during `snapshot list` should NOT block the
// command, it should degrade to "no remote snapshots reachable" with a
// warning. Write operations (Put/PutWithProgress) call MkdirAll for the
// write. Write operations (Put/PutWithProgress) call MkdirAll for the
// per-blob parent directory, which also covers basePath on first use.
// Get and Stat report a key under a missing basePath as ErrNotFound.
// List and ListStream fail on a missing basePath, because listing it as
// an empty store would make `prune` drop every local snapshot record.
//
// Uses the real OS filesystem by default; call SetFilesystem to
// override for testing.
@@ -119,13 +119,20 @@ func (f *FileStorer) Delete(_ context.Context, key string) error {
return nil
}
// List returns all keys with the given prefix.
// List returns all keys with the given prefix. It fails when the
// destination directory is missing; a missing prefix under it is an
// empty listing.
func (f *FileStorer) List(ctx context.Context, prefix string) ([]string, error) {
var keys []string
_, err := f.fs.Stat(f.basePath)
if err != nil {
return nil, fmt.Errorf("checking destination directory: %w", err)
}
basePath := f.fullPath(prefix)
// Check if base path exists
// Check if the prefix exists
exists, err := afero.Exists(f.fs, basePath)
if err != nil {
return nil, fmt.Errorf("checking path: %w", err)
@@ -167,16 +174,24 @@ func (f *FileStorer) List(ctx context.Context, prefix string) ([]string, error)
return keys, nil
}
// ListStream returns a channel of ObjectInfo for large result sets.
// ListStream returns a channel of ObjectInfo for large result sets. Like
// List, it sends an error when the destination directory is missing.
func (f *FileStorer) ListStream(ctx context.Context, prefix string) <-chan ObjectInfo {
ch := make(chan ObjectInfo)
go func() {
defer close(ch)
_, err := f.fs.Stat(f.basePath)
if err != nil {
ch <- ObjectInfo{Err: fmt.Errorf("checking destination directory: %w", err)}
return
}
basePath := f.fullPath(prefix)
// Check if base path exists
// Check if the prefix exists
exists, err := afero.Exists(f.fs, basePath)
if err != nil {
ch <- ObjectInfo{Err: fmt.Errorf("checking path: %w", err)}
+35
View File
@@ -1,6 +1,10 @@
package storage_test
import (
"context"
"errors"
"io/fs"
"path/filepath"
"testing"
"sneak.berlin/go/vaultik/internal/storage"
@@ -25,3 +29,34 @@ func TestFileStorer(t *testing.T) {
t.Parallel()
runStorerConformance(t, newFileStorer)
}
// TestFileStorerListMissingDestination checks that List and ListStream
// fail when the destination directory does not exist, instead of
// reporting an empty store.
func TestFileStorerListMissingDestination(t *testing.T) {
t.Parallel()
ctx := context.Background()
s, err := storage.NewFileStorer(filepath.Join(t.TempDir(), "unmounted"))
if err != nil {
t.Fatalf("NewFileStorer: %v", err)
}
keys, err := s.List(ctx, "metadata/")
if !errors.Is(err, fs.ErrNotExist) {
t.Errorf("List = %v, %v; want a not-exist error", keys, err)
}
var streamErr error
for object := range s.ListStream(ctx, "metadata/") {
if object.Err != nil {
streamErr = object.Err
}
}
if !errors.Is(streamErr, fs.ErrNotExist) {
t.Errorf("ListStream error = %v, want a not-exist error", streamErr)
}
}
+84
View File
@@ -4,11 +4,14 @@ import (
"context"
"errors"
"net/http/httptest"
"slices"
"strings"
"testing"
"github.com/johannesboyne/gofakes3"
"github.com/johannesboyne/gofakes3/backend/s3mem"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/s3"
"sneak.berlin/go/vaultik/internal/storage"
)
@@ -79,3 +82,84 @@ func TestS3StorerMissingKeyMapsToErrNotFound(t *testing.T) {
t.Errorf("Stat on missing key: got %v, want ErrNotFound", err)
}
}
// TestS3URLPrefixKeyLayout pins the bucket keys an s3:// URL reads and
// writes: the README's remote storage layout, with the prefix joined to
// each key by one "/". s3://b/p and s3://b/p/ must be the same
// destination, or a host that writes the URL the other way finds no
// snapshots. The listed object is put straight into the bucket, as
// another host would have written it.
func TestS3URLPrefixKeyLayout(t *testing.T) {
t.Parallel()
const (
blobKey = "blobs/aa/bb/aabbccdd"
manifestKey = "metadata/snap/manifest.json.zst"
manifestBody = "manifest"
)
cases := []struct {
urlPath string // URL path after the bucket name
keyPrefix string // what every key in the bucket must start with
}{
{urlPath: "/p", keyPrefix: "p/"},
{urlPath: "/p/", keyPrefix: "p/"},
{urlPath: "", keyPrefix: ""},
}
for _, tc := range cases {
storageURL := "s3://" + s3TestBucket + tc.urlPath
t.Run(storageURL, func(t *testing.T) {
t.Parallel()
backend := s3mem.New()
err := backend.CreateBucket(s3TestBucket)
if err != nil {
t.Fatalf("create bucket: %v", err)
}
srv := httptest.NewServer(gofakes3.New(backend).Server())
t.Cleanup(srv.Close)
storer, err := storage.NewStorer(&config.Config{
StorageURL: storageURL + "?endpoint=" + srv.URL,
S3: config.S3Config{
AccessKeyID: "key",
SecretAccessKey: "secret",
},
})
if err != nil {
t.Fatalf("NewStorer: %v", err)
}
ctx := context.Background()
err = storer.Put(ctx, blobKey, strings.NewReader("blob"))
if err != nil {
t.Fatalf("Put: %v", err)
}
_, err = backend.HeadObject(s3TestBucket, tc.keyPrefix+blobKey)
if err != nil {
t.Errorf("blob not stored at %q: %v", tc.keyPrefix+blobKey, err)
}
_, err = backend.PutObject(s3TestBucket, tc.keyPrefix+manifestKey,
nil, strings.NewReader(manifestBody), int64(len(manifestBody)))
if err != nil {
t.Fatalf("seed manifest: %v", err)
}
keys, err := storer.List(ctx, "metadata/")
if err != nil {
t.Fatalf("List: %v", err)
}
if !slices.Equal(keys, []string{manifestKey}) {
t.Errorf("List(metadata/) = %q, want [%q]", keys, manifestKey)
}
})
}
}
+1 -2
View File
@@ -161,8 +161,7 @@ func rejectUnknownParams(query url.Values, allowed ...string) error {
// *url.Error that url.Parse returns embeds the raw URL in its message, so
// wrapping it directly would echo a credential-bearing URL into logs.
func wrapParseError(err error) error {
var uerr *url.Error
if errors.As(err, &uerr) {
if uerr, ok := errors.AsType[*url.Error](err); ok {
return fmt.Errorf("invalid URL: %w", uerr.Err)
}
@@ -0,0 +1,234 @@
package vaultik_test
import (
"context"
"crypto/rand"
"path/filepath"
"slices"
"strings"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/chunker"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/storage/faultstore"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// These tests cover https://git.eeqj.de/sneak/vaultik/issues/214: once
// the only snapshot referencing a changed file's current blob is
// dropped, an older snapshot still keeps the file's row, and the next
// backup must re-chunk the file instead of treating it as unchanged.
//
// Each backup run uses its own snapshot name over the same source
// directory, so the snapshot IDs differ without waiting for their
// one-second timestamps to tick over. File rows are keyed by path, so
// the runs share them exactly as runs of one snapshot name would.
// changedFileConfig returns a full-backup config with the snapshot names
// "first", "second" and "third", all backing up dataDir.
func changedFileConfig(dataDir, dbPath string) *config.Config {
source := config.SnapshotConfig{Paths: []string{dataDir}}
cfg := faultTestConfig()
cfg.IndexPath = dbPath
cfg.ChunkSize = config.Size(faultChunkSize)
cfg.Snapshots = map[string]config.SnapshotConfig{
"first": source,
"second": source,
"third": source,
}
return cfg
}
// backUp runs a full backup of the named snapshot.
func backUp(v *vaultik.Vaultik, name string) error {
return v.CreateSnapshot(&vaultik.SnapshotCreateOptions{
Cron: true,
Snapshots: []string{name},
})
}
// appendedFileSize is the size of the file before the tests append to it.
// It is twice the largest chunk the chunker cuts, so the file's first
// chunk ends before the appended bytes and stays in a blob the first
// snapshot still references, while its new chunks are only in the blob
// that gets dropped.
const appendedFileSize = 2 * chunker.ChunkSizeSpread * faultChunkSize
// appendRandomBytes appends n random bytes to the file at path, creating
// the file if it does not exist, and records the new content in files.
// Random content gives the chunker real cut points.
func appendRandomBytes(
t *testing.T, fs afero.Fs, files map[string][]byte, path string, n int64,
) {
t.Helper()
added := make([]byte, n)
_, err := rand.Read(added)
require.NoError(t, err)
content := slices.Concat(files[path], added)
require.NoError(t, afero.WriteFile(fs, path, content, 0o644))
files[path] = content
}
// localSnapshotID returns the ID of the one snapshot in the local index
// that was backed up under name.
func localSnapshotID(
ctx context.Context, t *testing.T,
repos *database.Repositories, name string,
) string {
t.Helper()
snapshots, err := repos.Snapshots.ListRecent(ctx, listRecentTestLimit)
require.NoError(t, err)
var ids []string
for _, s := range snapshots {
if strings.Contains(s.ID.String(), "_"+name+"_") {
ids = append(ids, s.ID.String())
}
}
require.Lenf(t, ids, 1, "expected one local snapshot named %q", name)
return ids[0]
}
// assertThirdSnapshotRestores closes the local index, then restores the
// snapshot named "third" from store alone and byte-compares every file
// against files.
func assertThirdSnapshotRestores(
ctx context.Context, t *testing.T, cfg *config.Config,
store storage.Storer, repos *database.Repositories, db *database.DB,
fs afero.Fs, restoreDir string, files map[string][]byte,
) {
t.Helper()
id := localSnapshotID(ctx, t, repos, "third")
require.NoError(t, db.Close())
reader := newReaderVaultik(ctx, cfg, store, nil, fs)
require.NoError(t, reader.Restore(&vaultik.RestoreOptions{
SnapshotID: id,
TargetDir: restoreDir,
Verify: true,
}), "the backup after the drop must restore the changed file")
assertRestoredTree(t, fs, restoreDir, files)
}
// Trigger 1: bytes are appended to the file, a second snapshot backs it
// up, and that snapshot is removed. The first snapshot keeps the file row,
// which now lists the appended content's chunks, while removal drops the
// blob that held them.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestBackupAfterRemovingNewestSnapshotRestoresChangedFile(t *testing.T) {
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
storeDir := filepath.Join(tempDir, "remote")
restoreDir := filepath.Join(tempDir, "restored")
dbPath := filepath.Join(tempDir, "index.sqlite")
changedPath := filepath.Join(dataDir, "appended.bin")
ctx := context.Background()
files := writeFaultSourceTree(t, fs, dataDir)
cfg := changedFileConfig(dataDir, dbPath)
appendRandomBytes(t, fs, files, changedPath, appendedFileSize)
store, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
repos := database.NewRepositories(db)
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
require.NoError(t, backUp(v, "first"))
appendRandomBytes(t, fs, files, changedPath, faultChunkSize)
require.NoError(t, backUp(v, "second"))
_, err = v.RemoveSnapshot(localSnapshotID(ctx, t, repos, "second"),
&vaultik.RemoveOptions{Force: true})
require.NoError(t, err)
require.NoError(t, backUp(v, "third"))
assertThirdSnapshotRestores(
ctx, t, cfg, store, repos, db, fs, restoreDir, files)
}
// Trigger 2: bytes are appended to the file and the run that backs it up
// is interrupted at the manifest upload, after its blobs were uploaded.
// The next run's prune drops that incomplete snapshot and its blob, while
// the first snapshot keeps the file row, which now lists the appended
// content's chunks.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestBackupAfterInterruptedRunRestoresChangedFile(t *testing.T) {
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
storeDir := filepath.Join(tempDir, "remote")
restoreDir := filepath.Join(tempDir, "restored")
dbPath := filepath.Join(tempDir, "index.sqlite")
changedPath := filepath.Join(dataDir, "appended.bin")
ctx := context.Background()
files := writeFaultSourceTree(t, fs, dataDir)
cfg := changedFileConfig(dataDir, dbPath)
appendRandomBytes(t, fs, files, changedPath, appendedFileSize)
inner, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
failManifest := false
store := faultstore.New(inner)
store.OnPut = func(key string) faultstore.PutAction {
if failManifest && strings.HasSuffix(key, "manifest.json.zst") {
return faultstore.PutFail
}
return faultstore.PutNormal
}
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
repos := database.NewRepositories(db)
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
require.NoError(t, backUp(v, "first"))
appendRandomBytes(t, fs, files, changedPath, faultChunkSize)
failManifest = true
require.Error(t, backUp(v, "second"),
"the run must fail when its manifest upload fails")
failManifest = false
require.NoError(t, backUp(v, "third"))
assertThirdSnapshotRestores(
ctx, t, cfg, inner, repos, db, fs, restoreDir, files)
}
+22 -21
View File
@@ -301,6 +301,10 @@ func TestInterruptedBlobUploadRecordsNoUploadedBlob(t *testing.T) {
writeFaultSourceTree(t, fs, dataDir)
// No upload succeeds, so nothing creates the destination directory.
// It must exist for the listing below to show that no blob survived.
require.NoError(t, os.Mkdir(storeDir, 0o750))
inner, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
@@ -680,18 +684,10 @@ func faultScannerFactory(
// Scenario 5: the restore target runs out of space mid-file. Restore
// must fail with an out-of-space error, and must not leave a truncated
// file at the target path presenting as a complete restore. Restore
// today writes each file straight to its final path and does not remove
// it when a write fails, so the truncated file survives; deleting it is
// tracked by https://git.eeqj.de/sneak/vaultik/issues/163. Skipped until
// that lands, so the destination assertion below is recorded rather than
// dropped.
// file at the target path presenting as a complete restore.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestRestoreReportsDiskFull(t *testing.T) {
t.Skip("blocked on https://git.eeqj.de/sneak/vaultik/issues/163: " +
"a disk-full write leaves a truncated file at the target path " +
"instead of removing it")
log.Initialize(log.Config{})
osFS := afero.NewOsFs()
@@ -717,10 +713,10 @@ func TestRestoreReportsDiskFull(t *testing.T) {
id := fullFaultBackup(ctx, t, osFS, inner, cfg, repos, dataDir, dbPath, "diskfull")
require.NoError(t, db.Close())
// Restore onto a filesystem that allows only a few bytes of file
// content: enough to create files, far too little to hold them.
// Restore onto a target that allows only a few bytes of file content:
// enough to create files, far too little to hold them.
budget := int64(8)
quota := &quotaFS{Fs: osFS, remaining: &budget}
quota := &quotaFS{Fs: osFS, dir: restoreDir, remaining: &budget}
v := newReaderVaultik(ctx, cfg, inner, nil, quota)
err = v.Restore(&vaultik.RestoreOptions{SnapshotID: id, TargetDir: restoreDir})
@@ -754,21 +750,26 @@ func assertRestoredTree(
// budget is exhausted, mirroring a real ENOSPC.
var errNoSpace = errors.New("no space left on device")
// quotaFS is an afero.Fs whose files may write only a fixed total number
// of content bytes before failing, simulating a full restore target. It
// wraps the interface so every method except Create delegates to the
// real filesystem; only file writes are capped.
// quotaFS is an afero.Fs on which files opened under dir may write only
// a fixed total number of content bytes before failing, simulating a full
// restore target. Every other method, and every file outside dir, goes
// straight to the real filesystem: restore also writes the decrypted
// metadata database under $TMPDIR through this filesystem, and capping
// that would fail the restore before it wrote anything to the target.
type quotaFS struct {
afero.Fs
dir string
remaining *int64
}
//nolint:ireturn // afero.Fs.Create's signature requires returning afero.File.
func (q *quotaFS) Create(name string) (afero.File, error) {
f, err := q.Fs.Create(name)
if err != nil {
return nil, err
//nolint:ireturn // afero.Fs.OpenFile's signature requires returning afero.File.
func (q *quotaFS) OpenFile(
name string, flag int, perm os.FileMode,
) (afero.File, error) {
f, err := q.Fs.OpenFile(name, flag, perm)
if err != nil || !strings.HasPrefix(name, q.dir) {
return f, err
}
return &quotaFile{File: f, remaining: q.remaining}, nil
+15 -5
View File
@@ -140,24 +140,34 @@ func parseSnapshotName(snapshotID string) string {
// parseDuration parses a duration string with support for human-friendly units:
// d/day/days, w/week/weeks, mo/month/months, y/year/years, plus standard Go
// duration units. Following Go, m is minutes and mo is months. A bare number,
// an unknown unit, and a negative value are all rejected.
// an unknown unit, and a negative value are all rejected. Outside the Go
// units the input must be whole numbers each followed directly by a unit
// (2w3d), so 1.0y and 30 days are errors; decimals work only in Go units
// (1.5h).
func parseDuration(s string) (time.Duration, error) {
if strings.HasPrefix(strings.TrimSpace(s), "-") {
return 0, errNegativeDuration
}
// A bare number has no unit, but time.ParseDuration accepts 0, which
// would put the cutoff at now and select every snapshot.
_, err := strconv.ParseFloat(s, 64)
if err == nil {
return 0, fmt.Errorf("%w: %q", errInvalidDuration, s)
}
d, err := time.ParseDuration(s)
if err == nil {
return d, nil
}
re := regexp.MustCompile(`(\d+)\s*([a-zA-Z]+)`)
matches := re.FindAllStringSubmatch(s, -1)
if len(matches) == 0 {
if !regexp.MustCompile(`^(\d+[a-zA-Z]+)+$`).MatchString(s) {
return 0, fmt.Errorf("%w: %q", errInvalidDuration, s)
}
re := regexp.MustCompile(`(\d+)([a-zA-Z]+)`)
matches := re.FindAllStringSubmatch(s, -1)
var total time.Duration
for _, match := range matches {
+11
View File
@@ -59,6 +59,7 @@ func TestParseDuration(t *testing.T) {
{"30s", 30 * time.Second, false},
{"6m", 6 * time.Minute, false},
{"1h", time.Hour, false},
{"1.5h", 90 * time.Minute, false},
// Extended calendar units.
{"30d", 30 * 24 * time.Hour, false},
{"3days", 3 * 24 * time.Hour, false},
@@ -73,11 +74,21 @@ func TestParseDuration(t *testing.T) {
{"1y6mo", 365*24*time.Hour + 180*24*time.Hour, false},
// Rejected inputs.
{"6", 0, true}, // bare number, no unit
{"0", 0, true}, // bare number; time.ParseDuration accepts it
{"+0", 0, true}, // bare number; time.ParseDuration accepts it
{"5x", 0, true}, // unknown unit
{"-5d", 0, true}, // negative, extended unit
{"-5h", 0, true}, // negative, Go unit
{"", 0, true}, // empty
{"garbage", 0, true},
// Characters outside the number-and-unit parts must not be
// skipped: 1.0y read as 0y would select every snapshot.
{"1.0y", 0, true},
{"2.1w", 0, true},
{"1,5y", 0, true},
{"30 days", 0, true},
{"30 days ago", 0, true},
{"x7d", 0, true},
}
for _, tt := range tests {
@@ -0,0 +1,155 @@
package vaultik_test
import (
"bytes"
"context"
"io/fs"
"os"
"path/filepath"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/ui"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// These tests cover https://git.eeqj.de/sneak/vaultik/issues/220: a
// file:// destination whose directory is missing, such as a USB stick
// that is not plugged in, cannot be listed. It is not an empty store, so
// no command may conclude from it that the local snapshots are gone.
// backUpToFileDestination backs up the snapshot named "first" to a
// file:// destination at storeDir, which need not exist yet. Everything
// the returned Vaultik prints after the backup goes to the returned
// buffer.
func backUpToFileDestination(
ctx context.Context, t *testing.T, storeDir string,
) (*vaultik.Vaultik, *database.Repositories, *bytes.Buffer) {
t.Helper()
osFs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
dbPath := filepath.Join(tempDir, "index.sqlite")
writeFaultSourceTree(t, osFs, dataDir)
store, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
t.Cleanup(func() { _ = db.Close() })
repos := database.NewRepositories(db)
cfg := changedFileConfig(dataDir, dbPath)
v := newBackupVaultik(ctx, cfg, store, repos, db, osFs)
require.NoError(t, backUp(v, "first"))
out := &bytes.Buffer{}
v.Stdout = out
v.UI = ui.NewWithColor(out, false)
return v, repos, out
}
// backUpThenUnplug backs up to a file:// destination, then moves the
// destination directory away, as unplugging the volume it lives on would.
func backUpThenUnplug(
ctx context.Context, t *testing.T,
) (*vaultik.Vaultik, *database.Repositories, *bytes.Buffer) {
t.Helper()
storeDir := filepath.Join(t.TempDir(), "usbstick")
v, repos, out := backUpToFileDestination(ctx, t, storeDir)
require.NoError(t, os.Rename(storeDir, storeDir+"-unplugged"))
return v, repos, out
}
// TestFirstBackupCreatesDestinationDirectory checks that a first backup
// to a destination directory that does not exist yet creates it, and
// that the destination can be listed afterwards.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestFirstBackupCreatesDestinationDirectory(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
storeDir := filepath.Join(t.TempDir(), "volume", "backup")
v, _, out := backUpToFileDestination(ctx, t, storeDir)
require.NoError(t, v.ListSnapshots(false))
assert.NotContains(t, out.String(), "Could not list backup destination store")
assert.NotContains(t, out.String(), "not found in backup destination store")
}
// TestListSnapshotsWarnsWhenDestinationMissing checks that snapshot list
// warns and shows the local index alone, without reporting the local
// snapshot as missing from the destination.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestListSnapshotsWarnsWhenDestinationMissing(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
v, repos, out := backUpThenUnplug(ctx, t)
id := localSnapshotID(ctx, t, repos, "first")
require.NoError(t, v.ListSnapshots(false))
assert.Contains(t, out.String(), "Could not list backup destination store")
assert.Contains(t, out.String(), "Showing snapshots from the local index only.")
assert.Contains(t, out.String(), id)
assert.NotContains(t, out.String(), "not found in backup destination store")
}
// TestRemoveSnapshotWarnsWhenDestinationMissing checks that snapshot
// remove warns that the metadata could not be removed from the
// destination, instead of reporting that it was.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestRemoveSnapshotWarnsWhenDestinationMissing(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
v, repos, out := backUpThenUnplug(ctx, t)
result, err := v.RemoveSnapshot(localSnapshotID(ctx, t, repos, "first"),
&vaultik.RemoveOptions{Force: true})
require.NoError(t, err)
assert.False(t, result.RemoteRemoved)
assert.Contains(t, out.String(),
"Could not remove snapshot metadata from remote")
assert.NotContains(t, out.String(),
"Removed snapshot metadata from remote storage")
}
// TestPruneKeepsLocalRecordsWhenDestinationMissing checks that prune
// fails on a destination it cannot list and deletes no local snapshot
// record.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestPruneKeepsLocalRecordsWhenDestinationMissing(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
v, repos, _ := backUpThenUnplug(ctx, t)
err := v.Prune(&vaultik.PruneOptions{Force: true})
require.ErrorIs(t, err, fs.ErrNotExist)
require.ErrorContains(t, err, "listing remote snapshots")
snapshots, err := repos.Snapshots.ListRecent(ctx, listRecentTestLimit)
require.NoError(t, err)
assert.Len(t, snapshots, 1, "prune must delete no local snapshot record")
}
+27 -3
View File
@@ -355,7 +355,7 @@ func (v *Vaultik) runRestoreLoop(
fileID, ready := plan.popReady()
if !ready {
downloaded, err := session.downloadNextBlobSet(plan)
downloaded, err := session.downloadNextBlobSet(plan, filesByID)
if err != nil {
return err
}
@@ -411,7 +411,13 @@ func (v *Vaultik) runRestoreLoop(
// blob set and downloads its blobs; after each blob lands, the plan
// moves any pending file whose set just emptied onto the ready queue.
// Returns false when nothing is pending download (the caller stops).
func (s *restoreSession) downloadNextBlobSet(plan *restorePlan) (bool, error) {
//
// A blob that cannot be downloaded is reported through
// handleRestoreFileError for every pending file that references it, so
// it aborts the restore unless --skip-errors is set.
func (s *restoreSession) downloadNextBlobSet(
plan *restorePlan, filesByID map[types.FileID]*database.File,
) (bool, error) {
s.sweeper.sweep()
next, ok := plan.pickNextDownload()
@@ -433,7 +439,25 @@ func (s *restoreSession) downloadNextBlobSet(plan *restorePlan) (bool, error) {
err := s.downloadBlobToCache(hash, blob.CompressedSize, blob.UncompressedSize)
if err != nil {
return false, fmt.Errorf("downloading blob %s: %w", shortHash(hash), err)
err = fmt.Errorf("downloading blob %s: %w", shortHash(hash), err)
// On cancel the error says nothing about the blob, so it ends
// the restore instead of failing the files that need it.
if s.ctx.Err() != nil {
return false, err
}
for _, fileID := range plan.filesReferencingBlob(hash) {
fileErr := s.v.handleRestoreFileError(
plan, s.opts, s.result, filesByID[fileID], fileID, err)
if fileErr != nil {
return false, fileErr
}
}
// next is among the failed files, so the rest of its blob set
// is left for any other file that still needs it.
return true, nil
}
s.result.BlobsDownloaded++
@@ -10,6 +10,7 @@ import (
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/cli"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
@@ -27,10 +28,12 @@ import (
//
// The backup half writes a snapshot with one index and hostname. The
// restore half throws that index away entirely: a fresh, empty index and
// a config that shares nothing with the original but the storage location
// and the secret key. If restore or verify needed the original local
// index — or the human snapshot ID that only that index holds — this test
// could not run, because the recovery host can know neither.
// a config written by `config init` and `config set storage_url`, as in
// the README's steps for restoring on another machine, which shares
// nothing with the original but the storage location and the secret key.
// If restore or verify needed the original local index — or
// the human snapshot ID that only that index holds — this test could not
// run, because the recovery host can know neither.
func TestRestoreOnAnotherMachine(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
@@ -58,7 +61,12 @@ func TestRestoreOnAnotherMachine(t *testing.T) {
// Recovery host: a fresh empty index, a different hostname, and no
// age_recipients — only the secret key and the same storage location.
recovery, stdout := newRecoveryHost(ctx, t, fs, storer)
configPath := filepath.Join(tempDir, "config.yml")
runVaultikCommand(t, "--config", configPath, "config", "init")
runVaultikCommand(t, "--config", configPath,
"config", "set", "storage_url", "file://"+storeDir)
recovery, stdout := newRecoveryHost(ctx, t, fs, storer, configPath)
// The recovery index really is empty. This is the assertion that makes
// the test a guard against restore quietly depending on the original
@@ -93,6 +101,10 @@ func TestRestoreOnAnotherMachine(t *testing.T) {
remote.RemoteKey, &vaultik.VerifyOptions{Deep: true}))
assertRestoredTreeMatches(t, fs, restoreDir, sourceFiles)
// With no public key configured, a backup must refuse to start.
err = recovery.CreateSnapshot(&vaultik.SnapshotCreateOptions{Cron: true})
require.ErrorContains(t, err, "age_recipients")
}
// writeRecoverySourceTree writes a small source tree spanning several
@@ -118,14 +130,24 @@ func writeRecoverySourceTree(
}
// newRecoveryHost builds the Vaultik a replacement machine would run: an
// empty in-memory index, a hostname different from the backup host, no
// age_recipients, and only the secret key plus the shared storer. It
// returns the instance and the buffer its stdout is wired to.
// empty in-memory index, a hostname different from the backup host, and
// only the secret key plus the shared storer. Its config is read from
// configPath by config.Load, as every command reads it. It returns the
// instance and the buffer its stdout is wired to.
func newRecoveryHost(
ctx context.Context, t *testing.T, fs afero.Fs, storer storage.Storer,
configPath string,
) (*vaultik.Vaultik, *bytes.Buffer) {
t.Helper()
cfg, err := config.Load(configPath)
require.NoError(t, err)
// Set directly rather than through VAULTIK_AGE_SECRET_KEY, which a
// parallel test cannot change.
cfg.AgeSecretKey = testAgeSecretKey
cfg.Hostname = "recovery-host"
recoveryDB, err := database.New(ctx, ":memory:")
require.NoError(t, err)
t.Cleanup(func() { _ = recoveryDB.Close() })
@@ -133,10 +155,7 @@ func newRecoveryHost(
stdout := &bytes.Buffer{}
recovery := &vaultik.Vaultik{
Config: &config.Config{
AgeSecretKey: testAgeSecretKey,
Hostname: "recovery-host",
},
Config: cfg,
Storage: storer,
Fs: fs,
Repositories: database.NewRepositories(recoveryDB),
@@ -150,6 +169,20 @@ func newRecoveryHost(
return recovery, stdout
}
// runVaultikCommand runs one vaultik command line in-process and fails the
// test if it returns an error. The command writes the cli package's global
// flag variables, so it must not be called from two tests that run at the
// same time.
func runVaultikCommand(t *testing.T, args ...string) {
t.Helper()
cmd := cli.NewRootCommand()
cmd.SetArgs(args)
cmd.SetOut(io.Discard)
cmd.SetErr(io.Discard)
require.NoError(t, cmd.Execute())
}
// assertRestoredTreeMatches byte-compares every restored file against its
// source content.
func assertRestoredTreeMatches(
@@ -1,6 +1,7 @@
package vaultik //nolint:testpackage // sets ctx/cancel and inspects scratch files
import (
"bytes"
"context"
"io"
"path/filepath"
@@ -11,6 +12,7 @@ import (
"time"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
@@ -157,3 +159,51 @@ func scratchEntries(t *testing.T, dir string) []string {
return matches
}
// TestRestoreSkipErrorsCancelDuringBlobDownload cancels a SkipErrors
// restore while a blob download is in progress. The download fails only
// because of the cancel, so Restore must return context.Canceled without
// reporting the file that needs the blob as failed.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestRestoreSkipErrorsCancelDuringBlobDownload(t *testing.T) {
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
cfg, storer, snapshotID, srcPath := backupOneFile(context.Background(),
t, fs, tempDir, "a.txt", []byte("hello vaultik"), 0o644)
gate := newBlockingBlobStorer(storer)
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
var out bytes.Buffer
v := newRestoreVaultik(ctx, cfg, gate, fs)
v.UI = ui.NewWithColor(&out, false)
restoreErr := make(chan error, 1)
go func() {
restoreErr <- v.Restore(&RestoreOptions{
SnapshotID: snapshotID,
TargetDir: filepath.Join(tempDir, "restored"),
SkipErrors: true,
})
}()
select {
case <-gate.entered:
case <-time.After(30 * time.Second):
t.Fatal("restore never reached the blob-download phase")
}
cancel()
require.ErrorIs(t, <-restoreErr, context.Canceled)
assert.NotContains(t, out.String(), srcPath,
"the cancel was reported as a failed file")
}
+14
View File
@@ -214,6 +214,20 @@ func (p *restorePlan) blobsNeeded(fileID types.FileID) []string {
return out
}
// filesReferencingBlob returns the pending files that reference the
// named blob, in any order. The result is a copy, so the caller may
// finishFile each of them while ranging over it.
func (p *restorePlan) filesReferencingBlob(blobHash string) []types.FileID {
files := p.blobFiles[blobHash]
out := make([]types.FileID, 0, len(files))
for id := range files {
out = append(out, id)
}
return out
}
// hasPending reports whether any unfinished files remain.
func (p *restorePlan) hasPending() bool {
return len(p.fileBlobs) > 0
@@ -0,0 +1,194 @@
package vaultik_test
import (
"bytes"
"context"
"fmt"
"path/filepath"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/ui"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// A file no larger than the chunker's minimum chunk size (a quarter of
// faultChunkSize) is stored as one chunk, so it lives in exactly one blob.
// More of them than fit in faultMaxBlobSize make the snapshot span two
// blobs.
const (
missingBlobFileBytes = int(faultChunkSize / 4)
missingBlobFileCount = 20
)
// missingBlobBackup is a snapshot of single-chunk files spread over two
// blobs, with one of those blobs deleted from the store.
type missingBlobBackup struct {
fs afero.Fs
cfg *config.Config
storer storage.Storer
snapshotID string
restoreDir string
// files holds the original content by source path.
files map[string][]byte
// lost holds the source paths whose chunk is in the deleted blob.
lost map[string]bool
}
// TestRestoreSkipErrorsSkipsFilesOfMissingBlob restores with SkipErrors
// after one blob of a two-blob snapshot was deleted. Every file stored in
// that blob must be reported as failed and left absent, every other file
// must be restored intact, and Restore must still return an error.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestRestoreSkipErrorsSkipsFilesOfMissingBlob(t *testing.T) {
ctx := context.Background()
backup := backupThenDeleteOneBlob(ctx, t)
var out bytes.Buffer
v := newReaderVaultik(ctx, backup.cfg, backup.storer, nil, backup.fs)
v.UI = ui.NewWithColor(&out, false)
err := v.Restore(&vaultik.RestoreOptions{
SnapshotID: backup.snapshotID,
TargetDir: backup.restoreDir,
SkipErrors: true,
})
require.Error(t, err, "restore must fail when files were skipped")
assert.Contains(t, err.Error(),
fmt.Sprintf("%d file(s) failed to restore", len(backup.lost)))
for path, content := range backup.files {
restored := filepath.Join(backup.restoreDir, path)
if backup.lost[path] {
assert.Containsf(t, out.String(), path,
"%s needs the deleted blob and must be reported", path)
assert.NoFileExists(t, restored)
continue
}
got, err := afero.ReadFile(backup.fs, restored)
require.NoErrorf(t, err, "%s does not need the deleted blob", path)
assert.Equalf(t, content, got, "%s restored with wrong content", path)
}
}
// TestRestoreMissingBlobAbortsWithoutSkipErrors checks that a deleted blob
// still ends the restore with an error when SkipErrors is not set.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestRestoreMissingBlobAbortsWithoutSkipErrors(t *testing.T) {
ctx := context.Background()
backup := backupThenDeleteOneBlob(ctx, t)
v := newReaderVaultik(ctx, backup.cfg, backup.storer, nil, backup.fs)
err := v.Restore(&vaultik.RestoreOptions{
SnapshotID: backup.snapshotID,
TargetDir: backup.restoreDir,
})
require.ErrorIs(t, err, storage.ErrNotFound)
}
// backupThenDeleteOneBlob backs up missingBlobFileCount single-chunk
// files, then deletes from the store the blob holding the first of them.
func backupThenDeleteOneBlob(
ctx context.Context, t *testing.T,
) *missingBlobBackup {
t.Helper()
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
dbPath := filepath.Join(tempDir, "index.sqlite")
cfg := faultTestConfig()
require.NoError(t, fs.MkdirAll(dataDir, 0o755))
files := make(map[string][]byte, missingBlobFileCount)
for i := range missingBlobFileCount {
path := filepath.Join(dataDir, fmt.Sprintf("file-%02d.bin", i))
files[path] = bytesPattern(
fmt.Sprintf("file-%02d-", i), missingBlobFileBytes)
require.NoError(t, afero.WriteFile(fs, path, files[path], 0o644))
}
storer, err := storage.NewFileStorer(filepath.Join(tempDir, "remote"))
require.NoError(t, err)
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
repos := database.NewRepositories(db)
id := fullFaultBackup(
ctx, t, fs, storer, cfg, repos, dataDir, dbPath, "missingblob")
blobOfFile := make(map[string]string, len(files))
for path := range files {
blobOfFile[path] = blobHashOfSingleChunkFile(ctx, t, repos, path)
}
require.NoError(t, db.Close())
deleted := blobOfFile[filepath.Join(dataDir, "file-00.bin")]
require.NoError(t, storer.Delete(ctx, fmt.Sprintf(
"blobs/%s/%s/%s", deleted[:2], deleted[2:4], deleted)))
lost := make(map[string]bool)
for path, hash := range blobOfFile {
if hash == deleted {
lost[path] = true
}
}
require.Less(t, len(lost), len(files),
"the snapshot must span more than one blob")
return &missingBlobBackup{
fs: fs,
cfg: cfg,
storer: storer,
snapshotID: id,
restoreDir: filepath.Join(tempDir, "restored"),
files: files,
lost: lost,
}
}
// blobHashOfSingleChunkFile returns the hash of the blob holding the one
// chunk of the file at path, as recorded in the local index.
func blobHashOfSingleChunkFile(
ctx context.Context, t *testing.T,
repos *database.Repositories, path string,
) string {
t.Helper()
chunks, err := repos.FileChunks.GetByPath(ctx, path)
require.NoError(t, err)
require.Lenf(t, chunks, 1, "%s must be a single chunk", path)
blobChunk, err := repos.BlobChunks.GetByChunkHash(
ctx, chunks[0].ChunkHash.String())
require.NoError(t, err)
require.NotNilf(t, blobChunk, "chunk of %s is in no blob", path)
blob, err := repos.Blobs.GetByID(ctx, blobChunk.BlobID.String())
require.NoError(t, err)
return blob.Hash.String()
}
+9
View File
@@ -23,6 +23,9 @@ var (
errSnapshotVerifyFailed = errors.New("verification failed")
errRemoveAllNeedsForce = errors.New("--all requires --force")
errInvalidTableName = errors.New("invalid table name")
errNoAgeRecipients = errors.New(
"creating a snapshot needs at least one public key in " +
"age_recipients (generate a keypair with: age-keygen)")
)
// listRecentLimit caps how many snapshot rows are fetched from the
@@ -42,6 +45,12 @@ type SnapshotCreateOptions struct {
// CreateSnapshot executes the snapshot creation operation
func (v *Vaultik) CreateSnapshot(opts *SnapshotCreateOptions) error {
// config.Load accepts an empty list, since listing, verifying and
// restoring need no public key.
if len(v.Config.AgeRecipients) == 0 {
return errNoAgeRecipients
}
overallStartTime := time.Now()
log.Info("Starting snapshot creation",
+27 -20
View File
@@ -48,12 +48,12 @@ missing() {
! command -v "$1" >/dev/null 2>&1
}
# Docker is a hard requirement, not a nice-to-have: script/lint lints by
# building Dockerfile.lint, whose digest-pinned golangci-lint image is
# the only place the linter runs, and script/check and script/precommit
# both run script/lint. A bootstrap that prints "bootstrap complete" on a
# machine where `make check` cannot run is a false success, so this fails
# instead.
# Docker is a hard requirement, not a nice-to-have: script/lint and
# script/test build the lint and test phases of the Dockerfile, the only
# place the linter and the tests run, and script/check and
# script/precommit both run them. A bootstrap that prints "bootstrap
# complete" on a machine where `make check` cannot run is a false
# success, so this fails instead.
#
# Installing docker from here was considered and rejected: it needs root,
# a running daemon, and on macOS a GUI cask, so an attempt would itself
@@ -80,15 +80,15 @@ bootstrap: FAILED - $reason.
Docker is required to develop this repo. Without it these do not work:
script/lint builds Dockerfile.lint, which runs the linter as a
build step in a digest-pinned golangci-lint image.
That FROM line is the single source of truth for the
linter version
script/check runs script/lint
script/lint builds the lint phase of the Dockerfile, which runs
the linter as a build step in a digest-pinned
golangci-lint image
script/test builds the test phase of the Dockerfile
script/check runs script/test and script/lint
script/precommit runs script/check, so commits are blocked by the
pre-commit hook installed by script/setup
script/cibuild builds Dockerfile.lint and Dockerfile, which is what
CI runs
script/cibuild runs script/check and builds the image, which is
what CI runs
Install docker (and start the daemon, checking DOCKER_HOST and your
group membership), then re-run script/bootstrap. golangci-lint on PATH
@@ -104,15 +104,22 @@ main() {
if missing git; then pkg_install git git git git; fi
if missing make; then pkg_install gnumake make make make; fi
# Go toolchain
if missing go; then pkg_install go golang go go; fi
# Go toolchain: the host's own, or else the version go.mod names,
# hash-verified, in .tool/go. That directory is not on the caller's
# PATH; this script, the Makefile, script/fmt, script/fmt-check,
# script/precommit and script/release add it to theirs.
if missing go; then
"$ROOT/script/install-go"
PATH="$PATH:$ROOT/.tool/go/bin"
fi
# golangci-lint is deliberately NOT installed: script/lint lints by
# building Dockerfile.lint, whose digest-pinned image is the only
# place the linter runs, so whatever a package manager happens to
# ship would only be a shadow of the pinned version that could drift
# from CI. Nothing on the host is ever used as a linter, at any
# version, so installing one here would buy nothing.
# building the lint phase of the Dockerfile, whose digest-pinned
# image is the only place the linter runs, so whatever a package
# manager happens to ship would only be a shadow of the pinned
# version that could drift from CI. Nothing on the host is ever used
# as a linter, at any version, so installing one here would buy
# nothing.
# goreleaser, at the version pinned by script/install-goreleaser and
# verified against a hardcoded sha256. Package managers are not used
+3 -2
View File
@@ -1,7 +1,8 @@
#!/bin/sh
# script/check: run all checks (test, lint, fmt-check). Our own
# extension to scripts-to-rule-them-all. Must not modify any files.
# Generic: usually needs no adaptation.
# extension to scripts-to-rule-them-all. test and lint are Docker
# phases; fmt-check is native, because a formatter writes the working
# tree. Must not modify any files.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
+18 -68
View File
@@ -1,78 +1,28 @@
#!/bin/sh
# script/cibuild: run the CI build. This is the full gate, and it is two
# builds, in this order:
#
# Dockerfile.lint the linter, as a build step (a clean build IS a
# clean lint)
# Dockerfile `make fmt-check` and `make test` in the builder
# stage, then the product image
#
# Either one failing fails this script. Note what follows from the
# split: script/docker builds only the product image and so no longer
# lints -- this script and script/check (which runs script/lint) are the
# things that decide whether the tree is clean.
#
# Generic apart from the two Dockerfiles: the Gitea workflow runs this
# on push.
# script/cibuild: run the CI build. It bootstraps first: a CI runner
# checks out and runs this and nothing else, and script/fmt-check runs
# the formatter on the host, which a pristine checkout cannot do.
# --no-cache for the same reason as script/docker: the gate phases the
# final stage depends on are RUN steps, and a cached one is a check that
# did not run.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
# Both Dockerfiles key their check layers on CHECK_EPOCH, so a fresh
# value is what forces those layers to re-run: without it an
# unchanged tree replays them from cache, the checks never execute,
# and the build still exits 0. Each ARG sits immediately above the
# check RUNs, so dependency and module layers still cache. Both
# Dockerfiles also refuse to build at all when CHECK_EPOCH is empty,
# so a missing value fails loudly here rather than passing quietly.
#
# The value must be unique per invocation, not per second. `date +%s`
# is second-granular, so two concurrent invocations in the same
# second get identical epochs and the later one can be served from
# cache -- the original defect in miniature. `%N` alone does not fix
# it: busybox silently drops %N, exits 0, and hands back second
# granularity with no warning. `$$` is what makes this correct
# regardless, since concurrent invocations have different pids.
#
# Assign the epoch on its own line rather than inline in the
# argument. Under `set -eu` a command substitution that fails
# inside an argument does NOT abort the script: CHECK_EPOCH would
# become an empty string, an empty string is a constant, and a
# constant CHECK_EPOCH is exactly the cached-check false green this
# script exists to prevent -- so the guard would disarm itself and
# still exit 0. As a bare assignment, `set -e` catches a failing
# `date` and no build starts.
#
# A separate value per build, because they are separate builds: one
# `date` shared between them would still be fresh, but reusing it
# invites the two to be collapsed into a single value that is
# computed somewhere else and passed in.
epoch="$(date +%s%N)$$"
# cacheonly for the lint build: its verdict is the exit status and
# the image is never run, so exporting it is pure cost. See
# script/lint.
docker build --output=type=cacheonly \
--build-arg CHECK_EPOCH="$epoch" -f Dockerfile.lint .
# Version, commit and build date are computed here on the host, the
# same way script/docker does, and passed into the product build so
# the CI-built image reports its real source. The build context
# excludes .git (see .dockerignore), so the build cannot derive them
# itself; without these it would stamp the Dockerfile's dev/unknown
# fallbacks. VERSION comes from script/version, the source of truth
# shared with the Makefile.
version="$("$ROOT/script/version")"
commit="$(git rev-parse HEAD 2>/dev/null || echo unknown)"
commit_date="$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)"
epoch="$(date +%s%N)$$"
docker build --build-arg CHECK_EPOCH="$epoch" \
"$SCRIPT_DIR/bootstrap"
"$SCRIPT_DIR/check"
# Own line: a failing command substitution inside an argument does
# not trip `set -e`, so the inline form degrades silently to an
# empty constant. The VERSION build argument takes precedence over
# the version a build stage derives from the .git in the context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
docker build --no-cache \
--build-arg VERSION="$version" \
--build-arg COMMIT="$commit" \
--build-arg COMMIT_DATE="$commit_date" \
.
-t "$("$SCRIPT_DIR/projectname")" .
}
main "$@"
+9 -33
View File
@@ -1,14 +1,8 @@
#!/bin/sh
# script/docker: build the Docker image tagged with the project name.
# Identical in all repos; the tag comes from script/projectname.
# Generic: needs no adaptation.
#
# This builds the PRODUCT image only, and the product Dockerfile has no
# lint stage: linting lives in Dockerfile.lint and is run by
# script/lint. So a green here means `make fmt-check` and `make test`
# passed and the image built -- it says nothing about lint. The gates
# are script/check (which runs script/lint) and script/cibuild (which
# builds both files).
# --no-cache because the gate phases the final stage depends on are RUN
# steps, and a cached one is a check that did not run.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
@@ -16,32 +10,14 @@ ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
# Same CHECK_EPOCH contract as script/cibuild, for the same reason
# and with the same bare-assignment and `$$` requirements -- see the
# comments there. This script is not the CI gate, but a local build
# is almost always warm, so without this it would report a green the
# tree had not earned and the two entrypoints would disagree about
# whether the tree is clean. The Dockerfile now refuses to build
# without a non-empty value, so this is required, not optional.
epoch="$(date +%s%N)$$"
# Version, commit and build date are computed here on the host,
# where .git exists, and passed into the build. The build context
# excludes .git (see .dockerignore), so the container cannot derive
# them itself -- it used to try and always got "unknown", giving
# every image a "commit: unknown" it could not be traced from.
# VERSION comes from script/version, the source of truth shared with
# the Makefile, so a Docker build reports the same string (tag,
# dev-<sha>, or a -dirty variant) that a local build of the same
# tree would.
version="$("$SCRIPT_DIR/version")"
commit="$(git rev-parse HEAD 2>/dev/null || echo unknown)"
commit_date="$(git show -s --format=%cs HEAD 2>/dev/null || echo unknown)"
docker build --build-arg CHECK_EPOCH="$epoch" \
# Own line: a failing command substitution inside an argument does
# not trip `set -e`, so the inline form degrades silently to an
# empty constant. The VERSION build argument takes precedence over
# the version a build stage derives from the .git in the context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
docker build --no-cache \
--build-arg VERSION="$version" \
--build-arg COMMIT="$commit" \
--build-arg COMMIT_DATE="$commit_date" \
-t "$("$SCRIPT_DIR/projectname")" .
}
+2
View File
@@ -6,6 +6,8 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
# Where script/bootstrap installs Go when the host has none.
PATH="$PATH:$ROOT/.tool/go/bin"
go fmt ./...
}
+8 -1
View File
@@ -7,7 +7,14 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
unformatted="$(gofmt -l .)"
# Where script/bootstrap installs Go when the host has none.
PATH="$PATH:$ROOT/.tool/go/bin"
# .tool/go holds that toolchain's own sources, which are not ours.
if ! unformatted="$(find . -path ./.tool -prune -o -name '*.go' -type f \
-exec gofmt -l {} +)"; then
echo "fmt-check: gofmt failed" >&2
exit 1
fi
if [ -n "$unformatted" ]; then
echo "Files not formatted:" >&2
echo "$unformatted" >&2
+23 -18
View File
@@ -4,12 +4,13 @@
# own extension to scripts-to-rule-them-all. Idempotent: exits at once
# when the pinned toolchain is already installed.
#
# Only .gitea/workflows/release.yml calls this. goreleaser is not a
# .gitea/workflows/release.yml calls this, and so does script/bootstrap
# when the host has no Go, as on the check runner. goreleaser is not a
# compiler: it shells out to `go` for the `before:` hook and for every
# one of the four cross-compiles, so the release runner needs a Go
# toolchain on PATH. check.yml never does -- it builds inside the
# digest-pinned Dockerfile images -- so this is the release path's only
# host Go, and per REPO_POLICIES.md it must be pinned by hash.
# toolchain on PATH. The check runner compiles nothing on the host; it
# uses this Go only for bootstrap's `go mod download` and for gofmt in
# script/fmt-check. Per REPO_POLICIES.md a host Go is pinned by hash.
# actions/setup-go exposes no checksum input, so Go is installed the way
# script/install-goreleaser installs goreleaser: download the exact
# archive from go.dev and refuse it unless its sha256 matches the value
@@ -18,11 +19,11 @@
# The version is go.mod's `go` directive, the single source of truth for
# the toolchain. GO_VERSION below MUST equal it, and this script fails
# when they disagree -- so bumping Go is one reviewed change touching
# go.mod, the checksum here, and the Dockerfile golang digest together.
# go.mod, the checksums here, and the Dockerfile's two golang digests
# together.
#
# Linux only, because that is what the release runner is. A darwin dev
# building a snapshot uses their own Go; supporting an OS means adding
# its checksums.
# Linux and macOS, each on amd64 and arm64: the four archives whose
# checksums are committed below.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
@@ -33,6 +34,8 @@ ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
GO_VERSION="1.26.1"
SHA256_LINUX_AMD64="031f088e5d955bab8657ede27ad4e3bc5b7c1ba281f05f245bcc304f327c987a"
SHA256_LINUX_ARM64="a290581cfe4fe28ddd737dde3095f3dbeb7f2e4065cab4eae44dfc53b760c2f7"
SHA256_DARWIN_AMD64="65773dab2f8cc4cd23d93ba6d0a805de150ca0b78378879292be0b903b8cdd08"
SHA256_DARWIN_ARM64="353df43a7811ce284c8938b5f3c7df40b7bfb6f56cb165b150bc40b5e2dd541f"
GOROOT_DIR="$ROOT/.tool/go"
GOCMD="$GOROOT_DIR/bin/go"
@@ -101,26 +104,28 @@ main() {
arch="$(uname -m)"
case "$os" in
Linux) os="linux" ;;
Darwin) os="darwin" ;;
*)
echo "install-go: unsupported OS $os (release runner is Linux)" >&2
echo "install-go: unsupported OS $os" >&2
exit 1
;;
esac
case "$arch" in
x86_64 | amd64)
arch="amd64"
sum="$SHA256_LINUX_AMD64"
;;
arm64 | aarch64)
arch="arm64"
sum="$SHA256_LINUX_ARM64"
;;
x86_64 | amd64) arch="amd64" ;;
arm64 | aarch64) arch="arm64" ;;
*)
echo "install-go: no pinned checksum for architecture $arch" >&2
echo "install-go: unsupported architecture $arch" >&2
exit 1
;;
esac
case "${os}-${arch}" in
linux-amd64) sum="$SHA256_LINUX_AMD64" ;;
linux-arm64) sum="$SHA256_LINUX_ARM64" ;;
darwin-amd64) sum="$SHA256_DARWIN_AMD64" ;;
darwin-arm64) sum="$SHA256_DARWIN_ARM64" ;;
esac
archive="go${GO_VERSION}.${os}-${arch}.tar.gz"
url="https://go.dev/dl/${archive}"
+13 -98
View File
@@ -1,108 +1,23 @@
#!/bin/sh
# script/lint: run the linter.
# script/lint: run the linter. Linting is a phase of the Dockerfile and
# this builds that phase alone; the linter is never installed or run on
# a developer host, where a shared result cache and a host-global lock
# make its answer untrustworthy.
#
# The linter runs inside the image built by Dockerfile.lint, and it runs
# there as a BUILD STEP: a successful build of that file IS a clean
# lint. Nothing lints on the host, at any version, ever. That FROM line
# is the single source of truth for the linter version in this repo, so
# a local run and a CI run of the same tree cannot disagree.
#
# One container per run means one lint cache and one golangci-lint lock
# per run, both private to that run and thrown away with it. That is
# what makes concurrent runs on a shared host safe, and it is why this
# script no longer carries per-worktree cache directories, a lock-retry
# loop, or an output audit: there is no shared state left for them to
# defend (issue https://git.eeqj.de/sneak/vaultik/issues/113).
#
# To watch the linter execute, set BUILDKIT_PROGRESS=plain, which docker
# honours directly:
#
# BUILDKIT_PROGRESS=plain script/lint
#
# The check layers -- `golangci-lint config verify` and then
# `golangci-lint run` -- must appear as executing rather than CACHED on
# every run; see the CHECK_EPOCH comment in Dockerfile.lint.
# The phase is not the last stage in the file, so it is built only when
# --target names it. --no-cache because a cached lint layer is a lint
# that did not run. The tag makes each build replace the previous image
# instead of leaving a dangling one behind.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
DOCKERFILE="$ROOT/Dockerfile.lint"
require_docker() {
if ! command -v docker >/dev/null 2>&1; then
cat >&2 <<EOF
lint: docker is required to run the pinned linter.
lint image declared by: $DOCKERFILE
Install docker. Linting with any other golangci-lint is not supported:
it is what lets a local run pass while CI fails. A golangci-lint on
PATH is never used, whatever its version.
EOF
exit 1
fi
if ! docker info >/dev/null 2>&1; then
cat >&2 <<EOF
lint: the docker daemon is not reachable, so the pinned linter cannot
run.
lint image declared by: $DOCKERFILE
Start the daemon (and check DOCKER_HOST / your group membership). This
script will not fall back to a different linter version or to an
unpinned binary on PATH.
EOF
exit 1
fi
}
usage() {
cat >&2 <<EOF
usage: $(basename "$0")
script/lint takes no arguments. The linter runs as a build step, so
there is no command line to pass flags to; anything accepted here would
have to be silently dropped. To apply autofixes, use script/lint-fix,
which runs the same pinned image as a container for exactly this
reason.
EOF
exit 2
}
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
[ "$#" -eq 0 ] || usage
cd "$ROOT"
require_docker
# A fresh epoch per invocation is what forces the check layers to
# execute; the layers above the ARG in Dockerfile.lint still cache,
# so a run is not cold. The value must be unique per invocation, not
# per second: `date +%s` is second-granular, so two concurrent
# invocations in the same second would get identical epochs and the
# later one could be served from cache -- the false green in
# miniature. `%N` alone does not fix it either, because busybox
# silently drops %N, exits 0, and hands back second granularity with
# no warning. `$$` is what makes this correct regardless, since
# concurrent invocations have different pids.
#
# Assign it on its own line rather than inline in the argument.
# Under `set -eu` a command substitution that fails inside an
# argument does NOT abort the script: CHECK_EPOCH would become an
# empty string, an empty string is a constant, and a constant epoch
# is exactly the cached-lint false green this guards against. As a
# bare assignment, `set -e` catches a failing `date` and no build
# starts.
epoch="$(date +%s%N)$$"
# cacheonly: the lint verdict is the build's exit status, and the
# image it would otherwise produce is never run. Exporting it costs
# most of the wall time of a warm run and leaves a dangling image
# behind on every invocation, on a host that may be running many.
docker build \
--output=type=cacheonly \
--build-arg CHECK_EPOCH="$epoch" \
-f "$DOCKERFILE" \
"$ROOT"
docker build --no-cache \
--target lint \
-t "$("$SCRIPT_DIR/projectname")-lint" .
}
main "$@"
+9 -8
View File
@@ -5,27 +5,28 @@
#
# THIS IS A DEVELOPER CONVENIENCE AND NEVER A GATE. Nothing in
# script/check, script/precommit or script/cibuild calls it, and no gate
# reads its exit status. The gate is script/lint, which builds
# Dockerfile.lint; run that afterwards to find out whether the tree is
# actually clean.
# reads its exit status. The gate is script/lint, which builds the lint
# phase of the Dockerfile; run that afterwards to find out whether the
# tree is actually clean.
#
# Unlike script/lint this cannot be a build step: a build step writes
# into an image, and fixes have to land in the worktree. So it runs the
# same pinned image as a container with the tree bind-mounted, which
# means it needs a LOCAL docker daemon -- a remote daemon has no access
# to these files, and this script will appear to do nothing there. The
# image reference is parsed out of Dockerfile.lint's FROM line, so the
# image reference is parsed out of the lint phase's FROM line, so the
# autofixer is always the same version as the linter that gates; fixes
# written by a different version are not necessarily fixes for the
# version that decides.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
DOCKERFILE="$ROOT/Dockerfile.lint"
DOCKERFILE="$ROOT/Dockerfile"
# The image reference from Dockerfile.lint, tag and digest included.
# The image reference from the FROM line ending `AS lint`, tag and
# digest included.
lint_image() {
awk '$1 == "FROM" { print $2; exit }' "$DOCKERFILE"
awk '$1 == "FROM" && $NF == "lint" { print $2; exit }' "$DOCKERFILE"
}
main() {
@@ -33,7 +34,7 @@ main() {
image="$(lint_image)"
if [ -z "$image" ]; then
echo "lint-fix: no FROM line found in $DOCKERFILE" >&2
echo "lint-fix: no FROM ... AS lint line found in $DOCKERFILE" >&2
exit 1
fi
+2
View File
@@ -9,6 +9,8 @@ ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
# Where script/bootstrap installs Go when the host has none.
PATH="$PATH:$ROOT/.tool/go/bin"
go mod tidy
go fmt ./...
git diff --exit-code -- go.mod go.sum || {
+2
View File
@@ -44,6 +44,8 @@ resolve_goreleaser() {
main() {
cd "$ROOT"
# Where script/bootstrap installs Go when the host has none.
PATH="$PATH:$ROOT/.tool/go/bin"
if ! bin="$(resolve_goreleaser)"; then
cat >&2 <<EOF
+10 -60
View File
@@ -1,69 +1,19 @@
#!/bin/sh
# script/test: run the test suite. Quiet on success; on failure, rerun
# verbosely for full diagnostic output (the exit 1 ensures the rerun
# never turns a failure into a pass).
# script/test: run the test suite. Testing is a phase of the Dockerfile
# and this builds that phase alone, on the same terms as script/lint:
# --target because a phase that is not the last stage is built only when
# named, --no-cache because a cached test layer is a test that did not
# run, and a tag so each build replaces the previous image.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# The flags live in one function so the quiet run and the verbose rerun
# below cannot drift apart. A rerun that used different flags would
# diagnose a different program than the one that failed.
#
# -count=1 is the documented way to bypass Go's test result cache, and
# it is not optional here. Without it, a package whose inputs are
# unchanged prints `ok <pkg> (cached)`, and that line is
# indistinguishable -- to every check this repo performs -- from a
# package that actually ran. The whole suite reports its full set of
# `ok` lines in under half a second having executed nothing. That
# matters beyond the local inner loop: the Dockerfile's `RUN make test`
# is forced to re-execute by CHECK_EPOCH, but a GOCACHE baked into an
# earlier image layer survives into the re-executed step, so the step
# can re-run and still do no work. It is applied unconditionally rather
# than only in the containerised path because the pre-commit hook runs
# this same script; a gate that is honest only in CI is dishonest
# exactly where people lean on it most.
#
# -timeout is a hang backstop, not a performance budget: its job is to
# turn a deadlocked test into a stack dump instead of a wedged CI job,
# so it wants to sit far above the slowest legitimate runtime, not just
# above it. It is per test binary and covers test execution only -- the
# clock starts inside testing.M.Run, after compilation and linking, so
# build time is not charged against it. (Measured: a containerised run
# with an empty GOCACHE reports per-package durations within noise of a
# warm host run. A shell `timeout 30 go test ./...` would include
# compilation, but that is a different mechanism from this flag.)
#
# The 120s value DELIBERATELY DIVERGES from REPO_POLICIES.md:192, which
# mandates "Add a 30-second timeout", and from that file's canonical Go
# recipe at :212-214, which uses -timeout 30s. REPO_POLICIES.md is
# org-canonical and cannot be amended from this repo, so the divergence
# is recorded here instead, and issue #101 proposes amending the policy
# text upstream. Do not revert this to 30s without reading #101 first.
#
# Why it diverges: the slowest packages are internal/database and
# internal/vaultik, observed under -race at about 6.4s warm, 8.1s in a
# cold containerised run on a contended host, and 10.2s in an
# independent cold run on this same host. The worst case is not tightly
# characterised -- each fresh measurement has come in above the last --
# which is itself an argument for generous headroom. Against the 10.2s
# observation, 30s is only 2.9x: not a safety margin but a flake
# waiting for a slow day, whose failure mode is a timeout that looks
# like a real defect. 120s leaves about 12x while still bounding a hung
# package -- including the verbose rerun below -- to a few minutes. The
# cost of that choice, also recorded on #101: because of the rerun, a
# hung package pays the timeout twice.
run_tests() {
go test -race -timeout 120s -count=1 "$@" ./...
}
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
run_tests || {
echo "--- Rerunning with -v for details ---"
run_tests -v
exit 1
}
docker build --no-cache \
--target test \
-t "$("$SCRIPT_DIR/projectname")-test" .
}
main "$@"
+15 -60
View File
@@ -1,73 +1,28 @@
#!/bin/sh
# script/version: output the version string to bake into the binary.
# Our own extension to scripts-to-rule-them-all, and the single source
# of truth for the version: the Makefile's LDFLAGS call this rather
# than carrying a hardcoded constant, which is what used to make every
# local build claim to be 1.0.0-rc.1 regardless of git state.
# Our own extension to scripts-to-rule-them-all. The Makefile's LDFLAGS
# call this rather than carrying a hardcoded constant, which is what used
# to make every local build claim to be 1.0.0-rc.1 regardless of git
# state. script/docker and script/cibuild run the same `git describe`
# themselves, and fall back to "unknown" rather than "dev".
#
# The rules, in order:
# The version is `git describe --tags --always --dirty`: the tag on a
# tagged commit, tag-N-gHASH on a commit after one, the short commit
# when no tag is reachable, each with a "-dirty" suffix when tracked
# files have uncommitted changes. Untracked files are ignored: a stray
# scratch file does not change what was compiled. Outside a git checkout
# (release tarball, `go install`), or in one with no commits, it is
# "dev".
#
# HEAD is exactly on an annotated or lightweight tag
# -> that tag, with a leading "v" stripped
# anything else
# -> "dev-<12 chars of HEAD>"
# not a git checkout at all (release tarball, `go install`)
# -> "dev"
#
# Either of the first two gains a "-dirty" suffix when tracked files
# have uncommitted changes, because a modified checkout of v1.0.0 is
# not v1.0.0. Untracked files are ignored, matching `git describe
# --dirty`: a stray scratch file does not change what was compiled.
#
# The "v" is stripped so that a `make` build and a goreleaser build of
# the same tagged commit report the *same* string: goreleaser's
# {{ .Version }} is the tag without the prefix, and the release archive
# names are built from it. A tag named `v1.0.0` therefore produces
# `vaultik 1.0.0`, matching `vaultik_1.0.0_linux_amd64.tar.gz`.
#
# Nothing here ever invents a version number. An untagged build says so
# and names the commit it was built from; it does not round up to the
# nearest plausible release.
# Nothing here ever invents a version number.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Length of the commit prefix in a dev version. Matches
# globals.ShortCommit, so `vaultik version` shows the same 12 chars in
# its version line and its commit line.
SHORT_LEN=12
main() {
cd "$ROOT"
if ! git rev-parse --git-dir >/dev/null 2>&1; then
echo "dev"
return 0
fi
dirty=""
if [ -n "$(git status --porcelain --untracked-files=no 2>/dev/null)" ]; then
dirty="-dirty"
fi
# --exact-match so a *descendant* of a tag is not reported as that
# tag. Plain `git describe --tags` would call a commit 40 patches
# past v1.0.0 "v1.0.0-40-gabc1234", and the leading token of that is
# a released version the build is not.
tag="$(git describe --tags --exact-match HEAD 2>/dev/null || true)"
if [ -n "$tag" ]; then
echo "${tag#v}${dirty}"
return 0
fi
sha="$(git rev-parse "--short=$SHORT_LEN" HEAD 2>/dev/null || true)"
if [ -z "$sha" ]; then
# A repo with no commits at all.
echo "dev"
return 0
fi
echo "dev-${sha}${dirty}"
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
echo "${version:-dev}"
}
main "$@"