Compare commits

..
1 Commits
Author SHA1 Message Date
sneak 6dc134a50e Re-vendor the canonical files from sneak/prompts at dd4027b (closes #213)
check / check (push) Successful in 22m37s
Linting and testing become the lint and test phases of the Dockerfile,
and the build stage depends on both. Dockerfile.lint, CHECK_EPOCH and
the tests that checked them are removed. Every docker build in script/
passes --no-cache, and script/cibuild runs script/bootstrap first, which
now fetches apt package lists so a fresh CI runner can install Go. The
image takes its version from the VERSION build arg or git describe, and
still stamps the commit and its date from .git. The golangci-lint
v2.14.0 findings are fixed in the code. The rules in CLAUDE.md move into
AGENTS.md. IsDevVersion now counts "unknown", the version script/docker
stamps outside a git checkout.

Model: opus-5-5
2026-10-06 01:01:50 +00:00
14 changed files with 22 additions and 381 deletions
-3
View File
@@ -10,6 +10,3 @@ insert_final_newline = true
[Makefile] [Makefile]
indent_style = tab indent_style = tab
[*.go]
indent_style = tab
+5 -6
View File
@@ -16,12 +16,11 @@ jobs:
fetch-depth: 0 fetch-depth: 0
# goreleaser is not a compiler: it shells out to `go` for the # goreleaser is not a compiler: it shells out to `go` for the
# `before:` hook and for every one of the four cross-compiles. # `before:` hook and for every one of the four cross-compiles.
# Without this step the release either fails at the before-hook or, # Nothing else in this repo puts a Go toolchain on the runner --
# worse, ships binaries built by whatever Go the runner happens to # check.yml runs script/cibuild, which does all of its work inside
# carry. check.yml's script/cibuild also puts Go on its own runner, # the digest-pinned Dockerfile images -- so without this step the
# through script/bootstrap from the package manager at whatever # release either fails at the before-hook or, worse, ships binaries
# version that ships, but uses it only for `go mod download` and # built by whatever Go the runner happens to carry.
# gofmt; it compiles inside the digest-pinned Dockerfile images.
# #
# actions/setup-go would pin the action by commit sha, but the Go # actions/setup-go would pin the action by commit sha, but the Go
# tarball it downloads at runtime is verified against no value in # tarball it downloads at runtime is verified against no value in
-15
View File
@@ -45,18 +45,3 @@ node_modules/
[iI][dD]_[eE][cC][dD][sS][aA]_[sS][kK] [iI][dD]_[eE][cC][dD][sS][aA]_[sS][kK]
[iI][dD]_[eE][dD]25519 [iI][dD]_[eE][dD]25519
[iI][dD]_[eE][dD]25519_[sS][kK] [iI][dD]_[eE][dD]25519_[sS][kK]
# Go build and test output.
*.log
*.out
*.test
coverage.html
# This repo's own host-built artifacts.
/vaultik
/dist/
/.tool/
# Local configs for development; they hold storage credentials.
local-config.yaml
dev-config.yaml
+1
View File
@@ -128,3 +128,4 @@ Version: 2025-06-08
18. For estimates: backing up over 2.5Gbit/s ethernet to an S3 server 18. For estimates: backing up over 2.5Gbit/s ethernet to an S3 server
backed by 2000MB/sec SSD takes about 4 seconds per gigabyte. backed by 2000MB/sec SSD takes about 4 seconds per gigabyte.
+1 -1
View File
@@ -353,7 +353,7 @@ CreateSnapshot(opts)
## Deduplication Strategy ## Deduplication Strategy
1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid), unless the file lists a chunk that no uploaded blob holds; such a file is re-chunked 1. **File-level**: Files unchanged since last backup are skipped (metadata comparison: size, mtime, mode, uid, gid)
2. **Chunk-level**: Chunks are content-addressed by SHA256 hash. If a chunk hash already exists in the database, the chunk data is not re-uploaded. 2. **Chunk-level**: Chunks are content-addressed by SHA256 hash. If a chunk hash already exists in the database, the chunk data is not re-uploaded.
+2 -3
View File
@@ -43,10 +43,9 @@ COPY . .
# `git describe --tags --always` on the .git in the build context. With # `git describe --tags --always` on the .git in the build context. With
# .git present, a version that is still empty, dev or unknown fails the # .git present, a version that is still empty, dev or unknown fails the
# build: git is missing or could not read the checkout. The commit and # build: git is missing or could not read the checkout. The commit and
# its date always come from that .git. A context without .git, such as # its date always come from that .git, and are "unknown" without one.
# a source export, stamps "dev" and an "unknown" commit and date.
ARG VERSION ARG VERSION
RUN VERSION="${VERSION:-$(git describe --tags --always || echo dev)}"; \ RUN VERSION="${VERSION:-$(git describe --tags --always)}"; \
if [ -e .git ]; then \ if [ -e .git ]; then \
case "$VERSION" in ""|dev|unknown) \ case "$VERSION" in ""|dev|unknown) \
echo "version is '$VERSION' although .git is present" >&2; \ echo "version is '$VERSION' although .git is present" >&2; \
+3 -5
View File
@@ -40,8 +40,7 @@ Features:
* modern encryption ([age](https://age-encryption.org/), X25519 + ChaCha20-Poly1305) * modern encryption ([age](https://age-encryption.org/), X25519 + ChaCha20-Poly1305)
* content-defined chunking with deduplication (FastCDC) * content-defined chunking with deduplication (FastCDC)
* incremental backups (a file is re-chunked only when it changed or a * incremental backups (only changed files are re-chunked)
chunk it lists is held by no uploaded blob)
* multithreaded zstd compression at configurable levels * multithreaded zstd compression at configurable levels
* content-addressed immutable storage * content-addressed immutable storage
* local state tracking in SQLite (enables write-only incremental backups) * local state tracking in SQLite (enables write-only incremental backups)
@@ -500,9 +499,8 @@ format does and does not protect.
* Content-defined chunking using the FastCDC algorithm * Content-defined chunking using the FastCDC algorithm
* Average chunk size: configurable (default 10MB) * Average chunk size: configurable (default 10MB)
* Deduplication at file level (unchanged files skipped, unless a chunk * Deduplication at file level (unchanged files skipped) and chunk level
the file lists is held by no uploaded blob) and chunk level (identical (identical chunks across files stored once)
chunks across files stored once)
* Multiple chunks packed into blobs to reduce object count * Multiple chunks packed into blobs to reduce object count
### encryption ### encryption
-9
View File
@@ -31,15 +31,6 @@ the tag exists and is exercised; what is left is merging `next` to
is v2.14.0, with its new findings fixed in the code, and the rules in is v2.14.0, with its new findings fixed in the code, and the rules in
`CLAUDE.md` now live in `AGENTS.md`. `CLAUDE.md` now live in `AGENTS.md`.
- 2026-10-06: Made a backup re-chunk a known file that lists a chunk no
uploaded blob holds
([issue #214](https://git.eeqj.de/sneak/vaultik/issues/214)). File
rows are shared by every snapshot and updated in place, so removing or
pruning the only snapshot that referenced a changed file's current
blob left an older snapshot keeping a row that matched the disk while
no blob held its chunks. The next backup skipped the file and
completed a snapshot that could not restore it.
- 2026-09-22: Routed the last direct-to-stdout command output through - 2026-09-22: Routed the last direct-to-stdout command output through
`internal/ui` `internal/ui`
([issue #149](https://git.eeqj.de/sneak/vaultik/issues/149)). The ([issue #149](https://git.eeqj.de/sneak/vaultik/issues/149)). The
+6 -6
View File
@@ -45,9 +45,9 @@ func TestProductDockerfileTakesVersionAsBuildArg(t *testing.T) {
// TestProductDockerfileDerivesVersionFromGit fails unless a build given // TestProductDockerfileDerivesVersionFromGit fails unless a build given
// no VERSION, such as a plain `docker build .` of a clone, takes it from // no VERSION, such as a plain `docker build .` of a clone, takes it from
// `git describe` of the .git in its context, stamps "dev" when the // `git describe` of the .git in its context, and fails rather than
// context has no .git, and fails rather than stamp "dev" when that .git // stamp "dev" when that .git yields no version. The commit and its date
// yields no version. The commit and its date come from the same .git. // come from the same .git.
func TestProductDockerfileDerivesVersionFromGit(t *testing.T) { func TestProductDockerfileDerivesVersionFromGit(t *testing.T) {
t.Parallel() t.Parallel()
@@ -56,9 +56,9 @@ func TestProductDockerfileDerivesVersionFromGit(t *testing.T) {
buildAt := indexContaining(found, "go build") buildAt := indexContaining(found, "go build")
require.GreaterOrEqual(t, buildAt, 0, "%s must build", productDockerfile) require.GreaterOrEqual(t, buildAt, 0, "%s must build", productDockerfile)
assert.Contains(t, found[buildAt], "git describe --tags --always || echo dev", assert.Contains(t, found[buildAt], "git describe --tags --always",
"%s must derive the version from git when no VERSION is given,"+ "%s must derive the version from git when no VERSION is given",
" and stamp dev when the context has no .git", productDockerfile) productDockerfile)
assert.Contains(t, found[buildAt], "[ -e .git ]", assert.Contains(t, found[buildAt], "[ -e .git ]",
"%s must fail when the context carries .git but yields no version", "%s must fail when the context carries .git but yields no version",
productDockerfile) productDockerfile)
-1
View File
@@ -195,7 +195,6 @@ Tracks blob upload metrics.
1. **Change Detection** 1. **Change Detection**
- `SELECT * FROM files WHERE path = ?` - Get previous file metadata - `SELECT * FROM files WHERE path = ?` - Get previous file metadata
- Compare mtime, size, mode to detect changes - Compare mtime, size, mode to detect changes
- Re-chunk a file that lists a chunk no uploaded blob holds, even when its metadata is unchanged
- Skip unchanged files but still add to `snapshot_files` - Skip unchanged files but still add to `snapshot_files`
2. **Chunk Reuse** 2. **Chunk Reuse**
-50
View File
@@ -266,56 +266,6 @@ func (r *FileRepository) ListByPrefix(
return files, rows.Err() return files, rows.Err()
} }
// ListIDsWithChunksNotInUploadedBlobs returns the IDs of the files whose
// path starts with prefix and that list at least one chunk held by no
// blob whose upload has completed (uploaded_ts set). A new snapshot
// cannot reference such a chunk, so a backup must not treat the file as
// unchanged even when its metadata matches the file on disk.
func (r *FileRepository) ListIDsWithChunksNotInUploadedBlobs(
ctx context.Context, prefix string,
) ([]types.FileID, error) {
query := `
SELECT DISTINCT f.id
FROM files f
JOIN file_chunks fc ON fc.file_id = f.id
WHERE f.path LIKE ? || '%'
AND NOT EXISTS (
SELECT 1
FROM blob_chunks bc
JOIN blobs b ON bc.blob_id = b.id
WHERE bc.chunk_hash = fc.chunk_hash
AND b.uploaded_ts IS NOT NULL
)
`
rows, err := r.db.conn.QueryContext(ctx, query, prefix)
if err != nil {
return nil, fmt.Errorf("querying files: %w", err)
}
defer func() {
err := rows.Close()
if err != nil {
Fatalf("failed to close rows: %v", err)
}
}()
var ids []types.FileID
for rows.Next() {
var id types.FileID
err := rows.Scan(&id)
if err != nil {
return nil, fmt.Errorf("scanning file ID: %w", err)
}
ids = append(ids, id)
}
return ids, rows.Err()
}
// ListAll returns all files in the database // ListAll returns all files in the database
func (r *FileRepository) ListAll(ctx context.Context) ([]*File, error) { func (r *FileRepository) ListAll(ctx context.Context) ([]*File, error) {
query := ` query := `
-43
View File
@@ -73,11 +73,6 @@ type Scanner struct {
knownChunks map[string]struct{} knownChunks map[string]struct{}
knownChunksMu sync.RWMutex knownChunksMu sync.RWMutex
// filesToRechunk holds the IDs of known files that list a chunk no
// uploaded blob holds; they are re-chunked even when their metadata
// is unchanged.
filesToRechunk map[types.FileID]struct{}
// Pending chunk hashes - chunks that have been added to packer but not // Pending chunk hashes - chunks that have been added to packer but not
// yet committed to DB. When a blob finalizes, the committed chunks are // yet committed to DB. When a blob finalizes, the committed chunks are
// removed from this set. // removed from this set.
@@ -330,42 +325,9 @@ func (s *Scanner) loadDatabaseState(
s.ui.Completef("Loaded %s known chunks from local index database.", s.ui.Completef("Loaded %s known chunks from local index database.",
s.ui.Count(len(s.knownChunks))) s.ui.Count(len(s.knownChunks)))
err = s.loadFilesToRechunk(ctx, path)
if err != nil {
return nil, fmt.Errorf("loading files to re-chunk: %w", err)
}
return knownFiles, nil return knownFiles, nil
} }
// loadFilesToRechunk loads the IDs of known files under path that list a
// chunk no uploaded blob holds. A file row is shared by every snapshot
// that lists the file and is updated in place when the file changes,
// while a blob row is deleted once no snapshot references it. Removing
// the only snapshot that references a changed file's current blob
// therefore leaves an older snapshot keeping a file row whose metadata
// matches the disk while no blob the local index records as uploaded
// holds its chunks, so a new snapshot cannot reference them. The dropped
// blob can still be in remote storage until prune removes it.
func (s *Scanner) loadFilesToRechunk(ctx context.Context, path string) error {
ids, err := s.repos.Files.ListIDsWithChunksNotInUploadedBlobs(ctx, path)
if err != nil {
return fmt.Errorf("listing files: %w", err)
}
s.filesToRechunk = make(map[types.FileID]struct{}, len(ids))
for _, id := range ids {
s.filesToRechunk[id] = struct{}{}
}
if len(ids) > 0 {
log.Info("Re-chunking known files whose chunks are not all "+
"in uploaded blobs", "files", len(ids))
}
return nil
}
// repairInterruptedBlobs discards blob rows left by a previous run whose // repairInterruptedBlobs discards blob rows left by a previous run whose
// upload never completed. Such a blob has its chunks, blob_chunks, and // upload never completed. Such a blob has its chunks, blob_chunks, and
// blobs rows committed to the local index before the upload is attempted, // blobs rows committed to the local index before the upload is attempted,
@@ -1246,11 +1208,6 @@ func (s *Scanner) checkFileInMemory(
return file, true return file, true
} }
// No uploaded blob holds one of its chunks (see loadFilesToRechunk)
if _, rechunk := s.filesToRechunk[fileID]; rechunk {
return file, true
}
// Check if file has changed // Check if file has changed
if existingFile.Size != file.Size || if existingFile.Size != file.Size ||
existingFile.MTime.Unix() != file.MTime.Unix() || existingFile.MTime.Unix() != file.MTime.Unix() ||
@@ -1,234 +0,0 @@
package vaultik_test
import (
"context"
"crypto/rand"
"path/filepath"
"slices"
"strings"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/chunker"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/storage/faultstore"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// These tests cover https://git.eeqj.de/sneak/vaultik/issues/214: once
// the only snapshot referencing a changed file's current blob is
// dropped, an older snapshot still keeps the file's row, and the next
// backup must re-chunk the file instead of treating it as unchanged.
//
// Each backup run uses its own snapshot name over the same source
// directory, so the snapshot IDs differ without waiting for their
// one-second timestamps to tick over. File rows are keyed by path, so
// the runs share them exactly as runs of one snapshot name would.
// changedFileConfig returns a full-backup config with the snapshot names
// "first", "second" and "third", all backing up dataDir.
func changedFileConfig(dataDir, dbPath string) *config.Config {
source := config.SnapshotConfig{Paths: []string{dataDir}}
cfg := faultTestConfig()
cfg.IndexPath = dbPath
cfg.ChunkSize = config.Size(faultChunkSize)
cfg.Snapshots = map[string]config.SnapshotConfig{
"first": source,
"second": source,
"third": source,
}
return cfg
}
// backUp runs a full backup of the named snapshot.
func backUp(v *vaultik.Vaultik, name string) error {
return v.CreateSnapshot(&vaultik.SnapshotCreateOptions{
Cron: true,
Snapshots: []string{name},
})
}
// appendedFileSize is the size of the file before the tests append to it.
// It is twice the largest chunk the chunker cuts, so the file's first
// chunk ends before the appended bytes and stays in a blob the first
// snapshot still references, while its new chunks are only in the blob
// that gets dropped.
const appendedFileSize = 2 * chunker.ChunkSizeSpread * faultChunkSize
// appendRandomBytes appends n random bytes to the file at path, creating
// the file if it does not exist, and records the new content in files.
// Random content gives the chunker real cut points.
func appendRandomBytes(
t *testing.T, fs afero.Fs, files map[string][]byte, path string, n int64,
) {
t.Helper()
added := make([]byte, n)
_, err := rand.Read(added)
require.NoError(t, err)
content := slices.Concat(files[path], added)
require.NoError(t, afero.WriteFile(fs, path, content, 0o644))
files[path] = content
}
// localSnapshotID returns the ID of the one snapshot in the local index
// that was backed up under name.
func localSnapshotID(
ctx context.Context, t *testing.T,
repos *database.Repositories, name string,
) string {
t.Helper()
snapshots, err := repos.Snapshots.ListRecent(ctx, listRecentTestLimit)
require.NoError(t, err)
var ids []string
for _, s := range snapshots {
if strings.Contains(s.ID.String(), "_"+name+"_") {
ids = append(ids, s.ID.String())
}
}
require.Lenf(t, ids, 1, "expected one local snapshot named %q", name)
return ids[0]
}
// assertThirdSnapshotRestores closes the local index, then restores the
// snapshot named "third" from store alone and byte-compares every file
// against files.
func assertThirdSnapshotRestores(
ctx context.Context, t *testing.T, cfg *config.Config,
store storage.Storer, repos *database.Repositories, db *database.DB,
fs afero.Fs, restoreDir string, files map[string][]byte,
) {
t.Helper()
id := localSnapshotID(ctx, t, repos, "third")
require.NoError(t, db.Close())
reader := newReaderVaultik(ctx, cfg, store, nil, fs)
require.NoError(t, reader.Restore(&vaultik.RestoreOptions{
SnapshotID: id,
TargetDir: restoreDir,
Verify: true,
}), "the backup after the drop must restore the changed file")
assertRestoredTree(t, fs, restoreDir, files)
}
// Trigger 1: bytes are appended to the file, a second snapshot backs it
// up, and that snapshot is removed. The first snapshot keeps the file row,
// which now lists the appended content's chunks, while removal drops the
// blob that held them.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestBackupAfterRemovingNewestSnapshotRestoresChangedFile(t *testing.T) {
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
storeDir := filepath.Join(tempDir, "remote")
restoreDir := filepath.Join(tempDir, "restored")
dbPath := filepath.Join(tempDir, "index.sqlite")
changedPath := filepath.Join(dataDir, "appended.bin")
ctx := context.Background()
files := writeFaultSourceTree(t, fs, dataDir)
cfg := changedFileConfig(dataDir, dbPath)
appendRandomBytes(t, fs, files, changedPath, appendedFileSize)
store, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
repos := database.NewRepositories(db)
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
require.NoError(t, backUp(v, "first"))
appendRandomBytes(t, fs, files, changedPath, faultChunkSize)
require.NoError(t, backUp(v, "second"))
_, err = v.RemoveSnapshot(localSnapshotID(ctx, t, repos, "second"),
&vaultik.RemoveOptions{Force: true})
require.NoError(t, err)
require.NoError(t, backUp(v, "third"))
assertThirdSnapshotRestores(
ctx, t, cfg, store, repos, db, fs, restoreDir, files)
}
// Trigger 2: bytes are appended to the file and the run that backs it up
// is interrupted at the manifest upload, after its blobs were uploaded.
// The next run's prune drops that incomplete snapshot and its blob, while
// the first snapshot keeps the file row, which now lists the appended
// content's chunks.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestBackupAfterInterruptedRunRestoresChangedFile(t *testing.T) {
log.Initialize(log.Config{})
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "src")
storeDir := filepath.Join(tempDir, "remote")
restoreDir := filepath.Join(tempDir, "restored")
dbPath := filepath.Join(tempDir, "index.sqlite")
changedPath := filepath.Join(dataDir, "appended.bin")
ctx := context.Background()
files := writeFaultSourceTree(t, fs, dataDir)
cfg := changedFileConfig(dataDir, dbPath)
appendRandomBytes(t, fs, files, changedPath, appendedFileSize)
inner, err := storage.NewFileStorer(storeDir)
require.NoError(t, err)
failManifest := false
store := faultstore.New(inner)
store.OnPut = func(key string) faultstore.PutAction {
if failManifest && strings.HasSuffix(key, "manifest.json.zst") {
return faultstore.PutFail
}
return faultstore.PutNormal
}
db, err := database.New(ctx, dbPath)
require.NoError(t, err)
repos := database.NewRepositories(db)
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
require.NoError(t, backUp(v, "first"))
appendRandomBytes(t, fs, files, changedPath, faultChunkSize)
failManifest = true
require.Error(t, backUp(v, "second"),
"the run must fail when its manifest upload fails")
failManifest = false
require.NoError(t, backUp(v, "third"))
assertThirdSnapshotRestores(
ctx, t, cfg, inner, repos, db, fs, restoreDir, files)
}
+4 -5
View File
@@ -7,11 +7,10 @@
# Only .gitea/workflows/release.yml calls this. goreleaser is not a # Only .gitea/workflows/release.yml calls this. goreleaser is not a
# compiler: it shells out to `go` for the `before:` hook and for every # compiler: it shells out to `go` for the `before:` hook and for every
# one of the four cross-compiles, so the release runner needs a Go # one of the four cross-compiles, so the release runner needs a Go
# toolchain on PATH. check.yml compiles nothing on the host: it uses # toolchain on PATH. check.yml compiles nothing on the host: its one
# whatever Go script/bootstrap finds or installs only for that script's # host use of Go is gofmt in script/fmt-check, with whatever Go
# `go mod download` and for gofmt in script/fmt-check. So this is the # script/bootstrap finds or installs. So this is the release path's only
# only host Go that builds anything, and per REPO_POLICIES.md it must be # host Go, and per REPO_POLICIES.md it must be pinned by hash.
# pinned by hash.
# actions/setup-go exposes no checksum input, so Go is installed the way # actions/setup-go exposes no checksum input, so Go is installed the way
# script/install-goreleaser installs goreleaser: download the exact # script/install-goreleaser installs goreleaser: download the exact
# archive from go.dev and refuse it unless its sha256 matches the value # archive from go.dev and refuse it unless its sha256 matches the value