Author SHA1 Message Date
sneak 8baa11b6cb Report a prune count that could not be read as unknown, not 0 (closes #96)
check / check (pull_request) Successful in 4m8s
PruneDatabase read seven table counts with the error discarded, so a
query that could not run silently became 0 and the before/after delta
computed from it looked like real work.

Each read now goes through a helper that logs at warn on failure and
returns nil; nil renders as "unknown", never "0", so an empty table is
distinguishable from one that could not be queried. Deltas built from an
unknown count are themselves unknown. The failure is surfaced, not
propagated, so the command's failure conditions are unchanged. These
counts have no --json output — under --json the summary is suppressed
entirely — so nothing there can show a false 0.

Model: opus-4-8
2026-09-21 19:26:36 +00:00
10 changed files with 38 additions and 412 deletions
+3 -8
View File
@@ -104,12 +104,7 @@ Version: 2025-06-08
13. Pre-1.0: NEVER write database migrations. There are no live databases
anywhere — every user's local index can be rebuilt from a fresh full
backup. To change the schema, edit `internal/database/schema/001.sql`
(and any code that touches the affected tables) directly; do not add new
numbered schema files. Those numbered files and the `schema_migrations`
table they populate only bootstrap a fresh database — they are not an
upgrade path. The local index is disposable until 1.0 ships and is
tagged; once 1.0 is tagged that clause expires and the question of
upgrading existing indexes returns. See [`docs/DATAMODEL.md`](docs/DATAMODEL.md)
for the full explanation.
backup. When the schema changes, just change `schema.sql` (and any code
that touches the affected tables). The local index is disposable until
1.0 ships and is tagged.
+11 -70
View File
@@ -84,57 +84,6 @@ VAULTIK_AGE_SECRET_KEY='AGE-SECRET-KEY-...' vaultik snapshot restore <snapshot-i
# 0 3 * * * vaultik snapshot create --cron --prune --keep-newer-than 4w
```
## restoring on another machine
Restoring on a host that never ran the backup — a replacement machine
after the original is gone — is the case vaultik is built for. That host
needs only three things: the `vaultik` binary, the age **private** key,
and the storage credentials for the destination. It does **not** need the
local index, the original config file, or the original hostname.
```sh
# install
go install sneak.berlin/go/vaultik/cmd/vaultik@latest
# create a config and point it at the ORIGINAL backup destination
vaultik config init
vaultik config set storage_url "s3://bucket/prefix?endpoint=https://s3.example.com"
vaultik config set s3.access_key_id "..."
vaultik config set s3.secret_access_key "..."
# see what is on the destination store
vaultik snapshot list
```
`snapshot list` reads the destination store without the private key. A
snapshot that is not in this host's (empty) local index is shown as
remote-only: its row is identified by `<remote only:...>` rather than by
a `hostname_name_timestamp` name, because the name lives only in the
local index and the encrypted database and cannot be recovered from the
store. Its timestamp and compressed size are real. (See the `snapshot
list` description under [command details](#command-details) for the full
explanation.)
Use that remote key — the hex printed inside `<remote only:...>`, or the
full `remote_key` from `snapshot list --json` — to restore and verify:
```sh
# restore everything to /tmp/restored, then check every restored file's
# chunk hashes
VAULTIK_AGE_SECRET_KEY='AGE-SECRET-KEY-...' \
vaultik snapshot restore --verify <remote-key> /tmp/restored
# optionally, deep-verify the snapshot against the store (downloads and
# cryptographically checks every blob)
VAULTIK_AGE_SECRET_KEY='AGE-SECRET-KEY-...' \
vaultik snapshot verify --deep <remote-key>
```
`age_recipients` (the public key) is not needed to restore — only the
private key in `VAULTIK_AGE_SECRET_KEY`. Both the abbreviated key printed
in the table and the full 64-character key from `--json` are accepted; a
leading part of the key is enough as long as it is unambiguous.
---
## cli
@@ -296,8 +245,6 @@ local index alone, and still exits zero.
* Default (shallow): checks that all blobs referenced in the manifest exist in storage
* `--deep`: Downloads and decrypts each blob, verifies chunk hashes against the
encrypted metadata database
* Accepts the same identifiers as `snapshot restore`: a snapshot ID, or a
remote-only snapshot's remote key (or an unambiguous leading part of it)
* `--json`: Output results as JSON
**`snapshot purge`**: Remove old snapshots based on criteria. Retention is
@@ -328,10 +275,6 @@ on the destination in one go, use `vaultik remote nuke --force`.
**`snapshot restore`**: Restore files from a backup snapshot.
* Requires `VAULTIK_AGE_SECRET_KEY` environment variable
* Accepts a snapshot ID, or — for a snapshot only on the destination
store — its remote key (or an unambiguous leading part of it) as shown
by `snapshot list`. See
[restoring on another machine](#restoring-on-another-machine).
* Optional path arguments to restore specific files/directories (default: all)
* Preserves file permissions, timestamps, ownership (ownership requires root),
symlinks, and empty directories
@@ -514,13 +457,9 @@ Key fields:
sequentially. Restore speed is bound by single-stream throughput.
* **Device nodes, named pipes, and sockets are silently skipped.** Only
regular files, directories, and symlinks are backed up.
* **No upgrade path between versions.** There is no supported way to carry
an existing local index across a schema change; if the local SQLite
schema changes between versions, delete the local database (`vaultik
database delete`) and run a full backup. Remote storage is unaffected.
(The binary does embed numbered schema files and a `schema_migrations`
table to bootstrap a fresh database — see [`docs/DATAMODEL.md`](docs/DATAMODEL.md)
— but that is not an upgrade path.)
* **No database migrations.** If the local SQLite schema changes between
versions, delete the local database (`vaultik database delete`) and run
a full backup. Remote storage is unaffected.
* **Files that change during backup may be inconsistent.** There is no
filesystem snapshot or freeze. If a file is modified between the scan
and chunk phases, the backed-up copy may reflect a partial write.
@@ -586,12 +525,14 @@ priority.
### infrastructure
* **Cross-version schema upgrades.** There is no upgrade path between
released versions — pre-1.0 schema changes are handled by `vaultik
database delete` plus a full re-scan (see
[`docs/DATAMODEL.md`](docs/DATAMODEL.md)). Post-1.0 we'll need a
migration story to keep existing index databases usable across
upgrades.
* **Cross-machine restore documentation.** The "restore from
another host" workflow works but isn't documented as a
first-class operation in this README. Worth a dedicated section
once it's settled.
* **Schema migrations.** Currently nonexistent — pre-1.0 schema
changes are handled by `vaultik database delete` plus a full
re-scan. Post-1.0 we'll need a migration story to keep existing
index databases usable across upgrades.
* **Storage backend coverage tests.** S3, file://, and rclone://
all share the Storer interface but the rclone path is the least
exercised in CI.
+5 -24
View File
@@ -5,30 +5,11 @@
Vaultik uses a local SQLite database to track file metadata, chunk mappings, and blob associations during the backup process. This database serves as an index for incremental backups and enables efficient deduplication.
**Important Notes:**
This section is the authoritative explanation of the schema/migration story;
other documents (the README and `AGENTS.md`) link here.
- **No upgrade path between versions (pre-1.0)**: Vaultik has no supported way to
carry an existing local index across a schema change. The index is disposable
— if the on-disk schema changes between versions, delete the local SQLite
database (`vaultik database delete`) and run a full backup. Remote storage is
unaffected; the new index re-deduplicates against existing remote blobs. This
is the standing project policy, and it is separate from the schema bootstrap
described next.
- **Schema bootstrap**: a fresh database is populated from numbered SQL files
embedded in the binary under `internal/database/schema/`. `000.sql` creates the
`schema_migrations` table; `001.sql` creates the application tables. On opening
a database the code applies each numbered file that has not yet run and records
its version in `schema_migrations`. This bootstraps a new database; it does not
upgrade an existing one between released versions.
- **Changing the schema (pre-1.0)**: edit `internal/database/schema/001.sql` (and
the code that touches the affected tables) directly. Do not add new numbered
files — there is no installed base to migrate.
- **Disposability expires at 1.0**: the index is treated as disposable only until
1.0 ships and is tagged. Once 1.0 is tagged that clause expires and the
question of upgrading existing indexes returns. It is deliberately left open
here.
- **No Migration Support (pre-1.0)**: Vaultik does not support database schema
migrations. The local index is treated as disposable — if the schema changes,
delete the local SQLite database (`vaultik database delete`) and run a full
backup. The remote storage is unaffected; the new index will re-deduplicate
against existing remote blobs.
- **Version Compatibility**: In rare cases, you may need to use the same version
of Vaultik to restore a backup as was used to create it. This ensures
compatibility with the metadata format stored in S3.
+1 -4
View File
@@ -221,10 +221,7 @@ func newSnapshotVerifyCommand() *cobra.Command {
cmd := &cobra.Command{
Use: "verify <snapshot-id>",
Short: "Verify snapshot integrity",
Long: "Verifies that all blobs referenced in a snapshot exist.\n\n" +
"The snapshot may be named by its ID or, on a host with no local\n" +
"index, by the remote key that 'snapshot list' prints for a\n" +
"remote-only snapshot (an unambiguous leading part is enough).",
Long: "Verifies that all blobs referenced in a snapshot exist",
Args: requireSnapshotIDArg,
RunE: func(cmd *cobra.Command, args []string) error {
snapshotID := args[0]
-4
View File
@@ -48,10 +48,6 @@ target directory.
If no paths are specified, all files are restored.
If paths are specified, only matching files/directories are restored.
The snapshot may be named by its ID or, when restoring on a host with no
local index, by the remote key that 'snapshot list' prints for a
remote-only snapshot (an unambiguous leading part is enough).
Requires the VAULTIK_AGE_SECRET_KEY environment variable to be set with
the age private key.
+5 -10
View File
@@ -18,6 +18,7 @@ import (
"sneak.berlin/go/vaultik/internal/blobgen"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/snapshot"
"sneak.berlin/go/vaultik/internal/types"
)
@@ -576,20 +577,14 @@ func (v *Vaultik) handleRestoreVerification(
}
// downloadSnapshotDB downloads and decrypts the snapshot metadata
// database. The identifier is resolved to the snapshot's remote key: a
// human ID is hashed, and a remote key (or its abbreviation, as printed
// for a remote-only snapshot) is used as-is, so a host with no local
// index can restore the snapshots it can only see on the store.
// database. The snapshotID is the human ID; we hash it to the remote
// key for the storage path.
func (v *Vaultik) downloadSnapshotDB(
snapshotID string, identity age.Identity,
) (*database.DB, error) {
remoteKey, err := v.resolveSnapshotRemoteKey(snapshotID)
if err != nil {
return nil, err
}
// Download encrypted database from storage
dbKey := fmt.Sprintf("metadata/%s/db.zst.age", remoteKey)
dbKey := fmt.Sprintf("metadata/%s/db.zst.age",
snapshot.RemoteSnapshotKey(snapshotID))
reader, err := v.Storage.Get(v.ctx, dbKey)
if err != nil {
@@ -1,167 +0,0 @@
package vaultik_test
import (
"bytes"
"context"
"io"
"path/filepath"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/config"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
"sneak.berlin/go/vaultik/internal/snapshot"
"sneak.berlin/go/vaultik/internal/storage"
"sneak.berlin/go/vaultik/internal/ui"
"sneak.berlin/go/vaultik/internal/vaultik"
)
// TestRestoreOnAnotherMachine proves the disaster-recovery path: a host
// that has only the vaultik binary, the age secret key, and the storage
// credentials — no local index, a different hostname, and no
// age_recipients configured — can list, restore, and verify a snapshot
// straight from the destination store.
//
// The backup half writes a snapshot with one index and hostname. The
// restore half throws that index away entirely: a fresh, empty index and
// a config that shares nothing with the original but the storage location
// and the secret key. If restore or verify needed the original local
// index — or the human snapshot ID that only that index holds — this test
// could not run, because the recovery host can know neither.
func TestRestoreOnAnotherMachine(t *testing.T) {
log.Initialize(log.Config{})
t.Parallel()
fs := afero.NewOsFs()
tempDir := t.TempDir()
dataDir := filepath.Join(tempDir, "source")
storeDir := filepath.Join(tempDir, "remote")
restoreDir := filepath.Join(tempDir, "restored")
dbPath := filepath.Join(tempDir, "index.sqlite")
chunkSize := int64(64 * 1024)
maxBlobSize := int64(512 * 1024)
sourceFiles := writeRecoverySourceTree(t, fs, dataDir, chunkSize)
ctx := context.Background()
// Backup host: one index, hostname test-host, age_recipients set.
// runFileStorageBackup closes the index before returning, so nothing
// below can lean on it.
_, storer, originalID := runFileStorageBackup(
ctx, t, fs, dataDir, storeDir, dbPath, chunkSize, maxBlobSize)
// Recovery host: a fresh empty index, a different hostname, and no
// age_recipients — only the secret key and the same storage location.
recovery, stdout := newRecoveryHost(ctx, t, fs, storer)
// The recovery index really is empty. This is the assertion that makes
// the test a guard against restore quietly depending on the original
// index: if it did, an empty index would make restore fail.
localSnaps, err := recovery.Repositories.Snapshots.ListRecent(ctx, 100)
require.NoError(t, err)
require.Empty(t, localSnaps, "recovery host must start with no local index")
// List: the snapshot shows up as remote-only, identified by its remote
// key, with no recoverable human ID.
require.NoError(t, recovery.ListSnapshots(true))
rows := decodeListJSON(t, stdout.String())
require.Len(t, rows, 1)
remote := rows[0]
assert.False(t, remote.LocallyTracked, "snapshot must be remote-only here")
assert.Empty(t, remote.ID, "the human ID is unknown to the recovery host")
require.Len(t, remote.RemoteKey, 64)
assert.Equal(t, snapshot.RemoteSnapshotKey(originalID), remote.RemoteKey,
"the listed key is the hashed snapshot ID")
// Restore driven by the abbreviated identifier the table prints (the
// first 12 hex of the remote key), then deep-verify from the store
// keyed by the full remote key. Both are what a recovery host can know.
require.NoError(t, recovery.Restore(&vaultik.RestoreOptions{
SnapshotID: remote.RemoteKey[:12],
TargetDir: restoreDir,
Verify: true,
}))
require.NoError(t, recovery.RunDeepVerify(
remote.RemoteKey, &vaultik.VerifyOptions{Deep: true}))
assertRestoredTreeMatches(t, fs, restoreDir, sourceFiles)
}
// writeRecoverySourceTree writes a small source tree spanning several
// chunks (so restore reassembles real multi-chunk files) and returns the
// content keyed by absolute path.
func writeRecoverySourceTree(
t *testing.T, fs afero.Fs, dataDir string, chunkSize int64,
) map[string][]byte {
t.Helper()
sourceFiles := map[string][]byte{
filepath.Join(dataDir, "notes.txt"): []byte("recover me"),
filepath.Join(dataDir, "sub", "big.bin"): bytesPattern("big-", int(chunkSize*3)),
filepath.Join(dataDir, "sub", "small.bin"): bytesPattern("small-", 128),
}
for path, content := range sourceFiles {
require.NoError(t, fs.MkdirAll(filepath.Dir(path), 0o755))
require.NoError(t, afero.WriteFile(fs, path, content, 0o644))
}
return sourceFiles
}
// newRecoveryHost builds the Vaultik a replacement machine would run: an
// empty in-memory index, a hostname different from the backup host, no
// age_recipients, and only the secret key plus the shared storer. It
// returns the instance and the buffer its stdout is wired to.
func newRecoveryHost(
ctx context.Context, t *testing.T, fs afero.Fs, storer storage.Storer,
) (*vaultik.Vaultik, *bytes.Buffer) {
t.Helper()
recoveryDB, err := database.New(ctx, ":memory:")
require.NoError(t, err)
t.Cleanup(func() { _ = recoveryDB.Close() })
stdout := &bytes.Buffer{}
recovery := &vaultik.Vaultik{
Config: &config.Config{
AgeSecretKey: testAgeSecretKey,
Hostname: "recovery-host",
},
Storage: storer,
Fs: fs,
Repositories: database.NewRepositories(recoveryDB),
DB: recoveryDB,
Stdout: stdout,
Stderr: io.Discard,
UI: ui.NewWithColor(io.Discard, false),
}
recovery.SetContext(ctx)
return recovery, stdout
}
// assertRestoredTreeMatches byte-compares every restored file against its
// source content.
func assertRestoredTreeMatches(
t *testing.T, fs afero.Fs, restoreDir string, sourceFiles map[string][]byte,
) {
t.Helper()
for origPath, expected := range sourceFiles {
restored := filepath.Join(restoreDir, origPath)
got, err := afero.ReadFile(fs, restored)
require.NoErrorf(t, err, "restored file missing: %s", restored)
require.Truef(t, bytes.Equal(got, expected),
"byte mismatch for %s", origPath)
}
}
+3 -5
View File
@@ -670,11 +670,9 @@ func (v *Vaultik) VerifySnapshotWithOptions(
v.printVerifyHeader(snapshotID, opts)
// Resolve the identifier to the snapshot's remote key and download the
// manifest. A human ID is hashed; a remote key (or its abbreviation,
// as printed for a remote-only snapshot) is used as-is, so a host with
// no local index can verify a snapshot it can only see on the store.
manifest, err := v.resolveAndDownloadManifest(snapshotID)
// Download and parse manifest. The caller supplies a human
// snapshot ID; we hash it to address remote storage.
manifest, err := v.downloadManifestByKey(snapshot.RemoteSnapshotKey(snapshotID))
if err != nil {
if opts.JSON {
result.Status = verifyStatusFailed
-101
View File
@@ -1,101 +0,0 @@
package vaultik
import (
"errors"
"fmt"
"strings"
"sneak.berlin/go/vaultik/internal/snapshot"
)
// remoteKeyHexLen is the length of a full remote snapshot key: a SHA256
// digest rendered as lowercase hex.
const remoteKeyHexLen = 64
// Sentinel errors for resolving a snapshot identifier against the store.
var (
errSnapshotKeyNotFound = errors.New(
"no snapshot on the destination store matches this identifier")
errSnapshotKeyAmbiguous = errors.New(
"identifier matches more than one snapshot on the destination store")
)
// resolveSnapshotRemoteKey turns a snapshot identifier supplied on the
// command line into the remote key that names the snapshot's metadata
// directory on the destination store. Every remote path a restore or
// verify reads is built from that key.
//
// Two forms are accepted, matching the two things a host can know:
//
// - A human snapshot ID (hostname_name_timestamp), which a host holding
// the local index has. It is hashed to its remote key; the store is
// not consulted.
// - A remote key, or the leading part of one, which is all a host with
// no local index can know — it is exactly what `snapshot list` prints
// for a remote-only snapshot (see formatRemoteOnlyID). It is resolved
// against the destination store's metadata listing; an identifier that
// matches no snapshot, or more than one, is an error.
//
// The two are told apart by shape: a remote key is lowercase hex, and a
// human snapshot ID never is (it carries a hostname, underscores, and an
// RFC3339 timestamp).
func (v *Vaultik) resolveSnapshotRemoteKey(identifier string) (string, error) {
if !isRemoteKeyOrPrefix(identifier) {
return snapshot.RemoteSnapshotKey(identifier), nil
}
keys, err := v.listAllRemoteSnapshotKeys()
if err != nil {
return "", fmt.Errorf(
"listing destination store to resolve %q: %w", identifier, err)
}
var matches []string
for _, key := range keys {
if strings.HasPrefix(key, identifier) {
matches = append(matches, key)
}
}
switch len(matches) {
case 1:
return matches[0], nil
case 0:
return "", fmt.Errorf("%w: %s", errSnapshotKeyNotFound, identifier)
default:
return "", fmt.Errorf("%w: %s (%d matches)",
errSnapshotKeyAmbiguous, identifier, len(matches))
}
}
// resolveAndDownloadManifest resolves a snapshot identifier to its remote
// key (see resolveSnapshotRemoteKey) and downloads that snapshot's
// manifest.
func (v *Vaultik) resolveAndDownloadManifest(
identifier string,
) (*snapshot.Manifest, error) {
remoteKey, err := v.resolveSnapshotRemoteKey(identifier)
if err != nil {
return nil, err
}
return v.downloadManifestByKey(remoteKey)
}
// isRemoteKeyOrPrefix reports whether s is a full remote key or the
// leading part of one: 1 to 64 lowercase hex characters. A human snapshot
// ID is never all hex, so this shape test is enough to tell the two apart.
func isRemoteKeyOrPrefix(s string) bool {
if s == "" || len(s) > remoteKeyHexLen {
return false
}
for _, r := range s {
if (r < '0' || r > '9') && (r < 'a' || r > 'f') {
return false
}
}
return true
}
+9 -18
View File
@@ -138,15 +138,8 @@ func (v *Vaultik) RunDeepVerify(snapshotID string, opts *VerifyOptions) error {
func (v *Vaultik) loadVerificationData(
snapshotID string, opts *VerifyOptions, result *VerifyResult,
) (*snapshot.Manifest, *tempDB, []snapshot.BlobInfo, error) {
// Resolve the identifier to the snapshot's remote key. A human ID is
// hashed; a remote key (or its abbreviation, as printed for a
// remote-only snapshot) is used as-is, so a host with no local index
// can verify a snapshot it can only see on the store.
remoteKey, err := v.resolveSnapshotRemoteKey(snapshotID)
if err != nil {
return nil, nil, nil, v.deepVerifyFailure(result, opts,
fmt.Sprintf("resolving snapshot identifier: %v", err), err)
}
// All remote paths use the hashed key derived from the human ID.
remoteKey := snapshot.RemoteSnapshotKey(snapshotID)
// Download manifest. downloadManifestByKey is the single reader for
// remote manifests; see its doc comment.
@@ -193,7 +186,7 @@ func (v *Vaultik) loadVerificationData(
fmt.Errorf("failed to decrypt database: %w", err))
}
dbBlobs, err := v.getBlobsFromDatabase(tdb.DB)
dbBlobs, err := v.getBlobsFromDatabase(snapshotID, tdb.DB)
if err != nil {
_ = tdb.Close()
@@ -508,21 +501,19 @@ func (v *Vaultik) verifyBlobFinalIntegrity(
return nil
}
// getBlobsFromDatabase gets all blobs for the snapshot from the database.
//
// The exported per-snapshot database holds exactly one snapshot's data
// (see cleanSnapshotDB), so every row in snapshot_blobs belongs to it.
// We select them directly rather than filtering by the human snapshot ID,
// which a host restoring from the store alone does not have.
func (v *Vaultik) getBlobsFromDatabase(db *sql.DB) ([]snapshot.BlobInfo, error) {
// getBlobsFromDatabase gets all blobs for the snapshot from the database
func (v *Vaultik) getBlobsFromDatabase(
snapshotID string, db *sql.DB,
) ([]snapshot.BlobInfo, error) {
query := `
SELECT b.blob_hash, b.compressed_size
FROM snapshot_blobs sb
JOIN blobs b ON sb.blob_hash = b.blob_hash
WHERE sb.snapshot_id = ?
ORDER BY b.blob_hash
`
rows, err := db.QueryContext(v.ctx, query)
rows, err := db.QueryContext(v.ctx, query, snapshotID)
if err != nil {
return nil, fmt.Errorf("failed to query snapshot blobs: %w", err)
}