Commit Graph
305 Commits
Author SHA1 Message Date
clawbot 7696f83258 Leave remote info orphan figures unknown when a manifest is unreadable (closes #228)
check / check (push) Waiting to run
When a manifest could not be read, remote info skipped it, counted that
snapshot's blobs as orphaned and advised running prune. The orphan
figures are now unknown in that case, with no prune advice; --json gives
them as null and lists the unreadable remote keys in
unreadable_manifests. Only a listed manifest.json.zst is read, so a
directory without one, as an interrupted backup leaves, keeps the
figures known.

Names under metadata/ were used unchecked and printed raw. A name that
is not a remote key is now skipped with a warning. A manifest under it
is not read either, so it also leaves the figures unknown; --json counts
such manifests in skipped_manifest_count.

Model: opus-5-5
2026-10-07 10:59:26 +02:00
clawbot 5d1118d143 Quote a string setting that YAML would read as a number (closes #229)
check / check (push) Waiting to run
config set wrote every value as an unquoted YAML scalar, and config.Load
reads the file through untyped YAML, so an access key 00112233 loaded as
38043 and a hostname 007 as 7.

config set now looks the key up in config.Config by the fields' yaml
tags. A string setting is tagged !!str, which the encoder quotes
wherever YAML would read a number or a boolean. Other settings stay
unquoted, so compression_level 9 is still a number. A value that is not
valid UTF-8 stays untagged and is written as !!binary, which loads back
unchanged.

Judgement call: the type comes from reflection over config.Config.

Model: opus-5-5
2026-10-07 09:29:19 +02:00
clawbot b57ce2277d Store and compare file mtimes to the nanosecond (closes #226)
check / check (push) Waiting to run
The files table held mtime in whole seconds and the scanner compared
whole seconds. A file rewritten with its size unchanged and a new mtime
in the same second as the indexed one was treated as unchanged, and
every later snapshot restored the old content. A new mtime_nsec column
now holds the nanoseconds within the second that mtime holds, and the
scanner compares the full mtime.

A local index created before this change lacks the column and is
rebuilt with `vaultik database delete` and a full backup. A snapshot
made before it cannot be restored by this version.

Model: opus-5-5
2026-10-07 07:12:12 +02:00
clawbot 8496404d8b Make the process-wide lock atomic with flock (closes #227)
check / check (push) Waiting to run
Acquire read vaultik.pid, checked whether that PID was alive, then
wrote its own, so two writers started together could both pass the
check and both run. The lock is now an flock on vaultik.pid, held
while the file stays open; the kernel drops it when the process exits,
so the stale-PID check is gone.

Release empties the file instead of deleting it. Deleting it would let
a process that opened the old file a moment earlier lock it while
another creates and locks a new one.

The new concurrent test fails against the old code only when the race
is hit, not on every run; against the fix it cannot admit two callers.

Model: opus-5-5
2026-10-07 06:12:10 +02:00
clawbot cdc60c4dfa Write only the document to stdout from snapshot remove --json (closes #251)
check / check (push) Waiting to run
When the destination store cannot be reached, `snapshot remove` still
removes the snapshot from the local index and warns. The warning went
through the UI, which writes to stdout, so under `--json` it landed
ahead of the document and `| jq` failed on a command that exited 0.
Under `--json` the UI warning is now skipped; the logger's warning on
stderr carries the follow-up, with the snapshot ID as a field.

The follow-up, in the warning, the README and the command's help, said
`vaultik prune` would finish the cleanup, but `prune` never removes
snapshot metadata. They now say to run `vaultik snapshot remove` for the
snapshot again once the destination store is reachable.

Model: opus-5-5
2026-10-07 04:46:13 +02:00
clawbot 49eed7a3e5 Count each file, byte and upload once in backup statistics (closes #225)
check / check (push) Waiting to run
The scanner added a changed file's bytes again for each new chunk and
counted a file as unchanged for each chunk already stored. It now
counts files and bytes once, in the scan phase, and counts its own
uploads, so a --cron run, which has no progress reporter, records
them. The blob count no longer adds earlier paths' blobs again.

The snapshots row now stores the size of all files in total_size and
the referenced blobs' sizes in blob_size, blob_uncompressed_size and
compression_ratio, as docs/DATAMODEL.md says. Those sizes come from
one query, and a failed query fails the snapshot. DATAMODEL.md now
says chunk_count and blob_count count what the run added.

Removed UpdateSnapshotStats and GetCountBySnapshot, which nothing
calls any more.

Model: opus-5-5
2026-10-07 02:59:27 +02:00
clawbot 85d4ef118d Keep command output to the README's stdout and stderr rules (closes #224)
check / check (push) Successful in 16m49s
The startup banner moves from stdout to stderr, so a `completion`
script, a `config get` value and the hidden `__complete` command print
only their own output. `--quiet`, `--cron` and `--json` still suppress
it.

A failing `remote info`, `prune` or `snapshot remove` under `--json`
now reports its error on stderr. Their reporters returned early under
`--json`, so the failure reached neither stream.

`snapshot verify --quiet` writes no report. A failure is still
returned and printed on stderr, with the same exit status.

Judgement call: the banner's stream, posted on the issue for the owner.

Model: opus-5-5
2026-10-07 01:12:15 +02:00
clawbot b06f992152 Start the progress reporter once per snapshot, not once per path (closes #253)
check / check (push) Successful in 19m54s
A backup without --cron of a snapshot with two or more paths panicked
with "close of closed channel". Scan runs once per path, and it started
the progress reporter and deferred its Stop each time; Stop closes the
reporter's signal channel, so the second path's Stop panicked.
scanAllDirectories now starts the reporter before the first path and
stops it after the last, and Scan no longer starts or stops it. A new
test backs up a two-path snapshot with the reporter on and restores
both paths.

Model: opus-5-5
2026-10-07 00:12:10 +02:00
clawbot 5d685f03ce List the files under a path by path, not by string prefix (closes #223)
check / check (push) Successful in 12m16s
FileRepository.ListByPrefix matched with SQL LIKE: a plain string
prefix that ignores ASCII case and treats _ and % as wildcards.
Restoring /home/u/doc also restored doc2, DOC and doc.txt.bak, and a
backup counted the files of a longer sibling path as deleted. It is
now ListUnderPath, which returns the file at the path and every file
whose path starts with the path plus a slash, compared exactly. A
trailing slash is ignored, so "/" still lists every file.

Three tests relied on string-prefix matching and now name a full path
or call ListAll.

ListIDsWithChunksNotInUploadedBlobs keeps its LIKE: it only adds file
IDs the scan never looks up.

Model: opus-5-5
2026-10-06 22:12:16 +02:00
clawbot f59086c0e5 Return an error, not a panic, on a malformed snapshot database (closes #231)
check / check (push) Successful in 13m39s
Restore cut chunk hashes from the snapshot database to 16 characters
for its error messages, so a shorter hash panicked. Those messages now
use shortHash. Under --verify, a file_chunks row with no chunks row was
dereferenced, and the chunk size from the database was allocated in
one piece, so a negative or huge size panicked. A missing row is now an
error, a negative size is rejected, and each chunk is hashed by
streaming it from the restored file.

A restored file shorter than its chunks now fails verify as a short
read instead of an unexpected EOF.

Model: opus-5-5
2026-10-06 21:16:28 +02:00
clawbot d276d891da Join the S3 prefix to every key with one slash (closes #222)
check / check (push) Successful in 10m59s
The S3 client built each key as prefix + key, and the URL parser keeps
the prefix as written, so s3://bucket/p stored p + "blobs/..." with no
slash while s3://bucket/p/ stored p/blobs/.... A recovery host that wrote
the URL the other way found no snapshots.

NewClient now strips trailing slashes from the prefix and adds one back
when anything is left, giving the README layout for both URL forms; an
empty prefix stays at the bucket root. The s3.prefix config setting
goes through the same client and gets the same join.

A new test writes through each URL shape against an in-process S3
server, checks the key in the bucket, and lists through both List and
ListStream, which every snapshot listing uses.

Model: opus-5-5
2026-10-06 19:46:09 +02:00
clawbot 5aa5ba5544 Apply restored owners, modes and times in an order that keeps them (closes #219)
check / check (push) Successful in 11m53s
A directory got its stored mode and mtime before its contents were
written, so a read-only directory came back without its files and a
non-empty one carried the time of the restore. Directories are now
created owner-only (0700) and get their stored owner, mode and mtime
after the restore loop, each before its parent, skipping any whose
place a symlink has since taken. A file's mode is now applied after its
chown, which on Linux clears setuid and setgid. A symlink gets its
stored owner (as root) and mtime on the link itself, through
golang.org/x/sys/unix, now a direct dependency.

An interrupted restore leaves its directories at 0700.

Model: opus-5-5
2026-10-06 18:29:09 +02:00
clawbot 315b6483b8 Mark github.com/spf13/pflag as a direct dependency in go.mod (closes #246)
check / check (push) Successful in 12m34s
internal/cli/snapshot_restore_test.go imports github.com/spf13/pflag
directly, but go.mod still marked it // indirect. script/precommit runs
go mod tidy and fails when that changes go.mod, so the pre-commit hook
stopped every commit. This is the go mod tidy output: pflag moves to the
direct require block, and go.sum does not change. script/cibuild does
not run the tidy, which is why the gate stayed green.

Model: opus-5-5
2026-10-06 16:59:26 +02:00
clawbot 14fc4c9893 Load a config that has no age recipient (closes #221)
check / check (push) Successful in 13m47s
The README's steps for restoring on another machine failed at the first
command: `config init` wrote a placeholder recipient, and `config.Load`
rejects any recipient that does not parse. `config init` now writes an
empty `age_recipients` list, `config.Load` accepts an empty list, and
`snapshot create` refuses to start without a recipient. A malformed
recipient is still rejected at load.

The recovery-host test now builds its config with `config init` and
`config set` and reads it through `config.Load`, so it imports
`internal/cli`. On a fresh file, `config set age_recipients.0` writes the
list in flow style (`[age1...]`).

Model: opus-5-5
2026-10-06 15:46:13 +02:00
clawbot 66c80a70e3 Skip files of an unreadable blob under restore --skip-errors (closes #218)
check / check (push) Successful in 12m34s
A blob that failed to download ended `snapshot restore` even with
--skip-errors, after restoring whichever files came first. The download
error now goes through the same per-file handling as any other restore
error, once for every pending file that references the blob. With
--skip-errors those files are reported as failed, the rest are
restored, and the command still exits non-zero. Without the flag the
restore still aborts; the error now also names one affected file and
suggests --skip-errors. A cancelled restore still ends at once. The
--skip-errors help text and README now limit the packing and storage
caveat to snapshot creation.

Model: opus-5-5
2026-10-06 14:46:05 +02:00
clawbot 81f83b29f6 Re-vendor the canonical files from sneak/prompts at dd4027b (closes #213)
check / check (push) Successful in 12m47s
Linting and testing become the lint and test phases of the Dockerfile,
and the build stage depends on both. Dockerfile.lint, CHECK_EPOCH and
the tests that checked them are removed. Every docker build in script/
passes --no-cache, and script/cibuild runs script/bootstrap first. A
host without Go gets the go.mod version from script/install-go in
.tool/go, which bootstrap, the Makefile, fmt, fmt-check, precommit and
release add to PATH; fmt-check skips .tool. The image takes its version
from the VERSION build arg or git describe, dev without .git. This
repo's own entries follow the canonical content in .gitignore and
.editorconfig. The golangci-lint v2.14.0 findings are fixed. The rules
in CLAUDE.md move into AGENTS.md. IsDevVersion counts "unknown".

Model: opus-5-5
2026-10-06 12:46:13 +02:00
clawbot c4adb72d80 Run the local index in WAL mode with a busy timeout (closes #217)
check / check (push) Successful in 5m10s
check / check (pull_request) Successful in 5m38s
The connection settings were passed as `_journal_mode=`-style
parameters, which the SQLite driver drops without an error, so the
index ran in rollback-journal mode with no busy timeout. `snapshot
list` or `info` reading during a backup could make the backup's next
write fail with "database is locked". Both open paths now pass
`_pragma=` parameters; foreign keys moved there too.

With WAL on, rows committed to the open index can still be in the
-wal file, which a copy of the main file misses. The metadata export
now copies the index with VACUUM INTO, into an empty 0600 file.

The retry after a failed open no longer claims a TRUNCATE recovery; it
retries with the same settings.

Model: opus-5-5
2026-10-06 11:29:18 +02:00
clawbot 4a167e153a Record the real uid and gid of backed-up files (closes #216)
check / check (push) Successful in 4m58s
check / check (pull_request) Successful in 6m21s
The scanner read uid and gid by asserting the stat result to an
interface with Uid() and Gid() methods. *syscall.Stat_t has Uid and Gid
fields, not methods, so the assertion never matched and every file,
directory and symlink was stored as 0:0; a restore as root then gave
everything to root. The scanner now reads the fields of
*syscall.Stat_t.

The first backup after this change re-reads every file not owned by
root, because its stored uid and gid no longer match the disk.

When the tests run as root, as in the Docker build, the new test
compares 0 with 0 and cannot catch the defect; a non-root run does.

Model: opus-5-5
2026-10-06 09:46:16 +02:00
clawbot ea72697992 List a missing file:// destination directory as an error (closes #220)
check / check (push) Successful in 5m53s
check / check (pull_request) Successful in 4m53s
The file backend listed a destination directory that does not exist as
an empty store. With the volume unplugged, snapshot list reported every
local snapshot as missing from the store, snapshot remove said it had
removed metadata it never reached, and prune dropped every local
snapshot record. List and ListStream now fail when the destination
directory is missing, so those commands take their existing path for a
store that cannot be listed. A missing prefix under an existing
directory is still an empty listing, and a first backup still creates
the directory.

Three tests listed a file:// destination nothing had created; they now
create it.

Model: opus-5-5
2026-10-06 08:46:17 +02:00
clawbot 713be502bd Reject a duration with characters outside its parts (closes #215)
check / check (push) Successful in 7m28s
check / check (pull_request) Successful in 5m39s
parseDuration fell back to an unanchored search for number-and-unit
pieces when time.ParseDuration failed, and skipped everything in
between. 1.0y became 0, so `snapshot create --prune --keep-newer-than
1.0y` deleted every snapshot of the backed-up names, the new one
included. 2.1w became one week and 1,5y five years. The fallback now
requires the whole input to be whole-number-and-unit parts with nothing
between them. A bare number is rejected before time.ParseDuration sees
it, since Go reads 0 and +0 as zero with no unit.

Judgement call: a space between number and unit (`30 days`) was
accepted and is now an error, matching Go's own units.

Model: opus-5-5
2026-10-06 06:12:07 +02:00
clawbot 35cf985c18 Re-chunk a known file whose chunks no uploaded blob holds (closes #214)
check / check (push) Successful in 6m53s
check / check (pull_request) Successful in 6m17s
File rows are shared by every snapshot and updated in place, while a
blob row is deleted once no snapshot references it. Removing the newest
snapshot, or the prune after an interrupted run, could drop the only
blob holding a changed file's current chunks while an older snapshot
kept the file row. The next backup compared metadata only, skipped the
file, and completed a snapshot that could not restore it.

The scanner now loads the IDs of known files that list a chunk no
uploaded blob holds and re-chunks them even when their metadata is
unchanged.

The tests append to a file, so the file keeps its first chunk in a blob
the first snapshot still references. Each backup run gets its own
snapshot name, so the second-precision snapshot IDs differ without
sleeping.

Model: opus-5-5
2026-10-06 04:46:16 +02:00
clawbot 070090124a Stamp the tag or short commit in a plain docker build (closes #211)
check / check (push) Successful in 3m35s
check / check (pull_request) Successful in 3m16s
A plain `docker build .` stamped `dev`: `.dockerignore` left out `.git`
and the Dockerfile defaulted VERSION to `dev`. `.dockerignore` now
sends `.git` without `.git/config`. Given no build arguments, the
builder stamps `git describe --tags --always` and the commit and date
from git, and fails if `.git` is present but yields no version. The
empty CHECK_EPOCH refusal is gone so the plain build succeeds.
`script/version` now prints `git describe --tags --always --dirty`, so
make, the scripts and a plain build agree. `vaultik version` treats
the short commit, tag-N-gHASH forms and any version ending in `-dirty`
as development builds, so they keep the development-build notice.

Model: opus-5-5
2026-10-02 10:04:14 +02:00
clawbot 584444b619 List only after-1.0 work in the README roadmap (closes #208)
check / check (push) Successful in 3m35s
check / check (pull_request) Successful in 3m36s
The README roadmap and the TODO.md Next Step still described finished
1.0 work as remaining. The roadmap now lists only work planned after
1.0. Its security item says the code was reviewed before 1.0, every bug
found was fixed, and the accepted risks are listed; an outside audit
stays as after-1.0 work. The error-condition item is gone because every
failure case it listed has a fault-injection test. Daemon mode is added.
TODO.md says the 1.0 work is complete on next and that merging and
tagging are the owner's.

Judgement call: dropped the human-readable size flags item; no command
flag takes a raw-integer size.

Model: opus-5-5
2026-10-01 21:41:41 +02:00
clawbot b30e79ee45 Run the disk-full restore test again (closes #207)
check / check (push) Successful in 4m52s
check / check (pull_request) Successful in 5m19s
TestRestoreReportsDiskFull was skipped pending
#163, which is closed. The skip
and its "skipped until" wording are removed.

Restore is unchanged. With the skip removed the test failed because
restore succeeded: its simulated full disk capped only Create, but
restore now opens each file with OpenFile, so nothing was capped. It now
caps OpenFile, and only for files under the restore target: restore also
writes the decrypted metadata database under $TMPDIR through the same
filesystem, and capping that would fail the restore before any file
reached the target.

Judgement call: the test was corrected, not restore; both assertions are
unchanged.

Model: opus-5-5
2026-10-01 20:24:30 +02:00
sneak d886a9026f Merge branch 'main' into next
check / check (push) Successful in 4m8s
check / check (pull_request) Successful in 2m51s
2026-09-29 03:02:24 +02:00
clawbot 6e1f499048 Document that migrations are supported and none are added before 1.0 (closes #68)
check / check (push) Successful in 3m6s
check / check (pull_request) Successful in 3m8s
The docs now say vaultik supports migrations. The numbered files in `internal/database/schema/` are migrations: `schema_migrations` records which have run, and opening a database applies any that have not. None are added before 1.0 because nothing is installed anywhere yet, so a schema change edits `001.sql` directly. After 1.0 each change is a new numbered file, and an existing local database is migrated when vaultik is updated.

`docs/DATAMODEL.md` owns the explanation. The README caveat and roadmap entry and `AGENTS.md` policy 13 link to it. This replaces the wording from #146, which said there was no upgrade path.

Disclosure: `CLAUDE.md` line 33, the owner's file, changes from "do not need to support migrations" to "do not add migrations before 1.0".

Model: opus-5-5
2026-09-28 20:21:38 +02:00
sneak 15e6506e5d next: integrate accumulated work into main (#114)
check / check (push) Successful in 2m35s
Reviewed-on: #114
2026-09-23 13:03:02 +02:00
clawbot d24f5dc33c Adopt the canonical golangci-lint config (closes #90)
check / check (push) Successful in 3m9s
check / check (pull_request) Successful in 1m29s
The lint config is now the canonical file from `prompts`, which replaces the deprecated `gomodguard` with `gomodguard_v2`, so lint prints no deprecation warnings. It also turns on the `depguard` `test-support` rule. The one difference from canonical is that the deny list names vaultik's own test-only package `internal/storage/faultstore`, so shipped code cannot import it. The new config found nothing to fix in the source.

Issues and PRs that pin the old `.golangci.yml` sha256 as an untouched-file check need the new one: `7122fcf0dd0ea57441374f98ebd98bb3da23decb67f9209fee5175170838fbd1`.

Model: opus-5-5
2026-09-23 02:14:40 +02:00
clawbot eed117fe25 Route direct-stdout command output through internal/ui (closes #149)
check / check (pull_request) Successful in 1m27s
check / check (push) Successful in 3m18s
version, info, remote info, config, and database delete wrote plain text straight to stdout, so they were unstyled and ignored --quiet. Output is now governed by internal/ui in two buckets. Status lines and confirmations (config init, config set, the database delete prompt) go through the ui message methods and are silenced by --quiet. The data a command exists to produce is written plain -- the version/info/remote-info reports, the snapshot list table, config get values, and the --json documents -- and is never suppressed, since a script depends on it and a marker would corrupt a table or document. The database delete confirmation prompt is always shown. Pure-cli commands reach ui through a small commandUI helper.

Model: opus-4-8
2026-09-22 20:28:50 +02:00
clawbot dd7a610c23 Mark a snapshot complete only after its metadata export succeeds (closes #177)
check / check (pull_request) Successful in 1m42s
check / check (push) Successful in 3m35s
finalizeSnapshotMetadata marked the snapshot complete and then exported its metadata. A crash after completion but before/during the export left the local index showing the snapshot complete while the destination had no manifest or database, and PruneDatabase (which drops only NULL completed_at rows) kept it: a silently unrestorable snapshot.

Reorder so completion is recorded last. CompleteSnapshot is split into PopulateSnapshotBlobs (before the export) and MarkSnapshotComplete (after it). An interrupted export now leaves the snapshot incomplete, so the next run PruneDatabase drops it and re-backs-up the data; the reverse tiny window leaves a restorable snapshot the index reports honestly as remote-only. Update REPOSTRUCTURE.md guarantee 4 and the ARCHITECTURE.md flow. Add a fault-injection test driving the full create path.

Model: opus-4-8
2026-09-22 20:00:41 +02:00
clawbot 1548c0f933 Correct the security claims in docs and comments, and record the accepted risks (closes #171)
check / check (pull_request) Successful in 1m58s
check / check (push) Successful in 2m49s
Docs and comments only; no behaviour change. Corrects ten overclaims the security review found: snapshot names are hashed but the hash uses no secret, so a guessed hostname and name can be confirmed; a blob is named by hex(SHA256(SHA256(uncompressed contents))), stated once in docs/REPOSTRUCTURE.md and referenced elsewhere; double hashing does not hide known content (blob packing does); age uses ChaCha20-Poly1305, not XChaCha20; encryption is required, not optional; a snapshot is marked complete before its metadata is uploaded; the export comment now matches its only caller; deep verify detects corruption, not authorship; adding a recipient does not reach existing data; restore examples target a user-owned directory.

Adds an Accepted Risks subsection under Security Considerations with the seven documented risks, cross-referenced from the README.

Model: opus-4-8
2026-09-22 19:28:31 +02:00
clawbot c3bec7d3aa Reject a metadata database truncated to the age header and nonce (closes #152)
check / check (push) Successful in 2m5s
check / check (pull_request) Successful in 2m28s
An object holding just the age header and its 16-byte nonce decrypts without error: the truncated read surfaces as io.ErrUnexpectedEOF at the age layer, which the zstd decoder maps to a clean EOF at frame start. blobgen then reported zero bytes and no error, so a truncated stream was indistinguishable from a valid empty one, and the metadata database export slipped through -- restore built a fresh schema on the empty file and reported success.

blobgen.Reader.Read now, on EOF, reads once more from the age reader and surfaces io.ErrUnexpectedEOF unless that read is (0, io.EOF), the state a genuine end leaves. downloadSnapshotDB additionally rejects a zero-length decrypted database before any schema is built.

Model: opus-4-8
2026-09-22 18:11:57 +02:00
clawbot 1244c9e48d Bound download expansion and escape control chars on the terminal (closes #164)
check / check (pull_request) Successful in 1m47s
check / check (push) Successful in 3m11s
Objects fetched from the store are untrusted; several decode paths let one expand or print without limit.

- blobgen.LimitReader errors past a byte cap (not io.LimitReader silent EOF). DecodeManifest reads through caps on both compressed input and decompressed output, far above any real manifest, so json.Decode cannot buffer a compressible bomb. FetchAndDecryptBlob bounds decompression to the blob recorded uncompressed_size (not the restoring host blob_size_limit).
- downloadSnapshotDB streams straight from storage to its temp file with io.Copy, replacing two ReadAll calls that held the whole database twice.
- FetchBlob drops the per-blob Stat round-trip, its expectedSize parameter and returned size, all of which only fed a debug log.
- TTYHandler and ui.Writer escape control characters in messages, attribute keys/values, and rendered identifiers/paths before colour codes are applied, so a crafted value cannot drive the terminal.

Model: opus-4-8
2026-09-22 17:00:35 +02:00
clawbot 82c51a5337 Validate blob hashes, offsets and lengths from the destination (closes #155)
check / check (pull_request) Successful in 1m22s
check / check (push) Successful in 3m30s
A blob hash read back from the downloaded snapshot database or the store listing was trusted unchecked. A hostile remote could set a hash such as "aa/../../etc" and have a decrypted blob written outside the cache directory, or feed a short or negative value that panicked a command.

blobDiskCache.path now refuses any key with a path separator, and ReadAt rejects a negative offset or length, bounding so a sum cannot overflow past the check. A new isBlobHash helper gates FetchBlob, shallow and deep verify, and restore: buildBlobIndexes rejects every hash from the snapshot database before any fetch. The blobs/ and metadata/ listings skip a non-conforming name, and short-hash prefixes in log and error text go through a panic-safe shortHash helper.

Model: opus-4-8
2026-09-22 16:11:28 +02:00
clawbot d88ed64489 Parse the age identity key once and accept every identity in it (closes #165)
check / check (pull_request) Successful in 2m41s
check / check (push) Successful in 2m56s
Restore and verify --deep now parse the configured age secret key a single time through a new helper that uses age.ParseIdentities and hands every identity to age.Decrypt. A key file with several identities (a whole age-keygen file) is fully accepted, so a blob encrypted to any of its recipients decrypts, not just the first.

The helper is the first step of both commands, so a missing or unparseable key fails before anything is downloaded. Its error names the config source and never echoes the key value. config.extractAgeSecretKey and its silent fallback are removed; the key is stored raw and parsed only where decryption happens. README, the restore help, and the missing-key error now read the key from a file with \$(cat ...) rather than typed literally, keeping it out of shell history.

Model: opus-4-8
2026-09-22 15:45:27 +02:00
clawbot bd9656dbd4 Reject a decrypted snapshot database that is not the requested one (closes #156)
check / check (push) Successful in 1m26s
check / check (pull_request) Successful in 2m51s
Restore and deep verify downloaded and decrypted metadata/<key>/db.zst.age by object name alone. age decryption proves the database is readable, not that it is the snapshot that was asked for: an attacker who swaps in another valid db.zst.age could redirect the operation, and deep verify with a swapped database plus an empty manifest reported success with zero blobs verified.

After the database is opened, both paths now confirm its identity: an exported per-snapshot database holds one snapshot row, and a snapshot remote key derives from that row ID, so the database is the requested one exactly when its sole snapshot hashes back to the remote key fetched. The shared check lives in verifySnapshotDBIdentity, backed by a new SnapshotRepository.GetOnlySnapshot.

Model: opus-4-8
2026-09-22 15:01:02 +02:00
clawbot 7e611b95db Check blob sizes and the database in shallow verify (closes #169)
check / check (push) Successful in 1m21s
check / check (pull_request) Successful in 2m37s
Shallow snapshot verify only checked that each blob object existed and then reported "All blobs verified", overstating what it did.

It now compares each blob stored size against the manifest compressed_size, using the same comparison as the deep path, and checks that the snapshot encrypted database (db.zst.age) is present. A blob of the wrong size no longer counts as verified. The final line reports only what was checked: presence and size, not contents.

The README verify description and the CLI short/long text are corrected to match. Removed the now-unused resolveAndDownloadManifest helper and errBlobsMissing sentinel.

Model: opus-4-8
2026-09-22 14:28:44 +02:00
clawbot 4f27608560 Add negative and boundary tests for blobgen and types (closes #170)
check / check (push) Successful in 1m22s
check / check (pull_request) Successful in 1m18s
Test-only. internal/blobgen and internal/types had no negative or boundary coverage. Adds, in package blobgen_test: Writer-to-Reader round trips at the 64 KiB age-segment edges for random and compressible data, checking plaintext, byte counts and the reader/writer hashes by decrypting; a wrong-identity open; truncation and single-byte corruption of a multi-segment blob at every region; trailing bytes, empty input and garbage; rejected and accepted compression levels; nil, empty and invalid recipients; and a failing destination. In package types_test: Value/Scan round trips, NULL, wrong-type and malformed Scan, Parse and IsZero for FileID and BlobID.

The "cut right after the age header and nonce" truncation is excluded: it reads as valid and empty today and belongs to #152.

Model: opus-4-8
2026-09-22 14:28:34 +02:00
clawbot ae6aaaa388 Wait for the interrupted operation to clean up before exit (closes #159)
check / check (pull_request) Successful in 1m22s
check / check (push) Successful in 3m13s
On SIGINT/SIGTERM the process could exit before the interrupted command cleanup defers ran, leaving decrypted data in the temp directory (the blob cache and the decrypted snapshot database).

RunApp now mirrors fx run sequence: start, block on app.Wait(), then app.Stop(), returning only after Stop completes. fx delivers both an OS interrupt and the finished operation Shutdowner.Shutdown() on one channel. Stop runs the OnStop hooks; the operation hook cancels the command and waits for its goroutine to return (bounded by shutdownTimeout) before exit. The old code returned as soon as app.Done fired, without Stop, so a real interrupt unwound to os.Exit while cleanup still ran. Restore loops check the context between chunks and blobs so the wait ends promptly. A cli test drives RunApp through the OnStop hook.

Model: opus-4-8
2026-09-22 14:00:49 +02:00
clawbot f788668287 Remove the unused crypto path and write the blob-ID hash step once (closes #151)
check / check (push) Successful in 1m21s
check / check (pull_request) Successful in 2m39s
Production encryption and decryption already run through blobgen; the crypto package (Encryptor, Decryptor, UpdateRecipients, the fx Module) and Vaultik.GetEncryptor/GetDecryptor had no production caller. Delete crypto and route verify --deep through the same blobgen reader restore uses, parsing the age key once.

The second blob-ID hash step is now one exported blobgen.DoubleSHA256; Writer.Sum256 (the double hash) becomes Writer.ContentID so it no longer collides with Reader.Sum256 (the single plaintext hash). Also delete the never-adopted internal/types newtypes and the uncalled CleanupIncompleteSnapshots and its now-dead deleteSnapshot caller, and correct ARCHITECTURE.md. No production behavior changes.

Model: opus-4-8
2026-09-22 13:46:02 +02:00
clawbot 238ce3985f Scrub example config of real credentials and internal hosts (closes #172)
check / check (push) Successful in 1m20s
check / check (pull_request) Successful in 2m41s
config.example.yml carried a real-looking 20-char S3 access key id and 40-char secret, a private-address http:// endpoint, and a storage_url naming an internal rclone remote and pool path. Replace them with the same neutral placeholders the config init template uses: YOUR_ACCESS_KEY / YOUR_SECRET_KEY, a https://s3.example.com endpoint, a mybucket bucket, and rclone://myremote/path/to/backups. No behavior or other keys change.

The credentials live in the commented-out s3 block, which the loader never parses, so the new test reads the file raw text to assert the placeholders are present and no http:// endpoint remains, and also loads it to confirm the active storage_url still parses.

Model: opus-4-8
2026-09-22 13:45:52 +02:00
clawbot 548a7ae156 Give the local index and its export copy an explicit 0600 mode (closes #168)
check / check (pull_request) Successful in 1m23s
check / check (push) Successful in 2m57s
The local index lists every backed-up path and chunk hash, but its file mode was left to the SQLite driver and the umask, so under a typical 022 umask a fresh index (and its -wal/-shm side files) landed world-readable. The snapshot export copied the index to snapshot.db with a permissive create as well.

provideDatabase now calls ensureIndexFileMode before opening the driver: it creates the index 0600 if missing and chmods an existing one to 0600. Doing this before the driver opens the file matters because SQLite gives its -wal and -shm files the mode of the main database file. The export copy is now created 0600. Tests under umask 022 cover a fresh index, an existing 0644 index, and the export copy.

Model: opus-4-8
2026-09-22 13:12:07 +02:00
clawbot 3a58377127 Parse age_recipients at config load and never echo the entry (closes #153)
check / check (push) Successful in 1m23s
check / check (pull_request) Successful in 1m18s
Config.Validate now parses every age_recipients entry with age.ParseX25519Recipient, so a bad recipient fails at config load instead of deep in a backup after the snapshot row and tree walk. On failure the error names the position (age_recipients[N]) and never the value: a recipient string can itself be a secret key an operator pasted by mistake, and age's own error quotes its input. An entry starting with AGE-SECRET-KEY- gets a specific message.

The remaining parse sites (blobgen.NewWriter, crypto NewEncryptor and UpdateRecipients), reachable by callers that skip config.Load, likewise drop the value and age's wrapped error, naming only the position.

Model: opus-4-8
2026-09-22 13:01:00 +02:00
clawbot a6434de57f Open the downloaded snapshot database read-only, on a private temp dir (closes #162)
check / check (pull_request) Successful in 1m21s
check / check (push) Successful in 3m4s
Restore and deep verify used to open the decrypted snapshot database read-write through the local-index constructor, which applied migrations against whatever the file carried, and left the decrypted file in the shared temp directory. A forged file could redefine what restore queries return, and an interrupted open left decrypted metadata on disk.

Add database.OpenReadOnly: opens the file read-only (mode=ro) with query_only and trusted_schema=OFF, never applies schema files, and refuses a file whose schema carries a trigger, view or virtual table or lacks an expected table. Restore and deep verify now both use it, each inside its own private (0700) temp directory removed on every return path. pickNextDownload returns (FileID, bool) so a genuine nil-UUID file is not mistaken for "nothing left".

Model: opus-4-8
2026-09-22 12:45:53 +02:00
clawbot b4654f8e52 Abort the run when packing fails, even under --skip-errors (closes #161)
check / check (push) Successful in 1m22s
check / check (pull_request) Successful in 3m2s
A chunk is registered as pending (known, scanner-pending, packer pending-row) before it is packed. Under --skip-errors the scanner skipped a file on any processing error, including a failure inside addChunkToPacker (packing, database, encryption, upload). The pending chunk then stayed queued and a later blob finalize inserted it into the chunks table with no blob_chunks row, so a snapshot could complete holding a file whose chunk is in no blob and cannot be restored.

Errors from addChunkToPacker are now marked and abort the run regardless of --skip-errors; only open and read errors are skipped. The bookkeeping order is unchanged. Flag help and comments now say only unreadable files are skipped.

Model: opus-4-8
2026-09-22 12:28:44 +02:00
clawbot 39aef1c47c Stop config set echoing secrets; reject credential-bearing storage URLs (closes #166)
check / check (push) Successful in 1m21s
check / check (pull_request) Successful in 1m18s
config set now prints only the key name after a write, never the value: a value may be a secret such as s3.secret_access_key, and echoing it leaks into captured stdout and pasted terminals. The set logic moves into writeConfigSet so this is testable.

config set also tightens a pre-existing group- or world-readable config to 0600 after writing; the previous stat-and-preserve-mode block had no effect (os.WriteFile does not change an existing file mode) and is removed.

ParseStorageURL now rejects s3:// and rclone:// URLs that carry credentials in the userinfo or an unknown query parameter, naming s3.access_key_id and s3.secret_access_key as where credentials belong. On a url.Parse failure only the inner cause is wrapped, so the raw URL is not echoed. file:// is unchanged.

Model: opus-4-8
2026-09-22 12:28:32 +02:00
clawbot 96ebcd40d7 Reconcile purge against remote by hashed key, not human ID (closes #160)
check / check (pull_request) Successful in 1m20s
check / check (push) Successful in 2m51s
syncWithRemote compared human snapshot IDs against the hashed metadata/<key>/ directory names, which never match, so it deleted every local snapshot record; the purge that followed then found nothing to remove remotely. Reconcile via listAllRemoteSnapshotKeys and RemoteSnapshotKey(id), matching CleanupLocalSnapshots, so a row still backed by remote metadata is kept.

The purge tests only passed because their stubs used the human-ID layout production never writes; they now write metadata under the hashed remote key. New tests prove remotely-backed local rows survive the reconcile and that a purge removes the local row and remote metadata together.

Model: opus-4-8
2026-09-22 12:11:49 +02:00
clawbot d9f0220f94 Restore files at 0600 and make the blob hash check unskippable (closes #163)
check / check (pull_request) Successful in 1m21s
check / check (push) Successful in 2m42s
Regular files are now created with O_EXCL at mode 0600 and given their stored mode only after the content is written and closed, so a file whose stored mode is restrictive is never briefly readable by other local users mid-restore. A file whose write or close fails is removed rather than left partial, and a chmod failure is a user-visible warning instead of a debug line.

hashVerifyReader.Close now errors when closed before EOF, so a short read or early close can never obtain a blob whose hash was not verified; downloadBlobToCache drops the cache entry on any such failure.

verifyFile (--verify) now rejects a restored file with bytes past its last chunk. Tests cover each behaviour under umask 022.

Model: opus-4-8
2026-09-22 11:45:52 +02:00
clawbot 4c83e82543 Reject a blob_size_limit below the largest possible chunk (closes #167)
check / check (push) Successful in 1m20s
check / check (pull_request) Successful in 1m16s
Validate only rejected blob_size_limit below chunk_size, but the chunker can emit chunks up to chunk_size times the FastCDC size spread (four times), and the packer puts a single chunk of any size into an otherwise empty blob. A limit between one and four times chunk_size therefore let a blob reach four times the configured maximum, with most blobs holding a single chunk and so exposing individual chunk lengths to anyone who can list the destination.

Validate now rejects blob_size_limit below chunk_size times the spread, reusing the chunker's one constant (now exported as ChunkSizeSpread) instead of a second literal. The rule is stated in the error text, the Validate comment, the README config table, config.example.yml, and the generated config template.

Model: opus-4-8
2026-09-22 11:45:41 +02:00
clawbot 86361c8b50 Fail closed on unreadable manifests instead of losing blobs (closes #157)
check / check (pull_request) Successful in 1m20s
check / check (push) Successful in 2m42s
Prune learned which blobs are in use by reading every snapshot's manifest, but merely logged and skipped one it could not download or decode. Blobs referenced only by that snapshot then looked unreferenced and were deleted, with a zero exit -- and snapshot create --prune runs this unattended. collectReferencedBlobs now errors, naming the remote key, so prune deletes nothing and exits non-zero.

Manifest generation likewise skipped a blob whose lookup failed or was missing, yielding a manifest short of what the snapshot needs; it now fails. Deep verify only warned when the manifest omitted a database blob; it now fails on any divergence. Docs corrected.

Model: opus-4-8
2026-09-22 11:45:30 +02:00