Commit Graph
209 Commits
Author SHA1 Message Date
sneak c113e120d6 Leave remote info orphan figures unknown when a manifest is unreadable (closes #228)
check / check (push) Waiting to run
When a manifest could not be read, remote info skipped it, counted that
snapshot's blobs as orphaned and advised running prune. The orphan
figures are now unknown in that case: the report says how many
manifests could not be read and gives no prune advice, and --json gives
orphaned_blob_count and orphaned_blob_size as null and lists the remote
keys in unreadable_manifests.

Directory names under metadata/ were used unchecked and printed raw, so
control characters in one reached the terminal. A name that is not a
remote key (64 lowercase hex characters) is now skipped with a warning,
as the snapshot listing already does.

The closing log line of remote info now carries the unreadable manifest
count in place of the orphan count.

Model: opus-5-5
2026-10-07 04:36:04 +00:00
clawbot 8496404d8b Make the process-wide lock atomic with flock (closes #227)
check / check (push) Waiting to run
Acquire read vaultik.pid, checked whether that PID was alive, then
wrote its own, so two writers started together could both pass the
check and both run. The lock is now an flock on vaultik.pid, held
while the file stays open; the kernel drops it when the process exits,
so the stale-PID check is gone.

Release empties the file instead of deleting it. Deleting it would let
a process that opened the old file a moment earlier lock it while
another creates and locks a new one.

The new concurrent test fails against the old code only when the race
is hit, not on every run; against the fix it cannot admit two callers.

Model: opus-5-5
2026-10-07 06:12:10 +02:00
clawbot cdc60c4dfa Write only the document to stdout from snapshot remove --json (closes #251)
check / check (push) Waiting to run
When the destination store cannot be reached, `snapshot remove` still
removes the snapshot from the local index and warns. The warning went
through the UI, which writes to stdout, so under `--json` it landed
ahead of the document and `| jq` failed on a command that exited 0.
Under `--json` the UI warning is now skipped; the logger's warning on
stderr carries the follow-up, with the snapshot ID as a field.

The follow-up, in the warning, the README and the command's help, said
`vaultik prune` would finish the cleanup, but `prune` never removes
snapshot metadata. They now say to run `vaultik snapshot remove` for the
snapshot again once the destination store is reachable.

Model: opus-5-5
2026-10-07 04:46:13 +02:00
clawbot 49eed7a3e5 Count each file, byte and upload once in backup statistics (closes #225)
check / check (push) Waiting to run
The scanner added a changed file's bytes again for each new chunk and
counted a file as unchanged for each chunk already stored. It now
counts files and bytes once, in the scan phase, and counts its own
uploads, so a --cron run, which has no progress reporter, records
them. The blob count no longer adds earlier paths' blobs again.

The snapshots row now stores the size of all files in total_size and
the referenced blobs' sizes in blob_size, blob_uncompressed_size and
compression_ratio, as docs/DATAMODEL.md says. Those sizes come from
one query, and a failed query fails the snapshot. DATAMODEL.md now
says chunk_count and blob_count count what the run added.

Removed UpdateSnapshotStats and GetCountBySnapshot, which nothing
calls any more.

Model: opus-5-5
2026-10-07 02:59:27 +02:00
clawbot 85d4ef118d Keep command output to the README's stdout and stderr rules (closes #224)
check / check (push) Successful in 16m49s
The startup banner moves from stdout to stderr, so a `completion`
script, a `config get` value and the hidden `__complete` command print
only their own output. `--quiet`, `--cron` and `--json` still suppress
it.

A failing `remote info`, `prune` or `snapshot remove` under `--json`
now reports its error on stderr. Their reporters returned early under
`--json`, so the failure reached neither stream.

`snapshot verify --quiet` writes no report. A failure is still
returned and printed on stderr, with the same exit status.

Judgement call: the banner's stream, posted on the issue for the owner.

Model: opus-5-5
2026-10-07 01:12:15 +02:00
clawbot b06f992152 Start the progress reporter once per snapshot, not once per path (closes #253)
check / check (push) Successful in 19m54s
A backup without --cron of a snapshot with two or more paths panicked
with "close of closed channel". Scan runs once per path, and it started
the progress reporter and deferred its Stop each time; Stop closes the
reporter's signal channel, so the second path's Stop panicked.
scanAllDirectories now starts the reporter before the first path and
stops it after the last, and Scan no longer starts or stops it. A new
test backs up a two-path snapshot with the reporter on and restores
both paths.

Model: opus-5-5
2026-10-07 00:12:10 +02:00
clawbot 5d685f03ce List the files under a path by path, not by string prefix (closes #223)
check / check (push) Successful in 12m16s
FileRepository.ListByPrefix matched with SQL LIKE: a plain string
prefix that ignores ASCII case and treats _ and % as wildcards.
Restoring /home/u/doc also restored doc2, DOC and doc.txt.bak, and a
backup counted the files of a longer sibling path as deleted. It is
now ListUnderPath, which returns the file at the path and every file
whose path starts with the path plus a slash, compared exactly. A
trailing slash is ignored, so "/" still lists every file.

Three tests relied on string-prefix matching and now name a full path
or call ListAll.

ListIDsWithChunksNotInUploadedBlobs keeps its LIKE: it only adds file
IDs the scan never looks up.

Model: opus-5-5
2026-10-06 22:12:16 +02:00
clawbot f59086c0e5 Return an error, not a panic, on a malformed snapshot database (closes #231)
check / check (push) Successful in 13m39s
Restore cut chunk hashes from the snapshot database to 16 characters
for its error messages, so a shorter hash panicked. Those messages now
use shortHash. Under --verify, a file_chunks row with no chunks row was
dereferenced, and the chunk size from the database was allocated in
one piece, so a negative or huge size panicked. A missing row is now an
error, a negative size is rejected, and each chunk is hashed by
streaming it from the restored file.

A restored file shorter than its chunks now fails verify as a short
read instead of an unexpected EOF.

Model: opus-5-5
2026-10-06 21:16:28 +02:00
clawbot d276d891da Join the S3 prefix to every key with one slash (closes #222)
check / check (push) Successful in 10m59s
The S3 client built each key as prefix + key, and the URL parser keeps
the prefix as written, so s3://bucket/p stored p + "blobs/..." with no
slash while s3://bucket/p/ stored p/blobs/.... A recovery host that wrote
the URL the other way found no snapshots.

NewClient now strips trailing slashes from the prefix and adds one back
when anything is left, giving the README layout for both URL forms; an
empty prefix stays at the bucket root. The s3.prefix config setting
goes through the same client and gets the same join.

A new test writes through each URL shape against an in-process S3
server, checks the key in the bucket, and lists through both List and
ListStream, which every snapshot listing uses.

Model: opus-5-5
2026-10-06 19:46:09 +02:00
clawbot 5aa5ba5544 Apply restored owners, modes and times in an order that keeps them (closes #219)
check / check (push) Successful in 11m53s
A directory got its stored mode and mtime before its contents were
written, so a read-only directory came back without its files and a
non-empty one carried the time of the restore. Directories are now
created owner-only (0700) and get their stored owner, mode and mtime
after the restore loop, each before its parent, skipping any whose
place a symlink has since taken. A file's mode is now applied after its
chown, which on Linux clears setuid and setgid. A symlink gets its
stored owner (as root) and mtime on the link itself, through
golang.org/x/sys/unix, now a direct dependency.

An interrupted restore leaves its directories at 0700.

Model: opus-5-5
2026-10-06 18:29:09 +02:00
clawbot 14fc4c9893 Load a config that has no age recipient (closes #221)
check / check (push) Successful in 13m47s
The README's steps for restoring on another machine failed at the first
command: `config init` wrote a placeholder recipient, and `config.Load`
rejects any recipient that does not parse. `config init` now writes an
empty `age_recipients` list, `config.Load` accepts an empty list, and
`snapshot create` refuses to start without a recipient. A malformed
recipient is still rejected at load.

The recovery-host test now builds its config with `config init` and
`config set` and reads it through `config.Load`, so it imports
`internal/cli`. On a fresh file, `config set age_recipients.0` writes the
list in flow style (`[age1...]`).

Model: opus-5-5
2026-10-06 15:46:13 +02:00
clawbot 66c80a70e3 Skip files of an unreadable blob under restore --skip-errors (closes #218)
check / check (push) Successful in 12m34s
A blob that failed to download ended `snapshot restore` even with
--skip-errors, after restoring whichever files came first. The download
error now goes through the same per-file handling as any other restore
error, once for every pending file that references the blob. With
--skip-errors those files are reported as failed, the rest are
restored, and the command still exits non-zero. Without the flag the
restore still aborts; the error now also names one affected file and
suggests --skip-errors. A cancelled restore still ends at once. The
--skip-errors help text and README now limit the packing and storage
caveat to snapshot creation.

Model: opus-5-5
2026-10-06 14:46:05 +02:00
clawbot 81f83b29f6 Re-vendor the canonical files from sneak/prompts at dd4027b (closes #213)
check / check (push) Successful in 12m47s
Linting and testing become the lint and test phases of the Dockerfile,
and the build stage depends on both. Dockerfile.lint, CHECK_EPOCH and
the tests that checked them are removed. Every docker build in script/
passes --no-cache, and script/cibuild runs script/bootstrap first. A
host without Go gets the go.mod version from script/install-go in
.tool/go, which bootstrap, the Makefile, fmt, fmt-check, precommit and
release add to PATH; fmt-check skips .tool. The image takes its version
from the VERSION build arg or git describe, dev without .git. This
repo's own entries follow the canonical content in .gitignore and
.editorconfig. The golangci-lint v2.14.0 findings are fixed. The rules
in CLAUDE.md move into AGENTS.md. IsDevVersion counts "unknown".

Model: opus-5-5
2026-10-06 12:46:13 +02:00
clawbot c4adb72d80 Run the local index in WAL mode with a busy timeout (closes #217)
check / check (push) Successful in 5m10s
check / check (pull_request) Successful in 5m38s
The connection settings were passed as `_journal_mode=`-style
parameters, which the SQLite driver drops without an error, so the
index ran in rollback-journal mode with no busy timeout. `snapshot
list` or `info` reading during a backup could make the backup's next
write fail with "database is locked". Both open paths now pass
`_pragma=` parameters; foreign keys moved there too.

With WAL on, rows committed to the open index can still be in the
-wal file, which a copy of the main file misses. The metadata export
now copies the index with VACUUM INTO, into an empty 0600 file.

The retry after a failed open no longer claims a TRUNCATE recovery; it
retries with the same settings.

Model: opus-5-5
2026-10-06 11:29:18 +02:00
clawbot 4a167e153a Record the real uid and gid of backed-up files (closes #216)
check / check (push) Successful in 4m58s
check / check (pull_request) Successful in 6m21s
The scanner read uid and gid by asserting the stat result to an
interface with Uid() and Gid() methods. *syscall.Stat_t has Uid and Gid
fields, not methods, so the assertion never matched and every file,
directory and symlink was stored as 0:0; a restore as root then gave
everything to root. The scanner now reads the fields of
*syscall.Stat_t.

The first backup after this change re-reads every file not owned by
root, because its stored uid and gid no longer match the disk.

When the tests run as root, as in the Docker build, the new test
compares 0 with 0 and cannot catch the defect; a non-root run does.

Model: opus-5-5
2026-10-06 09:46:16 +02:00
clawbot ea72697992 List a missing file:// destination directory as an error (closes #220)
check / check (push) Successful in 5m53s
check / check (pull_request) Successful in 4m53s
The file backend listed a destination directory that does not exist as
an empty store. With the volume unplugged, snapshot list reported every
local snapshot as missing from the store, snapshot remove said it had
removed metadata it never reached, and prune dropped every local
snapshot record. List and ListStream now fail when the destination
directory is missing, so those commands take their existing path for a
store that cannot be listed. A missing prefix under an existing
directory is still an empty listing, and a first backup still creates
the directory.

Three tests listed a file:// destination nothing had created; they now
create it.

Model: opus-5-5
2026-10-06 08:46:17 +02:00
clawbot 713be502bd Reject a duration with characters outside its parts (closes #215)
check / check (push) Successful in 7m28s
check / check (pull_request) Successful in 5m39s
parseDuration fell back to an unanchored search for number-and-unit
pieces when time.ParseDuration failed, and skipped everything in
between. 1.0y became 0, so `snapshot create --prune --keep-newer-than
1.0y` deleted every snapshot of the backed-up names, the new one
included. 2.1w became one week and 1,5y five years. The fallback now
requires the whole input to be whole-number-and-unit parts with nothing
between them. A bare number is rejected before time.ParseDuration sees
it, since Go reads 0 and +0 as zero with no unit.

Judgement call: a space between number and unit (`30 days`) was
accepted and is now an error, matching Go's own units.

Model: opus-5-5
2026-10-06 06:12:07 +02:00
clawbot 35cf985c18 Re-chunk a known file whose chunks no uploaded blob holds (closes #214)
check / check (push) Successful in 6m53s
check / check (pull_request) Successful in 6m17s
File rows are shared by every snapshot and updated in place, while a
blob row is deleted once no snapshot references it. Removing the newest
snapshot, or the prune after an interrupted run, could drop the only
blob holding a changed file's current chunks while an older snapshot
kept the file row. The next backup compared metadata only, skipped the
file, and completed a snapshot that could not restore it.

The scanner now loads the IDs of known files that list a chunk no
uploaded blob holds and re-chunks them even when their metadata is
unchanged.

The tests append to a file, so the file keeps its first chunk in a blob
the first snapshot still references. Each backup run gets its own
snapshot name, so the second-precision snapshot IDs differ without
sleeping.

Model: opus-5-5
2026-10-06 04:46:16 +02:00
clawbot 070090124a Stamp the tag or short commit in a plain docker build (closes #211)
check / check (push) Successful in 3m35s
check / check (pull_request) Successful in 3m16s
A plain `docker build .` stamped `dev`: `.dockerignore` left out `.git`
and the Dockerfile defaulted VERSION to `dev`. `.dockerignore` now
sends `.git` without `.git/config`. Given no build arguments, the
builder stamps `git describe --tags --always` and the commit and date
from git, and fails if `.git` is present but yields no version. The
empty CHECK_EPOCH refusal is gone so the plain build succeeds.
`script/version` now prints `git describe --tags --always --dirty`, so
make, the scripts and a plain build agree. `vaultik version` treats
the short commit, tag-N-gHASH forms and any version ending in `-dirty`
as development builds, so they keep the development-build notice.

Model: opus-5-5
2026-10-02 10:04:14 +02:00
clawbot b30e79ee45 Run the disk-full restore test again (closes #207)
check / check (push) Successful in 4m52s
check / check (pull_request) Successful in 5m19s
TestRestoreReportsDiskFull was skipped pending
#163, which is closed. The skip
and its "skipped until" wording are removed.

Restore is unchanged. With the skip removed the test failed because
restore succeeded: its simulated full disk capped only Create, but
restore now opens each file with OpenFile, so nothing was capped. It now
caps OpenFile, and only for files under the restore target: restore also
writes the decrypted metadata database under $TMPDIR through the same
filesystem, and capping that would fail the restore before any file
reached the target.

Judgement call: the test was corrected, not restore; both assertions are
unchanged.

Model: opus-5-5
2026-10-01 20:24:30 +02:00
clawbot eed117fe25 Route direct-stdout command output through internal/ui (closes #149)
check / check (pull_request) Successful in 1m27s
check / check (push) Successful in 3m18s
version, info, remote info, config, and database delete wrote plain text straight to stdout, so they were unstyled and ignored --quiet. Output is now governed by internal/ui in two buckets. Status lines and confirmations (config init, config set, the database delete prompt) go through the ui message methods and are silenced by --quiet. The data a command exists to produce is written plain -- the version/info/remote-info reports, the snapshot list table, config get values, and the --json documents -- and is never suppressed, since a script depends on it and a marker would corrupt a table or document. The database delete confirmation prompt is always shown. Pure-cli commands reach ui through a small commandUI helper.

Model: opus-4-8
2026-09-22 20:28:50 +02:00
clawbot dd7a610c23 Mark a snapshot complete only after its metadata export succeeds (closes #177)
check / check (pull_request) Successful in 1m42s
check / check (push) Successful in 3m35s
finalizeSnapshotMetadata marked the snapshot complete and then exported its metadata. A crash after completion but before/during the export left the local index showing the snapshot complete while the destination had no manifest or database, and PruneDatabase (which drops only NULL completed_at rows) kept it: a silently unrestorable snapshot.

Reorder so completion is recorded last. CompleteSnapshot is split into PopulateSnapshotBlobs (before the export) and MarkSnapshotComplete (after it). An interrupted export now leaves the snapshot incomplete, so the next run PruneDatabase drops it and re-backs-up the data; the reverse tiny window leaves a restorable snapshot the index reports honestly as remote-only. Update REPOSTRUCTURE.md guarantee 4 and the ARCHITECTURE.md flow. Add a fault-injection test driving the full create path.

Model: opus-4-8
2026-09-22 20:00:41 +02:00
clawbot 1548c0f933 Correct the security claims in docs and comments, and record the accepted risks (closes #171)
check / check (pull_request) Successful in 1m58s
check / check (push) Successful in 2m49s
Docs and comments only; no behaviour change. Corrects ten overclaims the security review found: snapshot names are hashed but the hash uses no secret, so a guessed hostname and name can be confirmed; a blob is named by hex(SHA256(SHA256(uncompressed contents))), stated once in docs/REPOSTRUCTURE.md and referenced elsewhere; double hashing does not hide known content (blob packing does); age uses ChaCha20-Poly1305, not XChaCha20; encryption is required, not optional; a snapshot is marked complete before its metadata is uploaded; the export comment now matches its only caller; deep verify detects corruption, not authorship; adding a recipient does not reach existing data; restore examples target a user-owned directory.

Adds an Accepted Risks subsection under Security Considerations with the seven documented risks, cross-referenced from the README.

Model: opus-4-8
2026-09-22 19:28:31 +02:00
clawbot c3bec7d3aa Reject a metadata database truncated to the age header and nonce (closes #152)
check / check (push) Successful in 2m5s
check / check (pull_request) Successful in 2m28s
An object holding just the age header and its 16-byte nonce decrypts without error: the truncated read surfaces as io.ErrUnexpectedEOF at the age layer, which the zstd decoder maps to a clean EOF at frame start. blobgen then reported zero bytes and no error, so a truncated stream was indistinguishable from a valid empty one, and the metadata database export slipped through -- restore built a fresh schema on the empty file and reported success.

blobgen.Reader.Read now, on EOF, reads once more from the age reader and surfaces io.ErrUnexpectedEOF unless that read is (0, io.EOF), the state a genuine end leaves. downloadSnapshotDB additionally rejects a zero-length decrypted database before any schema is built.

Model: opus-4-8
2026-09-22 18:11:57 +02:00
clawbot 1244c9e48d Bound download expansion and escape control chars on the terminal (closes #164)
check / check (pull_request) Successful in 1m47s
check / check (push) Successful in 3m11s
Objects fetched from the store are untrusted; several decode paths let one expand or print without limit.

- blobgen.LimitReader errors past a byte cap (not io.LimitReader silent EOF). DecodeManifest reads through caps on both compressed input and decompressed output, far above any real manifest, so json.Decode cannot buffer a compressible bomb. FetchAndDecryptBlob bounds decompression to the blob recorded uncompressed_size (not the restoring host blob_size_limit).
- downloadSnapshotDB streams straight from storage to its temp file with io.Copy, replacing two ReadAll calls that held the whole database twice.
- FetchBlob drops the per-blob Stat round-trip, its expectedSize parameter and returned size, all of which only fed a debug log.
- TTYHandler and ui.Writer escape control characters in messages, attribute keys/values, and rendered identifiers/paths before colour codes are applied, so a crafted value cannot drive the terminal.

Model: opus-4-8
2026-09-22 17:00:35 +02:00
clawbot 82c51a5337 Validate blob hashes, offsets and lengths from the destination (closes #155)
check / check (pull_request) Successful in 1m22s
check / check (push) Successful in 3m30s
A blob hash read back from the downloaded snapshot database or the store listing was trusted unchecked. A hostile remote could set a hash such as "aa/../../etc" and have a decrypted blob written outside the cache directory, or feed a short or negative value that panicked a command.

blobDiskCache.path now refuses any key with a path separator, and ReadAt rejects a negative offset or length, bounding so a sum cannot overflow past the check. A new isBlobHash helper gates FetchBlob, shallow and deep verify, and restore: buildBlobIndexes rejects every hash from the snapshot database before any fetch. The blobs/ and metadata/ listings skip a non-conforming name, and short-hash prefixes in log and error text go through a panic-safe shortHash helper.

Model: opus-4-8
2026-09-22 16:11:28 +02:00
clawbot d88ed64489 Parse the age identity key once and accept every identity in it (closes #165)
check / check (pull_request) Successful in 2m41s
check / check (push) Successful in 2m56s
Restore and verify --deep now parse the configured age secret key a single time through a new helper that uses age.ParseIdentities and hands every identity to age.Decrypt. A key file with several identities (a whole age-keygen file) is fully accepted, so a blob encrypted to any of its recipients decrypts, not just the first.

The helper is the first step of both commands, so a missing or unparseable key fails before anything is downloaded. Its error names the config source and never echoes the key value. config.extractAgeSecretKey and its silent fallback are removed; the key is stored raw and parsed only where decryption happens. README, the restore help, and the missing-key error now read the key from a file with \$(cat ...) rather than typed literally, keeping it out of shell history.

Model: opus-4-8
2026-09-22 15:45:27 +02:00
clawbot bd9656dbd4 Reject a decrypted snapshot database that is not the requested one (closes #156)
check / check (push) Successful in 1m26s
check / check (pull_request) Successful in 2m51s
Restore and deep verify downloaded and decrypted metadata/<key>/db.zst.age by object name alone. age decryption proves the database is readable, not that it is the snapshot that was asked for: an attacker who swaps in another valid db.zst.age could redirect the operation, and deep verify with a swapped database plus an empty manifest reported success with zero blobs verified.

After the database is opened, both paths now confirm its identity: an exported per-snapshot database holds one snapshot row, and a snapshot remote key derives from that row ID, so the database is the requested one exactly when its sole snapshot hashes back to the remote key fetched. The shared check lives in verifySnapshotDBIdentity, backed by a new SnapshotRepository.GetOnlySnapshot.

Model: opus-4-8
2026-09-22 15:01:02 +02:00
clawbot 7e611b95db Check blob sizes and the database in shallow verify (closes #169)
check / check (push) Successful in 1m21s
check / check (pull_request) Successful in 2m37s
Shallow snapshot verify only checked that each blob object existed and then reported "All blobs verified", overstating what it did.

It now compares each blob stored size against the manifest compressed_size, using the same comparison as the deep path, and checks that the snapshot encrypted database (db.zst.age) is present. A blob of the wrong size no longer counts as verified. The final line reports only what was checked: presence and size, not contents.

The README verify description and the CLI short/long text are corrected to match. Removed the now-unused resolveAndDownloadManifest helper and errBlobsMissing sentinel.

Model: opus-4-8
2026-09-22 14:28:44 +02:00
clawbot 4f27608560 Add negative and boundary tests for blobgen and types (closes #170)
check / check (push) Successful in 1m22s
check / check (pull_request) Successful in 1m18s
Test-only. internal/blobgen and internal/types had no negative or boundary coverage. Adds, in package blobgen_test: Writer-to-Reader round trips at the 64 KiB age-segment edges for random and compressible data, checking plaintext, byte counts and the reader/writer hashes by decrypting; a wrong-identity open; truncation and single-byte corruption of a multi-segment blob at every region; trailing bytes, empty input and garbage; rejected and accepted compression levels; nil, empty and invalid recipients; and a failing destination. In package types_test: Value/Scan round trips, NULL, wrong-type and malformed Scan, Parse and IsZero for FileID and BlobID.

The "cut right after the age header and nonce" truncation is excluded: it reads as valid and empty today and belongs to #152.

Model: opus-4-8
2026-09-22 14:28:34 +02:00
clawbot ae6aaaa388 Wait for the interrupted operation to clean up before exit (closes #159)
check / check (pull_request) Successful in 1m22s
check / check (push) Successful in 3m13s
On SIGINT/SIGTERM the process could exit before the interrupted command cleanup defers ran, leaving decrypted data in the temp directory (the blob cache and the decrypted snapshot database).

RunApp now mirrors fx run sequence: start, block on app.Wait(), then app.Stop(), returning only after Stop completes. fx delivers both an OS interrupt and the finished operation Shutdowner.Shutdown() on one channel. Stop runs the OnStop hooks; the operation hook cancels the command and waits for its goroutine to return (bounded by shutdownTimeout) before exit. The old code returned as soon as app.Done fired, without Stop, so a real interrupt unwound to os.Exit while cleanup still ran. Restore loops check the context between chunks and blobs so the wait ends promptly. A cli test drives RunApp through the OnStop hook.

Model: opus-4-8
2026-09-22 14:00:49 +02:00
clawbot f788668287 Remove the unused crypto path and write the blob-ID hash step once (closes #151)
check / check (push) Successful in 1m21s
check / check (pull_request) Successful in 2m39s
Production encryption and decryption already run through blobgen; the crypto package (Encryptor, Decryptor, UpdateRecipients, the fx Module) and Vaultik.GetEncryptor/GetDecryptor had no production caller. Delete crypto and route verify --deep through the same blobgen reader restore uses, parsing the age key once.

The second blob-ID hash step is now one exported blobgen.DoubleSHA256; Writer.Sum256 (the double hash) becomes Writer.ContentID so it no longer collides with Reader.Sum256 (the single plaintext hash). Also delete the never-adopted internal/types newtypes and the uncalled CleanupIncompleteSnapshots and its now-dead deleteSnapshot caller, and correct ARCHITECTURE.md. No production behavior changes.

Model: opus-4-8
2026-09-22 13:46:02 +02:00
clawbot 238ce3985f Scrub example config of real credentials and internal hosts (closes #172)
check / check (push) Successful in 1m20s
check / check (pull_request) Successful in 2m41s
config.example.yml carried a real-looking 20-char S3 access key id and 40-char secret, a private-address http:// endpoint, and a storage_url naming an internal rclone remote and pool path. Replace them with the same neutral placeholders the config init template uses: YOUR_ACCESS_KEY / YOUR_SECRET_KEY, a https://s3.example.com endpoint, a mybucket bucket, and rclone://myremote/path/to/backups. No behavior or other keys change.

The credentials live in the commented-out s3 block, which the loader never parses, so the new test reads the file raw text to assert the placeholders are present and no http:// endpoint remains, and also loads it to confirm the active storage_url still parses.

Model: opus-4-8
2026-09-22 13:45:52 +02:00
clawbot 548a7ae156 Give the local index and its export copy an explicit 0600 mode (closes #168)
check / check (pull_request) Successful in 1m23s
check / check (push) Successful in 2m57s
The local index lists every backed-up path and chunk hash, but its file mode was left to the SQLite driver and the umask, so under a typical 022 umask a fresh index (and its -wal/-shm side files) landed world-readable. The snapshot export copied the index to snapshot.db with a permissive create as well.

provideDatabase now calls ensureIndexFileMode before opening the driver: it creates the index 0600 if missing and chmods an existing one to 0600. Doing this before the driver opens the file matters because SQLite gives its -wal and -shm files the mode of the main database file. The export copy is now created 0600. Tests under umask 022 cover a fresh index, an existing 0644 index, and the export copy.

Model: opus-4-8
2026-09-22 13:12:07 +02:00
clawbot 3a58377127 Parse age_recipients at config load and never echo the entry (closes #153)
check / check (push) Successful in 1m23s
check / check (pull_request) Successful in 1m18s
Config.Validate now parses every age_recipients entry with age.ParseX25519Recipient, so a bad recipient fails at config load instead of deep in a backup after the snapshot row and tree walk. On failure the error names the position (age_recipients[N]) and never the value: a recipient string can itself be a secret key an operator pasted by mistake, and age's own error quotes its input. An entry starting with AGE-SECRET-KEY- gets a specific message.

The remaining parse sites (blobgen.NewWriter, crypto NewEncryptor and UpdateRecipients), reachable by callers that skip config.Load, likewise drop the value and age's wrapped error, naming only the position.

Model: opus-4-8
2026-09-22 13:01:00 +02:00
clawbot a6434de57f Open the downloaded snapshot database read-only, on a private temp dir (closes #162)
check / check (pull_request) Successful in 1m21s
check / check (push) Successful in 3m4s
Restore and deep verify used to open the decrypted snapshot database read-write through the local-index constructor, which applied migrations against whatever the file carried, and left the decrypted file in the shared temp directory. A forged file could redefine what restore queries return, and an interrupted open left decrypted metadata on disk.

Add database.OpenReadOnly: opens the file read-only (mode=ro) with query_only and trusted_schema=OFF, never applies schema files, and refuses a file whose schema carries a trigger, view or virtual table or lacks an expected table. Restore and deep verify now both use it, each inside its own private (0700) temp directory removed on every return path. pickNextDownload returns (FileID, bool) so a genuine nil-UUID file is not mistaken for "nothing left".

Model: opus-4-8
2026-09-22 12:45:53 +02:00
clawbot b4654f8e52 Abort the run when packing fails, even under --skip-errors (closes #161)
check / check (push) Successful in 1m22s
check / check (pull_request) Successful in 3m2s
A chunk is registered as pending (known, scanner-pending, packer pending-row) before it is packed. Under --skip-errors the scanner skipped a file on any processing error, including a failure inside addChunkToPacker (packing, database, encryption, upload). The pending chunk then stayed queued and a later blob finalize inserted it into the chunks table with no blob_chunks row, so a snapshot could complete holding a file whose chunk is in no blob and cannot be restored.

Errors from addChunkToPacker are now marked and abort the run regardless of --skip-errors; only open and read errors are skipped. The bookkeeping order is unchanged. Flag help and comments now say only unreadable files are skipped.

Model: opus-4-8
2026-09-22 12:28:44 +02:00
clawbot 39aef1c47c Stop config set echoing secrets; reject credential-bearing storage URLs (closes #166)
check / check (push) Successful in 1m21s
check / check (pull_request) Successful in 1m18s
config set now prints only the key name after a write, never the value: a value may be a secret such as s3.secret_access_key, and echoing it leaks into captured stdout and pasted terminals. The set logic moves into writeConfigSet so this is testable.

config set also tightens a pre-existing group- or world-readable config to 0600 after writing; the previous stat-and-preserve-mode block had no effect (os.WriteFile does not change an existing file mode) and is removed.

ParseStorageURL now rejects s3:// and rclone:// URLs that carry credentials in the userinfo or an unknown query parameter, naming s3.access_key_id and s3.secret_access_key as where credentials belong. On a url.Parse failure only the inner cause is wrapped, so the raw URL is not echoed. file:// is unchanged.

Model: opus-4-8
2026-09-22 12:28:32 +02:00
clawbot 96ebcd40d7 Reconcile purge against remote by hashed key, not human ID (closes #160)
check / check (pull_request) Successful in 1m20s
check / check (push) Successful in 2m51s
syncWithRemote compared human snapshot IDs against the hashed metadata/<key>/ directory names, which never match, so it deleted every local snapshot record; the purge that followed then found nothing to remove remotely. Reconcile via listAllRemoteSnapshotKeys and RemoteSnapshotKey(id), matching CleanupLocalSnapshots, so a row still backed by remote metadata is kept.

The purge tests only passed because their stubs used the human-ID layout production never writes; they now write metadata under the hashed remote key. New tests prove remotely-backed local rows survive the reconcile and that a purge removes the local row and remote metadata together.

Model: opus-4-8
2026-09-22 12:11:49 +02:00
clawbot d9f0220f94 Restore files at 0600 and make the blob hash check unskippable (closes #163)
check / check (pull_request) Successful in 1m21s
check / check (push) Successful in 2m42s
Regular files are now created with O_EXCL at mode 0600 and given their stored mode only after the content is written and closed, so a file whose stored mode is restrictive is never briefly readable by other local users mid-restore. A file whose write or close fails is removed rather than left partial, and a chmod failure is a user-visible warning instead of a debug line.

hashVerifyReader.Close now errors when closed before EOF, so a short read or early close can never obtain a blob whose hash was not verified; downloadBlobToCache drops the cache entry on any such failure.

verifyFile (--verify) now rejects a restored file with bytes past its last chunk. Tests cover each behaviour under umask 022.

Model: opus-4-8
2026-09-22 11:45:52 +02:00
clawbot 4c83e82543 Reject a blob_size_limit below the largest possible chunk (closes #167)
check / check (push) Successful in 1m20s
check / check (pull_request) Successful in 1m16s
Validate only rejected blob_size_limit below chunk_size, but the chunker can emit chunks up to chunk_size times the FastCDC size spread (four times), and the packer puts a single chunk of any size into an otherwise empty blob. A limit between one and four times chunk_size therefore let a blob reach four times the configured maximum, with most blobs holding a single chunk and so exposing individual chunk lengths to anyone who can list the destination.

Validate now rejects blob_size_limit below chunk_size times the spread, reusing the chunker's one constant (now exported as ChunkSizeSpread) instead of a second literal. The rule is stated in the error text, the Validate comment, the README config table, config.example.yml, and the generated config template.

Model: opus-4-8
2026-09-22 11:45:41 +02:00
clawbot 86361c8b50 Fail closed on unreadable manifests instead of losing blobs (closes #157)
check / check (pull_request) Successful in 1m20s
check / check (push) Successful in 2m42s
Prune learned which blobs are in use by reading every snapshot's manifest, but merely logged and skipped one it could not download or decode. Blobs referenced only by that snapshot then looked unreferenced and were deleted, with a zero exit -- and snapshot create --prune runs this unattended. collectReferencedBlobs now errors, naming the remote key, so prune deletes nothing and exits non-zero.

Manifest generation likewise skipped a blob whose lookup failed or was missing, yielding a manifest short of what the snapshot needs; it now fails. Deep verify only warned when the manifest omitted a database blob; it now fails on any divergence. Docs corrected.

Model: opus-4-8
2026-09-22 11:45:30 +02:00
clawbot d77663d039 Default a scheme-less s3.* endpoint to TLS (closes #158)
check / check (pull_request) Successful in 1m19s
check / check (push) Successful in 3m15s
With the s3.* config form and an endpoint written without a scheme, use_ssl being omitted built an http:// endpoint, while config.example.yml documented use_ssl as defaulting to true. Over plain HTTP a network observer sees manifests, object names, sizes and the access key id, and can alter responses.

use_ssl is now *bool: omitted (nil) means the default, TLS; only an explicit use_ssl: false forces plain HTTP. This matches the s3:// URL form, which already defaults to TLS. The config init template dropped its misleading use_ssl line from the s3:// block (that key is never read for URLs; ?ssl=false controls TLS there) and points at ?ssl=false instead.

Model: opus-4-8
2026-09-22 11:11:34 +02:00
clawbot 3abe9cbd9e Scope the PID lock to mutating commands (closes #150)
check / check (push) Successful in 1m19s
check / check (pull_request) Successful in 1m16s
RunWithApp took the process-wide PID lock for every fx-backed command, so read-only commands (info, snapshot list, snapshot verify, remote info) failed with "already running" while a backup held it.

AppOptions now carries a lockMode declared at each call site: only mutating commands (snapshot create, snapshot purge, snapshot remove, prune, remote nuke) acquire the lock; read-only ones run without it. snapshot restore is classified read-only -- it writes only to its target directory, not the local index or remote store. The decision moves to a small acquireLockIfMutating helper, with a test that a read-only command runs while the lock is held and two mutators still exclude. The README locking section is rewritten to match.

Model: opus-4-8
2026-09-22 11:01:29 +02:00
clawbot 76a6917a35 Keep restore writes inside the target directory (closes #154)
check / check (pull_request) Successful in 1m21s
check / check (push) Successful in 2m51s
restoreFile and verifyRestoredFiles joined the stored path onto the target with no containment check, so a ".." segment or an absolute path escaped the target, and a restored symlink could redirect a later child write anywhere on disk. Since age decryption proves a snapshot is readable but not honest, and restore usually runs as root, a forged snapshot became an arbitrary file write.

Both call sites now go through containedRestorePath: it rejects a stored path unless filepath.IsLocal accepts it with the leading separator removed (barring "..", absolute, and empty paths), then Lstats each existing ancestor below the target and refuses to descend through a symlink. The target directory itself may be a symlink, and honest symlinks pointing outside the tree are still written verbatim.

Model: opus-4-8
2026-09-22 11:01:01 +02:00
clawbot 38ebfd843a Trust only uploaded blobs for deduplication (closes #148)
check / check (pull_request) Successful in 1m23s
check / check (push) Successful in 3m10s
An interrupted blob upload left the blob's chunks, blob_chunks, and blobs rows committed before the upload was attempted, so a later run deduplicated against data that never reached storage and produced a snapshot that reported success but could not be restored.

Fix (issue option b): a chunk counts as known only when a blob holding it has uploaded_ts set, and each run drops un-uploaded blob rows and the chunks they orphan at startup, so the affected data is re-chunked and re-uploaded. A blob recorded with no remote backend is marked uploaded so the invariant holds uniformly.

The reproduction is the interrupted-upload test from #72: its t.Skip is removed and it passes against this fix, and this branch's earlier duplicate copy is dropped. The interrupted metadata-export case is split to #177.

Model: opus-4-8
2026-09-22 10:29:21 +02:00
clawbot 6b7517a4dc Add fault-injection tests for interruption and corruption (closes #72)
check / check (push) Successful in 1m22s
check / check (pull_request) Successful in 2m42s
Adds internal/storage/faultstore, a storage.Storer wrapper that injects faults through the storage seam without patching production code: an upload that dies mid-stream, a backend reporting success while storing nothing, and reads returning corrupt or truncated bytes. Covers all six scenarios from the issue, each asserting the observable end state (index, destination, and what the user is told), not merely that an error returned. Scenario 1b (retry after an interrupted upload) exposed a real dedup defect and is skipped with a pointer to #148, which also owns the half-exported-state repair. Tests run serially because each calls log.Initialize on the global logger. No production behavior changes.

Model: opus-4-8
2026-09-22 09:46:46 +02:00
clawbot 994e5de613 Quiet only the stdout UI under --json, not the log level (closes #112)
check / check (push) Successful in 2m23s
check / check (pull_request) Successful in 1m24s
Per the decision on the issue (option 1), --json no longer implies Quiet. Folding --json into Quiet pinned the stderr log level to WARN, so prune --json gave a machine consumer no record of the local index rows it deleted. The two effects are now split: a JSON field on log.Options drives only the stdout UI-quiet in setupGlobals, keeping the JSON document clean, while the stderr log level follows --verbose/--debug again (diagnostics have gone to stderr since #82). The same coupling is removed for snapshot verify, snapshot remove and remote info; snapshot list was already decoupled.

Model: opus-4-8
2026-09-22 09:05:52 +02:00
clawbot 343129f891 Accept a remote key for restore and verify, and document it (closes #124)
check / check (push) Failing after 1s
check / check (pull_request) Failing after 1s
A machine restoring after the original is gone has no local index and cannot know a snapshot's human ID; snapshot list shows such snapshots only by their remote key, but restore and verify accepted only the human ID, so recovery could not be done as documented.

Restore and verify now also accept a remote key, or an unambiguous leading part of it as snapshot list prints it, resolved against the store's metadata listing. Human IDs are never pure hex, which tells the two forms apart. Deep verify reads the single snapshot in the downloaded per-snapshot database. A new README section walks the recovery end to end; a test backs up, then lists, restores and deep-verifies with an empty index, another hostname and no age_recipients.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)
2026-09-22 01:01:25 +02:00
clawbot a50e3fa038 Add tests for internal/storage: URL parsing, backends, shared conformance suite (closes #66)
check / check (push) Failing after 1s
check / check (pull_request) Successful in 3m25s
internal/storage, the package that parses store URLs and selects the backend, had no tests.

Adds table-driven tests for URL parsing (each scheme, query parameters, malformed input, unknown scheme, backend type chosen); one shared conformance suite for the Storer interface, run against the file backend in a temp directory and the s3 backend on the in-process harness internal/s3 already uses, so a new backend inherits it; and rclone construction and argument tests using its in-process local backend. A comment records that rclone data operations need a configured remote and are not unit-tested. No production code changed and no defect surfaced.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)
2026-09-22 00:58:27 +02:00