Commit Graph
77 Commits
Author SHA1 Message Date
sneak 1adb856599 Report unknown blob figures for a snapshot whose manifest cannot be read (closes #272)
check / check (push) Waiting to run
When remote info could not read a snapshot's manifest, the orphan
figures were unknown but the snapshot's row still gave 0 blobs and 0 B,
in the table and in --json. The row's blob count and blob size are now
unknown, and null in --json.

A directory with no manifest, as an interrupted backup leaves, still
shows 0: the orphan figures count its blobs as orphaned, so it
references none.

The constant holding the "unknown" text is renamed from countUnknown to
unknownText, since it now also stands for a size.

Model: opus-5-5
2026-10-08 05:46:45 +00:00
clawbot c06c4d201b Keep the progress line of a multi-path snapshot within 100% (closes #271)
check / check (push) Waiting to run
One progress reporter spans every path of a snapshot. Its counts of
files and bytes processed run across all the paths, while the totals
they were divided by were reset to each path's own, so a second path
smaller than the first showed more than 100% and an ETA of `unknown`.
The totals now add up over the paths scanned so far. The rate behind
the ETA restarts when each path's scan phase ends, so that phase does
not lower the rate while the path is processed; until then the rate is
the previous path's. SetTotalSize is renamed AddTotalSize because it
now adds.

The issue says the ETA went negative; it was computed negative and
printed as `unknown`.

Model: opus-5-5
2026-10-08 07:29:11 +02:00
clawbot 70f008a21a Document that --json skips the confirmation prompt (closes #268)
check / check (push) Waiting to run
snapshot remove and prune delete without asking under --json, because a
prompt on stdout would break the JSON document. Their --json help and
README entries named only --force as skipping the prompt. Both now say
--json skips it too. One test checks the --json help of both commands;
another runs prune with --json, no --force and empty stdin, and checks
that the unreferenced blob is deleted and stdout holds only the JSON
document.

Model: opus-5-5
2026-10-08 04:59:30 +02:00
clawbot fce253fa39 Give a second snapshot create in the same second its own ID (closes #270)
check / check (push) Waiting to run
The timestamp in a snapshot ID is in whole seconds, so a `snapshot
create` that started in the same second as the previous run of that
snapshot name got the same ID, and inserting its row failed with
`UNIQUE constraint failed: snapshots.id`. CreateSnapshotWithName now
looks the ID up in the local index first and, if it is taken, waits a
second and takes a new timestamp. The ID format is unchanged.

Only the local index is checked. That is where the insert fails, and
the process lock serializes runs, so nothing takes the ID between the
lookup and the insert.

Model: opus-5-5
2026-10-08 03:29:07 +02:00
clawbot e459a66099 Stop a backup on a symlink whose target cannot be read (closes #269)
check / check (push) Waiting to run
When readlink failed, the scanner logged at debug level and left the
symlink out of the snapshot, and the run reported success even without
--skip-errors. The error now goes through the same handling as any
other entry the walk cannot read: the run aborts, or with --skip-errors
the symlink is skipped with the usual "Failed to access" error line.

A symlink removed between the walk's lstat and the readlink also
aborts the run, as an entry that vanishes during the walk already does.

Model: opus-5-5
2026-10-08 00:12:13 +02:00
clawbot 32c4a46b49 Write rclone uploads under a temporary name and move them into place (closes #266)
check / check (push) Waiting to run
The rclone backend wrote each object straight to its key, so killing an
upload to a local or sftp remote left a truncated object there that the
next backup trusted.

On every remote with a server-side move, an object is now written under
a name ending in `.partial` and moved onto its key with rclone's
operations.Move, which removes an object already at the key first;
drive, dropbox and others will not move onto one. Listings skip
`.partial` names. Rclone's own copy also requires the PartialUploads
flag; this does not, because hdfs shows a file while it is written
without setting it. Remotes without a move are written in place.

Model: opus-5-5
2026-10-07 22:59:26 +02:00
clawbot b5389a62b5 Exit 130 and say so when a command is interrupted (closes #267)
check / check (push) Waiting to run
Ctrl-C or SIGTERM during snapshot create, restore or verify exited 0
with no error line (1, also silent, under snapshot verify --json), so
an unfinished --cron backup looked like a success. RunOperation now
records whether op returned while the Vaultik context was still live;
a run where it had not by the time RunWithApp returned is interrupted,
whatever op returned. It used to look for context.Canceled in op's
error, which verify --json does not return. Entry prints "interrupted
before the command finished" on stderr for it and returns 130, under
--cron and --json too.

SIGTERM also gives 130, as the issue asks, not 143.
The test does not cover an op still running when the 30s shutdown
timeout ends.

Model: opus-5-5
2026-10-07 21:59:27 +02:00
clawbot d87202fb70 Correct doc and help sentences that are false about the code (closes #233)
check / check (push) Waiting to run
A blob is written whole to a temporary file and uploaded once finished,
not streamed to storage. The README, ARCHITECTURE.md and
config.example.yml now say a backup needs free temporary space for each
blob (twice that for an rclone destination that cannot stream uploads)
and for the metadata export's copies of the local index, in $TMPDIR, or
partly in /var/tmp when TMPDIR is unset. Also corrected: the snapshot ID
format, what restore reads and how incomplete snapshots are removed in
docs/DATAMODEL.md, what source_path holds, the index_path default, the
config search order in the snapshot create help, what snapshot remove
cleans up in the prune help, how the release installs Go, and the
script/release and script/fmt-check comments.

Model: opus-5-5
2026-10-07 17:59:27 +02:00
clawbot e161343eac Speed up the vaultik and database package tests (closes #235)
check / check (push) Waiting to run
In the Dockerfile test phase, internal/vaultik spent 23.6s of 31.9s in
36 tests run one at a time. 24 of them were serial only because they
call log.Initialize; they now call it before t.Parallel(), so the
logger is replaced before any parallel test runs. The 12 still serial
change the umask, TMPDIR, os.Stderr or time.Local.

TestLargeDatasets took 8.4s of the 9.8s internal/database run by
committing each of its 1,500 inserts on its own; it now makes them in
one transaction. TestDedupOnlySnapshotRestores gives its second backup
its own snapshot name instead of sleeping 1.1s for a new snapshot ID.

Model: opus-5-5
2026-10-07 17:12:07 +02:00
clawbot d53202eb86 Fix two misleading messages (closes #240)
check / check (push) Waiting to run
The warning for a config file that others can read always said the
file contained S3 credentials, so a file:// config with none got a
false claim. When s3.access_key_id or s3.secret_access_key is set it
now says the file may contain them, because Load sees the values only
after smartconfig has replaced any ${...} reference, so a set
credential need not be in the file. Otherwise it says the file is
readable by others.

snapshot purge wrapped the listing error, which already starts with
"listing remote snapshots:", in that prefix a second time.
syncWithRemote now returns it unwrapped, as CleanupLocalSnapshots
does.

Model: opus-5-5
2026-10-07 15:29:07 +02:00
clawbot 8b22ae8d42 Pass s3.part_size to the multipart uploader (closes #232)
check / check (push) Canceled after 0s
s3.part_size was loaded and defaulted but never reached the S3 client,
whose uploader used a fixed 10MiB part. The client now takes the part
size from the config, for storage_url and for the s3.* fields, and
config load rejects a value below 5MiB or above 5GiB, an explicit 0
included. A blob too large for S3's limit of 10,000 parts at that size
is uploaded in larger parts, since the uploader cannot learn the size of
the reader it is given. The docs gave the default as 5MB, which the
config file reads as 5,000,000 bytes, below the minimum; they now say
5MiB.

Judgement call: the 5GiB maximum is enforced along with the 5MiB
minimum the issue names.

Model: opus-5-5
2026-10-07 14:29:10 +02:00
clawbot 3fc8a8f2f4 Read the snapshot name using the stored hostname (closes #230)
check / check (push) Canceled after 0s
A snapshot ID is hostname_name_timestamp, and purge took the name to be
everything between the first and the last underscore. With a hostname
such as my_host the name home came out as host_home, so
`snapshot purge --keep-latest --snapshot home` found nothing to delete
and `snapshot create --prune` purged nothing without a message. The name
is now read by removing the hostname stored with the snapshot, in the
short form the ID uses, so both may contain underscores. This was chosen
over rejecting underscores in `hostname` when the config loads, which
would also stop restores on such a host.

The purge consistency test stored a hostname that did not match its
snapshot IDs; it now matches, as it always does in production.

Model: opus-5-5
2026-10-07 12:12:08 +02:00
clawbot 7696f83258 Leave remote info orphan figures unknown when a manifest is unreadable (closes #228)
check / check (push) Canceled after 0s
When a manifest could not be read, remote info skipped it, counted that
snapshot's blobs as orphaned and advised running prune. The orphan
figures are now unknown in that case, with no prune advice; --json gives
them as null and lists the unreadable remote keys in
unreadable_manifests. Only a listed manifest.json.zst is read, so a
directory without one, as an interrupted backup leaves, keeps the
figures known.

Names under metadata/ were used unchecked and printed raw. A name that
is not a remote key is now skipped with a warning. A manifest under it
is not read either, so it also leaves the figures unknown; --json counts
such manifests in skipped_manifest_count.

Model: opus-5-5
2026-10-07 10:59:26 +02:00
clawbot 5d1118d143 Quote a string setting that YAML would read as a number (closes #229)
check / check (push) Canceled after 0s
config set wrote every value as an unquoted YAML scalar, and config.Load
reads the file through untyped YAML, so an access key 00112233 loaded as
38043 and a hostname 007 as 7.

config set now looks the key up in config.Config by the fields' yaml
tags. A string setting is tagged !!str, which the encoder quotes
wherever YAML would read a number or a boolean. Other settings stay
unquoted, so compression_level 9 is still a number. A value that is not
valid UTF-8 stays untagged and is written as !!binary, which loads back
unchanged.

Judgement call: the type comes from reflection over config.Config.

Model: opus-5-5
2026-10-07 09:29:19 +02:00
clawbot b57ce2277d Store and compare file mtimes to the nanosecond (closes #226)
check / check (push) Canceled after 0s
The files table held mtime in whole seconds and the scanner compared
whole seconds. A file rewritten with its size unchanged and a new mtime
in the same second as the indexed one was treated as unchanged, and
every later snapshot restored the old content. A new mtime_nsec column
now holds the nanoseconds within the second that mtime holds, and the
scanner compares the full mtime.

A local index created before this change lacks the column and is
rebuilt with `vaultik database delete` and a full backup. A snapshot
made before it cannot be restored by this version.

Model: opus-5-5
2026-10-07 07:12:12 +02:00
clawbot 8496404d8b Make the process-wide lock atomic with flock (closes #227)
check / check (push) Canceled after 0s
Acquire read vaultik.pid, checked whether that PID was alive, then
wrote its own, so two writers started together could both pass the
check and both run. The lock is now an flock on vaultik.pid, held
while the file stays open; the kernel drops it when the process exits,
so the stale-PID check is gone.

Release empties the file instead of deleting it. Deleting it would let
a process that opened the old file a moment earlier lock it while
another creates and locks a new one.

The new concurrent test fails against the old code only when the race
is hit, not on every run; against the fix it cannot admit two callers.

Model: opus-5-5
2026-10-07 06:12:10 +02:00
clawbot cdc60c4dfa Write only the document to stdout from snapshot remove --json (closes #251)
check / check (push) Canceled after 0s
When the destination store cannot be reached, `snapshot remove` still
removes the snapshot from the local index and warns. The warning went
through the UI, which writes to stdout, so under `--json` it landed
ahead of the document and `| jq` failed on a command that exited 0.
Under `--json` the UI warning is now skipped; the logger's warning on
stderr carries the follow-up, with the snapshot ID as a field.

The follow-up, in the warning, the README and the command's help, said
`vaultik prune` would finish the cleanup, but `prune` never removes
snapshot metadata. They now say to run `vaultik snapshot remove` for the
snapshot again once the destination store is reachable.

Model: opus-5-5
2026-10-07 04:46:13 +02:00
clawbot 49eed7a3e5 Count each file, byte and upload once in backup statistics (closes #225)
check / check (push) Canceled after 0s
The scanner added a changed file's bytes again for each new chunk and
counted a file as unchanged for each chunk already stored. It now
counts files and bytes once, in the scan phase, and counts its own
uploads, so a --cron run, which has no progress reporter, records
them. The blob count no longer adds earlier paths' blobs again.

The snapshots row now stores the size of all files in total_size and
the referenced blobs' sizes in blob_size, blob_uncompressed_size and
compression_ratio, as docs/DATAMODEL.md says. Those sizes come from
one query, and a failed query fails the snapshot. DATAMODEL.md now
says chunk_count and blob_count count what the run added.

Removed UpdateSnapshotStats and GetCountBySnapshot, which nothing
calls any more.

Model: opus-5-5
2026-10-07 02:59:27 +02:00
clawbot 85d4ef118d Keep command output to the README's stdout and stderr rules (closes #224)
check / check (push) Successful in 16m49s
The startup banner moves from stdout to stderr, so a `completion`
script, a `config get` value and the hidden `__complete` command print
only their own output. `--quiet`, `--cron` and `--json` still suppress
it.

A failing `remote info`, `prune` or `snapshot remove` under `--json`
now reports its error on stderr. Their reporters returned early under
`--json`, so the failure reached neither stream.

`snapshot verify --quiet` writes no report. A failure is still
returned and printed on stderr, with the same exit status.

Judgement call: the banner's stream, posted on the issue for the owner.

Model: opus-5-5
2026-10-07 01:12:15 +02:00
clawbot b06f992152 Start the progress reporter once per snapshot, not once per path (closes #253)
check / check (push) Successful in 19m54s
A backup without --cron of a snapshot with two or more paths panicked
with "close of closed channel". Scan runs once per path, and it started
the progress reporter and deferred its Stop each time; Stop closes the
reporter's signal channel, so the second path's Stop panicked.
scanAllDirectories now starts the reporter before the first path and
stops it after the last, and Scan no longer starts or stops it. A new
test backs up a two-path snapshot with the reporter on and restores
both paths.

Model: opus-5-5
2026-10-07 00:12:10 +02:00
clawbot 5d685f03ce List the files under a path by path, not by string prefix (closes #223)
check / check (push) Successful in 12m16s
FileRepository.ListByPrefix matched with SQL LIKE: a plain string
prefix that ignores ASCII case and treats _ and % as wildcards.
Restoring /home/u/doc also restored doc2, DOC and doc.txt.bak, and a
backup counted the files of a longer sibling path as deleted. It is
now ListUnderPath, which returns the file at the path and every file
whose path starts with the path plus a slash, compared exactly. A
trailing slash is ignored, so "/" still lists every file.

Three tests relied on string-prefix matching and now name a full path
or call ListAll.

ListIDsWithChunksNotInUploadedBlobs keeps its LIKE: it only adds file
IDs the scan never looks up.

Model: opus-5-5
2026-10-06 22:12:16 +02:00
clawbot f59086c0e5 Return an error, not a panic, on a malformed snapshot database (closes #231)
check / check (push) Successful in 13m39s
Restore cut chunk hashes from the snapshot database to 16 characters
for its error messages, so a shorter hash panicked. Those messages now
use shortHash. Under --verify, a file_chunks row with no chunks row was
dereferenced, and the chunk size from the database was allocated in
one piece, so a negative or huge size panicked. A missing row is now an
error, a negative size is rejected, and each chunk is hashed by
streaming it from the restored file.

A restored file shorter than its chunks now fails verify as a short
read instead of an unexpected EOF.

Model: opus-5-5
2026-10-06 21:16:28 +02:00
clawbot d276d891da Join the S3 prefix to every key with one slash (closes #222)
check / check (push) Successful in 10m59s
The S3 client built each key as prefix + key, and the URL parser keeps
the prefix as written, so s3://bucket/p stored p + "blobs/..." with no
slash while s3://bucket/p/ stored p/blobs/.... A recovery host that wrote
the URL the other way found no snapshots.

NewClient now strips trailing slashes from the prefix and adds one back
when anything is left, giving the README layout for both URL forms; an
empty prefix stays at the bucket root. The s3.prefix config setting
goes through the same client and gets the same join.

A new test writes through each URL shape against an in-process S3
server, checks the key in the bucket, and lists through both List and
ListStream, which every snapshot listing uses.

Model: opus-5-5
2026-10-06 19:46:09 +02:00
clawbot 5aa5ba5544 Apply restored owners, modes and times in an order that keeps them (closes #219)
check / check (push) Successful in 11m53s
A directory got its stored mode and mtime before its contents were
written, so a read-only directory came back without its files and a
non-empty one carried the time of the restore. Directories are now
created owner-only (0700) and get their stored owner, mode and mtime
after the restore loop, each before its parent, skipping any whose
place a symlink has since taken. A file's mode is now applied after its
chown, which on Linux clears setuid and setgid. A symlink gets its
stored owner (as root) and mtime on the link itself, through
golang.org/x/sys/unix, now a direct dependency.

An interrupted restore leaves its directories at 0700.

Model: opus-5-5
2026-10-06 18:29:09 +02:00
clawbot 315b6483b8 Mark github.com/spf13/pflag as a direct dependency in go.mod (closes #246)
check / check (push) Successful in 12m34s
internal/cli/snapshot_restore_test.go imports github.com/spf13/pflag
directly, but go.mod still marked it // indirect. script/precommit runs
go mod tidy and fails when that changes go.mod, so the pre-commit hook
stopped every commit. This is the go mod tidy output: pflag moves to the
direct require block, and go.sum does not change. script/cibuild does
not run the tidy, which is why the gate stayed green.

Model: opus-5-5
2026-10-06 16:59:26 +02:00
clawbot 14fc4c9893 Load a config that has no age recipient (closes #221)
check / check (push) Successful in 13m47s
The README's steps for restoring on another machine failed at the first
command: `config init` wrote a placeholder recipient, and `config.Load`
rejects any recipient that does not parse. `config init` now writes an
empty `age_recipients` list, `config.Load` accepts an empty list, and
`snapshot create` refuses to start without a recipient. A malformed
recipient is still rejected at load.

The recovery-host test now builds its config with `config init` and
`config set` and reads it through `config.Load`, so it imports
`internal/cli`. On a fresh file, `config set age_recipients.0` writes the
list in flow style (`[age1...]`).

Model: opus-5-5
2026-10-06 15:46:13 +02:00
clawbot 66c80a70e3 Skip files of an unreadable blob under restore --skip-errors (closes #218)
check / check (push) Successful in 12m34s
A blob that failed to download ended `snapshot restore` even with
--skip-errors, after restoring whichever files came first. The download
error now goes through the same per-file handling as any other restore
error, once for every pending file that references the blob. With
--skip-errors those files are reported as failed, the rest are
restored, and the command still exits non-zero. Without the flag the
restore still aborts; the error now also names one affected file and
suggests --skip-errors. A cancelled restore still ends at once. The
--skip-errors help text and README now limit the packing and storage
caveat to snapshot creation.

Model: opus-5-5
2026-10-06 14:46:05 +02:00
clawbot 81f83b29f6 Re-vendor the canonical files from sneak/prompts at dd4027b (closes #213)
check / check (push) Successful in 12m47s
Linting and testing become the lint and test phases of the Dockerfile,
and the build stage depends on both. Dockerfile.lint, CHECK_EPOCH and
the tests that checked them are removed. Every docker build in script/
passes --no-cache, and script/cibuild runs script/bootstrap first. A
host without Go gets the go.mod version from script/install-go in
.tool/go, which bootstrap, the Makefile, fmt, fmt-check, precommit and
release add to PATH; fmt-check skips .tool. The image takes its version
from the VERSION build arg or git describe, dev without .git. This
repo's own entries follow the canonical content in .gitignore and
.editorconfig. The golangci-lint v2.14.0 findings are fixed. The rules
in CLAUDE.md move into AGENTS.md. IsDevVersion counts "unknown".

Model: opus-5-5
2026-10-06 12:46:13 +02:00
clawbot c4adb72d80 Run the local index in WAL mode with a busy timeout (closes #217)
check / check (push) Successful in 5m10s
check / check (pull_request) Successful in 5m38s
The connection settings were passed as `_journal_mode=`-style
parameters, which the SQLite driver drops without an error, so the
index ran in rollback-journal mode with no busy timeout. `snapshot
list` or `info` reading during a backup could make the backup's next
write fail with "database is locked". Both open paths now pass
`_pragma=` parameters; foreign keys moved there too.

With WAL on, rows committed to the open index can still be in the
-wal file, which a copy of the main file misses. The metadata export
now copies the index with VACUUM INTO, into an empty 0600 file.

The retry after a failed open no longer claims a TRUNCATE recovery; it
retries with the same settings.

Model: opus-5-5
2026-10-06 11:29:18 +02:00
clawbot 4a167e153a Record the real uid and gid of backed-up files (closes #216)
check / check (push) Successful in 4m58s
check / check (pull_request) Successful in 6m21s
The scanner read uid and gid by asserting the stat result to an
interface with Uid() and Gid() methods. *syscall.Stat_t has Uid and Gid
fields, not methods, so the assertion never matched and every file,
directory and symlink was stored as 0:0; a restore as root then gave
everything to root. The scanner now reads the fields of
*syscall.Stat_t.

The first backup after this change re-reads every file not owned by
root, because its stored uid and gid no longer match the disk.

When the tests run as root, as in the Docker build, the new test
compares 0 with 0 and cannot catch the defect; a non-root run does.

Model: opus-5-5
2026-10-06 09:46:16 +02:00
clawbot ea72697992 List a missing file:// destination directory as an error (closes #220)
check / check (push) Successful in 5m53s
check / check (pull_request) Successful in 4m53s
The file backend listed a destination directory that does not exist as
an empty store. With the volume unplugged, snapshot list reported every
local snapshot as missing from the store, snapshot remove said it had
removed metadata it never reached, and prune dropped every local
snapshot record. List and ListStream now fail when the destination
directory is missing, so those commands take their existing path for a
store that cannot be listed. A missing prefix under an existing
directory is still an empty listing, and a first backup still creates
the directory.

Three tests listed a file:// destination nothing had created; they now
create it.

Model: opus-5-5
2026-10-06 08:46:17 +02:00
clawbot 713be502bd Reject a duration with characters outside its parts (closes #215)
check / check (push) Successful in 7m28s
check / check (pull_request) Successful in 5m39s
parseDuration fell back to an unanchored search for number-and-unit
pieces when time.ParseDuration failed, and skipped everything in
between. 1.0y became 0, so `snapshot create --prune --keep-newer-than
1.0y` deleted every snapshot of the backed-up names, the new one
included. 2.1w became one week and 1,5y five years. The fallback now
requires the whole input to be whole-number-and-unit parts with nothing
between them. A bare number is rejected before time.ParseDuration sees
it, since Go reads 0 and +0 as zero with no unit.

Judgement call: a space between number and unit (`30 days`) was
accepted and is now an error, matching Go's own units.

Model: opus-5-5
2026-10-06 06:12:07 +02:00
clawbot 35cf985c18 Re-chunk a known file whose chunks no uploaded blob holds (closes #214)
check / check (push) Successful in 6m53s
check / check (pull_request) Successful in 6m17s
File rows are shared by every snapshot and updated in place, while a
blob row is deleted once no snapshot references it. Removing the newest
snapshot, or the prune after an interrupted run, could drop the only
blob holding a changed file's current chunks while an older snapshot
kept the file row. The next backup compared metadata only, skipped the
file, and completed a snapshot that could not restore it.

The scanner now loads the IDs of known files that list a chunk no
uploaded blob holds and re-chunks them even when their metadata is
unchanged.

The tests append to a file, so the file keeps its first chunk in a blob
the first snapshot still references. Each backup run gets its own
snapshot name, so the second-precision snapshot IDs differ without
sleeping.

Model: opus-5-5
2026-10-06 04:46:16 +02:00
clawbot 584444b619 List only after-1.0 work in the README roadmap (closes #208)
check / check (push) Successful in 3m35s
check / check (pull_request) Successful in 3m36s
The README roadmap and the TODO.md Next Step still described finished
1.0 work as remaining. The roadmap now lists only work planned after
1.0. Its security item says the code was reviewed before 1.0, every bug
found was fixed, and the accepted risks are listed; an outside audit
stays as after-1.0 work. The error-condition item is gone because every
failure case it listed has a fault-injection test. Daemon mode is added.
TODO.md says the 1.0 work is complete on next and that merging and
tagging are the owner's.

Judgement call: dropped the human-readable size flags item; no command
flag takes a raw-integer size.

Model: opus-5-5
2026-10-01 21:41:41 +02:00
clawbot eed117fe25 Route direct-stdout command output through internal/ui (closes #149)
check / check (pull_request) Successful in 1m27s
check / check (push) Successful in 3m18s
version, info, remote info, config, and database delete wrote plain text straight to stdout, so they were unstyled and ignored --quiet. Output is now governed by internal/ui in two buckets. Status lines and confirmations (config init, config set, the database delete prompt) go through the ui message methods and are silenced by --quiet. The data a command exists to produce is written plain -- the version/info/remote-info reports, the snapshot list table, config get values, and the --json documents -- and is never suppressed, since a script depends on it and a marker would corrupt a table or document. The database delete confirmation prompt is always shown. Pure-cli commands reach ui through a small commandUI helper.

Model: opus-4-8
2026-09-22 20:28:50 +02:00
clawbot 82c51a5337 Validate blob hashes, offsets and lengths from the destination (closes #155)
check / check (pull_request) Successful in 1m22s
check / check (push) Successful in 3m30s
A blob hash read back from the downloaded snapshot database or the store listing was trusted unchecked. A hostile remote could set a hash such as "aa/../../etc" and have a decrypted blob written outside the cache directory, or feed a short or negative value that panicked a command.

blobDiskCache.path now refuses any key with a path separator, and ReadAt rejects a negative offset or length, bounding so a sum cannot overflow past the check. A new isBlobHash helper gates FetchBlob, shallow and deep verify, and restore: buildBlobIndexes rejects every hash from the snapshot database before any fetch. The blobs/ and metadata/ listings skip a non-conforming name, and short-hash prefixes in log and error text go through a panic-safe shortHash helper.

Model: opus-4-8
2026-09-22 16:11:28 +02:00
clawbot 76a6917a35 Keep restore writes inside the target directory (closes #154)
check / check (pull_request) Successful in 1m21s
check / check (push) Successful in 2m51s
restoreFile and verifyRestoredFiles joined the stored path onto the target with no containment check, so a ".." segment or an absolute path escaped the target, and a restored symlink could redirect a later child write anywhere on disk. Since age decryption proves a snapshot is readable but not honest, and restore usually runs as root, a forged snapshot became an arbitrary file write.

Both call sites now go through containedRestorePath: it rejects a stored path unless filepath.IsLocal accepts it with the leading separator removed (barring "..", absolute, and empty paths), then Lstats each existing ancestor below the target and refuses to descend through a symlink. The target directory itself may be a symlink, and honest symlinks pointing outside the tree are still written verbatim.

Model: opus-4-8
2026-09-22 11:01:01 +02:00
clawbot 38ebfd843a Trust only uploaded blobs for deduplication (closes #148)
check / check (pull_request) Successful in 1m23s
check / check (push) Successful in 3m10s
An interrupted blob upload left the blob's chunks, blob_chunks, and blobs rows committed before the upload was attempted, so a later run deduplicated against data that never reached storage and produced a snapshot that reported success but could not be restored.

Fix (issue option b): a chunk counts as known only when a blob holding it has uploaded_ts set, and each run drops un-uploaded blob rows and the chunks they orphan at startup, so the affected data is re-chunked and re-uploaded. A blob recorded with no remote backend is marked uploaded so the invariant holds uniformly.

The reproduction is the interrupted-upload test from #72: its t.Skip is removed and it passes against this fix, and this branch's earlier duplicate copy is dropped. The interrupted metadata-export case is split to #177.

Model: opus-4-8
2026-09-22 10:29:21 +02:00
clawbot 994e5de613 Quiet only the stdout UI under --json, not the log level (closes #112)
check / check (push) Successful in 2m23s
check / check (pull_request) Successful in 1m24s
Per the decision on the issue (option 1), --json no longer implies Quiet. Folding --json into Quiet pinned the stderr log level to WARN, so prune --json gave a machine consumer no record of the local index rows it deleted. The two effects are now split: a JSON field on log.Options drives only the stdout UI-quiet in setupGlobals, keeping the JSON document clean, while the stderr log level follows --verbose/--debug again (diagnostics have gone to stderr since #82). The same coupling is removed for snapshot verify, snapshot remove and remote info; snapshot list was already decoupled.

Model: opus-4-8
2026-09-22 09:05:52 +02:00
clawbot c355ef4d25 Report a prune count that could not be read as unknown, not 0 (closes #96)
check / check (pull_request) Failing after 1s
check / check (push) Successful in 2m46s
Prune read table row counts before and after to report how many orphaned files, chunks and blobs it removed, and discarded the error from every read. A failed query therefore reported as a count of 0, and the summary showed plausible wrong numbers.

A count that cannot be read is now logged as a warning (on stderr, also under --json) and shown as "unknown"; a difference computed from an unknown count is itself unknown. 0 still means the table was empty. No --json document carries these counts, so none can show a false 0.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)
2026-09-21 21:41:57 +02:00
clawbot 9ca962969a Map s3 not-found to storage.ErrNotFound in Get and Stat (closes #129)
check / check (push) Failing after 1s
check / check (pull_request) Successful in 3m50s
The Storer interface documents that Get and Stat return storage.ErrNotFound for a missing object. The file and rclone backends did; the s3 backend returned the raw SDK error, so callers testing for ErrNotFound behaved differently on s3.

S3Storer.Get and Stat now wrap ErrNotFound when the SDK reports a missing object and leave every other error untouched. The SDK reports a missing key two ways (NoSuchKey from Get, NotFound from Head); both are recognised in one helper, s3.IsNotFound, which HeadObject now also uses. The mapping lives in the storage package because internal/s3 cannot import it.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)
2026-09-21 21:07:35 +02:00
clawbot 89ebfc78e2 Use one duration parser and fix the --older-than months example (closes #123)
check / check (push) Failing after 1s
check / check (pull_request) Failing after 1s
Two parseDuration functions existed with different grammars; only the one in internal/vaultik/helpers.go was reachable from a flag. The unused copy in internal/cli/duration.go is deleted, so no flag accepts anything it did not before.

The README gave 6m as the six-months example for snapshot purge --older-than, but m is minutes: that command removed every snapshot older than six minutes. The example is now 6mo, and the help for --older-than and --keep-newer-than states that m is minutes and mo is months.

The parser now rejects negative durations, which it used to accept or silently make positive.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)

Co-authored-by: clawbot <clawbot@noreply.example.org>
2026-09-21 20:58:30 +02:00
clawbot 07ef3a1c78 Drop the lint-guard shell scanner, keep the Dockerfile.lint checks (closes #121)
check / check (push) Failing after 0s
check / check (pull_request) Successful in 2m37s
The guard test in cmd/vaultik/lintdocker_test.go tried to prove that no script runs the linter outside the container by parsing shell scripts with a hand-written scanner. Four reviews each found another spelling it missed; such a parser cannot be complete, and nobody could follow it in one reading.

The scanner, its helpers and their tests are deleted. The plain Dockerfile.lint assertions stay: the linter image is pinned by digest, config verify runs before run, and the per-run value reaches both steps. TODO.md no longer claims a test proves the property; script/lint is the only lint entry point, and keeping it so is a review matter.

Judgement call: this drops a guard two reviewers asked to harden.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (decision, merge)
2026-09-21 20:41:33 +02:00
clawbot c423d13191 Hash the plaintext, not the encrypted bytes, in verify --deep (closes #131)
check / check (push) Failing after 0s
check / check (pull_request) Failing after 0s
The last step of verify --deep hashed the encrypted bytes it downloaded once with SHA256 and compared the result to the blob name. The name is the double SHA256 of the blob plaintext, so the two could never match and deep verification failed on every healthy blob with "blob hash mismatch".

It now hashes the decompressed plaintext as chunk verification streams it and compares the double SHA256 of that to the blob name, the same derivation the writer uses. A new test backs up a real snapshot, runs deep verify on it, then flips one byte in a stored blob and expects failure.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)
2026-09-21 20:24:37 +02:00
clawbot 75a10d3a22 Hash-verify the Go toolchain in the release workflow (closes #105)
check / check (push) Failing after 0s
check / check (pull_request) Successful in 4m13s
The release workflow installed Go with actions/setup-go, which pins the action but not the Go archive it downloads, so the compiler that builds the published binaries was verified against nothing in this repo.

New script/install-go, modelled on script/install-goreleaser, downloads the go.dev archive for the version in go.mod and refuses it unless its sha256 matches the value committed in the script. It fails if its version disagrees with go.mod, and on any OS or architecture other than the Linux release runners. GOTOOLCHAIN=local on the release step keeps the verified toolchain from switching itself.

Judgement call: release path only; script/bootstrap still uses the host Go.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)
2026-09-21 19:48:52 +02:00
clawbot 3d56dd7eb0 VACUUM snapshot metadata through the sqlite driver, not a CLI (closes #120)
check / check (push) Failing after 1s
check / check (pull_request) Successful in 2m57s
snapshot create compacted the metadata database by running a sqlite3 command-line binary, after every blob had already been uploaded. On a host without that binary, which includes anyone who installed with go install, the backup failed at the last step, and two tests failed the same way.

VACUUM now runs through the Go sqlite driver the program already uses, and its error is returned to the caller. The runtime Docker image no longer installs the sqlite package, since nothing in the binary calls it.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)
2026-09-21 19:41:28 +02:00
clawbot d2a0510cb4 Trigger CI on next, not only main (closes #122)
check / check (pull_request) Failing after 0s
check / check (push) Successful in 3m33s
`check.yml` ran only on push to `main` and on pull requests against `main`. Every unit is a PR based on `next`, and `next` is pushed on each squash-merge, so no unit PR and no push to `next` ever ran CI; a broken `next` would first surface on the milestone PR. `next` is added to both branch lists; nothing else in the workflow changes. The README Entrypoints section now says where CI runs.

Disclosure: the CI run on the PR itself fired (the proof the trigger works) but was red because the runner had no disk space left before any check step ran; the local gate was green.

Model: opus-4-8 (implementation and review)
model: claude-fable-5
2026-09-21 14:55:59 +02:00
sneak d257f8f658 Lint in a container as a build step, via Dockerfile.lint (closes #113)
check / check (pull_request) Successful in 2m58s
Every lint run now happens inside its own container, invoked through
script/lint, and linting is a build step rather than a container
command: a successful build of the new root Dockerfile.lint IS a clean
lint. That shape also works where the docker daemon is remote and bind
mounts are impossible.

Its FROM line -- golangci/golangci-lint:v2.12.2, pinned by digest -- is
now the only pin of the linter version in this repo.

A container per run has its own lint cache and its own golangci-lint
lock, both discarded with it, so neither cross-worktree contamination
nor lock contention exists any more. The machinery that defended
against them is therefore gone: the per-worktree cache directories, the
lock-retry loop, and script/lint-audit, which existed to catch findings
replayed from a cache that no longer exists. So is the host lint path
in its entirety -- the native escape hatch, its version detection, and
VAULTIK_LINT_IN_CONTAINER in both script/lint and the Dockerfile.
Nothing lints on the host, at any version.

A cached build lints nothing, so the CHECK_EPOCH mechanism the product
Dockerfile already used is what makes a green mean something:
ARG CHECK_EPOCH with no default, placed below the module layers so
dependency caching survives, a `RUN [ -n "$CHECK_EPOCH" ] || exit 1`
guard so a build that withholds the arg fails instead of replaying, and
the value expanded into the lint command itself. script/lint computes
`epoch="$(date +%s%N)$$"` as a bare assignment on its own line, because
inline in the argument a failing substitution does not abort under
`set -eu` and yields a constant empty epoch -- which is exactly the
false green being prevented.

The product Dockerfile loses its lint stage rather than gaining a
second linter pin. That stage ran `make lint`, which is now
`docker build`: docker-in-docker inside a BuildKit step with no daemon.
Calling golangci-lint directly there instead would have meant two
independently bumpable digests for one tool. `make fmt-check` moves
beside `make test` in the builder stage, and script/cibuild now builds
Dockerfile.lint and then Dockerfile, each with its own fresh epoch,
failing on either. Consequence, stated in comments rather than left to
be discovered: script/docker builds the product image only and no
longer lints; script/check and script/cibuild are the gates.

`golangci-lint config verify` runs as its own epoch-keyed layer, above
the lint. It is not belt-and-braces: `golangci-lint run` rejects a
config it cannot PARSE but silently IGNORES an unknown top-level KEY.
Renaming .golangci.yml's `linters:` to `linterz:` -- one character --
discards `default: all`, the disable list and every threshold, leaves
only the small default linter set running, and exits 0 reporting
`0 issues.` on a tree the real config fails with an lll finding, in a
run whose lint layer demonstrably executed. That is a set-but-
ineffective config falling back to defaults instead of failing loudly,
sitting in the gate's own configuration. `config verify` catches it and
does so with the network genuinely off at this pin: under
`docker run --network none` against the pinned digest it exits 0 on
this repo's config and exits 3 on the `linterz:` variant. It is keyed
on CHECK_EPOCH like the lint itself, because a cached validation
validates nothing.

script/lint-fix is kept, reimplemented as a bind-mounted
docker run against the image parsed out of Dockerfile.lint -- a build
step cannot write fixes back to the worktree -- and its header states
outright that it is a developer convenience, never a gate, and needs a
local daemon.

cmd/vaultik/lintdocker_test.go parses both Dockerfiles and both scripts
and fails if any part of the mechanism is dropped: the digest pin, the
defaultless ARG below `go mod download`, the emptiness guard, the
expansion of the epoch into each check command, the bare per-invocation
epoch assignment in both scripts, cibuild building both files, the
config verification running before the lint, and -- structurally, not
by searching for one retired variable name -- that no script invokes
golangci-lint except through docker. Every one of those losses is
silent: the build still exits 0 and nothing is checked, which is why
they are asserted rather than trusted. The scanner behind the last of
those has its own test, because a structural check that goes blind
passes on every tree, including a broken one.

script/lint takes no arguments now, and says so instead of dropping
them: a build step has no command line to pass linter flags to.
2026-08-10 13:38:45 +00:00
clawbot 696ed9ab4d Gate prune's local-cleanup output on --json, and make make build build (closes #108)
check / check (push) Successful in 2m16s
Closes #110.

CleanupLocalSnapshots wrote three prose lines to stdout with no --json
awareness, covering every branch, so `vaultik prune --json | jq` failed
on any input. -q never helped either: printlnStdout and stdoutf write
straight to v.Stdout and never consult v.UI, which is what SetQuiet
affects. It now takes *PruneOptions, symmetric with its sibling phase
PruneBlobs, and gates all three writes.

Threading opts.JSON was chosen over moving the lines to log.Info,
because internal/log/log.go defaults the level to Warn: log.Info would
not have relocated them to stderr, it would have deleted them from a
plain `vaultik prune`, and "Removing stale local record" narrates the
deletion of local index rows. The stale-record count is deliberately not
added to PruneBlobsResult - every field there is blob-scoped and produced
by the phase that runs after this reconciliation, so adding it would
change a published --json schema as a side effect of a stream fix.

Note for anyone reading the --json contract: under --json the
stale-record removal now produces no signal in either stream. stdout is
correctly gated, stderr is level-pinned to Warn because --json sets
Quiet, and the count is not in the document. That is inherited behaviour
- PruneBlobs' own log.Info calls are equally invisible under --json - not
something this change introduced, and it is tracked separately.

make build exited 0 and produced nothing: .PHONY listed build with no
build: rule, and a phony target with no prerequisites and no recipe is
considered already satisfied, which turns what would be a hard error into
a silent success. In a repo where `make build` is the documented way to
build, a caller checking the exit code concluded the build worked. Now
`build: vaultik`, verified in both directions - a clean build produces
the binary, a deliberately broken one exits non-zero and produces none.

All 19 .PHONY names were audited; build was the only one lacking a rule.
TestPhonyTargetsAllHaveRules keeps that true for names added later, so
the class is closed rather than the instance.
2026-08-09 19:53:46 +02:00
clawbot f21e7c9e70 Suppress the startup banner for --json (closes #106)
check / check (push) Successful in 2m31s
The banner is printed to stdout before cobra parses, and
bannerSuppressedInArgs recognised only --quiet, -q and --cron. So every
--json document was preceded by two banner lines and a blank one, and
`vaultik snapshot list --json | jq` failed. Passing opts.JSON as
extraQuiet did not help: that calls UI.SetQuiet in an fx OnStart hook,
long after Entry has printed.

The raw-argv scan is extended rather than the banner moved after
parsing. root.go documents that the banner must survive cobra rejecting
its arguments and --help, and no single post-parse location covers those
paths. The subcommand-versus-persistent distinction does not decide it:
--cron is already in the suppression list and is itself subcommand-only,
existing on snapshot create alone, so this adds another instance of an
accepted imprecision rather than a new kind. The error directions are
asymmetric - a false positive loses a decorative banner, a false negative
corrupts a document - so the scan errs toward suppression, which is also
why --json=false suppresses, exactly as --quiet=false already does.

Four of the five --json commands now pipe into jq cleanly with no other
flags: snapshot list, snapshot verify, snapshot remove, remote info.
prune does not, because pruneLocalSnapshots writes three prose lines to
stdout with no --json awareness. That reproduces identically before this
change and -q never suppressed it either, since printlnStdout and
stdoutf bypass v.UI entirely. Tracked as #108.

Also fixed: TTYHandler's human-readable byte formatting did not survive
grouping, because the key check compared against the bare attribute name
and a grouped record presents it qualified. AGENTS.md policy 9 keyed the
log format on stdout's TTY-ness, which #82 made false by moving the
logger to stderr; it now names the log stream. Vaultik.Stderr keeps its
field with the comment amended to say outright that nothing writes to
it, and the dead listEnv.stderr is removed.
2026-08-09 19:18:36 +02:00