4 Commits
Author SHA1 Message Date
sneak 8baa11b6cb Report a prune count that could not be read as unknown, not 0 (closes #96)
check / check (pull_request) Successful in 4m8s
PruneDatabase read seven table counts with the error discarded, so a
query that could not run silently became 0 and the before/after delta
computed from it looked like real work.

Each read now goes through a helper that logs at warn on failure and
returns nil; nil renders as "unknown", never "0", so an empty table is
distinguishable from one that could not be queried. Deltas built from an
unknown count are themselves unknown. The failure is surfaced, not
propagated, so the command's failure conditions are unchanged. These
counts have no --json output — under --json the summary is suppressed
entirely — so nothing there can show a false 0.

Model: opus-4-8
2026-09-21 19:26:36 +00:00
clawbot 5927e1aa3d Write file:// blobs atomically via temp file and rename (closes #130)
check / check (push) Failing after 0s
check / check (pull_request) Failing after 0s
The file:// backend streamed each object straight to its final key, so an upload cut off mid-stream left a truncated object there. The next backup saw that Stat succeeded, recorded the blob as complete, and produced a snapshot that reported success but could not be restored.

Writes now go to a temporary file with a .partial suffix in the destination directory, are synced, then renamed onto the key. List and ListStream skip .partial files, so a leftover is never trusted as a blob and is overwritten when the key is written again. S3 PutObject is already atomic.

Disclosure: the containing directory is not synced after the rename, so a host crash right after it could still lose the object on some filesystems.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)
2026-09-21 21:24:42 +02:00
clawbot 9ca962969a Map s3 not-found to storage.ErrNotFound in Get and Stat (closes #129)
check / check (push) Failing after 1s
check / check (pull_request) Successful in 3m50s
The Storer interface documents that Get and Stat return storage.ErrNotFound for a missing object. The file and rclone backends did; the s3 backend returned the raw SDK error, so callers testing for ErrNotFound behaved differently on s3.

S3Storer.Get and Stat now wrap ErrNotFound when the SDK reports a missing object and leave every other error untouched. The SDK reports a missing key two ways (NoSuchKey from Get, NotFound from Head); both are recognised in one helper, s3.IsNotFound, which HeadObject now also uses. The mapping lives in the storage package because internal/s3 cannot import it.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)
2026-09-21 21:07:35 +02:00
clawbot 89ebfc78e2 Use one duration parser and fix the --older-than months example (closes #123)
check / check (push) Failing after 1s
check / check (pull_request) Failing after 1s
Two parseDuration functions existed with different grammars; only the one in internal/vaultik/helpers.go was reachable from a flag. The unused copy in internal/cli/duration.go is deleted, so no flag accepts anything it did not before.

The README gave 6m as the six-months example for snapshot purge --older-than, but m is minutes: that command removed every snapshot older than six minutes. The example is now 6mo, and the help for --older-than and --keep-newer-than states that m is minutes and mo is months.

The parser now rejects negative durations, which it used to accept or silently make positive.

model: claude-opus-4-8 (implementation, review); claude-fable-5-1 (merge)

Co-authored-by: clawbot <clawbot@noreply.example.org>
2026-09-21 20:58:30 +02:00
14 changed files with 520 additions and 523 deletions
+2 -1
View File
@@ -251,7 +251,8 @@ local index alone, and still exits zero.
per-snapshot-name (`--keep-latest` keeps the latest of each name, not the per-snapshot-name (`--keep-latest` keeps the latest of each name, not the
latest globally). latest globally).
* `--keep-latest`: Keep only the most recent snapshot of each name * `--keep-latest`: Keep only the most recent snapshot of each name
* `--older-than <duration>`: Remove snapshots older than duration (e.g. `30d`, `6m`, `1y`) * `--older-than <duration>`: Remove snapshots older than duration (e.g. `30d`,
`4w`, `6mo`, `1y`; `m` is minutes, `mo` is months)
* `--snapshot <name>`: Restrict to specific snapshot names (repeat for multiple) * `--snapshot <name>`: Restrict to specific snapshot names (repeat for multiple)
* `--force`: Skip confirmation prompt * `--force`: Skip confirmation prompt
+36
View File
@@ -25,6 +25,24 @@ release" is exactly the contradiction
# Completed Steps # Completed Steps
- 2026-09-21: Stopped `prune` from reporting a failed row count as 0
([issue #96](https://git.eeqj.de/sneak/vaultik/issues/96)). The seven
`getTableCount` reads in `PruneDatabase` discarded their error, so a
query that could not run became a plausible `0` and the before/after
delta computed from it looked like real work. Each read now logs at
warn on failure and renders as `unknown`, never `0`, so an empty table
is distinguishable from one that could not be queried. The counts have
no `--json` representation — under `--json` the summary is suppressed
entirely — so nothing there can show a false `0`.
- 2026-09-21: Made the s3 storage backend report a missing object as
`storage.ErrNotFound`, like the `file` and `rclone` backends and as the
`Storer` interface documents. `S3Storer.Get` and `Stat` returned the raw
AWS SDK error, so `errors.Is(err, storage.ErrNotFound)` was false on s3
and callers branched differently per backend. Added a small `s3.IsNotFound`
helper (reused by `HeadObject`) and a test that a missing key maps to
`ErrNotFound`
([issue #129](https://git.eeqj.de/sneak/vaultik/issues/129)).
- 2026-09-21: Fixed `verify --deep` reporting healthy snapshots as - 2026-09-21: Fixed `verify --deep` reporting healthy snapshots as
corrupt. Its final blob-integrity check hashed the encrypted corrupt. Its final blob-integrity check hashed the encrypted
downloaded bytes with a single SHA256 and compared that to the blob downloaded bytes with a single SHA256 and compared that to the blob
@@ -60,6 +78,24 @@ release" is exactly the contradiction
keeps that exact compiler from auto-switching. Bumping Go now touches keeps that exact compiler from auto-switching. Bumping Go now touches
`go.mod`, the checksum, and the `Dockerfile` `golang` digest together. `go.mod`, the checksum, and the `Dockerfile` `golang` digest together.
- 2026-09-21: Collapsed the two duration parsers into one and fixed the
`--older-than` months example
([issue #123](https://git.eeqj.de/sneak/vaultik/issues/123)). Two
functions named `parseDuration` existed with different grammars;
`snapshot purge --older-than` and `--keep-newer-than` both already went
through the one in `internal/vaultik`, while the richer copy in
`internal/cli/duration.go` was reachable only from its own test. Kept
the live-path parser and deleted the unused one, so no flag's accepted
grammar changes. The trap the issue was filed over: `README.md`
documented `6m` as the months example for `--older-than`, but `m` is
minutes, so the documented command deleted every snapshot older than
six minutes on a destructive flag. Corrected the doc to `6mo` and put
both flags' help text on one example list that states `m` is minutes
and `mo` is months. The surviving parser now rejects negatives, which
it previously accepted (`-5h`) or silently made positive (`-5d`).
Table-driven tests cover every unit, `6m` as six minutes, `6mo` as 180
days, and rejection of a bare number, an unknown unit, and a negative.
- 2026-08-10: Moved every lint run into its own container, as a build - 2026-08-10: Moved every lint run into its own container, as a build
step ([issue #113](https://git.eeqj.de/sneak/vaultik/issues/113)). step ([issue #113](https://git.eeqj.de/sneak/vaultik/issues/113)).
New root `Dockerfile.lint`, built by `script/lint`, runs New root `Dockerfile.lint`, built by `script/lint`, runs
-126
View File
@@ -1,126 +0,0 @@
package cli
import (
"errors"
"fmt"
"regexp"
"strconv"
"strings"
"time"
)
// Approximate lengths of the extended calendar units accepted by
// parseDuration.
const (
durationDay = 24 * time.Hour
durationWeek = 7 * durationDay
durationMonth = 30 * durationDay
durationYear = 365 * durationDay
)
var (
errNegativeDuration = errors.New("negative durations are not supported")
errInvalidDuration = errors.New("invalid duration format")
errUnknownTimeUnit = errors.New("unknown time unit")
)
// parseDuration parses duration strings. Supports standard Go duration format
// (e.g., "3h30m", "1h45m30s") as well as extended units:
// - d: days (e.g., "30d", "7d")
// - w: weeks (e.g., "2w", "4w")
// - mo: months (30 days) (e.g., "6mo", "1mo")
// - y: years (365 days) (e.g., "1y", "2y")
//
// Can combine units: "1y6mo", "2w3d", "1d12h30m"
func parseDuration(s string) (time.Duration, error) {
// First try standard Go duration parsing
d, err := time.ParseDuration(s)
if err == nil {
return d, nil
}
// Extended duration parsing
// Check for negative values
if strings.HasPrefix(strings.TrimSpace(s), "-") {
return 0, errNegativeDuration
}
// Pattern matches: number + unit, repeated
re := regexp.MustCompile(`(\d+(?:\.\d+)?)\s*([a-zA-Z]+)`)
matches := re.FindAllStringSubmatch(s, -1)
if len(matches) == 0 {
return 0, fmt.Errorf("%w: %q", errInvalidDuration, s)
}
var total time.Duration
for _, match := range matches {
valueStr := match[1]
unit := strings.ToLower(match[2])
value, err := strconv.ParseFloat(valueStr, 64)
if err != nil {
return 0, fmt.Errorf("invalid number %q: %w", valueStr, err)
}
d, err := durationForUnit(value, unit)
if err != nil {
return 0, err
}
total += d
}
return total, nil
}
// durationForUnit converts a value with a (case-normalized) unit suffix
// into a time.Duration, accepting Go's standard units plus the extended
// calendar units.
func durationForUnit(value float64, unit string) (time.Duration, error) {
switch unit {
// Standard time units
case "ns", "nanosecond", "nanoseconds":
return time.Duration(value), nil
case "us", "µs", "microsecond", "microseconds":
return time.Duration(value * float64(time.Microsecond)), nil
case "ms", "millisecond", "milliseconds":
return time.Duration(value * float64(time.Millisecond)), nil
case "s", "sec", "second", "seconds":
return time.Duration(value * float64(time.Second)), nil
case "m", "min", "minute", "minutes":
return time.Duration(value * float64(time.Minute)), nil
case "h", "hr", "hour", "hours":
return time.Duration(value * float64(time.Hour)), nil
// Extended units
case "d", "day", "days":
return time.Duration(value * float64(durationDay)), nil
case "w", "week", "weeks":
return time.Duration(value * float64(durationWeek)), nil
case "mo", "month", "months":
// Using 30 days as approximation
return time.Duration(value * float64(durationMonth)), nil
case "y", "year", "years":
// Using 365 days as approximation
return time.Duration(value * float64(durationYear)), nil
default:
// Try parsing as standard Go duration unit
testStr := "1" + unit
_, err := time.ParseDuration(testStr)
if err != nil {
return 0, fmt.Errorf("%w: %q", errUnknownTimeUnit, unit)
}
// It's a valid Go duration unit, parse the full value
fullStr := fmt.Sprintf("%g%s", value, unit)
d, err := time.ParseDuration(fullStr)
if err != nil {
return 0, fmt.Errorf("invalid duration %q: %w", fullStr, err)
}
return d, nil
}
}
-299
View File
@@ -1,299 +0,0 @@
package cli //nolint:testpackage // needs access to unexported parseDuration
import (
"testing"
"time"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
type parseDurationCase struct {
name string
input string
expected time.Duration
wantErr bool
}
// runParseDurationCases executes a table of parseDuration cases as
// parallel subtests.
func runParseDurationCases(t *testing.T, tests []parseDurationCase) {
t.Helper()
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
got, err := parseDuration(tt.input)
if tt.wantErr {
require.Error(t, err, "expected error for input %q", tt.input)
return
}
require.NoError(t, err, "unexpected error for input %q", tt.input)
assert.Equal(t, tt.expected, got, "duration mismatch for input %q", tt.input)
})
}
}
func TestParseDurationStandard(t *testing.T) {
t.Parallel()
runParseDurationCases(t, []parseDurationCase{
{
name: "standard seconds",
input: "30s",
expected: 30 * time.Second,
},
{
name: "standard minutes",
input: "45m",
expected: 45 * time.Minute,
},
{
name: "standard hours",
input: "2h",
expected: 2 * time.Hour,
},
{
name: "standard combined",
input: "3h30m",
expected: 3*time.Hour + 30*time.Minute,
},
{
name: "standard complex",
input: "1h45m30s",
expected: 1*time.Hour + 45*time.Minute + 30*time.Second,
},
{
name: "standard with milliseconds",
input: "1s500ms",
expected: 1*time.Second + 500*time.Millisecond,
},
})
}
func TestParseDurationExtendedUnits(t *testing.T) {
t.Parallel()
runParseDurationCases(t, []parseDurationCase{
// Extended units - days
{
name: "single day",
input: "1d",
expected: 24 * time.Hour,
},
{
name: "multiple days",
input: "7d",
expected: 7 * 24 * time.Hour,
},
{
name: "fractional days",
input: "1.5d",
expected: 36 * time.Hour,
},
{
name: "days spelled out",
input: "3days",
expected: 3 * 24 * time.Hour,
},
// Extended units - weeks
{
name: "single week",
input: "1w",
expected: 7 * 24 * time.Hour,
},
{
name: "multiple weeks",
input: "4w",
expected: 4 * 7 * 24 * time.Hour,
},
{
name: "weeks spelled out",
input: "2weeks",
expected: 2 * 7 * 24 * time.Hour,
},
// Extended units - months
{
name: "single month",
input: "1mo",
expected: 30 * 24 * time.Hour,
},
{
name: "multiple months",
input: "6mo",
expected: 6 * 30 * 24 * time.Hour,
},
{
name: "months spelled out",
input: "3months",
expected: 3 * 30 * 24 * time.Hour,
},
// Extended units - years
{
name: "single year",
input: "1y",
expected: 365 * 24 * time.Hour,
},
{
name: "multiple years",
input: "2y",
expected: 2 * 365 * 24 * time.Hour,
},
{
name: "years spelled out",
input: "1year",
expected: 365 * 24 * time.Hour,
},
})
}
func TestParseDurationCombinedAndErrors(t *testing.T) {
t.Parallel()
runParseDurationCases(t, []parseDurationCase{
// Combined extended units
{
name: "weeks and days",
input: "2w3d",
expected: 2*7*24*time.Hour + 3*24*time.Hour,
},
{
name: "years and months",
input: "1y6mo",
expected: 365*24*time.Hour + 6*30*24*time.Hour,
},
{
name: "days and hours",
input: "1d12h",
expected: 24*time.Hour + 12*time.Hour,
},
{
name: "complex combination",
input: "1y2mo3w4d5h6m7s",
expected: 365*24*time.Hour + 2*30*24*time.Hour +
3*7*24*time.Hour + 4*24*time.Hour +
5*time.Hour + 6*time.Minute + 7*time.Second,
},
{
name: "with spaces",
input: "1d 12h 30m",
expected: 24*time.Hour + 12*time.Hour + 30*time.Minute,
},
// Edge cases
{
name: "zero duration",
input: "0s",
expected: 0,
},
{
name: "large duration",
input: "10y",
expected: 10 * 365 * 24 * time.Hour,
},
// Error cases
{
name: "empty string",
input: "",
wantErr: true,
},
{
name: "invalid format",
input: "abc",
wantErr: true,
},
{
name: "unknown unit",
input: "5x",
wantErr: true,
},
{
name: "invalid number",
input: "xyzd",
wantErr: true,
},
{
name: "negative not supported",
input: "-5d",
wantErr: true,
},
})
}
func TestParseDurationSpecialCases(t *testing.T) {
t.Parallel()
// Test that standard Go durations work exactly as expected
standardDurations := []string{
"300ms",
"1.5h",
"2h45m",
"72h",
"1us",
"1µs",
"1ns",
}
for _, d := range standardDurations {
expected, err := time.ParseDuration(d)
require.NoError(t, err)
got, err := parseDuration(d)
require.NoError(t, err)
assert.Equal(t, expected, got, "standard duration %q should parse identically", d)
}
}
func TestParseDurationRealWorldExamples(t *testing.T) {
t.Parallel()
// Test real-world snapshot purge scenarios
tests := []struct {
description string
input string
olderThan time.Duration
}{
{
description: "keep snapshots from last 30 days",
input: "30d",
olderThan: 30 * 24 * time.Hour,
},
{
description: "keep snapshots from last 6 months",
input: "6mo",
olderThan: 6 * 30 * 24 * time.Hour,
},
{
description: "keep snapshots from last year",
input: "1y",
olderThan: 365 * 24 * time.Hour,
},
{
description: "keep snapshots from last week and a half",
input: "1w3d",
olderThan: 10 * 24 * time.Hour,
},
{
description: "keep snapshots from last 90 days",
input: "90d",
olderThan: 90 * 24 * time.Hour,
},
}
for _, tt := range tests {
t.Run(tt.description, func(t *testing.T) {
t.Parallel()
got, err := parseDuration(tt.input)
require.NoError(t, err)
assert.Equal(t, tt.olderThan, got)
// Verify the duration makes sense for snapshot purging
assert.Greater(t, got, time.Hour,
"snapshot purge duration should be at least an hour")
})
}
}
+4 -2
View File
@@ -141,7 +141,8 @@ specifying a path using --config or by setting VAULTIK_CONFIG to a path.`,
"orphaned blobs") "orphaned blobs")
cmd.Flags().StringVar(&opts.KeepNewerThan, "keep-newer-than", "", cmd.Flags().StringVar(&opts.KeepNewerThan, "keep-newer-than", "",
"With --prune: keep snapshots newer than this duration "+ "With --prune: keep snapshots newer than this duration "+
"(e.g. 4w, 30d, 6mo) instead of only the latest") "(e.g. 30d, 4w, 6mo, 1y; m is minutes, mo is months) "+
"instead of only the latest")
return cmd return cmd
} }
@@ -204,7 +205,8 @@ restrict the operation to specific snapshot names.`,
cmd.Flags().BoolVar(&opts.KeepLatest, "keep-latest", false, cmd.Flags().BoolVar(&opts.KeepLatest, "keep-latest", false,
"Keep only the latest snapshot of each name") "Keep only the latest snapshot of each name")
cmd.Flags().StringVar(&opts.OlderThan, "older-than", "", cmd.Flags().StringVar(&opts.OlderThan, "older-than", "",
"Remove snapshots older than duration (e.g., 30d, 6m, 1y)") "Remove snapshots older than duration "+
"(e.g. 30d, 4w, 6mo, 1y; m is minutes, mo is months)")
cmd.Flags().BoolVar(&opts.Force, "force", false, "Skip confirmation prompt") cmd.Flags().BoolVar(&opts.Force, "force", false, "Skip confirmation prompt")
cmd.Flags().StringArrayVar(&opts.Names, "snapshot", nil, cmd.Flags().StringArrayVar(&opts.Names, "snapshot", nil,
"Restrict to snapshots with these names (repeat for multiple)") "Restrict to snapshots with these names (repeat for multiple)")
+13 -5
View File
@@ -219,11 +219,7 @@ func (c *Client) HeadObject(ctx context.Context, key string) (bool, error) {
Key: aws.String(fullKey), Key: aws.String(fullKey),
}) })
if err != nil { if err != nil {
var ( if IsNotFound(err) {
notFound *s3types.NotFound
noSuchKey *s3types.NoSuchKey
)
if errors.As(err, &notFound) || errors.As(err, &noSuchKey) {
return false, nil return false, nil
} }
@@ -233,6 +229,18 @@ func (c *Client) HeadObject(ctx context.Context, key string) (bool, error) {
return true, nil return true, nil
} }
// IsNotFound reports whether err indicates that an object does not exist.
// Head and Get requests surface a missing object as different SDK types,
// so both are checked here.
func IsNotFound(err error) bool {
var (
notFound *s3types.NotFound
noSuchKey *s3types.NoSuchKey
)
return errors.As(err, &notFound) || errors.As(err, &noSuchKey)
}
// ObjectInfo contains information about an S3 object. // ObjectInfo contains information about an S3 object.
// It is used by ListObjectsStream to return object metadata // It is used by ListObjectsStream to return object metadata
// along with any errors encountered during listing. // along with any errors encountered during listing.
+79 -54
View File
@@ -46,31 +46,18 @@ func (f *FileStorer) SetFilesystem(fs afero.Fs) {
// storage base path. // storage base path.
const storageDirPerm = 0o755 const storageDirPerm = 0o755
// tempSuffix marks a partially written object. writeAtomic streams into a
// temp file carrying this suffix and only renames it onto the real key once
// the whole object is on disk, so an interrupted write can never leave a
// truncated object at the key a later run would Stat and trust as a complete
// blob. List and ListStream skip these files, so a leftover from an
// interrupted write is never listed or trusted as a blob; it is otherwise
// harmless and is overwritten when the same key is written again.
const tempSuffix = ".partial"
// Put stores data at the specified key. // Put stores data at the specified key.
func (f *FileStorer) Put(_ context.Context, key string, data io.Reader) error { func (f *FileStorer) Put(_ context.Context, key string, data io.Reader) error {
path := f.fullPath(key) return f.writeAtomic(key, data, nil)
// Create parent directories
dir := filepath.Dir(path)
err := f.fs.MkdirAll(dir, storageDirPerm)
if err != nil {
return fmt.Errorf("creating directories: %w", err)
}
file, err := f.fs.Create(path)
if err != nil {
return fmt.Errorf("creating file: %w", err)
}
defer func() { _ = file.Close() }()
_, err = io.Copy(file, data)
if err != nil {
return fmt.Errorf("writing file: %w", err)
}
return nil
} }
// PutWithProgress stores data with progress reporting. // PutWithProgress stores data with progress reporting.
@@ -78,35 +65,7 @@ func (f *FileStorer) PutWithProgress(
_ context.Context, key string, data io.Reader, _ context.Context, key string, data io.Reader,
_ int64, progress ProgressCallback, _ int64, progress ProgressCallback,
) error { ) error {
path := f.fullPath(key) return f.writeAtomic(key, data, progress)
// Create parent directories
dir := filepath.Dir(path)
err := f.fs.MkdirAll(dir, storageDirPerm)
if err != nil {
return fmt.Errorf("creating directories: %w", err)
}
file, err := f.fs.Create(path)
if err != nil {
return fmt.Errorf("creating file: %w", err)
}
defer func() { _ = file.Close() }()
// Wrap with progress tracking
pw := &progressWriter{
writer: file,
callback: progress,
}
_, err = io.Copy(pw, data)
if err != nil {
return fmt.Errorf("writing file: %w", err)
}
return nil
} }
// Get retrieves data from the specified key. // Get retrieves data from the specified key.
@@ -188,7 +147,7 @@ func (f *FileStorer) List(ctx context.Context, prefix string) ([]string, error)
default: default:
} }
if !info.IsDir() { if !info.IsDir() && !strings.HasSuffix(info.Name(), tempSuffix) {
// Convert back to key (relative path from basePath) // Convert back to key (relative path from basePath)
relPath, err := filepath.Rel(f.basePath, path) relPath, err := filepath.Rel(f.basePath, path)
if err != nil { if err != nil {
@@ -245,7 +204,7 @@ func (f *FileStorer) ListStream(ctx context.Context, prefix string) <-chan Objec
return nil //nolint:nilerr // continue walking despite errors return nil //nolint:nilerr // continue walking despite errors
} }
if !info.IsDir() { if !info.IsDir() && !strings.HasSuffix(info.Name(), tempSuffix) {
relPath, err := filepath.Rel(f.basePath, path) relPath, err := filepath.Rel(f.basePath, path)
if err != nil { if err != nil {
ch <- ObjectInfo{Err: fmt.Errorf("computing relative path: %w", err)} ch <- ObjectInfo{Err: fmt.Errorf("computing relative path: %w", err)}
@@ -275,6 +234,72 @@ func (f *FileStorer) Info() Info {
} }
} }
// writeAtomic streams data into a temp file in the destination directory,
// fsyncs it, and renames it onto the final key. The key therefore appears
// only once the whole object has been durably written; a failure part-way
// leaves a temp file (removed here on the failing path) rather than a
// truncated object at the key.
func (f *FileStorer) writeAtomic(
key string, data io.Reader, progress ProgressCallback,
) error {
path := f.fullPath(key)
dir := filepath.Dir(path)
err := f.fs.MkdirAll(dir, storageDirPerm)
if err != nil {
return fmt.Errorf("creating directories: %w", err)
}
tmp, err := afero.TempFile(f.fs, dir, filepath.Base(path)+"-*"+tempSuffix)
if err != nil {
return fmt.Errorf("creating temp file: %w", err)
}
tmpPath := tmp.Name()
// Remove the temp file unless the rename below claims it. On the success
// path renamed is true, so the deferred Close and Remove are harmless
// no-ops on a name that no longer exists.
renamed := false
defer func() {
_ = tmp.Close()
if !renamed {
_ = f.fs.Remove(tmpPath)
}
}()
var w io.Writer = tmp
if progress != nil {
w = &progressWriter{writer: tmp, callback: progress}
}
_, err = io.Copy(w, data)
if err != nil {
return fmt.Errorf("writing file: %w", err)
}
err = tmp.Sync()
if err != nil {
return fmt.Errorf("syncing temp file: %w", err)
}
err = tmp.Close()
if err != nil {
return fmt.Errorf("closing temp file: %w", err)
}
err = f.fs.Rename(tmpPath, path)
if err != nil {
return fmt.Errorf("renaming temp file: %w", err)
}
renamed = true
return nil
}
// fullPath returns the full filesystem path for a key. // fullPath returns the full filesystem path for a key.
func (f *FileStorer) fullPath(key string) string { func (f *FileStorer) fullPath(key string) string {
return filepath.Join(f.basePath, key) return filepath.Join(f.basePath, key)
+119
View File
@@ -0,0 +1,119 @@
package storage_test
import (
"context"
"errors"
"os"
"path/filepath"
"strings"
"testing"
"sneak.berlin/go/vaultik/internal/storage"
)
// errStreamInterrupted stands in for an upload cut off mid-stream.
var errStreamInterrupted = errors.New("connection reset mid-upload")
// failingReader yields its data once, then fails.
type failingReader struct {
data []byte
done bool
}
func (r *failingReader) Read(p []byte) (int, error) {
if r.done {
return 0, errStreamInterrupted
}
n := copy(p, r.data)
r.done = true
return n, nil
}
// TestFileStorer_InterruptedWriteLeavesNoTrustedObject checks that a write
// cut off mid-stream leaves nothing at the destination key, so a later run
// cannot Stat a truncated object and trust it as a complete blob.
func TestFileStorer_InterruptedWriteLeavesNoTrustedObject(t *testing.T) {
t.Parallel()
f, err := storage.NewFileStorer(t.TempDir())
if err != nil {
t.Fatalf("NewFileStorer: %v", err)
}
ctx := context.Background()
key := "blobs/aa/bb/aabbccddeeff"
err = f.PutWithProgress(ctx, key, &failingReader{data: []byte("partial")}, 4096, nil)
if err == nil {
t.Fatal("expected the interrupted write to fail, got nil")
}
_, err = f.Stat(ctx, key)
if !errors.Is(err, storage.ErrNotFound) {
t.Fatalf("expected key absent after interrupted write, got Stat err %v", err)
}
keys, err := f.List(ctx, "blobs/")
if err != nil {
t.Fatalf("List: %v", err)
}
if len(keys) != 0 {
t.Fatalf("expected no keys listed after interrupted write, got %v", keys)
}
}
// TestFileStorer_ListSkipsPartialFiles checks that a leftover temp file (the
// storage layer names them with a ".partial" suffix) is never surfaced as a
// key by List or ListStream.
func TestFileStorer_ListSkipsPartialFiles(t *testing.T) {
t.Parallel()
base := t.TempDir()
f, err := storage.NewFileStorer(base)
if err != nil {
t.Fatalf("NewFileStorer: %v", err)
}
ctx := context.Background()
realKey := "blobs/aa/bb/aabbccddeeff"
err = f.Put(ctx, realKey, strings.NewReader("blob-bytes"))
if err != nil {
t.Fatalf("Put: %v", err)
}
// A stray temp file, as an interrupted write would leave behind.
leftover := filepath.Join(base, "blobs/aa/bb/aabbccddeeff-123456.partial")
err = os.WriteFile(leftover, []byte("half"), 0o600)
if err != nil {
t.Fatalf("writing leftover temp file: %v", err)
}
keys, err := f.List(ctx, "blobs/")
if err != nil {
t.Fatalf("List: %v", err)
}
if len(keys) != 1 || keys[0] != realKey {
t.Fatalf("List should return only the real key, got %v", keys)
}
var streamed []string
for obj := range f.ListStream(ctx, "blobs/") {
if obj.Err != nil {
t.Fatalf("ListStream: %v", obj.Err)
}
streamed = append(streamed, obj.Key)
}
if len(streamed) != 1 || streamed[0] != realKey {
t.Fatalf("ListStream should return only the real key, got %v", streamed)
}
}
+16 -1
View File
@@ -38,14 +38,29 @@ func (s *S3Storer) PutWithProgress(
} }
// Get retrieves data from the specified key. // Get retrieves data from the specified key.
// Returns ErrNotFound if the object does not exist.
func (s *S3Storer) Get(ctx context.Context, key string) (io.ReadCloser, error) { func (s *S3Storer) Get(ctx context.Context, key string) (io.ReadCloser, error) {
return s.client.GetObject(ctx, key) rc, err := s.client.GetObject(ctx, key)
if err != nil {
if s3.IsNotFound(err) {
return nil, fmt.Errorf("get %q: %w", key, ErrNotFound)
}
return nil, err
}
return rc, nil
} }
// Stat returns metadata about an object without retrieving its contents. // Stat returns metadata about an object without retrieving its contents.
// Returns ErrNotFound if the object does not exist.
func (s *S3Storer) Stat(ctx context.Context, key string) (*ObjectInfo, error) { func (s *S3Storer) Stat(ctx context.Context, key string) (*ObjectInfo, error) {
info, err := s.client.StatObject(ctx, key) info, err := s.client.StatObject(ctx, key)
if err != nil { if err != nil {
if s3.IsNotFound(err) {
return nil, fmt.Errorf("stat %q: %w", key, ErrNotFound)
}
return nil, err return nil, err
} }
+59
View File
@@ -0,0 +1,59 @@
package storage_test
import (
"context"
"errors"
"net/http/httptest"
"testing"
"github.com/johannesboyne/gofakes3"
"github.com/johannesboyne/gofakes3/backend/s3mem"
"sneak.berlin/go/vaultik/internal/s3"
"sneak.berlin/go/vaultik/internal/storage"
)
// TestS3StorerMissingKeyMapsToErrNotFound verifies that the s3 backend reports
// a missing object as storage.ErrNotFound, matching the file and rclone
// backends and the Storer contract. Without the mapping, Get and Stat leak the
// raw SDK error and errors.Is(err, storage.ErrNotFound) is false.
//
//nolint:paralleltest // shares an in-process S3 server via t.Cleanup
func TestS3StorerMissingKeyMapsToErrNotFound(t *testing.T) {
const bucket = "test-bucket"
backend := s3mem.New()
err := backend.CreateBucket(bucket)
if err != nil {
t.Fatalf("create bucket: %v", err)
}
srv := httptest.NewServer(gofakes3.New(backend).Server())
t.Cleanup(srv.Close)
ctx := context.Background()
client, err := s3.NewClient(ctx, s3.Config{
Endpoint: srv.URL,
Bucket: bucket,
AccessKeyID: "test",
SecretAccessKey: "test",
Region: "us-east-1",
})
if err != nil {
t.Fatalf("new client: %v", err)
}
storer := storage.NewS3Storer(client)
_, err = storer.Get(ctx, "does-not-exist")
if !errors.Is(err, storage.ErrNotFound) {
t.Errorf("Get on missing key: got %v, want ErrNotFound", err)
}
_, err = storer.Stat(ctx, "does-not-exist")
if !errors.Is(err, storage.ErrNotFound) {
t.Errorf("Stat on missing key: got %v, want ErrNotFound", err)
}
}
+7 -1
View File
@@ -35,6 +35,7 @@ var (
"invalid snapshot ID format: expected hostname_snapshotname_timestamp") "invalid snapshot ID format: expected hostname_snapshotname_timestamp")
errInvalidDuration = errors.New("invalid duration") errInvalidDuration = errors.New("invalid duration")
errUnknownTimeUnit = errors.New("unknown time unit") errUnknownTimeUnit = errors.New("unknown time unit")
errNegativeDuration = errors.New("negative durations are not supported")
) )
// Time-unit lengths used by parseDuration. // Time-unit lengths used by parseDuration.
@@ -138,8 +139,13 @@ func parseSnapshotName(snapshotID string) string {
// parseDuration parses a duration string with support for human-friendly units: // parseDuration parses a duration string with support for human-friendly units:
// d/day/days, w/week/weeks, mo/month/months, y/year/years, plus standard Go // d/day/days, w/week/weeks, mo/month/months, y/year/years, plus standard Go
// duration units (h, m, s). // duration units. Following Go, m is minutes and mo is months. A bare number,
// an unknown unit, and a negative value are all rejected.
func parseDuration(s string) (time.Duration, error) { func parseDuration(s string) (time.Duration, error) {
if strings.HasPrefix(strings.TrimSpace(s), "-") {
return 0, errNegativeDuration
}
d, err := time.ParseDuration(s) d, err := time.ParseDuration(s)
if err == nil { if err == nil {
return d, nil return d, nil
+25 -6
View File
@@ -51,13 +51,32 @@ func TestParseDuration(t *testing.T) {
want time.Duration want time.Duration
err bool err bool
}{ }{
{"30d", 30 * 24 * time.Hour, false}, // Go units, including the m-is-minutes / mo-is-months distinction
{"4w", 4 * 7 * 24 * time.Hour, false}, // that this parser exists to keep straight.
{"6mo", 6 * 30 * 24 * time.Hour, false}, {"10ns", 10 * time.Nanosecond, false},
{"1y", 365 * 24 * time.Hour, false}, {"10us", 10 * time.Microsecond, false},
{"2w3d", 2*7*24*time.Hour + 3*24*time.Hour, false}, {"500ms", 500 * time.Millisecond, false},
{"1h", time.Hour, false},
{"30s", 30 * time.Second, false}, {"30s", 30 * time.Second, false},
{"6m", 6 * time.Minute, false},
{"1h", time.Hour, false},
// Extended calendar units.
{"30d", 30 * 24 * time.Hour, false},
{"3days", 3 * 24 * time.Hour, false},
{"4w", 4 * 7 * 24 * time.Hour, false},
{"2weeks", 2 * 7 * 24 * time.Hour, false},
{"6mo", 180 * 24 * time.Hour, false},
{"1month", 30 * 24 * time.Hour, false},
{"1y", 365 * 24 * time.Hour, false},
{"2years", 2 * 365 * 24 * time.Hour, false},
// Combined units.
{"2w3d", 2*7*24*time.Hour + 3*24*time.Hour, false},
{"1y6mo", 365*24*time.Hour + 180*24*time.Hour, false},
// Rejected inputs.
{"6", 0, true}, // bare number, no unit
{"5x", 0, true}, // unknown unit
{"-5d", 0, true}, // negative, extended unit
{"-5h", 0, true}, // negative, Go unit
{"", 0, true}, // empty
{"garbage", 0, true}, {"garbage", 0, true},
} }
+79
View File
@@ -0,0 +1,79 @@
package vaultik //nolint:testpackage // exercises unexported count helpers
import (
"context"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/vaultik/internal/database"
"sneak.berlin/go/vaultik/internal/log"
)
// TestTableCountForReportSurfacesReadFailure is the regression guard for
// the discarded-error bug: getTableCount for a table its query cannot
// resolve must not silently become 0. A count that could not be read is
// reported as unknown, which a reader can tell apart from an empty table.
//
//nolint:paralleltest // installs the global logger via log.Initialize
func TestTableCountForReportSurfacesReadFailure(t *testing.T) {
log.Initialize(log.Config{})
ctx := context.Background()
db, err := database.New(ctx, ":memory:")
require.NoError(t, err)
t.Cleanup(func() { _ = db.Close() })
v := &Vaultik{DB: db}
v.SetContext(ctx)
// A table present in the schema reads as a real count.
blobs := v.tableCountForReport("blobs")
require.NotNil(t, blobs, "an existing table must read as a real count")
assert.Equal(t, int64(0), *blobs)
// A syntactically valid name the sanitizer accepts but whose table
// the query cannot resolve is the exact shape #96 describes: a
// would-be loud failure that used to be discarded into a 0.
_, err = v.getTableCount("snapshots_missing")
require.Error(t, err, "a query against a nonexistent table must fail")
missing := v.tableCountForReport("snapshots_missing")
assert.Nil(t, missing, "a failed read is unknown, not a count")
// The rendered count for a failed read must say unknown, never 0.
assert.Equal(t, countUnknown, countText(missing))
assert.NotEqual(t, "0", countText(missing))
}
// TestCountTextDistinguishesEmptyFromUnknown pins the distinction the
// output has to preserve: 0 means the table was empty, "unknown" means
// the count could not be read.
func TestCountTextDistinguishesEmptyFromUnknown(t *testing.T) {
t.Parallel()
zero := int64(0)
seven := int64(7)
assert.Equal(t, "0", countText(&zero))
assert.Equal(t, "7", countText(&seven))
assert.Equal(t, countUnknown, countText(nil))
}
// TestCountDiffUnknownWhenEitherSideUnknown checks that a delta computed
// from an unreadable count is itself unknown rather than a plausible
// number.
func TestCountDiffUnknownWhenEitherSideUnknown(t *testing.T) {
t.Parallel()
before := int64(10)
after := int64(3)
require.NotNil(t, countDiff(&before, &after))
assert.Equal(t, int64(7), *countDiff(&before, &after))
assert.Nil(t, countDiff(nil, &after), "unknown before yields unknown delta")
assert.Nil(t, countDiff(&before, nil), "unknown after yields unknown delta")
assert.Nil(t, countDiff(nil, nil))
}
+79 -26
View File
@@ -8,6 +8,7 @@ import (
"path/filepath" "path/filepath"
"regexp" "regexp"
"sort" "sort"
"strconv"
"strings" "strings"
"time" "time"
@@ -1540,12 +1541,17 @@ func (v *Vaultik) outputRemoveJSON(result *RemoveResult) error {
return encoder.Encode(result) return encoder.Encode(result)
} }
// PruneResult contains statistics about the prune operation // PruneResult contains statistics about the prune operation.
// SnapshotsDeleted counts snapshots actually deleted. FilesDeleted,
// ChunksDeleted, and BlobsDeleted are derived from before/after row
// counts of the local index; each is nil when a count could not be read,
// so an unreadable count is reported as unknown rather than silently
// as 0.
type PruneResult struct { type PruneResult struct {
SnapshotsDeleted int64 SnapshotsDeleted int64
FilesDeleted int64 FilesDeleted *int64
ChunksDeleted int64 ChunksDeleted *int64
BlobsDeleted int64 BlobsDeleted *int64
} }
// PruneDatabase removes incomplete snapshots and orphaned files, chunks, // PruneDatabase removes incomplete snapshots and orphaned files, chunks,
@@ -1560,7 +1566,7 @@ func (v *Vaultik) PruneDatabase() (*PruneResult, error) {
result := &PruneResult{} result := &PruneResult{}
// Snapshot counts before deletion of incompletes. // Snapshot counts before deletion of incompletes.
snapshotCountBefore, _ := v.getTableCount("snapshots") snapshotCountBefore := v.tableCountForReport("snapshots")
// First, delete any incomplete snapshots // First, delete any incomplete snapshots
incompleteSnapshots, err := v.Repositories.Snapshots.GetIncompleteSnapshots(v.ctx) incompleteSnapshots, err := v.Repositories.Snapshots.GetIncompleteSnapshots(v.ctx)
@@ -1575,9 +1581,9 @@ func (v *Vaultik) PruneDatabase() (*PruneResult, error) {
} }
// Get counts before cleanup for reporting // Get counts before cleanup for reporting
fileCountBefore, _ := v.getTableCount("files") fileCountBefore := v.tableCountForReport("files")
chunkCountBefore, _ := v.getTableCount("chunks") chunkCountBefore := v.tableCountForReport("chunks")
blobCountBefore, _ := v.getTableCount("blobs") blobCountBefore := v.tableCountForReport("blobs")
// Run the cleanup // Run the cleanup
err = v.SnapshotManager.CleanupOrphanedData(v.ctx) err = v.SnapshotManager.CleanupOrphanedData(v.ctx)
@@ -1586,36 +1592,83 @@ func (v *Vaultik) PruneDatabase() (*PruneResult, error) {
} }
// Get counts after cleanup // Get counts after cleanup
fileCountAfter, _ := v.getTableCount("files") fileCountAfter := v.tableCountForReport("files")
chunkCountAfter, _ := v.getTableCount("chunks") chunkCountAfter := v.tableCountForReport("chunks")
blobCountAfter, _ := v.getTableCount("blobs") blobCountAfter := v.tableCountForReport("blobs")
result.FilesDeleted = fileCountBefore - fileCountAfter result.FilesDeleted = countDiff(fileCountBefore, fileCountAfter)
result.ChunksDeleted = chunkCountBefore - chunkCountAfter result.ChunksDeleted = countDiff(chunkCountBefore, chunkCountAfter)
result.BlobsDeleted = blobCountBefore - blobCountAfter result.BlobsDeleted = countDiff(blobCountBefore, blobCountAfter)
log.Info("Local database prune complete", log.Info("Local database prune complete",
"incomplete_snapshots", result.SnapshotsDeleted, "incomplete_snapshots", result.SnapshotsDeleted,
"orphaned_files", result.FilesDeleted, "orphaned_files", countText(result.FilesDeleted),
"orphaned_chunks", result.ChunksDeleted, "orphaned_chunks", countText(result.ChunksDeleted),
"orphaned_blobs", result.BlobsDeleted, "orphaned_blobs", countText(result.BlobsDeleted),
) )
snapshotCountAfter := snapshotCountBefore - result.SnapshotsDeleted // Snapshots remaining after removing the incomplete ones; unknown if
// the pre-prune snapshot count could not be read.
snapshotsRemain := countDiff(snapshotCountBefore, &result.SnapshotsDeleted)
v.UI.Completef("Pruned local index database.") v.UI.Completef("Pruned local index database.")
v.UI.Detailf("Incomplete snapshots: %d removed (%d remain).", v.UI.Detailf("Incomplete snapshots: %s removed (%s remain).",
result.SnapshotsDeleted, snapshotCountAfter) countText(&result.SnapshotsDeleted), countText(snapshotsRemain))
v.UI.Detailf("Orphaned files: %d removed (%d remain).", v.UI.Detailf("Orphaned files: %s removed (%s remain).",
result.FilesDeleted, fileCountAfter) countText(result.FilesDeleted), countText(fileCountAfter))
v.UI.Detailf("Orphaned chunks: %d removed (%d remain).", v.UI.Detailf("Orphaned chunks: %s removed (%s remain).",
result.ChunksDeleted, chunkCountAfter) countText(result.ChunksDeleted), countText(chunkCountAfter))
v.UI.Detailf("Orphaned blobs: %d removed (%d remain).", v.UI.Detailf("Orphaned blobs: %s removed (%s remain).",
result.BlobsDeleted, blobCountAfter) countText(result.BlobsDeleted), countText(blobCountAfter))
return result, nil return result, nil
} }
// countUnknown is what a count reads as when its query could not be run,
// distinct from "0", which means the table really was empty.
const countUnknown = "unknown"
// tableCountForReport returns the row count of a table for the prune
// summary, or nil if the count could not be read. A read failure is
// logged at warn — visible even under --json, which routes warnings to
// stderr — and then rendered as unknown rather than silently becoming 0,
// so a broken query is a visible failure instead of a plausible wrong
// number.
func (v *Vaultik) tableCountForReport(tableName string) *int64 {
count, err := v.getTableCount(tableName)
if err != nil {
log.Warn("could not read table row count for prune summary",
"table", tableName, "error", err)
return nil
}
return &count
}
// countDiff returns before-after, or nil if either count is unknown so
// that an unreadable count does not collapse into a plausible delta.
func countDiff(before, after *int64) *int64 {
if before == nil || after == nil {
return nil
}
diff := *before - *after
return &diff
}
// countText renders a count that may be unknown: nil (the read failed)
// becomes "unknown", never "0", so a reader can tell an empty table from
// one that could not be queried.
func countText(count *int64) string {
if count == nil {
return countUnknown
}
return strconv.FormatInt(*count, 10)
}
// validTableNameRe matches table names containing only lowercase // validTableNameRe matches table names containing only lowercase
// alphanumeric characters and underscores. // alphanumeric characters and underscores.
var validTableNameRe = regexp.MustCompile(`^[a-z0-9_]+$`) var validTableNameRe = regexp.MustCompile(`^[a-z0-9_]+$`)