Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
ac27768edf |
+5
-9
@@ -54,12 +54,10 @@ The database tracks five primary entities and their relationships:
|
||||
|
||||
#### File (`database.File`)
|
||||
Represents a file, directory, or symlink in the backup system. Stores metadata needed for restoration:
|
||||
- Path, mtime
|
||||
- Path, source_path (for restore path stripping), mtime
|
||||
- Size, mode, ownership (uid, gid)
|
||||
- Symlink target (if applicable)
|
||||
|
||||
It also stores `source_path`, the source directory the scan found it under, made absolute and with symlinks resolved. Restore does not read it.
|
||||
|
||||
#### Chunk (`database.Chunk`)
|
||||
A content-addressed unit of data. Files are split into variable-size chunks using the FastCDC algorithm:
|
||||
- `ChunkHash`: SHA256 hash of chunk content (primary key)
|
||||
@@ -84,11 +82,9 @@ The final storage unit uploaded to S3. Contains many compressed and encrypted ch
|
||||
Blob creation process:
|
||||
1. Chunks are accumulated (up to MaxBlobSize, typically 10GB)
|
||||
2. As each chunk is added, its uncompressed bytes are fed to a running SHA-256
|
||||
3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and written to a temporary file in `$TMPDIR` (`/tmp` when unset)
|
||||
3. Concurrently, the same bytes are compressed with zstd, then encrypted with age (recipients configured in config), and streamed to storage
|
||||
4. On finalize, the blob's name is the double SHA-256 of the uncompressed contents — `hex(SHA256(SHA256(...)))` — not a hash of the compressed, encrypted bytes
|
||||
5. The finished file is uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}` and then deleted
|
||||
|
||||
The metadata export at the end of a backup also uses `$TMPDIR`: it copies the local index there and runs `VACUUM` on the copy, which writes another temporary copy and a write-ahead log. A backup therefore needs free space in `$TMPDIR` of the larger of `blob_size_limit` and about three times the size of the local index.
|
||||
5. Uploaded to `blobs/{hash[0:2]}/{hash[2:4]}/{hash}`
|
||||
|
||||
#### BlobChunk (`database.BlobChunk`)
|
||||
Maps chunks to their position within blobs:
|
||||
@@ -339,10 +335,10 @@ CreateSnapshot(opts)
|
||||
│ │
|
||||
│ └─► Accumulate statistics
|
||||
│
|
||||
├─► SnapshotManager.PopulateSnapshotBlobs() // record referenced blobs
|
||||
│
|
||||
├─► SnapshotManager.UpdateSnapshotStatsExtended()
|
||||
│
|
||||
├─► SnapshotManager.PopulateSnapshotBlobs() // record referenced blobs
|
||||
│
|
||||
├─► SnapshotManager.ExportSnapshotMetadata()
|
||||
│ │
|
||||
│ ├─► Copy database to temp file
|
||||
|
||||
@@ -352,9 +352,8 @@ may hold snapshots this host doesn't know about), which is what
|
||||
prune` invocation to run as a follow-up. Local row cleanup (files,
|
||||
chunks, blobs the snapshot was the last referrer for) runs
|
||||
automatically. If the destination store is unreachable, the local-DB
|
||||
removal still completes and a warning is emitted; run `vaultik snapshot
|
||||
remove <snapshot-id>` again once the store is reachable to remove the
|
||||
snapshot's metadata from it (`vaultik prune` does not). To wipe everything
|
||||
removal still completes and a warning is emitted; rerun `vaultik prune`
|
||||
once the store is reachable to finish remote cleanup. To wipe everything
|
||||
on the destination in one go, use `vaultik remote nuke --force`.
|
||||
* `--local-only`: Skip remote cleanup; only touch the local index
|
||||
* `--dry-run`: Show what would be deleted without deleting
|
||||
@@ -390,13 +389,7 @@ recipients, and local database statistics.
|
||||
|
||||
**`remote info`**: Show storage backend type and location plus detailed
|
||||
remote storage inventory: per-snapshot metadata sizes, blob counts, and
|
||||
orphaned blob detection. A name under `metadata/` that is not a remote
|
||||
key is skipped with a warning and is not printed. If a listed
|
||||
`manifest.json.zst` cannot be read, or sits under a skipped name, the
|
||||
orphaned blob figures are reported as unknown; `--json` gives them as
|
||||
`null`, lists the remote key of each unreadable manifest in
|
||||
`unreadable_manifests` and counts the manifests under skipped names in
|
||||
`skipped_manifest_count`.
|
||||
orphaned blob detection.
|
||||
* `--json`: Output as JSON
|
||||
|
||||
**`remote nuke`**: Delete every snapshot's metadata and every blob from the
|
||||
@@ -546,7 +539,7 @@ complete annotated example also lives in
|
||||
| `s3.*` | | Legacy S3 configuration (endpoint, bucket, credentials) |
|
||||
| `exclude` | | Global exclude patterns (applied to all snapshots) |
|
||||
| `chunk_size` | `10MB` | Average chunk size for content-defined chunking |
|
||||
| `blob_size_limit` | `10GB` | Maximum blob size before splitting. Must be at least four times `chunk_size` (the largest chunk the chunker can emit), otherwise a single-chunk blob could exceed the limit. Each blob is written in full to a temporary file in `$TMPDIR` (`/tmp` when unset) before it is uploaded, and the metadata export at the end of a backup works on a copy of the local index there. A backup needs free space there of the larger of this limit and about three times the size of the local index |
|
||||
| `blob_size_limit` | `10GB` | Maximum blob size before splitting. Must be at least four times `chunk_size` (the largest chunk the chunker can emit), otherwise a single-chunk blob could exceed the limit |
|
||||
| `compression_level` | `3` | zstd compression level (1-19) |
|
||||
| `hostname` | system hostname | Hostname used in snapshot IDs |
|
||||
| `index_path` | platform data dir | Local SQLite index path |
|
||||
@@ -805,8 +798,7 @@ them. We provide:
|
||||
in the `Dockerfile` together.
|
||||
* `script/release` — cross-compile and publish the release artifacts
|
||||
with the pinned `goreleaser`. Refuses a `goreleaser` on `PATH` whose
|
||||
version is not the pinned one, because a different version would build
|
||||
a different release from the same tag.
|
||||
version is not the pinned one, on the same reasoning as `script/lint`.
|
||||
* `script/release-snapshot` — the same build with no publishing and no
|
||||
tagging, into `./dist`
|
||||
* `script/test` — run the test suite by building the `test` phase of
|
||||
@@ -916,14 +908,14 @@ It is passed to `goreleaser` as `GITEA_TOKEN`. The runner's automatic
|
||||
token is deliberately not used: it is not guaranteed to carry release
|
||||
write access.
|
||||
|
||||
The Go toolchain that compiles the released binaries is installed by
|
||||
`script/install-go`, which downloads the version named by `go.mod`
|
||||
(currently `1.26.1`, the same version the `Dockerfile` builder stage
|
||||
pins by digest) and refuses the archive unless its sha256 matches the
|
||||
value committed in the script. `goreleaser` shells out to `go` for every
|
||||
The Go toolchain that compiles the released binaries comes from an
|
||||
`actions/setup-go` step pinned by commit sha, reading its version from
|
||||
`go.mod` (currently `1.26.1`, the same version the `Dockerfile` builder
|
||||
stage pins by digest). `goreleaser` shells out to `go` for every
|
||||
cross-compile, so without that step the release would either fail
|
||||
outright or ship binaries built by whatever unpinned toolchain the
|
||||
runner happened to carry.
|
||||
runner happened to carry — the one unpinned thing in an otherwise
|
||||
hash-pinned release path.
|
||||
|
||||
To rehearse the whole build without publishing or tagging anything:
|
||||
|
||||
|
||||
@@ -22,120 +22,6 @@ the tag exists and is exercised; what is left is merging `next` to
|
||||
|
||||
# Completed Steps
|
||||
|
||||
- 2026-10-07: Corrected documentation, help text and comments that were
|
||||
false about the code
|
||||
([issue #233](https://git.eeqj.de/sneak/vaultik/issues/233)). A blob
|
||||
is not streamed to storage; it is written in full to a temporary file
|
||||
in `$TMPDIR` and uploaded once finished, and the metadata export works
|
||||
on a copy of the local index there. The README and
|
||||
`config.example.yml` now say a backup needs free space there of the
|
||||
larger of `blob_size_limit` and about three times the size of the
|
||||
local index. Also corrected: the snapshot ID format, what restore
|
||||
reads and how incomplete snapshots are removed in `docs/DATAMODEL.md`,
|
||||
what `source_path` holds, the `index_path` and config file defaults,
|
||||
what `snapshot remove` cleans up, how the release gets its Go
|
||||
toolchain, and the `script/release` and `script/fmt-check` comments.
|
||||
|
||||
- 2026-10-07: Made two messages say only what is true
|
||||
([issue #240](https://git.eeqj.de/sneak/vaultik/issues/240)). A config
|
||||
file that others can read was warned about as containing S3
|
||||
credentials even when it set none, as a `file://` config does. The
|
||||
warning now says the file may contain S3 credentials only when
|
||||
`s3.access_key_id` or `s3.secret_access_key` is set, since either may
|
||||
come from a `${...}` reference rather than the file, and otherwise
|
||||
says the file is readable by others. `snapshot purge` against a
|
||||
destination store it could not list gave an error with
|
||||
`listing remote snapshots:` in it twice; the prefix now appears once.
|
||||
|
||||
- 2026-10-07: Made `s3.part_size` set the multipart upload part size
|
||||
([issue #232](https://git.eeqj.de/sneak/vaultik/issues/232)). It was
|
||||
loaded and defaulted but never passed to the S3 client, whose uploader
|
||||
used a fixed 10MiB part. It now reaches the uploader for `storage_url`
|
||||
and for the `s3.*` fields, and a part size S3 refuses, below 5MiB or
|
||||
above 5GiB, `0` included, fails at config load. A blob too large for
|
||||
S3's limit of 10,000 parts at the configured size is uploaded in larger
|
||||
parts. The docs gave the default as `5MB`, which the config file reads
|
||||
as 5,000,000 bytes, below the minimum; they now say `5MiB`.
|
||||
|
||||
- 2026-10-07: Made per-name retention work when the hostname contains `_`
|
||||
([issue #230](https://git.eeqj.de/sneak/vaultik/issues/230)). A
|
||||
snapshot ID is `hostname_name_timestamp`, and the name was read as
|
||||
everything between the first and the last `_`, so with
|
||||
`hostname: my_host` the name `home` came out as `host_home`.
|
||||
`snapshot purge --keep-latest --snapshot home` then printed "No
|
||||
snapshots to delete", and `snapshot create --prune` purged nothing
|
||||
without a message. The name is now read using the hostname the
|
||||
`snapshots` table stores with each snapshot, cut at its first `.` as it
|
||||
is in the ID.
|
||||
|
||||
- 2026-10-07: Made `remote info` stop reporting a snapshot's blobs as
|
||||
orphaned when its manifest cannot be read, and stop printing raw
|
||||
names from under `metadata/`
|
||||
([issue #228](https://git.eeqj.de/sneak/vaultik/issues/228)). A
|
||||
manifest it failed to read was skipped, so that snapshot's blobs were
|
||||
counted as orphaned and the report advised running `vaultik prune`.
|
||||
The orphan figures are now unknown in that case, with no prune
|
||||
advice, and `--json` gives them as `null` with the unreadable remote
|
||||
keys in `unreadable_manifests`. A name under `metadata/` that is not
|
||||
64 lowercase hex characters is now skipped with a warning instead of
|
||||
being printed, control characters included. A manifest under a
|
||||
skipped name is then not read either, so it also leaves the orphan
|
||||
figures unknown, and `--json` counts such manifests in
|
||||
`skipped_manifest_count`. A directory with no manifest in it, as left
|
||||
by an interrupted backup, leaves the figures known.
|
||||
|
||||
- 2026-10-07: Made `config set` keep a string that looks like a number
|
||||
([issue #229](https://git.eeqj.de/sneak/vaultik/issues/229)). It wrote
|
||||
every value unquoted, and `config.Load` reads the file through untyped
|
||||
YAML, so an access key `00112233` loaded as `38043` and a hostname `007`
|
||||
as `7`. A value for a string setting in `config.Config` is now tagged as
|
||||
a YAML string, which the file quotes wherever YAML would read a number or
|
||||
a boolean; other settings are still written unquoted.
|
||||
|
||||
- 2026-10-07: Made a backup notice a file rewritten with its size
|
||||
unchanged and a new mtime in the same second as the one in the index
|
||||
([issue #226](https://git.eeqj.de/sneak/vaultik/issues/226)). The
|
||||
`files` table held mtime in whole seconds and the scanner compared
|
||||
whole seconds, so every later snapshot kept the old content. A new
|
||||
`mtime_nsec` column holds the nanoseconds within the second that
|
||||
`mtime` holds, and the scanner compares the full mtime. A local index
|
||||
created before the change lacks the column and is rebuilt with
|
||||
`vaultik database delete` and a full backup.
|
||||
|
||||
- 2026-10-07: Made taking the process-wide lock atomic
|
||||
([issue #227](https://git.eeqj.de/sneak/vaultik/issues/227)). The lock
|
||||
read `vaultik.pid`, checked whether that PID was alive and then wrote
|
||||
its own, so two writers started together could both pass the check and
|
||||
both run. It is now an `flock` on `vaultik.pid`, held until the run
|
||||
ends; the kernel drops it when the process exits, so a crash leaves no
|
||||
lock behind. A clean exit now empties the file instead of deleting it,
|
||||
because deleting it would let two later runs each lock a different
|
||||
file.
|
||||
|
||||
- 2026-10-06: Made `snapshot remove --json` write only its document to
|
||||
stdout when the destination store cannot be reached
|
||||
([issue #251](https://git.eeqj.de/sneak/vaultik/issues/251)). Its
|
||||
warning that the snapshot's metadata was left on the destination store
|
||||
went to stdout ahead of the document, breaking `| jq` on a command that
|
||||
exited 0. Under `--json` the warning now reaches stderr only, through
|
||||
the logger. The warning, the README and the command's help said
|
||||
`vaultik prune` would finish the cleanup, but `prune` never removes
|
||||
snapshot metadata; they now say to run `vaultik snapshot remove` for the
|
||||
snapshot again once the destination store is reachable.
|
||||
|
||||
- 2026-10-06: Made the backup summary and the `snapshots` row count each
|
||||
file, byte and upload once
|
||||
([issue #225](https://git.eeqj.de/sneak/vaultik/issues/225)). The
|
||||
scanner added a file's bytes again for each new chunk and counted a
|
||||
file as unchanged for each chunk already stored, so a first backup
|
||||
reported twice its size and "backed up" could go negative. Upload
|
||||
figures came from the progress reporter, which `--cron` turns off, and
|
||||
`blob_count` counted earlier paths' blobs again for each later path.
|
||||
The scanner now counts uploads itself; `blob_size`,
|
||||
`blob_uncompressed_size` and `compression_ratio` describe the blobs
|
||||
the snapshot references, and `docs/DATAMODEL.md` now says
|
||||
`chunk_count` and `blob_count` count what the run added.
|
||||
|
||||
- 2026-10-06: Made command output follow the README's stdout and stderr
|
||||
rules ([issue #224](https://git.eeqj.de/sneak/vaultik/issues/224)). The
|
||||
startup banner went to stdout, so a `completion` script or a
|
||||
@@ -145,14 +31,6 @@ the tag exists and is exercised; what is left is merging `next` to
|
||||
`snapshot verify --quiet` printed its whole report; it now prints
|
||||
none, and a failure still reaches stderr with the same exit status.
|
||||
|
||||
- 2026-10-06: Made a backup without `--cron` of a snapshot with two or
|
||||
more `paths` complete instead of panicking with `close of closed
|
||||
channel` ([issue #253](https://git.eeqj.de/sneak/vaultik/issues/253)).
|
||||
`Scan` runs once per path and started and stopped the progress
|
||||
reporter each time, and a second stop panics. The reporter is now
|
||||
started and stopped once per snapshot, around the scans of all its
|
||||
paths.
|
||||
|
||||
- 2026-10-06: Made a restore path argument select only that path and
|
||||
what is beneath it
|
||||
([issue #223](https://git.eeqj.de/sneak/vaultik/issues/223)). The
|
||||
|
||||
+6
-14
@@ -287,18 +287,15 @@ storage_url: "rclone://myremote/path/to/backups"
|
||||
# #use_ssl: true
|
||||
#
|
||||
# # Part size for multipart uploads
|
||||
# # Minimum 5MiB, maximum 5GiB; affects memory usage during upload
|
||||
# # A blob too large for 10,000 parts of this size gets larger parts
|
||||
# # Supports: 10MB, 16MiB, 100MiB, etc. (5MB is below the minimum)
|
||||
# # Default: 5MiB
|
||||
# #part_size: 5MiB
|
||||
# # Minimum 5MB, affects memory usage during upload
|
||||
# # Supports: 5MB, 10M, 100MiB, etc.
|
||||
# # Default: 5MB
|
||||
# #part_size: 5MB
|
||||
|
||||
# Path to local SQLite index database
|
||||
# This database tracks file state for incremental backups
|
||||
# Default: the platform data directory, e.g.
|
||||
# macOS: ~/Library/Application Support/vaultik/index.sqlite
|
||||
# Linux: ~/.local/share/vaultik/index.sqlite
|
||||
#index_path: /path/to/index.sqlite
|
||||
# Default: /var/lib/vaultik/index.sqlite
|
||||
#index_path: /var/lib/vaultik/index.sqlite
|
||||
|
||||
# Average chunk size for content-defined chunking
|
||||
# Smaller chunks = better deduplication but more metadata
|
||||
@@ -313,11 +310,6 @@ storage_url: "rclone://myremote/path/to/backups"
|
||||
# Chunking uses no secret (the FastCDC parameters are fixed and public). At a
|
||||
# large limit a blob holds hundreds of chunks, so individual chunk lengths are
|
||||
# not visible in its size; lowering the limit toward chunk_size exposes them.
|
||||
# Each blob is written in full to a temporary file in $TMPDIR (/tmp when
|
||||
# unset) before it is uploaded, and the metadata export at the end of a
|
||||
# backup works on a copy of the local index there. A backup needs free space
|
||||
# there of the larger of this limit and about three times the size of the
|
||||
# local index.
|
||||
# Supports: 1GB, 10G, 500MB, 1GiB, etc.
|
||||
# Default: 10GB
|
||||
#blob_size_limit: 10GB
|
||||
|
||||
+10
-10
@@ -36,8 +36,7 @@ Stores metadata about files in the filesystem being backed up.
|
||||
**Columns:**
|
||||
- `id` (TEXT PRIMARY KEY) - UUID for the file record
|
||||
- `path` (TEXT NOT NULL UNIQUE) - Absolute file path
|
||||
- `mtime` (INTEGER NOT NULL) - Modification time, whole seconds since the Unix epoch
|
||||
- `mtime_nsec` (INTEGER NOT NULL) - Nanoseconds within that second, 0 to 999999999
|
||||
- `mtime` (INTEGER NOT NULL) - Modification time as Unix timestamp
|
||||
- `size` (INTEGER NOT NULL) - File size in bytes
|
||||
- `mode` (INTEGER NOT NULL) - Unix file permissions and type
|
||||
- `uid` (INTEGER NOT NULL) - User ID of file owner
|
||||
@@ -111,17 +110,17 @@ Maps chunks to the blobs that contain them.
|
||||
Tracks backup snapshots.
|
||||
|
||||
**Columns:**
|
||||
- `id` (TEXT PRIMARY KEY) - Snapshot ID (format: `hostname_name_timestamp`, e.g. `server1_home_2025-06-01T12:00:00Z`: the hostname up to its first `.`, the snapshot name, and an RFC 3339 UTC timestamp)
|
||||
- `id` (TEXT PRIMARY KEY) - Snapshot ID (format: hostname-YYYYMMDD-HHMMSSZ)
|
||||
- `hostname` (TEXT) - Hostname where backup was created
|
||||
- `vaultik_version` (TEXT) - Version of Vaultik used
|
||||
- `vaultik_git_revision` (TEXT) - Git revision of Vaultik used
|
||||
- `started_at` (INTEGER) - Start timestamp
|
||||
- `completed_at` (INTEGER) - Completion timestamp (NULL if in progress)
|
||||
- `file_count` (INTEGER) - Number of files in snapshot
|
||||
- `chunk_count` (INTEGER) - Number of chunks this snapshot stored that were not stored before
|
||||
- `blob_count` (INTEGER) - Number of blobs this snapshot created
|
||||
- `chunk_count` (INTEGER) - Number of unique chunks
|
||||
- `blob_count` (INTEGER) - Number of blobs referenced
|
||||
- `total_size` (INTEGER) - Total size of all files
|
||||
- `blob_size` (INTEGER) - Total compressed size of all referenced blobs
|
||||
- `blob_size` (INTEGER) - Total size of all blobs (compressed)
|
||||
- `blob_uncompressed_size` (INTEGER) - Total uncompressed size of all referenced blobs
|
||||
- `compression_ratio` (REAL) - Compression ratio achieved
|
||||
- `compression_level` (INTEGER) - Compression level used for this snapshot
|
||||
@@ -218,8 +217,8 @@ The `{remote-key}` directory name is a one-way hash of the human snapshot ID, so
|
||||
### 4. Restore Process
|
||||
|
||||
The restore process doesn't use the local database. Instead:
|
||||
1. Downloads and decrypts the snapshot's metadata database (`db.zst.age`) from S3
|
||||
2. Downloads the blobs holding the chunks of the files being restored, found through that database's `blob_chunks` table; the manifest is not read
|
||||
1. Downloads snapshot metadata from S3
|
||||
2. Downloads required blobs based on manifest
|
||||
3. Reconstructs files from decrypted and decompressed chunks
|
||||
|
||||
### 5. Pruning
|
||||
@@ -232,8 +231,9 @@ The restore process doesn't use the local database. Instead:
|
||||
|
||||
Before each backup:
|
||||
1. Query incomplete snapshots (where `completed_at IS NULL`)
|
||||
2. Delete each one and all its associations, without checking S3 for its metadata
|
||||
3. Clean up orphaned files, chunks, and blobs
|
||||
2. Check if metadata exists in S3
|
||||
3. If no metadata, delete snapshot and all associations
|
||||
4. Clean up orphaned files, chunks, and blobs
|
||||
|
||||
## Repository Pattern
|
||||
|
||||
|
||||
+1
-56
@@ -7,14 +7,11 @@ import (
|
||||
"os"
|
||||
"os/exec"
|
||||
"path/filepath"
|
||||
"reflect"
|
||||
"strconv"
|
||||
"strings"
|
||||
"unicode/utf8"
|
||||
|
||||
"github.com/spf13/cobra"
|
||||
"gopkg.in/yaml.v3"
|
||||
"sneak.berlin/go/vaultik/internal/config"
|
||||
"sneak.berlin/go/vaultik/internal/ui"
|
||||
)
|
||||
|
||||
@@ -34,9 +31,6 @@ const configDirMode = 0o755
|
||||
// yaml.Marshal's 4-space default.
|
||||
const configYAMLIndent = 2
|
||||
|
||||
// yamlStringTag is YAML's tag for a string scalar.
|
||||
const yamlStringTag = "!!str"
|
||||
|
||||
var (
|
||||
errConfigExists = errors.New("config file already exists")
|
||||
errEmptyConfig = errors.New("empty config file")
|
||||
@@ -205,7 +199,7 @@ storage_url: ""
|
||||
# access_key_id: YOUR_ACCESS_KEY
|
||||
# secret_access_key: YOUR_SECRET_KEY
|
||||
# # region: us-east-1 # Default: us-east-1
|
||||
# # part_size: 5MiB # Upload part size, 5MiB to 5GiB. Default: 5MiB
|
||||
# # part_size: 5MB # Multipart upload part size. Default: 5MB
|
||||
# # For the s3:// form, disable TLS with ?ssl=false in the URL, not use_ssl.
|
||||
|
||||
# ─── OPTIONAL ────────────────────────────────────────────────────────────────
|
||||
@@ -589,58 +583,9 @@ func yamlPathSet(root *yaml.Node, keys []string, value string) error {
|
||||
}
|
||||
}
|
||||
|
||||
// config.Load reads the file through untyped YAML, which turns an
|
||||
// unquoted 00112233 into the number 38043 and 1e5 into 100000. Tagging
|
||||
// a string setting as a string makes the encoder quote such a value.
|
||||
// Other settings stay unquoted, so compression_level 9 is a number.
|
||||
// The encoder refuses to write a value that is not valid UTF-8 as a
|
||||
// string. Left untagged, such a value is written as base64 !!binary and
|
||||
// loads back unchanged.
|
||||
if configKeyIsString(keys) && utf8.ValidString(value) {
|
||||
node.Tag = yamlStringTag
|
||||
}
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// configKeyIsString reports whether the dotted key names a string in
|
||||
// config.Config, following the fields' yaml tags, as s3.access_key_id and
|
||||
// snapshots.home.exclude.0 do.
|
||||
func configKeyIsString(keys []string) bool {
|
||||
typ := reflect.TypeFor[config.Config]()
|
||||
|
||||
for _, key := range keys {
|
||||
switch {
|
||||
case typ.Kind() == reflect.Map || typ.Kind() == reflect.Slice:
|
||||
// The key is a snapshot name or a list index.
|
||||
typ = typ.Elem()
|
||||
case typ.Kind() == reflect.Struct:
|
||||
field, ok := yamlField(typ, key)
|
||||
if !ok {
|
||||
return false
|
||||
}
|
||||
|
||||
typ = field.Type
|
||||
default:
|
||||
return false
|
||||
}
|
||||
}
|
||||
|
||||
return typ.Kind() == reflect.String
|
||||
}
|
||||
|
||||
// yamlField returns the field of struct type typ whose yaml tag names key.
|
||||
func yamlField(typ reflect.Type, key string) (reflect.StructField, bool) {
|
||||
for field := range typ.Fields() {
|
||||
name, _, _ := strings.Cut(field.Tag.Get("yaml"), ",")
|
||||
if name == key {
|
||||
return field, true
|
||||
}
|
||||
}
|
||||
|
||||
return reflect.StructField{}, false
|
||||
}
|
||||
|
||||
// yamlSetInMapping resolves (creating if needed) the value node for key
|
||||
// within a mapping node, setting it to value when it is the final path
|
||||
// element, and returns the node to descend into.
|
||||
|
||||
@@ -4,7 +4,6 @@ import (
|
||||
"bytes"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
@@ -95,110 +94,6 @@ func TestConfigSetRecipientOnFreshConfig(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestConfigSetStringLooksLikeNumber sets string settings to values that
|
||||
// YAML reads as numbers or booleans when they are unquoted, and checks that
|
||||
// config.Load returns each one unchanged.
|
||||
func TestConfigSetStringLooksLikeNumber(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
tests := []struct {
|
||||
key string
|
||||
value string
|
||||
field func(cfg *config.Config) string
|
||||
}{
|
||||
{"s3.access_key_id", "00112233",
|
||||
func(cfg *config.Config) string { return cfg.S3.AccessKeyID }},
|
||||
{"s3.secret_access_key", "12345678901234567890123456789012",
|
||||
func(cfg *config.Config) string { return cfg.S3.SecretAccessKey }},
|
||||
{"hostname", "007",
|
||||
func(cfg *config.Config) string { return cfg.Hostname }},
|
||||
{"s3.prefix", "1e5",
|
||||
func(cfg *config.Config) string { return cfg.S3.Prefix }},
|
||||
{"s3.bucket", "true",
|
||||
func(cfg *config.Config) string { return cfg.S3.Bucket }},
|
||||
{"s3.region", "FALSE",
|
||||
func(cfg *config.Config) string { return cfg.S3.Region }},
|
||||
{"snapshots.home.exclude.0", "1.10",
|
||||
func(cfg *config.Config) string { return cfg.Snapshots["home"].Exclude[0] }},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.key+"="+tt.value, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
cfg := loadAfterConfigSet(t, tt.key, tt.value)
|
||||
|
||||
got := tt.field(cfg)
|
||||
if got != tt.value {
|
||||
t.Errorf("%s = %q after config set %q", tt.key, got, tt.value)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestConfigSetNonUTF8Path checks that config set still accepts a value that
|
||||
// is not valid UTF-8, such as a path with a Latin-1 file name, and that
|
||||
// config.Load returns it unchanged.
|
||||
func TestConfigSetNonUTF8Path(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const dir = "/srv/caf\xe9"
|
||||
|
||||
cfg := loadAfterConfigSet(t, "snapshots.home.paths.0", dir)
|
||||
|
||||
got := cfg.Snapshots["home"].Paths[0]
|
||||
if got != dir {
|
||||
t.Errorf("snapshots.home.paths.0 = %q, want %q", got, dir)
|
||||
}
|
||||
}
|
||||
|
||||
// TestConfigSetNumberStaysNumber checks that a number set for an integer
|
||||
// setting is still read as a number, not as a quoted string.
|
||||
func TestConfigSetNumberStaysNumber(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const level = 9
|
||||
|
||||
cfg := loadAfterConfigSet(t, "compression_level", strconv.Itoa(level))
|
||||
|
||||
if cfg.CompressionLevel != level {
|
||||
t.Errorf("compression_level = %d, want %d", cfg.CompressionLevel, level)
|
||||
}
|
||||
}
|
||||
|
||||
// loadAfterConfigSet writes the file `config init` writes, sets storage_url
|
||||
// to a local directory so that the file passes validation, applies
|
||||
// `config set key value` and returns what config.Load reads back.
|
||||
func loadAfterConfigSet(t *testing.T, key, value string) *config.Config {
|
||||
t.Helper()
|
||||
|
||||
path := filepath.Join(t.TempDir(), "config.yml")
|
||||
|
||||
err := os.WriteFile(path, []byte(defaultConfigTemplate), configFileMode)
|
||||
if err != nil {
|
||||
t.Fatalf("write config: %v", err)
|
||||
}
|
||||
|
||||
out := ui.NewWithColor(&bytes.Buffer{}, false)
|
||||
|
||||
err = writeConfigSet(out, path, "storage_url", "file:///mnt/backups")
|
||||
if err != nil {
|
||||
t.Fatalf("config set storage_url: %v", err)
|
||||
}
|
||||
|
||||
err = writeConfigSet(out, path, key, value)
|
||||
if err != nil {
|
||||
t.Fatalf("config set %s: %v", key, err)
|
||||
}
|
||||
|
||||
cfg, err := config.Load(path)
|
||||
if err != nil {
|
||||
t.Fatalf("config.Load: %v", err)
|
||||
}
|
||||
|
||||
return cfg
|
||||
}
|
||||
|
||||
const testYAML = `# top comment
|
||||
compression_level: 3
|
||||
age_recipients:
|
||||
|
||||
@@ -2,9 +2,7 @@ package cli //nolint:testpackage // shares hermeticConfig and the capture helper
|
||||
|
||||
import (
|
||||
"context"
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"log/slog"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
@@ -89,45 +87,12 @@ func TestEntryJSONFailureIsReportedOnStderr(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestEntrySnapshotRemoveJSONWarningIsOnStderr runs `snapshot remove
|
||||
// --json` on a snapshot in the local index, against a destination
|
||||
// directory that does not exist. The command removes the snapshot from
|
||||
// the local index and still exits 0. Its stdout must hold the document
|
||||
// alone, with the warning about the destination store on stderr: the
|
||||
// command to run again once it is reachable, and the snapshot's ID in
|
||||
// the record's snapshot_id field.
|
||||
//
|
||||
//nolint:paralleltest // replaces os.Args, os.Stdout, os.Stderr and the xdg globals
|
||||
func TestEntrySnapshotRemoveJSONWarningIsOnStderr(t *testing.T) {
|
||||
configPath, indexPath := writeMissingDestinationConfig(t)
|
||||
seedStaleSnapshotRecord(t, indexPath)
|
||||
|
||||
code, stdout, stderr := runEntry(t, flagConfig, configPath,
|
||||
cmdSnapshot, cmdRemove, stalePruneSnapshotID, flagJSON)
|
||||
|
||||
require.Equal(t, 0, code)
|
||||
requireExactlyOneJSONDocument(t, stdout)
|
||||
|
||||
// stderr is a pipe here, so the logger writes one JSON record a line.
|
||||
var warning map[string]any
|
||||
|
||||
for line := range strings.Lines(stderr) {
|
||||
if strings.Contains(line,
|
||||
"Could not remove snapshot metadata from remote storage") {
|
||||
require.NoError(t, json.Unmarshal([]byte(line), &warning))
|
||||
}
|
||||
}
|
||||
|
||||
require.NotNil(t, warning, "the warning must reach stderr")
|
||||
assert.Contains(t, warning[slog.MessageKey],
|
||||
"run 'vaultik snapshot remove' with the snapshot's ID again")
|
||||
assert.Equal(t, stalePruneSnapshotID, warning["snapshot_id"])
|
||||
}
|
||||
|
||||
// writeMissingDestinationConfig builds a config whose destination
|
||||
// directory does not exist. Returns the config path and the path of
|
||||
// its local index, which is not created here.
|
||||
func writeMissingDestinationConfig(t *testing.T) (string, string) {
|
||||
// writeUnusableDestinationConfig builds a config whose destination
|
||||
// directory does not exist, which fails `remote info`, and whose local
|
||||
// index is bound to another destination, which fails `prune` and
|
||||
// `snapshot remove` (a missing destination alone only makes `snapshot
|
||||
// remove` warn). Returns the config path.
|
||||
func writeUnusableDestinationConfig(t *testing.T) string {
|
||||
t.Helper()
|
||||
|
||||
dir := t.TempDir()
|
||||
@@ -149,19 +114,6 @@ func writeMissingDestinationConfig(t *testing.T) (string, string) {
|
||||
xdg.Reload()
|
||||
t.Cleanup(xdg.Reload)
|
||||
|
||||
return configPath, indexPath
|
||||
}
|
||||
|
||||
// writeUnusableDestinationConfig builds a config whose destination
|
||||
// directory does not exist, which fails `remote info`, and whose local
|
||||
// index is bound to another destination, which fails `prune` and
|
||||
// `snapshot remove` (a missing destination alone only makes `snapshot
|
||||
// remove` warn). Returns the config path.
|
||||
func writeUnusableDestinationConfig(t *testing.T) string {
|
||||
t.Helper()
|
||||
|
||||
configPath, indexPath := writeMissingDestinationConfig(t)
|
||||
|
||||
ctx := context.Background()
|
||||
|
||||
db, err := database.New(ctx, indexPath)
|
||||
@@ -170,7 +122,7 @@ func writeUnusableDestinationConfig(t *testing.T) string {
|
||||
defer func() { require.NoError(t, db.Close()) }()
|
||||
|
||||
require.NoError(t, database.NewRepositories(db).LocalMeta.Set(ctx,
|
||||
database.LocalMetaKeyStorageURL, "file://"+t.TempDir()))
|
||||
database.LocalMetaKeyStorageURL, "file://"+filepath.Join(dir, "other")))
|
||||
|
||||
return configPath
|
||||
}
|
||||
|
||||
@@ -22,11 +22,9 @@ scans every snapshot manifest in the destination store, builds the
|
||||
set of still-referenced blob hashes, and deletes any blob not in that
|
||||
set.
|
||||
|
||||
Snapshot create --prune runs the same cleanup automatically; this
|
||||
command is the manual entry point for the same work (e.g. after a
|
||||
crashed backup or to reclaim storage). Snapshot remove leaves blobs in
|
||||
place; run this command afterwards to delete the ones no longer
|
||||
referenced.`,
|
||||
Snapshot create --prune and snapshot remove run the same cleanup
|
||||
automatically; this command is the manual entry point for the same
|
||||
work (e.g. after a crashed backup or to reclaim storage).`,
|
||||
Args: cobra.NoArgs,
|
||||
RunE: func(cmd *cobra.Command, _ []string) error {
|
||||
// Use unified config resolution
|
||||
|
||||
@@ -66,9 +66,8 @@ func newSnapshotCreateCommand() *cobra.Command {
|
||||
If snapshot names are provided, only those snapshots are created.
|
||||
If no names are provided, all configured snapshots are created.
|
||||
|
||||
The config is read from the path given by --config or VAULTIK_CONFIG;
|
||||
otherwise from the platform config directory (~/.config/vaultik/config.yml
|
||||
on Linux), then /etc/vaultik/config.yml.`,
|
||||
Config is located at /etc/vaultik/config.yml by default, but can be overridden by
|
||||
specifying a path using --config or by setting VAULTIK_CONFIG to a path.`,
|
||||
Args: cobra.ArbitraryArgs,
|
||||
RunE: func(cmd *cobra.Command, args []string) error {
|
||||
// Pass snapshot names from args
|
||||
@@ -259,9 +258,8 @@ Use --local-only to skip the remote half (e.g. when you want to forget a
|
||||
snapshot locally without touching the destination store).
|
||||
|
||||
If the remote is unreachable, the local-database removal still completes
|
||||
and a warning is emitted; run 'vaultik snapshot remove <snapshot-id>' again
|
||||
once the destination store is reachable to remove the snapshot's metadata
|
||||
from it ('vaultik prune' does not).
|
||||
and a warning is emitted; rerun 'vaultik prune' once the destination store
|
||||
is reachable to finish remote cleanup.
|
||||
|
||||
To wipe the entire destination store and start over, use 'vaultik remote
|
||||
nuke --force' — it is the single supported entry point for that.`,
|
||||
|
||||
@@ -33,14 +33,11 @@ const secretKeyPrefix = "AGE-SECRET-KEY-"
|
||||
const (
|
||||
defaultBlobSizeLimit = Size(10 * 1024 * 1024 * 1024) // 10GB
|
||||
defaultChunkSize = Size(10 * 1024 * 1024) // 10MB
|
||||
defaultS3PartSize = Size(5 * 1024 * 1024) // 5MiB
|
||||
defaultS3PartSize = Size(5 * 1024 * 1024) // 5MB
|
||||
defaultCompressionLevel = 3
|
||||
minChunkSize = 1024 * 1024 // 1MB
|
||||
minCompressionLevel = 1
|
||||
maxCompressionLevel = 19
|
||||
// S3 accepts a multipart upload part from 5MiB to 5GiB.
|
||||
minS3PartSize = 5 * 1024 * 1024
|
||||
maxS3PartSize = 5 * 1024 * 1024 * 1024
|
||||
)
|
||||
|
||||
// Sentinel validation errors.
|
||||
@@ -58,7 +55,6 @@ var (
|
||||
"blob_size_limit must be at least the largest chunk the chunker can " +
|
||||
"emit (chunk_size times the FastCDC size spread)")
|
||||
errBadCompression = errors.New("compression_level must be between 1 and 19")
|
||||
errBadS3PartSize = errors.New("s3.part_size must be between 5MiB and 5GiB")
|
||||
errBadStorageScheme = errors.New(
|
||||
"storage_url must start with s3://, file://, or rclone://")
|
||||
errStorageNotConfigured = errors.New(
|
||||
@@ -248,7 +244,6 @@ func Load(path string) (*Config, error) {
|
||||
ChunkSize: defaultChunkSize,
|
||||
IndexPath: filepath.Join(xdg.DataHome, appName, "index.sqlite"),
|
||||
CompressionLevel: defaultCompressionLevel,
|
||||
S3: S3Config{PartSize: defaultS3PartSize},
|
||||
}
|
||||
|
||||
// Convert smartconfig data to YAML then unmarshal
|
||||
@@ -299,13 +294,17 @@ func Load(path string) (*Config, error) {
|
||||
cfg.S3.Region = "us-east-1"
|
||||
}
|
||||
|
||||
if cfg.S3.PartSize == 0 {
|
||||
cfg.S3.PartSize = defaultS3PartSize
|
||||
}
|
||||
|
||||
// Check config file permissions (warn if world or group readable)
|
||||
//nolint:gosec // G703: config path is operator-supplied by design
|
||||
info, statErr := os.Stat(path)
|
||||
if statErr == nil {
|
||||
mode := info.Mode().Perm()
|
||||
if mode&0044 != 0 { // group or world readable
|
||||
log.Warn(cfg.readableByOthersWarning(),
|
||||
log.Warn("Config file has insecure permissions (contains S3 credentials)",
|
||||
"path", path,
|
||||
"mode", fmt.Sprintf("%04o", mode),
|
||||
"recommendation", "chmod 600 "+path)
|
||||
@@ -333,7 +332,6 @@ func Load(path string) (*Config, error) {
|
||||
// (chunk_size times chunker.ChunkSizeSpread), so a single-chunk blob never
|
||||
// exceeds the configured limit
|
||||
// - Compression level must be between 1 and 19
|
||||
// - S3 part size must be between 5MiB and 5GiB, the part sizes S3 accepts
|
||||
//
|
||||
// Returns an error describing the first validation failure encountered.
|
||||
func (c *Config) Validate() error {
|
||||
@@ -378,11 +376,6 @@ func (c *Config) Validate() error {
|
||||
return errBadCompression
|
||||
}
|
||||
|
||||
if c.S3.PartSize.Int64() < minS3PartSize ||
|
||||
c.S3.PartSize.Int64() > maxS3PartSize {
|
||||
return errBadS3PartSize
|
||||
}
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
@@ -418,18 +411,6 @@ func (c *Config) setAgeSecretKey() {
|
||||
}
|
||||
}
|
||||
|
||||
// readableByOthersWarning is the warning Load logs when others can read
|
||||
// the config file. It says "may contain" because the S3 credentials are
|
||||
// seen only after smartconfig has replaced any ${...} reference in the
|
||||
// file with its value, so a set credential need not be in the file.
|
||||
func (c *Config) readableByOthersWarning() string {
|
||||
if c.S3.AccessKeyID != "" || c.S3.SecretAccessKey != "" {
|
||||
return "Config file is readable by others and may contain S3 credentials"
|
||||
}
|
||||
|
||||
return "Config file is readable by others"
|
||||
}
|
||||
|
||||
// validateStorage validates storage configuration.
|
||||
// If StorageURL is set, it takes precedence. S3 URLs require credentials.
|
||||
// File URLs don't require any S3 configuration.
|
||||
|
||||
@@ -8,7 +8,6 @@ import (
|
||||
"testing"
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/chunker"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
)
|
||||
|
||||
const (
|
||||
@@ -167,7 +166,6 @@ func TestValidateBlobSizeLimit(t *testing.T) {
|
||||
ChunkSize: chunkSize,
|
||||
BlobSizeLimit: blobLimit,
|
||||
CompressionLevel: 3,
|
||||
S3: S3Config{PartSize: defaultS3PartSize},
|
||||
}
|
||||
}
|
||||
|
||||
@@ -223,120 +221,6 @@ func TestValidateBlobSizeLimit(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestValidateS3PartSize checks that s3.part_size is held to the part sizes
|
||||
// S3 accepts, 5MiB to 5GiB, by changing only the part size of the test
|
||||
// config. "5MB" in the config file is 5,000,000 bytes, below the minimum.
|
||||
func TestValidateS3PartSize(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
base, err := Load(os.Getenv("VAULTIK_CONFIG"))
|
||||
if err != nil {
|
||||
t.Fatalf("Failed to load config: %v", err)
|
||||
}
|
||||
|
||||
tests := []struct {
|
||||
name string
|
||||
partSize Size
|
||||
wantErr bool
|
||||
}{
|
||||
{
|
||||
name: "5MB is rejected",
|
||||
partSize: 5_000_000,
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "one byte below 5MiB is rejected",
|
||||
partSize: minS3PartSize - 1,
|
||||
wantErr: true,
|
||||
},
|
||||
{
|
||||
name: "5MiB is accepted",
|
||||
partSize: minS3PartSize,
|
||||
wantErr: false,
|
||||
},
|
||||
{
|
||||
name: "5GiB is accepted",
|
||||
partSize: maxS3PartSize,
|
||||
wantErr: false,
|
||||
},
|
||||
{
|
||||
name: "one byte above 5GiB is rejected",
|
||||
partSize: maxS3PartSize + 1,
|
||||
wantErr: true,
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
cfg := *base
|
||||
cfg.S3.PartSize = tt.partSize
|
||||
|
||||
err := cfg.Validate()
|
||||
if tt.wantErr {
|
||||
if !errors.Is(err, errBadS3PartSize) {
|
||||
t.Fatalf("Validate() error = %v, want errBadS3PartSize", err)
|
||||
}
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
if err != nil {
|
||||
t.Fatalf("Validate() unexpected error: %v", err)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// TestLoadS3PartSize checks that a config file without s3.part_size loads
|
||||
// with the 5MiB default, and that an explicit 0 fails at load like any other
|
||||
// part size S3 refuses.
|
||||
func TestLoadS3PartSize(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const withoutPartSize = "snapshots:\n" +
|
||||
" test:\n" +
|
||||
" paths: [/tmp/vaultik-test-source]\n" +
|
||||
"storage_url: file:///tmp/vaultik-test-storage\n"
|
||||
|
||||
writeConfig := func(t *testing.T, text string) string {
|
||||
t.Helper()
|
||||
|
||||
path := filepath.Join(t.TempDir(), "config.yml")
|
||||
|
||||
err := os.WriteFile(path, []byte(text), 0o600)
|
||||
if err != nil {
|
||||
t.Fatalf("write config: %v", err)
|
||||
}
|
||||
|
||||
return path
|
||||
}
|
||||
|
||||
t.Run("absent loads as 5MiB", func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
cfg, err := Load(writeConfig(t, withoutPartSize))
|
||||
if err != nil {
|
||||
t.Fatalf("Load() unexpected error: %v", err)
|
||||
}
|
||||
|
||||
if cfg.S3.PartSize != defaultS3PartSize {
|
||||
t.Errorf("s3.part_size = %d, want %d",
|
||||
cfg.S3.PartSize, defaultS3PartSize)
|
||||
}
|
||||
})
|
||||
|
||||
t.Run("0 is rejected", func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
_, err := Load(writeConfig(t, withoutPartSize+"s3:\n part_size: 0\n"))
|
||||
if !errors.Is(err, errBadS3PartSize) {
|
||||
t.Fatalf("Load() error = %v, want errBadS3PartSize", err)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
// TestValidateAgeRecipients checks that recipients are parsed at config load
|
||||
// (a bad entry fails immediately, not mid-backup) and that no invalid entry —
|
||||
// least of all a pasted secret key — is echoed in the error. An empty list
|
||||
@@ -352,7 +236,6 @@ func TestValidateAgeRecipients(t *testing.T) {
|
||||
ChunkSize: Size(10 * 1024 * 1024),
|
||||
BlobSizeLimit: Size(10 * 1024 * 1024 * 1024),
|
||||
CompressionLevel: 3,
|
||||
S3: S3Config{PartSize: defaultS3PartSize},
|
||||
}
|
||||
}
|
||||
|
||||
@@ -460,125 +343,3 @@ func TestAgeSecretKeySourceName(t *testing.T) {
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// loadReadableConfig writes configYAML to a file that others can read,
|
||||
// loads it, and returns what the logger wrote to stderr meanwhile. The
|
||||
// logger writes to the os.Stderr it finds when it is initialized, so
|
||||
// os.Stderr is pointed at a file first. Not parallel-safe: os.Stderr and
|
||||
// the logger are process-global.
|
||||
func loadReadableConfig(t *testing.T, configYAML string) string {
|
||||
t.Helper()
|
||||
|
||||
dir := t.TempDir()
|
||||
configPath := filepath.Join(dir, "config.yml")
|
||||
stderrPath := filepath.Join(dir, "stderr")
|
||||
|
||||
err := os.WriteFile(configPath, []byte(configYAML), 0o600)
|
||||
if err != nil {
|
||||
t.Fatalf("writing config: %v", err)
|
||||
}
|
||||
|
||||
//nolint:gosec // G302: the test needs a config file others can read
|
||||
err = os.Chmod(configPath, 0o644)
|
||||
if err != nil {
|
||||
t.Fatalf("chmod config: %v", err)
|
||||
}
|
||||
|
||||
stderrFile, err := os.Create(stderrPath) //nolint:gosec // G304: test temp path
|
||||
if err != nil {
|
||||
t.Fatalf("creating stderr file: %v", err)
|
||||
}
|
||||
|
||||
previous := os.Stderr
|
||||
os.Stderr = stderrFile
|
||||
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
_, loadErr := Load(configPath)
|
||||
|
||||
os.Stderr = previous
|
||||
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
_ = stderrFile.Close()
|
||||
|
||||
if loadErr != nil {
|
||||
t.Fatalf("Load() error = %v", loadErr)
|
||||
}
|
||||
|
||||
captured, err := os.ReadFile(stderrPath) //nolint:gosec // G304: test temp path
|
||||
if err != nil {
|
||||
t.Fatalf("reading stderr file: %v", err)
|
||||
}
|
||||
|
||||
return string(captured)
|
||||
}
|
||||
|
||||
// TestLoadWarnsReadableConfigWithoutS3Credentials checks that a config
|
||||
// file others can read, holding no S3 credentials, is warned about
|
||||
// without a claim that it holds them.
|
||||
//
|
||||
//nolint:paralleltest // loadReadableConfig replaces os.Stderr
|
||||
func TestLoadWarnsReadableConfigWithoutS3Credentials(t *testing.T) {
|
||||
stderr := loadReadableConfig(t, `
|
||||
storage_url: file:///var/backups/vaultik
|
||||
snapshots:
|
||||
home:
|
||||
paths:
|
||||
- /home
|
||||
`)
|
||||
|
||||
if !strings.Contains(stderr, "Config file is readable by others") {
|
||||
t.Errorf("expected a warning that the file is readable by others, got %q",
|
||||
stderr)
|
||||
}
|
||||
|
||||
if strings.Contains(stderr, "S3 credentials") {
|
||||
t.Errorf("warning names S3 credentials the file does not set: %q", stderr)
|
||||
}
|
||||
}
|
||||
|
||||
// TestLoadWarnsReadableConfigWithS3Credentials checks that a config file
|
||||
// others can read and that sets S3 credentials, as values or as ${ENV:...}
|
||||
// references, is warned about as one that may contain them.
|
||||
//
|
||||
//nolint:paralleltest // loadReadableConfig replaces os.Stderr
|
||||
func TestLoadWarnsReadableConfigWithS3Credentials(t *testing.T) {
|
||||
t.Setenv("VAULTIK_TEST_ACCESS_KEY_ID", "test-access-key")
|
||||
t.Setenv("VAULTIK_TEST_SECRET_ACCESS_KEY", "test-secret-key")
|
||||
|
||||
configs := map[string]string{
|
||||
"values": `
|
||||
storage_url: s3://bucket/prefix?endpoint=s3.example.com
|
||||
s3:
|
||||
access_key_id: test-access-key
|
||||
secret_access_key: test-secret-key
|
||||
snapshots:
|
||||
home:
|
||||
paths:
|
||||
- /home
|
||||
`,
|
||||
"references": `
|
||||
storage_url: s3://bucket/prefix?endpoint=s3.example.com
|
||||
s3:
|
||||
access_key_id: ${ENV:VAULTIK_TEST_ACCESS_KEY_ID}
|
||||
secret_access_key: ${ENV:VAULTIK_TEST_SECRET_ACCESS_KEY}
|
||||
snapshots:
|
||||
home:
|
||||
paths:
|
||||
- /home
|
||||
`,
|
||||
}
|
||||
|
||||
for name, configYAML := range configs {
|
||||
t.Run(name, func(t *testing.T) {
|
||||
stderr := loadReadableConfig(t, configYAML)
|
||||
|
||||
if !strings.Contains(stderr,
|
||||
"Config file is readable by others and may contain S3 credentials") {
|
||||
t.Errorf("expected a warning naming the S3 credentials, got %q",
|
||||
stderr)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
+19
-28
@@ -33,13 +33,11 @@ func (r *FileRepository) Create(ctx context.Context, tx *sql.Tx, file *File) err
|
||||
}
|
||||
|
||||
query := `
|
||||
INSERT INTO files
|
||||
(id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
INSERT INTO files (id, path, source_path, mtime, size, mode, uid, gid, link_target)
|
||||
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?)
|
||||
ON CONFLICT(path) DO UPDATE SET
|
||||
source_path = excluded.source_path,
|
||||
mtime = excluded.mtime,
|
||||
mtime_nsec = excluded.mtime_nsec,
|
||||
size = excluded.size,
|
||||
mode = excluded.mode,
|
||||
uid = excluded.uid,
|
||||
@@ -56,19 +54,16 @@ func (r *FileRepository) Create(ctx context.Context, tx *sql.Tx, file *File) err
|
||||
if tx != nil {
|
||||
LogSQL("Execute", query,
|
||||
file.ID.String(), file.Path.String(), file.SourcePath.String(),
|
||||
file.MTime.Unix(), file.MTime.Nanosecond(),
|
||||
file.Size, file.Mode, file.UID, file.GID,
|
||||
file.MTime.Unix(), file.Size, file.Mode, file.UID, file.GID,
|
||||
file.LinkTarget.String())
|
||||
err = tx.QueryRowContext(ctx, query,
|
||||
file.ID.String(), file.Path.String(), file.SourcePath.String(),
|
||||
file.MTime.Unix(), file.MTime.Nanosecond(),
|
||||
file.Size, file.Mode, file.UID, file.GID,
|
||||
file.MTime.Unix(), file.Size, file.Mode, file.UID, file.GID,
|
||||
file.LinkTarget.String()).Scan(&idStr)
|
||||
} else {
|
||||
err = r.db.QueryRowWithLog(ctx, query,
|
||||
file.ID.String(), file.Path.String(), file.SourcePath.String(),
|
||||
file.MTime.Unix(), file.MTime.Nanosecond(),
|
||||
file.Size, file.Mode, file.UID, file.GID,
|
||||
file.MTime.Unix(), file.Size, file.Mode, file.UID, file.GID,
|
||||
file.LinkTarget.String()).Scan(&idStr)
|
||||
}
|
||||
|
||||
@@ -89,7 +84,7 @@ func (r *FileRepository) Create(ctx context.Context, tx *sql.Tx, file *File) err
|
||||
// in the index.
|
||||
func (r *FileRepository) GetByPath(ctx context.Context, path string) (*File, error) {
|
||||
query := `
|
||||
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
|
||||
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
|
||||
FROM files
|
||||
WHERE path = ?
|
||||
`
|
||||
@@ -109,7 +104,7 @@ func (r *FileRepository) GetByPath(ctx context.Context, path string) (*File, err
|
||||
// GetByID retrieves a file by its UUID
|
||||
func (r *FileRepository) GetByID(ctx context.Context, id types.FileID) (*File, error) {
|
||||
query := `
|
||||
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
|
||||
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
|
||||
FROM files
|
||||
WHERE id = ?
|
||||
`
|
||||
@@ -132,7 +127,7 @@ func (r *FileRepository) GetByPathTx(
|
||||
ctx context.Context, tx *sql.Tx, path string,
|
||||
) (*File, error) {
|
||||
query := `
|
||||
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
|
||||
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
|
||||
FROM files
|
||||
WHERE path = ?
|
||||
`
|
||||
@@ -163,14 +158,13 @@ func (r *FileRepository) ListModifiedSince(
|
||||
ctx context.Context, since time.Time,
|
||||
) ([]*File, error) {
|
||||
query := `
|
||||
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
|
||||
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
|
||||
FROM files
|
||||
WHERE (mtime, mtime_nsec) >= (?, ?)
|
||||
WHERE mtime >= ?
|
||||
ORDER BY path
|
||||
`
|
||||
|
||||
rows, err := r.db.conn.QueryContext(ctx, query,
|
||||
since.Unix(), since.Nanosecond())
|
||||
rows, err := r.db.conn.QueryContext(ctx, query, since.Unix())
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("querying files: %w", err)
|
||||
}
|
||||
@@ -245,7 +239,7 @@ func (r *FileRepository) ListUnderPath(
|
||||
|
||||
// LIKE would ignore ASCII case and treat _ and % in path as wildcards.
|
||||
query := `
|
||||
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
|
||||
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
|
||||
FROM files
|
||||
WHERE path = ? OR substr(path, 1, length(?)) = ?
|
||||
ORDER BY path
|
||||
@@ -330,7 +324,7 @@ func (r *FileRepository) ListIDsWithChunksNotInUploadedBlobs(
|
||||
// ListAll returns all files in the database
|
||||
func (r *FileRepository) ListAll(ctx context.Context) ([]*File, error) {
|
||||
query := `
|
||||
SELECT id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target
|
||||
SELECT id, path, source_path, mtime, size, mode, uid, gid, link_target
|
||||
FROM files
|
||||
ORDER BY path
|
||||
`
|
||||
@@ -371,7 +365,7 @@ func (r *FileRepository) CreateBatch(
|
||||
}
|
||||
|
||||
// Each files row binds this many SQL variables.
|
||||
const fileCols = 10
|
||||
const fileCols = 9
|
||||
|
||||
// Batch at 100 rows to be safe with SQLite's variable limit.
|
||||
const batchSize = 100
|
||||
@@ -382,7 +376,7 @@ func (r *FileRepository) CreateBatch(
|
||||
batch := files[i:end]
|
||||
|
||||
query := `INSERT INTO files
|
||||
(id, path, source_path, mtime, mtime_nsec, size, mode, uid, gid, link_target)
|
||||
(id, path, source_path, mtime, size, mode, uid, gid, link_target)
|
||||
VALUES `
|
||||
|
||||
args := make([]any, 0, len(batch)*fileCols)
|
||||
@@ -394,12 +388,11 @@ func (r *FileRepository) CreateBatch(
|
||||
querySb325.WriteString(", ")
|
||||
}
|
||||
|
||||
querySb325.WriteString("(?, ?, ?, ?, ?, ?, ?, ?, ?, ?)")
|
||||
querySb325.WriteString("(?, ?, ?, ?, ?, ?, ?, ?, ?)")
|
||||
|
||||
args = append(args,
|
||||
f.ID.String(), f.Path.String(), f.SourcePath.String(),
|
||||
f.MTime.Unix(), f.MTime.Nanosecond(),
|
||||
f.Size, f.Mode, f.UID, f.GID,
|
||||
f.MTime.Unix(), f.Size, f.Mode, f.UID, f.GID,
|
||||
f.LinkTarget.String())
|
||||
}
|
||||
|
||||
@@ -408,7 +401,6 @@ func (r *FileRepository) CreateBatch(
|
||||
query += ` ON CONFLICT(path) DO UPDATE SET
|
||||
source_path = excluded.source_path,
|
||||
mtime = excluded.mtime,
|
||||
mtime_nsec = excluded.mtime_nsec,
|
||||
size = excluded.size,
|
||||
mode = excluded.mode,
|
||||
uid = excluded.uid,
|
||||
@@ -468,7 +460,7 @@ func (r *FileRepository) scanFileFrom(row fileRowScanner) (*File, error) {
|
||||
var (
|
||||
file File
|
||||
idStr, pathStr, sourcePathStr string
|
||||
mtimeUnix, mtimeNsec int64
|
||||
mtimeUnix int64
|
||||
linkTarget sql.NullString
|
||||
)
|
||||
|
||||
@@ -477,7 +469,6 @@ func (r *FileRepository) scanFileFrom(row fileRowScanner) (*File, error) {
|
||||
&pathStr,
|
||||
&sourcePathStr,
|
||||
&mtimeUnix,
|
||||
&mtimeNsec,
|
||||
&file.Size,
|
||||
&file.Mode,
|
||||
&file.UID,
|
||||
@@ -496,7 +487,7 @@ func (r *FileRepository) scanFileFrom(row fileRowScanner) (*File, error) {
|
||||
file.Path = types.FilePath(pathStr)
|
||||
file.SourcePath = types.SourcePath(sourcePathStr)
|
||||
|
||||
file.MTime = time.Unix(mtimeUnix, mtimeNsec).UTC()
|
||||
file.MTime = time.Unix(mtimeUnix, 0).UTC()
|
||||
if linkTarget.Valid {
|
||||
file.LinkTarget = types.FilePath(linkTarget.String)
|
||||
}
|
||||
|
||||
@@ -252,115 +252,6 @@ func TestFileRepositorySymlink(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// An mtime after 2262 or before 1678 does not fit in int64 nanoseconds
|
||||
// since the epoch, and must still come back from the database unchanged.
|
||||
func TestFileRepositoryMTimeOutsideInt64NanosecondRange(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
db, cleanup := setupTestDB(t)
|
||||
defer cleanup()
|
||||
|
||||
ctx := context.Background()
|
||||
repo := database.NewFileRepository(db)
|
||||
|
||||
mtimes := []time.Time{
|
||||
time.Date(2300, time.January, 1, 0, 0, 0, 123456789, time.UTC),
|
||||
time.Date(1601, time.January, 1, 0, 0, 0, 987654321, time.UTC),
|
||||
}
|
||||
|
||||
for _, mtime := range mtimes {
|
||||
created := &database.File{
|
||||
Path: types.FilePath("/created-" + mtime.Format(time.RFC3339Nano)),
|
||||
MTime: mtime,
|
||||
}
|
||||
|
||||
err := repo.Create(ctx, nil, created)
|
||||
if err != nil {
|
||||
t.Fatalf("failed to create file: %v", err)
|
||||
}
|
||||
|
||||
batched := &database.File{
|
||||
ID: types.NewFileID(),
|
||||
Path: types.FilePath("/batched-" + mtime.Format(time.RFC3339Nano)),
|
||||
MTime: mtime,
|
||||
}
|
||||
|
||||
err = repo.CreateBatch(ctx, nil, []*database.File{batched})
|
||||
if err != nil {
|
||||
t.Fatalf("failed to batch create file: %v", err)
|
||||
}
|
||||
|
||||
for _, path := range []types.FilePath{created.Path, batched.Path} {
|
||||
retrieved, err := repo.GetByPath(ctx, path.String())
|
||||
if err != nil {
|
||||
t.Fatalf("failed to get file: %v", err)
|
||||
}
|
||||
|
||||
if !retrieved.MTime.Equal(mtime) {
|
||||
t.Errorf("%s: mtime got %v, want %v",
|
||||
path, retrieved.MTime, mtime)
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// A file already in the index and rewritten within the same second must get
|
||||
// its new nanoseconds stored, through both Create and CreateBatch.
|
||||
func TestFileRepositoryUpsertMTimeInSameSecond(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
db, cleanup := setupTestDB(t)
|
||||
defer cleanup()
|
||||
|
||||
ctx := context.Background()
|
||||
repo := database.NewFileRepository(db)
|
||||
|
||||
indexed := time.Date(2026, time.October, 7, 12, 0, 0, 100000000, time.UTC)
|
||||
rewritten := indexed.Add(800 * time.Millisecond)
|
||||
|
||||
tests := []struct {
|
||||
name string
|
||||
upsert func(file *database.File) error
|
||||
}{
|
||||
{"Create", func(file *database.File) error {
|
||||
return repo.Create(ctx, nil, file)
|
||||
}},
|
||||
{"CreateBatch", func(file *database.File) error {
|
||||
return repo.CreateBatch(ctx, nil, []*database.File{file})
|
||||
}},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
file := &database.File{
|
||||
ID: types.NewFileID(),
|
||||
Path: types.FilePath("/" + tt.name),
|
||||
MTime: indexed,
|
||||
}
|
||||
|
||||
err := tt.upsert(file)
|
||||
if err != nil {
|
||||
t.Fatalf("%s: failed to create file: %v", tt.name, err)
|
||||
}
|
||||
|
||||
file.MTime = rewritten
|
||||
|
||||
err = tt.upsert(file)
|
||||
if err != nil {
|
||||
t.Fatalf("%s: failed to update file: %v", tt.name, err)
|
||||
}
|
||||
|
||||
retrieved, err := repo.GetByPath(ctx, file.Path.String())
|
||||
if err != nil {
|
||||
t.Fatalf("%s: failed to get file: %v", tt.name, err)
|
||||
}
|
||||
|
||||
if !retrieved.MTime.Equal(rewritten) {
|
||||
t.Errorf("%s: mtime got %v, want %v",
|
||||
tt.name, retrieved.MTime, rewritten)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
func TestFileRepositoryTransaction(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
|
||||
@@ -14,7 +14,8 @@ type File struct {
|
||||
ID types.FileID // UUID primary key
|
||||
Path types.FilePath // Absolute path of the file
|
||||
|
||||
// SourcePath is the source directory this file came from.
|
||||
// SourcePath is the source directory this file came from (used for
|
||||
// restore path stripping).
|
||||
SourcePath types.SourcePath
|
||||
MTime time.Time
|
||||
Size int64
|
||||
@@ -98,8 +99,8 @@ type Snapshot struct {
|
||||
StartedAt time.Time
|
||||
CompletedAt *time.Time // nil if still in progress
|
||||
FileCount int64
|
||||
ChunkCount int64 // Chunks this snapshot stored that were not stored before
|
||||
BlobCount int64 // Blobs this snapshot created
|
||||
ChunkCount int64
|
||||
BlobCount int64
|
||||
TotalSize int64 // Total size of all referenced files
|
||||
|
||||
// BlobSize is the total size of all referenced blobs (compressed and
|
||||
|
||||
@@ -606,7 +606,8 @@ func TestTimezoneHandling(t *testing.T) {
|
||||
t.Skip("timezone not available")
|
||||
}
|
||||
|
||||
nyTime := time.Now().In(loc)
|
||||
// Use Truncate to remove sub-second precision since we store as Unix timestamps
|
||||
nyTime := time.Now().In(loc).Truncate(time.Second)
|
||||
file := &File{
|
||||
Path: "/timezone-test.txt",
|
||||
MTime: nyTime,
|
||||
|
||||
@@ -5,9 +5,8 @@
|
||||
CREATE TABLE IF NOT EXISTS files (
|
||||
id TEXT PRIMARY KEY, -- UUID
|
||||
path TEXT NOT NULL UNIQUE,
|
||||
source_path TEXT NOT NULL DEFAULT '', -- The source directory this file came from
|
||||
mtime INTEGER NOT NULL, -- whole seconds since the Unix epoch
|
||||
mtime_nsec INTEGER NOT NULL, -- nanoseconds within that second, 0 to 999999999
|
||||
source_path TEXT NOT NULL DEFAULT '', -- The source directory this file came from (for restore path stripping)
|
||||
mtime INTEGER NOT NULL,
|
||||
size INTEGER NOT NULL,
|
||||
mode INTEGER NOT NULL,
|
||||
uid INTEGER NOT NULL,
|
||||
|
||||
@@ -127,7 +127,6 @@ func (r *SnapshotRepository) UpdateExtendedStats(
|
||||
snapshotID string,
|
||||
blobUncompressedSize int64,
|
||||
compressionLevel int,
|
||||
uploadBytes int64,
|
||||
uploadDurationMs int64,
|
||||
) error {
|
||||
compressionRatio, err := r.extendedCompressionRatio(
|
||||
@@ -142,7 +141,7 @@ func (r *SnapshotRepository) UpdateExtendedStats(
|
||||
SET blob_uncompressed_size = ?,
|
||||
compression_ratio = ?,
|
||||
compression_level = ?,
|
||||
upload_bytes = ?,
|
||||
upload_bytes = blob_size,
|
||||
upload_duration_ms = ?
|
||||
WHERE id = ?
|
||||
`
|
||||
@@ -150,11 +149,11 @@ func (r *SnapshotRepository) UpdateExtendedStats(
|
||||
if tx != nil {
|
||||
_, err = tx.ExecContext(ctx, query,
|
||||
blobUncompressedSize, compressionRatio, compressionLevel,
|
||||
uploadBytes, uploadDurationMs, snapshotID)
|
||||
uploadDurationMs, snapshotID)
|
||||
} else {
|
||||
_, err = r.db.ExecWithLog(ctx, query,
|
||||
blobUncompressedSize, compressionRatio, compressionLevel,
|
||||
uploadBytes, uploadDurationMs, snapshotID)
|
||||
uploadDurationMs, snapshotID)
|
||||
}
|
||||
|
||||
if err != nil {
|
||||
@@ -544,30 +543,6 @@ func (r *SnapshotRepository) GetSnapshotTotalCompressedSize(
|
||||
return totalSize, nil
|
||||
}
|
||||
|
||||
// GetSnapshotBlobSizes returns the total compressed and uncompressed sizes
|
||||
// of all blobs referenced by a snapshot.
|
||||
func (r *SnapshotRepository) GetSnapshotBlobSizes(
|
||||
ctx context.Context, snapshotID string,
|
||||
) (int64, int64, error) {
|
||||
query := `
|
||||
SELECT COALESCE(SUM(b.compressed_size), 0),
|
||||
COALESCE(SUM(b.uncompressed_size), 0)
|
||||
FROM snapshot_blobs sb
|
||||
JOIN blobs b ON sb.blob_hash = b.blob_hash
|
||||
WHERE sb.snapshot_id = ?
|
||||
`
|
||||
|
||||
var compressed, uncompressed int64
|
||||
|
||||
err := r.db.conn.QueryRowContext(ctx, query, snapshotID).Scan(
|
||||
&compressed, &uncompressed)
|
||||
if err != nil {
|
||||
return 0, 0, fmt.Errorf("querying snapshot blob sizes: %w", err)
|
||||
}
|
||||
|
||||
return compressed, uncompressed, nil
|
||||
}
|
||||
|
||||
// GetSnapshotUncompressedChunkSize returns the sum of plaintext sizes of all unique
|
||||
// chunks referenced by a snapshot (via snapshot_files → file_chunks → chunks).
|
||||
func (r *SnapshotRepository) GetSnapshotUncompressedChunkSize(
|
||||
|
||||
@@ -145,65 +145,6 @@ func TestSnapshotRepositoryUpdateCounts(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// GetSnapshotBlobSizes totals the blobs the snapshot references, and only
|
||||
// those.
|
||||
func TestSnapshotRepositoryGetSnapshotBlobSizes(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
db, cleanup := setupTestDB(t)
|
||||
defer cleanup()
|
||||
|
||||
ctx := context.Background()
|
||||
repos := database.NewRepositories(db)
|
||||
|
||||
snapshot := &database.Snapshot{
|
||||
ID: "2024-01-03T12:00:00Z",
|
||||
Hostname: testHostname,
|
||||
VaultikVersion: testVersion,
|
||||
StartedAt: time.Now().Truncate(time.Second),
|
||||
}
|
||||
|
||||
err := repos.Snapshots.Create(ctx, nil, snapshot)
|
||||
if err != nil {
|
||||
t.Fatalf("failed to create snapshot: %v", err)
|
||||
}
|
||||
|
||||
blobs := []*database.Blob{
|
||||
{Hash: "referenced-1", CompressedSize: 10, UncompressedSize: 100},
|
||||
{Hash: "referenced-2", CompressedSize: 20, UncompressedSize: 200},
|
||||
{Hash: "unreferenced", CompressedSize: 40, UncompressedSize: 400},
|
||||
}
|
||||
|
||||
for _, blob := range blobs {
|
||||
blob.ID = types.NewBlobID()
|
||||
blob.CreatedTS = time.Now().Truncate(time.Second)
|
||||
|
||||
err = repos.Blobs.Create(ctx, nil, blob)
|
||||
if err != nil {
|
||||
t.Fatalf("failed to create blob %s: %v", blob.Hash, err)
|
||||
}
|
||||
}
|
||||
|
||||
for _, blob := range blobs[:2] {
|
||||
err = repos.Snapshots.AddBlob(ctx, nil, snapshot.ID.String(),
|
||||
blob.ID, blob.Hash)
|
||||
if err != nil {
|
||||
t.Fatalf("failed to add blob %s to snapshot: %v", blob.Hash, err)
|
||||
}
|
||||
}
|
||||
|
||||
compressed, uncompressed, err := repos.Snapshots.GetSnapshotBlobSizes(
|
||||
ctx, snapshot.ID.String())
|
||||
if err != nil {
|
||||
t.Fatalf("failed to get snapshot blob sizes: %v", err)
|
||||
}
|
||||
|
||||
if compressed != 30 || uncompressed != 300 {
|
||||
t.Errorf("blob sizes: got %d and %d, want 30 and 300",
|
||||
compressed, uncompressed)
|
||||
}
|
||||
}
|
||||
|
||||
func TestSnapshotRepositoryListRecent(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
|
||||
@@ -158,3 +158,19 @@ type UploadStats struct {
|
||||
MinDurationMs int64
|
||||
MaxDurationMs int64
|
||||
}
|
||||
|
||||
// GetCountBySnapshot returns the count of uploads for a specific snapshot
|
||||
func (r *UploadRepository) GetCountBySnapshot(
|
||||
ctx context.Context, snapshotID string,
|
||||
) (int64, error) {
|
||||
query := `SELECT COUNT(*) FROM uploads WHERE snapshot_id = ?`
|
||||
|
||||
var count int64
|
||||
|
||||
err := r.conn.QueryRowContext(ctx, query, snapshotID).Scan(&count)
|
||||
if err != nil {
|
||||
return 0, err
|
||||
}
|
||||
|
||||
return count, nil
|
||||
}
|
||||
|
||||
+49
-64
@@ -10,18 +10,15 @@ import (
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
"golang.org/x/sys/unix"
|
||||
"syscall"
|
||||
)
|
||||
|
||||
// ErrAlreadyRunning indicates another vaultik instance is running.
|
||||
var ErrAlreadyRunning = errors.New("another vaultik instance is already running")
|
||||
|
||||
// Lock represents an acquired PID lock: an flock(2) on the PID file,
|
||||
// held while the file stays open. The kernel drops it when the process
|
||||
// exits, however it exits, so a crashed run never leaves the lock held.
|
||||
// Lock represents an acquired PID lock.
|
||||
type Lock struct {
|
||||
file *os.File
|
||||
path string
|
||||
}
|
||||
|
||||
const (
|
||||
@@ -32,9 +29,10 @@ const (
|
||||
)
|
||||
|
||||
// Acquire attempts to acquire a PID lock in the specified directory.
|
||||
// If another process holds the lock, it returns ErrAlreadyRunning with
|
||||
// that process's PID. On success, it writes the current PID to the lock
|
||||
// file and returns a Lock that must be released with Release().
|
||||
// If the lock file exists and the process is still running, it returns
|
||||
// ErrAlreadyRunning with details about the existing process.
|
||||
// On success, it writes the current PID to the lock file and returns
|
||||
// a Lock that must be released with Release().
|
||||
func Acquire(lockDir string) (*Lock, error) {
|
||||
// Ensure lock directory exists
|
||||
err := os.MkdirAll(lockDir, lockDirPerm)
|
||||
@@ -44,82 +42,56 @@ func Acquire(lockDir string) (*Lock, error) {
|
||||
|
||||
lockPath := filepath.Join(lockDir, "vaultik.pid")
|
||||
|
||||
// No O_TRUNC: the file may hold the PID of the process that has the
|
||||
// lock, which the error below reports.
|
||||
file, err := os.OpenFile( //nolint:gosec // G304: path is our own lock file
|
||||
lockPath, os.O_RDWR|os.O_CREATE, pidFilePerm)
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("opening PID file: %w", err)
|
||||
}
|
||||
|
||||
err = unix.Flock(int(file.Fd()), unix.LOCK_EX|unix.LOCK_NB)
|
||||
if err != nil {
|
||||
_ = file.Close()
|
||||
|
||||
if errors.Is(err, unix.EWOULDBLOCK) {
|
||||
return nil, alreadyRunningError(lockPath)
|
||||
// Check for existing lock
|
||||
existingPID, err := readPIDFile(lockPath)
|
||||
if err == nil {
|
||||
// Lock file exists, check if process is running
|
||||
if isProcessRunning(existingPID) {
|
||||
return nil, fmt.Errorf("%w (PID %d)", ErrAlreadyRunning, existingPID)
|
||||
}
|
||||
|
||||
return nil, fmt.Errorf("locking PID file: %w", err)
|
||||
// Process is not running, stale lock file - we can take over
|
||||
}
|
||||
|
||||
err = writePID(file)
|
||||
// Write our PID
|
||||
pid := os.Getpid()
|
||||
|
||||
err = os.WriteFile(lockPath, []byte(strconv.Itoa(pid)), pidFilePerm)
|
||||
if err != nil {
|
||||
_ = file.Close()
|
||||
|
||||
return nil, err
|
||||
return nil, fmt.Errorf("writing PID file: %w", err)
|
||||
}
|
||||
|
||||
return &Lock{file: file}, nil
|
||||
return &Lock{path: lockPath}, nil
|
||||
}
|
||||
|
||||
// Release empties the PID file and closes it, which drops the lock.
|
||||
// Release removes the PID lock file.
|
||||
// It is safe to call Release multiple times.
|
||||
func (l *Lock) Release() error {
|
||||
if l == nil || l.file == nil {
|
||||
if l == nil || l.path == "" {
|
||||
return nil
|
||||
}
|
||||
|
||||
file := l.file
|
||||
l.file = nil
|
||||
|
||||
// Do not remove the file here. A process that opened it a moment
|
||||
// earlier could then lock the removed file while another creates and
|
||||
// locks a new one, and both would run.
|
||||
truncateErr := file.Truncate(0)
|
||||
closeErr := file.Close()
|
||||
|
||||
return errors.Join(truncateErr, closeErr)
|
||||
}
|
||||
|
||||
// writePID replaces the contents of the locked PID file with the current
|
||||
// PID.
|
||||
func writePID(file *os.File) error {
|
||||
err := file.Truncate(0)
|
||||
// Verify we still own the lock (our PID is in the file)
|
||||
existingPID, err := readPIDFile(l.path)
|
||||
if err != nil {
|
||||
return fmt.Errorf("truncating PID file: %w", err)
|
||||
// File already gone or unreadable - that's fine
|
||||
return nil //nolint:nilerr // unreadable lock file means nothing to release
|
||||
}
|
||||
|
||||
_, err = file.WriteAt([]byte(strconv.Itoa(os.Getpid())), 0)
|
||||
if err != nil {
|
||||
return fmt.Errorf("writing PID file: %w", err)
|
||||
if existingPID != os.Getpid() {
|
||||
// Someone else wrote to our lock file - don't remove it
|
||||
return nil
|
||||
}
|
||||
|
||||
err = os.Remove(l.path)
|
||||
if err != nil && !os.IsNotExist(err) {
|
||||
return fmt.Errorf("removing PID file: %w", err)
|
||||
}
|
||||
|
||||
l.path = "" // Prevent double-release
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// alreadyRunningError reports that another process holds the lock,
|
||||
// naming its PID when the file holds one. The holder writes its PID just
|
||||
// after it locks, so the file can briefly be empty.
|
||||
func alreadyRunningError(lockPath string) error {
|
||||
pid, err := readPIDFile(lockPath)
|
||||
if err != nil {
|
||||
return ErrAlreadyRunning
|
||||
}
|
||||
|
||||
return fmt.Errorf("%w (PID %d)", ErrAlreadyRunning, pid)
|
||||
}
|
||||
|
||||
// readPIDFile reads and parses the PID from a lock file.
|
||||
func readPIDFile(path string) (int, error) {
|
||||
data, err := os.ReadFile(path) //nolint:gosec // G304: path is our own lock file
|
||||
@@ -134,3 +106,16 @@ func readPIDFile(path string) (int, error) {
|
||||
|
||||
return pid, nil
|
||||
}
|
||||
|
||||
// isProcessRunning checks if a process with the given PID is running.
|
||||
func isProcessRunning(pid int) bool {
|
||||
process, err := os.FindProcess(pid)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
|
||||
// On Unix, FindProcess always succeeds. We need to send signal 0 to check.
|
||||
err = process.Signal(syscall.Signal(0))
|
||||
|
||||
return err == nil
|
||||
}
|
||||
|
||||
@@ -4,7 +4,6 @@ import (
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strconv"
|
||||
"sync"
|
||||
"testing"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
@@ -34,10 +33,9 @@ func TestAcquireAndRelease(t *testing.T) {
|
||||
err = lock.Release()
|
||||
require.NoError(t, err)
|
||||
|
||||
// Verify PID file is empty
|
||||
data, err = os.ReadFile(pidPath) //nolint:gosec // G304: test's own temp file
|
||||
require.NoError(t, err)
|
||||
assert.Empty(t, data)
|
||||
// Verify PID file is gone
|
||||
_, err = os.Stat(pidPath)
|
||||
assert.True(t, os.IsNotExist(err))
|
||||
}
|
||||
|
||||
func TestAcquireBlocksSecondInstance(t *testing.T) {
|
||||
@@ -57,64 +55,6 @@ func TestAcquireBlocksSecondInstance(t *testing.T) {
|
||||
lock2, err := pidlock.Acquire(tmpDir)
|
||||
require.ErrorIs(t, err, pidlock.ErrAlreadyRunning)
|
||||
assert.Nil(t, lock2)
|
||||
|
||||
// Once the first lock is released, the next Acquire succeeds
|
||||
require.NoError(t, lock1.Release())
|
||||
|
||||
lock3, err := pidlock.Acquire(tmpDir)
|
||||
require.NoError(t, err)
|
||||
require.NoError(t, lock3.Release())
|
||||
}
|
||||
|
||||
// TestConcurrentAcquireAdmitsOne starts many Acquire calls at the same
|
||||
// moment, as two cron entries firing together would, and checks that
|
||||
// exactly one of them gets the lock.
|
||||
func TestConcurrentAcquireAdmitsOne(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const callers = 50
|
||||
|
||||
tmpDir := t.TempDir()
|
||||
start := make(chan struct{})
|
||||
|
||||
var (
|
||||
mu sync.Mutex
|
||||
acquired []*pidlock.Lock
|
||||
failures []error
|
||||
wg sync.WaitGroup
|
||||
)
|
||||
|
||||
for range callers {
|
||||
wg.Go(func() {
|
||||
<-start
|
||||
|
||||
lock, err := pidlock.Acquire(tmpDir)
|
||||
|
||||
mu.Lock()
|
||||
defer mu.Unlock()
|
||||
|
||||
if err != nil {
|
||||
failures = append(failures, err)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
acquired = append(acquired, lock)
|
||||
})
|
||||
}
|
||||
|
||||
close(start)
|
||||
wg.Wait()
|
||||
|
||||
for _, lock := range acquired {
|
||||
require.NoError(t, lock.Release())
|
||||
}
|
||||
|
||||
assert.Len(t, acquired, 1, "exactly one caller should hold the lock")
|
||||
|
||||
for _, err := range failures {
|
||||
require.ErrorIs(t, err, pidlock.ErrAlreadyRunning)
|
||||
}
|
||||
}
|
||||
|
||||
func TestAcquireWithStaleLock(t *testing.T) {
|
||||
|
||||
+5
-23
@@ -27,12 +27,10 @@ type Client struct {
|
||||
bucket string
|
||||
prefix string
|
||||
endpoint string
|
||||
partSize int64
|
||||
}
|
||||
|
||||
// Config contains S3 client configuration.
|
||||
// All fields are required except Prefix, which defaults to an empty string,
|
||||
// and PartSize, where zero means the SDK default of 5 MiB.
|
||||
// All fields are required except Prefix, which defaults to an empty string.
|
||||
// A non-empty Prefix is joined to every key with one "/", whether or not
|
||||
// it ends with one.
|
||||
// The Endpoint field should include the protocol (http:// or https://).
|
||||
@@ -43,9 +41,6 @@ type Config struct {
|
||||
AccessKeyID string
|
||||
SecretAccessKey string
|
||||
Region string
|
||||
// PartSize is the size in bytes of each part of a multipart upload.
|
||||
// An upload too large for S3's limit of 10,000 parts gets larger parts.
|
||||
PartSize int64
|
||||
}
|
||||
|
||||
// nopLogger is a logger that discards all output.
|
||||
@@ -95,7 +90,6 @@ func NewClient(ctx context.Context, cfg Config) (*Client, error) {
|
||||
bucket: cfg.Bucket,
|
||||
prefix: prefix,
|
||||
endpoint: cfg.Endpoint,
|
||||
partSize: cfg.PartSize,
|
||||
}, nil
|
||||
}
|
||||
|
||||
@@ -129,9 +123,12 @@ func (c *Client) PutObjectWithProgress(
|
||||
) error {
|
||||
fullKey := c.prefix + key
|
||||
|
||||
// uploadPartSize is 10MB for better progress granularity.
|
||||
const uploadPartSize = 10 * 1024 * 1024
|
||||
|
||||
// Create an uploader with the S3 client
|
||||
uploader := manager.NewUploader(c.s3Client, func(u *manager.Uploader) {
|
||||
u.PartSize = uploadPartSize(c.partSize, size)
|
||||
u.PartSize = uploadPartSize
|
||||
})
|
||||
|
||||
// Create a progress reader that tracks upload progress
|
||||
@@ -152,21 +149,6 @@ func (c *Client) PutObjectWithProgress(
|
||||
return err
|
||||
}
|
||||
|
||||
// uploadPartSize returns the part size for an upload of size bytes: the
|
||||
// configured part size (the SDK default when zero), raised where needed so
|
||||
// the upload fits in S3's limit of 10,000 parts. The uploader cannot raise
|
||||
// it itself, because it cannot seek the progress reader to learn its size.
|
||||
func uploadPartSize(configured, size int64) int64 {
|
||||
if configured == 0 {
|
||||
configured = manager.DefaultUploadPartSize
|
||||
}
|
||||
|
||||
maxParts := int64(manager.MaxUploadParts)
|
||||
smallestThatFits := (size + maxParts - 1) / maxParts // rounded up
|
||||
|
||||
return max(configured, smallestThatFits)
|
||||
}
|
||||
|
||||
// GetObject downloads an object from S3 with the specified key.
|
||||
// The key is automatically prefixed with the configured prefix.
|
||||
// Returns a ReadCloser containing the object data. The caller must
|
||||
|
||||
@@ -1,55 +0,0 @@
|
||||
package s3
|
||||
|
||||
import "testing"
|
||||
|
||||
// TestUploadPartSize checks that an upload too large for 10,000 parts of the
|
||||
// configured size gets parts just large enough to fit in 10,000.
|
||||
func TestUploadPartSize(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const mib = 1024 * 1024
|
||||
|
||||
tests := []struct {
|
||||
name string
|
||||
configured int64
|
||||
size int64
|
||||
want int64
|
||||
}{
|
||||
{
|
||||
name: "an upload that fits keeps the configured size",
|
||||
configured: 5 * mib,
|
||||
size: 10 * 1024 * mib,
|
||||
want: 5 * mib,
|
||||
},
|
||||
{
|
||||
name: "exactly 10,000 parts keeps the configured size",
|
||||
configured: 6 * mib,
|
||||
size: 10_000 * 6 * mib,
|
||||
want: 6 * mib,
|
||||
},
|
||||
{
|
||||
name: "one byte more than 10,000 parts adds a byte to each",
|
||||
configured: 6 * mib,
|
||||
size: 10_000*6*mib + 1,
|
||||
want: 6*mib + 1,
|
||||
},
|
||||
{
|
||||
name: "zero means the SDK default of 5MiB",
|
||||
configured: 0,
|
||||
size: 1,
|
||||
want: 5 * mib,
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
got := uploadPartSize(tt.configured, tt.size)
|
||||
if got != tt.want {
|
||||
t.Errorf("uploadPartSize(%d, %d) = %d, want %d",
|
||||
tt.configured, tt.size, got, tt.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
@@ -28,7 +28,6 @@ func provideClient(lc fx.Lifecycle, cfg *config.Config) (*Client, error) {
|
||||
AccessKeyID: cfg.S3.AccessKeyID,
|
||||
SecretAccessKey: cfg.S3.SecretAccessKey,
|
||||
Region: cfg.S3.Region,
|
||||
PartSize: cfg.S3.PartSize.Int64(),
|
||||
})
|
||||
if err != nil {
|
||||
return nil, err
|
||||
|
||||
@@ -66,6 +66,7 @@ type ProgressStats struct {
|
||||
BlobsCreated atomic.Int64
|
||||
BlobsUploaded atomic.Int64
|
||||
BytesUploaded atomic.Int64
|
||||
UploadDurationMs atomic.Int64 // Total milliseconds spent uploading
|
||||
CurrentFile atomic.Value // stores string
|
||||
TotalSize atomic.Int64 // Total size to process (set after scan phase)
|
||||
TotalFiles atomic.Int64 // Total files to process in phase 2
|
||||
@@ -230,6 +231,9 @@ func (pr *ProgressReporter) ReportUploadComplete(
|
||||
// Clear current upload
|
||||
pr.stats.CurrentUpload.Store((*UploadInfo)(nil))
|
||||
|
||||
// Add to total upload duration
|
||||
pr.stats.UploadDurationMs.Add(duration.Milliseconds())
|
||||
|
||||
// Calculate speed
|
||||
if duration < time.Millisecond {
|
||||
duration = time.Millisecond
|
||||
|
||||
@@ -58,8 +58,8 @@ type Scanner struct {
|
||||
compressionLevel int
|
||||
ageRecipient string
|
||||
snapshotID string // Current snapshot being processed
|
||||
// currentSourcePath is the source directory being scanned, stored with
|
||||
// each file record.
|
||||
// currentSourcePath is the source directory being scanned (used for
|
||||
// restore path stripping).
|
||||
currentSourcePath string
|
||||
exclude []string // Glob patterns for files/directories to exclude
|
||||
compiledExclude []compiledPattern // Compiled glob patterns
|
||||
@@ -92,6 +92,9 @@ type Scanner struct {
|
||||
|
||||
// Mutex for coordinating blob creation
|
||||
packerMu sync.Mutex // Blocks chunk production during blob creation
|
||||
|
||||
// Context for cancellation
|
||||
scanCtx context.Context //nolint:containedctx // set per-Scan for packer callbacks
|
||||
}
|
||||
|
||||
// Periodic status output intervals and thresholds for the scan and
|
||||
@@ -131,23 +134,18 @@ type ScannerConfig struct {
|
||||
SkipErrors bool
|
||||
}
|
||||
|
||||
// ScanResult contains the results of a scan operation. Files and bytes
|
||||
// are counted per file: BytesScanned is the size of the new and changed
|
||||
// files, BytesSkipped that of the unchanged ones.
|
||||
// ScanResult contains the results of a scan operation
|
||||
type ScanResult struct {
|
||||
FilesScanned int
|
||||
FilesSkipped int
|
||||
FilesDeleted int
|
||||
BytesScanned int64
|
||||
BytesSkipped int64
|
||||
BytesDeleted int64
|
||||
ChunksCreated int
|
||||
BlobsCreated int
|
||||
BlobsUploaded int
|
||||
BytesUploaded int64
|
||||
UploadDuration time.Duration
|
||||
StartTime time.Time
|
||||
EndTime time.Time
|
||||
FilesScanned int
|
||||
FilesSkipped int
|
||||
FilesDeleted int
|
||||
BytesScanned int64
|
||||
BytesSkipped int64
|
||||
BytesDeleted int64
|
||||
ChunksCreated int
|
||||
BlobsCreated int
|
||||
StartTime time.Time
|
||||
EndTime time.Time
|
||||
}
|
||||
|
||||
// NewScanner creates a new scanner instance
|
||||
@@ -211,7 +209,9 @@ func (s *Scanner) Scan(
|
||||
ctx context.Context, path string, snapshotID string,
|
||||
) (*ScanResult, error) {
|
||||
s.snapshotID = snapshotID
|
||||
// Store source path for file records (used during restore)
|
||||
s.currentSourcePath = path
|
||||
s.scanCtx = ctx
|
||||
result := &ScanResult{
|
||||
StartTime: time.Now().UTC(),
|
||||
}
|
||||
@@ -219,13 +219,17 @@ func (s *Scanner) Scan(
|
||||
// Set blob handler for concurrent upload
|
||||
if s.storage != nil {
|
||||
log.Debug("Setting blob handler for storage uploads")
|
||||
s.packer.SetBlobHandler(func(blobWithReader *blob.WithReader) error {
|
||||
return s.handleBlobReady(ctx, blobWithReader, result)
|
||||
})
|
||||
s.packer.SetBlobHandler(s.handleBlobReady)
|
||||
} else {
|
||||
log.Debug("No storage configured, blobs will not be uploaded")
|
||||
}
|
||||
|
||||
// Start progress reporting if enabled
|
||||
if s.progress != nil {
|
||||
s.progress.Start()
|
||||
defer s.progress.Stop()
|
||||
}
|
||||
|
||||
// Phase 0: Repair any state left by an interrupted previous run, then
|
||||
// load known files and chunks from the database into memory for fast
|
||||
// lookup.
|
||||
@@ -290,14 +294,13 @@ func (s *Scanner) Scan(
|
||||
log.Info("Phase 2/3: Skipping (no files need processing, metadata-only snapshot)")
|
||||
}
|
||||
|
||||
result.EndTime = time.Now().UTC()
|
||||
// Finalize result with blob statistics
|
||||
s.finalizeScanResult(ctx, result)
|
||||
|
||||
return result, nil
|
||||
}
|
||||
|
||||
// GetProgress returns the progress reporter for this scanner, or nil when
|
||||
// progress is off. Scan neither starts nor stops it: the caller does,
|
||||
// once for all the paths it scans, because a second Stop panics.
|
||||
// GetProgress returns the progress reporter for this scanner
|
||||
func (s *Scanner) GetProgress() *ProgressReporter {
|
||||
return s.progress
|
||||
}
|
||||
@@ -432,6 +435,27 @@ func (s *Scanner) summarizeScanPhase(
|
||||
s.ui.Completef("%s.", msg)
|
||||
}
|
||||
|
||||
// finalizeScanResult populates final blob statistics in the scan result
|
||||
// by querying the packer and database for blob/upload counts
|
||||
func (s *Scanner) finalizeScanResult(ctx context.Context, result *ScanResult) {
|
||||
blobs := s.packer.GetFinishedBlobs()
|
||||
result.BlobsCreated += len(blobs)
|
||||
|
||||
// Query database for actual blob count created during this snapshot
|
||||
// The database is authoritative, especially for concurrent blob uploads
|
||||
// We count uploads rather than all snapshot_blobs to get only NEW blobs
|
||||
if s.snapshotID != "" {
|
||||
uploadCount, err := s.repos.Uploads.GetCountBySnapshot(ctx, s.snapshotID)
|
||||
if err != nil {
|
||||
log.Warn("Failed to query upload count from database", "error", err)
|
||||
} else {
|
||||
result.BlobsCreated = int(uploadCount)
|
||||
}
|
||||
}
|
||||
|
||||
result.EndTime = time.Now().UTC()
|
||||
}
|
||||
|
||||
// loadKnownFiles loads the known files at and beneath path from the
|
||||
// database into a map for fast lookup. Every loaded file the scan does
|
||||
// not find is counted as deleted. This avoids per-file database queries
|
||||
@@ -1197,8 +1221,9 @@ func (s *Scanner) checkFileInMemory(
|
||||
}
|
||||
|
||||
file := &database.File{
|
||||
ID: fileID,
|
||||
Path: types.FilePath(path),
|
||||
ID: fileID,
|
||||
Path: types.FilePath(path),
|
||||
// Store source directory for restore path stripping
|
||||
SourcePath: types.SourcePath(s.currentSourcePath),
|
||||
MTime: info.ModTime(),
|
||||
Size: info.Size(),
|
||||
@@ -1219,7 +1244,7 @@ func (s *Scanner) checkFileInMemory(
|
||||
|
||||
// Check if file has changed
|
||||
if existingFile.Size != file.Size ||
|
||||
!existingFile.MTime.Equal(file.MTime) ||
|
||||
existingFile.MTime.Unix() != file.MTime.Unix() ||
|
||||
existingFile.Mode != file.Mode ||
|
||||
existingFile.UID != file.UID ||
|
||||
existingFile.GID != file.GID {
|
||||
@@ -1490,24 +1515,24 @@ func (s *Scanner) finalizeProcessPhase(ctx context.Context, result *ScanResult)
|
||||
}
|
||||
|
||||
// handleBlobReady is called by the packer when a blob is finalized
|
||||
func (s *Scanner) handleBlobReady(
|
||||
ctx context.Context, blobWithReader *blob.WithReader, result *ScanResult,
|
||||
) error {
|
||||
func (s *Scanner) handleBlobReady(blobWithReader *blob.WithReader) error {
|
||||
startTime := time.Now().UTC()
|
||||
finishedBlob := blobWithReader.FinishedBlob
|
||||
|
||||
result.BlobsCreated++
|
||||
|
||||
if s.progress != nil {
|
||||
s.progress.ReportUploadStart(finishedBlob.Hash, finishedBlob.Compressed)
|
||||
s.progress.GetStats().BlobsCreated.Add(1)
|
||||
}
|
||||
|
||||
ctx := s.scanCtx
|
||||
if ctx == nil {
|
||||
ctx = context.Background()
|
||||
}
|
||||
|
||||
blobPath := fmt.Sprintf("blobs/%s/%s/%s",
|
||||
finishedBlob.Hash[:2], finishedBlob.Hash[2:4], finishedBlob.Hash)
|
||||
|
||||
blobExists, err := s.uploadBlobIfNeeded(
|
||||
ctx, blobPath, blobWithReader, startTime, result)
|
||||
blobExists, err := s.uploadBlobIfNeeded(ctx, blobPath, blobWithReader, startTime)
|
||||
if err != nil {
|
||||
s.cleanupBlobTempFile(blobWithReader)
|
||||
|
||||
@@ -1542,7 +1567,6 @@ func (s *Scanner) uploadBlobIfNeeded(
|
||||
blobPath string,
|
||||
blobWithReader *blob.WithReader,
|
||||
startTime time.Time,
|
||||
result *ScanResult,
|
||||
) (bool, error) {
|
||||
finishedBlob := blobWithReader.FinishedBlob
|
||||
|
||||
@@ -1578,10 +1602,6 @@ func (s *Scanner) uploadBlobIfNeeded(
|
||||
uploadDuration := time.Since(startTime)
|
||||
uploadSpeedBps := float64(finishedBlob.Compressed) / uploadDuration.Seconds()
|
||||
|
||||
result.BlobsUploaded++
|
||||
result.BytesUploaded += finishedBlob.Compressed
|
||||
result.UploadDuration += uploadDuration
|
||||
|
||||
s.ui.Completef("Uploaded blob %s (%s) in %s at %s.",
|
||||
s.ui.Hex(finishedBlob.Hash),
|
||||
s.ui.Size(finishedBlob.Compressed),
|
||||
@@ -1797,9 +1817,9 @@ func (s *Scanner) processFileStreaming(
|
||||
size: chunk.Size,
|
||||
})
|
||||
|
||||
if !chunkExists {
|
||||
s.updateChunkStats(chunk.Size, result)
|
||||
s.updateChunkStats(chunkExists, chunk.Size, result)
|
||||
|
||||
if !chunkExists {
|
||||
err := s.addChunkToPacker(ctx, chunk)
|
||||
if err != nil {
|
||||
// Mark as a packer error so --skip-errors cannot swallow it:
|
||||
@@ -1827,16 +1847,26 @@ func (s *Scanner) processFileStreaming(
|
||||
return nil
|
||||
}
|
||||
|
||||
// updateChunkStats counts a chunk that was not already stored. The scan
|
||||
// result's file counts, BytesScanned and BytesSkipped are not touched
|
||||
// here: the scan phase counts each file once.
|
||||
func (s *Scanner) updateChunkStats(chunkSize int64, result *ScanResult) {
|
||||
result.ChunksCreated++
|
||||
// updateChunkStats updates scan result and progress stats for a processed chunk
|
||||
func (s *Scanner) updateChunkStats(
|
||||
chunkExists bool, chunkSize int64, result *ScanResult,
|
||||
) {
|
||||
if chunkExists {
|
||||
result.FilesSkipped++
|
||||
|
||||
if s.progress != nil {
|
||||
s.progress.GetStats().ChunksCreated.Add(1)
|
||||
s.progress.GetStats().BytesProcessed.Add(chunkSize)
|
||||
s.progress.UpdateChunkingActivity()
|
||||
result.BytesSkipped += chunkSize
|
||||
if s.progress != nil {
|
||||
s.progress.GetStats().BytesSkipped.Add(chunkSize)
|
||||
}
|
||||
} else {
|
||||
result.ChunksCreated++
|
||||
result.BytesScanned += chunkSize
|
||||
|
||||
if s.progress != nil {
|
||||
s.progress.GetStats().ChunksCreated.Add(1)
|
||||
s.progress.GetStats().BytesProcessed.Add(chunkSize)
|
||||
s.progress.UpdateChunkingActivity()
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
|
||||
@@ -105,21 +105,17 @@ func (sm *SnapshotManager) CreateSnapshot(
|
||||
return sm.CreateSnapshotWithName(ctx, hostname, "", version, gitRevision)
|
||||
}
|
||||
|
||||
// ShortHostname returns hostname up to its first dot. A snapshot ID starts
|
||||
// with this form, while the snapshots table stores the full hostname.
|
||||
func ShortHostname(hostname string) string {
|
||||
short, _, _ := strings.Cut(hostname, ".")
|
||||
|
||||
return short
|
||||
}
|
||||
|
||||
// CreateSnapshotWithName creates a new snapshot record with an optional
|
||||
// snapshot name. The snapshot ID format is: hostname_name_timestamp or
|
||||
// hostname_timestamp if name is empty.
|
||||
func (sm *SnapshotManager) CreateSnapshotWithName(
|
||||
ctx context.Context, hostname, name, version, gitRevision string,
|
||||
) (string, error) {
|
||||
shortHostname := ShortHostname(hostname)
|
||||
// Use short hostname (strip domain if present)
|
||||
shortHostname := hostname
|
||||
if before, _, ok := strings.Cut(hostname, "."); ok {
|
||||
shortHostname = before
|
||||
}
|
||||
|
||||
// Build snapshot ID with optional name
|
||||
timestamp := time.Now().UTC().Format("2006-01-02T15:04:05Z")
|
||||
@@ -158,6 +154,26 @@ func (sm *SnapshotManager) CreateSnapshotWithName(
|
||||
return snapshotID, nil
|
||||
}
|
||||
|
||||
// UpdateSnapshotStats updates the statistics for a snapshot during backup
|
||||
func (sm *SnapshotManager) UpdateSnapshotStats(
|
||||
ctx context.Context, snapshotID string, stats BackupStats,
|
||||
) error {
|
||||
err := sm.repos.WithTx(ctx, func(ctx context.Context, tx *sql.Tx) error {
|
||||
return sm.repos.Snapshots.UpdateCounts(ctx, tx, snapshotID,
|
||||
int64(stats.FilesScanned),
|
||||
int64(stats.ChunksCreated),
|
||||
int64(stats.BlobsCreated),
|
||||
stats.BytesScanned,
|
||||
stats.BytesUploaded,
|
||||
)
|
||||
})
|
||||
if err != nil {
|
||||
return fmt.Errorf("updating snapshot stats: %w", err)
|
||||
}
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
// UpdateSnapshotStatsExtended updates snapshot statistics with extended metrics.
|
||||
// This includes compression level, uncompressed blob size, and upload duration.
|
||||
func (sm *SnapshotManager) UpdateSnapshotStatsExtended(
|
||||
@@ -169,8 +185,8 @@ func (sm *SnapshotManager) UpdateSnapshotStatsExtended(
|
||||
int64(stats.FilesScanned),
|
||||
int64(stats.ChunksCreated),
|
||||
int64(stats.BlobsCreated),
|
||||
stats.TotalSize,
|
||||
stats.BlobSize,
|
||||
stats.BytesScanned,
|
||||
stats.BytesUploaded,
|
||||
)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -180,7 +196,6 @@ func (sm *SnapshotManager) UpdateSnapshotStatsExtended(
|
||||
return sm.repos.Snapshots.UpdateExtendedStats(ctx, tx, snapshotID,
|
||||
stats.BlobUncompressedSize,
|
||||
stats.CompressionLevel,
|
||||
stats.BytesUploaded,
|
||||
stats.UploadDurationMs,
|
||||
)
|
||||
})
|
||||
@@ -875,7 +890,7 @@ func (sm *SnapshotManager) getFileSize(path string) int64 {
|
||||
// BackupStats contains statistics from a backup operation
|
||||
type BackupStats struct {
|
||||
FilesScanned int
|
||||
TotalSize int64 // Total size of all files examined
|
||||
BytesScanned int64
|
||||
ChunksCreated int
|
||||
BlobsCreated int
|
||||
BytesUploaded int64
|
||||
@@ -885,7 +900,6 @@ type BackupStats struct {
|
||||
type ExtendedBackupStats struct {
|
||||
BackupStats
|
||||
|
||||
BlobSize int64 // Total compressed size of all referenced blobs
|
||||
BlobUncompressedSize int64 // Total uncompressed size of all referenced blobs
|
||||
CompressionLevel int // Compression level used for this snapshot
|
||||
UploadDurationMs int64 // Total milliseconds spent uploading to S3
|
||||
|
||||
@@ -99,7 +99,6 @@ func storerFromParsedS3URL(parsed *URL, cfg *config.Config) (Storer, error) {
|
||||
AccessKeyID: cfg.S3.AccessKeyID,
|
||||
SecretAccessKey: cfg.S3.SecretAccessKey,
|
||||
Region: region,
|
||||
PartSize: cfg.S3.PartSize.Int64(),
|
||||
})
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("creating S3 client: %w", err)
|
||||
@@ -135,7 +134,6 @@ func storerFromLegacyS3Config(cfg *config.Config) (Storer, error) {
|
||||
AccessKeyID: cfg.S3.AccessKeyID,
|
||||
SecretAccessKey: cfg.S3.SecretAccessKey,
|
||||
Region: region,
|
||||
PartSize: cfg.S3.PartSize.Int64(),
|
||||
})
|
||||
if err != nil {
|
||||
return nil, fmt.Errorf("creating S3 client: %w", err)
|
||||
|
||||
@@ -1,14 +1,11 @@
|
||||
package storage_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"errors"
|
||||
"net/http"
|
||||
"net/http/httptest"
|
||||
"slices"
|
||||
"strings"
|
||||
"sync/atomic"
|
||||
"testing"
|
||||
|
||||
"github.com/johannesboyne/gofakes3"
|
||||
@@ -22,13 +19,6 @@ import (
|
||||
// s3TestBucket is the bucket created for each in-process S3 server.
|
||||
const s3TestBucket = "test-bucket"
|
||||
|
||||
// Credentials for the tests that build a storer from a config.Config. The
|
||||
// in-process S3 server accepts any.
|
||||
const (
|
||||
s3TestAccessKeyID = "key"
|
||||
s3TestSecretAccessKey = "secret"
|
||||
)
|
||||
|
||||
// newS3Storer builds an s3:// backend backed by a fresh in-process
|
||||
// S3 server. It reuses the same in-memory S3 harness (gofakes3 + s3mem
|
||||
// over httptest) that internal/s3 and the not-found regression test use,
|
||||
@@ -138,8 +128,8 @@ func TestS3URLPrefixKeyLayout(t *testing.T) {
|
||||
storer, err := storage.NewStorer(&config.Config{
|
||||
StorageURL: storageURL + "?endpoint=" + srv.URL,
|
||||
S3: config.S3Config{
|
||||
AccessKeyID: s3TestAccessKeyID,
|
||||
SecretAccessKey: s3TestSecretAccessKey,
|
||||
AccessKeyID: "key",
|
||||
SecretAccessKey: "secret",
|
||||
},
|
||||
})
|
||||
if err != nil {
|
||||
@@ -181,90 +171,6 @@ func TestS3URLPrefixKeyLayout(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// TestS3UploadUsesConfiguredPartSize checks that s3.part_size reaches the
|
||||
// multipart uploader, through storage_url and through the s3.* fields. An
|
||||
// object three parts long must arrive as three parts; at the SDK's default
|
||||
// of 5 MiB it would arrive as four.
|
||||
func TestS3UploadUsesConfiguredPartSize(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
const (
|
||||
partSize = 6 * 1024 * 1024
|
||||
wantParts = 3
|
||||
)
|
||||
|
||||
backend := s3mem.New()
|
||||
|
||||
err := backend.CreateBucket(s3TestBucket)
|
||||
if err != nil {
|
||||
t.Fatalf("create bucket: %v", err)
|
||||
}
|
||||
|
||||
// Every part of a multipart upload is one request with a partNumber.
|
||||
var parts atomic.Int32
|
||||
|
||||
fake := gofakes3.New(backend).Server()
|
||||
srv := httptest.NewServer(http.HandlerFunc(
|
||||
func(w http.ResponseWriter, r *http.Request) {
|
||||
if r.URL.Query().Has("partNumber") {
|
||||
parts.Add(1)
|
||||
}
|
||||
|
||||
fake.ServeHTTP(w, r)
|
||||
}))
|
||||
t.Cleanup(srv.Close)
|
||||
|
||||
cases := []struct {
|
||||
name string
|
||||
cfg *config.Config
|
||||
}{
|
||||
{
|
||||
name: "storage_url",
|
||||
cfg: &config.Config{
|
||||
StorageURL: "s3://" + s3TestBucket + "?endpoint=" + srv.URL,
|
||||
S3: config.S3Config{
|
||||
AccessKeyID: s3TestAccessKeyID,
|
||||
SecretAccessKey: s3TestSecretAccessKey,
|
||||
PartSize: partSize,
|
||||
},
|
||||
},
|
||||
},
|
||||
{
|
||||
name: "s3.endpoint",
|
||||
cfg: &config.Config{
|
||||
S3: config.S3Config{
|
||||
Endpoint: srv.URL,
|
||||
Bucket: s3TestBucket,
|
||||
AccessKeyID: s3TestAccessKeyID,
|
||||
SecretAccessKey: s3TestSecretAccessKey,
|
||||
PartSize: partSize,
|
||||
},
|
||||
},
|
||||
},
|
||||
}
|
||||
|
||||
for _, tc := range cases {
|
||||
parts.Store(0)
|
||||
|
||||
storer, err := storage.NewStorer(tc.cfg)
|
||||
if err != nil {
|
||||
t.Fatalf("%s: NewStorer: %v", tc.name, err)
|
||||
}
|
||||
|
||||
data := bytes.NewReader(make([]byte, wantParts*partSize))
|
||||
|
||||
err = storer.PutWithProgress(
|
||||
context.Background(), "blob", data, data.Size(), nil)
|
||||
if err != nil {
|
||||
t.Fatalf("%s: PutWithProgress: %v", tc.name, err)
|
||||
}
|
||||
|
||||
if got := parts.Load(); got != wantParts {
|
||||
t.Errorf("%s: uploaded in %d parts, want %d", tc.name, got, wantParts)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// listStreamKeys returns the keys ListStream yields under a prefix, and
|
||||
// fails the test on a listing error.
|
||||
func listStreamKeys(t *testing.T, s storage.Storer, prefix string) []string {
|
||||
|
||||
@@ -155,8 +155,8 @@ type BlobHash string
|
||||
// FilePath represents an absolute path to a file or directory.
|
||||
type FilePath string
|
||||
|
||||
// SourcePath is the source directory a scan found a file under, made
|
||||
// absolute and with symlinks resolved.
|
||||
// SourcePath represents the root directory from which files are backed up.
|
||||
// Used during restore to strip the source prefix from paths.
|
||||
type SourcePath string
|
||||
|
||||
// Hostname identifies a host machine.
|
||||
|
||||
+18
-23
@@ -9,7 +9,6 @@ import (
|
||||
"time"
|
||||
|
||||
"github.com/dustin/go-humanize"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
"sneak.berlin/go/vaultik/internal/types"
|
||||
)
|
||||
|
||||
@@ -47,9 +46,12 @@ const (
|
||||
year = 365 * day
|
||||
)
|
||||
|
||||
// A snapshot ID split on "_" has at least a hostname and a trailing
|
||||
// timestamp.
|
||||
const minSnapshotIDParts = 2
|
||||
// Snapshot IDs split on "_" into hostname, optional name parts, and a
|
||||
// trailing timestamp.
|
||||
const (
|
||||
minSnapshotIDParts = 2
|
||||
minSnapshotIDNameParts = 3
|
||||
)
|
||||
|
||||
// SnapshotInfo contains information about a snapshot.
|
||||
//
|
||||
@@ -119,27 +121,20 @@ func parseSnapshotTimestamp(snapshotID string) (time.Time, error) {
|
||||
return timestamp.UTC(), nil
|
||||
}
|
||||
|
||||
// parseSnapshotName extracts the snapshot name from a snapshot ID of the
|
||||
// form hostname_name_timestamp, given the hostname stored with that
|
||||
// snapshot. The hostname and the name may both contain underscores, so the
|
||||
// name is what is left after removing the short hostname and its "_" from
|
||||
// the front and the last "_" and the timestamp from the end. Returns "" for
|
||||
// an ID with no name (hostname_timestamp), and for an ID that does not start
|
||||
// with that hostname, which CreateSnapshotWithName never writes.
|
||||
func parseSnapshotName(snapshotID, hostname string) string {
|
||||
prefix := snapshot.ShortHostname(hostname) + "_"
|
||||
|
||||
rest, ok := strings.CutPrefix(snapshotID, prefix)
|
||||
if !ok {
|
||||
// parseSnapshotName extracts the snapshot name from a snapshot ID.
|
||||
// Format: hostname_snapshotname_timestamp — the middle part(s) between hostname
|
||||
// and the RFC3339 timestamp are the snapshot name (may contain underscores).
|
||||
// Returns the snapshot name, or empty string if the ID is malformed.
|
||||
func parseSnapshotName(snapshotID string) string {
|
||||
parts := strings.Split(snapshotID, "_")
|
||||
if len(parts) < minSnapshotIDNameParts {
|
||||
// Format: hostname_timestamp — no snapshot name
|
||||
return ""
|
||||
}
|
||||
|
||||
end := strings.LastIndex(rest, "_")
|
||||
if end < 0 {
|
||||
return ""
|
||||
}
|
||||
|
||||
return rest[:end]
|
||||
// Format: hostname_name_timestamp — middle parts are the name.
|
||||
// The last part is the RFC3339 timestamp, the first part is the hostname,
|
||||
// everything in between is the snapshot name (which may itself contain underscores).
|
||||
return strings.Join(parts[1:len(parts)-1], "_")
|
||||
}
|
||||
|
||||
// parseDuration parses a duration string with support for human-friendly units:
|
||||
|
||||
@@ -11,55 +11,33 @@ func TestParseSnapshotName(t *testing.T) {
|
||||
tests := []struct {
|
||||
name string
|
||||
snapshotID string
|
||||
hostname string
|
||||
want string
|
||||
}{
|
||||
{
|
||||
name: "standard format with name",
|
||||
snapshotID: "myhost_home_2026-01-12T14:41:15Z",
|
||||
hostname: "myhost",
|
||||
want: "home",
|
||||
},
|
||||
{
|
||||
name: "standard format with different name",
|
||||
snapshotID: "server1_system_2026-02-15T09:30:00Z",
|
||||
hostname: "server1",
|
||||
want: "system",
|
||||
},
|
||||
{
|
||||
name: "name with underscores",
|
||||
snapshotID: "myhost_my_special_backup_2026-03-01T00:00:00Z",
|
||||
hostname: "myhost",
|
||||
want: "my_special_backup",
|
||||
},
|
||||
{
|
||||
name: "hostname with underscores",
|
||||
snapshotID: "my_host_docs_2026-03-01T00:00:00Z",
|
||||
hostname: "my_host",
|
||||
want: "docs",
|
||||
},
|
||||
{
|
||||
name: "stored hostname with domain",
|
||||
snapshotID: "my_host_mail_2026-03-01T00:00:00Z",
|
||||
hostname: "my_host.example.com",
|
||||
want: "mail",
|
||||
},
|
||||
{
|
||||
name: "no name",
|
||||
snapshotID: "my_host_2026-03-01T00:00:00Z",
|
||||
hostname: "my_host",
|
||||
want: "",
|
||||
},
|
||||
}
|
||||
|
||||
for _, tt := range tests {
|
||||
t.Run(tt.name, func(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
got := parseSnapshotName(tt.snapshotID, tt.hostname)
|
||||
got := parseSnapshotName(tt.snapshotID)
|
||||
if got != tt.want {
|
||||
t.Errorf("parseSnapshotName(%q, %q) = %q, want %q",
|
||||
tt.snapshotID, tt.hostname, got, tt.want)
|
||||
t.Errorf("parseSnapshotName(%q) = %q, want %q",
|
||||
tt.snapshotID, got, tt.want)
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
+24
-108
@@ -183,11 +183,6 @@ type SnapshotMetadataInfo struct {
|
||||
TotalSize int64 `json:"total_size"`
|
||||
BlobCount int `json:"blob_count"`
|
||||
BlobsSize int64 `json:"blobs_size"`
|
||||
|
||||
// Set when the listing holds this snapshot's manifest.json.zst. A
|
||||
// backup interrupted before its manifest upload leaves a directory
|
||||
// without one, which prune does not treat as a snapshot.
|
||||
hasManifest bool
|
||||
}
|
||||
|
||||
// RemoteInfoResult contains all remote storage information
|
||||
@@ -211,20 +206,9 @@ type RemoteInfoResult struct {
|
||||
ReferencedBlobCount int `json:"referenced_blob_count"`
|
||||
ReferencedBlobSize int64 `json:"referenced_blob_size"`
|
||||
|
||||
// Orphaned blobs. Both stay nil (null in the JSON) when a manifest
|
||||
// was listed but not read, since that snapshot's blobs would be
|
||||
// counted as orphaned.
|
||||
OrphanedBlobCount *int `json:"orphaned_blob_count"`
|
||||
OrphanedBlobSize *int64 `json:"orphaned_blob_size"`
|
||||
|
||||
// Remote key of each snapshot whose manifest could not be read
|
||||
UnreadableManifests []string `json:"unreadable_manifests,omitempty"`
|
||||
|
||||
// Number of manifests not read because the name above them under
|
||||
// metadata/ is not a remote key. The names themselves are not
|
||||
// reported: they come from the destination store and may hold
|
||||
// control characters.
|
||||
SkippedManifestCount int `json:"skipped_manifest_count,omitempty"`
|
||||
// Orphaned blobs
|
||||
OrphanedBlobCount int `json:"orphaned_blob_count"`
|
||||
OrphanedBlobSize int64 `json:"orphaned_blob_size"`
|
||||
}
|
||||
|
||||
// RemoteInfo displays information about remote storage
|
||||
@@ -250,28 +234,16 @@ func (v *Vaultik) RemoteInfo(jsonOutput bool) error {
|
||||
v.stdoutf("Scanning snapshot metadata...\n")
|
||||
}
|
||||
|
||||
snapshotMetadata, snapshotIDs, skippedManifestCount, err := v.collectSnapshotMetadata()
|
||||
snapshotMetadata, snapshotIDs, err := v.collectSnapshotMetadata()
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
result.SkippedManifestCount = skippedManifestCount
|
||||
|
||||
if showText {
|
||||
manifestCount := 0
|
||||
|
||||
for _, info := range snapshotMetadata {
|
||||
if info.hasManifest {
|
||||
manifestCount++
|
||||
}
|
||||
}
|
||||
|
||||
v.stdoutf("Downloading %d manifest(s)...\n", manifestCount)
|
||||
v.stdoutf("Downloading %d manifest(s)...\n", len(snapshotIDs))
|
||||
}
|
||||
|
||||
referencedBlobs, unreadableManifests := v.collectReferencedBlobsFromManifests(
|
||||
snapshotIDs, snapshotMetadata)
|
||||
result.UnreadableManifests = unreadableManifests
|
||||
referencedBlobs := v.collectReferencedBlobsFromManifests(snapshotIDs, snapshotMetadata)
|
||||
|
||||
v.populateRemoteInfoResult(result, snapshotMetadata, snapshotIDs, referencedBlobs)
|
||||
|
||||
@@ -284,7 +256,7 @@ func (v *Vaultik) RemoteInfo(jsonOutput bool) error {
|
||||
"snapshots", result.TotalMetadataCount,
|
||||
"total_blobs", result.TotalBlobCount,
|
||||
"referenced_blobs", result.ReferencedBlobCount,
|
||||
"unreadable_manifests", len(result.UnreadableManifests))
|
||||
"orphaned_blobs", result.OrphanedBlobCount)
|
||||
|
||||
if jsonOutput {
|
||||
enc := json.NewEncoder(v.Stdout)
|
||||
@@ -301,18 +273,16 @@ func (v *Vaultik) RemoteInfo(jsonOutput bool) error {
|
||||
}
|
||||
|
||||
// collectSnapshotMetadata scans remote metadata and returns
|
||||
// per-snapshot info, sorted IDs and the number of manifests it skipped
|
||||
// because the name above them is not a remote key.
|
||||
// per-snapshot info and sorted IDs.
|
||||
func (v *Vaultik) collectSnapshotMetadata() (
|
||||
map[string]*SnapshotMetadataInfo, []string, int, error,
|
||||
map[string]*SnapshotMetadataInfo, []string, error,
|
||||
) {
|
||||
snapshotMetadata := make(map[string]*SnapshotMetadataInfo)
|
||||
skippedManifestCount := 0
|
||||
|
||||
metadataCh := v.Storage.ListStream(v.ctx, "metadata/")
|
||||
for obj := range metadataCh {
|
||||
if obj.Err != nil {
|
||||
return nil, nil, 0, fmt.Errorf("listing metadata: %w", obj.Err)
|
||||
return nil, nil, fmt.Errorf("listing metadata: %w", obj.Err)
|
||||
}
|
||||
|
||||
parts := strings.Split(obj.Key, "/")
|
||||
@@ -321,22 +291,6 @@ func (v *Vaultik) collectSnapshotMetadata() (
|
||||
}
|
||||
|
||||
snapshotID := parts[1]
|
||||
filename := parts[2]
|
||||
isManifest := filename == "manifest.json.zst"
|
||||
|
||||
// The name comes from the destination store, which is not
|
||||
// trusted, and is printed in the report. Accept it only in the
|
||||
// form of a remote key.
|
||||
if !isBlobHash(snapshotID) {
|
||||
log.Warn("Skipping non-conforming key under metadata/",
|
||||
"key", obj.Key)
|
||||
|
||||
if isManifest {
|
||||
skippedManifestCount++
|
||||
}
|
||||
|
||||
continue
|
||||
}
|
||||
|
||||
if _, exists := snapshotMetadata[snapshotID]; !exists {
|
||||
snapshotMetadata[snapshotID] = &SnapshotMetadataInfo{SnapshotID: snapshotID}
|
||||
@@ -344,10 +298,7 @@ func (v *Vaultik) collectSnapshotMetadata() (
|
||||
|
||||
info := snapshotMetadata[snapshotID]
|
||||
|
||||
if isManifest {
|
||||
info.hasManifest = true
|
||||
}
|
||||
|
||||
filename := parts[2]
|
||||
if strings.HasPrefix(filename, "manifest") {
|
||||
info.ManifestSize = obj.Size
|
||||
} else if strings.HasPrefix(filename, "db") {
|
||||
@@ -364,25 +315,17 @@ func (v *Vaultik) collectSnapshotMetadata() (
|
||||
|
||||
sort.Strings(snapshotIDs)
|
||||
|
||||
return snapshotMetadata, snapshotIDs, skippedManifestCount, nil
|
||||
return snapshotMetadata, snapshotIDs, nil
|
||||
}
|
||||
|
||||
// collectReferencedBlobsFromManifests downloads the listed manifests
|
||||
// and returns referenced blob hashes with sizes, and the remote keys
|
||||
// of the manifests it could not read.
|
||||
// collectReferencedBlobsFromManifests downloads manifests and returns
|
||||
// referenced blob hashes with sizes.
|
||||
func (v *Vaultik) collectReferencedBlobsFromManifests(
|
||||
snapshotIDs []string, snapshotMetadata map[string]*SnapshotMetadataInfo,
|
||||
) (map[string]int64, []string) {
|
||||
) map[string]int64 {
|
||||
referencedBlobs := make(map[string]int64)
|
||||
|
||||
var unreadable []string
|
||||
|
||||
for _, snapshotID := range snapshotIDs {
|
||||
info := snapshotMetadata[snapshotID]
|
||||
if !info.hasManifest {
|
||||
continue
|
||||
}
|
||||
|
||||
// snapshotIDs here are remote keys, taken straight from the
|
||||
// metadata/ listing. downloadManifestByKey is the single reader
|
||||
// for remote manifests; see its doc comment.
|
||||
@@ -390,11 +333,10 @@ func (v *Vaultik) collectReferencedBlobsFromManifests(
|
||||
if err != nil {
|
||||
log.Warn("Failed to read manifest", "snapshot", snapshotID, "error", err)
|
||||
|
||||
unreadable = append(unreadable, snapshotID)
|
||||
|
||||
continue
|
||||
}
|
||||
|
||||
info := snapshotMetadata[snapshotID]
|
||||
info.BlobCount = manifest.BlobCount
|
||||
|
||||
var blobsSize int64
|
||||
@@ -407,7 +349,7 @@ func (v *Vaultik) collectReferencedBlobsFromManifests(
|
||||
info.BlobsSize = blobsSize
|
||||
}
|
||||
|
||||
return referencedBlobs, unreadable
|
||||
return referencedBlobs
|
||||
}
|
||||
|
||||
// populateRemoteInfoResult fills in the result's snapshot and
|
||||
@@ -436,9 +378,8 @@ func (v *Vaultik) populateRemoteInfoResult(
|
||||
}
|
||||
|
||||
// scanRemoteBlobStorage lists all blobs on remote and computes orphan
|
||||
// stats when every listed manifest was read. showText is true only
|
||||
// when the human report is being printed (not --json, not --quiet),
|
||||
// gating the progress line.
|
||||
// stats. showText is true only when the human report is being printed
|
||||
// (not --json, not --quiet), gating the progress line.
|
||||
func (v *Vaultik) scanRemoteBlobStorage(
|
||||
result *RemoteInfoResult, referencedBlobs map[string]int64, showText bool,
|
||||
) error {
|
||||
@@ -465,28 +406,13 @@ func (v *Vaultik) scanRemoteBlobStorage(
|
||||
result.TotalBlobSize += obj.Size
|
||||
}
|
||||
|
||||
// A blob named only by a manifest that could not be read, or by one
|
||||
// under a skipped name, would be counted as orphaned, so the orphan
|
||||
// figures stay unknown.
|
||||
if len(result.UnreadableManifests) > 0 || result.SkippedManifestCount > 0 {
|
||||
return nil
|
||||
}
|
||||
|
||||
var (
|
||||
orphanedCount int
|
||||
orphanedSize int64
|
||||
)
|
||||
|
||||
for hash, size := range allBlobs {
|
||||
if _, referenced := referencedBlobs[hash]; !referenced {
|
||||
orphanedCount++
|
||||
orphanedSize += size
|
||||
result.OrphanedBlobCount++
|
||||
result.OrphanedBlobSize += size
|
||||
}
|
||||
}
|
||||
|
||||
result.OrphanedBlobCount = &orphanedCount
|
||||
result.OrphanedBlobSize = &orphanedSize
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
@@ -539,21 +465,11 @@ func (v *Vaultik) printRemoteInfoTable(result *RemoteInfoResult) {
|
||||
v.stdoutf("Referenced by snapshots: %s (%s)\n",
|
||||
humanize.Comma(int64(result.ReferencedBlobCount)),
|
||||
ubytes(result.ReferencedBlobSize))
|
||||
|
||||
if result.OrphanedBlobCount == nil {
|
||||
v.stdoutf("Orphaned (unreferenced): unknown "+
|
||||
"(%d manifest(s) could not be read, "+
|
||||
"%d manifest(s) under a non-conforming name skipped)\n",
|
||||
len(result.UnreadableManifests), result.SkippedManifestCount)
|
||||
|
||||
return
|
||||
}
|
||||
|
||||
v.stdoutf("Orphaned (unreferenced): %s (%s)\n",
|
||||
humanize.Comma(int64(*result.OrphanedBlobCount)),
|
||||
ubytes(*result.OrphanedBlobSize))
|
||||
humanize.Comma(int64(result.OrphanedBlobCount)),
|
||||
ubytes(result.OrphanedBlobSize))
|
||||
|
||||
if *result.OrphanedBlobCount > 0 {
|
||||
if result.OrphanedBlobCount > 0 {
|
||||
v.stdoutf("\nRun 'vaultik prune' to remove orphaned blobs.\n")
|
||||
}
|
||||
}
|
||||
|
||||
@@ -6,7 +6,6 @@ import (
|
||||
"io/fs"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
@@ -154,23 +153,3 @@ func TestPruneKeepsLocalRecordsWhenDestinationMissing(t *testing.T) {
|
||||
require.NoError(t, err)
|
||||
assert.Len(t, snapshots, 1, "prune must delete no local snapshot record")
|
||||
}
|
||||
|
||||
// TestPurgeSaysListingFailedOnceWhenDestinationMissing checks that
|
||||
// snapshot purge fails on a destination it cannot list, with an error
|
||||
// that says "listing remote snapshots" once.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestPurgeSaysListingFailedOnceWhenDestinationMissing(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
ctx := context.Background()
|
||||
v, _, _ := backUpThenUnplug(ctx, t)
|
||||
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
KeepLatest: true,
|
||||
Force: true,
|
||||
})
|
||||
require.ErrorIs(t, err, fs.ErrNotExist)
|
||||
assert.Equal(t, 1, strings.Count(err.Error(), "listing remote snapshots"),
|
||||
err.Error())
|
||||
}
|
||||
|
||||
@@ -42,7 +42,7 @@ func setupConsistencyTest(
|
||||
completedAt := startedAt.Add(5 * time.Minute)
|
||||
snap := &database.Snapshot{
|
||||
ID: types.SnapshotID(id),
|
||||
Hostname: snapHostname,
|
||||
Hostname: testHostname,
|
||||
VaultikVersion: testLabel,
|
||||
StartedAt: startedAt,
|
||||
CompletedAt: &completedAt,
|
||||
|
||||
@@ -17,10 +17,8 @@ import (
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
|
||||
// Snapshot IDs reused across the purge tests, and the hostname they were
|
||||
// taken on.
|
||||
// Snapshot IDs reused across the purge tests.
|
||||
const (
|
||||
snapHostname = "testhost"
|
||||
snapSystemT0 = "testhost_system_2026-01-01T00:00:00Z"
|
||||
snapHomeT0 = "testhost_home_2026-01-01T00:00:00Z"
|
||||
snapHomeT1 = "testhost_home_2026-01-01T01:00:00Z"
|
||||
@@ -28,12 +26,9 @@ const (
|
||||
)
|
||||
|
||||
// setupPurgeTest creates a Vaultik instance with an in-memory database and mock
|
||||
// storage pre-populated with the given snapshot IDs, all taken on hostname.
|
||||
// Each snapshot is marked as completed. Remote metadata stubs are created so
|
||||
// syncWithRemote keeps them.
|
||||
func setupPurgeTest(
|
||||
t *testing.T, hostname string, snapshotIDs []string,
|
||||
) *vaultik.Vaultik {
|
||||
// storage pre-populated with the given snapshot IDs. Each snapshot is marked as
|
||||
// completed. Remote metadata stubs are created so syncWithRemote keeps them.
|
||||
func setupPurgeTest(t *testing.T, snapshotIDs []string) *vaultik.Vaultik {
|
||||
t.Helper()
|
||||
|
||||
ctx := context.Background()
|
||||
@@ -56,7 +51,7 @@ func setupPurgeTest(
|
||||
completedAt := startedAt.Add(5 * time.Minute)
|
||||
snap := &database.Snapshot{
|
||||
ID: types.SnapshotID(id),
|
||||
Hostname: types.Hostname(hostname),
|
||||
Hostname: "testhost",
|
||||
VaultikVersion: testLabel,
|
||||
StartedAt: startedAt,
|
||||
CompletedAt: &completedAt,
|
||||
@@ -125,7 +120,7 @@ func TestPurgeKeepLatest_PerName(t *testing.T) {
|
||||
"testhost_system_2026-01-01T04:00:00Z",
|
||||
}
|
||||
|
||||
v := setupPurgeTest(t, snapHostname, snapshotIDs)
|
||||
v := setupPurgeTest(t, snapshotIDs)
|
||||
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
KeepLatest: true,
|
||||
@@ -153,7 +148,7 @@ func TestPurgeKeepLatest_SingleName(t *testing.T) {
|
||||
"testhost_home_2026-01-01T02:00:00Z",
|
||||
}
|
||||
|
||||
v := setupPurgeTest(t, snapHostname, snapshotIDs)
|
||||
v := setupPurgeTest(t, snapshotIDs)
|
||||
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
KeepLatest: true,
|
||||
@@ -181,7 +176,7 @@ func TestPurgeKeepLatest_WithNameFilter(t *testing.T) {
|
||||
"testhost_home_2026-01-01T04:00:00Z",
|
||||
}
|
||||
|
||||
v := setupPurgeTest(t, snapHostname, snapshotIDs)
|
||||
v := setupPurgeTest(t, snapshotIDs)
|
||||
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
KeepLatest: true,
|
||||
@@ -203,7 +198,7 @@ func TestPurgeKeepLatest_NoSnapshots(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
v := setupPurgeTest(t, snapHostname, nil)
|
||||
v := setupPurgeTest(t, nil)
|
||||
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
KeepLatest: true,
|
||||
@@ -221,7 +216,7 @@ func TestPurgeKeepLatest_NameFilterNoMatch(t *testing.T) {
|
||||
"testhost_system_2026-01-01T01:00:00Z",
|
||||
}
|
||||
|
||||
v := setupPurgeTest(t, snapHostname, snapshotIDs)
|
||||
v := setupPurgeTest(t, snapshotIDs)
|
||||
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
KeepLatest: true,
|
||||
@@ -248,7 +243,7 @@ func TestPurgeOlderThan_WithNameFilter(t *testing.T) {
|
||||
snapHomeT0,
|
||||
}
|
||||
|
||||
v := setupPurgeTest(t, snapHostname, snapshotIDs)
|
||||
v := setupPurgeTest(t, snapshotIDs)
|
||||
|
||||
// Purge only "home" snapshots older than 365 days
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
@@ -282,7 +277,7 @@ func TestPurgeKeepLatest_ThreeNames(t *testing.T) {
|
||||
"testhost_home_2026-01-01T06:00:00Z",
|
||||
}
|
||||
|
||||
v := setupPurgeTest(t, snapHostname, snapshotIDs)
|
||||
v := setupPurgeTest(t, snapshotIDs)
|
||||
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
KeepLatest: true,
|
||||
@@ -296,28 +291,3 @@ func TestPurgeKeepLatest_ThreeNames(t *testing.T) {
|
||||
assert.Contains(t, remaining, "testhost_system_2026-01-01T04:00:00Z")
|
||||
assert.Contains(t, remaining, "testhost_media_2026-01-01T05:00:00Z")
|
||||
}
|
||||
|
||||
// A hostname may contain underscores, so the snapshot name cannot be found
|
||||
// by splitting the ID at them. A purge by name must still select "docs".
|
||||
func TestPurgeKeepLatest_HostnameWithUnderscore(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
const (
|
||||
system = "my_host_system_2026-01-01T00:00:00Z"
|
||||
docsT1 = "my_host_docs_2026-01-01T01:00:00Z"
|
||||
docsT2 = "my_host_docs_2026-01-01T02:00:00Z"
|
||||
)
|
||||
|
||||
v := setupPurgeTest(t, "my_host", []string{system, docsT1, docsT2})
|
||||
|
||||
err := v.PurgeSnapshotsWithOptions(&vaultik.SnapshotPurgeOptions{
|
||||
KeepLatest: true,
|
||||
Force: true,
|
||||
Names: []string{"docs"},
|
||||
})
|
||||
require.NoError(t, err)
|
||||
|
||||
assert.ElementsMatch(t, []string{system, docsT2},
|
||||
listRemainingSnapshots(t, v))
|
||||
}
|
||||
|
||||
@@ -1,159 +0,0 @@
|
||||
package vaultik_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"encoding/json"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
)
|
||||
|
||||
// testBlobHashB is a blob that the manifest written by addRemote does
|
||||
// not reference.
|
||||
const testBlobHashB = "bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb" +
|
||||
"bbbbbbbbbbbbbbbbbbbbbbbbbbbbbbbb"
|
||||
|
||||
// TestRemoteInfo_UnreadableManifestLeavesOrphansUnknown checks that a
|
||||
// manifest remote info cannot read makes the orphan figures unknown. A
|
||||
// blob referenced only by that snapshot would otherwise be counted as
|
||||
// orphaned, and the report would advise running prune.
|
||||
func TestRemoteInfo_UnreadableManifestLeavesOrphansUnknown(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
env := newListEnv(t)
|
||||
|
||||
// The readable manifest references blob A only.
|
||||
env.addRemote(t, listRemoteID, time.Date(2026, 3, 2, 0, 0, 0, 0, time.UTC))
|
||||
addBlob(t, env.store.testStorer, testBlobHashA)
|
||||
addBlob(t, env.store.testStorer, testBlobHashB)
|
||||
|
||||
// With every manifest readable, blob B is orphaned.
|
||||
require.NoError(t, env.v.RemoteInfo(true))
|
||||
|
||||
var doc map[string]any
|
||||
|
||||
require.NoError(t, json.Unmarshal(env.stdout.Bytes(), &doc))
|
||||
assert.InDelta(t, 1, doc["orphaned_blob_count"], 0)
|
||||
|
||||
// A second snapshot whose manifest cannot be decoded. Blob B may be
|
||||
// one of its blobs.
|
||||
unreadableKey := snapshot.RemoteSnapshotKey(listLocalID)
|
||||
require.NoError(t, env.store.Put(context.Background(),
|
||||
"metadata/"+unreadableKey+"/manifest.json.zst",
|
||||
bytes.NewReader([]byte("not a valid manifest"))))
|
||||
|
||||
env.stdout.Reset()
|
||||
require.NoError(t, env.v.RemoteInfo(false))
|
||||
|
||||
text := env.stdout.String()
|
||||
assert.Contains(t, text, "Orphaned (unreferenced): unknown "+
|
||||
"(1 manifest(s) could not be read, "+
|
||||
"0 manifest(s) under a non-conforming name skipped)")
|
||||
assert.NotContains(t, text, "vaultik prune")
|
||||
|
||||
env.stdout.Reset()
|
||||
require.NoError(t, env.v.RemoteInfo(true))
|
||||
|
||||
doc = nil
|
||||
require.NoError(t, json.Unmarshal(env.stdout.Bytes(), &doc))
|
||||
assert.Contains(t, doc, "orphaned_blob_count")
|
||||
assert.Nil(t, doc["orphaned_blob_count"])
|
||||
assert.Contains(t, doc, "orphaned_blob_size")
|
||||
assert.Nil(t, doc["orphaned_blob_size"])
|
||||
assert.Equal(t, []any{unreadableKey}, doc["unreadable_manifests"])
|
||||
}
|
||||
|
||||
// TestRemoteInfo_SkipsNonConformingMetadataName checks that a directory
|
||||
// under metadata/ whose name is not a remote key is left out of the
|
||||
// report, and that the orphan figures are unknown when it holds a
|
||||
// manifest. The name comes from the destination store; printed raw, its
|
||||
// control characters would reach the terminal. Its manifest is not
|
||||
// read, so a blob only it references would otherwise be counted as
|
||||
// orphaned.
|
||||
func TestRemoteInfo_SkipsNonConformingMetadataName(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
env := newListEnv(t)
|
||||
env.addRemote(t, listRemoteID, time.Date(2026, 3, 2, 0, 0, 0, 0, time.UTC))
|
||||
addBlob(t, env.store.testStorer, testBlobHashA)
|
||||
addBlob(t, env.store.testStorer, testBlobHashB)
|
||||
require.NoError(t, env.store.Put(context.Background(),
|
||||
"metadata/\x1b[31mred/manifest.json.zst",
|
||||
bytes.NewReader([]byte("not a valid manifest"))))
|
||||
|
||||
require.NoError(t, env.v.RemoteInfo(false))
|
||||
|
||||
text := env.stdout.String()
|
||||
assert.NotContains(t, text, "\x1b")
|
||||
assert.NotContains(t, text, "31mred")
|
||||
assert.Contains(t, text, "Total (1 snapshots)")
|
||||
assert.Contains(t, text, "Orphaned (unreferenced): unknown "+
|
||||
"(0 manifest(s) could not be read, "+
|
||||
"1 manifest(s) under a non-conforming name skipped)")
|
||||
assert.NotContains(t, text, "vaultik prune")
|
||||
|
||||
env.stdout.Reset()
|
||||
require.NoError(t, env.v.RemoteInfo(true))
|
||||
|
||||
out := env.stdout.String()
|
||||
assert.NotContains(t, out, "31mred")
|
||||
|
||||
var doc map[string]any
|
||||
|
||||
require.NoError(t, json.Unmarshal([]byte(out), &doc))
|
||||
assert.Contains(t, doc, "orphaned_blob_count")
|
||||
assert.Nil(t, doc["orphaned_blob_count"])
|
||||
assert.Contains(t, doc, "orphaned_blob_size")
|
||||
assert.Nil(t, doc["orphaned_blob_size"])
|
||||
assert.InDelta(t, 1, doc["skipped_manifest_count"], 0)
|
||||
assert.NotContains(t, doc, "unreadable_manifests")
|
||||
}
|
||||
|
||||
// TestRemoteInfo_DirectoryWithoutManifestLeavesOrphansKnown checks that
|
||||
// a directory under metadata/ holding no manifest.json.zst, such as one
|
||||
// left by a backup interrupted before its manifest upload, leaves the
|
||||
// orphan figures known. prune does not treat such a directory as a
|
||||
// snapshot and deletes the blobs the report lists as orphaned.
|
||||
func TestRemoteInfo_DirectoryWithoutManifestLeavesOrphansKnown(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
t.Parallel()
|
||||
|
||||
env := newListEnv(t)
|
||||
env.addRemote(t, listRemoteID, time.Date(2026, 3, 2, 0, 0, 0, 0, time.UTC))
|
||||
addBlob(t, env.store.testStorer, testBlobHashA)
|
||||
addBlob(t, env.store.testStorer, testBlobHashB)
|
||||
|
||||
// One directory under a remote key and one under a non-conforming
|
||||
// name, each holding only a database.
|
||||
names := []string{snapshot.RemoteSnapshotKey(listLocalID), "\x1b[31mred"}
|
||||
for _, name := range names {
|
||||
require.NoError(t, env.store.Put(context.Background(),
|
||||
"metadata/"+name+"/db.zst.age",
|
||||
bytes.NewReader([]byte("not a valid database"))))
|
||||
}
|
||||
|
||||
require.NoError(t, env.v.RemoteInfo(false))
|
||||
|
||||
text := env.stdout.String()
|
||||
assert.NotContains(t, text, "\x1b")
|
||||
assert.Contains(t, text, "Downloading 1 manifest(s)...")
|
||||
assert.Contains(t, text, "Orphaned (unreferenced): 1 (")
|
||||
assert.Contains(t, text, "Run 'vaultik prune' to remove orphaned blobs.")
|
||||
|
||||
env.stdout.Reset()
|
||||
require.NoError(t, env.v.RemoteInfo(true))
|
||||
|
||||
var doc map[string]any
|
||||
|
||||
require.NoError(t, json.Unmarshal(env.stdout.Bytes(), &doc))
|
||||
assert.InDelta(t, 1, doc["orphaned_blob_count"], 0)
|
||||
assert.NotContains(t, doc, "unreadable_manifests")
|
||||
assert.NotContains(t, doc, "skipped_manifest_count")
|
||||
}
|
||||
@@ -1,73 +0,0 @@
|
||||
package vaultik_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
|
||||
// A file rewritten with its size unchanged and a new mtime in the same
|
||||
// second as the mtime the index holds must still be backed up. See
|
||||
// https://git.eeqj.de/sneak/vaultik/issues/226.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestBackupOfSameSecondRewriteRestoresNewContent(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
dataDir := filepath.Join(tempDir, "src")
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
restoreDir := filepath.Join(tempDir, "restored")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
rewrittenPath := filepath.Join(dataDir, "small.txt")
|
||||
|
||||
ctx := context.Background()
|
||||
files := writeFaultSourceTree(t, fs, dataDir)
|
||||
cfg := changedFileConfig(dataDir, dbPath)
|
||||
|
||||
firstMTime := time.Date(2026, time.January, 2, 3, 4, 5, 0, time.UTC).
|
||||
Add(100 * time.Millisecond)
|
||||
secondMTime := firstMTime.Add(800 * time.Millisecond)
|
||||
|
||||
require.NoError(t, fs.Chtimes(rewrittenPath, firstMTime, firstMTime))
|
||||
|
||||
store, err := storage.NewFileStorer(storeDir)
|
||||
require.NoError(t, err)
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
|
||||
|
||||
require.NoError(t, backUp(v, "first"))
|
||||
|
||||
// Upper-casing ASCII text keeps its size.
|
||||
files[rewrittenPath] = bytes.ToUpper(files[rewrittenPath])
|
||||
require.NoError(t, afero.WriteFile(fs, rewrittenPath, files[rewrittenPath], 0o644))
|
||||
require.NoError(t, fs.Chtimes(rewrittenPath, secondMTime, secondMTime))
|
||||
|
||||
require.NoError(t, backUp(v, "second"))
|
||||
|
||||
id := localSnapshotID(ctx, t, repos, "second")
|
||||
require.NoError(t, db.Close())
|
||||
|
||||
reader := newReaderVaultik(ctx, cfg, store, nil, fs)
|
||||
require.NoError(t, reader.Restore(&vaultik.RestoreOptions{
|
||||
SnapshotID: id,
|
||||
TargetDir: restoreDir,
|
||||
Verify: true,
|
||||
}))
|
||||
|
||||
assertRestoredTree(t, fs, restoreDir, files)
|
||||
}
|
||||
@@ -14,7 +14,6 @@ import (
|
||||
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/snapshot"
|
||||
"sneak.berlin/go/vaultik/internal/types"
|
||||
)
|
||||
|
||||
// Sentinel errors for snapshot management.
|
||||
@@ -190,11 +189,6 @@ type snapshotStats struct {
|
||||
totalBytesUploaded int64
|
||||
totalBlobsUploaded int
|
||||
uploadDuration time.Duration
|
||||
|
||||
// The sizes of all blobs the snapshot references, set by
|
||||
// finalizeSnapshotMetadata once snapshot_blobs is populated.
|
||||
blobSize int64
|
||||
blobUncompressedSize int64
|
||||
}
|
||||
|
||||
// createNamedSnapshot creates a single named snapshot
|
||||
@@ -234,6 +228,8 @@ func (v *Vaultik) createNamedSnapshot(
|
||||
return err
|
||||
}
|
||||
|
||||
v.collectUploadStats(scanner, stats)
|
||||
|
||||
err = v.finalizeSnapshotMetadata(snapshotID, stats)
|
||||
if err != nil {
|
||||
return err
|
||||
@@ -285,11 +281,6 @@ func (v *Vaultik) resolveSnapshotPaths(snapName string) ([]string, error) {
|
||||
func (v *Vaultik) scanAllDirectories(
|
||||
scanner *snapshot.Scanner, resolvedDirs []string, snapshotID string,
|
||||
) (*snapshotStats, error) {
|
||||
if progress := scanner.GetProgress(); progress != nil {
|
||||
progress.Start()
|
||||
defer progress.Stop()
|
||||
}
|
||||
|
||||
stats := &snapshotStats{}
|
||||
|
||||
for i, dir := range resolvedDirs {
|
||||
@@ -318,9 +309,6 @@ func (v *Vaultik) scanAllDirectories(
|
||||
stats.totalBytesSkipped += result.BytesSkipped
|
||||
stats.totalFilesDeleted += result.FilesDeleted
|
||||
stats.totalBytesDeleted += result.BytesDeleted
|
||||
stats.totalBlobsUploaded += result.BlobsUploaded
|
||||
stats.totalBytesUploaded += result.BytesUploaded
|
||||
stats.uploadDuration += result.UploadDuration
|
||||
|
||||
log.Info("Directory scan complete",
|
||||
"path", dir,
|
||||
@@ -336,6 +324,18 @@ func (v *Vaultik) scanAllDirectories(
|
||||
return stats, nil
|
||||
}
|
||||
|
||||
// collectUploadStats gathers upload statistics from the scanner's
|
||||
// progress reporter.
|
||||
func (v *Vaultik) collectUploadStats(scanner *snapshot.Scanner, stats *snapshotStats) {
|
||||
if s := scanner.GetProgress(); s != nil {
|
||||
progressStats := s.GetStats()
|
||||
stats.totalBytesUploaded = progressStats.BytesUploaded.Load()
|
||||
stats.totalBlobsUploaded = int(progressStats.BlobsUploaded.Load())
|
||||
stats.uploadDuration = time.Duration(
|
||||
progressStats.UploadDurationMs.Load()) * time.Millisecond
|
||||
}
|
||||
}
|
||||
|
||||
// finalizeSnapshotMetadata updates stats, exports metadata, and only then
|
||||
// marks the snapshot complete. Recording completion last is deliberate: an
|
||||
// export interrupted by a crash leaves the snapshot incomplete rather than
|
||||
@@ -345,39 +345,31 @@ func (v *Vaultik) scanAllDirectories(
|
||||
func (v *Vaultik) finalizeSnapshotMetadata(
|
||||
snapshotID string, stats *snapshotStats,
|
||||
) error {
|
||||
// snapshot_blobs must be populated before the blob sizes below, which
|
||||
// total the snapshot's blobs, and before the export, which builds the
|
||||
// manifest and the trimmed metadata database from it.
|
||||
err := v.SnapshotManager.PopulateSnapshotBlobs(v.ctx, snapshotID)
|
||||
if err != nil {
|
||||
return fmt.Errorf("populating snapshot blobs: %w", err)
|
||||
}
|
||||
|
||||
stats.blobSize, stats.blobUncompressedSize, err =
|
||||
v.Repositories.Snapshots.GetSnapshotBlobSizes(v.ctx, snapshotID)
|
||||
if err != nil {
|
||||
return fmt.Errorf("getting snapshot blob sizes: %w", err)
|
||||
}
|
||||
|
||||
extStats := snapshot.ExtendedBackupStats{
|
||||
BackupStats: snapshot.BackupStats{
|
||||
FilesScanned: stats.totalFiles,
|
||||
TotalSize: stats.totalBytes + stats.totalBytesSkipped,
|
||||
BytesScanned: stats.totalBytes,
|
||||
ChunksCreated: stats.totalChunks,
|
||||
BlobsCreated: stats.totalBlobs,
|
||||
BytesUploaded: stats.totalBytesUploaded,
|
||||
},
|
||||
BlobSize: stats.blobSize,
|
||||
BlobUncompressedSize: stats.blobUncompressedSize,
|
||||
BlobUncompressedSize: 0,
|
||||
CompressionLevel: v.Config.CompressionLevel,
|
||||
UploadDurationMs: stats.uploadDuration.Milliseconds(),
|
||||
}
|
||||
|
||||
err = v.SnapshotManager.UpdateSnapshotStatsExtended(v.ctx, snapshotID, extStats)
|
||||
err := v.SnapshotManager.UpdateSnapshotStatsExtended(v.ctx, snapshotID, extStats)
|
||||
if err != nil {
|
||||
return fmt.Errorf("updating snapshot stats: %w", err)
|
||||
}
|
||||
|
||||
// snapshot_blobs must be populated before the export, which builds the
|
||||
// manifest and the trimmed metadata database from it.
|
||||
err = v.SnapshotManager.PopulateSnapshotBlobs(v.ctx, snapshotID)
|
||||
if err != nil {
|
||||
return fmt.Errorf("populating snapshot blobs: %w", err)
|
||||
}
|
||||
|
||||
err = v.SnapshotManager.ExportSnapshotMetadata(
|
||||
v.ctx, v.Config.IndexPath, snapshotID)
|
||||
if err != nil {
|
||||
@@ -412,10 +404,12 @@ func (v *Vaultik) printSnapshotSummary(
|
||||
totalFilesChanged := stats.totalFiles - stats.totalFilesSkipped
|
||||
totalBytesAll := stats.totalBytes + stats.totalBytesSkipped
|
||||
|
||||
// Get total blob sizes from database
|
||||
compressedSize, uncompressedSize := v.getSnapshotBlobSizes(snapshotID)
|
||||
|
||||
var compressionRatio float64
|
||||
if stats.blobUncompressedSize > 0 {
|
||||
compressionRatio = float64(stats.blobSize) /
|
||||
float64(stats.blobUncompressedSize)
|
||||
if uncompressedSize > 0 {
|
||||
compressionRatio = float64(compressedSize) / float64(uncompressedSize)
|
||||
} else {
|
||||
compressionRatio = 1.0
|
||||
}
|
||||
@@ -443,8 +437,8 @@ func (v *Vaultik) printSnapshotSummary(
|
||||
|
||||
if stats.totalBlobsUploaded > 0 {
|
||||
v.UI.Detailf("Storage: %s compressed from %s (%.2fx ratio).",
|
||||
v.UI.Size(stats.blobSize),
|
||||
v.UI.Size(stats.blobUncompressedSize),
|
||||
v.UI.Size(compressedSize),
|
||||
v.UI.Size(uncompressedSize),
|
||||
compressionRatio)
|
||||
v.UI.Detailf("Upload: %d blobs, %s in %s (%s).",
|
||||
stats.totalBlobsUploaded,
|
||||
@@ -456,6 +450,27 @@ func (v *Vaultik) printSnapshotSummary(
|
||||
v.UI.Detailf("Snapshot create duration: %s.", v.UI.Duration(snapshotDuration))
|
||||
}
|
||||
|
||||
// getSnapshotBlobSizes returns total compressed and uncompressed blob
|
||||
// sizes for a snapshot.
|
||||
func (v *Vaultik) getSnapshotBlobSizes(snapshotID string) (int64, int64) {
|
||||
var compressed, uncompressed int64
|
||||
|
||||
blobHashes, err := v.Repositories.Snapshots.GetBlobHashes(v.ctx, snapshotID)
|
||||
if err != nil {
|
||||
return 0, 0
|
||||
}
|
||||
|
||||
for _, hash := range blobHashes {
|
||||
blob, err := v.Repositories.Blobs.GetByHash(v.ctx, hash)
|
||||
if err == nil && blob != nil {
|
||||
compressed += blob.CompressedSize
|
||||
uncompressed += blob.UncompressedSize
|
||||
}
|
||||
}
|
||||
|
||||
return compressed, uncompressed
|
||||
}
|
||||
|
||||
// SnapshotPurgeOptions contains options for the snapshot purge command.
|
||||
type SnapshotPurgeOptions struct {
|
||||
KeepLatest bool // Keep only the most recent snapshot per name
|
||||
@@ -496,23 +511,19 @@ func (v *Vaultik) PurgeSnapshotsWithOptions(opts *SnapshotPurgeOptions) error {
|
||||
nameFilter[n] = struct{}{}
|
||||
}
|
||||
|
||||
// Collect completed snapshots and their names, applying the name filter.
|
||||
// Collect completed snapshots, applying the name filter.
|
||||
snapshots := make([]SnapshotInfo, 0, len(dbSnapshots))
|
||||
names := make(map[types.SnapshotID]string, len(dbSnapshots))
|
||||
|
||||
for _, s := range dbSnapshots {
|
||||
if s.CompletedAt == nil {
|
||||
continue
|
||||
}
|
||||
|
||||
name := parseSnapshotName(s.ID.String(), s.Hostname.String())
|
||||
if len(nameFilter) > 0 {
|
||||
if _, ok := nameFilter[name]; !ok {
|
||||
if _, ok := nameFilter[parseSnapshotName(s.ID.String())]; !ok {
|
||||
continue
|
||||
}
|
||||
}
|
||||
|
||||
names[s.ID] = name
|
||||
snapshots = append(snapshots, SnapshotInfo{
|
||||
ID: s.ID,
|
||||
Timestamp: s.StartedAt,
|
||||
@@ -525,7 +536,7 @@ func (v *Vaultik) PurgeSnapshotsWithOptions(opts *SnapshotPurgeOptions) error {
|
||||
return snapshots[i].Timestamp.After(snapshots[j].Timestamp)
|
||||
})
|
||||
|
||||
toDelete, err := selectSnapshotsToPurge(snapshots, names, opts)
|
||||
toDelete, err := selectSnapshotsToPurge(snapshots, opts)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
@@ -543,11 +554,9 @@ func (v *Vaultik) PurgeSnapshotsWithOptions(opts *SnapshotPurgeOptions) error {
|
||||
|
||||
// selectSnapshotsToPurge applies the purge retention criteria to the
|
||||
// newest-first sorted snapshot list and returns the deletion
|
||||
// candidates. names maps each snapshot's ID to its snapshot name.
|
||||
// candidates.
|
||||
func selectSnapshotsToPurge(
|
||||
snapshots []SnapshotInfo,
|
||||
names map[types.SnapshotID]string,
|
||||
opts *SnapshotPurgeOptions,
|
||||
snapshots []SnapshotInfo, opts *SnapshotPurgeOptions,
|
||||
) ([]SnapshotInfo, error) {
|
||||
var toDelete []SnapshotInfo
|
||||
|
||||
@@ -558,7 +567,7 @@ func selectSnapshotsToPurge(
|
||||
seen := make(map[string]bool)
|
||||
|
||||
for _, snap := range snapshots {
|
||||
name := names[snap.ID]
|
||||
name := parseSnapshotName(snap.ID.String())
|
||||
if seen[name] {
|
||||
toDelete = append(toDelete, snap)
|
||||
|
||||
@@ -1052,7 +1061,7 @@ func (v *Vaultik) syncWithRemote() error {
|
||||
// every local snapshot record (issue #160).
|
||||
remoteKeys, err := v.listAllRemoteSnapshotKeys()
|
||||
if err != nil {
|
||||
return err
|
||||
return fmt.Errorf("listing remote snapshots: %w", err)
|
||||
}
|
||||
|
||||
remoteKeySet := make(map[string]bool, len(remoteKeys))
|
||||
@@ -1117,12 +1126,6 @@ type RemoveResult struct {
|
||||
// just-removed snapshot left behind on the destination store.
|
||||
const pruneCommandHint = "vaultik prune"
|
||||
|
||||
// snapshotRemoveCommandHint is the command suggested, with the
|
||||
// snapshot's ID, when a remove could not reach the destination store:
|
||||
// running it again removes the snapshot's metadata there, which
|
||||
// `vaultik prune` never does.
|
||||
const snapshotRemoveCommandHint = "vaultik snapshot remove"
|
||||
|
||||
// RemoveSnapshot removes a snapshot from the local index database and,
|
||||
// unless LocalOnly is set, also strips the snapshot's metadata from the
|
||||
// destination store. Blobs are NOT touched: removing a snapshot's
|
||||
@@ -1159,7 +1162,7 @@ func (v *Vaultik) RemoveSnapshot(
|
||||
}
|
||||
|
||||
if !opts.LocalOnly {
|
||||
result.RemoteRemoved = v.removeSnapshotRemote(snapshotID, opts)
|
||||
result.RemoteRemoved = v.removeSnapshotRemote(snapshotID)
|
||||
}
|
||||
|
||||
if v.SnapshotManager != nil {
|
||||
@@ -1242,11 +1245,9 @@ func (v *Vaultik) confirmRemoveSnapshot(snapshotID string, opts *RemoveOptions)
|
||||
// removeSnapshotRemote strips the snapshot's metadata from the
|
||||
// destination store, warning and proceeding on failure: the local-DB
|
||||
// removal has already happened, so the user is told the remote half
|
||||
// didn't finish and to run `vaultik snapshot remove` for the snapshot
|
||||
// again once the destination store is reachable (`vaultik prune` never
|
||||
// removes snapshot metadata). Returns true when the remote removal
|
||||
// succeeded.
|
||||
func (v *Vaultik) removeSnapshotRemote(snapshotID string, opts *RemoveOptions) bool {
|
||||
// didn't finish and can retry with `vaultik prune` once the destination
|
||||
// store is reachable. Returns true when the remote removal succeeded.
|
||||
func (v *Vaultik) removeSnapshotRemote(snapshotID string) bool {
|
||||
log.Info("Removing snapshot metadata from remote storage",
|
||||
"snapshot_id", snapshotID)
|
||||
|
||||
@@ -1254,17 +1255,13 @@ func (v *Vaultik) removeSnapshotRemote(snapshotID string, opts *RemoveOptions) b
|
||||
|
||||
err := v.deleteRemoteSnapshotByKey(remoteKey)
|
||||
if err != nil {
|
||||
log.Warn("Could not remove snapshot metadata from remote storage; "+
|
||||
"run '"+snapshotRemoveCommandHint+"' with the snapshot's ID "+
|
||||
"again once the remote is reachable",
|
||||
"snapshot_id", snapshotID, "error", err)
|
||||
log.Warn("Could not remove snapshot metadata from remote storage",
|
||||
"error", err)
|
||||
|
||||
// The UI writes to stdout, which under --json holds only the
|
||||
// document; the log record above is the warning on stderr.
|
||||
if v.UI != nil && !opts.JSON {
|
||||
if v.UI != nil {
|
||||
v.UI.Warningf("Could not remove snapshot metadata from remote: "+
|
||||
"%v. Run '%s %s' again once the remote is reachable.",
|
||||
err, snapshotRemoveCommandHint, snapshotID)
|
||||
"%v. Run '%s' once the remote is reachable to finish cleanup.",
|
||||
err, pruneCommandHint)
|
||||
}
|
||||
|
||||
return false
|
||||
|
||||
@@ -1,285 +0,0 @@
|
||||
package vaultik_test
|
||||
|
||||
import (
|
||||
"bytes"
|
||||
"context"
|
||||
"fmt"
|
||||
"os"
|
||||
"path/filepath"
|
||||
"strings"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"github.com/stretchr/testify/assert"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/config"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
"sneak.berlin/go/vaultik/internal/storage/faultstore"
|
||||
"sneak.berlin/go/vaultik/internal/ui"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
|
||||
// These tests cover https://git.eeqj.de/sneak/vaultik/issues/225: the
|
||||
// summary printed after a backup, and the statistics stored in the
|
||||
// snapshots table, count each file, byte and upload once, and a --cron
|
||||
// run records its uploads.
|
||||
|
||||
// summaryUploadDelay slows every blob upload, so a run's upload time is
|
||||
// at least this long per blob even on a local store.
|
||||
const summaryUploadDelay = 20 * time.Millisecond
|
||||
|
||||
// summaryEnv is a backup setup whose user-facing output is kept in out.
|
||||
type summaryEnv struct {
|
||||
v *vaultik.Vaultik
|
||||
db *database.DB
|
||||
repos *database.Repositories
|
||||
out *bytes.Buffer
|
||||
|
||||
// aPath is a.bin, whose content copy.bin repeats; aSize is its size
|
||||
// and totalSize the size of all three source files.
|
||||
aPath string
|
||||
aSize int64
|
||||
totalSize int64
|
||||
}
|
||||
|
||||
// newSummaryEnv writes src/one/a.bin, src/one/small.txt and
|
||||
// src/two/copy.bin, a copy of a.bin. Every chunk of copy.bin is therefore
|
||||
// already stored by the time the backup reaches it.
|
||||
//
|
||||
// The snapshot names "first" and "second" back up src; "split" backs up
|
||||
// src/one and src/two as two paths.
|
||||
func newSummaryEnv(t *testing.T) *summaryEnv {
|
||||
t.Helper()
|
||||
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
srcDir := filepath.Join(tempDir, "src")
|
||||
dirOne := filepath.Join(srcDir, "one")
|
||||
dirTwo := filepath.Join(srcDir, "two")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
ctx := context.Background()
|
||||
|
||||
aContent := bytesPattern("a-", int(3*faultChunkSize))
|
||||
smallContent := []byte("hello vaultik")
|
||||
files := map[string][]byte{
|
||||
filepath.Join(dirOne, "a.bin"): aContent,
|
||||
filepath.Join(dirOne, "small.txt"): smallContent,
|
||||
filepath.Join(dirTwo, "copy.bin"): aContent,
|
||||
}
|
||||
|
||||
for path, content := range files {
|
||||
require.NoError(t, fs.MkdirAll(filepath.Dir(path), 0o755))
|
||||
require.NoError(t, afero.WriteFile(fs, path, content, 0o644))
|
||||
}
|
||||
|
||||
cfg := faultTestConfig()
|
||||
cfg.IndexPath = dbPath
|
||||
cfg.ChunkSize = config.Size(faultChunkSize)
|
||||
cfg.Snapshots = map[string]config.SnapshotConfig{
|
||||
"first": {Paths: []string{srcDir}},
|
||||
"second": {Paths: []string{srcDir}},
|
||||
"split": {Paths: []string{dirOne, dirTwo}},
|
||||
}
|
||||
|
||||
inner, err := storage.NewFileStorer(filepath.Join(tempDir, "remote"))
|
||||
require.NoError(t, err)
|
||||
|
||||
store := faultstore.New(inner)
|
||||
store.OnPut = func(key string) faultstore.PutAction {
|
||||
if strings.HasPrefix(key, "blobs/") {
|
||||
time.Sleep(summaryUploadDelay)
|
||||
}
|
||||
|
||||
return faultstore.PutNormal
|
||||
}
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
t.Cleanup(func() { _ = db.Close() })
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
out := &bytes.Buffer{}
|
||||
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
|
||||
v.UI = ui.NewWithColor(out, false)
|
||||
|
||||
return &summaryEnv{
|
||||
v: v,
|
||||
db: db,
|
||||
repos: repos,
|
||||
out: out,
|
||||
aPath: filepath.Join(dirOne, "a.bin"),
|
||||
aSize: int64(len(aContent)),
|
||||
totalSize: int64(2*len(aContent) + len(smallContent)),
|
||||
}
|
||||
}
|
||||
|
||||
// backUp runs a backup of the named snapshot and returns its output.
|
||||
func (e *summaryEnv) backUp(t *testing.T, name string, cron bool) string {
|
||||
t.Helper()
|
||||
|
||||
e.out.Reset()
|
||||
require.NoError(t, e.v.CreateSnapshot(&vaultik.SnapshotCreateOptions{
|
||||
Cron: cron,
|
||||
Snapshots: []string{name},
|
||||
}))
|
||||
|
||||
return e.out.String()
|
||||
}
|
||||
|
||||
// snapshot returns the local snapshots row of the snapshot named name.
|
||||
func (e *summaryEnv) snapshot(t *testing.T, name string) *database.Snapshot {
|
||||
t.Helper()
|
||||
|
||||
ctx := context.Background()
|
||||
snap, err := e.repos.Snapshots.GetByID(ctx,
|
||||
localSnapshotID(ctx, t, e.repos, name))
|
||||
require.NoError(t, err)
|
||||
require.NotNil(t, snap)
|
||||
|
||||
return snap
|
||||
}
|
||||
|
||||
// uploads returns how many blobs the snapshot uploaded and their
|
||||
// total size, as recorded in the uploads table.
|
||||
func (e *summaryEnv) uploads(t *testing.T, snapshotID string) (int64, int64) {
|
||||
t.Helper()
|
||||
|
||||
var count, size int64
|
||||
|
||||
err := e.db.Conn().QueryRowContext(context.Background(), `
|
||||
SELECT COUNT(*), COALESCE(SUM(size), 0)
|
||||
FROM uploads WHERE snapshot_id = ?`, snapshotID).Scan(&count, &size)
|
||||
require.NoError(t, err)
|
||||
|
||||
return count, size
|
||||
}
|
||||
|
||||
// referencedBlobSizes returns the compressed and uncompressed sizes of
|
||||
// all blobs the snapshot references.
|
||||
func (e *summaryEnv) referencedBlobSizes(
|
||||
t *testing.T, snapshotID string,
|
||||
) (int64, int64) {
|
||||
t.Helper()
|
||||
|
||||
var compressed, uncompressed int64
|
||||
|
||||
err := e.db.Conn().QueryRowContext(context.Background(), `
|
||||
SELECT COALESCE(SUM(b.compressed_size), 0),
|
||||
COALESCE(SUM(b.uncompressed_size), 0)
|
||||
FROM snapshot_blobs sb JOIN blobs b ON b.blob_hash = sb.blob_hash
|
||||
WHERE sb.snapshot_id = ?`, snapshotID).Scan(&compressed, &uncompressed)
|
||||
require.NoError(t, err)
|
||||
|
||||
return compressed, uncompressed
|
||||
}
|
||||
|
||||
// filesLine returns the summary's line of file counts.
|
||||
func filesLine(examined, backedUp, unchanged int) string {
|
||||
return fmt.Sprintf("Files: %d examined, %d backed up, %d unchanged.",
|
||||
examined, backedUp, unchanged)
|
||||
}
|
||||
|
||||
// dataLine returns the summary's line of byte counts.
|
||||
func (e *summaryEnv) dataLine(total, backedUp int64) string {
|
||||
return fmt.Sprintf("Data: %s total (%s backed up).",
|
||||
e.v.UI.Size(total), e.v.UI.Size(backedUp))
|
||||
}
|
||||
|
||||
// A first backup stores copy.bin's chunks while backing up a.bin, so
|
||||
// copy.bin's chunks are deduplicated within the run. Each file and byte
|
||||
// is still counted once.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestSnapshotSummaryFirstRun(t *testing.T) {
|
||||
env := newSummaryEnv(t)
|
||||
|
||||
summary := env.backUp(t, "first", false)
|
||||
|
||||
assert.Contains(t, summary, filesLine(3, 3, 0))
|
||||
assert.Contains(t, summary, env.dataLine(env.totalSize, env.totalSize))
|
||||
|
||||
snap := env.snapshot(t, "first")
|
||||
uploadCount, uploadBytes := env.uploads(t, snap.ID.String())
|
||||
require.Positive(t, uploadCount)
|
||||
|
||||
assert.Contains(t, summary, fmt.Sprintf("Upload: %d blobs, %s in ",
|
||||
uploadCount, env.v.UI.Size(uploadBytes)))
|
||||
|
||||
assert.Equal(t, int64(3), snap.FileCount)
|
||||
assert.Equal(t, env.totalSize, snap.TotalSize)
|
||||
assert.Equal(t, uploadCount, snap.BlobCount)
|
||||
assert.Equal(t, uploadBytes, snap.UploadBytes)
|
||||
}
|
||||
|
||||
// An incremental backup where a.bin's mtime changed but its content did
|
||||
// not: a.bin is backed up again and every one of its chunks is already
|
||||
// stored.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestSnapshotSummaryIncrementalRunWithDeduplicatedChunks(t *testing.T) {
|
||||
env := newSummaryEnv(t)
|
||||
|
||||
env.backUp(t, "first", false)
|
||||
|
||||
later := time.Now().Add(time.Hour)
|
||||
require.NoError(t, os.Chtimes(env.aPath, later, later))
|
||||
|
||||
summary := env.backUp(t, "second", false)
|
||||
|
||||
assert.Contains(t, summary, filesLine(3, 1, 2))
|
||||
assert.Contains(t, summary, env.dataLine(env.totalSize, env.aSize))
|
||||
assert.NotContains(t, summary, "Upload:")
|
||||
|
||||
snap := env.snapshot(t, "second")
|
||||
compressed, uncompressed := env.referencedBlobSizes(t, snap.ID.String())
|
||||
require.Positive(t, compressed)
|
||||
|
||||
assert.Equal(t, env.totalSize, snap.TotalSize)
|
||||
assert.Zero(t, snap.ChunkCount)
|
||||
assert.Zero(t, snap.BlobCount)
|
||||
assert.Zero(t, snap.UploadBytes)
|
||||
assert.Equal(t, compressed, snap.BlobSize,
|
||||
"blob_size must total the blobs the snapshot references")
|
||||
assert.Equal(t, uncompressed, snap.BlobUncompressedSize)
|
||||
}
|
||||
|
||||
// Under --cron the progress reporter is off; the upload figures must
|
||||
// still reach the summary and the snapshots row. The snapshot has two
|
||||
// paths, each backed up by its own scan.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestSnapshotSummaryCronRunRecordsUploads(t *testing.T) {
|
||||
env := newSummaryEnv(t)
|
||||
|
||||
summary := env.backUp(t, "split", true)
|
||||
|
||||
snap := env.snapshot(t, "split")
|
||||
uploadCount, uploadBytes := env.uploads(t, snap.ID.String())
|
||||
require.Positive(t, uploadCount)
|
||||
|
||||
assert.Contains(t, summary, filesLine(3, 3, 0))
|
||||
assert.Contains(t, summary, env.dataLine(env.totalSize, env.totalSize))
|
||||
assert.Contains(t, summary, fmt.Sprintf("Upload: %d blobs, %s in ",
|
||||
uploadCount, env.v.UI.Size(uploadBytes)))
|
||||
|
||||
assert.Equal(t, env.totalSize, snap.TotalSize)
|
||||
assert.Equal(t, uploadCount, snap.BlobCount,
|
||||
"blob_count must count each blob once, however many paths the "+
|
||||
"snapshot has")
|
||||
assert.Equal(t, uploadBytes, snap.UploadBytes)
|
||||
assert.GreaterOrEqual(t, snap.UploadDurationMs,
|
||||
uploadCount*summaryUploadDelay.Milliseconds())
|
||||
|
||||
compressed, uncompressed := env.referencedBlobSizes(t, snap.ID.String())
|
||||
require.Positive(t, uncompressed)
|
||||
|
||||
assert.Equal(t, compressed, snap.BlobSize)
|
||||
assert.Equal(t, uncompressed, snap.BlobUncompressedSize)
|
||||
assert.InDelta(t, float64(compressed)/float64(uncompressed),
|
||||
snap.CompressionRatio, 1e-9)
|
||||
}
|
||||
@@ -1,71 +0,0 @@
|
||||
package vaultik_test
|
||||
|
||||
import (
|
||||
"context"
|
||||
"maps"
|
||||
"path/filepath"
|
||||
"testing"
|
||||
|
||||
"github.com/spf13/afero"
|
||||
"github.com/stretchr/testify/require"
|
||||
"sneak.berlin/go/vaultik/internal/config"
|
||||
"sneak.berlin/go/vaultik/internal/database"
|
||||
"sneak.berlin/go/vaultik/internal/log"
|
||||
"sneak.berlin/go/vaultik/internal/storage"
|
||||
"sneak.berlin/go/vaultik/internal/vaultik"
|
||||
)
|
||||
|
||||
// A backup without --cron runs the progress reporter while one scanner
|
||||
// scans each path of the snapshot in turn. See
|
||||
// https://git.eeqj.de/sneak/vaultik/issues/253.
|
||||
//
|
||||
//nolint:paralleltest // installs the global logger via log.Initialize
|
||||
func TestBackupWithoutCronOfTwoPathSnapshotRestoresBothPaths(t *testing.T) {
|
||||
log.Initialize(log.Config{})
|
||||
|
||||
const snapshotName = "data"
|
||||
|
||||
fs := afero.NewOsFs()
|
||||
tempDir := t.TempDir()
|
||||
firstDir := filepath.Join(tempDir, "first")
|
||||
secondDir := filepath.Join(tempDir, "second")
|
||||
storeDir := filepath.Join(tempDir, "remote")
|
||||
restoreDir := filepath.Join(tempDir, "restored")
|
||||
dbPath := filepath.Join(tempDir, "index.sqlite")
|
||||
|
||||
ctx := context.Background()
|
||||
files := writeFaultSourceTree(t, fs, firstDir)
|
||||
maps.Copy(files, writeFaultSourceTree(t, fs, secondDir))
|
||||
|
||||
cfg := faultTestConfig()
|
||||
cfg.IndexPath = dbPath
|
||||
cfg.ChunkSize = config.Size(faultChunkSize)
|
||||
cfg.Snapshots = map[string]config.SnapshotConfig{
|
||||
snapshotName: {Paths: []string{firstDir, secondDir}},
|
||||
}
|
||||
|
||||
store, err := storage.NewFileStorer(storeDir)
|
||||
require.NoError(t, err)
|
||||
|
||||
db, err := database.New(ctx, dbPath)
|
||||
require.NoError(t, err)
|
||||
|
||||
repos := database.NewRepositories(db)
|
||||
v := newBackupVaultik(ctx, cfg, store, repos, db, fs)
|
||||
|
||||
require.NoError(t, v.CreateSnapshot(&vaultik.SnapshotCreateOptions{
|
||||
Snapshots: []string{snapshotName},
|
||||
}))
|
||||
|
||||
id := localSnapshotID(ctx, t, repos, snapshotName)
|
||||
require.NoError(t, db.Close())
|
||||
|
||||
reader := newReaderVaultik(ctx, cfg, store, nil, fs)
|
||||
require.NoError(t, reader.Restore(&vaultik.RestoreOptions{
|
||||
SnapshotID: id,
|
||||
TargetDir: restoreDir,
|
||||
Verify: true,
|
||||
}))
|
||||
|
||||
assertRestoredTree(t, fs, restoreDir, files)
|
||||
}
|
||||
+2
-5
@@ -1,9 +1,6 @@
|
||||
#!/bin/sh
|
||||
# script/fmt-check: check formatting (read-only). Fails instead of
|
||||
# writing. It checks every Go file outside .tool, which is more than
|
||||
# script/fmt formats: `go fmt ./...` skips `testdata` directories and
|
||||
# files and directories whose names start with `.` or `_`. Fix a file
|
||||
# only this reports with `gofmt -w`.
|
||||
# script/fmt-check: check formatting (read-only). Same scope as
|
||||
# script/fmt, but fails instead of writing.
|
||||
set -eu
|
||||
|
||||
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
|
||||
|
||||
+4
-3
@@ -21,9 +21,10 @@ goreleaser_version() {
|
||||
head -n 1
|
||||
}
|
||||
|
||||
# Resolve the goreleaser to run. A binary on PATH is accepted only when
|
||||
# it is exactly the pinned version, because a differently versioned tool
|
||||
# would produce a differently built release from the same tag. Anything else
|
||||
# Resolve the goreleaser to run, on the same rule script/lint uses for
|
||||
# golangci-lint: a binary on PATH is accepted only when it is exactly
|
||||
# the pinned version, because a differently versioned tool would
|
||||
# produce a differently built release from the same tag. Anything else
|
||||
# comes from .tool/bin, and a missing one is a loud failure naming the
|
||||
# script that installs it rather than a silent fallback.
|
||||
resolve_goreleaser() {
|
||||
|
||||
+1
-1
@@ -19,7 +19,7 @@ s3:
|
||||
secret_access_key: test-secret-key
|
||||
region: us-east-1
|
||||
use_ssl: true
|
||||
part_size: 5242880 # 5MiB
|
||||
part_size: 5242880 # 5MB
|
||||
index_path: /tmp/vaultik-test.sqlite
|
||||
chunk_size: 10MB
|
||||
blob_size_limit: 10GB
|
||||
|
||||
Reference in New Issue
Block a user