Compare commits

Author SHA1 Message Date
sneak b3ac072b8d Configure prettier and make fmt-check cover markdown (closes #69)
check / check (push) Successful in 37s
script/fmt ran prettier with default settings over root-level *.md and
*.json, swallowing every failure with `|| true`, while script/fmt-check
checked gofmt only. The formatter and the gate therefore disagreed
silently: `make fmt` rewrote markdown that `make check` never looked at,
including REPO_POLICIES.md, which is a verbatim copy of an authoritative
upstream document that local tooling must not touch.

Configuration:

- .prettierrc pins the two policy deviations from prettier defaults,
  four-space indents and proseWrap: always. Nothing else.
- .prettierignore excludes REPO_POLICIES.md so no local run can drift it
  from upstream again, plus .golangci.yml (user-owned, and listed even
  though the current file set does not reach it) and node_modules,
  vendor, bin.

One canonical file set:

- New script/prettier takes --write or --check and applies the same
  patterns in both modes, so script/fmt and script/fmt-check cannot
  drift apart by construction. The patterns are repo-wide (**/*.md,
  **/*.json) rather than root-only, so markdown in subdirectories such
  as a future docs/ is covered.
- No `|| true` anywhere, and no --no-error-on-unmatched-pattern: both
  patterns always match tracked files, so an empty match means the glob
  broke and prettier should say so instead of passing vacuously. A
  missing prettier is a hard error naming script/bootstrap, not a
  silent skip.

Pinned prettier:

- package.json/yarn.lock pin prettier 3.9.6; the lockfile carries the
  integrity hash, and --frozen-lockfile enforces it. script/prettier
  prefers node_modules/.bin/prettier and warns on stderr when it has to
  fall back to a PATH prettier of unknown version.
- script/bootstrap now installs node, yarn, and the locked JS deps. Its
  NODE_VERSION and YARN_VERSION pins already existed.

Docker gate:

- The golangci-lint image has no node, so the lint stage runs the new
  script/fmt-check-go (the Go half of fmt-check, extracted) instead of
  the whole thing.
- The markdown half gets its own stage on a digest-pinned node image
  shipping exactly the node and yarn versions bootstrap pins. The
  builder stage takes a COPY --from dependency on it, so BuildKit cannot
  skip it and a markdown violation fails `docker build .` rather than
  being skipped somewhere nobody looks.

Markdown files other than REPO_POLICIES.md are reformatted here for the
first time under the policy settings.
2026-08-09 05:04:38 +00:00
72 changed files with 2279 additions and 5323 deletions
+2 -62
View File
@@ -1,64 +1,4 @@
# .dockerignore does NOT use .gitignore semantics. Docker matches with
# moby/patternmatcher: filepath.Match plus `**`, so `*` does not cross
# `/` and an unprefixed pattern is anchored at the context root. Every
# depth-independent pattern therefore needs `**/`, or `config/.env` and
# `certs/server.key` still ship while this file reads as solved. Only
# genuinely root-anchored entries go unprefixed. Never transplant these
# into .gitignore, where `**/` is wrong.
#
# Matching is case-sensitive, so secrets use character ranges rather
# than an ALL-CAPS twin, which would still miss `Server.Key`.
#
# Extend with this repo's own host-built artifacts, written anchored:
# `/myapp`, never `**/myapp`, which also matches `cmd/myapp/` and
# deletes the package directory from the context.
# .git is sent without its config. Without a VERSION build argument the
# stage that compiles runs `git describe --tags --always` on .git, which
# does not need .git/config; that file can hold a credential, such as a
# password in a remote URL or the token the CI checkout step stores there.
.git/config
# Agent scratch: one full checkout of the repo per in-flight agent.
# Anchored because it occurs once where agents run at the repo root.
# KNOWN GAP: a repo running agents in subdirectories still ships
# `services/api/.claude/` and must add its own anchored entry.
.claude
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Re-include a committed template with a negation if the
# build needs one: `!docs/example.env`.
**/*.[eE][nN][vV]
**/.[eE][nN][vV].*
**/.[eE][nN][vV][rR][cC]
# Private keys and the bundles carrying them. Public certificates
# (*.crt, *.cer) are deliberately absent: they are legitimate inputs.
**/*.[pP][eE][mM]
**/*.[kK][eE][yY]
**/*.[pP]12
**/*.[pP][fF][xX]
**/[iI][dD]_[rR][sS][aA]
**/[iI][dD]_[dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]
**/[iI][dD]_[eE][dD]25519
# Dependencies: restored inside the image, never copied in.
**/node_modules
# OS metadata.
**/.DS_Store
**/Thumbs.db
# Editor state: never a build input, and it churns COPY.
**/*.swp
**/*.swo
**/*~
**/*.bak
**/.idea
**/.vscode
**/*.sublime-*
# This repo's own host-built archives (Makefile).
*.tmp
*.dockerimage
.git
node_modules
-12
View File
@@ -1,12 +0,0 @@
root = true
[*]
indent_style = space
indent_size = 4
end_of_line = lf
charset = utf-8
trim_trailing_whitespace = true
insert_final_newline = true
[Makefile]
indent_style = tab
+2 -21
View File
@@ -10,24 +10,5 @@ modcache.tzst
# Generated manifest files
.index.mf
# Secrets
.env
.env.*
*.key
*.pem
# OS files
.DS_Store
Thumbs.db
# Editor files
*.swp
*.swo
*~
.idea/
.vscode/
# Go build artifacts
*.log
*.out
*.test
# Stale files
.drone.yml
-35
View File
@@ -1,35 +0,0 @@
version: "2"
# Config schema uses the golangci-lint v2 layout (settings live under
# linters.settings, not top-level linters-settings) so that the
# thresholds below are actually applied by golangci-lint >= v2.
run:
timeout: 5m
modules-download-mode: readonly
linters:
default: all
disable:
# Genuinely incompatible with project patterns
- exhaustruct # Requires all struct fields
- depguard # Dependency allow/block lists
- godot # Requires comments to end with periods
- wsl # Deprecated, replaced by wsl_v5
- gomodguard # Deprecated, replaced by gomodguard_v2
- wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go
settings:
lll:
line-length: 88
funlen:
lines: 80
statements: 50
cyclop:
max-complexity: 15
dupl:
threshold: 100
issues:
max-issues-per-linter: 0
max-same-issues: 0
+2 -5
View File
@@ -27,8 +27,5 @@ source for coding standards, formatting, linting, and workflow rules.
- The proto definition is in `mfer/mf.proto`; generated `.pb.go` files are
committed (required for `go get` compatibility).
- The format specification is in `FORMAT.md`.
- Open work, open design questions included, is tracked only in the repo's
issues: https://git.eeqj.de/sneak/mfer/issues. There is no `TODO.md` and no
TODO list in `README.md`; do not add either. For this repo this overrides the
`REPO_POLICIES.md` rule to put the todo list in the README, per sneak's
ruling: https://git.eeqj.de/sneak/mfer/issues/76#issuecomment-118130.
- See the TODO section in `README.md` for the 1.0 implementation plan and open
design questions.
+4 -28
View File
@@ -1,6 +1,6 @@
# Lint stage — fast feedback on formatting and lint issues
# golangci/golangci-lint:v2.12.2 (Debian-based), 2026-08-07
FROM golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 AS lint
# golangci/golangci-lint:v2.0.2 (2026-03-14)
FROM golangci/golangci-lint@sha256:d55581f7797e7a0877a7c3aaa399b01bdc57d2874d6412601a046cc4062cb62e AS lint
WORKDIR /src
COPY go.mod go.sum ./
@@ -14,9 +14,7 @@ RUN touch mfer/mf.pb.go
# Go half of fmt-check only: this image has no node, so no prettier. The
# markdown half runs in the mdfmt stage below.
RUN make fmt-check-go
# The linter directly, not `make lint`: script/lint builds this stage, and
# there is no docker inside this build.
RUN golangci-lint run --config .golangci.yml ./...
RUN make lint
# Markdown/JSON format stage — prettier needs node, which the Go images
# do not have. node:22.17.0-bookworm-slim (2026-08-09); ships node
@@ -50,29 +48,7 @@ COPY . .
RUN touch mfer/mf.pb.go
RUN make test
# A build context sent as a tar archive, as upaas sends it, keeps its files'
# owners, and git refuses to read a checkout owned by another user.
RUN git config --system --add safe.directory /src
# The revision `mfer version` prints, stamped into main.Gitrev: the VERSION
# build argument when one is given (script/docker passes one), otherwise
# `git describe --tags --always` of the .git the build context carries: the
# tag on a tagged commit, tag-N-gHASH on a commit after one, the short commit
# when no tag is reachable. git ships in this base image. A context that
# carries .git and still yields no version fails the build.
ARG VERSION
RUN version="${VERSION:-$(git describe --tags --always)}"; \
if [ -e .git ] && { [ -z "$version" ] || [ "$version" = dev ] || \
[ "$version" = unknown ]; }; then \
echo "no version could be derived although the build context carries .git" >&2; \
exit 1; \
fi; \
cd cmd/mfer && \
CGO_ENABLED=0 go build -tags urfave_cli_no_docs -ldflags "-X main.Gitrev=$version" -o /mfer .
# Fail unless /mfer is statically linked: scratch has no C library to run it.
RUN ldd /mfer 2>&1 | grep -q 'not a dynamic executable'
RUN cd cmd/mfer && go build -tags urfave_cli_no_docs -o /mfer .
FROM scratch
COPY --from=builder /mfer /mfer
+1 -3
View File
@@ -47,9 +47,7 @@ allows verifying data integrity before decompression.
The `innerMessage` field is compressed with
[Zstandard (zstd)](https://facebook.github.io/zstd/). Implementations must
enforce a decompression size limit to prevent decompression bombs. The reference
implementation limits decompressed size to 256 MB. It writes zstd frames with a
window of at most 8 MiB, the largest window the zstd format recommends decoders
support, and refuses frames that ask for a larger one.
implementation limits decompressed size to 256 MB.
## Inner Message (`MFFile`)
+4 -4
View File
@@ -13,7 +13,7 @@ GOLDFLAGS += -X main.Version=$(VERSION)
GOLDFLAGS += -X main.Gitrev=$(GITREV_BUILD)
GOFLAGS := -ldflags "$(GOLDFLAGS)"
.PHONY: bootstrap setup docker default run ci test fuzz check lint fmt fmt-check fmt-check-go fmt-check-md hooks fixme
.PHONY: bootstrap setup docker default run ci test check lint fmt fmt-check fmt-check-go fmt-check-md hooks fixme
default: fmt test
@@ -32,9 +32,6 @@ ci: test
test:
@script/test
fuzz:
@script/fuzz
$(PROTOC_GEN_GO):
test -e $(PROTOC_GEN_GO) || go install -v google.golang.org/protobuf/cmd/protoc-gen-go@v1.28.1
@@ -58,6 +55,9 @@ fmt-check-md:
hooks:
@script/install-precommit
devprereqs:
which golangci-lint || go install -v github.com/golangci/golangci-lint/cmd/golangci-lint@v2.0.2
mfer/mf.pb.go: mfer/mf.proto
cd mfer && go generate .
+232 -99
View File
@@ -1,12 +1,12 @@
# mfer
[mfer](https://git.eeqj.de/sneak/mfer) is a [WTFPL](https://wtfpl.net)-licensed
(public domain) [Go](https://golang.org) library and command-line tool by
[@sneak](https://sneak.berlin) that specifies and generates `.mf` manifest files
over a directory tree to encapsulate metadata about the files — such as
cryptographic checksums and signatures over same — to aid in archiving,
downloading, streaming, and mirroring. It was first published in 2022. The
manifest files' data is serialized with Google's
[mfer](https://git.eeqj.de/sneak/mfer) is a reference implementation library and
thin wrapper command-line utility written in [Go](https://golang.org) and first
published in 2022 under the [WTFPL](https://wtfpl.net) (public domain) license.
It specifies and generates `.mf` manifest files over a directory tree of files
to encapsulate metadata about them (such as cryptographic checksums or
signatures over same) to aid in archiving, downloading, and streaming, or
mirroring. The manifest files' data is serialized with Google's
[protobuf serialization format](https://developers.google.com/protocol-buffers).
The structure of these files can be found
[in the format specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto)
@@ -21,41 +21,10 @@ This project was started by [@sneak](https://sneak.berlin) to scratch an itch in
as a de-facto standard and be incorporated into other software. A compatible
javascript library is planned.
# Getting Started
`mfer` builds from source with a Go 1.23+ toolchain. The generated protobuf code
is committed, so no `protoc` toolchain is required:
```sh
git clone https://git.eeqj.de/sneak/mfer.git
cd mfer
go build -o bin/mfer ./cmd/mfer
```
Generate a manifest for a directory tree, verify it later, and fetch a published
tree by URL:
```sh
# Write .index.mf, a manifest of the files under the current directory.
bin/mfer gen .
# Verify the files on disk against the manifest. Exits nonzero if any file
# is missing or corrupted.
bin/mfer check .index.mf
# Download and cryptographically verify a tree published over HTTP: mfer
# fetches <url>/index.mf, then downloads every file it lists.
bin/mfer fetch https://example.com/tree/
```
Run `bin/mfer help` for the full command list, or `bin/mfer <command> --help`
for a single command's options.
# Build Status
CI runs `script/cibuild`, which builds the Docker image with `--no-cache`, so
the formatting, lint and test steps in the `Dockerfile` run on every build. The
`main` branch must always be green.
CI runs via `script/cibuild` (`docker build .`), which executes `make check`
(formatting, linting, tests). The `main` branch must always be green.
# Entrypoints
@@ -65,23 +34,18 @@ standard: normalized scripts in `script/` are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call them. We
provide:
- `script/bootstrap` — install all dependencies (Go, Go module download, and
node/yarn plus the prettier version pinned in `package.json`/`yarn.lock`),
idempotently; golangci-lint is not installed, it runs only in Docker
- `script/bootstrap` — install all dependencies (Go, golangci-lint, Go module
download, and node/yarn plus the prettier version pinned in
`package.json`/`yarn.lock`), idempotently
- `script/setup` — make a fresh clone ready for development: runs
`script/bootstrap`, then `script/install-precommit`
- `script/projectname` — output the project name (`mfer`); used by other scripts
such as `script/docker`
- `script/test` — run the test suite (`go test`), regenerating the protobuf code
first if it is stale
- `script/fuzz` — fuzz the manifest parser for one minute; run by hand
(`make fuzz`), never by CI, while `script/test` runs its committed seed corpus
as ordinary tests
- `script/lint` — run `golangci-lint` in Docker: builds only the `lint` stage of
the `Dockerfile` (the Go format check, then the linter), uncached so it runs
every time, then removes the image
- `script/fmt` — format all code and docs (writes): `gofumpt` and
`script/prettier --write`
- `script/lint` — run `golangci-lint` and verify `gofmt` cleanliness
- `script/fmt` — format all code and docs (writes): `gofumpt`,
`golangci-lint run --fix`, and `script/prettier --write`
- `script/prettier` — run prettier over the repository's canonical file set
(Markdown and JSON, minus `.prettierignore`) in the given mode, `--write` or
`--check`; the single definition of that file set, so `script/fmt` and
@@ -92,8 +56,8 @@ provide:
Docker lint stage, whose image has no node
- `script/check` — run `script/test`, `script/lint`, and `script/fmt-check`
- `script/docker` — build the Docker image tagged with the project name
- `script/cibuild` — CI entrypoint: builds the image with the same command as
`script/docker`, uncached, so the checks in the Dockerfile run every time
- `script/cibuild` — CI entrypoint: `docker build .` (the Dockerfile runs the
checks)
- `script/precommit` — pre-commit checks: `go mod tidy` verification, then
`script/check`
- `script/install-precommit` — install the git pre-commit hook that runs
@@ -107,18 +71,17 @@ yet. Primary development happens on a privately-run Gitea instance at
[tracked there](https://git.eeqj.de/sneak/mfer/issues).
Changes must always be formatted with a standard `go fmt`, syntactically valid,
and must pass the linting defined in the repository's `.golangci.yml`, which
`make lint` runs in Docker. The `main` branch is protected and all changes must
be made via [pull requests](https://git.eeqj.de/sneak/mfer/pulls) and pass CI to
be merged. Any changes submitted to this project must also be
and must pass the linting defined in the repository (presently only the
`golangci-lint` defaults), which can be run with a `make lint`. The `main`
branch is protected and all changes must be made via
[pull requests](https://git.eeqj.de/sneak/mfer/pulls) and pass CI to be merged.
Any changes submitted to this project must also be
[WTFPL-licensed](https://wtfpl.net) to be considered.
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
tooling requirements, and workflow conventions.
# Rationale
## The problem
# Problem Statement
Given a plain URL, there is no standard way to safely and programmatically
download everything "under" that URL path. `wget -r` can traverse directory
@@ -142,7 +105,7 @@ Real issues I face:
- when I download a large file via HTTP, I have no way of knowing if the file
content is what it's supposed to be
## The solution
# Proposed Solution
A standard, a manifest file format, and a tool for generating same.
@@ -174,28 +137,6 @@ The manifest file would do several important things:
- maybe a bittorrent chunklist for torrent client compatibility? perhaps a
top-level infohash for the whole manifest?
# Design
The repository is split into a reusable library and a thin command-line wrapper
around it.
- `mfer/` is the reusable library and the heart of the project: it defines the
manifest format and implements building, scanning, checking, serialization,
and signing. The protobuf schema is `mfer/mf.proto`, and the generated code it
produces (`mfer/mf.pb.go`) is committed alongside it so the library builds
with `go get` and needs no `protoc` toolchain.
- `internal/cli/` holds the command implementations — `generate` (alias `gen`),
`check`, `freshen`, `export`, `list` (alias `ls`), `fetch`, and `version` —
that wire the library to the command-line interface.
- `internal/log/` provides the logging used by the commands and the library.
- `internal/bork/` provides the error the library returns when a manifest's
decompressed contents are not the size the manifest records.
- `cmd/mfer/` is the entrypoint: its `main` package runs `internal/cli` and
exits with the status it returns.
Everything under `internal/` is private to this repository; only the `mfer/`
package is intended for import by other software.
# Design Goals
- Replace SHASUMS/SHASUMS.asc files
@@ -215,23 +156,18 @@ package is intended for import by other software.
- metadata size should not be used as an excuse to sacrifice utility (such
as providing checksums over each chunk of a large file)
# Original Design Questions
These were the open questions when the project started; open design questions
are now tracked only in the [issues](https://git.eeqj.de/sneak/mfer/issues).
# Open Questions
- Should the manifest file include checksums of individual file chunks, or just
for the whole assembled file? If so, should the chunk size be fixed or
dynamic? Still open, on
[issue 81](https://git.eeqj.de/sneak/mfer/issues/81#issuecomment-118698).
for the whole assembled file?
- If so, should the chunksize be fixed or dynamic?
- Should the manifest signature format be GnuPG signatures, or those from
OpenBSD's signify (of which there is a good
[golang implementation](https://github.com/frankbraun/gosignify))? Still open,
as question 10 on [issue 82](https://git.eeqj.de/sneak/mfer/issues/82).
[golang implementation](https://github.com/frankbraun/gosignify)?
- Should the on-disk serialization format be proto3 or json? Settled: it is
proto3, see `FORMAT.md` and `mfer/mf.proto`.
- Should the on-disk serialization format be proto3 or json?
# Tool Examples
@@ -309,10 +245,207 @@ regardless of filesystem format.
Please email [`sneak@sneak.berlin`](mailto:sneak@sneak.berlin) with your desired
username for an account on this Gitea instance.
# TODO
# TODO: Remaining Work for 1.0
Open work, open design questions included, is tracked in this repo's issues:
[https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues).
## Design Questions (Owner Decision Required)
These require @sneak's input before implementation. Answers should be added
inline below each question.
### Format Design
**1. Should `MFFileChecksum` be simplified?** Currently it's a separate message
wrapping a single `bytes multiHash` field. Since multihash already
self-describes the algorithm, `repeated bytes hashes` directly on `MFFilePath`
would be simpler and reduce per-file protobuf overhead. Is the extra message
layer intentional (e.g. planning to add per-hash metadata like `verified_at`)?
> _answer:_
**2. Should file permissions/mode be stored?** The format stores mtime/ctime but
not Unix file permissions. For archival use this may not matter, but for
software distribution or filesystem restoration it's a gap. Should we reserve a
field now (e.g. `optional uint32 mode = 305`) even if we don't populate it yet?
> _answer:_
**3. Should `atime` be removed from the schema?** Access time is volatile,
non-deterministic, and often disabled (`noatime`). Including it means two
manifests of the same directory at different times will differ, which conflicts
with the determinism goal. Remove it, or document it as "never set by default"?
> _answer:_
**4. What are the path normalization rules?** The proto has `string path` with
no specification about: always forward-slash? Must be relative? No `..`
components allowed? UTF-8 NFC vs NFD normalization (macOS vs Linux)? Max path
length? This is a security issue (path traversal) and a cross-platform
compatibility issue. What rules should the spec mandate?
> _answer:_
**5. Should we add a version byte after the magic?** Currently `ZNAVSRFG` is
followed immediately by protobuf. Adding a version byte (`ZNAVSRFG\x01`) would
allow future framing changes without requiring protobuf parsing to detect the
version. `MFFileOuter.Version` serves this purpose but requires successful
deserialization to read. Worth the extra byte?
> _answer:_
**6. Should we add a length-prefix after the magic?** Protobuf is not
self-delimiting. If we ever want to concatenate manifests or append data after
the protobuf, the current framing is insufficient. Add a varint or fixed-width
length-prefix?
> _answer:_
### Signature Design
**7. What does the outer SHA-256 hash cover — compressed or uncompressed data?**
The code currently hashes compressed data (good for verifying before
decompression), but this should be explicitly documented. Which is the intended
behavior?
> _answer:_
**8. Should `signatureString()` sign raw bytes instead of a hex-encoded
string?** Currently the canonical string is `MAGIC-UUID-MULTIHASH` with hex
encoding, which adds a transformation layer. Signing the raw `sha256` bytes (or
compressed `innerMessage` directly) would be simpler. Keep the string format or
switch to raw bytes?
> _answer:_
**9. Should we support detached signature files (`.mf.sig`)?** Embedded
signatures are better for single-file distribution. Detached `.mf.sig` files
follow the familiar `SHASUMS`/`SHASUMS.asc` pattern and are simpler for HTTP
serving. Support both modes?
> _answer:_
**10. GPG vs pure-Go crypto for signatures?** Shelling out to `gpg` is fragile
(may not be installed, version-dependent output).
`github.com/ProtonMail/go-crypto` provides pure-Go OpenPGP, or we could use
Ed25519/signify (simpler, no key management). Which direction?
> _answer:_
### Implementation Design
**11. Should manifests be deterministic by default?** This means: sort file
entries by path, omit `createdAt` timestamp (or make it opt-in), no `atime`.
Should determinism be the default, with a `--include-timestamps` flag to opt in?
> _answer:_
**12. Should we consolidate or keep both scanner/checker implementations?**
There are two parallel implementations: `mfer/scanner.go` + `mfer/checker.go`
(typed with `FileSize`, `RelFilePath`) and `internal/scanner/` +
`internal/checker/` (raw `int64`, `string`). The `mfer/` versions are superior.
Delete the `internal/` versions?
> _answer:_
**13. Should the `manifest` type be exported?** Currently unexported with
exported constructors (`NewManifestFromReader`, `NewManifestFromFile`).
Consumers can't declare `var m *mfer.manifest`. Export the type, or define an
interface?
> _answer:_
**14. What should the Go module path be for 1.0?** Currently
`sneak.berlin/go/mfer` in `go.mod` but `git.eeqj.de/sneak/mfer/mfer` in the
proto `go_package` option. Which is canonical?
> _answer:_
## Implementation Tasks
### Repo Infrastructure
- [ ] Add `.golangci.yml` (fetch from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`)
- [ ] Add `.editorconfig`
- [ ] Add `.gitea/workflows/check.yml` that runs `docker build .`
### Format & Correctness
- [ ] Resolve proto `go_package` path inconsistency
(`git.eeqj.de/sneak/mfer/mfer` vs `sneak.berlin/go/mfer`)
- [ ] Specify path invariants — add proto comments requiring UTF-8,
forward-slash, relative paths, no `..`, no leading `/`; validate in
`Builder.AddFile` and `Builder.AddFileWithHash` (pending design question
answer)
- [ ] Remove or deprecate `atime` from proto (pending design question answer)
- [ ] Reserve `optional uint32 mode = 305` in `MFFilePath` for future file
permissions (pending design question answer)
- [ ] Add version byte after magic — `ZNAVSRFG\x01` for format version 1
(pending design question answer)
- [ ] Write format specification document — separate from README: magic, outer
structure, compression, inner structure, path invariants, signature
scheme, canonical serialization
### Library
- [ ] Delete `internal/scanner/` and `internal/checker/` — consolidate on
`mfer/` package versions; update CLI code (pending design question answer)
- [ ] Add deterministic file ordering — sort entries by path (lexicographic,
byte-order) in `Builder.Build()`; add test asserting byte-identical output
from two runs
- [ ] Add decompression size limit — `io.LimitReader` in `deserializeInner()`
with `m.pbOuter.Size` as bound
- [ ] Fix `errors.Is` dead code in checker — replace with `os.IsNotExist(err)`
or `errors.Is(err, fs.ErrNotExist)`
- [ ] Fix `AddFile` to verify size — check `totalRead == size` after reading,
return error on mismatch
- [ ] Export the `manifest` type or define a public interface (pending design
question answer) — currently consumers cannot hold a reference to a loaded
manifest in their own type declarations
- [ ] Replace GPG subprocess calls with pure-Go crypto (pending design question
answer) — current implementation shells out to `gpg` which may not be
installed
- [ ] Add timeout to any remaining subprocess calls
### CLI
- [ ] Fix flag naming — all CLI flags should use kebab-case as primary
(`--include-dotfiles`, `--follow-symlinks`)
- [ ] Fix URL construction in fetch — use `BaseURL.JoinPath()` or
`url.JoinPath()` instead of string concatenation
- [ ] Add progress rate-limiting to Checker — throttle to once per second,
matching Scanner
- [ ] Add `--deterministic` flag or make it default — omit `createdAt`, sort
files (pending design question answer)
- [ ] Wire `--version` flag properly (currently only a `version` subcommand
exists; top-level `--version` shows urfave/cli generic output)
- [ ] Add retry logic to `fetch` — currently no retries on transient HTTP
errors; needs exponential backoff
- [ ] `fetch` command uses bare `http.Get` with no timeout — needs `http.Client`
with configurable timeout
### Testing & Robustness
- [ ] Add fuzzing tests for `NewManifestFromReader` — protobuf deserialization
of untrusted input needs fuzz coverage
- [ ] Add integration test for `freshen` CLI command — current tests only verify
setup, not the actual freshen operation end-to-end
- [ ] Add test for `fetch` CLI command end-to-end (currently only `downloadFile`
is tested)
### Documentation
- [ ] Promote `FORMAT.md` as primary spec reference; README should link to it
more prominently
- [ ] Audit and update all error messages for consistency and helpfulness
- [ ] Document the signature scheme more thoroughly (canonical string format,
verification steps)
### Release
- [ ] Finalize Go module path
- [ ] Update version constant in `mfer/constants.go`
- [ ] Add `--version` output matching SemVer
- [ ] Tag `v1.0.0`
# See Also
@@ -329,9 +462,9 @@ Open work, open design questions included, is tracked in this repo's issues:
- Issues:
[https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues)
# Author
# Authors
- [@sneak](https://sneak.berlin)
- [@sneak &lt;sneak@sneak.berlin&gt;](mailto:sneak@sneak.berlin)
# License
+107
View File
@@ -0,0 +1,107 @@
# Workflow
- branch (from `main`)
- do the work in Next Step
- move Next Step to the top of Completed Steps
- move the top item of Future Steps into Next Step
- commit (`TODO.md` changes in the same commit as the work)
- merge to `main` if the branch is not protected, otherwise open a PR
- push
# Status
pre-1.0. No git tags. README section "TODO: Remaining Work for 1.0" lists open
design questions and implementation tasks; policy compliance work is in flight
and unmerged.
# Next Step
Land the in-flight compliance branch chore/align-repo-policies: finish and
commit the uncommitted work (32 modified Go files, new untracked .golangci.yml
and TODO.md), confirm `make check` is green, merge the branch (one commit ahead
of main as of 2026-07-03) to main, and push.
# Completed Steps
- 2026-08-09: added `.prettierrc`/`.prettierignore`, gave `script/fmt` and
`script/fmt-check` one shared prettier file set via `script/prettier`, dropped
the `|| true` that hid prettier failures, and added a node-based Dockerfile
stage so a markdown formatting violation fails `docker build .` (#69)
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints, Makefile
shims, README Entrypoints section
- 2026-07-03: aligned repo tooling, docs, and config with standardized policies
(7d9a138, on chore/align-repo-policies, unmerged)
- 2026-06-28: moved to standardized repo policies (#56, on main)
- 2026-04-07: added 1.0 roadmap as README TODO section, removed old TODO.md
(#54)
- 2026-03-20: added Gitea Actions CI workflow (#53)
- 2026-03-17: added REPO_POLICIES.md, renamed CLAUDE.md to AGENTS.md (#51);
removed committed .index.mf (#52)
- 2026-03-15: split Dockerfile with pre-built golangci-lint stage for faster CI
(#45)
- 2026-03-01: 1.0 quality polish: code review, tests, bug fixes, docs (#32)
- 2026-02-20: deterministic file ordering in Builder.Build() (#28); removed
committed vendor/modcache archives (#35)
- 2026-02-08: added --seed flag for deterministic manifest UUID
# Future Steps
- Compliance (fold of TODO.md audit 2026-07-02; verify which items the in-flight
branch already closes, then check off):
- Add .editorconfig (canonical copy from sneak/prompts)
- Add standardized .golangci.yml (present untracked on the branch;
user-owned, copy verbatim)
- Make .gitignore cover secrets (.env, _.key, _.pem), OS files (.DS_Store),
and editor files (_.swp, _~)
- Make fmt-check/lint verify with gofumpt, not gofmt -l, so `make check`
matches what `make fmt` writes
- Add README "Getting Started" section with copy-pasteable install/usage
block
- Move FORMAT.md from repo root to docs/ and update the AGENTS.md reference
- Pin Makefile-installed Go tools (protoc-gen-go@v1.28.1,
golangci-lint@v2.0.2) by module hash, not mutable tag
- Set `make test` timeout to 30s (currently 10s)
- Add explicit README "Rationale" heading (content exists under other
names); name the author in the README Description first line
- Reconcile root-level AGENTS.md with directory-hygiene policy (keep or
relocate)
- Add a `make build` target
- Rewrite `make hooks` to use printf or a heredoc instead of non-portable
`echo '...\n...'`
- Answer the 14 owner design questions in the README 1.0 roadmap:
- Format: simplify MFFileChecksum; store file mode; drop atime; specify path
normalization rules; version byte after magic; length-prefix after magic
- Signatures: hash covers compressed or uncompressed data; sign raw bytes vs
hex canonical string; detached .mf.sig support; GPG subprocess vs pure-Go
crypto
- Implementation: deterministic manifests by default; consolidate duplicate
scanner/checker implementations; export the manifest type; canonical Go
module path for 1.0
- Format and correctness:
- Resolve proto go_package vs go.mod module path inconsistency
- Specify and validate path invariants (UTF-8, forward-slash, relative, no
.., no leading /)
- Remove or deprecate atime; reserve mode field; add version byte (all
pending design answers)
- Write a standalone format specification document
- Library:
- Delete internal/scanner and internal/checker; consolidate on the mfer/
package versions (pending design answer)
- Add decompression size limit via io.LimitReader in deserializeInner()
- Fix errors.Is dead code in checker; make AddFile verify totalRead == size
- Export manifest type or define a public interface (pending)
- Replace GPG subprocess with pure-Go crypto (pending); add timeouts to
remaining subprocess calls
- CLI:
- Kebab-case primary flag names; fix fetch URL construction with
url.JoinPath; add http.Client timeout and retry with backoff to fetch;
rate-limit Checker progress output; add --deterministic flag or default;
wire top-level --version properly
- Testing:
- Fuzz NewManifestFromReader; end-to-end tests for freshen and fetch
- Documentation:
- Promote docs/FORMAT.md as primary spec reference; audit error messages;
document the signature scheme fully
- Release:
- Finalize module path, bump version constant, SemVer --version output, tag
v1.0.0
+6 -1
View File
@@ -1,7 +1,12 @@
#!/bin/bash
#
if [[ ! -z "$DRONE_COMMIT_SHA" ]]; then
echo "${DRONE_COMMIT_SHA:0:7}"
exit 0
fi
if [[ ! -z "$GITREV" ]]; then
echo $GITREV
else
git describe --tags --always --dirty=-dirty
git describe --always --dirty=-dirty
fi
+1 -7
View File
@@ -1,4 +1,3 @@
// Command mfer generates and verifies file manifests.
package main
import (
@@ -7,13 +6,8 @@ import (
"sneak.berlin/go/mfer/internal/cli"
)
// Appname is the name of this program.
const Appname = "mfer"
// Version and Gitrev are injected at build time via -ldflags.
//
//nolint:gochecknoglobals // set via ldflags at build time
var (
Appname string = "mfer"
Version string
Gitrev string
)
+2 -16
View File
@@ -6,20 +6,6 @@ import (
"github.com/stretchr/testify/assert"
)
// TestAppname pins the program name that main passes to cli.Run; it is
// the name that appears in usage output and in the log prefix. It also
// keeps this package compiled under `go test`.
func TestAppname(t *testing.T) {
t.Parallel()
assert.Equal(t, "mfer", Appname)
}
// TestVersionDefaults documents that Version and Gitrev are empty unless
// injected at build time via -ldflags.
func TestVersionDefaults(t *testing.T) {
t.Parallel()
assert.Empty(t, Version)
assert.Empty(t, Gitrev)
func TestBuild(t *testing.T) {
assert.True(t, true)
}
+6 -5
View File
@@ -1,14 +1,15 @@
// Package bork defines the sentinel errors used by the manifest
// reader and writer.
package bork
import (
"errors"
"fmt"
)
var (
// ErrMissingMagic indicates the input lacks the manifest magic bytes.
ErrMissingMagic = errors.New("missing magic bytes in file")
// ErrFileTruncated indicates the input ended before the expected length.
ErrMissingMagic = errors.New("missing magic bytes in file")
ErrFileTruncated = errors.New("file/stream is truncated abnormally")
)
func Newf(format string, args ...interface{}) error {
return fmt.Errorf(format, args...)
}
+2 -5
View File
@@ -1,14 +1,11 @@
package bork_test
package bork
import (
"testing"
"github.com/stretchr/testify/assert"
"sneak.berlin/go/mfer/internal/bork"
)
func TestBuild(t *testing.T) {
t.Parallel()
assert.Error(t, bork.ErrMissingMagic)
assert.NotNil(t, ErrMissingMagic)
}
+110 -259
View File
@@ -1,15 +1,10 @@
// Package cli implements the mfer command-line interface.
package cli
import (
"context"
"encoding/hex"
"errors"
"fmt"
"io"
"math"
"path/filepath"
"strconv"
"strings"
"time"
@@ -20,256 +15,21 @@ import (
"sneak.berlin/go/mfer/mfer"
)
// fingerprintHexLen is the length of a full GPG key fingerprint in hex
// characters.
const fingerprintHexLen = 40
var (
// errNoManifestFound indicates no manifest file was found in the
// searched directory.
errNoManifestFound = errors.New("no manifest found")
// errInvalidFingerprint indicates a malformed --require-signature
// fingerprint argument. The length is spliced in from
// fingerprintHexLen so the two cannot drift apart.
errInvalidFingerprint = errors.New(
"invalid fingerprint: must be exactly " +
strconv.Itoa(fingerprintHexLen) + " hex characters")
// errManifestNotSigned indicates a signature was required but the
// manifest is unsigned. It is wrapped mid-sentence so that the
// rendered message stays exactly as mfer has always printed it.
errManifestNotSigned = errors.New("manifest is not signed")
// errSignerMismatch indicates the embedded signing key fingerprint
// does not match the required signer. Its text is the mid-sentence
// fragment of the rendered message, which users grep for in CI and
// which must therefore not change; match it with errors.Is rather
// than by reading it.
errSignerMismatch = errors.New("does not match required")
)
// safeUint64 converts a non-negative int64 to uint64, clamping negative
// values to zero.
func safeUint64(n int64) uint64 {
if n < 0 {
return 0
}
return uint64(n)
}
// safeRateUint64 converts a bytes-per-second rate to uint64 for display.
//
// A rate is computed as bytes/elapsed, so it is +Inf when the elapsed
// time rounds to zero and NaN when zero bytes were processed in zero
// time. Neither has a defined conversion to uint64, and on amd64 +Inf
// converts to a number that renders as "8.0 EiB/s"; both display as zero
// instead.
func safeRateUint64(rate float64) uint64 {
if math.IsNaN(rate) || math.IsInf(rate, 0) || rate <= 0 {
return 0
}
if rate >= math.MaxUint64 {
return math.MaxUint64
}
return uint64(rate)
}
// findManifest looks for a manifest file in the given directory.
// It checks for index.mf and .index.mf, returning the first one found.
func findManifest(fs afero.Fs, dir string) (string, error) {
candidates := []string{"index.mf", ".index.mf"}
for _, name := range candidates {
path := filepath.Join(dir, name)
exists, err := afero.Exists(fs, path)
if err != nil {
return "", err
}
if exists {
return path, nil
}
}
return "", fmt.Errorf(
"%w in %s (looked for index.mf and .index.mf)", errNoManifestFound, dir)
}
// fetchManifestToTemp downloads a manifest URL to a temporary file and
// returns the temp file path. The caller is responsible for removing it.
func (mfa *CLIApp) fetchManifestToTemp(url string) (string, error) {
rc, fetchErr := mfa.openManifestReader(url)
if fetchErr != nil {
return "", fetchErr
}
tmpFile, tmpErr := afero.TempFile(mfa.Fs, "", "mfer-manifest-*.mf")
if tmpErr != nil {
_ = rc.Close()
return "", fmt.Errorf("failed to create temp file: %w", tmpErr)
}
tmpPath := tmpFile.Name()
_, cpErr := io.Copy(tmpFile, rc)
_ = rc.Close()
_ = tmpFile.Close()
if cpErr != nil {
_ = mfa.Fs.Remove(tmpPath)
return "", fmt.Errorf("failed to download manifest: %w", cpErr)
}
return tmpPath, nil
}
// verifyRequiredSigner enforces the --require-signature fingerprint
// against the manifest's embedded signing key.
func verifyRequiredSigner(
ctx context.Context, chk *mfer.Checker, requiredSigner string,
) error {
// Validate fingerprint format: must be exactly 40 hex characters
if len(requiredSigner) != fingerprintHexLen {
return fmt.Errorf("%w, got %d", errInvalidFingerprint, len(requiredSigner))
}
_, err := hex.DecodeString(requiredSigner)
if err != nil {
return fmt.Errorf("invalid fingerprint: must be valid hex: %w", err)
}
if !chk.IsSigned() {
return fmt.Errorf("%w, but signature from %s is required",
errManifestNotSigned, requiredSigner)
}
// Extract fingerprint from the embedded public key (not from the
// signer field). This validates the key is importable and gets its
// actual fingerprint.
embeddedFP, err := chk.ExtractEmbeddedSigningKeyFP(ctx)
if err != nil {
return fmt.Errorf(
"failed to extract fingerprint from embedded signing key: %w", err)
}
// Compare fingerprints - must be exact match (case-insensitive)
if !strings.EqualFold(embeddedFP, requiredSigner) {
return fmt.Errorf("embedded signing key fingerprint %s %w %s",
embeddedFP, errSignerMismatch, requiredSigner)
}
log.Infof("manifest signature verified (signer: %s)", embeddedFP)
return nil
}
// reportCheckProgress renders progress updates until the channel closes.
func reportCheckProgress(progress <-chan mfer.CheckStatus) {
for status := range progress {
if status.ETA > 0 {
log.Progressf("Checking: %d/%d files, %s/s, ETA %s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)),
status.ETA.Round(time.Second),
status.Failures)
} else {
log.Progressf("Checking: %d/%d files, %s/s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)),
status.Failures)
}
}
log.ProgressDone()
}
// countCheckFailures consumes check results, counting and logging
// failures, then closes done.
func countCheckFailures(
results <-chan mfer.Result, failures *int64, done chan<- struct{},
) {
for result := range results {
if result.Status != mfer.StatusOK {
*failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
} else {
log.Verbosef("%s: %s", result.Status, result.Path)
}
}
close(done)
}
// findExtraFiles reports files present on disk but absent from the
// manifest, counting each as a failure.
func findExtraFiles(ctx *cli.Context, chk *mfer.Checker, failures *int64) error {
extraResults := make(chan mfer.Result, 1)
extraDone := make(chan struct{})
go func() {
for result := range extraResults {
*failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
}
close(extraDone)
}()
err := chk.FindExtraFiles(ctx.Context, extraResults)
if err != nil {
return fmt.Errorf("failed to check for extra files: %w", err)
}
<-extraDone
return nil
}
// runCheck runs the manifest check with progress and result reporting
// and returns the number of failures.
func runCheck(ctx *cli.Context, chk *mfer.Checker, showProgress bool) (int64, error) {
// Set up results channel
results := make(chan mfer.Result, 1)
// Set up progress channel
var progress chan mfer.CheckStatus
if showProgress {
progress = make(chan mfer.CheckStatus, 1)
go reportCheckProgress(progress)
}
// Process results in a goroutine
var failures int64
done := make(chan struct{})
go countCheckFailures(results, &failures, done)
// Run check
err := chk.Check(ctx.Context, results, progress)
if err != nil {
return 0, fmt.Errorf("check failed: %w", err)
}
// Wait for results processing to complete
<-done
// Check for extra files if requested
if ctx.Bool("no-extra-files") {
err = findExtraFiles(ctx, chk, &failures)
if err != nil {
return 0, err
}
}
return failures, nil
return "", fmt.Errorf("no manifest found in %s (looked for index.mf and .index.mf)", dir)
}
func (mfa *CLIApp) checkManifestOperation(ctx *cli.Context) error {
@@ -282,13 +42,24 @@ func (mfa *CLIApp) checkManifestOperation(ctx *cli.Context) error {
// URL manifests need to be downloaded to a temp file for the checker
if isHTTPURL(manifestPath) {
tmpPath, tmpErr := mfa.fetchManifestToTemp(manifestPath)
rc, fetchErr := mfa.openManifestReader(manifestPath)
if fetchErr != nil {
return fmt.Errorf("check: %w", fetchErr)
}
tmpFile, tmpErr := afero.TempFile(mfa.Fs, "", "mfer-manifest-*.mf")
if tmpErr != nil {
return fmt.Errorf("check: %w", tmpErr)
_ = rc.Close()
return fmt.Errorf("check: failed to create temp file: %w", tmpErr)
}
tmpPath := tmpFile.Name()
_, cpErr := io.Copy(tmpFile, rc)
_ = rc.Close()
_ = tmpFile.Close()
if cpErr != nil {
_ = mfa.Fs.Remove(tmpPath)
return fmt.Errorf("check: failed to download manifest: %w", cpErr)
}
defer func() { _ = mfa.Fs.Remove(tmpPath) }()
manifestPath = tmpPath
}
@@ -306,31 +77,111 @@ func (mfa *CLIApp) checkManifestOperation(ctx *cli.Context) error {
// Check signature requirement
requiredSigner := ctx.String("require-signature")
if requiredSigner != "" {
err = verifyRequiredSigner(ctx.Context, chk, requiredSigner)
if err != nil {
return err
// Validate fingerprint format: must be exactly 40 hex characters
if len(requiredSigner) != 40 {
return fmt.Errorf("invalid fingerprint: must be exactly 40 hex characters, got %d", len(requiredSigner))
}
if _, err := hex.DecodeString(requiredSigner); err != nil {
return fmt.Errorf("invalid fingerprint: must be valid hex: %w", err)
}
if !chk.IsSigned() {
return fmt.Errorf("manifest is not signed, but signature from %s is required", requiredSigner)
}
// Extract fingerprint from the embedded public key (not from the signer field)
// This validates the key is importable and gets its actual fingerprint
embeddedFP, err := chk.ExtractEmbeddedSigningKeyFP()
if err != nil {
return fmt.Errorf("failed to extract fingerprint from embedded signing key: %w", err)
}
// Compare fingerprints - must be exact match (case-insensitive)
if !strings.EqualFold(embeddedFP, requiredSigner) {
return fmt.Errorf("embedded signing key fingerprint %s does not match required %s", embeddedFP, requiredSigner)
}
log.Infof("manifest signature verified (signer: %s)", embeddedFP)
}
log.Infof("manifest contains %d files, %s", chk.FileCount(),
humanize.IBytes(safeUint64(int64(chk.TotalBytes()))))
log.Infof("manifest contains %d files, %s", chk.FileCount(), humanize.IBytes(uint64(chk.TotalBytes())))
failures, err := runCheck(ctx, chk, showProgress)
// Set up results channel
results := make(chan mfer.Result, 1)
// Set up progress channel
var progress chan mfer.CheckStatus
if showProgress {
progress = make(chan mfer.CheckStatus, 1)
go func() {
for status := range progress {
if status.ETA > 0 {
log.Progressf("Checking: %d/%d files, %s/s, ETA %s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)),
status.ETA.Round(time.Second),
status.Failures)
} else {
log.Progressf("Checking: %d/%d files, %s/s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)),
status.Failures)
}
}
log.ProgressDone()
}()
}
// Process results in a goroutine
var failures int64
done := make(chan struct{})
go func() {
for result := range results {
if result.Status != mfer.StatusOK {
failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
} else {
log.Verbosef("%s: %s", result.Status, result.Path)
}
}
close(done)
}()
// Run check
err = chk.Check(ctx.Context, results, progress)
if err != nil {
return err
return fmt.Errorf("check failed: %w", err)
}
// Wait for results processing to complete
<-done
// Check for extra files if requested
if ctx.Bool("no-extra-files") {
extraResults := make(chan mfer.Result, 1)
extraDone := make(chan struct{})
go func() {
for result := range extraResults {
failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
}
close(extraDone)
}()
err = chk.FindExtraFiles(ctx.Context, extraResults)
if err != nil {
return fmt.Errorf("failed to check for extra files: %w", err)
}
<-extraDone
}
elapsed := time.Since(mfa.startupTime).Seconds()
rate := float64(chk.TotalBytes()) / elapsed
if failures == 0 {
log.Infof("checked %d files (%s) in %.1fs (%s/s): all OK",
chk.FileCount(), humanize.IBytes(safeUint64(int64(chk.TotalBytes()))),
elapsed, humanize.IBytes(safeRateUint64(rate)))
log.Infof("checked %d files (%s) in %.1fs (%s/s): all OK", chk.FileCount(), humanize.IBytes(uint64(chk.TotalBytes())), elapsed, humanize.IBytes(uint64(rate)))
} else {
log.Infof("checked %d files (%s) in %.1fs (%s/s): %d failed",
chk.FileCount(), humanize.IBytes(safeUint64(int64(chk.TotalBytes()))),
elapsed, humanize.IBytes(safeRateUint64(rate)), failures)
log.Infof("checked %d files (%s) in %.1fs (%s/s): %d failed", chk.FileCount(), humanize.IBytes(uint64(chk.TotalBytes())), elapsed, humanize.IBytes(uint64(rate)), failures)
}
if failures > 0 {
+7 -11
View File
@@ -7,18 +7,15 @@ import (
"github.com/spf13/afero"
)
// NoColor disables colored output when set. Automatically true if the
// NO_COLOR disables colored output when set. Automatically true if the
// NO_COLOR environment variable is present (per https://no-color.org/).
//
//nolint:gochecknoglobals // process-wide setting derived from the environment
var NoColor = noColorEnvSet()
var NO_COLOR bool
// noColorEnvSet reports whether the NO_COLOR environment variable is
// present.
func noColorEnvSet() bool {
_, exists := os.LookupEnv("NO_COLOR")
return exists
func init() {
NO_COLOR = false
if _, exists := os.LookupEnv("NO_COLOR"); exists {
NO_COLOR = true
}
}
// RunOptions contains all configuration for running the CLI application.
@@ -67,6 +64,5 @@ func RunWithOptions(opts *RunOptions) int {
}
m.run(opts.Args)
return m.exitCode
}
+206 -501
View File
File diff suppressed because it is too large Load Diff
-355
View File
@@ -1,355 +0,0 @@
//nolint:testpackage // white-box tests exercise unexported internals
package cli
import (
"bytes"
"context"
"flag"
"net/http"
"net/http/httptest"
"os"
"os/exec"
"path/filepath"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
urfcli "github.com/urfave/cli/v2"
"sneak.berlin/go/mfer/mfer"
)
// These tests pin the exact rendered text of the CLI's user-visible error
// messages. The messages are grepped for in CI pipelines and quoted in bug
// reports, so a reword is a deliberate change, never a refactoring side
// effect.
//
// Every case drives the real function that emits the message and asserts on
// what it returns. No production format string is restated here: a test that
// only re-rendered a copied format string would keep passing after the real
// message changed, which is exactly the regression these tests exist to
// catch.
// Full 40-hex fingerprints used where a message embeds one.
const (
msgFpA = "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"
msgFpB = "BBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBB"
)
// runLocked runs fn while holding runMu, so operations that write to the
// process-global logger do not race the other CLI runs.
func runLocked(fn func() error) error {
runMu.Lock()
defer runMu.Unlock()
return fn()
}
// unsignedChecker builds a Checker over a freshly scanned, unsigned manifest.
func unsignedChecker(t *testing.T) *mfer.Checker {
t.Helper()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/d", 0o755))
require.NoError(t, afero.WriteFile(fs, "/d/f.txt", []byte("hi"), 0o644))
s := mfer.NewScannerWithOptions(&mfer.ScannerOptions{Fs: fs})
require.NoError(t, s.EnumeratePath("/d", nil))
var buf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &buf, nil))
require.NoError(t, afero.WriteFile(fs, "/d/index.mf", buf.Bytes(), 0o644))
chk, err := mfer.NewChecker("/d/index.mf", "/d", fs)
require.NoError(t, err)
require.False(t, chk.IsSigned())
return chk
}
func TestNoManifestFoundMessage(t *testing.T) {
t.Parallel()
_, err := findManifest(afero.NewMemMapFs(), "/tmp/x")
require.ErrorIs(t, err, errNoManifestFound)
assert.EqualError(t, err,
"no manifest found in /tmp/x (looked for index.mf and .index.mf)")
}
func TestVerifyRequiredSignerMessages(t *testing.T) {
t.Parallel()
t.Run("invalid fingerprint length", func(t *testing.T) {
t.Parallel()
err := verifyRequiredSigner(context.Background(),
unsignedChecker(t), "12345678")
require.ErrorIs(t, err, errInvalidFingerprint)
assert.EqualError(t, err,
"invalid fingerprint: must be exactly 40 hex characters, got 8")
})
t.Run("manifest not signed", func(t *testing.T) {
t.Parallel()
err := verifyRequiredSigner(context.Background(),
unsignedChecker(t), msgFpA)
require.ErrorIs(t, err, errManifestNotSigned)
assert.EqualError(t, err,
"manifest is not signed, but signature from "+msgFpA+" is required")
})
}
// TestSignerMismatchMessage drives verifyRequiredSigner against a real signed
// manifest. The embedded fingerprint is whatever the generated key produced,
// so it is read back from the checker and substituted into the expected
// string; the required signer is a fixed value that cannot match it. Requires
// gpg and is skipped where it is absent, as the other signing tests are.
//
//nolint:paralleltest // signedChecker calls t.Setenv, which bars t.Parallel
func TestSignerMismatchMessage(t *testing.T) {
chk := signedChecker(t)
embeddedFP, err := chk.ExtractEmbeddedSigningKeyFP(context.Background())
require.NoError(t, err)
err = verifyRequiredSigner(context.Background(), chk, msgFpB)
require.ErrorIs(t, err, errSignerMismatch)
assert.EqualError(t, err,
"embedded signing key fingerprint "+embeddedFP+
" does not match required "+msgFpB)
}
// signedChecker builds a Checker over a manifest signed by a throwaway GPG
// key generated in a temporary GNUPGHOME.
func signedChecker(t *testing.T) *mfer.Checker {
t.Helper()
_, err := exec.LookPath("gpg")
if err != nil {
t.Skip("gpg not installed, skipping signing test")
}
gpgHome := t.TempDir()
params := "%no-protection\n" +
"Key-Type: RSA\nKey-Length: 2048\n" +
"Name-Real: MFER Test Key\nName-Email: test@mfer.test\n" +
"Expire-Date: 0\n%commit\n"
paramsFile := filepath.Join(gpgHome, "key-params")
require.NoError(t, os.WriteFile(paramsFile, []byte(params), 0o600))
//nolint:gosec // paramsFile is a test-controlled path inside t.TempDir()
cmd := exec.CommandContext(context.Background(), "gpg",
"--batch", "--gen-key", paramsFile)
cmd.Env = append(os.Environ(), "GNUPGHOME="+gpgHome)
out, err := cmd.CombinedOutput()
if err != nil {
t.Skipf("failed to generate test GPG key: %v: %s", err, out)
}
t.Setenv("GNUPGHOME", gpgHome)
b := mfer.NewBuilder()
b.SetSigningOptions(&mfer.SigningOptions{KeyID: mfer.GPGKeyID("test@mfer.test")})
content := []byte("signed file")
_, err = b.AddFile("f.txt", mfer.FileSize(len(content)), mfer.ModTime{},
bytes.NewReader(content), nil)
require.NoError(t, err)
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/index.mf", buf.Bytes(), 0o644))
chk, err := mfer.NewChecker("/index.mf", "/", fs)
require.NoError(t, err)
require.True(t, chk.IsSigned())
return chk
}
func TestPathDoesNotExistMessage(t *testing.T) {
t.Parallel()
set := flag.NewFlagSet("gen", flag.ContinueOnError)
require.NoError(t, set.Parse([]string{"nope"}))
mfa := &CLIApp{Fs: afero.NewMemMapFs()}
ctx := urfcli.NewContext(nil, set, nil)
_, err := mfa.collectInputPaths(ctx.Args())
require.ErrorIs(t, err, errPathNotExist)
assert.EqualError(t, err, "path does not exist: nope")
}
func TestOutputFileExistsMessage(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/d", 0o755))
require.NoError(t, afero.WriteFile(fs, "/d/f.txt", []byte("hi"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/out.mf", []byte("old"), 0o644))
set := flag.NewFlagSet("gen", flag.ContinueOnError)
set.String("output", "", "")
set.Bool("force", false, "")
require.NoError(t, set.Parse([]string{"/d"}))
require.NoError(t, set.Set("output", "/out.mf"))
mfa := &CLIApp{Fs: fs}
ctx := urfcli.NewContext(nil, set, nil)
// generateManifestOperation writes to the process-global logger during
// enumeration, so serialize with the other CLI runs.
err := runLocked(func() error { return mfa.generateManifestOperation(ctx) })
require.ErrorIs(t, err, errOutputExists)
assert.EqualError(t, err,
"output file /out.mf already exists (use --force to overwrite)")
}
// TestUnknownCommandMessage drives the root command's action. run only logs
// the error that action returns, so the test lets run build the app with no
// command given and then runs that same app on an unknown command to get the
// error itself.
func TestUnknownCommandMessage(t *testing.T) {
t.Parallel()
mfa := &CLIApp{
appname: testApp,
Stdout: &bytes.Buffer{},
Stderr: &bytes.Buffer{},
Fs: afero.NewMemMapFs(),
}
// run points the process-global logger at this app's output, so
// serialize with the other CLI runs.
err := runLocked(func() error {
mfa.run([]string{testApp})
return mfa.app.Run([]string{testApp, "bogus"})
})
require.ErrorIs(t, err, errUnknownCommand)
assert.EqualError(t, err, `unknown command "bogus"`)
}
func TestManifestLoaderHTTPStatusMessage(t *testing.T) {
t.Parallel()
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusNotFound)
}))
defer server.Close()
mfa := &CLIApp{Fs: afero.NewMemMapFs()}
_, err := mfa.openManifestReader(server.URL + "/foo.mf")
require.ErrorIs(t, err, errHTTPStatus)
assert.EqualError(t, err,
"failed to fetch "+server.URL+"/foo.mf: HTTP 404")
}
func TestFetchManifestHTTPStatusMessage(t *testing.T) {
t.Parallel()
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusNotFound)
}))
defer server.Close()
set := flag.NewFlagSet("fetch", flag.ContinueOnError)
require.NoError(t, set.Parse([]string{server.URL}))
mfa := &CLIApp{Fs: afero.NewMemMapFs()}
ctx := urfcli.NewContext(nil, set, nil)
// fetchManifestOperation logs to the process-global logger.
err := runLocked(func() error { return mfa.fetchManifestOperation(ctx) })
require.ErrorIs(t, err, errHTTPStatus)
assert.EqualError(t, err, "failed to fetch manifest: HTTP 404")
}
func TestFetchFileHTTPStatusMessage(t *testing.T) {
t.Parallel()
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusInternalServerError)
}))
defer server.Close()
err := downloadFile(context.Background(), server.URL+"/x", "x",
&mfer.MFFilePath{}, nil)
require.ErrorIs(t, err, errHTTPStatus)
assert.EqualError(t, err, "HTTP 500")
}
func TestURLRequiredMessage(t *testing.T) {
t.Parallel()
set := flag.NewFlagSet("fetch", flag.ContinueOnError)
require.NoError(t, set.Parse([]string{}))
mfa := &CLIApp{Fs: afero.NewMemMapFs()}
ctx := urfcli.NewContext(nil, set, nil)
// fetchManifestOperation logs to the process-global logger.
err := runLocked(func() error { return mfa.fetchManifestOperation(ctx) })
require.ErrorIs(t, err, errURLRequired)
assert.EqualError(t, err, "URL argument required")
}
func TestSanitizePathMessages(t *testing.T) {
t.Parallel()
t.Run("empty", func(t *testing.T) {
t.Parallel()
_, err := sanitizePath("")
require.ErrorIs(t, err, errEmptyPath)
assert.EqualError(t, err, "empty path")
})
t.Run("absolute", func(t *testing.T) {
t.Parallel()
_, err := sanitizePath("/etc/passwd")
require.ErrorIs(t, err, errAbsolutePath)
assert.EqualError(t, err, "absolute path not allowed: /etc/passwd")
})
t.Run("traversal", func(t *testing.T) {
t.Parallel()
_, err := sanitizePath("../x")
require.ErrorIs(t, err, errPathTraversal)
assert.EqualError(t, err, "path traversal not allowed: ../x")
})
}
func TestSizeMismatchMessage(t *testing.T) {
t.Parallel()
// finishDownload returns the size-mismatch error before it touches the
// paths, digest, or entry, so those can be zero here.
err := finishDownload("", "", 9, 10, nil, nil, nil, nil)
require.ErrorIs(t, err, errSizeMismatch)
assert.EqualError(t, err, "size mismatch: expected 10 bytes, got 9")
}
func TestHashMismatchMessage(t *testing.T) {
t.Parallel()
// A 32-byte digest that matches none of the (empty) manifest hashes.
err := verifyDownloadedHash(make([]byte, 32), &mfer.MFFilePath{})
require.ErrorIs(t, err, errHashMismatch)
require.NotErrorIs(t, err, errSizeMismatch)
assert.EqualError(t, err, "hash mismatch")
}
+10 -15
View File
@@ -29,7 +29,6 @@ func (mfa *CLIApp) exportManifestOperation(ctx *cli.Context) error {
if err != nil {
return fmt.Errorf("export: %w", err)
}
defer func() { _ = rc.Close() }()
manifest, err := mfer.NewManifestFromReader(rc)
@@ -42,23 +41,21 @@ func (mfa *CLIApp) exportManifestOperation(ctx *cli.Context) error {
for _, f := range files {
entry := ExportEntry{
Path: f.GetPath(),
Size: f.GetSize(),
Hashes: make([]string, 0, len(f.GetHashes())),
Path: f.Path,
Size: f.Size,
Hashes: make([]string, 0, len(f.Hashes)),
}
for _, h := range f.GetHashes() {
entry.Hashes = append(entry.Hashes, hex.EncodeToString(h.GetMultiHash()))
for _, h := range f.Hashes {
entry.Hashes = append(entry.Hashes, hex.EncodeToString(h.MultiHash))
}
if mtime, ok := entryMtime(f); ok {
t := mtime.UTC().Format(time.RFC3339Nano)
if f.Mtime != nil {
t := time.Unix(f.Mtime.Seconds, int64(f.Mtime.Nanos)).UTC().Format(time.RFC3339Nano)
entry.Mtime = &t
}
if f.GetCtime() != nil {
t := time.Unix(f.GetCtime().GetSeconds(), int64(f.GetCtime().GetNanos())).
UTC().Format(time.RFC3339Nano)
if f.Ctime != nil {
t := time.Unix(f.Ctime.Seconds, int64(f.Ctime.Nanos)).UTC().Format(time.RFC3339Nano)
entry.Ctime = &t
}
@@ -67,9 +64,7 @@ func (mfa *CLIApp) exportManifestOperation(ctx *cli.Context) error {
enc := json.NewEncoder(mfa.Stdout)
enc.SetIndent("", " ")
err = enc.Encode(entries)
if err != nil {
if err := enc.Encode(entries); err != nil {
return fmt.Errorf("export: failed to encode JSON: %w", err)
}
+17 -36
View File
@@ -14,12 +14,9 @@ import (
"sneak.berlin/go/mfer/mfer"
)
const testCmdExport = "export"
// buildTestManifest creates a manifest from in-memory files and returns its bytes.
func buildTestManifest(t *testing.T, files map[string][]byte) []byte {
t.Helper()
sourceFs := afero.NewMemMapFs()
for path, content := range files {
require.NoError(t, sourceFs.MkdirAll("/", 0o755))
@@ -31,15 +28,11 @@ func buildTestManifest(t *testing.T, files map[string][]byte) []byte {
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
var buf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &buf, nil))
return buf.Bytes()
}
func TestExportManifestOperation(t *testing.T) {
t.Parallel()
testFiles := map[string][]byte{
"hello.txt": []byte("Hello, World!"),
"sub/file.txt": []byte("nested content"),
@@ -51,10 +44,9 @@ func TestExportManifestOperation(t *testing.T) {
require.NoError(t, afero.WriteFile(fs, "/test.mf", manifestData, 0o644))
var stdout, stderr bytes.Buffer
exitCode := runCLI(&RunOptions{
Appname: testApp,
Args: []string{testApp, testCmdExport, "/test.mf"},
exitCode := RunWithOptions(&RunOptions{
Appname: "mfer",
Args: []string{"mfer", "export", "/test.mf"},
Stdin: &bytes.Buffer{},
Stdout: &stdout,
Stderr: &stderr,
@@ -72,33 +64,28 @@ func TestExportManifestOperation(t *testing.T) {
for _, e := range entries {
pathSet[e.Path] = true
assert.NotEmpty(t, e.Hashes, "entry %s should have hashes", e.Path)
assert.Positive(t, e.Size, "entry %s should have positive size", e.Path)
assert.Greater(t, e.Size, int64(0), "entry %s should have positive size", e.Path)
}
assert.True(t, pathSet["hello.txt"])
assert.True(t, pathSet["sub/file.txt"])
}
func TestExportFromHTTPURL(t *testing.T) {
t.Parallel()
testFiles := map[string][]byte{
"a.txt": []byte("aaa"),
}
manifestData := buildTestManifest(t, testFiles)
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(manifestData)
}))
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(manifestData)
}))
defer server.Close()
var stdout, stderr bytes.Buffer
exitCode := runCLI(&RunOptions{
Appname: testApp,
Args: []string{testApp, testCmdExport, server.URL + "/index.mf"},
exitCode := RunWithOptions(&RunOptions{
Appname: "mfer",
Args: []string{"mfer", "export", server.URL + "/index.mf"},
Stdin: &bytes.Buffer{},
Stdout: &stdout,
Stderr: &stderr,
@@ -114,25 +101,21 @@ func TestExportFromHTTPURL(t *testing.T) {
}
func TestListFromHTTPURL(t *testing.T) {
t.Parallel()
testFiles := map[string][]byte{
"one.txt": []byte("1"),
"two.txt": []byte("22"),
}
manifestData := buildTestManifest(t, testFiles)
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
_, _ = w.Write(manifestData)
}))
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
_, _ = w.Write(manifestData)
}))
defer server.Close()
var stdout, stderr bytes.Buffer
exitCode := runCLI(&RunOptions{
Appname: testApp,
Args: []string{testApp, "list", server.URL + "/index.mf"},
exitCode := RunWithOptions(&RunOptions{
Appname: "mfer",
Args: []string{"mfer", "list", server.URL + "/index.mf"},
Stdin: &bytes.Buffer{},
Stdout: &stdout,
Stderr: &stderr,
@@ -146,8 +129,6 @@ func TestListFromHTTPURL(t *testing.T) {
}
func TestIsHTTPURL(t *testing.T) {
t.Parallel()
assert.True(t, isHTTPURL("http://example.com/manifest.mf"))
assert.True(t, isHTTPURL("https://example.com/manifest.mf"))
assert.False(t, isHTTPURL("/local/path.mf"))
+104 -316
View File
@@ -2,9 +2,7 @@ package cli
import (
"bytes"
"context"
"crypto/sha256"
"errors"
"fmt"
"io"
"net/http"
@@ -22,53 +20,6 @@ import (
"sneak.berlin/go/mfer/mfer"
)
const (
// progressChanBuffer is the buffer size of the download progress
// channel.
progressChanBuffer = 10
// bitsPerByte converts a bytes-per-second rate to bits per second.
bitsPerByte = 8
// dirPerms is the permission mode for directories created for
// downloaded files. Fetched trees are content that is normally
// published (served by a web server, read by another uid), so the
// traversal bit for group and other must stay set.
dirPerms os.FileMode = 0o755
// filePerms is the permission mode, before the umask, for downloaded
// files. It is the mode os.Create uses; like dirPerms, it keeps group
// and other read access.
filePerms os.FileMode = 0o666
// Bitrate unit thresholds in bits per second.
bpsPerGbps = 1e9
bpsPerMbps = 1e6
bpsPerKbps = 1e3
)
var (
// errURLRequired indicates the fetch command was run without a URL
// argument.
errURLRequired = errors.New("URL argument required")
// errEmptyPath indicates an empty file path in the manifest.
errEmptyPath = errors.New("empty path")
// errAbsolutePath indicates an absolute file path in the manifest.
errAbsolutePath = errors.New("absolute path not allowed")
// errPathTraversal indicates a manifest path escaping the target
// directory.
errPathTraversal = errors.New("path traversal not allowed")
// errSymlinkInPath indicates a manifest path running through a
// symlink that already exists in the target directory.
errSymlinkInPath = errors.New("symlink in path not allowed")
// errSizeMismatch indicates a downloaded file with an unexpected
// size.
errSizeMismatch = errors.New("size mismatch")
// errHashMismatch indicates a downloaded file whose hash matches no
// manifest hash.
errHashMismatch = errors.New("hash mismatch")
)
// DownloadProgress reports the progress of a single file download.
type DownloadProgress struct {
Path string // File path being downloaded
@@ -78,98 +29,14 @@ type DownloadProgress struct {
ETA time.Duration // Estimated time to completion
}
// httpGet issues a GET request for the given URL using the provided
// context and returns the response. The caller must close the body.
//
// Errors are returned unwrapped: this helper replaced direct http.Get
// calls, and each caller already supplies its own context string, so
// adding one here would change user-visible messages.
func httpGet(ctx context.Context, fileURL string) (*http.Response, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, fileURL, nil)
if err != nil {
return nil, err
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
return resp, nil
}
// reportDownloadProgress renders download progress until the channel
// closes, then closes done.
func reportDownloadProgress(progress <-chan DownloadProgress, done chan<- struct{}) {
defer close(done)
for p := range progress {
rate := formatBitrate(p.BytesPerSec * bitsPerByte)
if p.ETA > 0 {
log.Infof("%s: %s/%s, %s, ETA %s",
p.Path, humanize.IBytes(safeUint64(p.BytesRead)),
humanize.IBytes(safeUint64(p.TotalBytes)),
rate, p.ETA.Round(time.Second))
} else {
log.Infof("%s: %s/%s, %s",
p.Path, humanize.IBytes(safeUint64(p.BytesRead)),
humanize.IBytes(safeUint64(p.TotalBytes)), rate)
}
}
}
// manifestBaseURL returns the URL of the directory containing the
// manifest, with a trailing slash.
func manifestBaseURL(manifestURL string) (*url.URL, error) {
baseURL, err := url.Parse(manifestURL)
if err != nil {
return nil, fmt.Errorf("fetch: invalid manifest URL: %w", err)
}
baseURL.Path = path.Dir(baseURL.Path)
if !strings.HasSuffix(baseURL.Path, "/") {
baseURL.Path += "/"
}
return baseURL, nil
}
// downloadManifestFiles downloads every file in the manifest, reporting
// progress on the progress channel.
func downloadManifestFiles(
ctx context.Context,
baseURL *url.URL,
files []*mfer.MFFilePath,
progress chan<- DownloadProgress,
) error {
for _, f := range files {
// Sanitize the path to prevent path traversal attacks
localPath, err := sanitizePath(f.GetPath())
if err != nil {
return fmt.Errorf("invalid path in manifest: %w", err)
}
fileURL := baseURL.String() + encodeFilePath(f.GetPath())
log.Infof("fetching %s", f.GetPath())
err = downloadFile(ctx, fileURL, localPath, f, progress)
if err != nil {
return fmt.Errorf("failed to download %s: %w", f.GetPath(), err)
}
}
return nil
}
func (mfa *CLIApp) fetchManifestOperation(ctx *cli.Context) error {
log.Debug("fetchManifestOperation()")
if ctx.Args().Len() == 0 {
return errURLRequired
return fmt.Errorf("URL argument required")
}
inputURL := ctx.Args().Get(0)
manifestURL, err := resolveManifestURL(inputURL)
if err != nil {
return fmt.Errorf("invalid URL: %w", err)
@@ -178,16 +45,14 @@ func (mfa *CLIApp) fetchManifestOperation(ctx *cli.Context) error {
log.Infof("fetching manifest from %s", manifestURL)
// Fetch manifest
resp, err := httpGet(ctx.Context, manifestURL)
resp, err := http.Get(manifestURL)
if err != nil {
return fmt.Errorf("failed to fetch manifest: %w", err)
}
defer func() { _ = resp.Body.Close() }()
if resp.StatusCode != http.StatusOK {
return fmt.Errorf("failed to fetch manifest: %w %d",
errHTTPStatus, resp.StatusCode)
return fmt.Errorf("failed to fetch manifest: HTTP %d", resp.StatusCode)
}
// Parse manifest
@@ -200,43 +65,74 @@ func (mfa *CLIApp) fetchManifestOperation(ctx *cli.Context) error {
log.Infof("manifest contains %d files", len(files))
// Compute base URL (directory containing manifest)
baseURL, err := manifestBaseURL(manifestURL)
baseURL, err := url.Parse(manifestURL)
if err != nil {
return err
return fmt.Errorf("fetch: invalid manifest URL: %w", err)
}
baseURL.Path = path.Dir(baseURL.Path)
if !strings.HasSuffix(baseURL.Path, "/") {
baseURL.Path += "/"
}
// Calculate total bytes to download
var totalBytes int64
for _, f := range files {
totalBytes += f.GetSize()
totalBytes += f.Size
}
// Create progress channel and start progress reporter goroutine
progress := make(chan DownloadProgress, progressChanBuffer)
done := make(chan struct{})
// Create progress channel
progress := make(chan DownloadProgress, 10)
go reportDownloadProgress(progress, done)
// Start progress reporter goroutine
done := make(chan struct{})
go func() {
defer close(done)
for p := range progress {
rate := formatBitrate(p.BytesPerSec * 8)
if p.ETA > 0 {
log.Infof("%s: %s/%s, %s, ETA %s",
p.Path, humanize.IBytes(uint64(p.BytesRead)), humanize.IBytes(uint64(p.TotalBytes)),
rate, p.ETA.Round(time.Second))
} else {
log.Infof("%s: %s/%s, %s",
p.Path, humanize.IBytes(uint64(p.BytesRead)), humanize.IBytes(uint64(p.TotalBytes)), rate)
}
}
}()
// Track download start time
startTime := time.Now()
// Download each file
dlErr := downloadManifestFiles(ctx.Context, baseURL, files, progress)
for _, f := range files {
// Sanitize the path to prevent path traversal attacks
localPath, err := sanitizePath(f.Path)
if err != nil {
close(progress)
<-done
return fmt.Errorf("invalid path in manifest: %w", err)
}
fileURL := baseURL.String() + encodeFilePath(f.Path)
log.Infof("fetching %s", f.Path)
if err := downloadFile(fileURL, localPath, f, progress); err != nil {
close(progress)
<-done
return fmt.Errorf("failed to download %s: %w", f.Path, err)
}
}
close(progress)
<-done
if dlErr != nil {
return dlErr
}
// Print summary
elapsed := time.Since(startTime)
avgBytesPerSec := float64(totalBytes) / elapsed.Seconds()
avgRate := formatBitrate(avgBytesPerSec * bitsPerByte)
avgRate := formatBitrate(avgBytesPerSec * 8)
log.Infof("downloaded %d files (%s) in %.1fs (%s avg)",
len(files),
humanize.IBytes(safeUint64(totalBytes)),
humanize.IBytes(uint64(totalBytes)),
elapsed.Seconds(),
avgRate)
@@ -249,7 +145,6 @@ func encodeFilePath(p string) string {
for i, seg := range segments {
segments[i] = url.PathEscape(seg)
}
return strings.Join(segments, "/")
}
@@ -258,12 +153,12 @@ func encodeFilePath(p string) string {
func sanitizePath(p string) (string, error) {
// Reject empty paths
if p == "" {
return "", errEmptyPath
return "", fmt.Errorf("empty path")
}
// Reject absolute paths
if filepath.IsAbs(p) {
return "", fmt.Errorf("%w: %s", errAbsolutePath, p)
return "", fmt.Errorf("absolute path not allowed: %s", p)
}
// Clean the path to resolve . and ..
@@ -271,46 +166,17 @@ func sanitizePath(p string) (string, error) {
// Reject paths that escape the current directory
if strings.HasPrefix(cleaned, ".."+string(filepath.Separator)) || cleaned == ".." {
return "", fmt.Errorf("%w: %s", errPathTraversal, p)
return "", fmt.Errorf("path traversal not allowed: %s", p)
}
// Also check for absolute paths after cleaning (handles edge cases)
if filepath.IsAbs(cleaned) {
return "", fmt.Errorf("%w: %s", errAbsolutePath, p)
return "", fmt.Errorf("absolute path not allowed: %s", p)
}
return cleaned, nil
}
// checkNoSymlinks returns an error if any part of the relative path p
// already exists as a symlink. sanitizePath checks p only as text, so
// without this a symlink inside the target directory could send a write
// to p outside of it. Parts that do not exist yet are fine: fetch creates
// them as plain directories and files. Call it immediately before each
// write: a symlink created after it returns is not caught.
func checkNoSymlinks(p string) error {
current := ""
for _, part := range strings.Split(p, string(filepath.Separator)) {
current = filepath.Join(current, part)
info, err := os.Lstat(current)
if errors.Is(err, os.ErrNotExist) {
return nil
}
if err != nil {
return fmt.Errorf("failed to check %s for a symlink: %w", current, err)
}
if info.Mode()&os.ModeSymlink != 0 {
return fmt.Errorf("%w: %s", errSymlinkInPath, current)
}
}
return nil
}
// resolveManifestURL takes a URL and returns the manifest URL.
// If the URL already ends with .mf, it's returned as-is.
// Otherwise, index.mf is appended.
@@ -348,14 +214,10 @@ type progressWriter struct {
func (pw *progressWriter) Write(p []byte) (int, error) {
n, err := pw.w.Write(p)
pw.written += int64(n)
if pw.progress != nil {
var (
bytesPerSec float64
eta time.Duration
)
var bytesPerSec float64
var eta time.Duration
elapsed := time.Since(pw.startTime)
if elapsed > 0 && pw.written > 0 {
bytesPerSec = float64(pw.written) / elapsed.Seconds()
@@ -364,7 +226,6 @@ func (pw *progressWriter) Write(p []byte) (int, error) {
eta = time.Duration(float64(remainingBytes)/bytesPerSec) * time.Second
}
}
sendProgress(pw.progress, DownloadProgress{
Path: pw.path,
BytesRead: pw.written,
@@ -373,19 +234,18 @@ func (pw *progressWriter) Write(p []byte) (int, error) {
ETA: eta,
})
}
return n, err
}
// formatBitrate formats a bits-per-second value with appropriate unit prefix.
func formatBitrate(bps float64) string {
switch {
case bps >= bpsPerGbps:
return fmt.Sprintf("%.1f Gbps", bps/bpsPerGbps)
case bps >= bpsPerMbps:
return fmt.Sprintf("%.1f Mbps", bps/bpsPerMbps)
case bps >= bpsPerKbps:
return fmt.Sprintf("%.1f Kbps", bps/bpsPerKbps)
case bps >= 1e9:
return fmt.Sprintf("%.1f Gbps", bps/1e9)
case bps >= 1e6:
return fmt.Sprintf("%.1f Mbps", bps/1e6)
case bps >= 1e3:
return fmt.Sprintf("%.1f Kbps", bps/1e3)
default:
return fmt.Sprintf("%.0f bps", bps)
}
@@ -399,115 +259,53 @@ func sendProgress(ch chan<- DownloadProgress, p DownloadProgress) {
}
}
// tempPathFor computes the temporary download path for a local file.
// For dotfiles, just append .tmp (they're already hidden); for regular
// files, prefix with . and append .tmp.
func tempPathFor(localPath string) string {
// downloadFile downloads a URL to a local file path with hash verification.
// It downloads to a temporary file, verifies the hash, then renames to the final path.
// Progress is reported via the progress channel.
func downloadFile(fileURL, localPath string, entry *mfer.MFFilePath, progress chan<- DownloadProgress) error {
// Create parent directories if needed
dir := filepath.Dir(localPath)
base := filepath.Base(localPath)
if dir != "" && dir != "." {
if err := os.MkdirAll(dir, 0o755); err != nil {
return fmt.Errorf("failed to create directory %s: %w", dir, err)
}
}
// Compute temp file path in the same directory
// For dotfiles, just append .tmp (they're already hidden)
// For regular files, prefix with . and append .tmp
base := filepath.Base(localPath)
var tmpName string
if strings.HasPrefix(base, ".") {
tmpName = base + ".tmp"
} else {
tmpName = "." + base + ".tmp"
}
tmpPath := filepath.Join(dir, tmpName)
if dir == "" || dir == "." {
return tmpName
tmpPath = tmpName
}
return filepath.Join(dir, tmpName)
}
// verifyDownloadedHash checks the computed sha256 digest against the
// manifest entry's hashes; at least one must match.
func verifyDownloadedHash(digest []byte, entry *mfer.MFFilePath) error {
computed, err := multihash.Encode(digest, multihash.SHA2_256)
if err != nil {
return fmt.Errorf("failed to encode hash: %w", err)
}
for _, hash := range entry.GetHashes() {
if bytes.Equal(computed, hash.GetMultiHash()) {
return nil
}
}
return errHashMismatch
}
// downloadFile downloads a URL to a local file path with hash verification.
// It downloads to a temporary file, verifies the hash, then renames to the final path.
// Progress is reported via the progress channel.
func downloadFile(
ctx context.Context,
fileURL, localPath string,
entry *mfer.MFFilePath,
progress chan<- DownloadProgress,
) error {
// Enforce the path invariant here rather than relying on the caller,
// so every entry point to downloadFile gets the same treatment.
localPath, err := sanitizePath(localPath)
if err != nil {
return fmt.Errorf("invalid path: %w", err)
}
// Create parent directories if needed
dir := filepath.Dir(localPath)
if dir != "" && dir != "." {
err = checkNoSymlinks(dir)
if err != nil {
return err
}
err = os.MkdirAll(dir, dirPerms)
if err != nil {
return fmt.Errorf("failed to create directory %s: %w", dir, err)
}
}
tmpPath := tempPathFor(localPath)
// Fetch file
resp, err := httpGet(ctx, fileURL)
resp, err := http.Get(fileURL) //nolint:gosec // URL constructed from manifest base
if err != nil {
return fmt.Errorf("HTTP request failed: %w", err)
}
defer func() { _ = resp.Body.Close() }()
if resp.StatusCode != http.StatusOK {
return fmt.Errorf("%w %d", errHTTPStatus, resp.StatusCode)
return fmt.Errorf("HTTP %d", resp.StatusCode)
}
// Determine expected size
expectedSize := entry.GetSize()
expectedSize := entry.Size
totalBytes := resp.ContentLength
if totalBytes < 0 {
totalBytes = expectedSize
}
err = checkNoSymlinks(tmpPath)
if err != nil {
return err
}
// Remove whatever is at tmpPath, such as a leftover from an
// interrupted run, rather than write into it: it may be a hard link
// to a file outside the target directory, and removing a hard link
// removes only this name. If the removal fails, O_EXCL below makes
// the create fail.
_ = os.Remove(tmpPath)
// Create the temp file only if nothing is at tmpPath (O_EXCL).
//
// G304: tmpPath is a relative path that sanitizePath keeps inside the
// target directory as text, and checkNoSymlinks just found no symlink
// in it.
out, err := os.OpenFile( //nolint:gosec // G304: see comment above
tmpPath, os.O_RDWR|os.O_CREATE|os.O_EXCL, filePerms)
// Create temp file
out, err := os.Create(tmpPath)
if err != nil {
return fmt.Errorf("failed to create temp file: %w", err)
}
@@ -530,55 +328,45 @@ func downloadFile(
// Close file before checking errors (to flush writes)
closeErr := out.Close()
err = finishDownload(
tmpPath, localPath, written, expectedSize, h.Sum(nil), entry,
copyErr, closeErr)
if err != nil {
_ = os.Remove(tmpPath)
return err
}
return nil
}
// finishDownload validates the copy result, verifies size and hash, and
// moves the temp file into place. On error the caller removes tmpPath.
func finishDownload(
tmpPath, localPath string,
written, expectedSize int64,
digest []byte,
entry *mfer.MFFilePath,
copyErr, closeErr error,
) error {
// If copy failed, clean up temp file and return error
if copyErr != nil {
_ = os.Remove(tmpPath)
return copyErr
}
if closeErr != nil {
_ = os.Remove(tmpPath)
return closeErr
}
// Verify size
if written != expectedSize {
return fmt.Errorf("%w: expected %d bytes, got %d",
errSizeMismatch, expectedSize, written)
_ = os.Remove(tmpPath)
return fmt.Errorf("size mismatch: expected %d bytes, got %d", expectedSize, written)
}
// Encode computed hash as multihash
computed, err := multihash.Encode(h.Sum(nil), multihash.SHA2_256)
if err != nil {
_ = os.Remove(tmpPath)
return fmt.Errorf("failed to encode hash: %w", err)
}
// Verify hash against manifest (at least one must match)
err := verifyDownloadedHash(digest, entry)
if err != nil {
return err
hashMatch := false
for _, hash := range entry.Hashes {
if bytes.Equal(computed, hash.MultiHash) {
hashMatch = true
break
}
}
err = checkNoSymlinks(localPath)
if err != nil {
return err
if !hashMatch {
_ = os.Remove(tmpPath)
return fmt.Errorf("hash mismatch")
}
// Rename temp file to final path
err = os.Rename(tmpPath, localPath)
if err != nil {
if err := os.Rename(tmpPath, localPath); err != nil {
_ = os.Remove(tmpPath)
return fmt.Errorf("failed to rename temp file: %w", err)
}
+140 -274
View File
@@ -1,4 +1,3 @@
//nolint:testpackage // white-box tests exercise unexported internals
package cli
import (
@@ -17,25 +16,13 @@ import (
"sneak.berlin/go/mfer/mfer"
)
const (
testFileTxt = "file.txt"
testDirFile = "dir/file.txt"
testIndexMF = "https://example.com/path/index.mf"
// Exactly what url.Parse renders, with no wrapper of our own.
urlParseControlCharErr = `parse "http://example.com/\x7f": ` +
`net/url: invalid control character in URL`
)
func TestEncodeFilePath(t *testing.T) {
t.Parallel()
tests := []struct {
input string
expected string
}{
{testFileTxt, testFileTxt},
{testDirFile, testDirFile},
{"file.txt", "file.txt"},
{"dir/file.txt", "dir/file.txt"},
{"my file.txt", "my%20file.txt"},
{"dir/my file.txt", "dir/my%20file.txt"},
{"file#1.txt", "file%231.txt"},
@@ -46,8 +33,6 @@ func TestEncodeFilePath(t *testing.T) {
for _, tt := range tests {
t.Run(tt.input, func(t *testing.T) {
t.Parallel()
result := encodeFilePath(tt.input)
assert.Equal(t, tt.expected, result)
})
@@ -55,27 +40,23 @@ func TestEncodeFilePath(t *testing.T) {
}
func TestSanitizePath(t *testing.T) {
t.Parallel()
// Valid paths that should be accepted
validTests := []struct {
input string
expected string
}{
{testFileTxt, testFileTxt},
{testDirFile, testDirFile},
{"file.txt", "file.txt"},
{"dir/file.txt", "dir/file.txt"},
{"dir/subdir/file.txt", "dir/subdir/file.txt"},
{"./file.txt", testFileTxt},
{"./dir/file.txt", testDirFile},
{"dir/./file.txt", testDirFile},
{"./file.txt", "file.txt"},
{"./dir/file.txt", "dir/file.txt"},
{"dir/./file.txt", "dir/file.txt"},
}
for _, tt := range validTests {
t.Run("valid:"+tt.input, func(t *testing.T) {
t.Parallel()
result, err := sanitizePath(tt.input)
require.NoError(t, err)
assert.NoError(t, err)
assert.Equal(t, tt.expected, result)
})
}
@@ -97,8 +78,6 @@ func TestSanitizePath(t *testing.T) {
for _, tt := range invalidTests {
t.Run("invalid:"+tt.desc, func(t *testing.T) {
t.Parallel()
_, err := sanitizePath(tt.input)
assert.Error(t, err, "expected error for path: %s", tt.input)
})
@@ -106,115 +85,36 @@ func TestSanitizePath(t *testing.T) {
}
func TestResolveManifestURL(t *testing.T) {
t.Parallel()
tests := []struct {
input string
expected string
}{
// Already ends with .mf - use as-is
{testIndexMF, testIndexMF},
{"https://example.com/path/index.mf", "https://example.com/path/index.mf"},
{"https://example.com/path/custom.mf", "https://example.com/path/custom.mf"},
{"https://example.com/foo.mf", "https://example.com/foo.mf"},
// Directory with trailing slash - append index.mf
{"https://example.com/path/", testIndexMF},
{"https://example.com/path/", "https://example.com/path/index.mf"},
{"https://example.com/", "https://example.com/index.mf"},
// Directory without trailing slash - add slash and index.mf
{"https://example.com/path", testIndexMF},
{"https://example.com/path", "https://example.com/path/index.mf"},
{"https://example.com", "https://example.com/index.mf"},
// With query strings
{
"https://example.com/path?foo=bar",
"https://example.com/path/index.mf?foo=bar",
},
{"https://example.com/path?foo=bar", "https://example.com/path/index.mf?foo=bar"},
}
for _, tt := range tests {
t.Run(tt.input, func(t *testing.T) {
t.Parallel()
result, err := resolveManifestURL(tt.input)
require.NoError(t, err)
assert.NoError(t, err)
assert.Equal(t, tt.expected, result)
})
}
// The sole caller wraps this error as "invalid URL: %w", so
// resolveManifestURL must return url.Parse's error unadorned.
t.Run("invalid:control character", func(t *testing.T) {
t.Parallel()
_, err := resolveManifestURL("http://example.com/\x7f")
require.ErrorContains(t, err, urlParseControlCharErr)
assert.NotContains(t, err.Error(), "failed to parse URL")
})
}
// scanToManifest scans sourceFs and returns the serialized manifest bytes.
func scanToManifest(t *testing.T, sourceFs afero.Fs) []byte {
t.Helper()
s := mfer.NewScannerWithOptions(&mfer.ScannerOptions{Fs: sourceFs})
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
return manifestBuf.Bytes()
}
// chdirTemp switches the working directory to a fresh temp dir for the
// duration of the test and returns its path.
func chdirTemp(t *testing.T) string {
t.Helper()
destDir := t.TempDir()
origDir, err := os.Getwd()
require.NoError(t, err)
require.NoError(t, os.Chdir(destDir))
t.Cleanup(func() { _ = os.Chdir(origDir) })
return destDir
}
// fetchTestHandler serves the manifest at /index.mf and the given files
// at their paths.
func fetchTestHandler(
manifestData []byte, testFiles map[string][]byte,
) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
path := r.URL.Path
if path == "/index.mf" {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(manifestData)
return
}
// Strip leading slash
if len(path) > 0 && path[0] == '/' {
path = path[1:]
}
content, exists := testFiles[path]
if !exists {
http.NotFound(w, r)
return
}
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(content)
}
}
//nolint:paralleltest // changes the process-global working directory
func TestFetchFromHTTP(t *testing.T) {
// Create source filesystem with test files
sourceFs := afero.NewMemMapFs()
@@ -234,14 +134,51 @@ func TestFetchFromHTTP(t *testing.T) {
}
// Generate manifest using scanner
manifestData := scanToManifest(t, sourceFs)
opts := &mfer.ScannerOptions{
Fs: sourceFs,
}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
manifestData := manifestBuf.Bytes()
// Create HTTP server that serves the source filesystem
server := httptest.NewServer(fetchTestHandler(manifestData, testFiles))
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
path := r.URL.Path
if path == "/index.mf" {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(manifestData)
return
}
// Strip leading slash
if len(path) > 0 && path[0] == '/' {
path = path[1:]
}
content, exists := testFiles[path]
if !exists {
http.NotFound(w, r)
return
}
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(content)
}))
defer server.Close()
// Change to a fresh destination directory for the test
destDir := chdirTemp(t)
// Create destination directory
destDir, err := os.MkdirTemp("", "mfer-fetch-test-*")
require.NoError(t, err)
defer func() { _ = os.RemoveAll(destDir) }()
// Change to dest directory for the test
origDir, err := os.Getwd()
require.NoError(t, err)
require.NoError(t, os.Chdir(destDir))
defer func() { _ = os.Chdir(origDir) }()
// Parse the manifest to get file entries
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestData))
@@ -252,125 +189,132 @@ func TestFetchFromHTTP(t *testing.T) {
// Download each file using downloadFile
progress := make(chan DownloadProgress, 10)
go func() {
for p := range progress {
_ = p // drain progress channel
for range progress {
// Drain progress channel
}
}()
baseURL := server.URL + "/"
for _, f := range files {
localPath, err := sanitizePath(f.GetPath())
localPath, err := sanitizePath(f.Path)
require.NoError(t, err)
fileURL := baseURL + f.GetPath()
err = downloadFile(context.Background(), fileURL, localPath, f, progress)
require.NoError(t, err, "failed to download %s", f.GetPath())
fileURL := baseURL + f.Path
err = downloadFile(fileURL, localPath, f, progress)
require.NoError(t, err, "failed to download %s", f.Path)
}
close(progress)
// Verify downloaded files match originals
for path, expectedContent := range testFiles {
downloadedPath := filepath.Join(destDir, path)
//nolint:gosec // test-controlled path
downloadedContent, err := os.ReadFile(downloadedPath)
require.NoError(t, err, "failed to read downloaded file %s", path)
assert.Equal(t, expectedContent, downloadedContent,
"content mismatch for %s", path)
assert.Equal(t, expectedContent, downloadedContent, "content mismatch for %s", path)
}
}
//nolint:paralleltest // changes the process-global working directory
func TestFetchHashMismatch(t *testing.T) {
// Create source filesystem with a test file
sourceFs := afero.NewMemMapFs()
originalContent := []byte("Original content")
require.NoError(t, afero.WriteFile(sourceFs, "/file.txt", originalContent, 0o644))
// Generate and parse manifest
manifestData := scanToManifest(t, sourceFs)
// Generate manifest
opts := &mfer.ScannerOptions{Fs: sourceFs}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestData))
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
// Parse manifest
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestBuf.Bytes()))
require.NoError(t, err)
files := manifest.Files()
require.Len(t, files, 1)
// Create server that serves DIFFERENT content (to trigger hash mismatch)
tamperedContent := []byte("Tampered content!")
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(tamperedContent)
}))
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(tamperedContent)
}))
defer server.Close()
// Work in a fresh temp directory
chdirTemp(t)
// Create temp directory
destDir, err := os.MkdirTemp("", "mfer-fetch-hash-test-*")
require.NoError(t, err)
defer func() { _ = os.RemoveAll(destDir) }()
origDir, err := os.Getwd()
require.NoError(t, err)
require.NoError(t, os.Chdir(destDir))
defer func() { _ = os.Chdir(origDir) }()
// Try to download - should fail with hash mismatch
err = downloadFile(context.Background(),
server.URL+"/file.txt", testFileTxt, files[0], nil)
require.Error(t, err)
err = downloadFile(server.URL+"/file.txt", "file.txt", files[0], nil)
assert.Error(t, err)
assert.Contains(t, err.Error(), "mismatch")
// Verify temp file was cleaned up
_, err = os.Stat(".file.txt.tmp")
assert.True(t, os.IsNotExist(err),
"temp file should be cleaned up on hash mismatch")
assert.True(t, os.IsNotExist(err), "temp file should be cleaned up on hash mismatch")
// Verify final file was not created
_, err = os.Stat(testFileTxt)
assert.True(t, os.IsNotExist(err),
"final file should not exist on hash mismatch")
_, err = os.Stat("file.txt")
assert.True(t, os.IsNotExist(err), "final file should not exist on hash mismatch")
}
//nolint:paralleltest // changes the process-global working directory
func TestFetchSizeMismatch(t *testing.T) {
// Create source filesystem with a test file
sourceFs := afero.NewMemMapFs()
originalContent := []byte("Original content with specific size")
require.NoError(t, afero.WriteFile(sourceFs, "/file.txt", originalContent, 0o644))
// Generate and parse manifest
manifestData := scanToManifest(t, sourceFs)
// Generate manifest
opts := &mfer.ScannerOptions{Fs: sourceFs}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestData))
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
// Parse manifest
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestBuf.Bytes()))
require.NoError(t, err)
files := manifest.Files()
require.Len(t, files, 1)
// Create server that serves content with wrong size
wrongSizeContent := []byte("Short")
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(wrongSizeContent)
}))
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(wrongSizeContent)
}))
defer server.Close()
// Work in a fresh temp directory
chdirTemp(t)
// Create temp directory
destDir, err := os.MkdirTemp("", "mfer-fetch-size-test-*")
require.NoError(t, err)
defer func() { _ = os.RemoveAll(destDir) }()
origDir, err := os.Getwd()
require.NoError(t, err)
require.NoError(t, os.Chdir(destDir))
defer func() { _ = os.Chdir(origDir) }()
// Try to download - should fail with size mismatch
err = downloadFile(context.Background(),
server.URL+"/file.txt", testFileTxt, files[0], nil)
require.Error(t, err)
err = downloadFile(server.URL+"/file.txt", "file.txt", files[0], nil)
assert.Error(t, err)
assert.Contains(t, err.Error(), "size mismatch")
// Verify temp file was cleaned up
_, err = os.Stat(".file.txt.tmp")
assert.True(t, os.IsNotExist(err),
"temp file should be cleaned up on size mismatch")
assert.True(t, os.IsNotExist(err), "temp file should be cleaned up on size mismatch")
}
//nolint:paralleltest // changes the process-global working directory
func TestFetchProgress(t *testing.T) {
// Create source filesystem with a larger test file
sourceFs := afero.NewMemMapFs()
@@ -378,47 +322,53 @@ func TestFetchProgress(t *testing.T) {
content := bytes.Repeat([]byte("x"), 100*1024) // 100KB
require.NoError(t, afero.WriteFile(sourceFs, "/large.txt", content, 0o644))
// Generate and parse manifest
manifestData := scanToManifest(t, sourceFs)
// Generate manifest
opts := &mfer.ScannerOptions{Fs: sourceFs}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestData))
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
// Parse manifest
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestBuf.Bytes()))
require.NoError(t, err)
files := manifest.Files()
require.Len(t, files, 1)
// Create server that serves the content
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream")
w.Header().Set("Content-Length", "102400")
// Write in chunks to allow progress reporting
reader := bytes.NewReader(content)
_, _ = io.Copy(w, reader)
}))
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream")
w.Header().Set("Content-Length", "102400")
// Write in chunks to allow progress reporting
reader := bytes.NewReader(content)
_, _ = io.Copy(w, reader)
}))
defer server.Close()
// Work in a fresh temp directory
chdirTemp(t)
// Create temp directory
destDir, err := os.MkdirTemp("", "mfer-fetch-progress-test-*")
require.NoError(t, err)
defer func() { _ = os.RemoveAll(destDir) }()
origDir, err := os.Getwd()
require.NoError(t, err)
require.NoError(t, os.Chdir(destDir))
defer func() { _ = os.Chdir(origDir) }()
// Set up progress channel and collect updates
progress := make(chan DownloadProgress, 100)
var progressUpdates []DownloadProgress
done := make(chan struct{})
go func() {
for p := range progress {
progressUpdates = append(progressUpdates, p)
}
close(done)
}()
// Download
err = downloadFile(context.Background(),
server.URL+"/large.txt", "large.txt", files[0], progress)
err = downloadFile(server.URL+"/large.txt", "large.txt", files[0], progress)
close(progress)
<-done
@@ -430,8 +380,7 @@ func TestFetchProgress(t *testing.T) {
// Verify final progress shows complete
if len(progressUpdates) > 0 {
last := progressUpdates[len(progressUpdates)-1]
assert.Equal(t, int64(len(content)), last.BytesRead,
"final progress should show all bytes read")
assert.Equal(t, int64(len(content)), last.BytesRead, "final progress should show all bytes read")
assert.Equal(t, "large.txt", last.Path)
}
@@ -440,86 +389,3 @@ func TestFetchProgress(t *testing.T) {
require.NoError(t, err)
assert.Equal(t, content, downloaded)
}
// TestFetchRefusesSymlinks runs fetch into a destination directory that
// holds a symlink pointing outside it, in each of the three places fetch
// writes: a parent directory, the temp file, and the file itself, which
// the temp file is renamed onto; and once as a directory inside a plain
// directory. The fetch must fail and nothing outside may change.
//
//nolint:paralleltest // changes the process-global working directory
func TestFetchRefusesSymlinks(t *testing.T) {
tests := []struct {
name string
entry string // the manifest's only file
link string // symlink placed in the destination directory
target string // what link points to, relative to the outside directory
}{
{"parent directory", "sub/deeper/file.txt", "sub", "."},
{"directory inside a plain directory", "docs/data/passwd", "docs/data", "."},
{"temp file", testFileTxt, ".file.txt.tmp", "new.txt"},
{"file", testFileTxt, testFileTxt, "new.txt"},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
content := []byte("fetched")
sourceFs := afero.NewMemMapFs()
require.NoError(t, sourceFs.MkdirAll(filepath.Dir("/"+tt.entry), 0o755))
require.NoError(t, afero.WriteFile(sourceFs, "/"+tt.entry, content, 0o644))
server := httptest.NewServer(fetchTestHandler(
scanToManifest(t, sourceFs), map[string][]byte{tt.entry: content}))
defer server.Close()
outside := t.TempDir()
chdirTemp(t)
require.NoError(t, os.MkdirAll(filepath.Dir(tt.link), 0o750))
require.NoError(t, os.Symlink(filepath.Join(outside, tt.target), tt.link))
opts := testOpts([]string{testApp, cmdFetch, "-q", server.URL}, afero.NewOsFs())
assert.Equal(t, 1, runCLI(opts))
assert.Contains(t, testStderr(t, opts), "failed to download "+tt.entry+
": symlink in path not allowed: "+tt.link)
written, err := os.ReadDir(outside)
require.NoError(t, err)
assert.Empty(t, written, "fetch wrote outside the destination")
})
}
}
// TestFetchReplacesHardLinkAtTempName runs fetch into a destination
// directory that holds, at the temp file's name, a hard link to a file
// outside it. To fetch that is an ordinary leftover from an interrupted
// earlier run: it must replace it and succeed, and the outside file must
// not change.
//
//nolint:paralleltest // changes the process-global working directory
func TestFetchReplacesHardLinkAtTempName(t *testing.T) {
content := []byte("fetched")
sourceFs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(sourceFs, "/"+testFileTxt, content, 0o644))
server := httptest.NewServer(fetchTestHandler(
scanToManifest(t, sourceFs), map[string][]byte{testFileTxt: content}))
defer server.Close()
outsideFile := filepath.Join(t.TempDir(), "secret.txt")
require.NoError(t, os.WriteFile(outsideFile, []byte("outside"), 0o600))
chdirTemp(t)
require.NoError(t, os.Link(outsideFile, ".file.txt.tmp"))
opts := testOpts([]string{testApp, cmdFetch, "-q", server.URL}, afero.NewOsFs())
require.Equal(t, 0, runCLI(opts), testStderr(t, opts))
fetched, err := os.ReadFile(testFileTxt)
require.NoError(t, err)
assert.Equal(t, content, fetched)
outside, err := os.ReadFile(outsideFile) //nolint:gosec // test-controlled path
require.NoError(t, err)
assert.Equal(t, "outside", string(outside), "fetch wrote outside the destination")
}
+232 -461
View File
@@ -1,9 +1,7 @@
package cli
import (
"context"
"crypto/sha256"
"errors"
"fmt"
"io"
"io/fs"
@@ -18,19 +16,6 @@ import (
"sneak.berlin/go/mfer/mfer"
)
const (
// hashBufSize is the read buffer size used when hashing files.
hashBufSize = 64 * 1024
// scanProgressInterval is how many scanned files pass between
// progress updates.
scanProgressInterval = 100
)
// errEntryMissingMtime indicates a manifest entry that carries no
// modification time where one is required to carry it forward unchanged.
var errEntryMissingMtime = errors.New("manifest entry has no mtime")
// FreshenStatus contains progress information for the freshen operation.
type FreshenStatus struct {
Phase string // "scan" or "hash"
@@ -51,292 +36,195 @@ type freshenEntry struct {
existing *mfer.MFFilePath // existing manifest entry if unchanged
}
// freshenScanner walks the filesystem and compares it against the
// entries of an existing manifest.
type freshenScanner struct {
fs afero.Fs
absBase string
manifestBase string
includeDotfiles bool
followSymlinks bool
showProgress bool
existingByPath map[string]*mfer.MFFilePath
func (mfa *CLIApp) freshenManifestOperation(ctx *cli.Context) error {
log.Debug("freshenManifestOperation()")
entries []*freshenEntry
scanCount int64
changed int64
added int64
unchanged int64
}
basePath := ctx.String("base")
showProgress := ctx.Bool("progress")
includeDotfiles := ctx.Bool("include-dotfiles")
followSymlinks := ctx.Bool("follow-symlinks")
// resolveSymlink resolves a symlink to its target's FileInfo. The
// second return value is false when the entry should be skipped.
func (s *freshenScanner) resolveSymlink(path string) (fs.FileInfo, bool) {
if !s.followSymlinks {
return nil, false
}
// Find manifest file
var manifestPath string
var err error
realPath, err := filepath.EvalSymlinks(path)
if err != nil {
return nil, false // Skip broken symlinks
}
realInfo, err := s.fs.Stat(realPath)
if err != nil || realInfo.IsDir() {
return nil, false
}
return realInfo, true
}
// recordEntry classifies a scanned file as changed, unchanged, or added
// relative to the existing manifest.
func (s *freshenScanner) recordEntry(relPath string, info fs.FileInfo) {
existing, inManifest := s.existingByPath[relPath]
if !inManifest {
s.added++
log.Verbosef("A %s", relPath)
s.entries = append(s.entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
return
}
// Check if changed (size or mtime). An entry with no recorded mtime
// cannot be compared, so it counts as changed and gets re-hashed;
// silently treating the absent mtime as the Unix epoch would classify
// every such entry as changed without saying why.
existingMtime, haveMtime := entryMtime(existing)
if !haveMtime {
log.Debugf("%s: manifest entry has no mtime, treating as changed",
relPath)
}
if !haveMtime || existing.GetSize() != info.Size() ||
!existingMtime.Equal(info.ModTime()) {
s.changed++
log.Verbosef("M %s", relPath)
s.entries = append(s.entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
if ctx.Args().Len() > 0 {
arg := ctx.Args().Get(0)
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
manifestPath, err = findManifest(mfa.Fs, arg)
if err != nil {
return fmt.Errorf("freshen: %w", err)
}
} else {
manifestPath = arg
}
} else {
s.unchanged++
s.entries = append(s.entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: false,
existing: existing,
})
}
// Mark as seen
delete(s.existingByPath, relPath)
}
// walk is the afero.Walk callback for the scan phase.
func (s *freshenScanner) walk(path string, info fs.FileInfo, walkErr error) error {
if walkErr != nil {
return walkErr
manifestPath, err = findManifest(mfa.Fs, ".")
if err != nil {
return fmt.Errorf("freshen: %w", err)
}
}
// Get relative path
relPath, err := filepath.Rel(s.absBase, path)
log.Infof("loading manifest from %s", manifestPath)
// Load existing manifest
manifest, err := mfer.NewManifestFromFile(mfa.Fs, manifestPath)
if err != nil {
return fmt.Errorf(
"freshen: failed to compute relative path for %s: %w", path, err)
return fmt.Errorf("failed to load manifest: %w", err)
}
// Skip the manifest file itself
if relPath == s.manifestBase || relPath == "."+s.manifestBase {
return nil
existingFiles := manifest.Files()
log.Infof("manifest contains %d files", len(existingFiles))
// Build map of existing entries by path
existingByPath := make(map[string]*mfer.MFFilePath, len(existingFiles))
for _, f := range existingFiles {
existingByPath[f.Path] = f
}
// Handle dotfiles
if !s.includeDotfiles && mfer.IsHiddenPath(filepath.ToSlash(relPath)) {
if info.IsDir() {
return filepath.SkipDir
// Phase 1: Scan filesystem
log.Infof("scanning filesystem...")
startScan := time.Now()
var entries []*freshenEntry
var scanCount int64
var removed, changed, added, unchanged int64
absBase, err := filepath.Abs(basePath)
if err != nil {
return fmt.Errorf("freshen: invalid base path: %w", err)
}
err = afero.Walk(mfa.Fs, absBase, func(path string, info fs.FileInfo, walkErr error) error {
if walkErr != nil {
return walkErr
}
return nil
}
// Get relative path
relPath, err := filepath.Rel(absBase, path)
if err != nil {
return fmt.Errorf("freshen: failed to compute relative path for %s: %w", path, err)
}
// Skip directories
if info.IsDir() {
return nil
}
// Handle symlinks
if info.Mode()&fs.ModeSymlink != 0 {
realInfo, keep := s.resolveSymlink(path)
if !keep {
// Skip the manifest file itself
if relPath == filepath.Base(manifestPath) || relPath == "."+filepath.Base(manifestPath) {
return nil
}
info = realInfo
}
s.scanCount++
// Check against existing manifest
s.recordEntry(relPath, info)
// Report scan progress
if s.showProgress && s.scanCount%scanProgressInterval == 0 {
log.Progressf("Scanning: %d files found", s.scanCount)
}
return nil
}
// resolveFreshenManifestPath determines the manifest path from the CLI
// arguments, searching directories for a manifest where needed.
func (mfa *CLIApp) resolveFreshenManifestPath(ctx *cli.Context) (string, error) {
if ctx.Args().Len() == 0 {
return findManifest(mfa.Fs, ".")
}
arg := ctx.Args().Get(0)
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
return findManifest(mfa.Fs, arg)
}
return arg, nil
}
// freshenHasher hashes changed and added files and feeds all entries to
// a manifest builder.
type freshenHasher struct {
fs afero.Fs
absBase string
showProgress bool
totalHashBytes int64
filesToHash int64
startHash time.Time
builder *mfer.Builder
hashedFiles int64
hashedBytes int64
}
// reportProgress renders hashing progress for the current byte count.
func (h *freshenHasher) reportProgress(n int64) {
if !h.showProgress {
return
}
currentBytes := h.hashedBytes + n
elapsed := time.Since(h.startHash)
var (
rate float64
eta time.Duration
)
if elapsed > 0 && currentBytes > 0 {
rate = float64(currentBytes) / elapsed.Seconds()
remaining := h.totalHashBytes - currentBytes
if rate > 0 {
eta = time.Duration(float64(remaining)/rate) * time.Second
// Handle dotfiles
if !includeDotfiles && mfer.IsHiddenPath(filepath.ToSlash(relPath)) {
if info.IsDir() {
return filepath.SkipDir
}
return nil
}
}
if eta > 0 {
log.Progressf("Hashing: %d/%d files, %s/s, ETA %s",
h.hashedFiles, h.filesToHash, humanize.IBytes(safeRateUint64(rate)),
eta.Round(time.Second))
} else {
log.Progressf("Hashing: %d/%d files, %s/s",
h.hashedFiles, h.filesToHash, humanize.IBytes(safeRateUint64(rate)))
}
}
// Skip directories
if info.IsDir() {
return nil
}
// processEntry hashes the entry if needed and adds it to the builder.
func (h *freshenHasher) processEntry(e *freshenEntry) error {
if !e.needsHash {
// Use existing entry
err := addExistingToBuilder(h.builder, e.existing)
if err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
// Handle symlinks
if info.Mode()&fs.ModeSymlink != 0 {
if !followSymlinks {
return nil
}
realPath, err := filepath.EvalSymlinks(path)
if err != nil {
return nil // Skip broken symlinks
}
realInfo, err := mfa.Fs.Stat(realPath)
if err != nil || realInfo.IsDir() {
return nil
}
info = realInfo
}
scanCount++
// Check against existing manifest
existing, inManifest := existingByPath[relPath]
if inManifest {
// Check if changed (size or mtime)
existingMtime := time.Unix(existing.Mtime.Seconds, int64(existing.Mtime.Nanos))
if existing.Size != info.Size() || !existingMtime.Equal(info.ModTime()) {
changed++
log.Verbosef("M %s", relPath)
entries = append(entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
} else {
unchanged++
entries = append(entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: false,
existing: existing,
})
}
// Mark as seen
delete(existingByPath, relPath)
} else {
added++
log.Verbosef("A %s", relPath)
entries = append(entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
}
// Report scan progress
if showProgress && scanCount%100 == 0 {
log.Progressf("Scanning: %d files found", scanCount)
}
return nil
})
if showProgress {
log.ProgressDone()
}
// Need to read and hash the file
absPath := filepath.Join(h.absBase, e.path)
f, err := h.fs.Open(absPath)
if err != nil {
return fmt.Errorf("failed to open %s: %w", e.path, err)
}
hash, bytesRead, err := hashFile(f, h.reportProgress)
_ = f.Close()
if err != nil {
return fmt.Errorf("failed to hash %s: %w", e.path, err)
return fmt.Errorf("failed to scan filesystem: %w", err)
}
h.hashedBytes += bytesRead
h.hashedFiles++
// Add to builder with computed hash
err = addFileToBuilder(h.builder, e.path, e.size, e.mtime, hash)
if err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
// Remaining entries in existingByPath are removed files
removed = int64(len(existingByPath))
for path := range existingByPath {
log.Verbosef("D %s", path)
}
return nil
}
scanDuration := time.Since(startScan)
log.Infof("scan complete in %s: %d unchanged, %d changed, %d added, %d removed",
scanDuration.Round(time.Millisecond), unchanged, changed, added, removed)
// writeFreshenedManifest writes the manifest atomically (write to a
// temp file, then rename over the target).
func writeFreshenedManifest(
ctx context.Context, afs afero.Fs, builder *mfer.Builder, manifestPath string,
) error {
tmpPath := manifestPath + ".tmp"
outFile, err := afs.Create(tmpPath)
if err != nil {
return fmt.Errorf("failed to create temp file: %w", err)
// Calculate total bytes to hash
var totalHashBytes int64
var filesToHash int64
for _, e := range entries {
if e.needsHash {
totalHashBytes += e.size
filesToHash++
}
}
err = builder.Build(ctx, outFile)
_ = outFile.Close()
if err != nil {
_ = afs.Remove(tmpPath)
return fmt.Errorf("failed to write manifest: %w", err)
// Phase 2: Hash changed and new files
if filesToHash > 0 {
log.Infof("hashing %d files (%s)...", filesToHash, humanize.IBytes(uint64(totalHashBytes)))
}
// Rename temp to final
err = afs.Rename(tmpPath, manifestPath)
if err != nil {
_ = afs.Remove(tmpPath)
startHash := time.Now()
var hashedFiles int64
var hashedBytes int64
return fmt.Errorf("failed to rename manifest: %w", err)
}
return nil
}
// newFreshenBuilder constructs the manifest builder configured from CLI
// flags.
func newFreshenBuilder(ctx *cli.Context) *mfer.Builder {
builder := mfer.NewBuilder()
if ctx.Bool("include-timestamps") {
builder.SetIncludeTimestamps(true)
@@ -350,77 +238,6 @@ func newFreshenBuilder(ctx *cli.Context) *mfer.Builder {
log.Infof("signing manifest with GPG key: %s", signKey)
}
return builder
}
// freshenScan runs the scan phase against the loaded manifest entries
// and returns the populated scanner and the count of removed files.
func (mfa *CLIApp) freshenScan(
ctx *cli.Context, manifestPath, absBase string,
existingByPath map[string]*mfer.MFFilePath,
) (*freshenScanner, int64, error) {
log.Infof("scanning filesystem...")
startScan := time.Now()
showProgress := ctx.Bool("progress")
scanner := &freshenScanner{
fs: mfa.Fs,
absBase: absBase,
manifestBase: filepath.Base(manifestPath),
includeDotfiles: ctx.Bool("include-dotfiles"),
followSymlinks: ctx.Bool("follow-symlinks"),
showProgress: showProgress,
existingByPath: existingByPath,
}
err := afero.Walk(mfa.Fs, absBase, scanner.walk)
if showProgress {
log.ProgressDone()
}
if err != nil {
return nil, 0, fmt.Errorf("failed to scan filesystem: %w", err)
}
// Remaining entries in existingByPath are removed files
removed := int64(len(existingByPath))
for path := range existingByPath {
log.Verbosef("D %s", path)
}
scanDuration := time.Since(startScan)
log.Infof("scan complete in %s: %d unchanged, %d changed, %d added, %d removed",
scanDuration.Round(time.Millisecond), scanner.unchanged, scanner.changed,
scanner.added, removed)
return scanner, removed, nil
}
// hashTotals returns the total byte count and file count of entries
// that need hashing.
func hashTotals(entries []*freshenEntry) (int64, int64) {
var (
totalHashBytes int64
filesToHash int64
)
for _, e := range entries {
if e.needsHash {
totalHashBytes += e.size
filesToHash++
}
}
return totalHashBytes, filesToHash
}
// runFreshenHash processes every entry through the hasher, aborting if
// the context is canceled.
func runFreshenHash(
ctx *cli.Context, hasher *freshenHasher, entries []*freshenEntry,
) error {
for _, e := range entries {
select {
case <-ctx.Done():
@@ -428,154 +245,122 @@ func runFreshenHash(
default:
}
err := hasher.processEntry(e)
if err != nil {
return err
if e.needsHash {
// Need to read and hash the file
absPath := filepath.Join(absBase, e.path)
f, err := mfa.Fs.Open(absPath)
if err != nil {
return fmt.Errorf("failed to open %s: %w", e.path, err)
}
hash, bytesRead, err := hashFile(f, e.size, func(n int64) {
if showProgress {
currentBytes := hashedBytes + n
elapsed := time.Since(startHash)
var rate float64
var eta time.Duration
if elapsed > 0 && currentBytes > 0 {
rate = float64(currentBytes) / elapsed.Seconds()
remaining := totalHashBytes - currentBytes
if rate > 0 {
eta = time.Duration(float64(remaining)/rate) * time.Second
}
}
if eta > 0 {
log.Progressf("Hashing: %d/%d files, %s/s, ETA %s",
hashedFiles, filesToHash, humanize.IBytes(uint64(rate)), eta.Round(time.Second))
} else {
log.Progressf("Hashing: %d/%d files, %s/s",
hashedFiles, filesToHash, humanize.IBytes(uint64(rate)))
}
}
})
_ = f.Close()
if err != nil {
return fmt.Errorf("failed to hash %s: %w", e.path, err)
}
hashedBytes += bytesRead
hashedFiles++
// Add to builder with computed hash
if err := addFileToBuilder(builder, e.path, e.size, e.mtime, hash); err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
}
} else {
// Use existing entry
if err := addExistingToBuilder(builder, e.existing); err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
}
}
}
return nil
}
// loadExistingEntries loads the manifest and indexes its file entries
// by path.
func (mfa *CLIApp) loadExistingEntries(
manifestPath string,
) (map[string]*mfer.MFFilePath, error) {
log.Infof("loading manifest from %s", manifestPath)
// Load existing manifest
manifest, err := mfer.NewManifestFromFile(mfa.Fs, manifestPath)
if err != nil {
return nil, fmt.Errorf("failed to load manifest: %w", err)
}
existingFiles := manifest.Files()
log.Infof("manifest contains %d files", len(existingFiles))
// Build map of existing entries by path
existingByPath := make(map[string]*mfer.MFFilePath, len(existingFiles))
for _, f := range existingFiles {
existingByPath[f.GetPath()] = f
}
return existingByPath, nil
}
func (mfa *CLIApp) freshenManifestOperation(ctx *cli.Context) error {
log.Debug("freshenManifestOperation()")
basePath := ctx.String("base")
showProgress := ctx.Bool("progress")
// Find manifest file
manifestPath, err := mfa.resolveFreshenManifestPath(ctx)
if err != nil {
return fmt.Errorf("freshen: %w", err)
}
existingByPath, err := mfa.loadExistingEntries(manifestPath)
if err != nil {
return err
}
absBase, err := filepath.Abs(basePath)
if err != nil {
return fmt.Errorf("freshen: invalid base path: %w", err)
}
// Phase 1: Scan filesystem
scanner, removed, err := mfa.freshenScan(ctx, manifestPath, absBase,
existingByPath)
if err != nil {
return err
}
// Calculate total bytes to hash
totalHashBytes, filesToHash := hashTotals(scanner.entries)
// Phase 2: Hash changed and new files
if filesToHash > 0 {
log.Infof("hashing %d files (%s)...", filesToHash,
humanize.IBytes(safeUint64(totalHashBytes)))
}
hasher := &freshenHasher{
fs: mfa.Fs,
absBase: absBase,
showProgress: showProgress,
totalHashBytes: totalHashBytes,
filesToHash: filesToHash,
startHash: time.Now(),
builder: newFreshenBuilder(ctx),
}
err = runFreshenHash(ctx, hasher, scanner.entries)
if err != nil {
return err
}
if showProgress && filesToHash > 0 {
log.ProgressDone()
}
// Print summary
log.Infof("freshen complete: %d unchanged, %d changed, %d added, %d removed",
scanner.unchanged, scanner.changed, scanner.added, removed)
unchanged, changed, added, removed)
// Skip writing if nothing changed
if scanner.changed == 0 && scanner.added == 0 && removed == 0 {
if changed == 0 && added == 0 && removed == 0 {
log.Infof("manifest unchanged, skipping write")
return nil
}
// Write updated manifest atomically (write to temp, then rename)
err = writeFreshenedManifest(ctx.Context, mfa.Fs, hasher.builder, manifestPath)
tmpPath := manifestPath + ".tmp"
outFile, err := mfa.Fs.Create(tmpPath)
if err != nil {
return err
return fmt.Errorf("failed to create temp file: %w", err)
}
err = builder.Build(outFile)
_ = outFile.Close()
if err != nil {
_ = mfa.Fs.Remove(tmpPath)
return fmt.Errorf("failed to write manifest: %w", err)
}
// Rename temp to final
if err := mfa.Fs.Rename(tmpPath, manifestPath); err != nil {
_ = mfa.Fs.Remove(tmpPath)
return fmt.Errorf("failed to rename manifest: %w", err)
}
totalDuration := time.Since(mfa.startupTime)
if hasher.hashedBytes > 0 {
hashDuration := time.Since(hasher.startHash)
hashRate := float64(hasher.hashedBytes) / hashDuration.Seconds()
if hashedBytes > 0 {
hashDuration := time.Since(startHash)
hashRate := float64(hashedBytes) / hashDuration.Seconds()
log.Infof("hashed %s in %.1fs (%s/s)",
humanize.IBytes(safeUint64(hasher.hashedBytes)),
totalDuration.Seconds(), humanize.IBytes(safeRateUint64(hashRate)))
humanize.IBytes(uint64(hashedBytes)), totalDuration.Seconds(), humanize.IBytes(uint64(hashRate)))
}
log.Infof("wrote %d files to %s", len(scanner.entries), manifestPath)
log.Infof("wrote %d files to %s", len(entries), manifestPath)
return nil
}
// hashFile reads a file and computes its SHA256 multihash.
// Progress callback is called with bytes read so far.
func hashFile(r io.Reader, progress func(int64)) ([]byte, int64, error) {
func hashFile(r io.Reader, size int64, progress func(int64)) ([]byte, int64, error) {
h := sha256.New()
buf := make([]byte, hashBufSize)
buf := make([]byte, 64*1024)
var total int64
for {
n, err := r.Read(buf)
if n > 0 {
h.Write(buf[:n])
total += int64(n)
if progress != nil {
progress(total)
}
}
if err == io.EOF {
break
}
// Returned unwrapped: the caller renders this as
// "failed to hash <path>: <err>" and adding a second layer here
// would change that message.
if err != nil {
return nil, total, err
}
@@ -590,29 +375,15 @@ func hashFile(r io.Reader, progress func(int64)) ([]byte, int64, error) {
}
// addFileToBuilder adds a new file entry to the builder
func addFileToBuilder(
b *mfer.Builder, path string, size int64, mtime time.Time, hash []byte,
) error {
return b.AddFileWithHash(
mfer.RelFilePath(path), mfer.FileSize(size), mfer.ModTime(mtime), hash)
func addFileToBuilder(b *mfer.Builder, path string, size int64, mtime time.Time, hash []byte) error {
return b.AddFileWithHash(mfer.RelFilePath(path), mfer.FileSize(size), mfer.ModTime(mtime), hash)
}
// addExistingToBuilder adds an existing manifest entry to the builder.
//
// Entries reach this path only when recordEntry classified them as
// unchanged, which requires a recorded mtime, so an absent mtime here is
// an error rather than something to paper over with the Unix epoch.
// addExistingToBuilder adds an existing manifest entry to the builder
func addExistingToBuilder(b *mfer.Builder, entry *mfer.MFFilePath) error {
mtime, ok := entryMtime(entry)
if !ok {
return fmt.Errorf("%w: %s", errEntryMissingMtime, entry.GetPath())
}
if len(entry.GetHashes()) == 0 {
mtime := time.Unix(entry.Mtime.Seconds, int64(entry.Mtime.Nanos))
if len(entry.Hashes) == 0 {
return nil
}
return b.AddFileWithHash(mfer.RelFilePath(entry.GetPath()),
mfer.FileSize(entry.GetSize()), mfer.ModTime(mtime),
entry.GetHashes()[0].GetMultiHash())
return b.AddFileWithHash(mfer.RelFilePath(entry.Path), mfer.FileSize(entry.Size), mfer.ModTime(mtime), entry.Hashes[0].MultiHash)
}
+28 -146
View File
@@ -1,12 +1,9 @@
//nolint:testpackage // white-box tests exercise unexported internals
package cli
import (
"bytes"
"context"
"os"
"testing"
"time"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
@@ -14,48 +11,24 @@ import (
"sneak.berlin/go/mfer/mfer"
)
// stubFileInfo is a minimal fs.FileInfo for exercising recordEntry
// without touching a filesystem.
type stubFileInfo struct {
size int64
mtime time.Time
}
func TestFreshenUnchanged(t *testing.T) {
// Create filesystem with test files
fs := afero.NewMemMapFs()
func (s stubFileInfo) Name() string { return "stub" }
func (s stubFileInfo) Size() int64 { return s.size }
func (s stubFileInfo) Mode() os.FileMode { return 0 }
func (s stubFileInfo) ModTime() time.Time { return s.mtime }
func (s stubFileInfo) IsDir() bool { return false }
func (s stubFileInfo) Sys() any { return nil }
// setupFreshenDir populates /testdir with two files, scans it, and
// writes the resulting manifest to /testdir/.index.mf.
func setupFreshenDir(t *testing.T, fs afero.Fs) {
t.Helper()
require.NoError(t, fs.MkdirAll(testDir, 0o755))
writeTestFile(t, fs, testFile1, "content1")
writeTestFile(t, fs, "/testdir/file2.txt", "content2")
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644))
// Generate initial manifest
opts := &mfer.ScannerOptions{Fs: fs}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumeratePath(testDir, nil))
require.NoError(t, s.EnumeratePath("/testdir", nil))
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
// Write manifest to filesystem
require.NoError(t,
afero.WriteFile(fs, "/testdir/.index.mf", manifestBuf.Bytes(), 0o644))
}
func TestFreshenUnchanged(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
setupFreshenDir(t, fs)
require.NoError(t, afero.WriteFile(fs, "/testdir/.index.mf", manifestBuf.Bytes(), 0o644))
// Parse manifest to verify
manifest, err := mfer.NewManifestFromFile(fs, "/testdir/.index.mf")
@@ -64,10 +37,23 @@ func TestFreshenUnchanged(t *testing.T) {
}
func TestFreshenWithChanges(t *testing.T) {
t.Parallel()
// Create filesystem with test files
fs := afero.NewMemMapFs()
setupFreshenDir(t, fs)
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644))
// Generate initial manifest
opts := &mfer.ScannerOptions{Fs: fs}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumeratePath("/testdir", nil))
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
// Write manifest to filesystem
require.NoError(t, afero.WriteFile(fs, "/testdir/.index.mf", manifestBuf.Bytes(), 0o644))
// Verify initial manifest has 2 files
manifest, err := mfer.NewManifestFromFile(fs, "/testdir/.index.mf")
@@ -75,17 +61,17 @@ func TestFreshenWithChanges(t *testing.T) {
assert.Len(t, manifest.Files(), 2)
// Add a new file
writeTestFile(t, fs, "/testdir/file3.txt", "content3")
require.NoError(t, afero.WriteFile(fs, "/testdir/file3.txt", []byte("content3"), 0o644))
// Modify file2 (change content and size)
writeTestFile(t, fs, "/testdir/file2.txt", "modified content2")
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("modified content2"), 0o644))
// Remove file1
require.NoError(t, fs.Remove(testFile1))
require.NoError(t, fs.Remove("/testdir/file1.txt"))
// Note: The freshen operation would need to be run here
// For now, we just verify the test setup is correct
exists, _ := afero.Exists(fs, testFile1)
exists, _ := afero.Exists(fs, "/testdir/file1.txt")
assert.False(t, exists)
exists, _ = afero.Exists(fs, "/testdir/file3.txt")
@@ -94,107 +80,3 @@ func TestFreshenWithChanges(t *testing.T) {
content, _ := afero.ReadFile(fs, "/testdir/file2.txt")
assert.Equal(t, "modified content2", string(content))
}
// TestFreshenRecordEntryMtimePresence pins the behavior of recordEntry
// with respect to MFFilePath.Mtime, which is a message pointer with
// proto3 field presence and may legitimately be absent.
//
// An absent mtime must never be read as time.Unix(0, 0): that value
// never equals a real modification time, so every entry would be
// classified as changed, re-hashed, and the manifest rewritten
// unconditionally - the exact inverse of what freshen is for, and
// silent. An entry with no mtime is therefore "changed" because it
// cannot be compared, not because it looks like it dates from 1970.
func TestFreshenRecordEntryMtimePresence(t *testing.T) {
t.Parallel()
const relPath = "file1.txt"
// The scanned file's mtime is the Unix epoch. If recordEntry ever misreads
// an absent manifest mtime as the epoch, the "absent" case below would
// compare equal to this and be classified unchanged, so the test fails.
mtime := time.Unix(0, 0)
info := stubFileInfo{size: 8, mtime: mtime}
for _, tc := range []struct {
name string
entry *mfer.MFFilePath
needsHash bool
changed int64
unchanged int64
}{
{
name: "matching mtime and size is unchanged",
entry: &mfer.MFFilePath{
Path: relPath,
Size: 8,
Mtime: &mfer.Timestamp{Seconds: mtime.Unix()},
},
needsHash: false,
changed: 0,
unchanged: 1,
},
{
name: "absent mtime is changed, not epoch",
entry: &mfer.MFFilePath{
Path: relPath,
Size: 8,
Mtime: nil,
},
needsHash: true,
changed: 1,
unchanged: 0,
},
} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
s := &freshenScanner{
existingByPath: map[string]*mfer.MFFilePath{relPath: tc.entry},
}
s.recordEntry(relPath, info)
require.Len(t, s.entries, 1)
assert.Equal(t, tc.needsHash, s.entries[0].needsHash)
assert.Equal(t, tc.changed, s.changed)
assert.Equal(t, tc.unchanged, s.unchanged)
assert.Zero(t, s.added)
})
}
}
// TestFreshenAddExistingRejectsMissingMtime pins that an entry with no
// mtime is never carried forward into a rebuilt manifest with a
// fabricated epoch timestamp.
func TestFreshenAddExistingRejectsMissingMtime(t *testing.T) {
t.Parallel()
b := mfer.NewBuilder()
entry := &mfer.MFFilePath{
Path: "file1.txt",
Size: 8,
Mtime: nil,
Hashes: []*mfer.MFFileChecksum{
{MultiHash: []byte{0x12, 0x20}},
},
}
err := addExistingToBuilder(b, entry)
require.ErrorIs(t, err, errEntryMissingMtime)
assert.Contains(t, err.Error(), "file1.txt")
}
// TestEntryMtime pins the presence semantics the callers depend on.
func TestEntryMtime(t *testing.T) {
t.Parallel()
got, ok := entryMtime(&mfer.MFFilePath{Mtime: nil})
assert.False(t, ok)
assert.True(t, got.IsZero())
got, ok = entryMtime(&mfer.MFFilePath{
Mtime: &mfer.Timestamp{Seconds: 1_700_000_000, Nanos: 500},
})
assert.True(t, ok)
assert.Equal(t, time.Unix(1_700_000_000, 500), got)
}
+87 -191
View File
@@ -1,7 +1,6 @@
package cli
import (
"errors"
"fmt"
"os"
"os/signal"
@@ -17,78 +16,9 @@ import (
"sneak.berlin/go/mfer/mfer"
)
var (
// errPathNotExist indicates an input path that does not exist.
errPathNotExist = errors.New("path does not exist")
// errOutputExists indicates the output file already exists and
// --force was not given. It is wrapped mid-sentence so that the
// rendered message stays exactly as mfer has always printed it.
errOutputExists = errors.New(
"already exists (use --force to overwrite)")
)
func (mfa *CLIApp) generateManifestOperation(ctx *cli.Context) error {
log.Debug("generateManifestOperation()")
// reportEnumProgress renders enumeration progress until the channel
// closes.
func reportEnumProgress(progress <-chan mfer.EnumerateStatus, wg *sync.WaitGroup) {
defer wg.Done()
for status := range progress {
log.Progressf("Enumerating: %d files, %s",
status.FilesFound,
humanize.IBytes(safeUint64(int64(status.BytesFound))))
}
log.ProgressDone()
}
// reportScanProgress renders scan progress until the channel closes.
func reportScanProgress(progress <-chan mfer.ScanStatus, wg *sync.WaitGroup) {
defer wg.Done()
for status := range progress {
if status.ETA > 0 {
log.Progressf("Scanning: %d/%d files, %s/s, ETA %s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)),
status.ETA.Round(time.Second))
} else {
log.Progressf("Scanning: %d/%d files, %s/s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)))
}
}
log.ProgressDone()
}
// collectInputPaths validates the input path arguments and returns them
// as absolute paths.
func (mfa *CLIApp) collectInputPaths(args cli.Args) ([]string, error) {
paths := make([]string, 0, args.Len())
for i := range args.Len() {
inputPath := args.Get(i)
ap, err := filepath.Abs(inputPath)
if err != nil {
return nil, fmt.Errorf("generate: invalid path %q: %w", inputPath, err)
}
// Validate path exists before adding to list
if exists, _ := afero.Exists(mfa.Fs, ap); !exists {
return nil, fmt.Errorf("%w: %s", errPathNotExist, inputPath)
}
log.Debugf("enumerating path: %s", ap)
paths = append(paths, ap)
}
return paths, nil
}
// buildScannerOptions constructs scanner options from the CLI flags.
func (mfa *CLIApp) buildScannerOptions(ctx *cli.Context) *mfer.ScannerOptions {
opts := &mfer.ScannerOptions{
IncludeDotfiles: ctx.Bool("include-dotfiles"),
FollowSymLinks: ctx.Bool("follow-symlinks"),
@@ -99,7 +29,6 @@ func (mfa *CLIApp) buildScannerOptions(ctx *cli.Context) *mfer.ScannerOptions {
// Set seed for deterministic UUID if provided
if seed := ctx.String("seed"); seed != "" {
opts.Seed = seed
log.Infof("using deterministic seed for manifest UUID")
}
@@ -111,167 +40,136 @@ func (mfa *CLIApp) buildScannerOptions(ctx *cli.Context) *mfer.ScannerOptions {
log.Infof("signing manifest with GPG key: %s", signKey)
}
return opts
}
// enumerateInputs runs the enumeration phase over the argument paths,
// or the current directory when no arguments are given.
func (mfa *CLIApp) enumerateInputs(
s *mfer.Scanner, args cli.Args, enumProgress chan mfer.EnumerateStatus,
) error {
if args.Len() == 0 {
// Default to current directory
err := s.EnumeratePath(".", enumProgress)
if err != nil {
return fmt.Errorf(
"generate: failed to enumerate current directory: %w", err)
}
return nil
}
// Collect and validate all paths first
paths, err := mfa.collectInputPaths(args)
if err != nil {
return err
}
err = s.EnumeratePaths(enumProgress, paths...)
if err != nil {
return fmt.Errorf("generate: failed to enumerate paths: %w", err)
}
return nil
}
// cleanupOnSignal installs a handler that removes the temp output file
// and exits when the process is interrupted. It returns the signal
// channel so the caller can stop and close it when done.
func (mfa *CLIApp) cleanupOnSignal(outFile afero.File, tmpPath string) chan os.Signal {
sigChan := make(chan os.Signal, 1)
signal.Notify(sigChan, os.Interrupt, syscall.SIGTERM)
go func() {
sig, ok := <-sigChan
if !ok || sig == nil {
return // Channel closed normally, not a signal
}
_ = outFile.Close()
_ = mfa.Fs.Remove(tmpPath)
os.Exit(1)
}()
return sigChan
}
// runEnumeratePhase enumerates all input paths with optional progress
// reporting and logs the totals.
func (mfa *CLIApp) runEnumeratePhase(ctx *cli.Context, s *mfer.Scanner) error {
// Set up enumeration progress reporting
var (
enumProgress chan mfer.EnumerateStatus
enumWg sync.WaitGroup
)
if ctx.Bool("progress") {
enumProgress = make(chan mfer.EnumerateStatus, 1)
enumWg.Add(1)
go reportEnumProgress(enumProgress, &enumWg)
}
err := mfa.enumerateInputs(s, ctx.Args(), enumProgress)
if err != nil {
return err
}
enumWg.Wait()
log.Infof("enumerated %d files, %s total", s.FileCount(),
humanize.IBytes(safeUint64(int64(s.TotalBytes()))))
return nil
}
func (mfa *CLIApp) generateManifestOperation(ctx *cli.Context) error {
log.Debug("generateManifestOperation()")
s := mfer.NewScannerWithOptions(mfa.buildScannerOptions(ctx))
s := mfer.NewScannerWithOptions(opts)
// Phase 1: Enumeration - collect paths and stat files
err := mfa.runEnumeratePhase(ctx, s)
if err != nil {
return err
args := ctx.Args()
showProgress := ctx.Bool("progress")
// Set up enumeration progress reporting
var enumProgress chan mfer.EnumerateStatus
var enumWg sync.WaitGroup
if showProgress {
enumProgress = make(chan mfer.EnumerateStatus, 1)
enumWg.Add(1)
go func() {
defer enumWg.Done()
for status := range enumProgress {
log.Progressf("Enumerating: %d files, %s",
status.FilesFound,
humanize.IBytes(uint64(status.BytesFound)))
}
log.ProgressDone()
}()
}
showProgress := ctx.Bool("progress")
if args.Len() == 0 {
// Default to current directory
if err := s.EnumeratePath(".", enumProgress); err != nil {
return fmt.Errorf("generate: failed to enumerate current directory: %w", err)
}
} else {
// Collect and validate all paths first
paths := make([]string, 0, args.Len())
for i := 0; i < args.Len(); i++ {
inputPath := args.Get(i)
ap, err := filepath.Abs(inputPath)
if err != nil {
return fmt.Errorf("generate: invalid path %q: %w", inputPath, err)
}
// Validate path exists before adding to list
if exists, _ := afero.Exists(mfa.Fs, ap); !exists {
return fmt.Errorf("path does not exist: %s", inputPath)
}
log.Debugf("enumerating path: %s", ap)
paths = append(paths, ap)
}
if err := s.EnumeratePaths(enumProgress, paths...); err != nil {
return fmt.Errorf("generate: failed to enumerate paths: %w", err)
}
}
enumWg.Wait()
log.Infof("enumerated %d files, %s total", s.FileCount(), humanize.IBytes(uint64(s.TotalBytes())))
// Check if output file exists
outputPath := ctx.String("output")
if exists, _ := afero.Exists(mfa.Fs, outputPath); exists && !ctx.Bool("force") {
return fmt.Errorf("output file %s %w", outputPath, errOutputExists)
if exists, _ := afero.Exists(mfa.Fs, outputPath); exists {
if !ctx.Bool("force") {
return fmt.Errorf("output file %s already exists (use --force to overwrite)", outputPath)
}
}
// Create temp file for atomic write
tmpPath := outputPath + ".tmp"
outFile, err := mfa.Fs.Create(tmpPath)
if err != nil {
return fmt.Errorf("failed to create temp file: %w", err)
}
// Set up signal handler to clean up temp file on Ctrl-C
sigChan := mfa.cleanupOnSignal(outFile, tmpPath)
sigChan := make(chan os.Signal, 1)
signal.Notify(sigChan, os.Interrupt, syscall.SIGTERM)
go func() {
sig, ok := <-sigChan
if !ok || sig == nil {
return // Channel closed normally, not a signal
}
_ = outFile.Close()
_ = mfa.Fs.Remove(tmpPath)
os.Exit(1)
}()
// Clean up temp file on any error or interruption
success := false
defer func() {
signal.Stop(sigChan)
close(sigChan)
_ = outFile.Close()
if !success {
_ = mfa.Fs.Remove(tmpPath)
}
}()
// Phase 2: Scan - read file contents and generate manifest
var (
scanProgress chan mfer.ScanStatus
scanWg sync.WaitGroup
)
var scanProgress chan mfer.ScanStatus
var scanWg sync.WaitGroup
if showProgress {
scanProgress = make(chan mfer.ScanStatus, 1)
scanWg.Add(1)
go reportScanProgress(scanProgress, &scanWg)
go func() {
defer scanWg.Done()
for status := range scanProgress {
if status.ETA > 0 {
log.Progressf("Scanning: %d/%d files, %s/s, ETA %s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)),
status.ETA.Round(time.Second))
} else {
log.Progressf("Scanning: %d/%d files, %s/s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)))
}
}
log.ProgressDone()
}()
}
err = s.ToManifest(ctx.Context, outFile, scanProgress)
scanWg.Wait()
if err != nil {
return fmt.Errorf("failed to generate manifest: %w", err)
}
// Close file before rename to ensure all data is flushed
err = outFile.Close()
if err != nil {
if err := outFile.Close(); err != nil {
return fmt.Errorf("failed to close temp file: %w", err)
}
// Atomic rename
err = mfa.Fs.Rename(tmpPath, outputPath)
if err != nil {
if err := mfa.Fs.Rename(tmpPath, outputPath); err != nil {
return fmt.Errorf("failed to rename temp file: %w", err)
}
@@ -279,9 +177,7 @@ func (mfa *CLIApp) generateManifestOperation(ctx *cli.Context) error {
elapsed := time.Since(mfa.startupTime).Seconds()
rate := float64(s.TotalBytes()) / elapsed
log.Infof("wrote %d files (%s) to %s in %.1fs (%s/s)", s.FileCount(),
humanize.IBytes(safeUint64(int64(s.TotalBytes()))), outputPath, elapsed,
humanize.IBytes(safeRateUint64(rate)))
log.Infof("wrote %d files (%s) to %s in %.1fs (%s/s)", s.FileCount(), humanize.IBytes(uint64(s.TotalBytes())), outputPath, elapsed, humanize.IBytes(uint64(rate)))
return nil
}
+3 -11
View File
@@ -25,7 +25,6 @@ func (mfa *CLIApp) listManifestOperation(ctx *cli.Context) error {
if err != nil {
return fmt.Errorf("list: %w", err)
}
defer func() { _ = rc.Close() }()
manifest, err := mfer.NewManifestFromReader(rc)
@@ -43,17 +42,10 @@ func (mfa *CLIApp) listManifestOperation(ctx *cli.Context) error {
for _, f := range files {
if longFormat {
// An entry may legitimately carry no mtime; render that as
// mtimeAbsent rather than as the Unix epoch.
mtimeStr := mtimeAbsent
if mtime, ok := entryMtime(f); ok {
mtimeStr = mtime.Format(time.RFC3339)
}
_, _ = fmt.Fprintf(mfa.Stdout, "%d\t%s\t%s%s",
f.GetSize(), mtimeStr, f.GetPath(), lineEnd)
mtime := time.Unix(f.Mtime.Seconds, int64(f.Mtime.Nanos))
_, _ = fmt.Fprintf(mfa.Stdout, "%d\t%s\t%s%s", f.Size, mtime.Format(time.RFC3339), f.Path, lineEnd)
} else {
_, _ = fmt.Fprintf(mfa.Stdout, "%s%s", f.GetPath(), lineEnd)
_, _ = fmt.Fprintf(mfa.Stdout, "%s%s", f.Path, lineEnd)
}
}
+3 -33
View File
@@ -1,8 +1,6 @@
package cli
import (
"context"
"errors"
"fmt"
"io"
"net/http"
@@ -12,17 +10,6 @@ import (
"github.com/urfave/cli/v2"
)
// manifestFetchTimeout bounds HTTP requests made to fetch a manifest.
const manifestFetchTimeout = 30 * time.Second
// errHTTPStatus indicates an HTTP response with a non-OK status code.
//
// Its text is the literal "HTTP" prefix of the rendered "HTTP <code>"
// message that mfer has always printed, so that wrapping it does not
// change any user-visible output. Match it with errors.Is; do not read
// its message.
var errHTTPStatus = errors.New("HTTP")
// isHTTPURL returns true if the string starts with http:// or https://.
func isHTTPURL(s string) bool {
return strings.HasPrefix(s, "http://") || strings.HasPrefix(s, "https://")
@@ -32,35 +19,21 @@ func isHTTPURL(s string) bool {
// The caller must close the returned reader.
func (mfa *CLIApp) openManifestReader(pathOrURL string) (io.ReadCloser, error) {
if isHTTPURL(pathOrURL) {
client := &http.Client{Timeout: manifestFetchTimeout}
req, err := http.NewRequestWithContext(
context.Background(), http.MethodGet, pathOrURL, nil,
)
client := &http.Client{Timeout: 30 * time.Second}
resp, err := client.Get(pathOrURL) //nolint:gosec // user-provided URL is intentional
if err != nil {
return nil, fmt.Errorf("failed to fetch %s: %w", pathOrURL, err)
}
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("failed to fetch %s: %w", pathOrURL, err)
}
if resp.StatusCode != http.StatusOK {
_ = resp.Body.Close()
return nil, fmt.Errorf("failed to fetch %s: %w %d",
pathOrURL, errHTTPStatus, resp.StatusCode)
return nil, fmt.Errorf("failed to fetch %s: HTTP %d", pathOrURL, resp.StatusCode)
}
return resp.Body, nil
}
f, err := mfa.Fs.Open(pathOrURL)
if err != nil {
return nil, err
}
return f, nil
}
@@ -73,14 +46,11 @@ func (mfa *CLIApp) resolveManifestArg(ctx *cli.Context) (string, error) {
if isHTTPURL(arg) {
return arg, nil
}
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
return findManifest(mfa.Fs, arg)
}
return arg, nil
}
return findManifest(mfa.Fs, ".")
}
+193 -295
View File
@@ -1,7 +1,6 @@
package cli
import (
"errors"
"fmt"
"io"
"os"
@@ -13,26 +12,8 @@ import (
"sneak.berlin/go/mfer/mfer"
)
// Command and flag names shared across command definitions and tests.
const (
cmdGenerate = "generate"
cmdCheck = "check"
cmdExport = "export"
cmdFetch = "fetch"
cmdVersion = "version"
flagProgress = "progress"
manifestArgsUsage = "[manifest file]"
)
// errUnknownCommand indicates an unrecognized command argument.
var errUnknownCommand = errors.New("unknown command")
// CLIApp is the main CLI application container. It holds configuration,
// I/O streams, and filesystem abstraction to enable testing and flexibility.
//
//nolint:revive // established name used throughout the codebase and tests
type CLIApp struct {
appname string
version string
@@ -60,59 +41,34 @@ const banner = `
\ \:\ \ \:\ \ \::/ \ \:\
\__\/ \__\/ \__\/ \__\/`
func (mfa *CLIApp) printBanner() {
if log.GetLevel() <= log.InfoLevel {
_, _ = fmt.Fprintln(mfa.Stdout, banner)
_, _ = fmt.Fprintf(mfa.Stdout, " mfer by @sneak: v%s released %s\n", mfer.Version, mfer.ReleaseDate)
_, _ = fmt.Fprintln(mfa.Stdout, " https://sneak.berlin/go/mfer")
}
}
// VersionString returns the version and git revision formatted for display.
func (mfa *CLIApp) VersionString() string {
if mfa.gitrev != "" {
return fmt.Sprintf("%s (%s)", mfer.Version, mfa.gitrev)
}
return mfer.Version
}
// printVersion writes the version line shared by the --version flag and the
// version subcommand, so both produce identical output.
func (mfa *CLIApp) printVersion() {
_, _ = fmt.Fprintf(mfa.Stdout, "%s version %s\n", mfa.appname, mfa.VersionString())
}
func (mfa *CLIApp) printBanner() {
if log.GetLevel() <= log.InfoLevel {
_, _ = fmt.Fprintln(mfa.Stdout, banner)
_, _ = fmt.Fprintf(mfa.Stdout,
" mfer by @sneak: v%s released %s\n",
mfer.Version, mfer.ReleaseDate)
_, _ = fmt.Fprintln(mfa.Stdout, " https://sneak.berlin/go/mfer")
}
}
// setVerbosity sets the log level from -v and -q, given before the subcommand
// name (the root's copies), after it (the subcommand's own copies), or both.
// urfave/cli reads a flag from the nearest command that defines it, so each
// command in the lineage is asked. The highest -v count wins rather than the
// sum, because a subcommand without its own copies reads the root's.
func (mfa *CLIApp) setVerbosity(c *cli.Context) {
_, present := os.LookupEnv("MFER_DEBUG")
verbosity := 0
quiet := false
for _, ctx := range c.Lineage() {
verbosity = max(verbosity, ctx.Count("verbose"))
quiet = quiet || ctx.Bool("quiet")
}
switch {
case present:
if present {
log.EnableDebugLogging()
case quiet:
} else if c.Bool("quiet") {
log.SetLevel(log.ErrorLevel)
default:
log.SetLevelFromVerbosity(verbosity)
} else {
log.SetLevelFromVerbosity(c.Count("verbose"))
}
}
// commonFlags returns the -v and -q flags taken by the root and by the
// generate, check, freshen and fetch subcommands.
// commonFlags returns the flags shared by most commands (-v, -q)
func commonFlags() []cli.Flag {
return []cli.Flag{
&cli.BoolFlag{
@@ -129,217 +85,10 @@ func commonFlags() []cli.Flag {
}
}
func (mfa *CLIApp) generateCommand() *cli.Command {
return &cli.Command{
Name: cmdGenerate,
Aliases: []string{"gen"},
Usage: "Generate manifest file",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.generateManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.BoolFlag{
Name: "follow-symlinks",
Aliases: []string{"L"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "include-dotfiles",
Aliases: []string{"IncludeDotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.StringFlag{
Name: "output",
Value: "./.index.mf",
Aliases: []string{"o"},
Usage: "Specify output filename",
},
&cli.BoolFlag{
Name: "force",
Aliases: []string{"f"},
Usage: "Overwrite output file if it exists",
},
&cli.BoolFlag{
Name: flagProgress,
Aliases: []string{"P"},
Usage: "Show progress during enumeration and scanning",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
EnvVars: []string{"MFER_SIGN_KEY"},
},
&cli.StringFlag{
Name: "seed",
Usage: "Seed value for deterministic manifest UUID",
EnvVars: []string{"MFER_SEED"},
},
&cli.BoolFlag{
Name: "include-timestamps",
Usage: "Include createdAt timestamp in manifest " +
"(omitted by default for determinism)",
},
),
}
}
func (mfa *CLIApp) checkCommand() *cli.Command {
return &cli.Command{
Name: cmdCheck,
Usage: "Validate files using manifest file",
ArgsUsage: manifestArgsUsage,
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.checkManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths from manifest",
},
&cli.BoolFlag{
Name: flagProgress,
Aliases: []string{"P"},
Usage: "Show progress during checking",
},
&cli.BoolFlag{
Name: "no-extra-files",
Usage: "Fail if files exist in base directory that are not in manifest",
},
&cli.StringFlag{
Name: "require-signature",
Aliases: []string{"S"},
Usage: "Require manifest to be signed by the specified GPG key ID",
EnvVars: []string{"MFER_REQUIRE_SIGNATURE"},
},
),
}
}
func (mfa *CLIApp) freshenCommand() *cli.Command {
return &cli.Command{
Name: "freshen",
Usage: "Update manifest with changed, new, and removed files",
ArgsUsage: manifestArgsUsage,
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.freshenManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths",
},
&cli.BoolFlag{
Name: "follow-symlinks",
Aliases: []string{"L"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "include-dotfiles",
Aliases: []string{"IncludeDotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.BoolFlag{
Name: flagProgress,
Aliases: []string{"P"},
Usage: "Show progress during scanning and hashing",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
EnvVars: []string{"MFER_SIGN_KEY"},
},
&cli.BoolFlag{
Name: "include-timestamps",
Usage: "Include createdAt timestamp in manifest " +
"(omitted by default for determinism)",
},
),
}
}
func (mfa *CLIApp) exportCommand() *cli.Command {
return &cli.Command{
Name: cmdExport,
Usage: "Export manifest contents as JSON",
ArgsUsage: "[manifest file or URL]",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
return mfa.exportManifestOperation(c)
},
}
}
func (mfa *CLIApp) versionCommand() *cli.Command {
return &cli.Command{
Name: cmdVersion,
Usage: "Show version",
Action: func(_ *cli.Context) error {
mfa.printVersion()
return nil
},
}
}
func (mfa *CLIApp) listCommand() *cli.Command {
return &cli.Command{
Name: "list",
Aliases: []string{"ls"},
Usage: "List files in manifest",
ArgsUsage: manifestArgsUsage,
Action: func(c *cli.Context) error {
return mfa.listManifestOperation(c)
},
Flags: []cli.Flag{
&cli.BoolFlag{
Name: "long",
Aliases: []string{"l"},
Usage: "Show size and mtime",
},
&cli.BoolFlag{
Name: "print0",
Usage: "Separate entries with NUL character (for xargs -0)",
},
},
}
}
func (mfa *CLIApp) fetchCommand() *cli.Command {
return &cli.Command{
Name: cmdFetch,
Usage: "fetch manifest and referenced files",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.fetchManifestOperation(c)
},
Flags: commonFlags(),
}
}
func (mfa *CLIApp) run(args []string) {
mfa.startupTime = time.Now()
if NoColor {
if NO_COLOR {
// shoutout to rob pike who thinks it's juvenile
log.DisableStyling()
}
@@ -348,21 +97,6 @@ func (mfa *CLIApp) run(args []string) {
log.SetOutput(mfa.Stdout, mfa.Stderr)
log.Init()
// -v means verbose, not version. urfave/cli's built-in version flag
// claims -v by default, which made "mfer -v --version" fail to parse and
// gave -v a different meaning at the root than on the generate, check,
// freshen and fetch subcommands, where it means verbose. Verbose is the
// more common meaning of -v in tools that offer both, so -v means verbose
// at the root too and the version flag takes the capital -V.
// VersionFlag and VersionPrinter are urfave/cli package globals; run() is
// serialized in tests, so assigning them here is safe.
cli.VersionFlag = &cli.BoolFlag{
Name: cmdVersion,
Aliases: []string{"V"},
Usage: "print the version",
}
cli.VersionPrinter = func(_ *cli.Context) { mfa.printVersion() }
mfa.app = &cli.App{
Name: mfa.appname,
Usage: "Manifest generator",
@@ -370,34 +104,198 @@ func (mfa *CLIApp) run(args []string) {
EnableBashCompletion: true,
Writer: mfa.Stdout,
ErrWriter: mfa.Stderr,
Flags: commonFlags(),
Action: func(c *cli.Context) error {
if c.Args().Len() > 0 {
return fmt.Errorf("%w %q", errUnknownCommand, c.Args().First())
return fmt.Errorf("unknown command %q", c.Args().First())
}
mfa.setVerbosity(c)
mfa.printBanner()
return cli.ShowAppHelp(c)
},
Commands: []*cli.Command{
mfa.generateCommand(),
mfa.checkCommand(),
mfa.freshenCommand(),
mfa.exportCommand(),
mfa.versionCommand(),
mfa.listCommand(),
mfa.fetchCommand(),
{
Name: "generate",
Aliases: []string{"gen"},
Usage: "Generate manifest file",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.generateManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.BoolFlag{
Name: "follow-symlinks",
Aliases: []string{"L"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "include-dotfiles",
Aliases: []string{"IncludeDotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.StringFlag{
Name: "output",
Value: "./.index.mf",
Aliases: []string{"o"},
Usage: "Specify output filename",
},
&cli.BoolFlag{
Name: "force",
Aliases: []string{"f"},
Usage: "Overwrite output file if it exists",
},
&cli.BoolFlag{
Name: "progress",
Aliases: []string{"P"},
Usage: "Show progress during enumeration and scanning",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
EnvVars: []string{"MFER_SIGN_KEY"},
},
&cli.StringFlag{
Name: "seed",
Usage: "Seed value for deterministic manifest UUID",
EnvVars: []string{"MFER_SEED"},
},
&cli.BoolFlag{
Name: "include-timestamps",
Usage: "Include createdAt timestamp in manifest (omitted by default for determinism)",
},
),
},
{
Name: "check",
Usage: "Validate files using manifest file",
ArgsUsage: "[manifest file]",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.checkManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths from manifest",
},
&cli.BoolFlag{
Name: "progress",
Aliases: []string{"P"},
Usage: "Show progress during checking",
},
&cli.BoolFlag{
Name: "no-extra-files",
Usage: "Fail if files exist in base directory that are not in manifest",
},
&cli.StringFlag{
Name: "require-signature",
Aliases: []string{"S"},
Usage: "Require manifest to be signed by the specified GPG key ID",
EnvVars: []string{"MFER_REQUIRE_SIGNATURE"},
},
),
},
{
Name: "freshen",
Usage: "Update manifest with changed, new, and removed files",
ArgsUsage: "[manifest file]",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.freshenManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths",
},
&cli.BoolFlag{
Name: "follow-symlinks",
Aliases: []string{"L"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "include-dotfiles",
Aliases: []string{"IncludeDotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.BoolFlag{
Name: "progress",
Aliases: []string{"P"},
Usage: "Show progress during scanning and hashing",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
EnvVars: []string{"MFER_SIGN_KEY"},
},
&cli.BoolFlag{
Name: "include-timestamps",
Usage: "Include createdAt timestamp in manifest (omitted by default for determinism)",
},
),
},
{
Name: "export",
Usage: "Export manifest contents as JSON",
ArgsUsage: "[manifest file or URL]",
Action: func(c *cli.Context) error {
return mfa.exportManifestOperation(c)
},
},
{
Name: "version",
Usage: "Show version",
Action: func(c *cli.Context) error {
_, _ = fmt.Fprintln(mfa.Stdout, mfa.VersionString())
return nil
},
},
{
Name: "list",
Aliases: []string{"ls"},
Usage: "List files in manifest",
ArgsUsage: "[manifest file]",
Action: func(c *cli.Context) error {
return mfa.listManifestOperation(c)
},
Flags: []cli.Flag{
&cli.BoolFlag{
Name: "long",
Aliases: []string{"l"},
Usage: "Show size and mtime",
},
&cli.BoolFlag{
Name: "print0",
Usage: "Separate entries with NUL character (for xargs -0)",
},
},
},
{
Name: "fetch",
Usage: "fetch manifest and referenced files",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.fetchManifestOperation(c)
},
Flags: commonFlags(),
},
},
}
mfa.app.HideVersion = false
err := mfa.app.Run(args)
if err != nil {
mfa.exitCode = 1
log.Errorf("%s", err)
log.WithError(err).Debugf("exiting")
}
}
-28
View File
@@ -1,28 +0,0 @@
package cli
import (
"time"
"sneak.berlin/go/mfer/mfer"
)
// mtimeAbsent is printed in place of a modification time when a manifest
// entry does not carry one.
const mtimeAbsent = "-"
// entryMtime returns the modification time recorded for a manifest entry.
//
// MFFilePath.Mtime is a message pointer with proto3 field presence, so an
// absent mtime is a representable, on-the-wire-valid state. It must never
// be conflated with a recorded mtime of the Unix epoch: callers that
// compare mtimes have to treat "absent" as "unknown", not as
// 1970-01-01T00:00:00Z, or every entry compares as modified. ok reports
// whether an mtime was actually recorded.
func entryMtime(entry *mfer.MFFilePath) (time.Time, bool) {
ts := entry.GetMtime()
if ts == nil {
return time.Time{}, false
}
return time.Unix(ts.GetSeconds(), int64(ts.GetNanos())), true
}
+52 -64
View File
@@ -1,5 +1,3 @@
// Package log provides leveled logging with progress output helpers
// on top of apex/log and pterm.
package log
import (
@@ -54,11 +52,6 @@ func (l Level) String() string {
}
}
// callerSkip is the runtime.Caller stack depth from the public Debug
// helpers to the caller of the log package.
const callerSkip = 2
//nolint:gochecknoglobals // package-level logger state by design
var (
// mu protects the output writers and level
mu sync.RWMutex
@@ -67,7 +60,7 @@ var (
// stderr is the writer for log output
stderr io.Writer = os.Stderr
// currentLevel is our log level (includes Verbose)
currentLevel = InfoLevel
currentLevel Level = InfoLevel
)
// SetOutput configures the output writers for the log package.
@@ -75,10 +68,8 @@ var (
func SetOutput(out, err io.Writer) {
mu.Lock()
defer mu.Unlock()
stdout = out
stderr = err
pterm.SetDefaultOutput(out)
}
@@ -86,7 +77,6 @@ func SetOutput(out, err io.Writer) {
func GetStdout() io.Writer {
mu.RLock()
defer mu.RUnlock()
return stdout
}
@@ -94,7 +84,6 @@ func GetStdout() io.Writer {
func GetStderr() io.Writer {
mu.RLock()
defer mu.RUnlock()
return stderr
}
@@ -102,7 +91,6 @@ func GetStderr() io.Writer {
func DisableStyling() {
pterm.DisableColor()
pterm.DisableStyling()
pterm.Debug.Prefix.Text = ""
pterm.Info.Prefix.Text = ""
pterm.Success.Prefix.Text = ""
@@ -112,16 +100,11 @@ func DisableStyling() {
}
// Init initializes the logger with the CLI handler and default log level.
//
// It reconfigures the process-global apex/log logger under the write lock so
// the global is never mutated while another goroutine holds the read lock to
// read it in emit. Without this, parallel callers (e.g. the test suite) race
// Init's SetLevel/SetHandler against concurrent log calls.
func Init() {
mu.Lock()
defer mu.Unlock()
log.SetHandler(acli.New(stderr))
mu.RLock()
w := stderr
mu.RUnlock()
log.SetHandler(acli.New(w))
log.SetLevel(log.DebugLevel) // Let apex/log pass everything; we filter ourselves
}
@@ -129,108 +112,110 @@ func Init() {
func isEnabled(l Level) bool {
mu.RLock()
defer mu.RUnlock()
return l >= currentLevel
}
// emit calls fn while holding the read lock if messages at level l are
// enabled. Holding the read lock across the apex/log call keeps the global
// logger from being read while Init reconfigures it under the write lock.
func emit(l Level, fn func()) {
mu.RLock()
defer mu.RUnlock()
if l >= currentLevel {
fn()
}
}
// Fatalf logs a formatted message at fatal level.
func Fatalf(format string, args ...any) {
emit(FatalLevel, func() { log.Fatalf(format, args...) })
func Fatalf(format string, args ...interface{}) {
if isEnabled(FatalLevel) {
log.Fatalf(format, args...)
}
}
// Fatal logs a message at fatal level.
func Fatal(arg string) {
emit(FatalLevel, func() { log.Fatal(arg) })
if isEnabled(FatalLevel) {
log.Fatal(arg)
}
}
// Errorf logs a formatted message at error level.
func Errorf(format string, args ...any) {
emit(ErrorLevel, func() { log.Errorf(format, args...) })
func Errorf(format string, args ...interface{}) {
if isEnabled(ErrorLevel) {
log.Errorf(format, args...)
}
}
// Error logs a message at error level.
func Error(arg string) {
emit(ErrorLevel, func() { log.Error(arg) })
if isEnabled(ErrorLevel) {
log.Error(arg)
}
}
// Warnf logs a formatted message at warn level.
func Warnf(format string, args ...any) {
emit(WarnLevel, func() { log.Warnf(format, args...) })
func Warnf(format string, args ...interface{}) {
if isEnabled(WarnLevel) {
log.Warnf(format, args...)
}
}
// Warn logs a message at warn level.
func Warn(arg string) {
emit(WarnLevel, func() { log.Warn(arg) })
if isEnabled(WarnLevel) {
log.Warn(arg)
}
}
// Infof logs a formatted message at info level.
func Infof(format string, args ...any) {
emit(InfoLevel, func() { log.Infof(format, args...) })
func Infof(format string, args ...interface{}) {
if isEnabled(InfoLevel) {
log.Infof(format, args...)
}
}
// Info logs a message at info level.
func Info(arg string) {
emit(InfoLevel, func() { log.Info(arg) })
if isEnabled(InfoLevel) {
log.Info(arg)
}
}
// Verbosef logs a formatted message at verbose level.
func Verbosef(format string, args ...any) {
emit(VerboseLevel, func() { log.Infof(format, args...) })
func Verbosef(format string, args ...interface{}) {
if isEnabled(VerboseLevel) {
log.Infof(format, args...)
}
}
// Verbose logs a message at verbose level.
func Verbose(arg string) {
emit(VerboseLevel, func() { log.Info(arg) })
if isEnabled(VerboseLevel) {
log.Info(arg)
}
}
// Debugf logs a formatted message at debug level with caller location.
func Debugf(format string, args ...any) {
func Debugf(format string, args ...interface{}) {
if isEnabled(DebugLevel) {
DebugReal(fmt.Sprintf(format, args...), callerSkip)
DebugReal(fmt.Sprintf(format, args...), 2)
}
}
// Debug logs a message at debug level with caller location.
func Debug(arg string) {
if isEnabled(DebugLevel) {
DebugReal(arg, callerSkip)
DebugReal(arg, 2)
}
}
// DebugReal logs at debug level with caller info from the specified stack depth.
func DebugReal(arg string, cs int) {
mu.RLock()
defer mu.RUnlock()
if DebugLevel < currentLevel {
if !isEnabled(DebugLevel) {
return
}
_, callerFile, callerLine, ok := runtime.Caller(cs)
if !ok {
return
}
tag := fmt.Sprintf("%s:%d: ", filepath.Base(callerFile), callerLine)
log.Debug(tag + arg)
}
// Dump logs a spew dump of the arguments at debug level.
func Dump(args ...any) {
func Dump(args ...interface{}) {
if isEnabled(DebugLevel) {
DebugReal(spew.Sdump(args...), callerSkip)
DebugReal(spew.Sdump(args...), 2)
}
}
@@ -261,7 +246,6 @@ func SetLevelFromVerbosity(l int) {
func SetLevel(l Level) {
mu.Lock()
defer mu.Unlock()
currentLevel = l
}
@@ -269,13 +253,17 @@ func SetLevel(l Level) {
func GetLevel() Level {
mu.RLock()
defer mu.RUnlock()
return currentLevel
}
// WithError returns a log entry with the error attached.
func WithError(e error) *log.Entry {
return log.Log.WithError(e)
}
// Progressf prints a progress message that overwrites the current line.
// Use ProgressDone() when progress is complete to move to the next line.
func Progressf(format string, args ...any) {
func Progressf(format string, args ...interface{}) {
pterm.Printf("\r"+format, args...)
}
+4 -4
View File
@@ -1,12 +1,12 @@
package log_test
package log
import (
"testing"
"sneak.berlin/go/mfer/internal/log"
"github.com/stretchr/testify/assert"
)
func TestBuild(t *testing.T) {
t.Parallel()
log.Init()
Init()
assert.True(t, true)
}
+32 -82
View File
@@ -1,9 +1,6 @@
// Package mfer implements the mfer manifest file format: building,
// serializing, verifying, and checking manifests of file trees.
package mfer
import (
"context"
"crypto/sha256"
"errors"
"fmt"
@@ -17,27 +14,6 @@ import (
"github.com/multiformats/go-multihash"
)
// readChunkSize is the buffer size used when reading file contents for
// hashing.
const readChunkSize = 64 * 1024
// The errPath* sentinels below are worded as the trailing fragment of the
// message ValidatePath renders, because the offending path is quoted
// before them (`path %q ...`). Wrapping them mid-sentence keeps the
// rendered text exactly as mfer has always printed it. Match them with
// errors.Is rather than by reading their messages.
var (
errPathEmpty = errors.New("path cannot be empty")
errPathNotUTF8 = errors.New("is not valid UTF-8")
errPathBackslash = errors.New("contains backslash; use forward slashes only")
errPathAbsolute = errors.New("is absolute; must be relative")
errPathEmptySegment = errors.New("contains empty segment")
errPathDotDot = errors.New("contains '..' segment")
errSizeMismatch = errors.New("size mismatch")
errNegativeSize = errors.New("size cannot be negative")
errEmptyHash = errors.New("hash cannot be nil or empty")
)
// ValidatePath checks that a file path conforms to manifest path invariants:
// - Must be valid UTF-8
// - Must use forward slashes only (no backslashes)
@@ -47,31 +23,25 @@ var (
// - Must not be empty
func ValidatePath(p string) error {
if p == "" {
return errPathEmpty
return errors.New("path cannot be empty")
}
if !utf8.ValidString(p) {
return fmt.Errorf("path %q %w", p, errPathNotUTF8)
return fmt.Errorf("path %q is not valid UTF-8", p)
}
if strings.ContainsRune(p, '\\') {
return fmt.Errorf("path %q %w", p, errPathBackslash)
return fmt.Errorf("path %q contains backslash; use forward slashes only", p)
}
if strings.HasPrefix(p, "/") {
return fmt.Errorf("path %q %w", p, errPathAbsolute)
return fmt.Errorf("path %q is absolute; must be relative", p)
}
for _, seg := range strings.Split(p, "/") {
if seg == "" {
return fmt.Errorf("path %q %w", p, errPathEmptySegment)
return fmt.Errorf("path %q contains empty segment", p)
}
if seg == ".." {
return fmt.Errorf("path %q %w", p, errPathDotDot)
return fmt.Errorf("path %q contains '..' segment", p)
}
}
return nil
}
@@ -98,7 +68,11 @@ type UnixNanos int32
// Timestamp converts ModTime to a protobuf Timestamp.
func (m ModTime) Timestamp() *Timestamp {
return newTimestampFromTime(time.Time(m))
t := time.Time(m)
return &Timestamp{
Seconds: t.Unix(),
Nanos: int32(t.Nanosecond()),
}
}
// Multihash represents a multihash-encoded file hash (typically SHA2-256).
@@ -119,6 +93,14 @@ type Builder struct {
fixedUUID []byte // if set, use this UUID instead of generating one
}
// SetSeed derives a deterministic UUID from the given seed string.
// The seed is hashed once with SHA-256 and the first 16 bytes are used
// as a fixed UUID for the manifest.
func (b *Builder) SetSeed(seed string) {
hash := sha256.Sum256([]byte(seed))
b.fixedUUID = hash[:16]
}
// NewBuilder creates a new Builder.
func NewBuilder() *Builder {
return &Builder{
@@ -127,14 +109,6 @@ func NewBuilder() *Builder {
}
}
// SetSeed derives a deterministic UUID from the given seed string.
// The seed is hashed once with SHA-256 and the first 16 bytes are used
// as a fixed UUID for the manifest.
func (b *Builder) SetSeed(seed string) {
hash := sha256.Sum256([]byte(seed))
b.fixedUUID = hash[:uuidLength]
}
// AddFile reads file content from reader, computes hashes, and adds to manifest.
// Progress updates are sent to the progress channel (if non-nil) without blocking.
// Returns the number of bytes read.
@@ -145,8 +119,7 @@ func (b *Builder) AddFile(
reader io.Reader,
progress chan<- FileHashProgress,
) (FileSize, error) {
err := ValidatePath(string(path))
if err != nil {
if err := ValidatePath(string(path)); err != nil {
return 0, err
}
@@ -155,8 +128,7 @@ func (b *Builder) AddFile(
// Read file in chunks, updating hash and progress
var totalRead FileSize
buf := make([]byte, readChunkSize)
buf := make([]byte, 64*1024) // 64KB chunks
for {
n, err := reader.Read(buf)
@@ -165,11 +137,9 @@ func (b *Builder) AddFile(
totalRead += FileSize(n)
sendFileHashProgress(progress, FileHashProgress{BytesRead: totalRead})
}
if err == io.EOF {
break
}
if err != nil {
return totalRead, err
}
@@ -177,10 +147,7 @@ func (b *Builder) AddFile(
// Verify actual bytes read matches declared size
if totalRead != size {
return totalRead, fmt.Errorf(
"%w for %q: declared %d bytes but read %d bytes",
errSizeMismatch, path, size, totalRead,
)
return totalRead, fmt.Errorf("size mismatch for %q: declared %d bytes but read %d bytes", path, size, totalRead)
}
// Encode hash as multihash (SHA2-256)
@@ -211,7 +178,6 @@ func sendFileHashProgress(ch chan<- FileHashProgress, p FileHashProgress) {
if ch == nil {
return
}
select {
case ch <- p:
default:
@@ -222,30 +188,21 @@ func sendFileHashProgress(ch chan<- FileHashProgress, p FileHashProgress) {
func (b *Builder) FileCount() int {
b.mu.Lock()
defer b.mu.Unlock()
return len(b.files)
}
// AddFileWithHash adds a file entry with a pre-computed hash.
// This is useful when the hash is already known (e.g., from an existing manifest).
// Returns an error if path is empty, size is negative, or hash is nil/empty.
func (b *Builder) AddFileWithHash(
path RelFilePath,
size FileSize,
mtime ModTime,
hash Multihash,
) error {
err := ValidatePath(string(path))
if err != nil {
func (b *Builder) AddFileWithHash(path RelFilePath, size FileSize, mtime ModTime, hash Multihash) error {
if err := ValidatePath(string(path)); err != nil {
return fmt.Errorf("add file: %w", err)
}
if size < 0 {
return errNegativeSize
return errors.New("size cannot be negative")
}
if len(hash) == 0 {
return errEmptyHash
return errors.New("hash cannot be nil or empty")
}
entry := &MFFilePath{
@@ -260,7 +217,6 @@ func (b *Builder) AddFileWithHash(
b.mu.Lock()
b.files = append(b.files, entry)
b.mu.Unlock()
return nil
}
@@ -269,7 +225,6 @@ func (b *Builder) AddFileWithHash(
func (b *Builder) SetIncludeTimestamps(include bool) {
b.mu.Lock()
defer b.mu.Unlock()
b.includeTimestamps = include
}
@@ -278,19 +233,17 @@ func (b *Builder) SetIncludeTimestamps(include bool) {
func (b *Builder) SetSigningOptions(opts *SigningOptions) {
b.mu.Lock()
defer b.mu.Unlock()
b.signingOptions = opts
}
// Build finalizes the manifest and writes it to the writer. ctx bounds the
// gpg runs that sign the manifest when signing options are set.
func (b *Builder) Build(ctx context.Context, w io.Writer) error {
// Build finalizes the manifest and writes it to the writer.
func (b *Builder) Build(w io.Writer) error {
b.mu.Lock()
defer b.mu.Unlock()
// Sort files by path for deterministic output
sort.Slice(b.files, func(i, j int) bool {
return b.files[i].GetPath() < b.files[j].GetPath()
return b.files[i].Path < b.files[j].Path
})
// Create inner manifest
@@ -310,22 +263,19 @@ func (b *Builder) Build(ctx context.Context, w io.Writer) error {
}
// Generate outer wrapper
err := m.generateOuter(ctx)
if err != nil {
if err := m.generateOuter(); err != nil {
return fmt.Errorf("build: generate outer: %w", err)
}
// Generate final output
err = m.generate(ctx)
if err != nil {
if err := m.generate(); err != nil {
return fmt.Errorf("build: generate: %w", err)
}
// Write to output
_, err = w.Write(m.output.Bytes())
_, err := w.Write(m.output.Bytes())
if err != nil {
return fmt.Errorf("build: write output: %w", err)
}
return nil
}
+44 -170
View File
@@ -1,10 +1,7 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"bytes"
"context"
"fmt"
"strings"
"testing"
"time"
@@ -13,34 +10,24 @@ import (
"github.com/stretchr/testify/require"
)
const testFileName = "file.txt"
func TestNewBuilder(t *testing.T) {
t.Parallel()
b := NewBuilder()
assert.NotNil(t, b)
assert.Equal(t, 0, b.FileCount())
}
func TestBuilderAddFile(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("test content")
reader := bytes.NewReader(content)
bytesRead, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
bytesRead, err := b.AddFile("test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil)
require.NoError(t, err)
assert.Equal(t, FileSize(len(content)), bytesRead)
assert.Equal(t, 1, b.FileCount())
}
func TestBuilderAddFileWithHash(t *testing.T) {
t.Parallel()
b := NewBuilder()
hash := make([]byte, 34) // SHA256 multihash is 34 bytes
@@ -50,72 +37,55 @@ func TestBuilderAddFileWithHash(t *testing.T) {
}
func TestBuilderAddFileWithHashValidation(t *testing.T) {
t.Parallel()
t.Run("empty path", func(t *testing.T) {
t.Parallel()
b := NewBuilder()
hash := make([]byte, 34)
err := b.AddFileWithHash("", 100, ModTime(time.Now()), hash)
require.Error(t, err)
assert.Error(t, err)
assert.Contains(t, err.Error(), "path")
})
t.Run("negative size", func(t *testing.T) {
t.Parallel()
b := NewBuilder()
hash := make([]byte, 34)
err := b.AddFileWithHash("test.txt", -1, ModTime(time.Now()), hash)
require.Error(t, err)
assert.Error(t, err)
assert.Contains(t, err.Error(), "size")
})
t.Run("nil hash", func(t *testing.T) {
t.Parallel()
b := NewBuilder()
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), nil)
require.Error(t, err)
assert.Error(t, err)
assert.Contains(t, err.Error(), "hash")
})
t.Run("empty hash", func(t *testing.T) {
t.Parallel()
b := NewBuilder()
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), []byte{})
require.Error(t, err)
assert.Error(t, err)
assert.Contains(t, err.Error(), "hash")
})
t.Run("valid inputs", func(t *testing.T) {
t.Parallel()
b := NewBuilder()
hash := make([]byte, 34)
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), hash)
require.NoError(t, err)
assert.NoError(t, err)
assert.Equal(t, 1, b.FileCount())
})
}
func TestBuilderBuild(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("test content")
reader := bytes.NewReader(content)
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
_, err := b.AddFile("test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil)
require.NoError(t, err)
var buf bytes.Buffer
err = b.Build(context.Background(), &buf)
err = b.Build(&buf)
require.NoError(t, err)
// Should have magic bytes
@@ -123,8 +93,6 @@ func TestBuilderBuild(t *testing.T) {
}
func TestNewTimestampFromTimeExtremeDate(t *testing.T) {
t.Parallel()
// Regression test: newTimestampFromTime used UnixNano() which panics
// for dates outside ~1678-2262. Now uses Nanosecond() which is safe.
tests := []struct {
@@ -139,19 +107,15 @@ func TestNewTimestampFromTimeExtremeDate(t *testing.T) {
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
// Should not panic
ts := newTimestampFromTime(tt.time)
assert.Equal(t, tt.time.Unix(), ts.GetSeconds())
assert.Equal(t, tt.time.Nanosecond(), int(ts.GetNanos()))
assert.Equal(t, tt.time.Unix(), ts.Seconds)
assert.Equal(t, int32(tt.time.Nanosecond()), ts.Nanos)
})
}
}
func TestBuilderDeterministicOutput(t *testing.T) {
t.Parallel()
buildManifest := func() []byte {
b := NewBuilder()
// Use a fixed createdAt and UUID so output is reproducible
@@ -171,32 +135,24 @@ func TestBuilderDeterministicOutput(t *testing.T) {
}
for _, f := range files {
r := bytes.NewReader([]byte(f.content))
_, err := b.AddFile(
RelFilePath(f.path), FileSize(len(f.content)), mtime, r, nil,
)
_, err := b.AddFile(RelFilePath(f.path), FileSize(len(f.content)), mtime, r, nil)
require.NoError(t, err)
}
var buf bytes.Buffer
err := b.Build(context.Background(), &buf)
err := b.Build(&buf)
require.NoError(t, err)
return buf.Bytes()
}
out1 := buildManifest()
out2 := buildManifest()
assert.Equal(t, out1, out2,
"two builds with same input should produce byte-identical output")
assert.Equal(t, out1, out2, "two builds with same input should produce byte-identical output")
}
func TestSetSeedDeterministic(t *testing.T) {
t.Parallel()
b1 := NewBuilder()
b1.SetSeed("test-seed-value")
b2 := NewBuilder()
b2.SetSeed("test-seed-value")
assert.Equal(t, b1.fixedUUID, b2.fixedUUID, "same seed should produce same UUID")
@@ -204,24 +160,19 @@ func TestSetSeedDeterministic(t *testing.T) {
b3 := NewBuilder()
b3.SetSeed("different-seed")
assert.NotEqual(t, b1.fixedUUID, b3.fixedUUID,
"different seeds should produce different UUIDs")
assert.NotEqual(t, b1.fixedUUID, b3.fixedUUID, "different seeds should produce different UUIDs")
}
func TestValidatePath(t *testing.T) {
t.Parallel()
valid := []string{
testFileName,
"file.txt",
"dir/file.txt",
"a/b/c/d.txt",
"file with spaces.txt",
"日本語.txt", //nolint:gosmopolitan // deliberately tests non-ASCII UTF-8 paths
"日本語.txt",
}
for _, p := range valid {
t.Run("valid:"+p, func(t *testing.T) {
t.Parallel()
assert.NoError(t, ValidatePath(p))
})
}
@@ -240,54 +191,42 @@ func TestValidatePath(t *testing.T) {
}
for _, tt := range invalid {
t.Run("invalid:"+tt.desc, func(t *testing.T) {
t.Parallel()
assert.Error(t, ValidatePath(tt.path))
})
}
}
func TestBuilderAddFileSizeMismatch(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("short")
reader := bytes.NewReader(content)
// Declare wrong size
_, err := b.AddFile("test.txt", FileSize(100), ModTime(time.Now()), reader, nil)
require.Error(t, err)
assert.Error(t, err)
assert.Contains(t, err.Error(), "size mismatch")
}
func TestBuilderAddFileInvalidPath(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("data")
reader := bytes.NewReader(content)
_, err := b.AddFile("", FileSize(len(content)), ModTime(time.Now()), reader, nil)
require.Error(t, err)
assert.Error(t, err)
reader.Reset(content)
_, err = b.AddFile(
"/absolute", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
_, err = b.AddFile("/absolute", FileSize(len(content)), ModTime(time.Now()), reader, nil)
assert.Error(t, err)
}
func TestBuilderAddFileWithProgress(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := bytes.Repeat([]byte("x"), 1000)
reader := bytes.NewReader(content)
progress := make(chan FileHashProgress, 100)
bytesRead, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, progress,
)
bytesRead, err := b.AddFile("test.txt", FileSize(len(content)), ModTime(time.Now()), reader, progress)
close(progress)
require.NoError(t, err)
assert.Equal(t, FileSize(1000), bytesRead)
@@ -296,15 +235,12 @@ func TestBuilderAddFileWithProgress(t *testing.T) {
for p := range progress {
updates = append(updates, p)
}
assert.NotEmpty(t, updates)
// Last update should show all bytes
assert.Equal(t, FileSize(1000), updates[len(updates)-1].BytesRead)
}
func TestBuilderBuildRoundTrip(t *testing.T) {
t.Parallel()
// Build a manifest, deserialize it, verify all fields survive round-trip
b := NewBuilder()
now := time.Date(2025, 6, 15, 12, 0, 0, 0, time.UTC)
@@ -320,14 +256,12 @@ func TestBuilderBuildRoundTrip(t *testing.T) {
for _, f := range files {
reader := bytes.NewReader(f.content)
_, err := b.AddFile(
RelFilePath(f.path), FileSize(len(f.content)), ModTime(now), reader, nil,
)
_, err := b.AddFile(RelFilePath(f.path), FileSize(len(f.content)), ModTime(now), reader, nil)
require.NoError(t, err)
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
require.NoError(t, b.Build(&buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
@@ -336,80 +270,46 @@ func TestBuilderBuildRoundTrip(t *testing.T) {
require.Len(t, mfiles, 3)
// Verify sorted order
assert.Equal(t, "alpha.txt", mfiles[0].GetPath())
assert.Equal(t, "beta/delta.txt", mfiles[1].GetPath())
assert.Equal(t, "beta/gamma.txt", mfiles[2].GetPath())
assert.Equal(t, "alpha.txt", mfiles[0].Path)
assert.Equal(t, "beta/delta.txt", mfiles[1].Path)
assert.Equal(t, "beta/gamma.txt", mfiles[2].Path)
// Verify sizes
assert.Equal(t, int64(len("alpha content")), mfiles[0].GetSize())
assert.Equal(t, int64(len("alpha content")), mfiles[0].Size)
// Verify hashes are present
for _, f := range mfiles {
require.NotEmpty(t, f.GetHashes(), "file %s should have hashes", f.GetPath())
assert.NotEmpty(t, f.GetHashes()[0].GetMultiHash())
require.NotEmpty(t, f.Hashes, "file %s should have hashes", f.Path)
assert.NotEmpty(t, f.Hashes[0].MultiHash)
}
}
// A manifest whose payload is larger than one zstd block (128 KiB) is
// written as a frame asking for the writer's whole window, zstdWindowSize,
// which is the most the parser accepts.
func TestBuilderBuildRoundTripLargeManifest(t *testing.T) {
t.Parallel()
hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256
b := NewBuilder()
for i := range 4000 {
path := RelFilePath(fmt.Sprintf("dir/file-%05d.txt", i))
require.NoError(t, b.AddFileWithHash(path, FileSize(i), ModTime{}, hash))
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
require.Greater(t, m.pbOuter.GetSize(), int64(128<<10))
assert.Len(t, m.Files(), 4000)
}
func TestNewManifestFromReaderInvalidMagic(t *testing.T) {
t.Parallel()
_, err := NewManifestFromReader(bytes.NewReader([]byte("NOT_VALID")))
require.Error(t, err)
assert.Error(t, err)
assert.Contains(t, err.Error(), "invalid file format")
}
func TestNewManifestFromReaderEmpty(t *testing.T) {
t.Parallel()
_, err := NewManifestFromReader(bytes.NewReader([]byte{}))
assert.Error(t, err)
}
func TestNewManifestFromReaderTruncated(t *testing.T) {
t.Parallel()
// Just the magic with nothing after
_, err := NewManifestFromReader(bytes.NewReader([]byte(MAGIC)))
assert.Error(t, err)
}
func TestManifestString(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("test")
reader := bytes.NewReader(content)
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
_, err := b.AddFile("test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil)
require.NoError(t, err)
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
require.NoError(t, b.Build(&buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
@@ -417,13 +317,10 @@ func TestManifestString(t *testing.T) {
}
func TestBuilderBuildEmpty(t *testing.T) {
t.Parallel()
b := NewBuilder()
var buf bytes.Buffer
err := b.Build(context.Background(), &buf)
err := b.Build(&buf)
require.NoError(t, err)
// Should still produce valid manifest with 0 files
@@ -431,69 +328,48 @@ func TestBuilderBuildEmpty(t *testing.T) {
}
func TestBuilderOmitsCreatedAtByDefault(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("hello")
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()),
bytes.NewReader(content), nil,
)
_, err := b.AddFile("test.txt", FileSize(len(content)), ModTime(time.Now()), bytes.NewReader(content), nil)
require.NoError(t, err)
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
require.NoError(t, b.Build(&buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.Nil(t, m.pbInner.GetCreatedAt(),
"createdAt should be nil by default for deterministic output")
assert.Nil(t, m.pbInner.CreatedAt, "createdAt should be nil by default for deterministic output")
}
func TestBuilderIncludesCreatedAtWhenRequested(t *testing.T) {
t.Parallel()
b := NewBuilder()
b.SetIncludeTimestamps(true)
content := []byte("hello")
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()),
bytes.NewReader(content), nil,
)
_, err := b.AddFile("test.txt", FileSize(len(content)), ModTime(time.Now()), bytes.NewReader(content), nil)
require.NoError(t, err)
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
require.NoError(t, b.Build(&buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.NotNil(t, m.pbInner.GetCreatedAt(),
"createdAt should be set when IncludeTimestamps is true")
assert.NotNil(t, m.pbInner.CreatedAt, "createdAt should be set when IncludeTimestamps is true")
}
func TestBuilderDeterministicFileOrder(t *testing.T) {
t.Parallel()
// Two builds with same files in different order should produce same file ordering.
// Note: UUIDs differ per build, so we compare parsed file lists, not raw bytes.
buildAndParse := func(order []string) []*MFFilePath {
b := NewBuilder()
for _, name := range order {
content := []byte("content of " + name)
_, err := b.AddFile(
RelFilePath(name), FileSize(len(content)),
ModTime(time.Unix(1000, 0)), bytes.NewReader(content), nil,
)
_, err := b.AddFile(RelFilePath(name), FileSize(len(content)), ModTime(time.Unix(1000, 0)), bytes.NewReader(content), nil)
require.NoError(t, err)
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
require.NoError(t, b.Build(&buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
return m.Files()
}
@@ -502,12 +378,10 @@ func TestBuilderDeterministicFileOrder(t *testing.T) {
require.Len(t, files1, 2)
require.Len(t, files2, 2)
for i := range files1 {
assert.Equal(t, files1[i].GetPath(), files2[i].GetPath())
assert.Equal(t, files1[i].GetSize(), files2[i].GetSize())
assert.Equal(t, files1[i].Path, files2[i].Path)
assert.Equal(t, files1[i].Size, files2[i].Size)
}
assert.Equal(t, "a.txt", files1[0].GetPath())
assert.Equal(t, "b.txt", files1[1].GetPath())
assert.Equal(t, "a.txt", files1[0].Path)
assert.Equal(t, "b.txt", files1[1].Path)
}
+80 -112
View File
@@ -14,8 +14,6 @@ import (
"github.com/spf13/afero"
)
var errNoSigningPubKey = errors.New("manifest has no signing public key")
// Result represents the outcome of checking a single file.
type Result struct {
Path RelFilePath // Relative path from manifest
@@ -26,7 +24,6 @@ type Result struct {
// Status represents the verification status of a file.
type Status int
// Verification result statuses reported for each checked file.
const (
StatusOK Status = iota // File matches manifest (size and hash verified)
StatusMissing // File not found on disk
@@ -73,8 +70,7 @@ type Checker struct {
fs afero.Fs
// manifestPaths is a set of paths in the manifest for quick lookup
manifestPaths map[RelFilePath]struct{}
// manifestRelPath is the relative path of the manifest file from
// basePath (for exclusion)
// manifestRelPath is the relative path of the manifest file from basePath (for exclusion)
manifestRelPath RelFilePath
// signature info from the manifest
signature []byte
@@ -101,10 +97,9 @@ func NewChecker(manifestPath string, basePath string, fs afero.Fs) (*Checker, er
}
files := m.Files()
manifestPaths := make(map[RelFilePath]struct{}, len(files))
for _, f := range files {
manifestPaths[RelFilePath(f.GetPath())] = struct{}{}
manifestPaths[RelFilePath(f.Path)] = struct{}{}
}
// Compute manifest's relative path from basePath for exclusion in FindExtraFiles
@@ -112,7 +107,6 @@ func NewChecker(manifestPath string, basePath string, fs afero.Fs) (*Checker, er
if err != nil {
return nil, err
}
manifestRel, err := filepath.Rel(abs, absManifest)
if err != nil {
manifestRel = ""
@@ -124,9 +118,9 @@ func NewChecker(manifestPath string, basePath string, fs afero.Fs) (*Checker, er
fs: fs,
manifestPaths: manifestPaths,
manifestRelPath: RelFilePath(manifestRel),
signature: m.pbOuter.GetSignature(),
signer: m.pbOuter.GetSigner(),
signingPubKey: m.pbOuter.GetSigningPubKey(),
signature: m.pbOuter.Signature,
signer: m.pbOuter.Signer,
signingPubKey: m.pbOuter.SigningPubKey,
}, nil
}
@@ -139,9 +133,8 @@ func (c *Checker) FileCount() FileCount {
func (c *Checker) TotalBytes() FileSize {
var total FileSize
for _, f := range c.files {
total += FileSize(f.GetSize())
total += FileSize(f.Size)
}
return total
}
@@ -155,8 +148,7 @@ func (c *Checker) Signer() []byte {
return c.signer
}
// SigningPubKey returns the signing public key if the manifest is signed,
// nil otherwise.
// SigningPubKey returns the signing public key if the manifest is signed, nil otherwise.
func (c *Checker) SigningPubKey() []byte {
return c.signingPubKey
}
@@ -164,27 +156,21 @@ func (c *Checker) SigningPubKey() []byte {
// ExtractEmbeddedSigningKeyFP imports the manifest's embedded public key into a
// temporary keyring and extracts its fingerprint. This validates the key and
// returns its actual fingerprint from the key material itself.
func (c *Checker) ExtractEmbeddedSigningKeyFP(ctx context.Context) (string, error) {
func (c *Checker) ExtractEmbeddedSigningKeyFP() (string, error) {
if len(c.signingPubKey) == 0 {
return "", errNoSigningPubKey
return "", errors.New("manifest has no signing public key")
}
return gpgExtractPubKeyFingerprint(ctx, c.signingPubKey)
return gpgExtractPubKeyFingerprint(c.signingPubKey)
}
// Check verifies all files against the manifest.
// Results are sent to the results channel as files are checked.
// Progress updates are sent to the progress channel approximately once per second.
// Both channels are closed when the method returns.
func (c *Checker) Check(
ctx context.Context,
results chan<- Result,
progress chan<- CheckStatus,
) error {
func (c *Checker) Check(ctx context.Context, results chan<- Result, progress chan<- CheckStatus) error {
if results != nil {
defer close(results)
}
if progress != nil {
defer close(progress)
}
@@ -192,11 +178,9 @@ func (c *Checker) Check(
totalFiles := FileCount(len(c.files))
totalBytes := c.TotalBytes()
var (
checkedFiles FileCount
checkedBytes FileSize
failures FileCount
)
var checkedFiles FileCount
var checkedBytes FileSize
var failures FileCount
startTime := time.Now()
lastProgressTime := time.Now()
@@ -212,7 +196,6 @@ func (c *Checker) Check(
if result.Status != StatusOK {
failures++
}
checkedFiles++
if results != nil {
@@ -222,12 +205,19 @@ func (c *Checker) Check(
// Send progress at most once per second (rate-limited)
if progress != nil {
now := time.Now()
isLast := checkedFiles == totalFiles
if isLast || now.Sub(lastProgressTime) >= time.Second {
bytesPerSec, eta := computeRateETA(
time.Since(startTime), checkedBytes, totalBytes,
)
elapsed := time.Since(startTime)
var bytesPerSec float64
var eta time.Duration
if elapsed > 0 && checkedBytes > 0 {
bytesPerSec = float64(checkedBytes) / elapsed.Seconds()
remainingBytes := totalBytes - checkedBytes
if bytesPerSec > 0 {
eta = time.Duration(float64(remainingBytes)/bytesPerSec) * time.Second
}
}
sendCheckStatus(progress, CheckStatus{
TotalFiles: totalFiles,
@@ -238,7 +228,6 @@ func (c *Checker) Check(
ETA: eta,
Failures: failures,
})
lastProgressTime = now
}
}
@@ -247,6 +236,59 @@ func (c *Checker) Check(
return nil
}
func (c *Checker) checkFile(entry *MFFilePath, checkedBytes *FileSize) Result {
absPath := filepath.Join(string(c.basePath), entry.Path)
relPath := RelFilePath(entry.Path)
// Check if file exists
info, err := c.fs.Stat(absPath)
if err != nil {
if errors.Is(err, os.ErrNotExist) || errors.Is(err, afero.ErrFileNotFound) {
return Result{Path: relPath, Status: StatusMissing, Message: "file not found"}
}
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check size
if info.Size() != entry.Size {
*checkedBytes += FileSize(info.Size())
return Result{
Path: relPath,
Status: StatusSizeMismatch,
Message: "size mismatch",
}
}
// Open and hash file
f, err := c.fs.Open(absPath)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
defer func() { _ = f.Close() }()
h := sha256.New()
n, err := io.Copy(h, f)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
*checkedBytes += FileSize(n)
// Encode as multihash and compare
computed, err := multihash.Encode(h.Sum(nil), multihash.SHA2_256)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check against all hashes in manifest (at least one must match)
for _, hash := range entry.Hashes {
if bytes.Equal(computed, hash.MultiHash) {
return Result{Path: relPath, Status: StatusOK}
}
}
return Result{Path: relPath, Status: StatusHashMismatch, Message: "hash mismatch"}
}
// FindExtraFiles walks the filesystem and reports files not in the manifest.
// Results are sent to the results channel. The channel is closed when done.
// Hidden files/directories (starting with .) are skipped, as they are excluded
@@ -256,7 +298,7 @@ func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) err
defer close(results)
}
walkFn := func(walkPath string, info os.FileInfo, err error) error {
return afero.Walk(c.fs, string(c.basePath), func(walkPath string, info os.FileInfo, err error) error {
if err != nil {
return err
}
@@ -278,7 +320,6 @@ func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) err
if info.IsDir() {
return filepath.SkipDir
}
return nil
}
@@ -306,79 +347,7 @@ func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) err
}
return nil
}
return afero.Walk(c.fs, string(c.basePath), walkFn)
}
func (c *Checker) checkFile(entry *MFFilePath, checkedBytes *FileSize) Result {
// entry.GetPath() is safe to join here: a manifest's entry paths are
// validated against the path invariants when it is loaded (see
// deserializeInner) or built (see Builder.AddFile), so a traversal or
// absolute path can never reach this point.
absPath := filepath.Join(string(c.basePath), entry.GetPath())
relPath := RelFilePath(entry.GetPath())
// Check if file exists
info, err := c.fs.Stat(absPath)
if err != nil {
if errors.Is(err, os.ErrNotExist) || errors.Is(err, afero.ErrFileNotFound) {
return Result{
Path: relPath,
Status: StatusMissing,
Message: "file not found",
}
}
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check size
if info.Size() != entry.GetSize() {
*checkedBytes += FileSize(info.Size())
return Result{
Path: relPath,
Status: StatusSizeMismatch,
Message: "size mismatch",
}
}
// Open and hash file
f, err := c.fs.Open(absPath)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
defer func() { _ = f.Close() }()
h := sha256.New()
n, err := io.Copy(h, f)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
*checkedBytes += FileSize(n)
// Encode as multihash and compare
computed, err := multihash.Encode(h.Sum(nil), multihash.SHA2_256)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check against all hashes in manifest (at least one must match)
for _, hash := range entry.GetHashes() {
if bytes.Equal(computed, hash.GetMultiHash()) {
return Result{Path: relPath, Status: StatusOK}
}
}
return Result{
Path: relPath,
Status: StatusHashMismatch,
Message: "hash mismatch",
}
})
}
// sendCheckStatus sends a status update without blocking.
@@ -386,7 +355,6 @@ func sendCheckStatus(ch chan<- CheckStatus, status CheckStatus) {
if ch == nil {
return
}
select {
case ch <- status:
default:
+63 -144
View File
@@ -1,4 +1,3 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
@@ -13,15 +12,7 @@ import (
"github.com/stretchr/testify/require"
)
const (
testFile1 = "file1.txt"
testFile2 = "file2.txt"
testExistsFile = "exists.txt"
)
func TestStatusString(t *testing.T) {
t.Parallel()
tests := []struct {
status Status
expected string
@@ -37,41 +28,31 @@ func TestStatusString(t *testing.T) {
for _, tt := range tests {
t.Run(tt.expected, func(t *testing.T) {
t.Parallel()
assert.Equal(t, tt.expected, tt.status.String())
})
}
}
// createTestManifest creates a manifest file in the filesystem with the given files.
func createTestManifest(
t *testing.T, fs afero.Fs, manifestPath string, files map[string][]byte,
) {
func createTestManifest(t *testing.T, fs afero.Fs, manifestPath string, files map[string][]byte) {
t.Helper()
builder := NewBuilder()
for path, content := range files {
reader := bytes.NewReader(content)
_, err := builder.AddFile(
RelFilePath(path), FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
_, err := builder.AddFile(RelFilePath(path), FileSize(len(content)), ModTime(time.Now()), reader, nil)
require.NoError(t, err)
}
var buf bytes.Buffer
require.NoError(t, builder.Build(context.Background(), &buf))
require.NoError(t, builder.Build(&buf))
require.NoError(t, afero.WriteFile(fs, manifestPath, buf.Bytes(), 0o644))
}
// createFilesOnDisk creates the given files on the filesystem under
// /data.
func createFilesOnDisk(t *testing.T, fs afero.Fs, files map[string][]byte) {
// createFilesOnDisk creates the given files on the filesystem.
func createFilesOnDisk(t *testing.T, fs afero.Fs, basePath string, files map[string][]byte) {
t.Helper()
basePath := "/data"
for path, content := range files {
fullPath := basePath + "/" + path
require.NoError(t, fs.MkdirAll(basePath, 0o755))
@@ -80,15 +61,11 @@ func createFilesOnDisk(t *testing.T, fs afero.Fs, files map[string][]byte) {
}
func TestNewChecker(t *testing.T) {
t.Parallel()
t.Run("valid manifest", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
testFile1: []byte("hello"),
testFile2: []byte("world"),
"file1.txt": []byte("hello"),
"file2.txt": []byte("world"),
}
createTestManifest(t, fs, "/manifest.mf", files)
@@ -99,16 +76,12 @@ func TestNewChecker(t *testing.T) {
})
t.Run("missing manifest", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
_, err := NewChecker("/nonexistent.mf", "/", fs)
assert.Error(t, err)
})
t.Run("invalid manifest", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/bad.mf", []byte("not a manifest"), 0o644))
_, err := NewChecker("/bad.mf", "/", fs)
@@ -117,8 +90,6 @@ func TestNewChecker(t *testing.T) {
}
func TestCheckerFileCountAndTotalBytes(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
"small.txt": []byte("hi"),
@@ -135,15 +106,13 @@ func TestCheckerFileCountAndTotalBytes(t *testing.T) {
}
func TestCheckAllFilesOK(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
testFile1: []byte("content one"),
testFile2: []byte("content two"),
"file1.txt": []byte("content one"),
"file2.txt": []byte("content two"),
}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, files)
createFilesOnDisk(t, fs, "/data", files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err)
@@ -158,24 +127,21 @@ func TestCheckAllFilesOK(t *testing.T) {
}
assert.Len(t, resultList, 2)
for _, r := range resultList {
assert.Equal(t, StatusOK, r.Status, "file %s should be OK", r.Path)
}
}
func TestCheckMissingFile(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
testExistsFile: []byte("I exist"),
"missing.txt": []byte("I don't exist on disk"),
"exists.txt": []byte("I exist"),
"missing.txt": []byte("I don't exist on disk"),
}
createTestManifest(t, fs, "/manifest.mf", files)
// Only create one file
createFilesOnDisk(t, fs, map[string][]byte{
testExistsFile: []byte("I exist"),
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"exists.txt": []byte("I exist"),
})
chk, err := NewChecker("/manifest.mf", "/data", fs)
@@ -186,17 +152,13 @@ func TestCheckMissingFile(t *testing.T) {
require.NoError(t, err)
var okCount, missingCount int
for r := range results {
switch r.Status {
case StatusOK:
okCount++
case StatusMissing:
missingCount++
assert.Equal(t, RelFilePath("missing.txt"), r.Path)
case StatusSizeMismatch, StatusHashMismatch, StatusExtra, StatusError:
// Not expected in this test; counted assertions below will fail.
}
}
@@ -205,16 +167,14 @@ func TestCheckMissingFile(t *testing.T) {
}
func TestCheckSizeMismatch(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
testFileName: []byte("original content"),
"file.txt": []byte("original content"),
}
createTestManifest(t, fs, "/manifest.mf", files)
// Create file with different size
createFilesOnDisk(t, fs, map[string][]byte{
testFileName: []byte("short"),
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"file.txt": []byte("short"),
})
chk, err := NewChecker("/manifest.mf", "/data", fs)
@@ -226,23 +186,21 @@ func TestCheckSizeMismatch(t *testing.T) {
r := <-results
assert.Equal(t, StatusSizeMismatch, r.Status)
assert.Equal(t, RelFilePath(testFileName), r.Path)
assert.Equal(t, RelFilePath("file.txt"), r.Path)
}
func TestCheckHashMismatch(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
originalContent := []byte("original content")
files := map[string][]byte{
testFileName: originalContent,
"file.txt": originalContent,
}
createTestManifest(t, fs, "/manifest.mf", files)
// Create file with same size but different content
differentContent := []byte("different contnt") // same length (16 bytes) but different
require.Len(t, differentContent, len(originalContent), "test requires same length")
createFilesOnDisk(t, fs, map[string][]byte{
testFileName: differentContent,
require.Equal(t, len(originalContent), len(differentContent), "test requires same length")
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"file.txt": differentContent,
})
chk, err := NewChecker("/manifest.mf", "/data", fs)
@@ -254,19 +212,17 @@ func TestCheckHashMismatch(t *testing.T) {
r := <-results
assert.Equal(t, StatusHashMismatch, r.Status)
assert.Equal(t, RelFilePath(testFileName), r.Path)
assert.Equal(t, RelFilePath("file.txt"), r.Path)
}
func TestCheckWithProgress(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
testFile1: bytes.Repeat([]byte("a"), 100),
testFile2: bytes.Repeat([]byte("b"), 200),
"file1.txt": bytes.Repeat([]byte("a"), 100),
"file2.txt": bytes.Repeat([]byte("b"), 200),
}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, files)
createFilesOnDisk(t, fs, "/data", files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err)
@@ -277,7 +233,9 @@ func TestCheckWithProgress(t *testing.T) {
err = chk.Check(context.Background(), results, progress)
require.NoError(t, err)
// results is fully buffered and closed; no draining needed
// Drain results
for range results {
}
// Check progress was sent
var progressUpdates []CheckStatus
@@ -296,17 +254,14 @@ func TestCheckWithProgress(t *testing.T) {
}
func TestCheckContextCancellation(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
// Create many files to ensure we have time to cancel
files := make(map[string][]byte)
for i := range 100 {
for i := 0; i < 100; i++ {
files[string(rune('a'+i%26))+".txt"] = bytes.Repeat([]byte("x"), 1000)
}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, files)
createFilesOnDisk(t, fs, "/data", files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err)
@@ -320,19 +275,17 @@ func TestCheckContextCancellation(t *testing.T) {
}
func TestFindExtraFiles(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
// Manifest only contains file1
manifestFiles := map[string][]byte{
testFile1: []byte("in manifest"),
"file1.txt": []byte("in manifest"),
}
createTestManifest(t, fs, "/manifest.mf", manifestFiles)
// Disk has file1 and file2
createFilesOnDisk(t, fs, map[string][]byte{
testFile1: []byte("in manifest"),
testFile2: []byte("extra file"),
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"file1.txt": []byte("in manifest"),
"file2.txt": []byte("extra file"),
})
chk, err := NewChecker("/manifest.mf", "/data", fs)
@@ -348,21 +301,19 @@ func TestFindExtraFiles(t *testing.T) {
}
assert.Len(t, extras, 1)
assert.Equal(t, RelFilePath(testFile2), extras[0].Path)
assert.Equal(t, RelFilePath("file2.txt"), extras[0].Path)
assert.Equal(t, StatusExtra, extras[0].Status)
assert.Equal(t, "not in manifest", extras[0].Message)
}
func TestFindExtraFilesSkipsManifestAndDotfiles(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
manifestFiles := map[string][]byte{
testFile1: []byte("in manifest"),
"file1.txt": []byte("in manifest"),
}
createTestManifest(t, fs, "/data/.index.mf", manifestFiles)
createFilesOnDisk(t, fs, map[string][]byte{
testFile1: []byte("in manifest"),
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"file1.txt": []byte("in manifest"),
})
// Create dotfile and manifest that should be skipped
require.NoError(t, afero.WriteFile(fs, "/data/.hidden", []byte("hidden"), 0o644))
@@ -387,21 +338,17 @@ func TestFindExtraFilesSkipsManifestAndDotfiles(t *testing.T) {
for _, e := range extras {
t.Logf("extra: %s", e.Path)
}
assert.Len(t, extras, 1)
if len(extras) > 0 {
assert.Equal(t, RelFilePath("extra.txt"), extras[0].Path)
}
}
func TestFindExtraFilesContextCancellation(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{testFileName: []byte("data")}
files := map[string][]byte{"file.txt": []byte("data")}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, files)
createFilesOnDisk(t, fs, "/data", files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err)
@@ -415,12 +362,10 @@ func TestFindExtraFilesContextCancellation(t *testing.T) {
}
func TestCheckNilChannels(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{testFileName: []byte("data")}
files := map[string][]byte{"file.txt": []byte("data")}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, files)
createFilesOnDisk(t, fs, "/data", files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err)
@@ -431,12 +376,10 @@ func TestCheckNilChannels(t *testing.T) {
}
func TestFindExtraFilesNilChannel(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{testFileName: []byte("data")}
files := map[string][]byte{"file.txt": []byte("data")}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, files)
createFilesOnDisk(t, fs, "/data", files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err)
@@ -447,8 +390,6 @@ func TestFindExtraFilesNilChannel(t *testing.T) {
}
func TestCheckSubdirectories(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
"dir1/file1.txt": []byte("content1"),
@@ -460,7 +401,6 @@ func TestCheckSubdirectories(t *testing.T) {
// Create files with full directory structure
for path, content := range files {
fullPath := "/data/" + path
require.NoError(t, fs.MkdirAll("/data/dir1/dir2/dir3", 0o755))
require.NoError(t, afero.WriteFile(fs, fullPath, content, 0o644))
}
@@ -473,30 +413,25 @@ func TestCheckSubdirectories(t *testing.T) {
require.NoError(t, err)
var okCount int
for r := range results {
assert.Equal(t, StatusOK, r.Status, "file %s should be OK", r.Path)
okCount++
}
assert.Equal(t, 3, okCount)
}
func TestCheckMissingFileDetectedWithoutFallback(t *testing.T) {
t.Parallel()
// Regression test: errors.Is(err, errors.New("...")) never matches because
// errors.New creates a new value each time. The fix uses os.ErrNotExist instead.
fs := afero.NewMemMapFs()
files := map[string][]byte{
testExistsFile: []byte("here"),
"missing.txt": []byte("not on disk"),
"exists.txt": []byte("here"),
"missing.txt": []byte("not on disk"),
}
createTestManifest(t, fs, "/manifest.mf", files)
// Only create one file on disk
createFilesOnDisk(t, fs, map[string][]byte{
testExistsFile: []byte("here"),
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"exists.txt": []byte("here"),
})
chk, err := NewChecker("/manifest.mf", "/data", fs)
@@ -513,29 +448,25 @@ func TestCheckMissingFileDetectedWithoutFallback(t *testing.T) {
assert.Equal(t, RelFilePath("missing.txt"), r.Path)
}
}
assert.Equal(t, 1, statusCounts[StatusOK], "one file should be OK")
assert.Equal(t, 1, statusCounts[StatusMissing], "one file should be MISSING")
assert.Equal(t, 0, statusCounts[StatusError], "no files should be ERROR")
}
func TestFindExtraFilesSkipsDotfiles(t *testing.T) {
t.Parallel()
// Regression test for #16: FindExtraFiles should not report dotfiles
// or the manifest file itself as extra files.
fs := afero.NewMemMapFs()
files := map[string][]byte{
testFile1: []byte("in manifest"),
"file1.txt": []byte("in manifest"),
}
createTestManifest(t, fs, "/data/.index.mf", files)
createFilesOnDisk(t, fs, files)
createFilesOnDisk(t, fs, "/data", files)
// Add dotfiles and manifest file on disk
require.NoError(t, afero.WriteFile(fs, "/data/.hidden", []byte("dotfile"), 0o644))
require.NoError(t, fs.MkdirAll("/data/.git", 0o755))
require.NoError(t,
afero.WriteFile(fs, "/data/.git/config", []byte("git config"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/data/.git/config", []byte("git config"), 0o644))
chk, err := NewChecker("/data/.index.mf", "/data", fs)
require.NoError(t, err)
@@ -550,21 +481,17 @@ func TestFindExtraFilesSkipsDotfiles(t *testing.T) {
}
// Should report NO extra files — dotfiles and manifest should be skipped
assert.Empty(t, extras,
"FindExtraFiles should not report dotfiles or manifest file as extra; got: %v",
extras)
assert.Empty(t, extras, "FindExtraFiles should not report dotfiles or manifest file as extra; got: %v", extras)
}
func TestFindExtraFilesSkipsManifestFile(t *testing.T) {
t.Parallel()
// The manifest file itself should never be reported as extra
fs := afero.NewMemMapFs()
files := map[string][]byte{
testFile1: []byte("content"),
"file1.txt": []byte("content"),
}
createTestManifest(t, fs, "/data/index.mf", files)
createFilesOnDisk(t, fs, files)
createFilesOnDisk(t, fs, "/data", files)
chk, err := NewChecker("/data/index.mf", "/data", fs)
require.NoError(t, err)
@@ -578,13 +505,10 @@ func TestFindExtraFilesSkipsManifestFile(t *testing.T) {
extras = append(extras, r)
}
assert.Empty(t, extras,
"manifest file should not be reported as extra; got: %v", extras)
assert.Empty(t, extras, "manifest file should not be reported as extra; got: %v", extras)
}
func TestCheckEmptyManifest(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
// Create manifest with no files
createTestManifest(t, fs, "/manifest.mf", map[string][]byte{})
@@ -603,26 +527,21 @@ func TestCheckEmptyManifest(t *testing.T) {
for range results {
count++
}
assert.Equal(t, 0, count)
}
func TestCheckProgressRateLimited(t *testing.T) {
t.Parallel()
// Create many small files - progress should be rate-limited, not one per file.
// With rate-limiting to once per second, we should get far fewer progress
// updates than files (plus one final update).
fs := afero.NewMemMapFs()
files := make(map[string][]byte, 100)
for i := range 100 {
for i := 0; i < 100; i++ {
name := fmt.Sprintf("file%03d.txt", i)
files[name] = []byte("content")
}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, files)
createFilesOnDisk(t, fs, "/data", files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err)
@@ -632,7 +551,9 @@ func TestCheckProgressRateLimited(t *testing.T) {
err = chk.Check(context.Background(), results, progress)
require.NoError(t, err)
// results is fully buffered and closed; no draining needed
// Drain results
for range results {
}
// Count progress updates
var progressCount int
@@ -642,8 +563,6 @@ func TestCheckProgressRateLimited(t *testing.T) {
// Should be far fewer than 100 (rate-limited to once per second)
// At minimum we get the final update
assert.GreaterOrEqual(t, progressCount, 1,
"should get at least the final progress update")
assert.Less(t, progressCount, 100,
"progress should be rate-limited, not one per file")
assert.GreaterOrEqual(t, progressCount, 1, "should get at least the final progress update")
assert.Less(t, progressCount, 100, "progress should be rate-limited, not one per file")
}
+1 -10
View File
@@ -1,20 +1,11 @@
package mfer
const (
// Version is the current mfer release version.
Version = "0.1.0"
// ReleaseDate is the date on which Version was released.
Version = "0.1.0"
ReleaseDate = "2025-12-17"
// MaxDecompressedSize is the maximum allowed size of decompressed manifest
// data (256 MB). This prevents decompression bombs from consuming excessive
// memory.
MaxDecompressedSize int64 = 256 * 1024 * 1024
// zstdWindowSize is the zstd window zstd.SpeedBestCompression gives mfer's writer.
zstdWindowSize = 8 << 20
// uuidLength is the length in bytes of a binary UUID.
uuidLength = 16
)
+47 -158
View File
@@ -2,7 +2,6 @@ package mfer
import (
"bytes"
"context"
"crypto/sha256"
"errors"
"fmt"
@@ -16,201 +15,105 @@ import (
"sneak.berlin/go/mfer/internal/log"
)
var (
errInvalidUUIDLength = errors.New("invalid UUID length")
errInvalidUUIDFormat = errors.New("invalid UUID format")
errUnknownVersion = errors.New("unknown version")
errUnknownCompression = errors.New("unknown compression type")
errCompressedHashWrong = errors.New("compressed data hash mismatch")
errSignatureNoPubKey = errors.New("signature present but no public key")
errDecompressedTooLarge = errors.New("decompressed data exceeds maximum allowed size")
errUUIDMismatch = errors.New("outer and inner UUID mismatch")
errInvalidFileFormat = errors.New("invalid file format")
errInvalidManifestPath = errors.New("manifest contains invalid path")
)
// validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable).
func validateUUID(data []byte) error {
if len(data) != uuidLength {
return errInvalidUUIDLength
if len(data) != 16 {
return errors.New("invalid UUID length")
}
// Try to parse as UUID to validate format
_, err := uuid.FromBytes(data)
if err != nil {
return errInvalidUUIDFormat
return errors.New("invalid UUID format")
}
return nil
}
// validateOuterHeader checks the outer message's version, compression
// type, and UUID.
func (m *manifest) validateOuterHeader() error {
if m.pbOuter.GetVersion() != MFFileOuter_VERSION_ONE {
return errUnknownVersion
func (m *manifest) deserializeInner() error {
if m.pbOuter.Version != MFFileOuter_VERSION_ONE {
return errors.New("unknown version")
}
if m.pbOuter.GetCompressionType() != MFFileOuter_COMPRESSION_ZSTD {
return errUnknownCompression
if m.pbOuter.CompressionType != MFFileOuter_COMPRESSION_ZSTD {
return errors.New("unknown compression type")
}
// Validate outer UUID before any decompression
err := validateUUID(m.pbOuter.GetUuid())
if err != nil {
return fmt.Errorf("outer UUID invalid: %w", err)
if err := validateUUID(m.pbOuter.Uuid); err != nil {
return errors.New("outer UUID invalid: " + err.Error())
}
return nil
}
// verifyOuterIntegrity checks the hash of the compressed payload and,
// if a signature is present, verifies it against the embedded public key.
func (m *manifest) verifyOuterIntegrity() error {
// Verify hash of compressed data before decompression
h := sha256.New()
_, err := h.Write(m.pbOuter.GetInnerMessage())
if err != nil {
if _, err := h.Write(m.pbOuter.InnerMessage); err != nil {
return fmt.Errorf("deserialize: hash write: %w", err)
}
sha256Hash := h.Sum(nil)
if !bytes.Equal(sha256Hash, m.pbOuter.GetSha256()) {
return errCompressedHashWrong
if !bytes.Equal(sha256Hash, m.pbOuter.Sha256) {
return errors.New("compressed data hash mismatch")
}
if len(m.pbOuter.GetSignature()) == 0 {
return nil
// Verify signature if present
if len(m.pbOuter.Signature) > 0 {
if len(m.pbOuter.SigningPubKey) == 0 {
return errors.New("signature present but no public key")
}
sigString, err := m.signatureString()
if err != nil {
return fmt.Errorf("failed to generate signature string for verification: %w", err)
}
if err := gpgVerify([]byte(sigString), m.pbOuter.Signature, m.pbOuter.SigningPubKey); err != nil {
return fmt.Errorf("signature verification failed: %w", err)
}
log.Infof("signature verified successfully")
}
if len(m.pbOuter.GetSigningPubKey()) == 0 {
return errSignatureNoPubKey
}
bb := bytes.NewBuffer(m.pbOuter.InnerMessage)
sigString, err := m.signatureString()
zr, err := zstd.NewReader(bb)
if err != nil {
return fmt.Errorf(
"failed to generate signature string for verification: %w", err,
)
}
// Loading a manifest takes no context; gpgTimeout still bounds gpg.
err = gpgVerify(
context.Background(),
[]byte(sigString),
m.pbOuter.GetSignature(),
m.pbOuter.GetSigningPubKey(),
)
if err != nil {
return fmt.Errorf("signature verification failed: %w", err)
}
log.Infof("signature verified successfully")
return nil
}
// decompressInner decompresses the inner payload, enforcing size limits
// to prevent decompression bombs.
func (m *manifest) decompressInner() ([]byte, error) {
bb := bytes.NewBuffer(m.pbOuter.GetInnerMessage())
// By default the decoder decodes a payload under 128 KiB in full,
// each frame up to the decoder's limit, before the LimitReader below
// reads any of it. Decoding synchronously and never in full makes it
// decode only what the LimitReader asks for. It also sets aside a new
// buffer for each frame that asks for a larger window than the frames
// before it, even a frame holding no data. Refusing windows above
// zstdWindowSize keeps each buffer to a little over zstdWindowSize, and
// the buffers of frames holding no data to about 16 times it in total.
zr, err := zstd.NewReader(bb,
zstd.WithDecoderConcurrency(1),
zstd.WithDecodeBuffersBelow(0),
zstd.WithDecoderMaxWindow(zstdWindowSize))
if err != nil {
return nil, fmt.Errorf("deserialize: zstd reader: %w", err)
return fmt.Errorf("deserialize: zstd reader: %w", err)
}
defer zr.Close()
// Limit decompressed size to prevent decompression bombs.
// Use declared size + 1 byte to detect overflow, capped at MaxDecompressedSize.
maxSize := MaxDecompressedSize
if m.pbOuter.GetSize() > 0 && m.pbOuter.GetSize() < maxSize {
maxSize = m.pbOuter.GetSize() + 1
if m.pbOuter.Size > 0 && m.pbOuter.Size < int64(maxSize) {
maxSize = int64(m.pbOuter.Size) + 1
}
limitedReader := io.LimitReader(zr, maxSize)
dat, err := io.ReadAll(limitedReader)
if err != nil {
return nil, fmt.Errorf("deserialize: decompress: %w", err)
return fmt.Errorf("deserialize: decompress: %w", err)
}
if int64(len(dat)) >= MaxDecompressedSize {
return nil, fmt.Errorf(
"%w of %d bytes", errDecompressedTooLarge, MaxDecompressedSize,
)
}
return dat, nil
}
func (m *manifest) deserializeInner() error {
err := m.validateOuterHeader()
if err != nil {
return err
}
err = m.verifyOuterIntegrity()
if err != nil {
return err
}
dat, err := m.decompressInner()
if err != nil {
return err
return fmt.Errorf("decompressed data exceeds maximum allowed size of %d bytes", MaxDecompressedSize)
}
isize := len(dat)
if int64(isize) != m.pbOuter.GetSize() {
log.Debugf("truncated data, got %d expected %d", isize, m.pbOuter.GetSize())
if int64(isize) != m.pbOuter.Size {
log.Debugf("truncated data, got %d expected %d", isize, m.pbOuter.Size)
return bork.ErrFileTruncated
}
// Deserialize inner message
m.pbInner = new(MFFile)
err = proto.Unmarshal(dat, m.pbInner)
if err != nil {
if err := proto.Unmarshal(dat, m.pbInner); err != nil {
return fmt.Errorf("deserialize: unmarshal inner: %w", err)
}
// Validate inner UUID
err = validateUUID(m.pbInner.GetUuid())
if err != nil {
return fmt.Errorf("inner UUID invalid: %w", err)
if err := validateUUID(m.pbInner.Uuid); err != nil {
return errors.New("inner UUID invalid: " + err.Error())
}
// Verify UUIDs match
if !bytes.Equal(m.pbOuter.GetUuid(), m.pbInner.GetUuid()) {
return errUUIDMismatch
if !bytes.Equal(m.pbOuter.Uuid, m.pbInner.Uuid) {
return errors.New("outer and inner UUID mismatch")
}
// Enforce the manifest path invariants on every entry as it is loaded,
// so that no consumer of a manifest — Checker today, any restore or
// extract path tomorrow — acts on a traversal or absolute path from an
// untrusted .mf. Reject loudly on the first offender rather than
// dropping entries, which would let a hostile manifest hide files from a
// check.
for _, f := range m.pbInner.GetFiles() {
err = ValidatePath(f.GetPath())
if err != nil {
return fmt.Errorf("%w: %w", errInvalidManifestPath, err)
}
}
log.Infof("loaded manifest with %d files", len(m.pbInner.GetFiles()))
log.Infof("loaded manifest with %d files", len(m.pbInner.Files))
return nil
}
@@ -219,26 +122,20 @@ func validateMagic(dat []byte) bool {
if len(dat) < ml {
return false
}
got := dat[0:ml]
expected := []byte(MAGIC)
return bytes.Equal(got, expected)
}
// NewManifestFromReader reads a manifest from an io.Reader.
//
//nolint:revive // unexported-return: exporting manifest is owner question 13
func NewManifestFromReader(input io.Reader) (*manifest, error) {
m := &manifest{}
dat, err := io.ReadAll(input)
if err != nil {
return nil, err
}
if !validateMagic(dat) {
return nil, errInvalidFileFormat
return nil, errors.New("invalid file format")
}
// remove magic bytes prefix:
@@ -248,15 +145,12 @@ func NewManifestFromReader(input io.Reader) (*manifest, error) {
// deserialize outer:
m.pbOuter = new(MFFileOuter)
err = proto.Unmarshal(dat, m.pbOuter)
if err != nil {
if err := proto.Unmarshal(dat, m.pbOuter); err != nil {
return nil, err
}
// deserialize inner:
err = m.deserializeInner()
if err != nil {
if err := m.deserializeInner(); err != nil {
return nil, err
}
@@ -265,19 +159,14 @@ func NewManifestFromReader(input io.Reader) (*manifest, error) {
// NewManifestFromFile reads a manifest from a file path using the given filesystem.
// If fs is nil, the real filesystem (OsFs) is used.
//
//nolint:revive // unexported-return: exporting manifest is owner question 13
func NewManifestFromFile(fs afero.Fs, path string) (*manifest, error) {
if fs == nil {
fs = afero.NewOsFs()
}
f, err := fs.Open(path)
if err != nil {
return nil, err
}
defer func() { _ = f.Close() }()
return NewManifestFromReader(f)
}
-83
View File
@@ -1,83 +0,0 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"bytes"
"runtime"
"testing"
"google.golang.org/protobuf/proto"
)
// FuzzNewManifestFromReader feeds arbitrary bytes to the manifest parser.
// `make test` runs it on the seed corpus in
// testdata/fuzz/FuzzNewManifestFromReader; `make fuzz` searches for new
// inputs.
//
// For every input the parser must return a manifest or an error, not both
// and not neither, and must not allocate more than a fixed multiple of its
// input and of the decompressed data it may read, plus room for the
// decoder's window buffers. A panic or a hang fails the test on its own.
func FuzzNewManifestFromReader(f *testing.F) {
// A signed manifest makes the parser write the key and signature to a
// temporary directory and run gpg on them. With gpg off the PATH and
// temporary files kept in the test's own directory, no process is
// started and nothing is written elsewhere; such input ends in an
// error instead.
f.Setenv("PATH", "")
f.Setenv("TMPDIR", f.TempDir())
f.Fuzz(func(t *testing.T, data []byte) {
var before, after runtime.MemStats
runtime.ReadMemStats(&before)
m, err := NewManifestFromReader(bytes.NewReader(data))
runtime.ReadMemStats(&after)
if (m == nil) == (err == nil) {
t.Fatalf("got manifest %p and error %v, want exactly one", m, err)
}
// The parser reads at most the declared size plus one byte of
// decompressed data, and never more than MaxDecompressedSize.
decompressed := uint64(MaxDecompressedSize)
outer := new(MFFileOuter)
if validateMagic(data) &&
proto.Unmarshal(data[len(MAGIC):], outer) == nil {
size := outer.GetSize()
if size > 0 && size < MaxDecompressedSize {
decompressed = uint64(size) + 1
}
}
// It also keeps a few copies of its input. Buffers grow by
// copying, so reaching those sizes allocates a few times them in
// total: sixteen times the input and the decompressed data leaves
// room for that.
//
// The decoder also sets aside a new buffer of one to two times the
// window for each frame that asks for a larger window than the
// frames before it, and refuses windows above zstdWindowSize. A
// frame that gives its content size instead of a window has that
// size as its window, refused above zstdWindowSize like any other.
// Frames asking for every window size up to that make it set aside
// about 16 times zstdWindowSize in total; 24 times leaves room.
//
// The seed whose frame claims 8 GiB fails if the decoder sets that
// size aside; the seed whose frames ask for ever larger windows
// fails if the decoder accepts windows of twice zstdWindowSize; the
// seed whose two frames together exceed MaxDecompressedSize fails
// if the decoder decodes them in full instead of stopping at the
// declared size.
limit := 16*(uint64(len(data))+decompressed) + 24*zstdWindowSize
allocated := after.TotalAlloc - before.TotalAlloc
if allocated > limit {
t.Fatalf("allocated %d bytes for %d bytes of input, limit %d",
allocated, len(data), limit)
}
})
}
-146
View File
@@ -1,146 +0,0 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"bytes"
"context"
"crypto/sha256"
"fmt"
"testing"
"github.com/google/uuid"
"github.com/klauspost/compress/zstd"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"google.golang.org/protobuf/encoding/protowire"
"google.golang.org/protobuf/proto"
)
// craftInnerBytes builds the wire bytes of an inner MFFile holding a single
// file entry whose path is exactly pathBytes. It writes the wire form by hand
// so a hostile path — including one that is not valid UTF-8 — can be embedded
// without proto.Marshal's own UTF-8 enforcement rejecting it first.
func craftInnerBytes(id uuid.UUID, pathBytes string) []byte {
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
entry = protowire.AppendString(entry, pathBytes)
inner := protowire.AppendTag(nil, 100, protowire.VarintType) // MFFile.version
inner = protowire.AppendVarint(inner, uint64(MFFile_VERSION_ONE))
inner = protowire.AppendTag(inner, 101, protowire.BytesType) // MFFile.files
inner = protowire.AppendBytes(inner, entry)
inner = protowire.AppendTag(inner, 102, protowire.BytesType) // MFFile.uuid
inner = protowire.AppendBytes(inner, id[:])
return inner
}
// wrapInner wraps inner MFFile wire bytes in a complete, well-formed .mf
// envelope (magic prefix, zstd-compressed payload, matching hash and UUID) so
// that deserialization reaches path validation rather than failing earlier on
// an integrity check.
func wrapInner(t *testing.T, id uuid.UUID, innerData []byte) []byte {
t.Helper()
var cbuf bytes.Buffer
zw, err := zstd.NewWriter(&cbuf, zstd.WithEncoderLevel(zstd.SpeedBestCompression))
require.NoError(t, err)
_, err = zw.Write(innerData)
require.NoError(t, err)
require.NoError(t, zw.Close())
compressed := cbuf.Bytes()
sum := sha256.Sum256(compressed)
outer := &MFFileOuter{
InnerMessage: compressed,
Size: int64(len(innerData)),
Sha256: sum[:],
Uuid: id[:],
Version: MFFileOuter_VERSION_ONE,
CompressionType: MFFileOuter_COMPRESSION_ZSTD,
}
ob, err := proto.Marshal(outer)
require.NoError(t, err)
return append([]byte(MAGIC), ob...)
}
func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
t.Parallel()
tests := []struct {
name string
path string
}{
{"parent traversal", "../escape"},
{"interior traversal", "a/../../escape"},
{"absolute path", "/etc/passwd"},
{"backslash path", `a\b`},
{"double slash", "a//b"},
{"empty path", ""},
{"invalid utf-8", "abc\xff"},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
id := uuid.New()
data := wrapInner(t, id, craftInnerBytes(id, tt.path))
_, err := NewManifestFromReader(bytes.NewReader(data))
require.Error(t, err)
if tt.path == "abc\xff" {
// A path that is not valid UTF-8 cannot survive the proto3
// string decoder, which rejects it before path validation
// runs; the manifest is still refused at load time.
return
}
require.ErrorIs(t, err, errInvalidManifestPath)
if tt.path != "" {
// ValidatePath quotes the path with %q; assert against the
// same rendering so escaped characters (e.g. a backslash)
// still match.
assert.Contains(t, err.Error(), fmt.Sprintf("%q", tt.path),
"error must name the offending path")
}
})
}
}
func TestDeserializeValidManifestRoundTrips(t *testing.T) {
t.Parallel()
hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256
b := NewBuilder()
require.NoError(t, b.AddFileWithHash("dir/file.txt", 123, ModTime{}, hash))
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(bytes.NewReader(buf.Bytes()))
require.NoError(t, err)
files := m.Files()
require.Len(t, files, 1)
assert.Equal(t, "dir/file.txt", files[0].GetPath())
assert.Equal(t, int64(123), files[0].GetSize())
}
// TestValidatePathRejectsInvalidUTF8 pins the ValidatePath rule that a manifest
// path must be valid UTF-8, independent of the proto decoder that also enforces
// it on the wire.
func TestValidatePathRejectsInvalidUTF8(t *testing.T) {
t.Parallel()
err := ValidatePath("abc\xff")
require.ErrorIs(t, err, errPathNotUTF8)
assert.Contains(t, err.Error(), "UTF-8")
}
-87
View File
@@ -1,87 +0,0 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"context"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// TestValidatePathMessagesVerbatim pins the exact rendered text of every
// ValidatePath rejection.
//
// These strings are user-visible and are assembled by wrapping static
// sentinels mid-sentence, which makes them easy to reword by accident
// while refactoring for errors.Is matchability. Changing one is a
// deliberate change, not a refactoring side effect.
func TestValidatePathMessagesVerbatim(t *testing.T) {
t.Parallel()
for _, tc := range []struct {
name string
path string
want string
is error
}{
{
name: "empty",
path: "",
want: "path cannot be empty",
is: errPathEmpty,
},
{
name: "not utf8",
path: "a\xffb",
want: `path "a\xffb" is not valid UTF-8`,
is: errPathNotUTF8,
},
{
name: "backslash",
path: `a\b`,
want: `path "a\\b" contains backslash; ` +
"use forward slashes only",
is: errPathBackslash,
},
{
name: "absolute",
path: "/a/b",
want: `path "/a/b" is absolute; must be relative`,
is: errPathAbsolute,
},
{
name: "empty segment",
path: "a//b",
want: `path "a//b" contains empty segment`,
is: errPathEmptySegment,
},
{
name: "dotdot segment",
path: "a/../b",
want: `path "a/../b" contains '..' segment`,
is: errPathDotDot,
},
} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
err := ValidatePath(tc.path)
require.Error(t, err)
assert.Equal(t, tc.want, err.Error())
require.ErrorIs(t, err, tc.is)
})
}
}
// TestSerializeInternalErrorMessagesVerbatim pins the two distinct
// "internal error" messages, which differ between generate and
// generateOuter and have always done so.
func TestSerializeInternalErrorMessagesVerbatim(t *testing.T) {
t.Parallel()
m := &manifest{}
require.EqualError(t, m.generate(context.Background()),
"internal error: pbInner not set")
require.EqualError(t, m.generateOuter(context.Background()), "internal error")
}
+89 -186
View File
@@ -2,55 +2,11 @@ package mfer
import (
"bytes"
"context"
"errors"
"fmt"
"io"
"os"
"os/exec"
"path/filepath"
"strings"
"time"
)
const (
// gpgTimeout bounds every gpg run, which can otherwise wait forever on
// a passphrase prompt or a stalled gpg-agent. A minute leaves a person
// time to type a passphrase or touch a smartcard.
gpgTimeout = time.Minute
// gpgWaitDelay is how long a gpg run keeps waiting for gpg's stdout
// and stderr to close once gpg has been killed or has exited. Reading
// what gpg itself wrote takes far less; only a process gpg left behind
// holds them open longer.
gpgWaitDelay = time.Second
// privateDirPerms is the permission mode for temporary GPG home
// directories.
privateDirPerms os.FileMode = 0o700
// privateFilePerms is the permission mode for temporary key,
// signature, and data files.
privateFilePerms os.FileMode = 0o600
// gpgFingerprintField is the record type tag for fingerprint lines
// in gpg --with-colons output.
gpgFingerprintField = "fpr"
// gpgFingerprintMinFields is the minimum number of colon-separated
// fields in a gpg fingerprint record (the fingerprint is field 10).
gpgFingerprintMinFields = 10
// gpg option names used from more than one call site.
gpgOptArmor = "--armor"
gpgOptHomedir = "--homedir"
gpgOptVerify = "--verify"
)
var (
errGPGKeyNotFound = errors.New("gpg key not found")
errFingerprintNotFound = errors.New("fingerprint not found for key")
errImportedFPRNotFound = errors.New("fingerprint not found in imported key")
)
// GPGKeyID represents a GPG key identifier (fingerprint or key ID).
@@ -61,90 +17,22 @@ type SigningOptions struct {
KeyID GPGKeyID
}
// gpgArgs builds a gpg argument list from opts followed by positional
// arguments, separated by an explicit "--" end-of-options marker.
//
// This matters because key IDs reach gpg as bare positional arguments
// (from --sign-key / MFER_SIGN_KEY) and gpg would otherwise parse a value
// beginning with "-" as one of its own options. Callers must route every
// non-option argument through here.
func gpgArgs(opts []string, positional ...string) []string {
args := make([]string, 0, len(opts)+1+len(positional))
args = append(args, opts...)
args = append(args, "--")
args = append(args, positional...)
// gpgSign creates a detached signature of the data using the specified key.
// Returns the armored detached signature.
func gpgSign(data []byte, keyID GPGKeyID) ([]byte, error) {
cmd := exec.Command("gpg", "--batch", "--no-tty",
"--detach-sign",
"--armor",
"--local-user", string(keyID),
)
return args
}
// runGPG runs the gpg binary in batch mode with the given arguments and
// optional stdin, returning captured stdout and stderr. gpg is killed when
// ctx ends or gpgTimeout passes, whichever comes first.
func runGPG(
ctx context.Context, stdin io.Reader, args ...string,
) (*bytes.Buffer, *bytes.Buffer, error) {
// exec.CommandContext kills only gpg itself. A gpg-agent that gpg
// starts runs detached and holds none of gpg's output, but another
// process gpg leaves behind (a wrapper script that runs the real gpg
// without exec, for example) can keep gpg's stdout or stderr open, and
// Run would wait for it to exit. WaitDelay stops that wait
// gpgWaitDelay after the kill; that process is left running.
ctx, cancel := context.WithTimeout(ctx, gpgTimeout)
defer cancel()
fullArgs := append([]string{"--batch", "--no-tty"}, args...)
// G204: the executable name is a compile-time constant. The arguments
// are not, so the guarantee that matters is placement: every
// caller-supplied value is passed either as the value of a named
// option or after the "--" end-of-options marker inserted by gpgArgs,
// and therefore cannot be reinterpreted by gpg as an option.
cmd := exec.CommandContext( //nolint:gosec // G204: see comment above
ctx, "gpg", fullArgs...)
cmd.WaitDelay = gpgWaitDelay
cmd.Stdin = stdin
cmd.Stdin = bytes.NewReader(data)
var stdout, stderr bytes.Buffer
cmd.Stdout = &stdout
cmd.Stderr = &stderr
err := cmd.Run()
if err != nil && ctx.Err() != nil {
// gpg was killed because ctx ended, which Run reports only as
// "signal: killed"; return the reason instead.
err = ctx.Err()
if errors.Is(err, context.DeadlineExceeded) {
err = fmt.Errorf("gpg timed out: %w", err)
}
}
return &stdout, &stderr, err
}
// parseFingerprint extracts the first fingerprint from gpg --with-colons
// output, or returns ok=false if none is present.
func parseFingerprint(colonOutput string) (string, bool) {
for _, line := range strings.Split(colonOutput, "\n") {
fields := strings.Split(line, ":")
if len(fields) >= gpgFingerprintMinFields &&
fields[0] == gpgFingerprintField {
return fields[9], true
}
}
return "", false
}
// gpgSign creates a detached signature of the data using the specified key.
// Returns the armored detached signature.
func gpgSign(ctx context.Context, data []byte, keyID GPGKeyID) ([]byte, error) {
stdout, stderr, err := runGPG(ctx, bytes.NewReader(data),
"--detach-sign",
gpgOptArmor,
"--local-user", string(keyID),
)
if err != nil {
if err := cmd.Run(); err != nil {
return nil, fmt.Errorf("gpg sign failed: %w: %s", err, stderr.String())
}
@@ -153,156 +41,171 @@ func gpgSign(ctx context.Context, data []byte, keyID GPGKeyID) ([]byte, error) {
// gpgExportPublicKey exports the public key for the specified key ID.
// Returns the armored public key.
func gpgExportPublicKey(ctx context.Context, keyID GPGKeyID) ([]byte, error) {
stdout, stderr, err := runGPG(ctx, nil,
gpgArgs([]string{"--export", gpgOptArmor}, string(keyID))...,
func gpgExportPublicKey(keyID GPGKeyID) ([]byte, error) {
cmd := exec.Command("gpg", "--batch", "--no-tty",
"--export",
"--armor",
string(keyID),
)
if err != nil {
var stdout, stderr bytes.Buffer
cmd.Stdout = &stdout
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
return nil, fmt.Errorf("gpg export failed: %w: %s", err, stderr.String())
}
if stdout.Len() == 0 {
return nil, fmt.Errorf("%w: %s", errGPGKeyNotFound, keyID)
return nil, fmt.Errorf("gpg key not found: %s", keyID)
}
return stdout.Bytes(), nil
}
// gpgGetKeyFingerprint gets the full fingerprint for a key ID.
func gpgGetKeyFingerprint(ctx context.Context, keyID GPGKeyID) ([]byte, error) {
stdout, stderr, err := runGPG(ctx, nil,
gpgArgs([]string{"--with-colons", "--fingerprint"}, string(keyID))...,
func gpgGetKeyFingerprint(keyID GPGKeyID) ([]byte, error) {
cmd := exec.Command("gpg", "--batch", "--no-tty",
"--with-colons",
"--fingerprint",
string(keyID),
)
if err != nil {
return nil, fmt.Errorf(
"gpg fingerprint lookup failed: %w: %s", err, stderr.String(),
)
var stdout, stderr bytes.Buffer
cmd.Stdout = &stdout
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
return nil, fmt.Errorf("gpg fingerprint lookup failed: %w: %s", err, stderr.String())
}
fpr, ok := parseFingerprint(stdout.String())
if !ok {
return nil, fmt.Errorf("%w: %s", errFingerprintNotFound, keyID)
// Parse the colon-delimited output to find the fingerprint
lines := strings.Split(stdout.String(), "\n")
for _, line := range lines {
fields := strings.Split(line, ":")
if len(fields) >= 10 && fields[0] == "fpr" {
return []byte(fields[9]), nil
}
}
return []byte(fpr), nil
return nil, fmt.Errorf("fingerprint not found for key: %s", keyID)
}
// gpgExtractPubKeyFingerprint imports a public key into a temporary keyring
// and extracts its fingerprint. This verifies the key is valid and returns
// the actual fingerprint from the key material.
func gpgExtractPubKeyFingerprint(ctx context.Context, pubKey []byte) (string, error) {
func gpgExtractPubKeyFingerprint(pubKey []byte) (string, error) {
// Create temporary directory for GPG operations
tmpDir, err := os.MkdirTemp("", "mfer-gpg-fingerprint-*")
if err != nil {
return "", fmt.Errorf("failed to create temp dir: %w", err)
}
defer func() { _ = os.RemoveAll(tmpDir) }()
// Set restrictive permissions
err = os.Chmod(tmpDir, privateDirPerms)
if err != nil {
if err := os.Chmod(tmpDir, 0o700); err != nil {
return "", fmt.Errorf("failed to set temp dir permissions: %w", err)
}
// Write public key to temp file
pubKeyFile := filepath.Join(tmpDir, "pubkey.asc")
err = os.WriteFile(pubKeyFile, pubKey, privateFilePerms)
if err != nil {
if err := os.WriteFile(pubKeyFile, pubKey, 0o600); err != nil {
return "", fmt.Errorf("failed to write public key: %w", err)
}
// Import the public key into the temporary keyring
_, importStderr, err := runGPG(ctx, nil,
gpgArgs([]string{gpgOptHomedir, tmpDir, "--import"}, pubKeyFile)...,
importCmd := exec.Command("gpg", "--batch", "--no-tty",
"--homedir", tmpDir,
"--import",
pubKeyFile,
)
if err != nil {
return "", fmt.Errorf(
"failed to import public key: %w: %s", err, importStderr.String(),
)
var importStderr bytes.Buffer
importCmd.Stderr = &importStderr
if err := importCmd.Run(); err != nil {
return "", fmt.Errorf("failed to import public key: %w: %s", err, importStderr.String())
}
// List keys to get fingerprint
listStdout, listStderr, err := runGPG(ctx, nil,
listCmd := exec.Command("gpg", "--batch", "--no-tty",
"--homedir", tmpDir,
"--with-colons",
"--fingerprint",
)
if err != nil {
return "", fmt.Errorf(
"failed to list keys: %w: %s", err, listStderr.String(),
)
var listStdout, listStderr bytes.Buffer
listCmd.Stdout = &listStdout
listCmd.Stderr = &listStderr
if err := listCmd.Run(); err != nil {
return "", fmt.Errorf("failed to list keys: %w: %s", err, listStderr.String())
}
fpr, ok := parseFingerprint(listStdout.String())
if !ok {
return "", errImportedFPRNotFound
// Parse the colon-delimited output to find the fingerprint
lines := strings.Split(listStdout.String(), "\n")
for _, line := range lines {
fields := strings.Split(line, ":")
if len(fields) >= 10 && fields[0] == "fpr" {
return fields[9], nil
}
}
return fpr, nil
return "", fmt.Errorf("fingerprint not found in imported key")
}
// gpgVerify verifies a detached signature against data using the provided public key.
// It creates a temporary keyring to import the public key for verification.
func gpgVerify(ctx context.Context, data, signature, pubKey []byte) error {
func gpgVerify(data, signature, pubKey []byte) error {
// Create temporary directory for GPG operations
tmpDir, err := os.MkdirTemp("", "mfer-gpg-verify-*")
if err != nil {
return fmt.Errorf("failed to create temp dir: %w", err)
}
defer func() { _ = os.RemoveAll(tmpDir) }()
// Set restrictive permissions
err = os.Chmod(tmpDir, privateDirPerms)
if err != nil {
if err := os.Chmod(tmpDir, 0o700); err != nil {
return fmt.Errorf("failed to set temp dir permissions: %w", err)
}
// Write public key to temp file
pubKeyFile := filepath.Join(tmpDir, "pubkey.asc")
err = os.WriteFile(pubKeyFile, pubKey, privateFilePerms)
if err != nil {
if err := os.WriteFile(pubKeyFile, pubKey, 0o600); err != nil {
return fmt.Errorf("failed to write public key: %w", err)
}
// Write signature to temp file
sigFile := filepath.Join(tmpDir, "signature.asc")
err = os.WriteFile(sigFile, signature, privateFilePerms)
if err != nil {
if err := os.WriteFile(sigFile, signature, 0o600); err != nil {
return fmt.Errorf("failed to write signature: %w", err)
}
// Write data to temp file
dataFile := filepath.Join(tmpDir, "data")
err = os.WriteFile(dataFile, data, privateFilePerms)
if err != nil {
if err := os.WriteFile(dataFile, data, 0o600); err != nil {
return fmt.Errorf("failed to write data: %w", err)
}
// Import the public key into the temporary keyring
_, importStderr, err := runGPG(ctx, nil,
gpgArgs([]string{gpgOptHomedir, tmpDir, "--import"}, pubKeyFile)...,
importCmd := exec.Command("gpg", "--batch", "--no-tty",
"--homedir", tmpDir,
"--import",
pubKeyFile,
)
if err != nil {
return fmt.Errorf(
"failed to import public key: %w: %s", err, importStderr.String(),
)
var importStderr bytes.Buffer
importCmd.Stderr = &importStderr
if err := importCmd.Run(); err != nil {
return fmt.Errorf("failed to import public key: %w: %s", err, importStderr.String())
}
// Verify the signature
_, verifyStderr, err := runGPG(ctx, nil,
gpgArgs([]string{gpgOptHomedir, tmpDir, gpgOptVerify},
sigFile, dataFile)...,
verifyCmd := exec.Command("gpg", "--batch", "--no-tty",
"--homedir", tmpDir,
"--verify",
sigFile,
dataFile,
)
if err != nil {
return fmt.Errorf(
"signature verification failed: %w: %s", err, verifyStderr.String(),
)
var verifyStderr bytes.Buffer
verifyCmd.Stderr = &verifyStderr
if err := verifyCmd.Run(); err != nil {
return fmt.Errorf("signature verification failed: %w: %s", err, verifyStderr.String())
}
return nil
+83 -224
View File
@@ -1,18 +1,13 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"bytes"
"context"
"io"
"os"
"os/exec"
"path/filepath"
"strconv"
"strings"
"syscall"
"testing"
"time"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
@@ -20,20 +15,35 @@ import (
)
// testGPGEnv sets up a temporary GPG home directory with a test key.
// Returns the key ID and the GPG home directory; callers must point
// GNUPGHOME at the returned directory (via t.Setenv) before using the
// gpg helpers under test.
func testGPGEnv(t *testing.T) (GPGKeyID, string) {
// Returns the key ID and a cleanup function.
func testGPGEnv(t *testing.T) (GPGKeyID, func()) {
t.Helper()
// Check if gpg is installed
_, err := exec.LookPath("gpg")
if err != nil {
if _, err := exec.LookPath("gpg"); err != nil {
t.Skip("gpg not installed, skipping signing test")
return "", func() {}
}
// Create temporary GPG home directory (0700 by default)
gpgHome := t.TempDir()
// Create temporary GPG home directory
gpgHome, err := os.MkdirTemp("", "mfer-gpg-test-*")
require.NoError(t, err)
// Set restrictive permissions on GPG home
require.NoError(t, os.Chmod(gpgHome, 0o700))
// Save original GNUPGHOME and set new one
origGPGHome := os.Getenv("GNUPGHOME")
require.NoError(t, os.Setenv("GNUPGHOME", gpgHome))
cleanup := func() {
if origGPGHome == "" {
_ = os.Unsetenv("GNUPGHOME")
} else {
_ = os.Setenv("GNUPGHOME", origGPGHome)
}
_ = os.RemoveAll(gpgHome)
}
// Generate a test key with no passphrase
keyParams := `%no-protection
@@ -47,57 +57,48 @@ Expire-Date: 0
paramsFile := filepath.Join(gpgHome, "key-params")
require.NoError(t, os.WriteFile(paramsFile, []byte(keyParams), 0o600))
ctx, cancel := context.WithTimeout(context.Background(), gpgTimeout)
defer cancel()
//nolint:gosec // paramsFile is a test-controlled path inside t.TempDir()
cmd := exec.CommandContext(ctx, "gpg",
"--batch", "--gen-key", paramsFile)
cmd := exec.Command("gpg", "--batch", "--gen-key", paramsFile)
cmd.Env = append(os.Environ(), "GNUPGHOME="+gpgHome)
output, err := cmd.CombinedOutput()
if err != nil {
cleanup()
t.Skipf("failed to generate test GPG key: %v: %s", err, output)
return "", func() {}
}
// Get the key fingerprint
cmd = exec.CommandContext(ctx, "gpg",
"--list-keys", "--with-colons", "test@mfer.test")
cmd = exec.Command("gpg", "--list-keys", "--with-colons", "test@mfer.test")
cmd.Env = append(os.Environ(), "GNUPGHOME="+gpgHome)
output, err = cmd.Output()
if err != nil {
cleanup()
t.Fatalf("failed to list test key: %v", err)
}
// Parse fingerprint from output
var keyID string
for _, line := range strings.Split(string(output), "\n") {
fields := strings.Split(line, ":")
if len(fields) >= gpgFingerprintMinFields &&
fields[0] == gpgFingerprintField {
if len(fields) >= 10 && fields[0] == "fpr" {
keyID = fields[9]
break
}
}
if keyID == "" {
cleanup()
t.Fatal("failed to find test key fingerprint")
}
return GPGKeyID(keyID), gpgHome
return GPGKeyID(keyID), cleanup
}
func TestGPGSign(t *testing.T) {
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
keyID, cleanup := testGPGEnv(t)
defer cleanup()
data := []byte("test data to sign")
sig, err := gpgSign(context.Background(), data, keyID)
sig, err := gpgSign(data, keyID)
require.NoError(t, err)
assert.NotEmpty(t, sig)
assert.Contains(t, string(sig), "-----BEGIN PGP SIGNATURE-----")
@@ -105,10 +106,10 @@ func TestGPGSign(t *testing.T) {
}
func TestGPGExportPublicKey(t *testing.T) {
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
keyID, cleanup := testGPGEnv(t)
defer cleanup()
pubKey, err := gpgExportPublicKey(context.Background(), keyID)
pubKey, err := gpgExportPublicKey(keyID)
require.NoError(t, err)
assert.NotEmpty(t, pubKey)
assert.Contains(t, string(pubKey), "-----BEGIN PGP PUBLIC KEY BLOCK-----")
@@ -116,67 +117,29 @@ func TestGPGExportPublicKey(t *testing.T) {
}
func TestGPGGetKeyFingerprint(t *testing.T) {
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
keyID, cleanup := testGPGEnv(t)
defer cleanup()
fingerprint, err := gpgGetKeyFingerprint(context.Background(), keyID)
fingerprint, err := gpgGetKeyFingerprint(keyID)
require.NoError(t, err)
assert.NotEmpty(t, fingerprint)
// The fingerprint should be 40 hex chars
assert.Len(t, fingerprint, 40, "fingerprint should be 40 hex chars")
}
// TestGPGArgsSeparatesPositionals pins that caller-supplied values are
// placed after an end-of-options marker. Key IDs arrive from --sign-key
// and MFER_SIGN_KEY as bare positional arguments, so without the marker
// a value beginning with "-" would be parsed by gpg as one of its own
// options.
func TestGPGArgsSeparatesPositionals(t *testing.T) {
t.Parallel()
assert.Equal(t,
[]string{"--opt-a", "--opt-b", "--", "--version"},
gpgArgs([]string{"--opt-a", "--opt-b"}, "--version"))
assert.Equal(t,
[]string{"--opt-c", "--", "sig", "data"},
gpgArgs([]string{"--opt-c"}, "sig", "data"))
assert.Equal(t, []string{"--opt-d", "--"},
gpgArgs([]string{"--opt-d"}))
}
// TestGPGOptionLikeKeyIDIsNotAnOption drives real gpg with a key ID that
// looks like an option and asserts it is treated as a (nonexistent) key
// rather than executed as gpg's own --version.
func TestGPGOptionLikeKeyIDIsNotAnOption(t *testing.T) {
_, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
pubKey, err := gpgExportPublicKey(context.Background(), GPGKeyID("--version"))
require.Error(t, err)
require.ErrorIs(t, err, errGPGKeyNotFound)
assert.NotContains(t, string(pubKey), "gpg (GnuPG)")
fpr, err := gpgGetKeyFingerprint(context.Background(), GPGKeyID("--version"))
require.Error(t, err)
assert.NotContains(t, string(fpr), "gpg (GnuPG)")
}
func TestGPGSignInvalidKey(t *testing.T) {
// Set up test environment (we need GNUPGHOME set)
_, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
_, cleanup := testGPGEnv(t)
defer cleanup()
data := []byte("test data")
_, err := gpgSign(context.Background(), data,
GPGKeyID("NONEXISTENT_KEY_ID_12345"))
_, err := gpgSign(data, GPGKeyID("NONEXISTENT_KEY_ID_12345"))
assert.Error(t, err)
}
func TestBuilderWithSigning(t *testing.T) {
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
keyID, cleanup := testGPGEnv(t)
defer cleanup()
// Create a builder with signing options
b := NewBuilder()
@@ -192,8 +155,7 @@ func TestBuilderWithSigning(t *testing.T) {
// Build the manifest
var buf bytes.Buffer
err = b.Build(context.Background(), &buf)
err = b.Build(&buf)
require.NoError(t, err)
// Parse the manifest and verify signature fields are populated
@@ -201,32 +163,26 @@ func TestBuilderWithSigning(t *testing.T) {
require.NoError(t, err)
require.NotNil(t, manifest.pbOuter)
assert.NotEmpty(t, manifest.pbOuter.GetSignature(),
"signature should be populated")
assert.NotEmpty(t, manifest.pbOuter.GetSigner(), "signer should be populated")
assert.NotEmpty(t, manifest.pbOuter.GetSigningPubKey(),
"signing public key should be populated")
assert.NotEmpty(t, manifest.pbOuter.Signature, "signature should be populated")
assert.NotEmpty(t, manifest.pbOuter.Signer, "signer should be populated")
assert.NotEmpty(t, manifest.pbOuter.SigningPubKey, "signing public key should be populated")
// Verify signature is a valid PGP signature
assert.Contains(t, string(manifest.pbOuter.GetSignature()),
"-----BEGIN PGP SIGNATURE-----")
assert.Contains(t, string(manifest.pbOuter.Signature), "-----BEGIN PGP SIGNATURE-----")
// Verify public key is a valid PGP public key block
assert.Contains(t, string(manifest.pbOuter.GetSigningPubKey()),
"-----BEGIN PGP PUBLIC KEY BLOCK-----")
assert.Contains(t, string(manifest.pbOuter.SigningPubKey), "-----BEGIN PGP PUBLIC KEY BLOCK-----")
}
func TestScannerWithSigning(t *testing.T) {
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
keyID, cleanup := testGPGEnv(t)
defer cleanup()
// Create in-memory filesystem with test files
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644))
// Create scanner with signing options
opts := &ScannerOptions{
@@ -249,61 +205,61 @@ func TestScannerWithSigning(t *testing.T) {
manifest, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.NotEmpty(t, manifest.pbOuter.GetSignature())
assert.NotEmpty(t, manifest.pbOuter.GetSigner())
assert.NotEmpty(t, manifest.pbOuter.GetSigningPubKey())
assert.NotEmpty(t, manifest.pbOuter.Signature)
assert.NotEmpty(t, manifest.pbOuter.Signer)
assert.NotEmpty(t, manifest.pbOuter.SigningPubKey)
}
func TestGPGVerify(t *testing.T) {
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
keyID, cleanup := testGPGEnv(t)
defer cleanup()
data := []byte("test data to sign and verify")
sig, err := gpgSign(context.Background(), data, keyID)
sig, err := gpgSign(data, keyID)
require.NoError(t, err)
pubKey, err := gpgExportPublicKey(context.Background(), keyID)
pubKey, err := gpgExportPublicKey(keyID)
require.NoError(t, err)
// Verify the signature
err = gpgVerify(context.Background(), data, sig, pubKey)
err = gpgVerify(data, sig, pubKey)
require.NoError(t, err)
}
func TestGPGVerifyInvalidSignature(t *testing.T) {
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
keyID, cleanup := testGPGEnv(t)
defer cleanup()
data := []byte("test data to sign")
sig, err := gpgSign(context.Background(), data, keyID)
sig, err := gpgSign(data, keyID)
require.NoError(t, err)
pubKey, err := gpgExportPublicKey(context.Background(), keyID)
pubKey, err := gpgExportPublicKey(keyID)
require.NoError(t, err)
// Try to verify with different data - should fail
wrongData := []byte("different data")
err = gpgVerify(context.Background(), wrongData, sig, pubKey)
err = gpgVerify(wrongData, sig, pubKey)
assert.Error(t, err)
}
func TestGPGVerifyBadPublicKey(t *testing.T) {
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
keyID, cleanup := testGPGEnv(t)
defer cleanup()
data := []byte("test data")
sig, err := gpgSign(context.Background(), data, keyID)
sig, err := gpgSign(data, keyID)
require.NoError(t, err)
// Try to verify with invalid public key - should fail
badPubKey := []byte("not a valid public key")
err = gpgVerify(context.Background(), data, sig, badPubKey)
err = gpgVerify(data, sig, badPubKey)
assert.Error(t, err)
}
func TestManifestSignatureVerification(t *testing.T) {
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
keyID, cleanup := testGPGEnv(t)
defer cleanup()
// Create a builder with signing options
b := NewBuilder()
@@ -319,8 +275,7 @@ func TestManifestSignatureVerification(t *testing.T) {
// Build the manifest
var buf bytes.Buffer
err = b.Build(context.Background(), &buf)
err = b.Build(&buf)
require.NoError(t, err)
// Parse the manifest - signature should be verified during load
@@ -329,12 +284,12 @@ func TestManifestSignatureVerification(t *testing.T) {
require.NotNil(t, manifest)
// Signature should be present and valid
assert.NotEmpty(t, manifest.pbOuter.GetSignature())
assert.NotEmpty(t, manifest.pbOuter.Signature)
}
func TestManifestTamperedSignatureFails(t *testing.T) {
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
keyID, cleanup := testGPGEnv(t)
defer cleanup()
// Create a signed manifest
b := NewBuilder()
@@ -348,8 +303,7 @@ func TestManifestTamperedSignatureFails(t *testing.T) {
require.NoError(t, err)
var buf bytes.Buffer
err = b.Build(context.Background(), &buf)
err = b.Build(&buf)
require.NoError(t, err)
// Tamper with the signature by replacing some bytes
@@ -358,7 +312,6 @@ func TestManifestTamperedSignatureFails(t *testing.T) {
for i := range data {
if i > 100 && data[i] == 'A' {
data[i] = 'B'
break
}
}
@@ -369,8 +322,6 @@ func TestManifestTamperedSignatureFails(t *testing.T) {
}
func TestBuilderWithoutSigning(t *testing.T) {
t.Parallel()
// Create a builder without signing options
b := NewBuilder()
@@ -382,8 +333,7 @@ func TestBuilderWithoutSigning(t *testing.T) {
// Build the manifest
var buf bytes.Buffer
err = b.Build(context.Background(), &buf)
err = b.Build(&buf)
require.NoError(t, err)
// Parse the manifest and verify signature fields are empty
@@ -391,98 +341,7 @@ func TestBuilderWithoutSigning(t *testing.T) {
require.NoError(t, err)
require.NotNil(t, manifest.pbOuter)
assert.Empty(t, manifest.pbOuter.GetSignature(),
"signature should be empty when not signing")
assert.Empty(t, manifest.pbOuter.GetSigner(),
"signer should be empty when not signing")
assert.Empty(t, manifest.pbOuter.GetSigningPubKey(),
"signing public key should be empty when not signing")
}
// fakeGPGPath writes script as an executable named gpg into a temporary
// directory and returns a PATH value with that directory first.
func fakeGPGPath(t *testing.T, script string) string {
t.Helper()
binDir := t.TempDir()
//nolint:gosec // G306: the fake gpg has to be executable
require.NoError(t, os.WriteFile(filepath.Join(binDir, "gpg"),
[]byte(script), 0o700))
return binDir + string(os.PathListSeparator) + os.Getenv("PATH")
}
// TestGPGTimeoutKillsGPG puts a fake gpg that never finishes first on
// PATH and checks that a run past its deadline is killed and reported as
// a timeout of the named operation, instead of hanging.
func TestGPGTimeoutKillsGPG(t *testing.T) {
t.Setenv("PATH", fakeGPGPath(t, "#!/bin/sh\nexec sleep 10\n"))
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
_, err := gpgSign(ctx, []byte("data"), GPGKeyID("any"))
require.ErrorIs(t, err, context.DeadlineExceeded)
assert.Contains(t, err.Error(), "gpg sign failed: gpg timed out")
}
// TestGPGCancelWhenChildHoldsOutput uses a fake gpg that runs sleep as a
// child instead of exec-ing it, the way a wrapper script around the real
// gpg might. Killing the fake gpg leaves sleep holding its stdout and
// stderr open; the call must still return once ctx ends instead of waiting
// for sleep to exit. The fake gpg writes the process ID of sleep to a named
// pipe; the test ends ctx only after reading it, so sleep is running by
// then, and kills sleep before returning.
func TestGPGCancelWhenChildHoldsOutput(t *testing.T) {
pidPipe := filepath.Join(t.TempDir(), "sleep.pid")
require.NoError(t, syscall.Mkfifo(pidPipe, 0o600))
// sleep outlasts the 10 s wait below, so a call that waits for it fails.
t.Setenv("PATH", fakeGPGPath(t,
"#!/bin/sh\nsleep 60 &\necho $! >'"+pidPipe+"'\nwait\n"))
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
signErr := make(chan error, 1)
go func() {
_, err := gpgSign(ctx, []byte("data"), GPGKeyID("any"))
signErr <- err
}()
pid, err := os.ReadFile(pidPipe) //nolint:gosec // G304: path inside t.TempDir()
require.NoError(t, err)
n, err := strconv.Atoi(strings.TrimSpace(string(pid)))
require.NoError(t, err)
sleep, err := os.FindProcess(n)
require.NoError(t, err)
t.Cleanup(func() { require.NoError(t, sleep.Kill()) })
cancel()
// The call should return about gpgWaitDelay (one second) after the
// cancel. 10 s is far above that and well under the 30 s test timeout,
// which would abort the whole package before the cleanup kills sleep.
select {
case err := <-signErr:
require.ErrorIs(t, err, context.Canceled)
case <-time.After(10 * time.Second):
t.Fatal("the call waited for the child holding gpg's output to exit")
}
}
// TestBuildPassesContextToSigning checks that a caller can cancel the gpg
// runs that sign a manifest through the context given to Build.
func TestBuildPassesContextToSigning(t *testing.T) {
t.Parallel()
b := NewBuilder()
b.SetSigningOptions(&SigningOptions{KeyID: "any"})
ctx, cancel := context.WithCancel(context.Background())
cancel()
require.ErrorIs(t, b.Build(ctx, io.Discard), context.Canceled)
assert.Empty(t, manifest.pbOuter.Signature, "signature should be empty when not signing")
assert.Empty(t, manifest.pbOuter.Signer, "signer should be empty when not signing")
assert.Empty(t, manifest.pbOuter.SigningPubKey, "signing public key should be empty when not signing")
}
+13 -28
View File
@@ -9,18 +9,9 @@ import (
"github.com/multiformats/go-multihash"
)
var (
errOuterNotSet = errors.New("pbOuter not set")
errUUIDNotSet = errors.New("UUID not set")
errSHA256NotSet = errors.New("SHA256 hash not set")
)
// manifest holds the internal representation of a manifest file.
// Use NewManifestFromFile or NewManifestFromReader to load an existing
// manifest, or use Builder to create a new one.
//
// Whether this type should be exported is an open design question owned by
// the repository owner; see README design question 13.
// Use NewManifestFromFile or NewManifestFromReader to load an existing manifest,
// or use Builder to create a new one.
type manifest struct {
pbInner *MFFile
pbOuter *MFFileOuter
@@ -32,9 +23,8 @@ type manifest struct {
func (m *manifest) String() string {
count := 0
if m.pbInner != nil {
count = len(m.pbInner.GetFiles())
count = len(m.pbInner.Files)
}
return fmt.Sprintf("<Manifest count=%d>", count)
}
@@ -43,8 +33,7 @@ func (m *manifest) Files() []*MFFilePath {
if m.pbInner == nil {
return nil
}
return m.pbInner.GetFiles()
return m.pbInner.Files
}
// signatureString generates the canonical string used for signing/verification.
@@ -52,24 +41,20 @@ func (m *manifest) Files() []*MFFilePath {
// Requires pbOuter to be set with Uuid and Sha256 fields.
func (m *manifest) signatureString() (string, error) {
if m.pbOuter == nil {
return "", errOuterNotSet
return "", errors.New("pbOuter not set")
}
if len(m.pbOuter.Uuid) == 0 {
return "", errors.New("UUID not set")
}
if len(m.pbOuter.Sha256) == 0 {
return "", errors.New("SHA256 hash not set")
}
if len(m.pbOuter.GetUuid()) == 0 {
return "", errUUIDNotSet
}
if len(m.pbOuter.GetSha256()) == 0 {
return "", errSHA256NotSet
}
mh, err := multihash.Encode(m.pbOuter.GetSha256(), multihash.SHA2_256)
mh, err := multihash.Encode(m.pbOuter.Sha256, multihash.SHA2_256)
if err != nil {
return "", fmt.Errorf("failed to encode multihash: %w", err)
}
uuidStr := hex.EncodeToString(m.pbOuter.GetUuid())
uuidStr := hex.EncodeToString(m.pbOuter.Uuid)
mhStr := hex.EncodeToString(mh)
return fmt.Sprintf("%s-%s-%s", MAGIC, uuidStr, mhStr), nil
}
+159 -310
View File
@@ -43,20 +43,12 @@ type ScanStatus struct {
// ScannerOptions configures scanner behavior.
type ScannerOptions struct {
// IncludeDotfiles includes files and directories starting with a dot
// (default: exclude).
IncludeDotfiles bool
// FollowSymLinks resolves symlinks instead of skipping them.
FollowSymLinks bool
// IncludeTimestamps includes a createdAt timestamp in the manifest
// (default: omit for determinism).
IncludeTimestamps bool
// Fs is the filesystem to use, defaults to OsFs if nil.
Fs afero.Fs
// SigningOptions holds GPG signing options (nil = no signing).
SigningOptions *SigningOptions
// Seed, if set, derives a deterministic UUID from this seed.
Seed string
IncludeDotfiles bool // Include files and directories starting with a dot (default: exclude)
FollowSymLinks bool // Resolve symlinks instead of skipping them
IncludeTimestamps bool // Include createdAt timestamp in manifest (default: omit for determinism)
Fs afero.Fs // Filesystem to use, defaults to OsFs if nil
SigningOptions *SigningOptions // GPG signing options (nil = no signing)
Seed string // If set, derive a deterministic UUID from this seed
}
// FileEntry represents a file that has been enumerated.
@@ -87,12 +79,10 @@ func NewScannerWithOptions(opts *ScannerOptions) *Scanner {
if opts == nil {
opts = &ScannerOptions{}
}
fs := opts.Fs
if fs == nil {
fs = afero.NewOsFs()
}
return &Scanner{
files: make([]*FileEntry, 0),
options: opts,
@@ -106,63 +96,47 @@ func (s *Scanner) EnumerateFile(filePath string) error {
if err != nil {
return err
}
info, err := s.fs.Stat(abs)
if err != nil {
return err
}
// For single files, use the filename as the relative path
basePath := filepath.Dir(abs)
return s.enumerateFileWithInfo(filepath.Base(abs), basePath, info, nil)
}
// EnumeratePath walks a directory path and adds all files to the scanner.
// If progress is non-nil, status updates are sent as files are discovered.
// The progress channel is closed when the method returns.
func (s *Scanner) EnumeratePath(
inputPath string,
progress chan<- EnumerateStatus,
) error {
func (s *Scanner) EnumeratePath(inputPath string, progress chan<- EnumerateStatus) error {
if progress != nil {
defer close(progress)
}
abs, err := filepath.Abs(inputPath)
if err != nil {
return err
}
afs := afero.NewReadOnlyFs(afero.NewBasePathFs(s.fs, abs))
return s.enumerateFS(afs, abs, progress)
}
// EnumeratePaths walks multiple directory paths and adds all files to the scanner.
// If progress is non-nil, status updates are sent as files are discovered.
// The progress channel is closed when the method returns.
func (s *Scanner) EnumeratePaths(
progress chan<- EnumerateStatus,
inputPaths ...string,
) error {
func (s *Scanner) EnumeratePaths(progress chan<- EnumerateStatus, inputPaths ...string) error {
if progress != nil {
defer close(progress)
}
for _, p := range inputPaths {
abs, err := filepath.Abs(p)
if err != nil {
return err
}
afs := afero.NewReadOnlyFs(afero.NewBasePathFs(s.fs, abs))
err = s.enumerateFS(afs, abs, progress)
if err != nil {
if err := s.enumerateFS(afs, abs, progress); err != nil {
return err
}
}
return nil
}
@@ -170,230 +144,31 @@ func (s *Scanner) EnumeratePaths(
// If progress is non-nil, status updates are sent as files are discovered.
// The progress channel is closed when the method returns.
// basePath is used to compute absolute paths for file reading.
func (s *Scanner) EnumerateFS(
afs afero.Fs,
basePath string,
progress chan<- EnumerateStatus,
) error {
func (s *Scanner) EnumerateFS(afs afero.Fs, basePath string, progress chan<- EnumerateStatus) error {
if progress != nil {
defer close(progress)
}
return s.enumerateFS(afs, basePath, progress)
}
// Files returns a copy of all files added to the scanner.
func (s *Scanner) Files() []*FileEntry {
s.mu.RLock()
defer s.mu.RUnlock()
out := make([]*FileEntry, len(s.files))
copy(out, s.files)
return out
}
// FileCount returns the number of files in the scanner.
func (s *Scanner) FileCount() FileCount {
s.mu.RLock()
defer s.mu.RUnlock()
return FileCount(len(s.files))
}
// TotalBytes returns the total size of all files in the scanner.
func (s *Scanner) TotalBytes() FileSize {
s.mu.RLock()
defer s.mu.RUnlock()
return s.totalBytes
}
// ToManifest reads all file contents, computes hashes, and generates a manifest.
// If progress is non-nil, status updates are sent approximately once per second.
// The progress channel is closed when the method returns.
// The manifest is written to the provided io.Writer.
func (s *Scanner) ToManifest(
ctx context.Context, w io.Writer, progress chan<- ScanStatus,
) error {
if progress != nil {
defer close(progress)
}
s.mu.RLock()
files := make([]*FileEntry, len(s.files))
copy(files, s.files)
totalFiles := FileCount(len(files))
var totalBytes FileSize
for _, f := range files {
totalBytes += f.Size
}
s.mu.RUnlock()
builder := s.configureBuilder()
var (
scannedFiles FileCount
scannedBytes FileSize
)
lastProgressTime := time.Now()
startTime := time.Now()
pt := &scanProgressTracker{
progress: progress,
totalFiles: totalFiles,
totalBytes: totalBytes,
startTime: startTime,
lastProgress: &lastProgressTime,
}
for _, entry := range files {
// Check for cancellation
select {
case <-ctx.Done():
return ctx.Err()
default:
}
bytesRead, err := s.scanFile(builder, pt, entry, scannedFiles, scannedBytes)
if err != nil {
return err
}
scannedFiles++
scannedBytes += bytesRead
}
// Send final progress (ETA is 0 at completion; remaining bytes are 0,
// so computeRateETA yields eta 0 and the same average rate as before)
if progress != nil {
rate, _ := computeRateETA(time.Since(startTime), scannedBytes, totalBytes)
sendScanStatus(progress, ScanStatus{
TotalFiles: totalFiles,
ScannedFiles: scannedFiles,
TotalBytes: totalBytes,
ScannedBytes: scannedBytes,
BytesPerSec: rate,
ETA: 0,
})
}
// Build and write manifest
return builder.Build(ctx, w)
}
// configureBuilder constructs a manifest builder configured from the
// scanner options.
func (s *Scanner) configureBuilder() *Builder {
builder := NewBuilder()
if s.options.IncludeTimestamps {
builder.SetIncludeTimestamps(true)
}
if s.options.SigningOptions != nil {
builder.SetSigningOptions(s.options.SigningOptions)
}
if s.options.Seed != "" {
builder.SetSeed(s.options.Seed)
}
return builder
}
// scanFile hashes a single file into the builder, forwarding per-file
// progress updates, and returns the number of bytes read.
func (s *Scanner) scanFile(
builder *Builder,
pt *scanProgressTracker,
entry *FileEntry,
scannedFiles FileCount,
scannedBytes FileSize,
) (FileSize, error) {
// Open file
f, err := s.fs.Open(string(entry.AbsPath))
if err != nil {
return 0, err
}
// Create progress channel for this file
var (
fileProgress chan FileHashProgress
wg sync.WaitGroup
)
if pt.progress != nil {
fileProgress = make(chan FileHashProgress, 1)
wg.Add(1)
go func(base FileSize, done FileCount) {
defer wg.Done()
pt.forward(fileProgress, done, base)
}(scannedBytes, scannedFiles)
}
// Add to manifest with progress channel
bytesRead, err := builder.AddFile(
entry.Path,
entry.Size,
entry.Mtime,
f,
fileProgress,
)
_ = f.Close()
// Close channel and wait for goroutine to finish
if fileProgress != nil {
close(fileProgress)
wg.Wait()
}
if err != nil {
return 0, err
}
log.Verbosef("+ %s (%s)", entry.Path, humanize.IBytes(sizeToUint64(bytesRead)))
return bytesRead, nil
}
// enumerateFS is the internal implementation that doesn't close the
// progress channel.
func (s *Scanner) enumerateFS(
afs afero.Fs,
basePath string,
progress chan<- EnumerateStatus,
) error {
// enumerateFS is the internal implementation that doesn't close the progress channel.
func (s *Scanner) enumerateFS(afs afero.Fs, basePath string, progress chan<- EnumerateStatus) error {
return afero.Walk(afs, "/", func(p string, info fs.FileInfo, err error) error {
if err != nil {
return err
}
if !s.options.IncludeDotfiles && IsHiddenPath(p) {
if info.IsDir() {
return filepath.SkipDir
}
return nil
}
return s.enumerateFileWithInfo(p, basePath, info, progress)
})
}
// enumerateFileWithInfo adds a file with pre-existing fs.FileInfo.
func (s *Scanner) enumerateFileWithInfo(
filePath string,
basePath string,
info fs.FileInfo,
progress chan<- EnumerateStatus,
) error {
func (s *Scanner) enumerateFileWithInfo(filePath string, basePath string, info fs.FileInfo, progress chan<- EnumerateStatus) error {
if info.IsDir() {
// Manifests contain only files, directories are implied
return nil
@@ -418,13 +193,11 @@ func (s *Scanner) enumerateFileWithInfo(
realPath, err := filepath.EvalSymlinks(absPath)
if err != nil {
// Skip broken symlinks
return nil //nolint:nilerr // broken symlinks are skipped by design
return nil
}
realInfo, err := s.fs.Stat(realPath)
if err != nil {
// Skip symlinks whose target cannot be stat'd
return nil //nolint:nilerr // unreadable targets are skipped by design
return nil
}
// Skip if symlink points to a directory
if realInfo.IsDir() {
@@ -459,78 +232,160 @@ func (s *Scanner) enumerateFileWithInfo(
return nil
}
// scanProgressTracker carries the shared state needed to report rate-limited
// scan progress updates.
type scanProgressTracker struct {
progress chan<- ScanStatus
totalFiles FileCount
totalBytes FileSize
startTime time.Time
lastProgress *time.Time
// Files returns a copy of all files added to the scanner.
func (s *Scanner) Files() []*FileEntry {
s.mu.RLock()
defer s.mu.RUnlock()
out := make([]*FileEntry, len(s.files))
copy(out, s.files)
return out
}
// forward relays per-file hash progress to the scan progress channel,
// rate-limited to one update per second.
func (pt *scanProgressTracker) forward(
fileProgress <-chan FileHashProgress,
scannedFiles FileCount,
baseBytes FileSize,
) {
for p := range fileProgress {
// Send progress at most once per second
now := time.Now()
if now.Sub(*pt.lastProgress) < time.Second {
continue
// FileCount returns the number of files in the scanner.
func (s *Scanner) FileCount() FileCount {
s.mu.RLock()
defer s.mu.RUnlock()
return FileCount(len(s.files))
}
// TotalBytes returns the total size of all files in the scanner.
func (s *Scanner) TotalBytes() FileSize {
s.mu.RLock()
defer s.mu.RUnlock()
return s.totalBytes
}
// ToManifest reads all file contents, computes hashes, and generates a manifest.
// If progress is non-nil, status updates are sent approximately once per second.
// The progress channel is closed when the method returns.
// The manifest is written to the provided io.Writer.
func (s *Scanner) ToManifest(ctx context.Context, w io.Writer, progress chan<- ScanStatus) error {
if progress != nil {
defer close(progress)
}
s.mu.RLock()
files := make([]*FileEntry, len(s.files))
copy(files, s.files)
totalFiles := FileCount(len(files))
var totalBytes FileSize
for _, f := range files {
totalBytes += f.Size
}
s.mu.RUnlock()
builder := NewBuilder()
if s.options.IncludeTimestamps {
builder.SetIncludeTimestamps(true)
}
if s.options.SigningOptions != nil {
builder.SetSigningOptions(s.options.SigningOptions)
}
if s.options.Seed != "" {
builder.SetSeed(s.options.Seed)
}
var scannedFiles FileCount
var scannedBytes FileSize
lastProgressTime := time.Now()
startTime := time.Now()
for _, entry := range files {
// Check for cancellation
select {
case <-ctx.Done():
return ctx.Err()
default:
}
currentBytes := baseBytes + p.BytesRead
rate, eta := computeRateETA(now.Sub(pt.startTime), currentBytes, pt.totalBytes)
// Open file
f, err := s.fs.Open(string(entry.AbsPath))
if err != nil {
return err
}
sendScanStatus(pt.progress, ScanStatus{
TotalFiles: pt.totalFiles,
// Create progress channel for this file
var fileProgress chan FileHashProgress
var wg sync.WaitGroup
if progress != nil {
fileProgress = make(chan FileHashProgress, 1)
wg.Add(1)
go func(baseScannedBytes FileSize) {
defer wg.Done()
for p := range fileProgress {
// Send progress at most once per second
now := time.Now()
if now.Sub(lastProgressTime) >= time.Second {
elapsed := now.Sub(startTime).Seconds()
currentBytes := baseScannedBytes + p.BytesRead
var rate float64
var eta time.Duration
if elapsed > 0 && currentBytes > 0 {
rate = float64(currentBytes) / elapsed
remainingBytes := totalBytes - currentBytes
if rate > 0 {
eta = time.Duration(float64(remainingBytes)/rate) * time.Second
}
}
sendScanStatus(progress, ScanStatus{
TotalFiles: totalFiles,
ScannedFiles: scannedFiles,
TotalBytes: totalBytes,
ScannedBytes: currentBytes,
BytesPerSec: rate,
ETA: eta,
})
lastProgressTime = now
}
}
}(scannedBytes)
}
// Add to manifest with progress channel
bytesRead, err := builder.AddFile(
entry.Path,
entry.Size,
entry.Mtime,
f,
fileProgress,
)
_ = f.Close()
// Close channel and wait for goroutine to finish
if fileProgress != nil {
close(fileProgress)
wg.Wait()
}
if err != nil {
return err
}
log.Verbosef("+ %s (%s)", entry.Path, humanize.IBytes(uint64(bytesRead)))
scannedFiles++
scannedBytes += bytesRead
}
// Send final progress (ETA is 0 at completion)
if progress != nil {
elapsed := time.Since(startTime).Seconds()
var rate float64
if elapsed > 0 {
rate = float64(scannedBytes) / elapsed
}
sendScanStatus(progress, ScanStatus{
TotalFiles: totalFiles,
ScannedFiles: scannedFiles,
TotalBytes: pt.totalBytes,
ScannedBytes: currentBytes,
TotalBytes: totalBytes,
ScannedBytes: scannedBytes,
BytesPerSec: rate,
ETA: eta,
ETA: 0,
})
*pt.lastProgress = now
}
}
// computeRateETA returns the average throughput over elapsed time and the
// estimated time to process the remaining bytes at that rate.
func computeRateETA(
elapsed time.Duration,
done FileSize,
total FileSize,
) (float64, time.Duration) {
var (
rate float64
eta time.Duration
)
if elapsed > 0 && done > 0 {
rate = float64(done) / elapsed.Seconds()
remaining := total - done
if rate > 0 {
eta = time.Duration(float64(remaining)/rate) * time.Second
}
}
return rate, eta
}
// sizeToUint64 converts a FileSize to uint64 for display, clamping
// negative values to zero so the conversion cannot overflow.
func sizeToUint64(v FileSize) uint64 {
if v < 0 {
return 0
}
return uint64(v)
// Build and write manifest
return builder.Build(w)
}
// IsHiddenPath returns true if the path or any of its parent directories
@@ -541,21 +396,17 @@ func IsHiddenPath(p string) bool {
if tp == "." || tp == "/" {
return false
}
if strings.HasPrefix(tp, ".") {
return true
}
for {
d, f := path.Split(tp)
if strings.HasPrefix(f, ".") {
return true
}
if d == "" {
return false
}
tp = d[0 : len(d)-1] // trim trailing slash from dir
}
}
@@ -566,7 +417,6 @@ func sendEnumerateStatus(ch chan<- EnumerateStatus, status EnumerateStatus) {
if ch == nil {
return
}
select {
case ch <- status:
default:
@@ -580,7 +430,6 @@ func sendScanStatus(ch chan<- ScanStatus, status ScanStatus) {
if ch == nil {
return
}
select {
case ch <- status:
default:
+44 -84
View File
@@ -1,4 +1,3 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
@@ -13,8 +12,6 @@ import (
)
func TestNewScanner(t *testing.T) {
t.Parallel()
s := NewScanner()
assert.NotNil(t, s)
assert.Equal(t, FileCount(0), s.FileCount())
@@ -22,18 +19,12 @@ func TestNewScanner(t *testing.T) {
}
func TestNewScannerWithOptions(t *testing.T) {
t.Parallel()
t.Run("nil options", func(t *testing.T) {
t.Parallel()
s := NewScannerWithOptions(nil)
assert.NotNil(t, s)
})
t.Run("with options", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
opts := &ScannerOptions{
IncludeDotfiles: true,
@@ -46,8 +37,6 @@ func TestNewScannerWithOptions(t *testing.T) {
}
func TestScannerEnumerateFile(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("hello world"), 0o644))
@@ -65,8 +54,6 @@ func TestScannerEnumerateFile(t *testing.T) {
}
func TestScannerEnumerateFileMissing(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumerateFile("/nonexistent.txt")
@@ -74,14 +61,11 @@ func TestScannerEnumerateFileMissing(t *testing.T) {
}
func TestScannerEnumeratePath(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir/subdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("one"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("two"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/subdir/file3.txt", []byte("three"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/subdir/file3.txt", []byte("three"), 0o644))
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumeratePath("/testdir", nil)
@@ -92,8 +76,6 @@ func TestScannerEnumeratePath(t *testing.T) {
}
func TestScannerEnumeratePathWithProgress(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("one"), 0o644))
@@ -118,8 +100,6 @@ func TestScannerEnumeratePathWithProgress(t *testing.T) {
}
func TestScannerEnumeratePaths(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/dir1", 0o755))
require.NoError(t, fs.MkdirAll("/dir2", 0o755))
@@ -134,20 +114,13 @@ func TestScannerEnumeratePaths(t *testing.T) {
}
func TestScannerExcludeDotfiles(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir/.hidden", 0o755))
require.NoError(t,
afero.WriteFile(fs, "/testdir/visible.txt", []byte("visible"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/.hidden.txt", []byte("hidden"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/.hidden/inside.txt", []byte("inside"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/visible.txt", []byte("visible"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/.hidden.txt", []byte("hidden"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/.hidden/inside.txt", []byte("inside"), 0o644))
t.Run("exclude by default", func(t *testing.T) {
t.Parallel()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs, IncludeDotfiles: false})
err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err)
@@ -158,8 +131,6 @@ func TestScannerExcludeDotfiles(t *testing.T) {
})
t.Run("include when enabled", func(t *testing.T) {
t.Parallel()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs, IncludeDotfiles: true})
err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err)
@@ -169,43 +140,34 @@ func TestScannerExcludeDotfiles(t *testing.T) {
}
func TestScannerToManifest(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file1.txt", []byte("content one"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file2.txt", []byte("content two"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content one"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content two"), 0o644))
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err)
var buf bytes.Buffer
err = s.ToManifest(context.Background(), &buf, nil)
require.NoError(t, err)
// Manifest should have magic bytes
assert.Positive(t, buf.Len())
assert.True(t, buf.Len() > 0)
assert.Equal(t, MAGIC, string(buf.Bytes()[:8]))
}
func TestScannerToManifestWithProgress(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file.txt", bytes.Repeat([]byte("x"), 1000), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file.txt", bytes.Repeat([]byte("x"), 1000), 0o644))
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err)
var buf bytes.Buffer
progress := make(chan ScanStatus, 10)
err = s.ToManifest(context.Background(), &buf, progress)
@@ -226,15 +188,12 @@ func TestScannerToManifestWithProgress(t *testing.T) {
}
func TestScannerToManifestContextCancellation(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
// Create many files to ensure we have time to cancel
for i := range 100 {
for i := 0; i < 100; i++ {
name := string(rune('a'+i%26)) + string(rune('0'+i/26)) + ".txt"
require.NoError(t,
afero.WriteFile(fs, "/testdir/"+name, bytes.Repeat([]byte("x"), 100), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/"+name, bytes.Repeat([]byte("x"), 100), 0o644))
}
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
@@ -245,30 +204,24 @@ func TestScannerToManifestContextCancellation(t *testing.T) {
cancel() // Cancel immediately
var buf bytes.Buffer
err = s.ToManifest(ctx, &buf, nil)
assert.ErrorIs(t, err, context.Canceled)
}
func TestScannerToManifestEmptyScanner(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
var buf bytes.Buffer
err := s.ToManifest(context.Background(), &buf, nil)
require.NoError(t, err)
// Should still produce a valid manifest
assert.Positive(t, buf.Len())
assert.True(t, buf.Len() > 0)
assert.Equal(t, MAGIC, string(buf.Bytes()[:8]))
}
func TestScannerFilesCopiesSlice(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("hello"), 0o644))
@@ -283,13 +236,10 @@ func TestScannerFilesCopiesSlice(t *testing.T) {
}
func TestScannerEnumerateFS(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir/sub", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file.txt", []byte("hello"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/sub/nested.txt", []byte("world"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/sub/nested.txt", []byte("world"), 0o644))
// Create a basepath filesystem
baseFs := afero.NewBasePathFs(fs, "/testdir")
@@ -302,37 +252,51 @@ func TestScannerEnumerateFS(t *testing.T) {
}
func TestSendEnumerateStatusNonBlocking(t *testing.T) {
t.Parallel()
// Nobody receives, so a blocking send would hang the test into its timeout.
// Channel with no buffer - send should not block
ch := make(chan EnumerateStatus)
sendEnumerateStatus(ch, EnumerateStatus{FilesFound: 1})
// This should not block
done := make(chan bool)
go func() {
sendEnumerateStatus(ch, EnumerateStatus{FilesFound: 1})
done <- true
}()
select {
case <-done:
// Success - did not block
case <-time.After(100 * time.Millisecond):
t.Fatal("sendEnumerateStatus blocked on full channel")
}
}
func TestSendScanStatusNonBlocking(t *testing.T) {
t.Parallel()
// Nobody receives, so a blocking send would hang the test into its timeout.
// Channel with no buffer - send should not block
ch := make(chan ScanStatus)
sendScanStatus(ch, ScanStatus{ScannedFiles: 1})
done := make(chan bool)
go func() {
sendScanStatus(ch, ScanStatus{ScannedFiles: 1})
done <- true
}()
select {
case <-done:
// Success - did not block
case <-time.After(100 * time.Millisecond):
t.Fatal("sendScanStatus blocked on full channel")
}
}
func TestSendStatusNilChannel(t *testing.T) {
t.Parallel()
// Should not panic with nil channel
sendEnumerateStatus(nil, EnumerateStatus{})
sendScanStatus(nil, ScanStatus{})
}
func TestScannerFileEntryFields(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
now := time.Now().Truncate(time.Second)
require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("content"), 0o644))
require.NoError(t, fs.Chtimes("/test.txt", now, now))
@@ -351,13 +315,11 @@ func TestScannerFileEntryFields(t *testing.T) {
}
func TestScannerLargeFileEnumeration(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
// Create 100 files
for i := range 100 {
for i := 0; i < 100; i++ {
name := "/testdir/" + string(rune('a'+i%26)) + string(rune('0'+i/26%10)) + ".txt"
require.NoError(t, afero.WriteFile(fs, name, []byte("data"), 0o644))
}
@@ -368,20 +330,20 @@ func TestScannerLargeFileEnumeration(t *testing.T) {
err := s.EnumeratePath("/testdir", progress)
require.NoError(t, err)
// progress is fully buffered and closed; no draining needed
// Drain channel
for range progress {
}
assert.Equal(t, FileCount(100), s.FileCount())
assert.Equal(t, FileSize(400), s.TotalBytes()) // 100 * 4 bytes
}
func TestIsHiddenPath(t *testing.T) {
t.Parallel()
tests := []struct {
path string
hidden bool
}{
{testFileName, false},
{"file.txt", false},
{".hidden", true},
{"dir/file.txt", false},
{"dir/.hidden", true},
@@ -398,8 +360,6 @@ func TestIsHiddenPath(t *testing.T) {
for _, tt := range tests {
t.Run(tt.path, func(t *testing.T) {
t.Parallel()
assert.Equal(t, tt.hidden, IsHiddenPath(tt.path), "IsHiddenPath(%q)", tt.path)
})
}
+31 -83
View File
@@ -2,11 +2,9 @@ package mfer
import (
"bytes"
"context"
"crypto/sha256"
"errors"
"fmt"
"math"
"time"
"github.com/google/uuid"
@@ -17,80 +15,47 @@ import (
// MAGIC is the file format magic bytes prefix (rot13 of "MANIFEST").
const MAGIC string = "ZNAVSRFG"
var (
// errInnerNotSet is returned by generate when the inner manifest is
// missing.
errInnerNotSet = errors.New("internal error: pbInner not set")
// errInternal is returned by generateOuter for the same condition.
// The two messages differ, and both are load-bearing for callers that
// match on text, so they are kept distinct.
errInternal = errors.New("internal error")
)
// nanosecondsInt32 converts t's nanosecond component to int32.
// time.Time.Nanosecond is documented to return a value in [0, 999999999],
// so the conversion cannot overflow. This sits directly in the manifest
// content path: silently substituting a default would zero every entry's
// mtime nanos and change the serialized bytes and their hash, so an
// out-of-contract value is a programming error and panics rather than
// being papered over.
func nanosecondsInt32(t time.Time) int32 {
n := t.Nanosecond()
if n < 0 || n > math.MaxInt32 {
panic(fmt.Sprintf(
"mfer: time.Time.Nanosecond out of contract: %d", n))
}
return int32(n)
}
func newTimestampFromTime(t time.Time) *Timestamp {
return &Timestamp{
Seconds: t.Unix(),
Nanos: nanosecondsInt32(t),
Nanos: int32(t.Nanosecond()),
}
}
func (m *manifest) generate(ctx context.Context) error {
func (m *manifest) generate() error {
if m.pbInner == nil {
return errInnerNotSet
return errors.New("internal error: pbInner not set")
}
if m.pbOuter == nil {
e := m.generateOuter(ctx)
e := m.generateOuter()
if e != nil {
return e
}
}
dat, err := proto.MarshalOptions{Deterministic: true}.Marshal(m.pbOuter)
if err != nil {
return fmt.Errorf("serialize: marshal outer: %w", err)
}
m.output = bytes.NewBufferString(MAGIC)
m.output = bytes.NewBuffer([]byte(MAGIC))
_, err = m.output.Write(dat)
if err != nil {
return fmt.Errorf("serialize: write output: %w", err)
}
return nil
}
func (m *manifest) generateOuter(ctx context.Context) error {
func (m *manifest) generateOuter() error {
if m.pbInner == nil {
return errInternal
return errors.New("internal error")
}
// Use fixed UUID if provided, otherwise generate a new one
var manifestUUID uuid.UUID
if len(m.fixedUUID) == uuidLength {
if len(m.fixedUUID) == 16 {
copy(manifestUUID[:], m.fixedUUID)
} else {
manifestUUID = uuid.New()
}
m.pbInner.Uuid = manifestUUID[:]
innerData, err := proto.MarshalOptions{Deterministic: true}.Marshal(m.pbInner)
@@ -100,29 +65,23 @@ func (m *manifest) generateOuter(ctx context.Context) error {
// Compress the inner data
idc := new(bytes.Buffer)
zw, err := zstd.NewWriter(idc, zstd.WithEncoderLevel(zstd.SpeedBestCompression))
if err != nil {
return fmt.Errorf("serialize: create compressor: %w", err)
}
_, err = zw.Write(innerData)
if err != nil {
return fmt.Errorf("serialize: compress: %w", err)
}
_ = zw.Close()
compressedData := idc.Bytes()
// Hash the compressed data for integrity verification before decompression
h := sha256.New()
_, err = h.Write(compressedData)
if err != nil {
if _, err := h.Write(compressedData); err != nil {
return fmt.Errorf("serialize: hash write: %w", err)
}
sha256Hash := h.Sum(nil)
m.pbOuter = &MFFileOuter{
@@ -136,40 +95,29 @@ func (m *manifest) generateOuter(ctx context.Context) error {
// Sign the manifest if signing options are provided
if m.signingOptions != nil && m.signingOptions.KeyID != "" {
return m.signOuter(ctx)
sigString, err := m.signatureString()
if err != nil {
return fmt.Errorf("failed to generate signature string: %w", err)
}
sig, err := gpgSign([]byte(sigString), m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to sign manifest: %w", err)
}
m.pbOuter.Signature = sig
fingerprint, err := gpgGetKeyFingerprint(m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to get key fingerprint: %w", err)
}
m.pbOuter.Signer = fingerprint
pubKey, err := gpgExportPublicKey(m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to export public key: %w", err)
}
m.pbOuter.SigningPubKey = pubKey
}
return nil
}
// signOuter signs the outer message with the configured GPG key and
// embeds the signature, signer fingerprint, and public key.
func (m *manifest) signOuter(ctx context.Context) error {
sigString, err := m.signatureString()
if err != nil {
return fmt.Errorf("failed to generate signature string: %w", err)
}
sig, err := gpgSign(ctx, []byte(sigString), m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to sign manifest: %w", err)
}
m.pbOuter.Signature = sig
fingerprint, err := gpgGetKeyFingerprint(ctx, m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to get key fingerprint: %w", err)
}
m.pbOuter.Signer = fingerprint
pubKey, err := gpgExportPublicKey(ctx, m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to export public key: %w", err)
}
m.pbOuter.SigningPubKey = pubKey
return nil
}
-2
View File
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("")
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFGy[i\x04\xe5O\x82A\x1d\xf4\xb0\xe2z7:U\xee\xa3\xf9\xd6m\xacZ\x9b\xce\x1d\xd9/{@\x1d\xa5y[i\x04\xe5O\x82A\x1d\xf4\xb0\xe2z7:U\xee\xa3\xf9\xd6m\xacZ\x9b\xce\x1d\xd9/{@\x1d\xa5")
-2
View File
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 \xa3\xf7\x97\xa7\xf3\x87:\x90)\\ӊj\xb9\xf7\xfaTJ\x1b\xe7:\xee\xbe\"V\xe0:\x8d(Z\xd6%\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\fc(\xb5/\xfd\x04\x00\xb1\x02\x00\xa0\x06\x01\xaa\x06=\n\x05a.txt\x10\x01\x1a$\n\"\x12 ʗ\x81\x12\xca\x1b\xbd\xca\xfa\xc21\xb3\x9a#\xdcM\xa7\x86\xef\xf8\x14|Nr\xb9\x80w\x85\xaf\xeeH\xbb\xf2\x12\v\b\x80\x92\xb8Ø\xfe\xff\xff\xff\x01\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3s\xeeG\x80")
-2
View File
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 \xa3\xf7\x97\xa7\xf3\x87:\x90)\\ӊj\xb9\xf7\xfaTJ\x1b\xe7:\xee\xbe\"V\xe0:\x8d(Z\xd6%\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\fc(\xb5/\xfd\x04\x00\xb1\x02\x00\xa0\x06\x01\xaa\x06=\n\x05a.txt\x10\x01\x1a$\n\"\x12 ʗ\x81\x12\xca\x1b\xbd\xca\xfa\xc21\xb3\x9a#\xdcM\xa7\x86\xef\xf8\x14|Nr\xb9\x80w\x85\xaf\xeeH\xbb\xf2\x12\v\b\x80\x92\xb8Ø\xfe\xff\xff\xff\x01\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3s\xeeG\x80\xca\f\xe8\x03-----BEGIN PGP SIGNATURE-----\n\niQEzBAABCgAdFiEET1Yr+4Y/3GtRtO6IhypRF2zvI64FAmrBIXAACgkQhypRF2zv\nI67BQAf/QrpX2MjY15YGMGkjR5oIhnx/YV96aGYZyZThzb+l/R/N75iVFVkhX21d\nZhQqdCsORrodTPAXic2g2UGVXP9PhNMh7n6Wm3LsvQYjrRQGrQnqtCkut+3tUt8K\n7pt4OAnnwRSieaVImA1COmzxIrQQKNOs6UkgmAstGuPV0XZoeDiSG8TUYJ/vieCn\np5hC0FFXtzfw4NtkxSmkewE0xBxIwFCA/RfSHCGH3m5K+tRz41vMEgGbL1iEp6+V\nuBaoEc4hqCgEt+Af2pA8VHfqeu2vKiwggOpYpaILXZKVqH9+tWHL1EBv9t0vTsYE\n9D57euuR9+kOdngYNPieP1yn5dOSHg==\n=kvgL\n-----END PGP SIGNATURE-----\n\xd2\f(4F562BFB863FDC6B51B4EE88872A51176CEF23AE\xda\f\xb5\a-----BEGIN PGP PUBLIC KEY BLOCK-----\n\nmQENBGrBIW8BCADESetN5EdxIe7Fafgxl99Yoo5cOexf7wJyYT0wfUYlRaxt3neR\nhir7LOfH4PZWWoDx7qghxCS4+vs7yGypl6JOm7jnJlhn4HneDa2zeIlgGW2TamyE\nua9KPWBQqkFOYmKPmzp+KnL6ncnBLR5mDkNKFyON812KVvteu6Dp/DNk4Meufe44\nWWr49LSFZa9gEbmRCoQGKby9F0H0yIi4FAc74VdQudy0+fMKcfkKjEvByMzlbBEK\n92Hq3sRFzWd3kvPliNjZTmlh5n5m9aBhMpoy3GkKy8gpDdFc6NLA9iAJe7oNMriR\nkVoa5EjQL1xCXAiAWTYA9NScFfU/574sCTxZABEBAAG0Hk1GRVIgVGVzdCBLZXkg\nPHRlc3RAbWZlci50ZXN0PokBTwQTAQoAORYhBE9WK/uGP9xrUbTuiIcqURds7yOu\nBQJqwSFvAxsvBAULCQgHAgYVCgkICwIEFgIDAQIeAQIXgAAKCRCHKlEXbO8jrjml\nCAC8wUK9wmvxq0+NZUpFyP+P29klLZYzBDaBrLPJFs0GjnG4kvfUAktWx0Ro80F7\ncjTJ4f44XjDj4glvSjbe2VaDnZl9FTfzUfG+xjD4462NgntQ4fHk/uG4F6d1ikWx\nkEoMpIn1PlSMas1jTQSGlxUr+zFwWuUbGq4n6hRxEnwLlwJlwQt/Aw1vPDYuPDE3\nOYDhJIAJyP+6e9W8ToaAG9byg/22KA1u1qxnNQqsx5Tped2VltAzdYub+yeCuNc8\nIUo5ILo/fQq3GM5sUEaHjPolv88WlDm3vcdbSbDoh5m2inDtg5zuUKJwu32UGu8l\nKoTjp4nxgQGy5WmeBzs4Hdm/\n=TCB8\n-----END PGP PUBLIC KEY BLOCK-----\n")
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06!\xc2\x06 \x91\x90*\xa5>\fݐ \x87\xbeaL\xc1\x05?\x0eR\xc18\xa4eՕ\xa95\xb9KʺoZ\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\f-(\xb5/\xfd\x04\x00\x01\x01\x00\xa0\x06\x01\xaa\x06\a\n\x05a.txt\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3a[k'")
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06\x1f\xc2\x06 \x91\x90*\xa5>\fݐ \x87\xbeaL\xc1\x05?\x0eR\xc18\xa4eՕ\xa95\xb9KʺoZ\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\f-(\xb5/\xfd\x04\x00\x01\x01\x00\xa0\x06\x01\xaa\x06\a\n\x05a.txt\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3a[k'")
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 \xa3\xf7\x97\xa7\xf3\x87:\x90)\\ӊj\xb9\xf7\xfaTJ\x1b\xe7:\xee\xbe\"V\xe0:\x8d(Z\xd6%\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\fc(\xb5/\xfd\x04\x00\xb1\x02\x00\xa0\x06\x01\xaa\x06=\n\x05a.txt\x10\x01\x1a$\n\"\x12 ʗ\x81\x12\xca\x1b\xbd\xca\xfa\xc21\xb3\x9a#\xdcM\xa7\x86\xef\xf8\x14|Nr\xb9\x80w\x85\xaf\xeeH\xbb\xf2\x12\v\b\x80\x92\xb8Ø\xfe\xff\xff\xff\x01\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3s\xeeG")
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 ")
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAV")
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFG")
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 \xa3\xf7\x97\xa7\xf3\x87:\x90)\\ӊj\xb9\xf7\xfaTJ\x1b\xe7:\xee\xbe\"V\xe0:\x8d(Z\xd6%\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\fc(\xb5/\xfd\x04\x00\xb1\x02\x00\xa0\x06\x01")
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFX\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 \xa3\xf7\x97\xa7\xf3\x87:\x90)\\ӊj\xb9\xf7\xfaTJ\x1b\xe7:\xee\xbe\"V\xe0:\x8d(Z\xd6%\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\fc(\xb5/\xfd\x04\x00\xb1\x02\x00\xa0\x06\x01\xaa\x06=\n\x05a.txt\x10\x01\x1a$\n\"\x12 ʗ\x81\x12\xca\x1b\xbd\xca\xfa\xc21\xb3\x9a#\xdcM\xa7\x86\xef\xf8\x14|Nr\xb9\x80w\x85\xaf\xeeH\xbb\xf2\x12\v\b\x80\x92\xb8Ø\xfe\xff\xff\xff\x01\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3s\xeeG\x80")
File diff suppressed because one or more lines are too long
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06\x01\xc2\x06 \xd6aQ\xa4\x85\xc0+\xbb\xca\x11\x11\a<\x019\x97\xb3\xbb3\xd0 \xd5U\xfa!\xaeAf<N@\x9d\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\f\x12(\xb5/\xfd\xc0\x00\x00\x00\x00\x00\x02\x00\x00\x00\v\x00\x00\x00")
@@ -1,2 +0,0 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06\x01\xc2\x06 \xc8g?\xa0\xc9\xc5]\xb57M\xe5O#\x04\xb2\x12\xc6)@=\xd2\xf3f/͉~\x10\x15\xfbkD\xca\x06\x10o\x1c*N;]L~\x9a\x8b\x1d.?@Qb\xba\f\x89\t(\xb5/\xfd\x00\x00\x01\x00\x00(\xb5/\xfd\x00\x01\x01\x00\x00(\xb5/\xfd\x00\x02\x01\x00\x00(\xb5/\xfd\x00\x03\x01\x00\x00(\xb5/\xfd\x00\x04\x01\x00\x00(\xb5/\xfd\x00\x05\x01\x00\x00(\xb5/\xfd\x00\x06\x01\x00\x00(\xb5/\xfd\x00\a\x01\x00\x00(\xb5/\xfd\x00\b\x01\x00\x00(\xb5/\xfd\x00\t\x01\x00\x00(\xb5/\xfd\x00\n\x01\x00\x00(\xb5/\xfd\x00\v\x01\x00\x00(\xb5/\xfd\x00\f\x01\x00\x00(\xb5/\xfd\x00\r\x01\x00\x00(\xb5/\xfd\x00\x0e\x01\x00\x00(\xb5/\xfd\x00\x0f\x01\x00\x00(\xb5/\xfd\x00\x10\x01\x00\x00(\xb5/\xfd\x00\x11\x01\x00\x00(\xb5/\xfd\x00\x12\x01\x00\x00(\xb5/\xfd\x00\x13\x01\x00\x00(\xb5/\xfd\x00\x14\x01\x00\x00(\xb5/\xfd\x00\x15\x01\x00\x00(\xb5/\xfd\x00\x16\x01\x00\x00(\xb5/\xfd\x00\x17\x01\x00\x00(\xb5/\xfd\x00\x18\x01\x00\x00(\xb5/\xfd\x00\x19\x01\x00\x00(\xb5/\xfd\x00\x1a\x01\x00\x00(\xb5/\xfd\x00\x1b\x01\x00\x00(\xb5/\xfd\x00\x1c\x01\x00\x00(\xb5/\xfd\x00\x1d\x01\x00\x00(\xb5/\xfd\x00\x1e\x01\x00\x00(\xb5/\xfd\x00\x1f\x01\x00\x00(\xb5/\xfd\x00 \x01\x00\x00(\xb5/\xfd\x00!\x01\x00\x00(\xb5/\xfd\x00\"\x01\x00\x00(\xb5/\xfd\x00#\x01\x00\x00(\xb5/\xfd\x00$\x01\x00\x00(\xb5/\xfd\x00%\x01\x00\x00(\xb5/\xfd\x00&\x01\x00\x00(\xb5/\xfd\x00'\x01\x00\x00(\xb5/\xfd\x00(\x01\x00\x00(\xb5/\xfd\x00)\x01\x00\x00(\xb5/\xfd\x00*\x01\x00\x00(\xb5/\xfd\x00+\x01\x00\x00(\xb5/\xfd\x00,\x01\x00\x00(\xb5/\xfd\x00-\x01\x00\x00(\xb5/\xfd\x00.\x01\x00\x00(\xb5/\xfd\x00/\x01\x00\x00(\xb5/\xfd\x000\x01\x00\x00(\xb5/\xfd\x001\x01\x00\x00(\xb5/\xfd\x002\x01\x00\x00(\xb5/\xfd\x003\x01\x00\x00(\xb5/\xfd\x004\x01\x00\x00(\xb5/\xfd\x005\x01\x00\x00(\xb5/\xfd\x006\x01\x00\x00(\xb5/\xfd\x007\x01\x00\x00(\xb5/\xfd\x008\x01\x00\x00(\xb5/\xfd\x009\x01\x00\x00(\xb5/\xfd\x00:\x01\x00\x00(\xb5/\xfd\x00;\x01\x00\x00(\xb5/\xfd\x00<\x01\x00\x00(\xb5/\xfd\x00=\x01\x00\x00(\xb5/\xfd\x00>\x01\x00\x00(\xb5/\xfd\x00?\x01\x00\x00(\xb5/\xfd\x00@\x01\x00\x00(\xb5/\xfd\x00A\x01\x00\x00(\xb5/\xfd\x00B\x01\x00\x00(\xb5/\xfd\x00C\x01\x00\x00(\xb5/\xfd\x00D\x01\x00\x00(\xb5/\xfd\x00E\x01\x00\x00(\xb5/\xfd\x00F\x01\x00\x00(\xb5/\xfd\x00G\x01\x00\x00(\xb5/\xfd\x00H\x01\x00\x00(\xb5/\xfd\x00I\x01\x00\x00(\xb5/\xfd\x00J\x01\x00\x00(\xb5/\xfd\x00K\x01\x00\x00(\xb5/\xfd\x00L\x01\x00\x00(\xb5/\xfd\x00M\x01\x00\x00(\xb5/\xfd\x00N\x01\x00\x00(\xb5/\xfd\x00O\x01\x00\x00(\xb5/\xfd\x00P\x01\x00\x00(\xb5/\xfd\x00Q\x01\x00\x00(\xb5/\xfd\x00R\x01\x00\x00(\xb5/\xfd\x00S\x01\x00\x00(\xb5/\xfd\x00T\x01\x00\x00(\xb5/\xfd\x00U\x01\x00\x00(\xb5/\xfd\x00V\x01\x00\x00(\xb5/\xfd\x00W\x01\x00\x00(\xb5/\xfd\x00X\x01\x00\x00(\xb5/\xfd\x00Y\x01\x00\x00(\xb5/\xfd\x00Z\x01\x00\x00(\xb5/\xfd\x00[\x01\x00\x00(\xb5/\xfd\x00\\\x01\x00\x00(\xb5/\xfd\x00]\x01\x00\x00(\xb5/\xfd\x00^\x01\x00\x00(\xb5/\xfd\x00_\x01\x00\x00(\xb5/\xfd\x00`\x01\x00\x00(\xb5/\xfd\x00a\x01\x00\x00(\xb5/\xfd\x00b\x01\x00\x00(\xb5/\xfd\x00c\x01\x00\x00(\xb5/\xfd\x00d\x01\x00\x00(\xb5/\xfd\x00e\x01\x00\x00(\xb5/\xfd\x00f\x01\x00\x00(\xb5/\xfd\x00g\x01\x00\x00(\xb5/\xfd\x00h\x01\x00\x00(\xb5/\xfd\x00i\x01\x00\x00(\xb5/\xfd\x00j\x01\x00\x00(\xb5/\xfd\x00k\x01\x00\x00(\xb5/\xfd\x00l\x01\x00\x00(\xb5/\xfd\x00m\x01\x00\x00(\xb5/\xfd\x00n\x01\x00\x00(\xb5/\xfd\x00o\x01\x00\x00(\xb5/\xfd\x00p\x01\x00\x00(\xb5/\xfd\x00q\x01\x00\x00(\xb5/\xfd\x00r\x01\x00\x00(\xb5/\xfd\x00s\x01\x00\x00(\xb5/\xfd\x00t\x01\x00\x00(\xb5/\xfd\x00u\x01\x00\x00(\xb5/\xfd\x00v\x01\x00\x00(\xb5/\xfd\x00w\x01\x00\x00(\xb5/\xfd\x00x\x01\x00\x00(\xb5/\xfd\x00y\x01\x00\x00(\xb5/\xfd\x00z\x01\x00\x00(\xb5/\xfd\x00{\x01\x00\x00(\xb5/\xfd\x00|\x01\x00\x00(\xb5/\xfd\x00}\x01\x00\x00(\xb5/\xfd\x00~\x01\x00\x00(\xb5/\xfd\x00\x7f\x01\x00\x00(\xb5/\xfd\x00\x80\x01\x00\x00")
File diff suppressed because one or more lines are too long
-2
View File
@@ -32,14 +32,12 @@ func (b BaseURL) JoinPath(path RelFilePath) (FileURL, error) {
for i, seg := range segments {
segments[i] = url.PathEscape(seg)
}
ref, err := url.Parse(strings.Join(segments, "/"))
if err != nil {
return "", err
}
resolved := base.ResolveReference(ref)
return FileURL(resolved.String()), nil
}
+3 -18
View File
@@ -1,4 +1,3 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
@@ -9,27 +8,19 @@ import (
)
func TestBaseURLJoinPath(t *testing.T) {
t.Parallel()
tests := []struct {
base BaseURL
path RelFilePath
expected string
}{
{"https://example.com/dir/", testFileName, "https://example.com/dir/file.txt"},
{"https://example.com/dir", testFileName, "https://example.com/dir/file.txt"},
{"https://example.com/dir/", "file.txt", "https://example.com/dir/file.txt"},
{"https://example.com/dir", "file.txt", "https://example.com/dir/file.txt"},
{"https://example.com/", "sub/file.txt", "https://example.com/sub/file.txt"},
{
"https://example.com/dir/",
"file with spaces.txt",
"https://example.com/dir/file%20with%20spaces.txt",
},
{"https://example.com/dir/", "file with spaces.txt", "https://example.com/dir/file%20with%20spaces.txt"},
}
for _, tt := range tests {
t.Run(string(tt.base)+"+"+string(tt.path), func(t *testing.T) {
t.Parallel()
result, err := tt.base.JoinPath(tt.path)
require.NoError(t, err)
assert.Equal(t, tt.expected, string(result))
@@ -38,22 +29,16 @@ func TestBaseURLJoinPath(t *testing.T) {
}
func TestBaseURLString(t *testing.T) {
t.Parallel()
b := BaseURL("https://example.com/")
assert.Equal(t, "https://example.com/", b.String())
}
func TestFileURLString(t *testing.T) {
t.Parallel()
f := FileURL("https://example.com/file.txt")
assert.Equal(t, "https://example.com/file.txt", f.String())
}
func TestManifestURLString(t *testing.T) {
t.Parallel()
m := ManifestURL("https://example.com/index.mf")
assert.Equal(t, "https://example.com/index.mf", m.String())
}
+6 -1
View File
@@ -140,7 +140,12 @@ main() {
# ---- Go repos ----
if missing go; then pkg_install go golang go go; fi
# No golangci-lint: script/lint runs it in Docker only.
# golangci-lint: packaged in nix, brew, and apk. On apt there is no
# package: download a specific release archive from GitHub and
# verify its hash (verify_sha256), never curl | sh.
if missing golangci-lint; then
pkg_install golangci-lint golangci-lint golangci-lint golangci-lint
fi
go mod download
# ---- Python repos ----
+5 -15
View File
@@ -1,24 +1,14 @@
#!/bin/sh
# script/cibuild: run the CI build; the Gitea workflow runs this on push.
# It builds the image with the same command as script/docker. --no-cache
# because the checks the final stage depends on are RUN steps, and a
# cached one is a check that did not run.
# script/cibuild: run the CI build. The Dockerfile runs script/check
# (via make check), so a successful build implies all checks pass.
# Generic: needs no adaptation. The Gitea workflow runs this on push.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
# Own line: a failing command substitution inside an argument does
# not trip `set -e`, so the inline form degrades silently to an
# empty constant. The VERSION build argument takes precedence over
# the version a build stage derives from the .git in the context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$("$SCRIPT_DIR/projectname")" .
docker build .
}
main "$@"
+2 -11
View File
@@ -1,8 +1,7 @@
#!/bin/sh
# script/docker: build the Docker image tagged with the project name.
# Identical in all repos; the tag comes from script/projectname.
# --no-cache because the gate phases the final stage depends on are RUN
# steps, and a cached one is a check that did not run.
# Generic: needs no adaptation.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
@@ -10,15 +9,7 @@ ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
# Own line: a failing command substitution inside an argument does
# not trip `set -e`, so the inline form degrades silently to an
# empty constant. The VERSION build argument takes precedence over
# the version a build stage derives from the .git in the context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$("$SCRIPT_DIR/projectname")" .
docker build -t "$("$SCRIPT_DIR/projectname")" .
}
main "$@"
+1
View File
@@ -19,6 +19,7 @@ main() {
cd "$ROOT"
ensure_pb
gofumpt -l -w mfer internal cmd
golangci-lint run --fix
# Markdown and JSON, over the same file set script/fmt-check verifies.
"$SCRIPT_DIR/prettier" --write
}
-17
View File
@@ -1,17 +0,0 @@
#!/bin/sh
# script/fuzz: fuzz the manifest parser for one minute. Run by hand only:
# script/test already runs the committed seed corpus as ordinary tests,
# and CI never fuzzes. An input that fails is written to
# mfer/testdata/fuzz/FuzzNewManifestFromReader/; once the parser is fixed,
# commit it there as a regression seed.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
go test -run '^$' -fuzz '^FuzzNewManifestFromReader$' \
-fuzztime 1m -parallel 2 ./mfer
}
main "$@"
+8 -11
View File
@@ -1,20 +1,17 @@
#!/bin/sh
# script/lint: run golangci-lint, in Docker only. Builds the lint stage of
# the Dockerfile, whose build runs the linter, so a successful build is a
# clean lint. --no-cache because a cached build runs no linter. The image
# is removed afterwards, whatever the outcome.
# script/lint: run the linter.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
# Tagged per run, so concurrent runs never remove each other's image.
image="$("$SCRIPT_DIR/projectname")-lint:$$"
# A failed build leaves no image, so there is nothing to remove then.
trap 'docker image rm "$image" >/dev/null 2>&1 || true' EXIT INT TERM
docker build --no-cache --target lint -t "$image" .
golangci-lint run
if [ -n "$(gofmt -l .)" ]; then
echo "gofmt: files need formatting:" >&2
gofmt -l . >&2
exit 1
fi
}
main "$@"
+1 -6
View File
@@ -17,12 +17,7 @@ ensure_pb() {
main() {
cd "$ROOT"
ensure_pb
go test -timeout 30s -race -cover ./... ||
{
echo "--- Rerunning with -v for details ---"
go test -timeout 30s -race -v ./...
exit 1
}
go test -v --timeout 10s ./...
}
main "$@"