25 Commits

Author SHA1 Message Date
5683d0f4ff Configure prettier and make fmt-check cover markdown (closes #69)
All checks were successful
check / check (push) Successful in 4s
Adds .prettierrc and .prettierignore, a single script/prettier entrypoint shared by fmt and fmt-check so the two cannot drift, and prettier 3.9.6 pinned by yarn.lock integrity hash.

Markdown formatting is now gated in the authoritative Docker build via a new mdfmt stage, since the golangci-lint image has no node. REPO_POLICIES.md is ignored so local tooling cannot drift it from upstream.

Removes the || true that made the previous prettier invocation unable to fail.
2026-08-10 16:16:42 +02:00
de476708e9 Update golangci-lint to v2.12.2 with canonical config (closes #60)
Some checks failed
check / check (push) Has been cancelled
Adopts golangci-lint v2.12.2 and the canonical .golangci.yml (default: all), and fixes all resulting findings across the tree.

Two intended behavior changes: absent MFFilePath.Mtime is handled explicitly in freshen, list and export rather than dereferenced (main panicked); gpg positional key IDs now follow an explicit -- end-of-options marker.

All twelve reworded user-visible error messages restored to byte-identical parity with main and pinned by tests.
2026-08-10 16:06:12 +02:00
6d19de74e7 scripts-to-rule-them-all (#58)
All checks were successful
check / check (push) Successful in 6s
Reviewed-on: #58
Co-authored-by: sneak <sneak@sneak.berlin>
Co-committed-by: sneak <sneak@sneak.berlin>
2026-07-07 02:13:35 +02:00
688ebd90a4 TODO (#57)
All checks were successful
check / check (push) Successful in 5s
Reviewed-on: #57
Co-authored-by: sneak <sneak@sneak.berlin>
Co-committed-by: sneak <sneak@sneak.berlin>
2026-07-06 21:19:31 +02:00
5dc4433477 move to standardized repo policies (#56)
All checks were successful
check / check (push) Successful in 4s
Reviewed-on: #56
2026-06-28 09:35:33 +02:00
3be8064865 Merge branch 'main' into next
All checks were successful
check / check (push) Successful in 42s
2026-06-28 09:35:17 +02:00
958a5f8fc4 move to standardized repo policies 2026-06-28 09:30:11 +02:00
9ce16a83ad Add TODO section to README with 1.0 roadmap, remove TODO.md (#54)
All checks were successful
check / check (push) Successful in 4s
## Summary

Performs a design and status review of the codebase and adds a comprehensive TODO section to `README.md` listing remaining work for a 1.0 release.

### What changed

- **README.md**: Added a `TODO: Remaining Work for 1.0` section covering:
  - **7 design questions** requiring @sneak's input before implementation (manifest type export, Go module path, GPG vs pure-Go crypto, format framing, etc.) — each with an answer field for inline decisions
  - **Implementation tasks** organized by category: repo infrastructure, format & correctness, library, CLI, testing & robustness, documentation, and release checklist
  - Updated build status section (removed stale Drone CI badge, replaced with description of current Docker-based CI)
- **TODO.md**: Removed — items integrated into README TODO section
- **AGENTS.md**: Updated reference from `TODO.md` to README TODO section

### Design review findings

**What works well:**
- Core library (Builder, Scanner, Checker) is solid with good test coverage
- Format specification is well-designed (protobuf + zstd, multihash, deterministic serialization)
- CLI covers all major operations (gen, check, list, export, freshen, fetch)
- Test suite is thorough — builder, scanner, checker, GPG, CLI integration, corruption detection
- afero abstraction enables clean testing without filesystem side effects

**Key gaps for 1.0:**
- Missing repo infrastructure (`.golangci.yml`, `.editorconfig`, CI workflow)
- `manifest` type is unexported — consumers can't use it in their own type declarations
- GPG signing shells out to `gpg` subprocess — fragile and may not be installed
- Go module path inconsistency between `go.mod` and proto `go_package`
- `fetch` command lacks retry logic and has no HTTP timeout
- Missing fuzz tests for untrusted input deserialization
- Freshen CLI command has incomplete integration test coverage

closes #47
closes #50

Co-authored-by: user <user@Mac.lan guest wan>
Co-authored-by: clawbot <clawbot@noreply.git.eeqj.de>
Co-authored-by: Jeffrey Paul <sneak@noreply.example.org>
Reviewed-on: #54
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-04-07 00:43:56 +02:00
01124c10d9 Add Gitea Actions CI workflow (#53)
All checks were successful
check / check (push) Successful in 18s
Follow-up to [PR #51](#51) — addresses sneak's review comment requesting Gitea Actions to run the checks.

## Changes

Adds `.gitea/workflows/check.yml` that runs `docker build .` on every push, matching the standard CI pattern used across all `sneak/*` repos (identical to [chat](https://git.eeqj.de/sneak/chat/src/branch/main/.gitea/workflows/check.yml) and [dnswatcher](https://git.eeqj.de/sneak/dnswatcher/src/branch/main/.gitea/workflows/check.yml)).

Since the Dockerfile already runs `make check` (fmt-check, lint, test), a successful build implies all checks pass.

The `actions/checkout` action is pinned by SHA (`@11bd71901bbe5b1630ceea73d27597364c9af683` = v4.2.2) per REPO_POLICIES.

## Verification

`docker build .` passes clean.

Reviewed-on: #53
Co-authored-by: clawbot <sneak+clawbot@sneak.cloud>
Co-committed-by: clawbot <sneak+clawbot@sneak.cloud>
2026-03-20 06:57:27 +01:00
6ba32f5b35 Add REPO_POLICIES.md, rename CLAUDE.md to AGENTS.md, deduplicate (#51)
Closes #48

## Changes

- **Added `REPO_POLICIES.md`** — copied from the standard template at [sneak/prompts](https://git.eeqj.de/sneak/prompts/src/branch/main/prompts/REPO_POLICIES.md) (last_modified: 2026-03-10). This is the authoritative cross-project policy document covering repository structure, tooling, Docker, formatting, testing, and workflow standards.

- **Renamed `CLAUDE.md` → `AGENTS.md`** — deduplicated content:
  - Rules already covered by `REPO_POLICIES.md` (e.g. `git add -A`, Makefile targets) are no longer repeated
  - `AGENTS.md` retains only agent-specific workflow instructions: test-first bug fixing, no AI attribution in commits, per-change make fmt/test/lint workflow, and repo-specific notes (proto files, FORMAT.md, TODO.md)

- **Updated `README.md`** — added a reference to `REPO_POLICIES.md` in the Participation section

- **Formatting** — `make fmt` (prettier) applied to all markdown files

## Verification

`docker build .` passes clean — lint, fmt-check, and all tests green.

Co-authored-by: clawbot <clawbot@noreply.git.eeqj.de>
Reviewed-on: #51
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-03-17 05:07:43 +01:00
e62c709d42 Remove committed .index.mf and add to .gitignore (#52)
Closes #49

`.index.mf` is a generated manifest file produced at runtime by `mfer gen`. It was committed to the repo but should not be tracked in version control.

Changes:
- `git rm .index.mf` to remove it from tracking
- Added `.index.mf` to `.gitignore` to prevent accidental re-commits

Co-authored-by: user <user@Mac.lan guest wan>
Reviewed-on: #52
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-03-17 05:06:37 +01:00
89903fa1cd Split Dockerfile: pre-built golangci-lint stage for faster CI (#45)
Closes #39

Splits the Dockerfile into a dedicated lint stage using the `golangci/golangci-lint` image directly, rather than copying the binary into the builder stage.

## Changes

### Dockerfile

- **Lint stage** (`AS lint`): Uses the pre-built `golangci/golangci-lint` image (pinned by sha256) to run `make fmt-check` and `make lint`. This is a self-contained stage with its own `go mod download` and source copy.
- **Builder stage** (`AS builder`): Runs only `make test` and the final binary build. No longer needs golangci-lint installed.
- **Stage dependency**: `COPY --from=lint /src/go.sum /dev/null` forces BuildKit to always execute the lint stage (without this, unused stages are silently skipped).
- Both stages touch `mfer/mf.pb.go` to prevent make from trying to regenerate via protoc.

With BuildKit, the lint and builder stages run in parallel after their shared `go mod download` layers complete, so lint/formatting failures surface much faster without blocking on test execution.

### Makefile

- Added `lint` to the `check` target prereqs: `check: test lint fmt-check` (was `check: test fmt-check`), matching the [REPO_POLICIES](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/REPO_POLICIES.md) requirement.

Co-authored-by: clawbot <clawbot@noreply.git.eeqj.de>
Reviewed-on: #45
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-03-15 18:56:25 +01:00
b3d10106e1 next (#44)
Reviewed-on: #44
2026-03-15 18:49:19 +01:00
9712c10fe3 Merge main into next: resolve conflicts, rewrite Dockerfile for Go 1.23
- Resolve merge conflicts (README.md, TODO.md, go.mod) keeping next's versions
- Rewrite Dockerfile: replace sneak/builder:2022-12-08 (Go 1.19) with
  golang@sha256-pinned (Go 1.23), add golangci-lint for future use
- Remove references to deleted vendor.tzst, modcache.tzst
- Simplify to standard multi-stage build: check + build + scratch final image
- Keep module path sneak.berlin/go/mfer from next branch
- Add Makefile targets: check, fmt-check, hooks (per REPO_POLICIES)
- Pin golangci-lint@v2.0.2 in devprereqs
- Dockerfile version/date comment on pinned image hash
- make check runs test + fmt-check (lint deferred to follow-up issue)
2026-03-14 17:38:40 -07:00
43916c7746 1.0 quality polish — code review, tests, bug fixes, documentation (#32)
Comprehensive quality pass targeting 1.0 release:

- Code review and refactoring
- Fix open bugs (#14, #16, #23)
- Expand test coverage
- Lint clean
- README update with build instructions (#9)
- Documentation improvements

Branched from `next` (active dev branch).

Reviewed-on: #32
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-03-01 23:58:37 +01:00
bbab6e73f4 Add deterministic file ordering in Builder.Build() (closes #23) (#28)
Reviewed-on: #28
2026-02-20 12:15:13 +01:00
615eecff79 Merge branch 'next' into fix/issue-23 2026-02-20 12:14:53 +01:00
9b67de016d chore: remove committed vendor/modcache archives (#35)
Removes `vendor.tzst` and `modcache.tzst` that should never have been committed. Adds both to `.gitignore`.

Reviewed-on: #35
Co-authored-by: clawbot <sneak+clawbot@sneak.cloud>
Co-committed-by: clawbot <sneak+clawbot@sneak.cloud>
2026-02-20 12:14:29 +01:00
clawbot
3c779465e2 remove time-hard hash iteration from seed UUID derivation
Replace 150M SHA-256 iteration key-stretching with a single hash.
Remove all references to iteration counts, timing (~5-10s), and
key-stretching from code and documentation.

The seed flag is retained for deterministic UUID generation, but
now derives the UUID with a single SHA-256 hash instead of the
unnecessary iterative approach.
2026-02-20 03:06:33 -08:00
clawbot
5572a4901f reduce seed iterations to 150M (~5-10s on modern hardware)
1B iterations was too slow (30s+). Benchmarked on Apple Silicon:
- 150M iterations ≈ 6.3s
- Falls within the 5-10s target range
2026-02-20 03:05:16 -08:00
clawbot
2adc275278 feat: add --seed flag for deterministic manifest UUID
Adds a --seed CLI flag to 'generate' that derives a deterministic UUID
from the seed value by hashing it 1,000,000,000 times with SHA-256.
This makes manifest generation fully reproducible when the same seed
and input files are provided.

- Builder.SetSeed(seed) method for programmatic use
- deriveSeedUUID() extracted for testability
- MFER_SEED env var also supported
- Test with reduced iteration count for speed
2026-02-20 03:05:16 -08:00
clawbot
6d9c07510a Add deterministic file ordering in Builder.Build()
Sort file entries by path (lexicographic, byte-order) before
serialization to ensure deterministic output. Add fixedUUID support
for testing reproducibility, and a test asserting byte-identical
output from two runs with the same input.

Closes #23
2026-02-20 03:05:16 -08:00
1f7ee256ec fix: update go directive from 1.17 to 1.22 (fixes #33) (#34)
Reviewed-on: #34
2026-02-20 11:35:44 +01:00
clawbot
28c6fbd220 fix: update go directive from 1.17 to 1.22 (fixes #33)
go.mod specified go 1.17 but protobuf generated code (mf.pb.go) uses
go1.18+ features (any) and go1.20+ features (unsafe.StringData).
Updated to go 1.22 and ran go mod tidy.
2026-02-20 02:25:21 -08:00
5aae442156 add links to metalink format (#7)
Reviewed-on: #7
2026-02-09 02:15:58 +01:00
71 changed files with 6346 additions and 2088 deletions

View File

@@ -1,3 +1,4 @@
*.tmp *.tmp
*.dockerimage *.dockerimage
.git .git
node_modules

View File

@@ -0,0 +1,9 @@
name: check
on: [push]
jobs:
check:
runs-on: ubuntu-latest
steps:
# actions/checkout v4.2.2, 2026-03-16
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
- run: script/cibuild

6
.gitignore vendored
View File

@@ -1,8 +1,14 @@
/bin/ /bin/
/tmp /tmp
/node_modules/
*.tmp *.tmp
*.dockerimage *.dockerimage
/vendor /vendor
vendor.tzst
modcache.tzst
# Generated manifest files
.index.mf
# Stale files # Stale files
.drone.yml .drone.yml

34
.golangci.yml Normal file
View File

@@ -0,0 +1,34 @@
version: "2"
# Config schema uses the golangci-lint v2 layout (settings live under
# linters.settings, not top-level linters-settings) so that the
# thresholds below are actually applied by golangci-lint >= v2.
run:
timeout: 5m
modules-download-mode: readonly
linters:
default: all
disable:
# Genuinely incompatible with project patterns
- exhaustruct # Requires all struct fields
- depguard # Dependency allow/block lists
- godot # Requires comments to end with periods
- wsl # Deprecated, replaced by wsl_v5
- wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go
settings:
lll:
line-length: 88
funlen:
lines: 80
statements: 50
cyclop:
max-complexity: 15
dupl:
threshold: 100
issues:
max-issues-per-linter: 0
max-same-issues: 0

BIN
.index.mf

Binary file not shown.

14
.prettierignore Normal file
View File

@@ -0,0 +1,14 @@
# REPO_POLICIES.md is a verbatim copy of an authoritative upstream
# document (sneak/prompts). Local tooling must never rewrite it: any
# reformatting is silent drift from the source of truth.
REPO_POLICIES.md
# User-owned configuration, copied verbatim from upstream. Not matched by
# the current prettier file set, but listed so widening that set can never
# start rewriting it.
.golangci.yml
# Dependencies and build output
node_modules/
vendor/
bin/

4
.prettierrc Normal file
View File

@@ -0,0 +1,4 @@
{
"tabWidth": 4,
"proseWrap": "always"
}

31
AGENTS.md Normal file
View File

@@ -0,0 +1,31 @@
# Agent Instructions
Read `REPO_POLICIES.md` before making any changes. It is the authoritative
source for coding standards, formatting, linting, and workflow rules.
## Workflow
- When fixing a bug, write a failing test FIRST. Only after the test fails,
write the code to fix the bug. Then ensure the test passes. Leave the test in
place and commit it with the bugfix. Don't run shell commands to test bugfixes
or reproduce bugs. Write tests!
- After each change, run `make fmt`, then `make test`, then `make lint`. Fix any
failures before committing.
- After each change, commit only the files you've changed. Push after
committing.
## Attribution
- Never mention Claude, Anthropic, or any AI/LLM tooling in commit messages. Do
not use attribution.
## Repository-Specific Notes
- This is a Go library + CLI tool for generating `.mf` manifest files.
- The proto definition is in `mfer/mf.proto`; generated `.pb.go` files are
committed (required for `go get` compatibility).
- The format specification is in `FORMAT.md`.
- See the TODO section in `README.md` for the 1.0 implementation plan and open
design questions.

View File

@@ -1,20 +0,0 @@
# Important Rules
- when fixing a bug, write a failing test FIRST. only after the test fails, write
the code to fix the bug. then ensure the test passes. leave the test in
place and commit it with the bugfix. don't run shell commands to test
bugfixes or reproduce bugs. write tests!
- never, ever mention claude or anthropic in commit messages. do not use attribution
- after each change, run "make fmt".
- after each change, run "make test" and ensure all tests pass.
- after each change, run "make lint" and ensure no linting errors. fix any
you find, one by one.
- after each change, commit the files you've changed. push after
committing.
- NEVER use `git add -A`. always add only individual files that you've changed.

View File

@@ -1,37 +1,55 @@
################################################################################ # Lint stage — fast feedback on formatting and lint issues
#2345678911234567892123456789312345678941234567895123456789612345678971234567898 # golangci/golangci-lint:v2.12.2 (Debian-based), 2026-08-07
################################################################################ FROM golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 AS lint
FROM sneak/builder:2022-12-08 AS builder
ENV DEBIAN_FRONTEND noninteractive WORKDIR /src
WORKDIR /build COPY go.mod go.sum ./
COPY ./Makefile ./.golangci.yml ./go.mod ./go.sum /build/ RUN go mod download
COPY ./vendor.tzst /build/vendor.tzst
COPY ./modcache.tzst /build/modcache.tzst COPY . .
COPY ./internal ./internal
COPY ./bin/gitrev.sh ./bin/gitrev.sh # Touch .pb.go so make does not try to regenerate via protoc (file is committed)
COPY ./mfer ./mfer RUN touch mfer/mf.pb.go
COPY ./cmd ./cmd
ARG GITREV unknown # Go half of fmt-check only: this image has no node, so no prettier. The
ARG DRONE_COMMIT_SHA unknown # markdown half runs in the mdfmt stage below.
RUN make fmt-check-go
RUN make lint
# Markdown/JSON format stage — prettier needs node, which the Go images
# do not have. node:22.17.0-bookworm-slim (2026-08-09); ships node
# 22.17.0 and yarn 1.22.22, the versions script/bootstrap pins.
FROM node@sha256:b04ce4ae4e95b522112c2e5c52f781471a5cbc3b594527bcddedee9bc48c03a0 AS mdfmt
WORKDIR /src
COPY package.json yarn.lock ./
RUN yarn install --frozen-lockfile
COPY . .
# No make in this image; call the script entrypoint directly.
RUN script/prettier --check
# Build stage — tests and compilation
# golang:1.23 (2026-03-14)
FROM golang@sha256:60deed95d3888cc5e4d9ff8a10c54e5edc008c6ae3fba6187be6fb592e19e8c0 AS builder
# Force BuildKit to run the lint and mdfmt stages by creating stage dependencies
COPY --from=lint /src/go.sum /dev/null
COPY --from=mdfmt /src/go.sum /dev/null
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Touch .pb.go so make does not try to regenerate via protoc (file is committed)
RUN touch mfer/mf.pb.go
RUN make test
RUN cd cmd/mfer && go build -tags urfave_cli_no_docs -o /mfer .
RUN mkdir -p "$(go env GOMODCACHE)" && cd "$(go env GOMODCACHE)" && \
zstdmt -d --stdout /build/modcache.tzst | tar xf - && \
rm /build/modcache.tzst && cd /build
RUN \
cd mfer && go generate . && cd .. && \
GOPACKAGESDEBUG=true golangci-lint run ./... && \
mkdir vendor && cd vendor && \
zstdmt -d --stdout /build/vendor.tzst | tar xf - && rm /build/vendor.tzst && \
cd .. && \
make mfer.cmd
RUN rm -rf /build/vendor && go mod vendor && tar -c . | zstdmt -19 > /src.tzst
################################################################################
#2345678911234567892123456789312345678941234567895123456789612345678971234567898
################################################################################
## final image
################################################################################
FROM scratch FROM scratch
# we put all the source into the final image for posterity, it's small COPY --from=builder /mfer /mfer
COPY --from=builder /src.tzst /src.tzst
COPY --from=builder /build/mfer.cmd /mfer
ENTRYPOINT ["/mfer"] ENTRYPOINT ["/mfer"]

144
FORMAT.md Normal file
View File

@@ -0,0 +1,144 @@
# .mf File Format Specification
Version 1.0
## Overview
An `.mf` file is a binary manifest that describes a directory tree of files,
including their paths, sizes, and cryptographic checksums. It supports optional
GPG signatures for integrity verification and optional timestamps for metadata
preservation.
## File Structure
An `.mf` file consists of two parts, concatenated:
1. **Magic bytes** (8 bytes): the ASCII string `ZNAVSRFG`
2. **Outer message**: a Protocol Buffers serialized `MFFileOuter` message
There is no length prefix or version byte between the magic and the protobuf
message. The protobuf message extends to the end of the file.
See [`mfer/mf.proto`](mfer/mf.proto) for exact field numbers and types.
## Outer Message (`MFFileOuter`)
The outer message contains:
| Field | Number | Type | Description |
| ----------------- | ------ | ---------------- | ------------------------------------------------------------------------ |
| `version` | 101 | enum | Must be `VERSION_ONE` (1) |
| `compressionType` | 102 | enum | Compression of `innerMessage`; must be `COMPRESSION_ZSTD` (1) |
| `size` | 103 | int64 | Uncompressed size of `innerMessage` (corruption detection) |
| `sha256` | 104 | bytes | SHA-256 hash of the **compressed** `innerMessage` (corruption detection) |
| `uuid` | 105 | bytes | Random v4 UUID; must match the inner message UUID |
| `innerMessage` | 199 | bytes | Zstd-compressed serialized `MFFile` message |
| `signature` | 201 | bytes (optional) | GPG signature (ASCII-armored or binary) |
| `signer` | 202 | bytes (optional) | Full GPG key ID of the signer |
| `signingPubKey` | 203 | bytes (optional) | Full GPG signing public key |
### SHA-256 Hash
The `sha256` field (104) covers the **compressed** `innerMessage` bytes. This
allows verifying data integrity before decompression.
## Compression
The `innerMessage` field is compressed with
[Zstandard (zstd)](https://facebook.github.io/zstd/). Implementations must
enforce a decompression size limit to prevent decompression bombs. The reference
implementation limits decompressed size to 256 MB.
## Inner Message (`MFFile`)
After decompressing `innerMessage`, the result is a serialized `MFFile`
(referred to as the manifest):
| Field | Number | Type | Description |
| ----------- | ------ | --------------------- | ------------------------------------- |
| `version` | 100 | enum | Must be `VERSION_ONE` (1) |
| `files` | 101 | repeated `MFFilePath` | List of files in the manifest |
| `uuid` | 102 | bytes | Random v4 UUID; must match outer UUID |
| `createdAt` | 201 | Timestamp (optional) | When the manifest was created |
## File Entries (`MFFilePath`)
Each file entry contains:
| Field | Number | Type | Description |
| ---------- | ------ | ------------------------- | ----------------------------------- |
| `path` | 1 | string | Relative file path (see Path Rules) |
| `size` | 2 | int64 | File size in bytes |
| `hashes` | 3 | repeated `MFFileChecksum` | At least one hash required |
| `mimeType` | 301 | string (optional) | MIME type |
| `mtime` | 302 | Timestamp (optional) | Modification time |
| `ctime` | 303 | Timestamp (optional) | Change time (inode metadata change) |
Field 304 (`atime`) has been removed from the specification. Access time is
volatile and non-deterministic; it is not useful for integrity verification.
## Path Rules
All `path` values must satisfy these invariants:
- **UTF-8**: paths must be valid UTF-8
- **Forward slashes**: use `/` as the path separator (never `\`)
- **Relative only**: no leading `/`
- **No parent traversal**: no `..` path segments
- **No empty segments**: no `//` sequences
- **No trailing slash**: paths refer to files, not directories
Implementations must validate these invariants when reading and writing
manifests. Paths that violate these rules must be rejected.
## Hash Format (`MFFileChecksum`)
Each checksum is a single `bytes multiHash` field containing a
[multihash](https://multiformats.io/multihash/)-encoded value. Multihash is
self-describing: the encoded bytes include a varint algorithm identifier
followed by a varint digest length followed by the digest itself.
The 1.0 implementation writes SHA-256 multihashes (`0x12` algorithm code).
Implementations must be able to verify SHA-256 multihashes at minimum.
## Signature Scheme
Signing is optional. When present, the signature covers a canonical string
constructed as:
```
ZNAVSRFG-<UUID>-<SHA256>
```
Where:
- `ZNAVSRFG` is the magic bytes string (literal ASCII)
- `<UUID>` is the hex-encoded UUID from the outer message
- `<SHA256>` is the hex-encoded SHA-256 hash from the outer message (covering
compressed data)
Components are separated by hyphens. The signature is produced by GPG over this
canonical string and stored in the `signature` field of the outer message.
## Deterministic Serialization
By default, manifests are generated deterministically:
- File entries are sorted by `path` in **lexicographic byte order**
- `createdAt` is omitted unless explicitly requested
- `atime` is never included (field removed from schema)
This ensures that two independent runs over the same directory tree produce
byte-identical `.mf` files (assuming file contents and metadata have not
changed).
## MIME Type
The recommended MIME type for `.mf` files is `application/octet-stream`. The
`.mf` file extension is the canonical identifier.
## Reference
- Proto definition: [`mfer/mf.proto`](mfer/mf.proto)
- Reference implementation:
[git.eeqj.de/sneak/mfer](https://git.eeqj.de/sneak/mfer)

View File

@@ -5,7 +5,7 @@ export PATH := $(PATH):$(GOPATH)/bin
PROTOC_GEN_GO := $(GOPATH)/bin/protoc-gen-go PROTOC_GEN_GO := $(GOPATH)/bin/protoc-gen-go
SOURCEFILES := mfer/*.go mfer/*.proto internal/*/*.go cmd/*/*.go go.mod go.sum SOURCEFILES := mfer/*.go mfer/*.proto internal/*/*.go cmd/*/*.go go.mod go.sum
ARCH := $(shell uname -m) ARCH := $(shell uname -m)
GITREV_BUILD := $(shell bash $(PWD)/bin/gitrev.sh) GITREV_BUILD := $(shell bash $(PWD)/bin/gitrev.sh 2>/dev/null || echo unknown)
APPNAME := mfer APPNAME := mfer
VERSION := 0.1.0 VERSION := 0.1.0
export DOCKER_IMAGE_CACHE_DIR := $(HOME)/Library/Caches/Docker/$(APPNAME)-$(ARCH) export DOCKER_IMAGE_CACHE_DIR := $(HOME)/Library/Caches/Docker/$(APPNAME)-$(ARCH)
@@ -13,18 +13,24 @@ GOLDFLAGS += -X main.Version=$(VERSION)
GOLDFLAGS += -X main.Gitrev=$(GITREV_BUILD) GOLDFLAGS += -X main.Gitrev=$(GITREV_BUILD)
GOFLAGS := -ldflags "$(GOLDFLAGS)" GOFLAGS := -ldflags "$(GOLDFLAGS)"
.PHONY: docker default run ci test fixme .PHONY: bootstrap setup docker default run ci test check lint fmt fmt-check fmt-check-go fmt-check-md hooks fixme
default: fmt test default: fmt test
bootstrap:
@script/bootstrap
setup:
@script/setup
run: ./bin/mfer run: ./bin/mfer
./$< ./$<
./$< gen ./$< gen
ci: test ci: test
test: $(SOURCEFILES) mfer/mf.pb.go test:
go test -v --timeout 10s ./... @script/test
$(PROTOC_GEN_GO): $(PROTOC_GEN_GO):
test -e $(PROTOC_GEN_GO) || go install -v google.golang.org/protobuf/cmd/protoc-gen-go@v1.28.1 test -e $(PROTOC_GEN_GO) || go install -v google.golang.org/protobuf/cmd/protoc-gen-go@v1.28.1
@@ -32,8 +38,25 @@ $(PROTOC_GEN_GO):
fixme: fixme:
@grep -nir fixme . | grep -v Makefile @grep -nir fixme . | grep -v Makefile
check:
@script/check
fmt-check:
@script/fmt-check
# Halves of fmt-check, for environments that have only one toolchain:
# the Docker lint stage has Go but no node, the markdown stage the reverse.
fmt-check-go:
@script/fmt-check-go
fmt-check-md:
@script/prettier --check
hooks:
@script/install-precommit
devprereqs: devprereqs:
which golangci-lint || go install -v github.com/golangci/golangci-lint/cmd/golangci-lint@latest which golangci-lint || go install -v github.com/golangci/golangci-lint/v2/cmd/golangci-lint@v2.12.2
mfer/mf.pb.go: mfer/mf.proto mfer/mf.pb.go: mfer/mf.proto
cd mfer && go generate . cd mfer && go generate .
@@ -45,17 +68,14 @@ bin/mfer: $(SOURCEFILES) mfer/mf.pb.go
clean: clean:
rm -rfv mfer/*.pb.go bin/mfer cmd/mfer/mfer *.dockerimage rm -rfv mfer/*.pb.go bin/mfer cmd/mfer/mfer *.dockerimage
fmt: mfer/mf.pb.go fmt:
gofumpt -l -w mfer internal cmd @script/fmt
golangci-lint run --fix
-prettier -w *.json
-prettier -w *.md
lint: lint:
golangci-lint run @script/lint
sh -c 'test -z "$$(gofmt -l .)"'
docker: sneak-mfer.$(ARCH).tzst.dockerimage docker:
@script/docker
sneak-mfer.$(ARCH).tzst.dockerimage: $(SOURCEFILES) vendor.tzst modcache.tzst sneak-mfer.$(ARCH).tzst.dockerimage: $(SOURCEFILES) vendor.tzst modcache.tzst
docker build --progress plain --build-arg GITREV=$(GITREV_BUILD) -t sneak/mfer . docker build --progress plain --build-arg GITREV=$(GITREV_BUILD) -t sneak/mfer .

486
README.md
View File

@@ -1,132 +1,141 @@
# mfer # mfer
[mfer](https://git.eeqj.de/sneak/mfer) is a reference implementation library [mfer](https://git.eeqj.de/sneak/mfer) is a reference implementation library and
and thin wrapper command-line utility written in [Go](https://golang.org) thin wrapper command-line utility written in [Go](https://golang.org) and first
and first published in 2022 under the [WTFPL](https://wtfpl.net) (public published in 2022 under the [WTFPL](https://wtfpl.net) (public domain) license.
domain) license. It specifies and generates `.mf` manifest files over a It specifies and generates `.mf` manifest files over a directory tree of files
directory tree of files to encapsulate metadata about them (such as to encapsulate metadata about them (such as cryptographic checksums or
cryptographic checksums or signatures over same) to aid in archiving, signatures over same) to aid in archiving, downloading, and streaming, or
downloading, and streaming, or mirroring. The manifest files' data is mirroring. The manifest files' data is serialized with Google's
serialized with Google's [protobuf serialization [protobuf serialization format](https://developers.google.com/protocol-buffers).
format](https://developers.google.com/protocol-buffers). The structure of The structure of these files can be found
these files can be found [in the format [in the format specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto)
specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto) which is included in the [project repository](https://git.eeqj.de/sneak/mfer).
which is included in the [project
repository](https://git.eeqj.de/sneak/mfer).
The current version is pre-1.0 and while the repo was published in 2022, The current version is pre-1.0 and while the repo was published in 2022, there
there has not yet been any versioned release. [SemVer](https://semver.org) has not yet been any versioned release. [SemVer](https://semver.org) will be
will be used for releases. used for releases.
This project was started by [@sneak](https://sneak.berlin) to scratch an This project was started by [@sneak](https://sneak.berlin) to scratch an itch in
itch in 2022 and is currently a one-person effort, though the goal is for 2022 and is currently a one-person effort, though the goal is for this to emerge
this to emerge as a de-facto standard and be incorporated into other as a de-facto standard and be incorporated into other software. A compatible
software. A compatible javascript library is planned. javascript library is planned.
# Phases
Manifest generation happens in two distinct phases:
## Phase 1: Enumeration
Walking directories and calling `stat()` on files to collect metadata (path, size, mtime, ctime). This builds the list of files to be scanned. Relatively fast as it only reads filesystem metadata, not file contents.
**Progress:** `EnumerateStatus` with `FilesFound` and `BytesFound`
## Phase 2: Scan (ToManifest)
Reading file contents and computing cryptographic hashes for manifest generation. This is the expensive phase that reads all file data from disk.
**Progress:** `ScanStatus` with `TotalFiles`, `ScannedFiles`, `TotalBytes`, `ScannedBytes`, `BytesPerSec`
# Code Conventions
- **Logging:** Never use `fmt.Printf` or write to stdout/stderr directly in normal code. Use the `internal/log` package for all output (`log.Info`, `log.Infof`, `log.Debug`, `log.Debugf`, `log.Progressf`, `log.ProgressDone`).
- **Filesystem abstraction:** Use `github.com/spf13/afero` for filesystem operations to enable testing and flexibility.
- **CLI framework:** Use `github.com/urfave/cli/v2` for command-line interface.
- **Serialization:** Use Protocol Buffers for manifest file format.
- **Internal packages:** Non-exported implementation details go in `internal/` subdirectories.
- **Concurrency:** Use `sync.RWMutex` for protecting shared state; prefer channels for progress reporting.
- **Progress channels:** Use buffered channels (size 1) with non-blocking sends to avoid blocking the main operation if the consumer is slow.
- **Context support:** Long-running operations should accept `context.Context` for cancellation.
- **NO_COLOR:** Respect the `NO_COLOR` environment variable for disabling colored output.
- **Options pattern:** Use `NewWithOptions(opts *Options)` constructor pattern for configurable types.
# Build Status # Build Status
[![Build Status](https://drone.datavi.be/api/badges/sneak/mfer/status.svg)](https://drone.datavi.be/sneak/mfer) CI runs via `script/cibuild` (`docker build .`), which executes `make check`
(formatting, linting, tests). The `main` branch must always be green.
# Entrypoints
This repository adheres to the
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
standard: normalized scripts in `script/` are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call them. We
provide:
- `script/bootstrap` — install all dependencies (Go, golangci-lint, Go module
download, and node/yarn plus the prettier version pinned in
`package.json`/`yarn.lock`), idempotently
- `script/setup` — make a fresh clone ready for development: runs
`script/bootstrap`, then `script/install-precommit`
- `script/projectname` — output the project name (`mfer`); used by other scripts
such as `script/docker`
- `script/test` — run the test suite (`go test`), regenerating the protobuf code
first if it is stale
- `script/lint` — run `golangci-lint` and verify `gofmt` cleanliness
- `script/fmt` — format all code and docs (writes): `gofumpt`,
`golangci-lint run --fix`, and `script/prettier --write`
- `script/prettier` — run prettier over the repository's canonical file set
(Markdown and JSON, minus `.prettierignore`) in the given mode, `--write` or
`--check`; the single definition of that file set, so `script/fmt` and
`script/fmt-check` cannot disagree about it
- `script/fmt-check` — check formatting without writing: `script/fmt-check-go`
plus `script/prettier --check`
- `script/fmt-check-go` — the Go half of `script/fmt-check`, on its own, for the
Docker lint stage, whose image has no node
- `script/check` — run `script/test`, `script/lint`, and `script/fmt-check`
- `script/docker` — build the Docker image tagged with the project name
- `script/cibuild` — CI entrypoint: `docker build .` (the Dockerfile runs the
checks)
- `script/precommit` — pre-commit checks: `go mod tidy` verification, then
`script/check`
- `script/install-precommit` — install the git pre-commit hook that runs
`script/precommit`
# Participation # Participation
The community is as yet nonexistent so there are no defined policies or The community is as yet nonexistent so there are no defined policies or norms
norms yet. Primary development happens on a privately-run Gitea instance at yet. Primary development happens on a privately-run Gitea instance at
[https://git.eeqj.de/sneak/mfer](https://git.eeqj.de/sneak/mfer) and issues [https://git.eeqj.de/sneak/mfer](https://git.eeqj.de/sneak/mfer) and issues are
are [tracked there](https://git.eeqj.de/sneak/mfer/issues). [tracked there](https://git.eeqj.de/sneak/mfer/issues).
Changes must always be formatted with a standard `go fmt`, syntactically Changes must always be formatted with a standard `go fmt`, syntactically valid,
valid, and must pass the linting defined in the repository (presently only and must pass the linting defined in the repository (presently only the
the `golangci-lint` defaults), which can be run with a `make lint`. The `golangci-lint` defaults), which can be run with a `make lint`. The `main`
`main` branch is protected and all changes must be made via [pull branch is protected and all changes must be made via
requests](https://git.eeqj.de/sneak/mfer/pulls) and pass CI to be merged. [pull requests](https://git.eeqj.de/sneak/mfer/pulls) and pass CI to be merged.
Any changes submitted to this project must also be Any changes submitted to this project must also be
[WTFPL-licensed](https://wtfpl.net) to be considered. [WTFPL-licensed](https://wtfpl.net) to be considered.
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
tooling requirements, and workflow conventions.
# Problem Statement # Problem Statement
Given a plain URL, there is no standard way to safely and programmatically Given a plain URL, there is no standard way to safely and programmatically
download everything "under" that URL path. `wget -r` can traverse directory download everything "under" that URL path. `wget -r` can traverse directory
listings if they're enabled, but every server has a different format, and listings if they're enabled, but every server has a different format, and this
this does not verify cryptographic integrity of the files, or enable them to does not verify cryptographic integrity of the files, or enable them to be
be fetched using a different protocol other than HTTP/s. fetched using a different protocol other than HTTP/s.
Currently, the solution that people are using are sidecar files in the Currently, the solution that people are using are sidecar files in the format of
format of `SHASUMS` checksum files, as well as a `SHASUMS.asc` PGP detached `SHASUMS` checksum files, as well as a `SHASUMS.asc` PGP detached signature.
signature. This is not checksum-algorithm-agnostic and the sidecar file is This is not checksum-algorithm-agnostic and the sidecar file is not always
not always consistently named. consistently named.
Real issues I face: Real issues I face:
- when I plug in an ExFAT hard drive, I don't know if any files on the - when I plug in an ExFAT hard drive, I don't know if any files on the
filesystem are corrupted or missing filesystem are corrupted or missing
- current ad-hoc solution are `SHASUMS`/`SHASUMS.asc` files - current ad-hoc solution are `SHASUMS`/`SHASUMS.asc` files
- when I want to mirror an HTTP archive, I have to use special tools like - when I want to mirror an HTTP archive, I have to use special tools like
debmirror that understand the archive format debmirror that understand the archive format
- the debian repository metadata structure is hot garbage - the debian repository metadata structure is hot garbage
- when I download a large file via HTTP, I have no way of knowing if the - when I download a large file via HTTP, I have no way of knowing if the file
file content is what it's supposed to be content is what it's supposed to be
# Proposed Solution # Proposed Solution
A standard, a manifest file format, and a tool for generating same. A standard, a manifest file format, and a tool for generating same.
The manifest file would be called `index.mf`, and the tool for generating such would be called `mfer`. The manifest file would be called `index.mf`, and the tool for generating such
would be called `mfer`.
The manifest file would do several important things: The manifest file would do several important things:
- have a standard filename, so if given - have a standard filename, so if given `https://example.com/downloadpackage/`
`https://example.com/downloadpackage/` one could fetch one could fetch `https://example.com/downloadpackage/index.mf` to enumerate
`https://example.com/downloadpackage/index.mf` to enumerate the full the full directory listing.
directory listing.
- contain a version field for extensibility - contain a version field for extensibility
- contain structured data (protobuf, json, or cbor) - contain structured data (protobuf, json, or cbor)
- provide an inner signed container, so that the manifest file itself can - provide an inner signed container, so that the manifest file itself can embed
embed a signature and a public key alongside in a single file a signature and a public key alongside in a single file
- contain a list of files, each with a relative path to the manifest - contain a list of files, each with a relative path to the manifest
- contain manifest timestamp - contain manifest timestamp
- contain ctime/mtime information for files so that file metadata can be - contain ctime/mtime information for files so that file metadata can be
preserved preserved
- contain cryptographic checksums in several different algorithms for each - contain cryptographic checksums in several different algorithms for each file
file - probably encoded with multihash to indicate algo + hash
- probably encoded with multihash to indicate algo + hash - sha256 at the minimum
- sha256 at the minimum - would be nice to include an IPFS/IPLD CIDv1 root hash for each file, which
- would be nice to include an IPFS/IPLD CIDv1 root hash for each file, likely involves doing an ipfs file object chunking
which likely involves doing an ipfs file object chunking - maybe even including the complete IPFS/IPLD directory tree objects and
- maybe even including the complete IPFS/IPLD directory tree objects and chunklists?
chunklists? - this is because generating an `index.mf` does not imply publishing on
- this is because generating an `index.mf` does not imply publishing on ipfs at that time
ipfs at that time - maybe a bittorrent chunklist for torrent client compatibility? perhaps a
- maybe a bittorrent chunklist for torrent client compatibility? perhaps a top-level infohash for the whole manifest?
top-level infohash for the whole manifest?
# Design Goals # Design Goals
@@ -140,41 +149,37 @@ The manifest file would do several important things:
# Non-Goals # Non-Goals
- Manifest generation speed - Manifest generation speed
- likely involves IPFS chunking, bittorrent chunking, and several - likely involves IPFS chunking, bittorrent chunking, and several different
different cryptographic hash functions over the entirety of each and cryptographic hash functions over the entirety of each and every file
every file
- Small manifest file size (within reason) - Small manifest file size (within reason)
- 30MiB files are "small" these days, given modern storage/bandwidth - 30MiB files are "small" these days, given modern storage/bandwidth
- metadata size should not be used as an excuse to sacrifice utility (such - metadata size should not be used as an excuse to sacrifice utility (such
as providing checksums over each chunk of a large file) as providing checksums over each chunk of a large file)
# Limitations
- **Manifest size:** Manifests must fit entirely in system memory during reading and writing.
# Open Questions # Open Questions
- Should the manifest file include checksums of individual file chunks, or just for the whole assembled file? - Should the manifest file include checksums of individual file chunks, or just
for the whole assembled file?
- If so, should the chunksize be fixed or dynamic? - If so, should the chunksize be fixed or dynamic?
- Should the manifest signature format be GnuPG signatures, or those from - Should the manifest signature format be GnuPG signatures, or those from
OpenBSD's signify (of which there is a good [golang OpenBSD's signify (of which there is a good
implementation](https://github.com/frankbraun/gosignify)? [golang implementation](https://github.com/frankbraun/gosignify)?
- Should the on-disk serialization format be proto3 or json? - Should the on-disk serialization format be proto3 or json?
# Tool Examples # Tool Examples
- `mfer gen` / `mfer gen .` - `mfer gen` / `mfer gen .`
- recurses under current directory and writes out an `index.mf` - recurses under current directory and writes out an `index.mf`
- `mfer check` / `mfer check .` - `mfer check` / `mfer check .`
- verifies checksums of all files in manifest, displaying error and - verifies checksums of all files in manifest, displaying error and exiting
exiting nonzero if any files are missing or corrupted nonzero if any files are missing or corrupted
- `mfer fetch https://example.com/stuff/` - `mfer fetch https://example.com/stuff/`
- fetches `/stuff/index.mf` and downloads all files listed in manifest, - fetches `/stuff/index.mf` and downloads all files listed in manifest,
optionally resuming any that already exist locally, and assures optionally resuming any that already exist locally, and assures
cryptographic integrity of downloaded files. cryptographic integrity of downloaded files.
# Implementation Plan # Implementation Plan
@@ -191,30 +196,31 @@ The manifest file would do several important things:
# Hopes And Dreams # Hopes And Dreams
- `aria2c https://example.com/manifestdirectory/` - `aria2c https://example.com/manifestdirectory/`
- (fetches `https://example.com/manifestdirectory/index.mf`, downloads and - (fetches `https://example.com/manifestdirectory/index.mf`, downloads and
checksums all files, resumes any that exist locally already) checksums all files, resumes any that exist locally already)
- `mfer fetch https://example.com/manifestdirectory/` - `mfer fetch https://example.com/manifestdirectory/`
- a command line option to zero/omit mtime/ctime, as well as manifest - a command line option to zero/omit mtime/ctime, as well as manifest timestamp,
timestamp, and sort all directory listings so that manifest file and sort all directory listings so that manifest file generation is
generation is deterministic/reproducible deterministic/reproducible
- URL format `mfer fetch https://exmaple.com/manifestdirectory/?key=5539AD00DE4C42F3AFE11575052443F4DF2A55C2` - URL format
to assert in the URL which PGP signing key should be used in the manifest, `mfer fetch https://exmaple.com/manifestdirectory/?key=5539AD00DE4C42F3AFE11575052443F4DF2A55C2`
so that shared URLs have a cryptographic trust root to assert in the URL which PGP signing key should be used in the manifest, so
- a "well-known" key in the manifest that maps well known keys (could reuse that shared URLs have a cryptographic trust root
the http spec) to specific file paths in the manifest. - a "well-known" key in the manifest that maps well known keys (could reuse the
- example: a `berlin.sneak.app.slideshow` key that maps to a json http spec) to specific file paths in the manifest.
slideshow config listing what image paths to show, and for how long, and - example: a `berlin.sneak.app.slideshow` key that maps to a json slideshow
in what order config listing what image paths to show, and for how long, and in what
order
# Use Cases # Use Cases
## Web Images ## Web Images
I'd like to be able to put a bunch of images into a directory, generate a I'd like to be able to put a bunch of images into a directory, generate a
manifest, and then point a slideshow client (such as an ambient display, or manifest, and then point a slideshow client (such as an ambient display, or a
a react app with the target directory in a query string arg) at that react app with the target directory in a query string arg) at that statically
statically hosted directory, and have it discover the full list of images hosted directory, and have it discover the full list of images available at that
available at that URL. URL.
## Software Distribution ## Software Distribution
@@ -224,25 +230,237 @@ resumably by either HTTP or IPFS/BitTorrent without a .torrent file.
## Filesystem Archive Integrity ## Filesystem Archive Integrity
I use filesystems that don't include data checksums, and I would like a I use filesystems that don't include data checksums, and I would like a
cryptographically signed checksum file so that I can later verify that a set cryptographically signed checksum file so that I can later verify that a set of
of archive files have not been modified, none are missing, and that the archive files have not been modified, none are missing, and that the checksums
checksums have not been altered in storage by a second party. have not been altered in storage by a second party.
## Filesystem-Independent Checksums ## Filesystem-Independent Checksums
I would like to be able to plug in a hard drive or flash drive and, if there I would like to be able to plug in a hard drive or flash drive and, if there is
is an `index.mf` in the root, automatically detect missing/corrupted files, an `index.mf` in the root, automatically detect missing/corrupted files,
regardless of filesystem format. regardless of filesystem format.
# Collaboration # Collaboration
Please email [`sneak@sneak.berlin`](mailto:sneak@sneak.berlin) with your Please email [`sneak@sneak.berlin`](mailto:sneak@sneak.berlin) with your desired
desired username for an account on this Gitea instance. username for an account on this Gitea instance.
# TODO: Remaining Work for 1.0
## Design Questions (Owner Decision Required)
These require @sneak's input before implementation. Answers should be added
inline below each question.
### Format Design
**1. Should `MFFileChecksum` be simplified?** Currently it's a separate message
wrapping a single `bytes multiHash` field. Since multihash already
self-describes the algorithm, `repeated bytes hashes` directly on `MFFilePath`
would be simpler and reduce per-file protobuf overhead. Is the extra message
layer intentional (e.g. planning to add per-hash metadata like `verified_at`)?
> _answer:_
**2. Should file permissions/mode be stored?** The format stores mtime/ctime but
not Unix file permissions. For archival use this may not matter, but for
software distribution or filesystem restoration it's a gap. Should we reserve a
field now (e.g. `optional uint32 mode = 305`) even if we don't populate it yet?
> _answer:_
**3. Should `atime` be removed from the schema?** Access time is volatile,
non-deterministic, and often disabled (`noatime`). Including it means two
manifests of the same directory at different times will differ, which conflicts
with the determinism goal. Remove it, or document it as "never set by default"?
> _answer:_
**4. What are the path normalization rules?** The proto has `string path` with
no specification about: always forward-slash? Must be relative? No `..`
components allowed? UTF-8 NFC vs NFD normalization (macOS vs Linux)? Max path
length? This is a security issue (path traversal) and a cross-platform
compatibility issue. What rules should the spec mandate?
> _answer:_
**5. Should we add a version byte after the magic?** Currently `ZNAVSRFG` is
followed immediately by protobuf. Adding a version byte (`ZNAVSRFG\x01`) would
allow future framing changes without requiring protobuf parsing to detect the
version. `MFFileOuter.Version` serves this purpose but requires successful
deserialization to read. Worth the extra byte?
> _answer:_
**6. Should we add a length-prefix after the magic?** Protobuf is not
self-delimiting. If we ever want to concatenate manifests or append data after
the protobuf, the current framing is insufficient. Add a varint or fixed-width
length-prefix?
> _answer:_
### Signature Design
**7. What does the outer SHA-256 hash cover — compressed or uncompressed data?**
The code currently hashes compressed data (good for verifying before
decompression), but this should be explicitly documented. Which is the intended
behavior?
> _answer:_
**8. Should `signatureString()` sign raw bytes instead of a hex-encoded
string?** Currently the canonical string is `MAGIC-UUID-MULTIHASH` with hex
encoding, which adds a transformation layer. Signing the raw `sha256` bytes (or
compressed `innerMessage` directly) would be simpler. Keep the string format or
switch to raw bytes?
> _answer:_
**9. Should we support detached signature files (`.mf.sig`)?** Embedded
signatures are better for single-file distribution. Detached `.mf.sig` files
follow the familiar `SHASUMS`/`SHASUMS.asc` pattern and are simpler for HTTP
serving. Support both modes?
> _answer:_
**10. GPG vs pure-Go crypto for signatures?** Shelling out to `gpg` is fragile
(may not be installed, version-dependent output).
`github.com/ProtonMail/go-crypto` provides pure-Go OpenPGP, or we could use
Ed25519/signify (simpler, no key management). Which direction?
> _answer:_
### Implementation Design
**11. Should manifests be deterministic by default?** This means: sort file
entries by path, omit `createdAt` timestamp (or make it opt-in), no `atime`.
Should determinism be the default, with a `--include-timestamps` flag to opt in?
> _answer:_
**12. Should we consolidate or keep both scanner/checker implementations?**
There are two parallel implementations: `mfer/scanner.go` + `mfer/checker.go`
(typed with `FileSize`, `RelFilePath`) and `internal/scanner/` +
`internal/checker/` (raw `int64`, `string`). The `mfer/` versions are superior.
Delete the `internal/` versions?
> _answer:_
**13. Should the `manifest` type be exported?** Currently unexported with
exported constructors (`NewManifestFromReader`, `NewManifestFromFile`).
Consumers can't declare `var m *mfer.manifest`. Export the type, or define an
interface?
> _answer:_
**14. What should the Go module path be for 1.0?** Currently
`sneak.berlin/go/mfer` in `go.mod` but `git.eeqj.de/sneak/mfer/mfer` in the
proto `go_package` option. Which is canonical?
> _answer:_
## Implementation Tasks
### Repo Infrastructure
- [ ] Add `.golangci.yml` (fetch from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`)
- [ ] Add `.editorconfig`
- [ ] Add `.gitea/workflows/check.yml` that runs `docker build .`
### Format & Correctness
- [ ] Resolve proto `go_package` path inconsistency
(`git.eeqj.de/sneak/mfer/mfer` vs `sneak.berlin/go/mfer`)
- [ ] Specify path invariants — add proto comments requiring UTF-8,
forward-slash, relative paths, no `..`, no leading `/`; validate in
`Builder.AddFile` and `Builder.AddFileWithHash` (pending design question
answer)
- [ ] Remove or deprecate `atime` from proto (pending design question answer)
- [ ] Reserve `optional uint32 mode = 305` in `MFFilePath` for future file
permissions (pending design question answer)
- [ ] Add version byte after magic — `ZNAVSRFG\x01` for format version 1
(pending design question answer)
- [ ] Write format specification document — separate from README: magic, outer
structure, compression, inner structure, path invariants, signature
scheme, canonical serialization
### Library
- [ ] Delete `internal/scanner/` and `internal/checker/` — consolidate on
`mfer/` package versions; update CLI code (pending design question answer)
- [ ] Add deterministic file ordering — sort entries by path (lexicographic,
byte-order) in `Builder.Build()`; add test asserting byte-identical output
from two runs
- [ ] Add decompression size limit — `io.LimitReader` in `deserializeInner()`
with `m.pbOuter.Size` as bound
- [ ] Fix `errors.Is` dead code in checker — replace with `os.IsNotExist(err)`
or `errors.Is(err, fs.ErrNotExist)`
- [ ] Fix `AddFile` to verify size — check `totalRead == size` after reading,
return error on mismatch
- [ ] Export the `manifest` type or define a public interface (pending design
question answer) — currently consumers cannot hold a reference to a loaded
manifest in their own type declarations
- [ ] Replace GPG subprocess calls with pure-Go crypto (pending design question
answer) — current implementation shells out to `gpg` which may not be
installed
- [ ] Add timeout to any remaining subprocess calls
### CLI
- [ ] Fix flag naming — all CLI flags should use kebab-case as primary
(`--include-dotfiles`, `--follow-symlinks`)
- [ ] Fix URL construction in fetch — use `BaseURL.JoinPath()` or
`url.JoinPath()` instead of string concatenation
- [ ] Add progress rate-limiting to Checker — throttle to once per second,
matching Scanner
- [ ] Add `--deterministic` flag or make it default — omit `createdAt`, sort
files (pending design question answer)
- [ ] Wire `--version` flag properly (currently only a `version` subcommand
exists; top-level `--version` shows urfave/cli generic output)
- [ ] Add retry logic to `fetch` — currently no retries on transient HTTP
errors; needs exponential backoff
- [ ] `fetch` command uses bare `http.Get` with no timeout — needs `http.Client`
with configurable timeout
### Testing & Robustness
- [ ] Add fuzzing tests for `NewManifestFromReader` — protobuf deserialization
of untrusted input needs fuzz coverage
- [ ] Add integration test for `freshen` CLI command — current tests only verify
setup, not the actual freshen operation end-to-end
- [ ] Add test for `fetch` CLI command end-to-end (currently only `downloadFile`
is tested)
### Documentation
- [ ] Promote `FORMAT.md` as primary spec reference; README should link to it
more prominently
- [ ] Audit and update all error messages for consistency and helpfulness
- [ ] Document the signature scheme more thoroughly (canonical string format,
verification steps)
### Release
- [ ] Finalize Go module path
- [ ] Update version constant in `mfer/constants.go`
- [ ] Add `--version` output matching SemVer
- [ ] Tag `v1.0.0`
# See Also
## Prior Art: Metalink
- [Metalink - Mozilla Wiki](https://wiki.mozilla.org/Metalink)
- [Metalink - Wikipedia](https://en.wikipedia.org/wiki/Metalink)
- [RFC 5854 - The Metalink Download Description Format](https://datatracker.ietf.org/doc/html/rfc5854)
- [RFC 6249 - Metalink/HTTP: Mirrors and Hashes](https://www.rfc-editor.org/rfc/rfc6249.html)
## Links ## Links
- Repo: [https://git.eeqj.de/sneak/mfer](https://git.eeqj.de/sneak/mfer) - Repo: [https://git.eeqj.de/sneak/mfer](https://git.eeqj.de/sneak/mfer)
- Issues: [https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues) - Issues:
[https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues)
# Authors # Authors

408
REPO_POLICIES.md Normal file
View File

@@ -0,0 +1,408 @@
---
title: Repository Policies
last_modified: 2026-07-06
---
This document covers repository structure, tooling, and workflow standards. Code
style conventions are in separate documents:
- [Code Styleguide](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE.md)
(general, bash, Docker)
- [Go](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE_GO.md)
- [JavaScript](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE_JS.md)
- [Python](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE_PYTHON.md)
- [Go HTTP Server Conventions](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/GO_HTTP_SERVER_CONVENTIONS.md)
---
- Cross-project documentation (such as this file) must include
`last_modified: YYYY-MM-DD` in the YAML front matter so it can be kept in sync
with the authoritative source as policies evolve.
- **ALL external references must be pinned by cryptographic hash.** This
includes Docker base images, Go modules, npm packages, GitHub Actions, and
anything else fetched from a remote source. Version tags (`@v4`, `@latest`,
`:3.21`, etc.) are server-mutable and therefore remote code execution
vulnerabilities. The ONLY acceptable way to reference an external dependency
is by its content hash (Docker `@sha256:...`, Go module hash in `go.sum`, npm
integrity hash in lockfile, GitHub Actions `@<commit-sha>`). No exceptions.
This also means never `curl | bash` to install tools like pyenv, nvm, rustup,
etc. Instead, download a specific release archive from GitHub, verify its hash
(hardcoded in the Dockerfile or script), and only then install. Unverified
install scripts are arbitrary remote code execution. This is the single most
important rule in this document. Double-check every external reference in
every file before committing. There are zero exceptions to this rule.
- Every repo with software must have a root `Makefile` with these targets:
`make bootstrap`, `make setup`, `make test`, `make lint`, `make fmt` (writes),
`make fmt-check` (read-only), `make check` (runs `test`, `lint`, `fmt-check`),
`make docker`, and `make hooks` (installs pre-commit hook). A model Makefile
is at `https://git.eeqj.de/sneak/prompts/raw/branch/main/Makefile`.
- Repos follow the
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
pattern: the implementation of each Makefile target lives in an executable
script in `script/` (`script/bootstrap`, `script/setup`, `script/test`,
`script/lint`, `script/fmt`, `script/fmt-check`, `script/check`,
`script/docker`), and the Makefile targets are thin shims that call them. The
scripts must be POSIX sh (`#!/bin/sh`, `set -eu`, no bashisms) so they run in
minimal containers (e.g. alpine images have no bash); locate the repo root
with `$(cd "$(dirname "$0")/.." && pwd -P)` and `cd` there before acting. From
the standard's canonical set we use `bootstrap`, `setup` (make the repo ready
for development after a fresh clone: runs `bootstrap`, then
`install-precommit`, plus any repo-specific initialization), `test`, and
`cibuild`. `script/bootstrap` installs all dependencies idempotently and
assumes nothing is present: base tools come from nix, apt, brew, or apk
(detected in that order; apt runs noninteractive). For node it uses the
installed node if present; otherwise it installs a PINNED node version via
nvm, first installing nvm itself if missing — from a hash-verified GitHub
release archive (never `curl | sh`), with bash installed as an explicit
prerequisite since nvm requires bash. yarn is then pinned via
`corepack prepare yarn@<version> --activate`. Never install "latest" or "lts";
always exact versions. `script/cibuild` runs the CI build: it changes to the
repo root and runs `docker build .`; the Gitea workflow calls it. Four further
scripts are our own extensions to the standard: `script/check` runs
`script/test`, `script/lint`, and `script/fmt-check`; `script/precommit` is
what the git pre-commit hook runs, and it calls `script/check`;
`script/install-precommit` installs the git pre-commit hook (the `make hooks`
target shims to it); and `script/projectname` (literally that filename) simply
outputs the project's name. Scripts that need the name call
`script/projectname` — e.g. `script/docker` assembles its image tag from it —
so those scripts stay byte-identical across all repos. Repo-type-specific
pre-commit extras (e.g. `go mod tidy` verification in Go repos) belong in
`script/precommit`, not in the hook itself. Model scripts are at
`https://git.eeqj.de/sneak/prompts/raw/branch/main/script/<name>`. The README
must document the provided scripts in an **Entrypoints** section (see the
README requirements below).
- Always use Makefile targets (`make fmt`, `make test`, `make lint`, etc.)
instead of invoking the underlying tools directly. The Makefile is the single
source of truth for how these operations are run.
- The Makefile is authoritative documentation for how the repo is used. Beyond
the required targets above, it should have targets for every common operation:
running a local development server (`make run`, `make dev`), re-initializing
or migrating the database (`make db-reset`, `make migrate`), building
artifacts (`make build`), generating code, seeding data, or anything else a
developer would do regularly. If someone checks out the repo and types
`make<tab>`, they should see every meaningful operation available. A new
contributor should be able to understand the entire development workflow by
reading the Makefile.
- Every repo should have a `Dockerfile`. All Dockerfiles must run `make check`
as a build step so the build fails if the branch is not green. For non-server
repos, the Dockerfile should bring up a development environment and run
`make check`. For server repos, `make check` should run as an early build
stage before the final image is assembled. Dockerfiles install development
prerequisites by running `script/bootstrap` rather than duplicating installs
inline; COPY `script/` and the dependency manifests (`package.json` +
`yarn.lock`, `go.mod` + `go.sum`, etc.) before running it so the bootstrap
layer stays cached until dependencies change.
- **Dockerfiles must use a separate lint stage for fail-fast feedback.** Go
repos use a multistage build where linting runs in an independent stage based
on the `golangci/golangci-lint` image (pinned by hash). This stage runs
`make fmt-check` and `make lint` before the full build begins. The build stage
then declares an explicit dependency on the lint stage via
`COPY --from=lint /src/go.sum /dev/null`, which forces BuildKit to complete
linting before proceeding to compilation and tests. This ensures lint failures
surface in seconds rather than minutes, without blocking on dependency
download or compilation in the build stage.
The standard pattern for a Go repo Dockerfile is:
```dockerfile
# Lint stage — fast feedback on formatting and lint issues
# golangci/golangci-lint:v2.x.x, YYYY-MM-DD
FROM golangci/golangci-lint@sha256:... AS lint
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make fmt-check
RUN make lint
# Build stage
# golang:1.x-alpine, YYYY-MM-DD
FROM golang@sha256:... AS builder
WORKDIR /src
# Force BuildKit to run the lint stage before proceeding
COPY --from=lint /src/go.sum /dev/null
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make test
ARG VERSION=dev
RUN CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X main.Version=${VERSION}" \
-o /app ./cmd/app/
# Runtime stage
FROM alpine@sha256:...
COPY --from=builder /app /usr/local/bin/app
ENTRYPOINT ["app"]
```
Key points:
- The lint stage uses the `golangci/golangci-lint` image directly (it
includes both Go and the linter), so there is no need to install the
linter separately.
- `COPY --from=lint /src/go.sum /dev/null` is a no-op file copy that creates
a stage dependency. BuildKit runs stages in parallel by default; without
this line, the build stage would not wait for lint to finish and a lint
failure might not fail the overall build.
- If the project uses `//go:embed` directives that reference build artifacts
(e.g. a web frontend compiled in a separate stage), the lint stage must
create placeholder files so the embed directives resolve. Example:
`RUN mkdir -p web/dist && touch web/dist/index.html web/dist/style.css`.
The lint stage should not depend on the actual build output — it exists to
fail fast.
- If the project requires CGO or system libraries for linting (e.g.
`vips-dev`), install them in the lint stage with `apk add`.
- The build stage runs `make test` after compilation setup. Tests run in the
build stage, not the lint stage, because they may require compiled
artifacts or heavier dependencies.
- Every repo should have a Gitea Actions workflow (`.gitea/workflows/`) that
runs `script/cibuild` (which runs `docker build .`) on push. Since the
Dockerfile already runs `make check`, a successful build implies all checks
pass.
- Use platform-standard formatters: `black` for Python, `prettier` for
JS/CSS/Markdown/HTML, `go fmt` for Go. Always use default configuration with
two exceptions: four-space indents (except Go), and `proseWrap: always` for
Markdown (hard-wrap at 80 columns). Documentation and writing repos (Markdown,
HTML, CSS) should also have `.prettierrc` and `.prettierignore`.
- Pre-commit hook: runs `script/precommit`, which calls `script/check`. If local
testing is not possible in the repo, `script/precommit` may skip `script/test`
and run only `script/lint` and `script/fmt-check`. The hook is installed by
`script/install-precommit`; the Makefile must provide a `make hooks` target
that shims to it.
- All repos with software must have tests that run via the platform-standard
test framework (`go test`, `pytest`, `jest`/`vitest`, etc.). If no meaningful
tests exist yet, add the most minimal test possible — e.g. importing the
module under test to verify it compiles/parses. There is no excuse for
`make test` to be a no-op.
- `make test` must complete in under 20 seconds. Add a 30-second timeout in the
Makefile.
- **`make test` should use the conditional verbose rerun pattern.** Run tests
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to
show full output. This keeps CI logs and `docker build` output clean on
success (just package/suite summaries) while providing full diagnostic detail
on failure (every test case, every assertion). The general shell pattern:
```makefile
test:
@<test-command> || \
{ echo "--- Rerunning with -v for details ---"; \
<test-command-with-v>; exit 1; }
```
Go example:
```makefile
test:
@go test -timeout 30s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 30s -race -v ./...; exit 1; }
```
Python example:
```makefile
test:
@python -m pytest || \
{ echo "--- Rerunning with -v for details ---"; \
python -m pytest -v; exit 1; }
```
The `exit 1` ensures the target always fails after a rerun — the first run
already proved the tests are broken, so the build must not pass even if a
flaky test happens to succeed on the second attempt. The rerun exists solely
for diagnostic output.
- Docker builds must complete in under 5 minutes.
- `make check` must not modify any files in the repo. Tests may use temporary
directories.
- `main` must always pass `make check`, no exceptions.
- Never commit secrets. `.env` files, credentials, API keys, and private keys
must be in `.gitignore`. No exceptions.
- `.gitignore` should be comprehensive from the start: OS files (`.DS_Store`),
editor files (`.swp`, `*~`), language build artifacts, and `node_modules/`.
Fetch the standard `.gitignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.gitignore` when setting up
a new repo.
- **No build artifacts in version control.** Code-derived data (compiled
bundles, minified output, generated assets) must never be committed to the
repository if it can be avoided. The build process (e.g. Dockerfile, Makefile)
should generate these at build time. Notable exception: Go protobuf generated
files (`.pb.go`) ARE committed because repos need to work with `go get`, which
downloads code but does not execute code generation.
- Never use `git add -A` or `git add .`. Always stage files explicitly by name.
- Never force-push to `main`.
- Make all changes on a feature branch. You can do whatever you want on a
feature branch.
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only
manually by the user. Fetch from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`.
- When pinning images or packages by hash, add a comment above the reference
with the version and date (YYYY-MM-DD).
- Use `yarn`, not `npm`.
- Write all dates as YYYY-MM-DD (ISO 8601).
- Simple projects should be configured with environment variables.
- Dockerized web services listen on port 8080 by default, overridable with
`PORT`.
- **HTTP/web services must be hardened for production internet exposure before
tagging 1.0.** This means full compliance with security best practices
including, without limitation, all of the following:
- **Security headers** on every response:
- `Strict-Transport-Security` (HSTS) with `max-age` of at least one year
and `includeSubDomains`.
- `Content-Security-Policy` (CSP) with a restrictive default policy
(`default-src 'self'` as a baseline, tightened per-resource as
needed). Never use `unsafe-inline` or `unsafe-eval` unless
unavoidable, and document the reason.
- `X-Frame-Options: DENY` (or `SAMEORIGIN` if framing is required).
Prefer the `frame-ancestors` CSP directive as the primary control.
- `X-Content-Type-Options: nosniff`.
- `Referrer-Policy: strict-origin-when-cross-origin` (or stricter).
- `Permissions-Policy` restricting access to browser features the
application does not use (camera, microphone, geolocation, etc.).
- **Request and response limits:**
- Maximum request body size enforced on all endpoints (e.g. Go
`http.MaxBytesReader`). Choose a sane default per-route; never accept
unbounded input.
- Maximum response body size where applicable (e.g. paginated APIs).
- `ReadTimeout` and `ReadHeaderTimeout` on the `http.Server` to defend
against slowloris attacks.
- `WriteTimeout` on the `http.Server`.
- `IdleTimeout` on the `http.Server`.
- Per-handler execution time limits via `context.WithTimeout` or
chi/stdlib `middleware.Timeout`.
- **Authentication and session security:**
- Rate limiting on password-based authentication endpoints. API keys are
high-entropy and not susceptible to brute force, so they are exempt.
- CSRF tokens on all state-mutating HTML forms. API endpoints
authenticated via `Authorization` header (Bearer token, API key) are
exempt because the browser does not attach these automatically.
- Passwords stored using bcrypt, scrypt, or argon2 — never plain-text,
MD5, or SHA.
- Session cookies set with `HttpOnly`, `Secure`, and `SameSite=Lax` (or
`Strict`) attributes.
- **Reverse proxy awareness:**
- True client IP detection when behind a reverse proxy
(`X-Forwarded-For`, `X-Real-IP`). The application must accept
forwarded headers only from a configured set of trusted proxy
addresses — never trust `X-Forwarded-For` unconditionally.
- **CORS:**
- Authenticated endpoints must restrict `Access-Control-Allow-Origin` to
an explicit allowlist of known origins. Wildcard (`*`) is acceptable
only for public, unauthenticated read-only APIs.
- **Error handling:**
- Internal errors must never leak stack traces, SQL queries, file paths,
or other implementation details to the client. Return generic error
messages in production; detailed errors only when `DEBUG` is enabled.
- **TLS:**
- Services never terminate TLS directly. They are always deployed behind
a TLS-terminating reverse proxy. The service itself listens on plain
HTTP. However, HSTS headers and `Secure` cookie flags must still be
set by the application so that the browser enforces HTTPS end-to-end.
This list is non-exhaustive. Apply defense-in-depth: if a standard security
hardening measure exists for HTTP services and is not listed here, it is
still expected. When in doubt, harden.
- `README.md` is the primary documentation. Required sections:
- **Description**: First line must include the project name, purpose,
category (web server, SPA, CLI tool, etc.), license, and author. Example:
"µPaaS is an MIT-licensed Go web application by @sneak that receives
git-frontend webhooks and deploys applications via Docker in realtime."
- **Getting Started**: Copy-pasteable install/usage code block.
- **Entrypoints**: Opens by stating that the repo adheres to the
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
standard (with that link), then documents each provided `script/`
entrypoint and its purpose.
- **Rationale**: Why does this exist?
- **Design**: How is the program structured?
- **TODO**: Update meticulously, even between commits. When planning, put
the todo list in the README so a new agent can pick up where the last one
left off.
- **License**: MIT, GPL, or WTFPL. Ask the user for new projects. Include a
`LICENSE` file in the repo root and a License section in the README.
- **Author**: [@sneak](https://sneak.berlin).
- First commit of a new repo should contain only `README.md`.
- Go module root: `sneak.berlin/go/<name>`. Always run `go mod tidy` before
committing.
- Use SemVer.
- Database migrations live in `internal/db/migrations/` and must be embedded in
the binary.
- `000_migration.sql` — contains ONLY the creation of the migrations
tracking table itself. Nothing else.
- `001_schema.sql` — the full application schema.
- **Pre-1.0.0:** never add additional migration files (002, 003, etc.).
There is no installed base to migrate. Edit `001_schema.sql` directly.
- **Post-1.0.0:** add new numbered migration files for each schema change.
Never edit existing migrations after release.
- All repos should have an `.editorconfig` enforcing the project's indentation
settings.
- Avoid putting files in the repo root unless necessary. Root should contain
only project-level config files (`README.md`, `Makefile`, `Dockerfile`,
`LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`, and
language-specific config). Everything else goes in a subdirectory. Canonical
subdirectory names:
- `bin/` — executable scripts and tools
- `cmd/` — Go command entrypoints
- `configs/` — configuration templates and examples
- `deploy/` — deployment manifests (k8s, compose, terraform)
- `docs/` — documentation and markdown (README.md stays in root)
- `internal/` — Go internal packages
- `internal/db/migrations/` — database migrations
- `pkg/` — Go library packages
- `share/` — systemd units, data files
- `static/` — static assets (images, fonts, etc.)
- `web/` — web frontend source
- When setting up a new repo, files from the `prompts` repo may be used as
templates. Fetch them from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/<path>`.
- New repos must contain at minimum:
- `README.md`, `.git`, `.gitignore`, `.editorconfig`
- `LICENSE`, `REPO_POLICIES.md` (copy from the `prompts` repo)
- `Makefile`
- `script/` entrypoints (`bootstrap`, `setup`, `projectname`, `test`,
`lint`, `fmt`, `fmt-check`, `check`, `docker`, `cibuild`, `precommit`,
`install-precommit`)
- `Dockerfile`, `.dockerignore`
- `.gitea/workflows/check.yml`
- Go: `go.mod`, `go.sum`, `.golangci.yml`
- JS: `package.json`, `yarn.lock`, `.prettierrc`, `.prettierignore`
- Python: `pyproject.toml`

213
TODO.md
View File

@@ -1,122 +1,109 @@
# TODO: mfer 1.0 # Workflow
## Design Questions - branch (from `main`)
- do the work in Next Step
- move Next Step to the top of Completed Steps
- move the top item of Future Steps into Next Step
- commit (`TODO.md` changes in the same commit as the work)
- merge to `main` if the branch is not protected, otherwise open a PR
- push
*sneak: please answer inline below each question. These are preserved for posterity.* # Status
### Format Design pre-1.0. No git tags. README section "TODO: Remaining Work for 1.0" lists open
design questions and implementation tasks; policy compliance work is in flight
and unmerged.
**1. Should `MFFileChecksum` be simplified?** # Next Step
Currently it's a separate message wrapping a single `bytes multiHash` field. Since multihash already self-describes the algorithm, `repeated bytes hashes` directly on `MFFilePath` would be simpler and reduce per-file protobuf overhead. Is the extra message layer intentional (e.g. planning to add per-hash metadata like `verified_at`)?
> *answer:* Work through the remaining compliance items folded from the 2026-07-02 audit
(the first group under Future Steps): `.editorconfig`, `.gitignore` coverage,
gofumpt-based `fmt-check`, README "Getting Started", and the rest.
`.golangci.yml` and `TODO.md` are tracked and committed as of 2026-08-07, so the
only thing left of the `chore/align-repo-policies` branch is the list below.
**2. Should file permissions/mode be stored?** # Completed Steps
The format stores mtime/ctime but not Unix file permissions. For archival use (ExFAT, filesystem-independent checksums) this may not matter, but for software distribution or filesystem restoration it's a gap. Should we reserve a field now (e.g. `optional uint32 mode = 305`) even if we don't populate it yet?
> *answer:* - 2026-08-09: added `.prettierrc`/`.prettierignore`, gave `script/fmt` and
`script/fmt-check` one shared prettier file set via `script/prettier`, dropped
the `|| true` that hid prettier failures, and added a node-based Dockerfile
stage so a markdown formatting violation fails `docker build .` (#69)
- 2026-08-07: updated golangci-lint to v2.12.2 everywhere it is pinned
(`Makefile`, `Dockerfile`), added the canonical `.golangci.yml`
(`default: all`), and fixed all resulting lint findings across the codebase
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints, Makefile
shims, README Entrypoints section
- 2026-07-03: aligned repo tooling, docs, and config with standardized policies
(7d9a138, on chore/align-repo-policies, unmerged)
- 2026-06-28: moved to standardized repo policies (#56, on main)
- 2026-04-07: added 1.0 roadmap as README TODO section, removed old TODO.md
(#54)
- 2026-03-20: added Gitea Actions CI workflow (#53)
- 2026-03-17: added REPO_POLICIES.md, renamed CLAUDE.md to AGENTS.md (#51);
removed committed .index.mf (#52)
- 2026-03-15: split Dockerfile with pre-built golangci-lint stage for faster CI
(#45)
- 2026-03-01: 1.0 quality polish: code review, tests, bug fixes, docs (#32)
- 2026-02-20: deterministic file ordering in Builder.Build() (#28); removed
committed vendor/modcache archives (#35)
- 2026-02-08: added --seed flag for deterministic manifest UUID
**3. Should `atime` be removed from the schema?** # Future Steps
Access time is volatile, non-deterministic, and often disabled (`noatime`). Including it means two manifests of the same directory at different times will differ, which conflicts with the determinism goal. Remove it, or document it as "never set by default"?
> *answer:* - Compliance (fold of TODO.md audit 2026-07-02; verify which items the in-flight
branch already closes, then check off):
**4. What are the path normalization rules?** - Add .editorconfig (canonical copy from sneak/prompts)
The proto has `string path` with no specification about: always forward-slash? Must be relative? No `..` components allowed? UTF-8 NFC vs NFD normalization (macOS vs Linux)? Max path length? This is a security issue (path traversal) and a cross-platform compatibility issue. What rules should the spec mandate? - Make .gitignore cover secrets (.env, _.key, _.pem), OS files (.DS_Store),
and editor files (_.swp, _~)
> *answer:* - Make fmt-check/lint verify with gofumpt, not gofmt -l, so `make check`
matches what `make fmt` writes
**5. Should we add a version byte after the magic?** - Add README "Getting Started" section with copy-pasteable install/usage
Currently `ZNAVSRFG` is followed immediately by protobuf. Adding a version byte (`ZNAVSRFG\x01`) would allow future framing changes without requiring protobuf parsing to detect the version. `MFFileOuter.Version` serves this purpose but requires successful deserialization to read. Worth the extra byte? block
- Move FORMAT.md from repo root to docs/ and update the AGENTS.md reference
> *answer:* - Pin Makefile-installed Go tools (`protoc-gen-go@v1.28.1`,
`golangci-lint@v2.12.2`) by module hash, not mutable tag
**6. Should we add a length-prefix after the magic?** - Set `make test` timeout to 30s (currently 10s)
Protobuf is not self-delimiting. If we ever want to concatenate manifests or append data after the protobuf, the current framing is insufficient. Add a varint or fixed-width length-prefix? - Add explicit README "Rationale" heading (content exists under other
names); name the author in the README Description first line
> *answer:* - Reconcile root-level AGENTS.md with directory-hygiene policy (keep or
relocate)
### Signature Design - Add a `make build` target
- Rewrite `make hooks` to use printf or a heredoc instead of non-portable
**7. What does the outer SHA-256 hash cover — compressed or uncompressed data?** `echo '...\n...'`
The review notes it currently hashes compressed data (good for verifying before decompression), but this should be explicitly documented. Which is the intended behavior? - Answer the 14 owner design questions in the README 1.0 roadmap:
- Format: simplify MFFileChecksum; store file mode; drop atime; specify path
> *answer:* normalization rules; version byte after magic; length-prefix after magic
- Signatures: hash covers compressed or uncompressed data; sign raw bytes vs
**8. Should `signatureString()` sign raw bytes instead of a hex-encoded string?** hex canonical string; detached .mf.sig support; GPG subprocess vs pure-Go
Currently the canonical string is `MAGIC-UUID-MULTIHASH` with hex encoding, which adds a transformation layer. Signing the raw `sha256` bytes (or compressed `innerMessage` directly) would be simpler. Keep the string format or switch to raw bytes? crypto
- Implementation: deterministic manifests by default; consolidate duplicate
> *answer:* scanner/checker implementations; export the manifest type; canonical Go
module path for 1.0
**9. Should we support detached signature files (`.mf.sig`)?** - Format and correctness:
Embedded signatures are better for single-file distribution. Detached `.mf.sig` files follow the familiar `SHASUMS`/`SHASUMS.asc` pattern and are simpler for HTTP serving. Support both modes? - Resolve proto go_package vs go.mod module path inconsistency
- Specify and validate path invariants (UTF-8, forward-slash, relative, no
> *answer:* .., no leading /)
- Remove or deprecate atime; reserve mode field; add version byte (all
**10. GPG vs pure-Go crypto for signatures?** pending design answers)
Shelling out to `gpg` is fragile (may not be installed, version-dependent output). `github.com/ProtonMail/go-crypto` provides pure-Go OpenPGP, or we could go Ed25519/signify (simpler, no key management). Which direction? - Write a standalone format specification document
- Library:
> *answer:* - Delete internal/scanner and internal/checker; consolidate on the mfer/
package versions (pending design answer)
### Implementation Design - Add decompression size limit via io.LimitReader in deserializeInner()
- Fix errors.Is dead code in checker; make AddFile verify totalRead == size
**11. Should manifests be deterministic by default?** - Export manifest type or define a public interface (pending)
This means: sort file entries by path, omit `createdAt` timestamp (or make it opt-in), no `atime`. Should determinism be the default, with a `--include-timestamps` flag to opt in? - Replace GPG subprocess with pure-Go crypto (pending); add timeouts to
remaining subprocess calls
> *answer:* - CLI:
- Kebab-case primary flag names; fix fetch URL construction with
**12. Should we consolidate or keep both scanner/checker implementations?** url.JoinPath; add http.Client timeout and retry with backoff to fetch;
There are two parallel implementations: `mfer/scanner.go` + `mfer/checker.go` (typed with `FileSize`, `RelFilePath`) and `internal/scanner/` + `internal/checker/` (raw `int64`, `string`). The `mfer/` versions are superior. Delete the `internal/` versions? rate-limit Checker progress output; add --deterministic flag or default;
wire top-level --version properly
> *answer:* - Testing:
- Fuzz NewManifestFromReader; end-to-end tests for freshen and fetch
**13. Should the `manifest` type be exported?** - Documentation:
Currently unexported with exported constructors (`New`, `NewFromPaths`, etc.). Consumers can't declare `var m *mfer.manifest`. Export the type, or define an interface? - Promote docs/FORMAT.md as primary spec reference; audit error messages;
document the signature scheme fully
> *answer:* - Release:
- Finalize module path, bump version constant, SemVer --version output, tag
**14. What should the Go module path be for 1.0?** v1.0.0
Currently mixed between `sneak.berlin/go/mfer` and `git.eeqj.de/sneak/mfer`. Which is canonical?
> *answer:*
---
## Implementation Plan
### Phase 1: Foundation (format correctness)
- [ ] Delete `internal/scanner/` and `internal/checker/` — consolidate on `mfer/` package versions; update CLI code
- [ ] Add deterministic file ordering — sort entries by path (lexicographic, byte-order) in `Builder.Build()`; add test asserting byte-identical output from two runs
- [ ] Add decompression size limit — `io.LimitReader` in `deserializeInner()` with `m.pbOuter.Size` as bound
- [ ] Fix `errors.Is` dead code in checker — replace with `os.IsNotExist(err)` or `errors.Is(err, fs.ErrNotExist)`
- [ ] Fix `AddFile` to verify size — check `totalRead == size` after reading, return error on mismatch
- [ ] Specify path invariants — add proto comments (UTF-8, forward-slash, relative, no `..`, no leading `/`); validate in `Builder.AddFile` and `Builder.AddFileWithHash`
### Phase 2: CLI polish
- [ ] Fix flag naming — all CLI flags use kebab-case as primary (`--include-dotfiles`, `--follow-symlinks`)
- [ ] Fix URL construction in fetch — use `BaseURL.JoinPath()` or `url.JoinPath()` instead of string concatenation
- [ ] Add progress rate-limiting to Checker — throttle to once per second, matching Scanner
- [ ] Add `--deterministic` flag (or make it default) — omit `createdAt`, sort files
### Phase 3: Robustness
- [ ] Replace GPG subprocess with pure-Go crypto — `github.com/ProtonMail/go-crypto` or Ed25519/signify
- [ ] Add timeout to any remaining subprocess calls
- [ ] Add fuzzing tests for `NewManifestFromReader`
- [ ] Add retry logic to fetch — exponential backoff for transient HTTP errors
### Phase 4: Format finalization
- [ ] Remove or deprecate `atime` from proto (pending design question answer)
- [ ] Reserve `optional uint32 mode = 305` in `MFFilePath` for future file permissions
- [ ] Add version byte after magic — `ZNAVSRFG\x01` for format version 1
- [ ] Write format specification document — separate from README: magic, outer structure, compression, inner structure, path invariants, signature scheme, canonical serialization
### Phase 5: Release prep
- [ ] Finalize Go module path
- [ ] Audit all error messages for consistency and helpfulness
- [ ] Add `--version` output matching SemVer
- [ ] Tag v1.0.0

View File

@@ -1,3 +1,4 @@
// Command mfer generates and verifies file manifests.
package main package main
import ( import (
@@ -6,8 +7,13 @@ import (
"sneak.berlin/go/mfer/internal/cli" "sneak.berlin/go/mfer/internal/cli"
) )
// Appname is the name of this program.
const Appname = "mfer"
// Version and Gitrev are injected at build time via -ldflags.
//
//nolint:gochecknoglobals // set via ldflags at build time
var ( var (
Appname string = "mfer"
Version string Version string
Gitrev string Gitrev string
) )

View File

@@ -6,6 +6,20 @@ import (
"github.com/stretchr/testify/assert" "github.com/stretchr/testify/assert"
) )
func TestBuild(t *testing.T) { // TestAppname pins the program name that main passes to cli.Run; it is
assert.True(t, true) // the name that appears in usage output and in the log prefix. It also
// keeps this package compiled under `go test`.
func TestAppname(t *testing.T) {
t.Parallel()
assert.Equal(t, "mfer", Appname)
}
// TestVersionDefaults documents that Version and Gitrev are empty unless
// injected at build time via -ldflags.
func TestVersionDefaults(t *testing.T) {
t.Parallel()
assert.Empty(t, Version)
assert.Empty(t, Gitrev)
} }

View File

@@ -1,15 +1,14 @@
// Package bork defines the sentinel errors used by the manifest
// reader and writer.
package bork package bork
import ( import (
"errors" "errors"
"fmt"
) )
var ( var (
ErrMissingMagic = errors.New("missing magic bytes in file") // ErrMissingMagic indicates the input lacks the manifest magic bytes.
ErrMissingMagic = errors.New("missing magic bytes in file")
// ErrFileTruncated indicates the input ended before the expected length.
ErrFileTruncated = errors.New("file/stream is truncated abnormally") ErrFileTruncated = errors.New("file/stream is truncated abnormally")
) )
func Newf(format string, args ...interface{}) error {
return fmt.Errorf(format, args...)
}

View File

@@ -1,11 +1,14 @@
package bork package bork_test
import ( import (
"testing" "testing"
"github.com/stretchr/testify/assert" "github.com/stretchr/testify/assert"
"sneak.berlin/go/mfer/internal/bork"
) )
func TestBuild(t *testing.T) { func TestBuild(t *testing.T) {
assert.NotNil(t, ErrMissingMagic) t.Parallel()
assert.Error(t, bork.ErrMissingMagic)
} }

View File

@@ -1,9 +1,14 @@
// Package cli implements the mfer command-line interface.
package cli package cli
import ( import (
"encoding/hex" "encoding/hex"
"errors"
"fmt" "fmt"
"io"
"math"
"path/filepath" "path/filepath"
"strconv"
"strings" "strings"
"time" "time"
@@ -14,49 +19,274 @@ import (
"sneak.berlin/go/mfer/mfer" "sneak.berlin/go/mfer/mfer"
) )
// fingerprintHexLen is the length of a full GPG key fingerprint in hex
// characters.
const fingerprintHexLen = 40
var (
// errNoManifestFound indicates no manifest file was found in the
// searched directory.
errNoManifestFound = errors.New("no manifest found")
// errInvalidFingerprint indicates a malformed --require-signature
// fingerprint argument. The length is spliced in from
// fingerprintHexLen so the two cannot drift apart.
errInvalidFingerprint = errors.New(
"invalid fingerprint: must be exactly " +
strconv.Itoa(fingerprintHexLen) + " hex characters")
// errManifestNotSigned indicates a signature was required but the
// manifest is unsigned. It is wrapped mid-sentence so that the
// rendered message stays exactly as mfer has always printed it.
errManifestNotSigned = errors.New("manifest is not signed")
// errSignerMismatch indicates the embedded signing key fingerprint
// does not match the required signer. Its text is the mid-sentence
// fragment of the rendered message, which users grep for in CI and
// which must therefore not change; match it with errors.Is rather
// than by reading it.
errSignerMismatch = errors.New("does not match required")
)
// safeUint64 converts a non-negative int64 to uint64, clamping negative
// values to zero.
func safeUint64(n int64) uint64 {
if n < 0 {
return 0
}
return uint64(n)
}
// safeRateUint64 converts a bytes-per-second rate to uint64 for display.
//
// A rate is computed as bytes/elapsed, so it is +Inf when the elapsed
// time rounds to zero and NaN when zero bytes were processed in zero
// time. Neither has a defined conversion to uint64, and on amd64 +Inf
// converts to a number that renders as "8.0 EiB/s"; both display as zero
// instead.
func safeRateUint64(rate float64) uint64 {
if math.IsNaN(rate) || math.IsInf(rate, 0) || rate <= 0 {
return 0
}
if rate >= math.MaxUint64 {
return math.MaxUint64
}
return uint64(rate)
}
// findManifest looks for a manifest file in the given directory. // findManifest looks for a manifest file in the given directory.
// It checks for index.mf and .index.mf, returning the first one found. // It checks for index.mf and .index.mf, returning the first one found.
func findManifest(fs afero.Fs, dir string) (string, error) { func findManifest(fs afero.Fs, dir string) (string, error) {
candidates := []string{"index.mf", ".index.mf"} candidates := []string{"index.mf", ".index.mf"}
for _, name := range candidates { for _, name := range candidates {
path := filepath.Join(dir, name) path := filepath.Join(dir, name)
exists, err := afero.Exists(fs, path) exists, err := afero.Exists(fs, path)
if err != nil { if err != nil {
return "", err return "", err
} }
if exists { if exists {
return path, nil return path, nil
} }
} }
return "", fmt.Errorf("no manifest found in %s (looked for index.mf and .index.mf)", dir)
return "", fmt.Errorf(
"%w in %s (looked for index.mf and .index.mf)", errNoManifestFound, dir)
}
// fetchManifestToTemp downloads a manifest URL to a temporary file and
// returns the temp file path. The caller is responsible for removing it.
func (mfa *CLIApp) fetchManifestToTemp(url string) (string, error) {
rc, fetchErr := mfa.openManifestReader(url)
if fetchErr != nil {
return "", fetchErr
}
tmpFile, tmpErr := afero.TempFile(mfa.Fs, "", "mfer-manifest-*.mf")
if tmpErr != nil {
_ = rc.Close()
return "", fmt.Errorf("failed to create temp file: %w", tmpErr)
}
tmpPath := tmpFile.Name()
_, cpErr := io.Copy(tmpFile, rc)
_ = rc.Close()
_ = tmpFile.Close()
if cpErr != nil {
_ = mfa.Fs.Remove(tmpPath)
return "", fmt.Errorf("failed to download manifest: %w", cpErr)
}
return tmpPath, nil
}
// verifyRequiredSigner enforces the --require-signature fingerprint
// against the manifest's embedded signing key.
func verifyRequiredSigner(chk *mfer.Checker, requiredSigner string) error {
// Validate fingerprint format: must be exactly 40 hex characters
if len(requiredSigner) != fingerprintHexLen {
return fmt.Errorf("%w, got %d", errInvalidFingerprint, len(requiredSigner))
}
_, err := hex.DecodeString(requiredSigner)
if err != nil {
return fmt.Errorf("invalid fingerprint: must be valid hex: %w", err)
}
if !chk.IsSigned() {
return fmt.Errorf("%w, but signature from %s is required",
errManifestNotSigned, requiredSigner)
}
// Extract fingerprint from the embedded public key (not from the
// signer field). This validates the key is importable and gets its
// actual fingerprint.
embeddedFP, err := chk.ExtractEmbeddedSigningKeyFP()
if err != nil {
return fmt.Errorf(
"failed to extract fingerprint from embedded signing key: %w", err)
}
// Compare fingerprints - must be exact match (case-insensitive)
if !strings.EqualFold(embeddedFP, requiredSigner) {
return fmt.Errorf("embedded signing key fingerprint %s %w %s",
embeddedFP, errSignerMismatch, requiredSigner)
}
log.Infof("manifest signature verified (signer: %s)", embeddedFP)
return nil
}
// reportCheckProgress renders progress updates until the channel closes.
func reportCheckProgress(progress <-chan mfer.CheckStatus) {
for status := range progress {
if status.ETA > 0 {
log.Progressf("Checking: %d/%d files, %s/s, ETA %s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)),
status.ETA.Round(time.Second),
status.Failures)
} else {
log.Progressf("Checking: %d/%d files, %s/s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)),
status.Failures)
}
}
log.ProgressDone()
}
// countCheckFailures consumes check results, counting and logging
// failures, then closes done.
func countCheckFailures(
results <-chan mfer.Result, failures *int64, done chan<- struct{},
) {
for result := range results {
if result.Status != mfer.StatusOK {
*failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
} else {
log.Verbosef("%s: %s", result.Status, result.Path)
}
}
close(done)
}
// findExtraFiles reports files present on disk but absent from the
// manifest, counting each as a failure.
func findExtraFiles(ctx *cli.Context, chk *mfer.Checker, failures *int64) error {
extraResults := make(chan mfer.Result, 1)
extraDone := make(chan struct{})
go func() {
for result := range extraResults {
*failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
}
close(extraDone)
}()
err := chk.FindExtraFiles(ctx.Context, extraResults)
if err != nil {
return fmt.Errorf("failed to check for extra files: %w", err)
}
<-extraDone
return nil
}
// runCheck runs the manifest check with progress and result reporting
// and returns the number of failures.
func runCheck(ctx *cli.Context, chk *mfer.Checker, showProgress bool) (int64, error) {
// Set up results channel
results := make(chan mfer.Result, 1)
// Set up progress channel
var progress chan mfer.CheckStatus
if showProgress {
progress = make(chan mfer.CheckStatus, 1)
go reportCheckProgress(progress)
}
// Process results in a goroutine
var failures int64
done := make(chan struct{})
go countCheckFailures(results, &failures, done)
// Run check
err := chk.Check(ctx.Context, results, progress)
if err != nil {
return 0, fmt.Errorf("check failed: %w", err)
}
// Wait for results processing to complete
<-done
// Check for extra files if requested
if ctx.Bool("no-extra-files") {
err = findExtraFiles(ctx, chk, &failures)
if err != nil {
return 0, err
}
}
return failures, nil
} }
func (mfa *CLIApp) checkManifestOperation(ctx *cli.Context) error { func (mfa *CLIApp) checkManifestOperation(ctx *cli.Context) error {
log.Debug("checkManifestOperation()") log.Debug("checkManifestOperation()")
var manifestPath string manifestPath, err := mfa.resolveManifestArg(ctx)
var err error if err != nil {
return fmt.Errorf("check: %w", err)
}
if ctx.Args().Len() > 0 { // URL manifests need to be downloaded to a temp file for the checker
arg := ctx.Args().Get(0) if isHTTPURL(manifestPath) {
// Check if arg is a directory or a file tmpPath, tmpErr := mfa.fetchManifestToTemp(manifestPath)
info, statErr := mfa.Fs.Stat(arg) if tmpErr != nil {
if statErr == nil && info.IsDir() { return fmt.Errorf("check: %w", tmpErr)
// It's a directory, look for manifest inside
manifestPath, err = findManifest(mfa.Fs, arg)
if err != nil {
return err
}
} else {
// Treat as a file path
manifestPath = arg
}
} else {
// No argument, look in current directory
manifestPath, err = findManifest(mfa.Fs, ".")
if err != nil {
return err
} }
defer func() { _ = mfa.Fs.Remove(tmpPath) }()
manifestPath = tmpPath
} }
basePath := ctx.String("base") basePath := ctx.String("base")
@@ -73,111 +303,31 @@ func (mfa *CLIApp) checkManifestOperation(ctx *cli.Context) error {
// Check signature requirement // Check signature requirement
requiredSigner := ctx.String("require-signature") requiredSigner := ctx.String("require-signature")
if requiredSigner != "" { if requiredSigner != "" {
// Validate fingerprint format: must be exactly 40 hex characters err = verifyRequiredSigner(chk, requiredSigner)
if len(requiredSigner) != 40 {
return fmt.Errorf("invalid fingerprint: must be exactly 40 hex characters, got %d", len(requiredSigner))
}
if _, err := hex.DecodeString(requiredSigner); err != nil {
return fmt.Errorf("invalid fingerprint: must be valid hex: %w", err)
}
if !chk.IsSigned() {
return fmt.Errorf("manifest is not signed, but signature from %s is required", requiredSigner)
}
// Extract fingerprint from the embedded public key (not from the signer field)
// This validates the key is importable and gets its actual fingerprint
embeddedFP, err := chk.ExtractEmbeddedSigningKeyFP()
if err != nil { if err != nil {
return fmt.Errorf("failed to extract fingerprint from embedded signing key: %w", err) return err
} }
// Compare fingerprints - must be exact match (case-insensitive)
if !strings.EqualFold(embeddedFP, requiredSigner) {
return fmt.Errorf("embedded signing key fingerprint %s does not match required %s", embeddedFP, requiredSigner)
}
log.Infof("manifest signature verified (signer: %s)", embeddedFP)
} }
log.Infof("manifest contains %d files, %s", chk.FileCount(), humanize.IBytes(uint64(chk.TotalBytes()))) log.Infof("manifest contains %d files, %s", chk.FileCount(),
humanize.IBytes(safeUint64(int64(chk.TotalBytes()))))
// Set up results channel failures, err := runCheck(ctx, chk, showProgress)
results := make(chan mfer.Result, 1)
// Set up progress channel
var progress chan mfer.CheckStatus
if showProgress {
progress = make(chan mfer.CheckStatus, 1)
go func() {
for status := range progress {
if status.ETA > 0 {
log.Progressf("Checking: %d/%d files, %s/s, ETA %s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)),
status.ETA.Round(time.Second),
status.Failures)
} else {
log.Progressf("Checking: %d/%d files, %s/s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)),
status.Failures)
}
}
log.ProgressDone()
}()
}
// Process results in a goroutine
var failures int64
done := make(chan struct{})
go func() {
for result := range results {
if result.Status != mfer.StatusOK {
failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
} else {
log.Verbosef("%s: %s", result.Status, result.Path)
}
}
close(done)
}()
// Run check
err = chk.Check(ctx.Context, results, progress)
if err != nil { if err != nil {
return fmt.Errorf("check failed: %w", err) return err
}
// Wait for results processing to complete
<-done
// Check for extra files if requested
if ctx.Bool("no-extra-files") {
extraResults := make(chan mfer.Result, 1)
extraDone := make(chan struct{})
go func() {
for result := range extraResults {
failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
}
close(extraDone)
}()
err = chk.FindExtraFiles(ctx.Context, extraResults)
if err != nil {
return fmt.Errorf("failed to check for extra files: %w", err)
}
<-extraDone
} }
elapsed := time.Since(mfa.startupTime).Seconds() elapsed := time.Since(mfa.startupTime).Seconds()
rate := float64(chk.TotalBytes()) / elapsed rate := float64(chk.TotalBytes()) / elapsed
if failures == 0 { if failures == 0 {
log.Infof("checked %d files (%s) in %.1fs (%s/s): all OK", chk.FileCount(), humanize.IBytes(uint64(chk.TotalBytes())), elapsed, humanize.IBytes(uint64(rate))) log.Infof("checked %d files (%s) in %.1fs (%s/s): all OK",
chk.FileCount(), humanize.IBytes(safeUint64(int64(chk.TotalBytes()))),
elapsed, humanize.IBytes(safeRateUint64(rate)))
} else { } else {
log.Infof("checked %d files (%s) in %.1fs (%s/s): %d failed", chk.FileCount(), humanize.IBytes(uint64(chk.TotalBytes())), elapsed, humanize.IBytes(uint64(rate)), failures) log.Infof("checked %d files (%s) in %.1fs (%s/s): %d failed",
chk.FileCount(), humanize.IBytes(safeUint64(int64(chk.TotalBytes()))),
elapsed, humanize.IBytes(safeRateUint64(rate)), failures)
} }
if failures > 0 { if failures > 0 {

View File

@@ -7,15 +7,18 @@ import (
"github.com/spf13/afero" "github.com/spf13/afero"
) )
// NO_COLOR disables colored output when set. Automatically true if the // NoColor disables colored output when set. Automatically true if the
// NO_COLOR environment variable is present (per https://no-color.org/). // NO_COLOR environment variable is present (per https://no-color.org/).
var NO_COLOR bool //
//nolint:gochecknoglobals // process-wide setting derived from the environment
var NoColor = noColorEnvSet()
func init() { // noColorEnvSet reports whether the NO_COLOR environment variable is
NO_COLOR = false // present.
if _, exists := os.LookupEnv("NO_COLOR"); exists { func noColorEnvSet() bool {
NO_COLOR = true _, exists := os.LookupEnv("NO_COLOR")
}
return exists
} }
// RunOptions contains all configuration for running the CLI application. // RunOptions contains all configuration for running the CLI application.
@@ -64,5 +67,6 @@ func RunWithOptions(opts *RunOptions) int {
} }
m.run(opts.Args) m.run(opts.Args)
return m.exitCode return m.exitCode
} }

View File

@@ -1,9 +1,13 @@
//nolint:testpackage // white-box tests exercise unexported internals
package cli package cli
import ( import (
"bytes" "bytes"
"errors"
"fmt" "fmt"
"math/rand" "math/rand"
"os"
"sync"
"testing" "testing"
"github.com/spf13/afero" "github.com/spf13/afero"
@@ -13,19 +17,53 @@ import (
"sneak.berlin/go/mfer/mfer" "sneak.berlin/go/mfer/mfer"
) )
func init() { const (
testApp = "mfer"
testDir = "/testdir"
testFile1 = "/testdir/file1.txt"
testMF = "/testdir/test.mf"
testOutput = "/output.mf"
testOutputTmp = "/output.mf.tmp"
testManifest = "/manifest.mf"
testFlagBase = "--base"
testFlagNoExtra = "--no-extra-files"
)
var errSimulatedWrite = errors.New("simulated write failure")
// runMu serializes CLI runs: RunWithOptions wires the process-global
// logger to the run's I/O streams, so parallel runs would cross-wire
// captured output between tests.
//
//nolint:gochecknoglobals // guards process-global logger state in tests
var runMu sync.Mutex
// runCLI invokes RunWithOptions while holding runMu so parallel tests
// capture their own output.
func runCLI(opts *RunOptions) int {
runMu.Lock()
defer runMu.Unlock()
return RunWithOptions(opts)
}
func TestMain(m *testing.M) {
// Prevent urfave/cli from calling os.Exit during tests // Prevent urfave/cli from calling os.Exit during tests
urfcli.OsExiter = func(code int) {} urfcli.OsExiter = func(_ int) {}
os.Exit(m.Run())
} }
func TestBuild(t *testing.T) { func TestBuild(t *testing.T) {
t.Parallel()
m := &CLIApp{} m := &CLIApp{}
assert.NotNil(t, m) assert.NotNil(t, m)
} }
func testOpts(args []string, fs afero.Fs) *RunOptions { func testOpts(args []string, fs afero.Fs) *RunOptions {
return &RunOptions{ return &RunOptions{
Appname: "mfer", Appname: testApp,
Version: "1.0.0", Version: "1.0.0",
Gitrev: "abc123", Gitrev: "abc123",
Args: args, Args: args,
@@ -36,374 +74,451 @@ func testOpts(args []string, fs afero.Fs) *RunOptions {
} }
} }
func TestVersionCommand(t *testing.T) { func testStdout(t *testing.T, opts *RunOptions) string {
fs := afero.NewMemMapFs() t.Helper()
opts := testOpts([]string{"mfer", "version"}, fs)
exitCode := RunWithOptions(opts) buf, ok := opts.Stdout.(*bytes.Buffer)
require.True(t, ok)
return buf.String()
}
func testStderr(t *testing.T, opts *RunOptions) string {
t.Helper()
buf, ok := opts.Stderr.(*bytes.Buffer)
require.True(t, ok)
return buf.String()
}
func writeTestFile(t *testing.T, fs afero.Fs, path, content string) {
t.Helper()
require.NoError(t, afero.WriteFile(fs, path, []byte(content), 0o644))
}
func TestVersionCommand(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
opts := testOpts([]string{testApp, "version"}, fs)
exitCode := runCLI(opts)
assert.Equal(t, 0, exitCode) assert.Equal(t, 0, exitCode)
stdout := opts.Stdout.(*bytes.Buffer).String()
stdout := testStdout(t, opts)
assert.Contains(t, stdout, mfer.Version) assert.Contains(t, stdout, mfer.Version)
assert.Contains(t, stdout, "abc123") assert.Contains(t, stdout, "abc123")
} }
func TestHelpCommand(t *testing.T) { func TestHelpCommand(t *testing.T) {
fs := afero.NewMemMapFs() t.Parallel()
opts := testOpts([]string{"mfer", "--help"}, fs)
exitCode := RunWithOptions(opts) fs := afero.NewMemMapFs()
opts := testOpts([]string{testApp, "--help"}, fs)
exitCode := runCLI(opts)
assert.Equal(t, 0, exitCode) assert.Equal(t, 0, exitCode)
stdout := opts.Stdout.(*bytes.Buffer).String()
assert.Contains(t, stdout, "generate") stdout := testStdout(t, opts)
assert.Contains(t, stdout, "check") assert.Contains(t, stdout, cmdGenerate)
assert.Contains(t, stdout, cmdCheck)
assert.Contains(t, stdout, "fetch") assert.Contains(t, stdout, "fetch")
} }
func TestGenerateCommand(t *testing.T) { func TestGenerateCommand(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test files in memory filesystem // Create test files in memory filesystem
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello world"), 0o644)) writeTestFile(t, fs, testFile1, "hello world")
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("test content"), 0o644)) writeTestFile(t, fs, "/testdir/file2.txt", "test content")
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/testdir/test.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testMF, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
assert.Equal(t, 0, exitCode, "stderr: %s", opts.Stderr.(*bytes.Buffer).String()) assert.Equal(t, 0, exitCode, "stderr: %s", testStderr(t, opts))
// Verify manifest was created // Verify manifest was created
exists, err := afero.Exists(fs, "/testdir/test.mf") exists, err := afero.Exists(fs, testMF)
require.NoError(t, err) require.NoError(t, err)
assert.True(t, exists) assert.True(t, exists)
} }
func TestGenerateAndCheckCommand(t *testing.T) { func TestGenerateAndCheckCommand(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test files with subdirectory // Create test files with subdirectory
require.NoError(t, fs.MkdirAll("/testdir/subdir", 0o755)) require.NoError(t, fs.MkdirAll("/testdir/subdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello world"), 0o644)) writeTestFile(t, fs, testFile1, "hello world")
require.NoError(t, afero.WriteFile(fs, "/testdir/subdir/file2.txt", []byte("test content"), 0o644)) writeTestFile(t, fs, "/testdir/subdir/file2.txt", "test content")
// Generate manifest // Generate manifest
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/testdir/test.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testMF, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode, "generate failed: %s", opts.Stderr.(*bytes.Buffer).String()) require.Equal(t, 0, exitCode, "generate failed: %s", testStderr(t, opts))
// Check manifest // Check manifest
opts = testOpts([]string{"mfer", "check", "-q", "--base", "/testdir", "/testdir/test.mf"}, fs) opts = testOpts([]string{testApp, cmdCheck, "-q", testFlagBase, testDir, testMF}, fs)
exitCode = RunWithOptions(opts) exitCode = runCLI(opts)
assert.Equal(t, 0, exitCode, "check failed: %s", opts.Stderr.(*bytes.Buffer).String()) assert.Equal(t, 0, exitCode, "check failed: %s", testStderr(t, opts))
} }
func TestCheckCommandWithMissingFile(t *testing.T) { func TestCheckCommandWithMissingFile(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test file // Create test file
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello world"), 0o644)) writeTestFile(t, fs, testFile1, "hello world")
// Generate manifest // Generate manifest
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/testdir/test.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testMF, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode, "generate failed: %s", opts.Stderr.(*bytes.Buffer).String()) require.Equal(t, 0, exitCode, "generate failed: %s", testStderr(t, opts))
// Delete the file // Delete the file
require.NoError(t, fs.Remove("/testdir/file1.txt")) require.NoError(t, fs.Remove(testFile1))
// Check manifest - should fail // Check manifest - should fail
opts = testOpts([]string{"mfer", "check", "-q", "--base", "/testdir", "/testdir/test.mf"}, fs) opts = testOpts([]string{testApp, cmdCheck, "-q", testFlagBase, testDir, testMF}, fs)
exitCode = RunWithOptions(opts) exitCode = runCLI(opts)
assert.Equal(t, 1, exitCode, "check should have failed for missing file") assert.Equal(t, 1, exitCode, "check should have failed for missing file")
} }
func TestCheckCommandWithCorruptedFile(t *testing.T) { func runCheckAfterRewrite(t *testing.T, rewritten, msg string) {
t.Helper()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test file // Create test file
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello world"), 0o644)) writeTestFile(t, fs, testFile1, "hello world")
// Generate manifest // Generate manifest
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/testdir/test.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testMF, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode, "generate failed: %s", opts.Stderr.(*bytes.Buffer).String()) require.Equal(t, 0, exitCode, "generate failed: %s", testStderr(t, opts))
// Rewrite the file, then check the manifest - it must fail
writeTestFile(t, fs, testFile1, rewritten)
opts = testOpts([]string{testApp, cmdCheck, "-q", testFlagBase, testDir, testMF}, fs)
exitCode = runCLI(opts)
assert.Equal(t, 1, exitCode, msg)
}
func TestCheckCommandWithCorruptedFile(t *testing.T) {
t.Parallel()
// Corrupt the file (change content but keep same size) // Corrupt the file (change content but keep same size)
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("HELLO WORLD"), 0o644)) runCheckAfterRewrite(t, "HELLO WORLD",
"check should have failed for corrupted file")
// Check manifest - should fail with hash mismatch
opts = testOpts([]string{"mfer", "check", "-q", "--base", "/testdir", "/testdir/test.mf"}, fs)
exitCode = RunWithOptions(opts)
assert.Equal(t, 1, exitCode, "check should have failed for corrupted file")
} }
func TestCheckCommandWithSizeMismatch(t *testing.T) { func TestCheckCommandWithSizeMismatch(t *testing.T) {
fs := afero.NewMemMapFs() t.Parallel()
// Create test file
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello world"), 0o644))
// Generate manifest
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/testdir/test.mf", "/testdir"}, fs)
exitCode := RunWithOptions(opts)
require.Equal(t, 0, exitCode, "generate failed: %s", opts.Stderr.(*bytes.Buffer).String())
// Change file size // Change file size
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("different size content here"), 0o644)) runCheckAfterRewrite(t, "different size content here",
"check should have failed for size mismatch")
// Check manifest - should fail with size mismatch
opts = testOpts([]string{"mfer", "check", "-q", "--base", "/testdir", "/testdir/test.mf"}, fs)
exitCode = RunWithOptions(opts)
assert.Equal(t, 1, exitCode, "check should have failed for size mismatch")
} }
func TestBannerOutput(t *testing.T) { func TestBannerOutput(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test file // Create test file
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
// Run without -q to see banner // Run without -q to see banner
opts := testOpts([]string{"mfer", "generate", "-o", "/testdir/test.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-o", testMF, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
assert.Equal(t, 0, exitCode) assert.Equal(t, 0, exitCode)
// Banner ASCII art should be in stdout // Banner ASCII art should be in stdout
stdout := opts.Stdout.(*bytes.Buffer).String() stdout := testStdout(t, opts)
assert.Contains(t, stdout, "___") assert.Contains(t, stdout, "___")
assert.Contains(t, stdout, "\\") assert.Contains(t, stdout, "\\")
} }
func TestUnknownCommand(t *testing.T) { func TestUnknownCommand(t *testing.T) {
fs := afero.NewMemMapFs() t.Parallel()
opts := testOpts([]string{"mfer", "unknown"}, fs)
exitCode := RunWithOptions(opts) fs := afero.NewMemMapFs()
opts := testOpts([]string{testApp, "unknown"}, fs)
exitCode := runCLI(opts)
assert.Equal(t, 1, exitCode) assert.Equal(t, 1, exitCode)
} }
func TestGenerateExcludesDotfilesByDefault(t *testing.T) { func TestGenerateExcludesDotfilesByDefault(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test files including dotfiles // Create test files including dotfiles
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
require.NoError(t, afero.WriteFile(fs, "/testdir/.hidden", []byte("secret"), 0o644)) writeTestFile(t, fs, "/testdir/.hidden", "secret")
// Generate manifest without --include-dotfiles (default excludes dotfiles) // Generate manifest without --include-dotfiles (default excludes dotfiles)
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/testdir/test.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testMF, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode) require.Equal(t, 0, exitCode)
// Check that manifest exists // Check that manifest exists
exists, _ := afero.Exists(fs, "/testdir/test.mf") exists, _ := afero.Exists(fs, testMF)
assert.True(t, exists) assert.True(t, exists)
// Verify manifest only has 1 file (the non-dotfile) // Verify manifest only has 1 file (the non-dotfile)
manifest, err := mfer.NewManifestFromFile(fs, "/testdir/test.mf") manifest, err := mfer.NewManifestFromFile(fs, testMF)
require.NoError(t, err) require.NoError(t, err)
assert.Len(t, manifest.Files(), 1) assert.Len(t, manifest.Files(), 1)
assert.Equal(t, "file1.txt", manifest.Files()[0].Path) assert.Equal(t, "file1.txt", manifest.Files()[0].GetPath())
} }
func TestGenerateWithIncludeDotfiles(t *testing.T) { func TestGenerateWithIncludeDotfiles(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test files including dotfiles // Create test files including dotfiles
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
require.NoError(t, afero.WriteFile(fs, "/testdir/.hidden", []byte("secret"), 0o644)) writeTestFile(t, fs, "/testdir/.hidden", "secret")
// Generate manifest with --include-dotfiles // Generate manifest with --include-dotfiles
opts := testOpts([]string{"mfer", "generate", "-q", "--include-dotfiles", "-o", "/testdir/test.mf", "/testdir"}, fs) opts := testOpts([]string{
exitCode := RunWithOptions(opts) testApp, cmdGenerate, "-q", "--include-dotfiles", "-o", testMF, testDir,
}, fs)
exitCode := runCLI(opts)
require.Equal(t, 0, exitCode) require.Equal(t, 0, exitCode)
// Verify manifest has 2 files (including dotfile) // Verify manifest has 2 files (including dotfile)
manifest, err := mfer.NewManifestFromFile(fs, "/testdir/test.mf") manifest, err := mfer.NewManifestFromFile(fs, testMF)
require.NoError(t, err) require.NoError(t, err)
assert.Len(t, manifest.Files(), 2) assert.Len(t, manifest.Files(), 2)
} }
func TestMultipleInputPaths(t *testing.T) { func TestMultipleInputPaths(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test files in multiple directories // Create test files in multiple directories
require.NoError(t, fs.MkdirAll("/dir1", 0o755)) require.NoError(t, fs.MkdirAll("/dir1", 0o755))
require.NoError(t, fs.MkdirAll("/dir2", 0o755)) require.NoError(t, fs.MkdirAll("/dir2", 0o755))
require.NoError(t, afero.WriteFile(fs, "/dir1/file1.txt", []byte("content1"), 0o644)) writeTestFile(t, fs, "/dir1/file1.txt", "content1")
require.NoError(t, afero.WriteFile(fs, "/dir2/file2.txt", []byte("content2"), 0o644)) writeTestFile(t, fs, "/dir2/file2.txt", "content2")
// Generate manifest from multiple paths // Generate manifest from multiple paths
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/output.mf", "/dir1", "/dir2"}, fs) opts := testOpts([]string{
exitCode := RunWithOptions(opts) testApp, cmdGenerate, "-q", "-o", testOutput, "/dir1", "/dir2",
assert.Equal(t, 0, exitCode, "stderr: %s", opts.Stderr.(*bytes.Buffer).String()) }, fs)
exitCode := runCLI(opts)
assert.Equal(t, 0, exitCode, "stderr: %s", testStderr(t, opts))
exists, _ := afero.Exists(fs, "/output.mf") exists, _ := afero.Exists(fs, testOutput)
assert.True(t, exists) assert.True(t, exists)
} }
func TestNoExtraFilesPass(t *testing.T) { func TestNoExtraFilesPass(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test files // Create test files
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("world"), 0o644)) writeTestFile(t, fs, "/testdir/file2.txt", "world")
// Generate manifest // Generate manifest
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/manifest.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testManifest, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode) require.Equal(t, 0, exitCode)
// Check with --no-extra-files (should pass - no extra files) // Check with --no-extra-files (should pass - no extra files)
opts = testOpts([]string{"mfer", "check", "-q", "--no-extra-files", "--base", "/testdir", "/manifest.mf"}, fs) opts = testOpts([]string{
exitCode = RunWithOptions(opts) testApp, cmdCheck, "-q", testFlagNoExtra, testFlagBase, testDir, testManifest,
}, fs)
exitCode = runCLI(opts)
assert.Equal(t, 0, exitCode) assert.Equal(t, 0, exitCode)
} }
func TestNoExtraFilesFail(t *testing.T) { func TestNoExtraFilesFail(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test files // Create test files
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
// Generate manifest // Generate manifest
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/manifest.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testManifest, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode) require.Equal(t, 0, exitCode)
// Add an extra file after manifest generation // Add an extra file after manifest generation
require.NoError(t, afero.WriteFile(fs, "/testdir/extra.txt", []byte("extra"), 0o644)) writeTestFile(t, fs, "/testdir/extra.txt", "extra")
// Check with --no-extra-files (should fail - extra file exists) // Check with --no-extra-files (should fail - extra file exists)
opts = testOpts([]string{"mfer", "check", "-q", "--no-extra-files", "--base", "/testdir", "/manifest.mf"}, fs) opts = testOpts([]string{
exitCode = RunWithOptions(opts) testApp, cmdCheck, "-q", testFlagNoExtra, testFlagBase, testDir, testManifest,
}, fs)
exitCode = runCLI(opts)
assert.Equal(t, 1, exitCode, "check should fail when extra files exist") assert.Equal(t, 1, exitCode, "check should fail when extra files exist")
} }
func TestNoExtraFilesWithSubdirectory(t *testing.T) { func TestNoExtraFilesWithSubdirectory(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test files with subdirectory // Create test files with subdirectory
require.NoError(t, fs.MkdirAll("/testdir/subdir", 0o755)) require.NoError(t, fs.MkdirAll("/testdir/subdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
require.NoError(t, afero.WriteFile(fs, "/testdir/subdir/file2.txt", []byte("world"), 0o644)) writeTestFile(t, fs, "/testdir/subdir/file2.txt", "world")
// Generate manifest // Generate manifest
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/manifest.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testManifest, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode) require.Equal(t, 0, exitCode)
// Add extra file in subdirectory // Add extra file in subdirectory
require.NoError(t, afero.WriteFile(fs, "/testdir/subdir/extra.txt", []byte("extra"), 0o644)) writeTestFile(t, fs, "/testdir/subdir/extra.txt", "extra")
// Check with --no-extra-files (should fail) // Check with --no-extra-files (should fail)
opts = testOpts([]string{"mfer", "check", "-q", "--no-extra-files", "--base", "/testdir", "/manifest.mf"}, fs) opts = testOpts([]string{
exitCode = RunWithOptions(opts) testApp, cmdCheck, "-q", testFlagNoExtra, testFlagBase, testDir, testManifest,
assert.Equal(t, 1, exitCode, "check should fail when extra files exist in subdirectory") }, fs)
exitCode = runCLI(opts)
assert.Equal(t, 1, exitCode,
"check should fail when extra files exist in subdirectory")
} }
func TestCheckWithoutNoExtraFilesIgnoresExtra(t *testing.T) { func TestCheckWithoutNoExtraFilesIgnoresExtra(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test file // Create test file
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
// Generate manifest // Generate manifest
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/manifest.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testManifest, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode) require.Equal(t, 0, exitCode)
// Add extra file // Add extra file
require.NoError(t, afero.WriteFile(fs, "/testdir/extra.txt", []byte("extra"), 0o644)) writeTestFile(t, fs, "/testdir/extra.txt", "extra")
// Check WITHOUT --no-extra-files (should pass - extra files ignored) // Check WITHOUT --no-extra-files (should pass - extra files ignored)
opts = testOpts([]string{"mfer", "check", "-q", "--base", "/testdir", "/manifest.mf"}, fs) opts = testOpts([]string{
exitCode = RunWithOptions(opts) testApp, cmdCheck, "-q", testFlagBase, testDir, testManifest,
assert.Equal(t, 0, exitCode, "check without --no-extra-files should ignore extra files") }, fs)
exitCode = runCLI(opts)
assert.Equal(t, 0, exitCode,
"check without --no-extra-files should ignore extra files")
} }
func TestGenerateAtomicWriteNoTempFileOnSuccess(t *testing.T) { func TestGenerateAtomicWriteNoTempFileOnSuccess(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test file // Create test file
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
// Generate manifest // Generate manifest
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/output.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testOutput, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode) require.Equal(t, 0, exitCode)
// Verify output file exists // Verify output file exists
exists, err := afero.Exists(fs, "/output.mf") exists, err := afero.Exists(fs, testOutput)
require.NoError(t, err) require.NoError(t, err)
assert.True(t, exists, "output file should exist") assert.True(t, exists, "output file should exist")
// Verify temp file does NOT exist // Verify temp file does NOT exist
tmpExists, err := afero.Exists(fs, "/output.mf.tmp") tmpExists, err := afero.Exists(fs, testOutputTmp)
require.NoError(t, err) require.NoError(t, err)
assert.False(t, tmpExists, "temp file should not exist after successful generation") assert.False(t, tmpExists,
"temp file should not exist after successful generation")
} }
func TestGenerateAtomicWriteOverwriteWithForce(t *testing.T) { func TestGenerateAtomicWriteOverwriteWithForce(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test file // Create test file
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
// Create existing manifest with different content // Create existing manifest with different content
require.NoError(t, afero.WriteFile(fs, "/output.mf", []byte("old content"), 0o644)) writeTestFile(t, fs, testOutput, "old content")
// Generate manifest with --force // Generate manifest with --force
opts := testOpts([]string{"mfer", "generate", "-q", "-f", "-o", "/output.mf", "/testdir"}, fs) opts := testOpts([]string{
exitCode := RunWithOptions(opts) testApp, cmdGenerate, "-q", "-f", "-o", testOutput, testDir,
}, fs)
exitCode := runCLI(opts)
require.Equal(t, 0, exitCode) require.Equal(t, 0, exitCode)
// Verify output file exists and was overwritten // Verify output file exists and was overwritten
content, err := afero.ReadFile(fs, "/output.mf") content, err := afero.ReadFile(fs, testOutput)
require.NoError(t, err) require.NoError(t, err)
assert.NotEqual(t, "old content", string(content), "manifest should be overwritten") assert.NotEqual(t, "old content", string(content),
"manifest should be overwritten")
// Verify temp file does NOT exist // Verify temp file does NOT exist
tmpExists, err := afero.Exists(fs, "/output.mf.tmp") tmpExists, err := afero.Exists(fs, testOutputTmp)
require.NoError(t, err) require.NoError(t, err)
assert.False(t, tmpExists, "temp file should not exist after successful generation") assert.False(t, tmpExists,
"temp file should not exist after successful generation")
} }
func TestGenerateFailsWithoutForceWhenOutputExists(t *testing.T) { func TestGenerateFailsWithoutForceWhenOutputExists(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test file // Create test file
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
// Create existing manifest // Create existing manifest
require.NoError(t, afero.WriteFile(fs, "/output.mf", []byte("existing"), 0o644)) writeTestFile(t, fs, testOutput, "existing")
// Generate manifest WITHOUT --force (should fail) // Generate manifest WITHOUT --force (should fail)
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/output.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testOutput, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
assert.Equal(t, 1, exitCode, "should fail when output exists without --force") assert.Equal(t, 1, exitCode, "should fail when output exists without --force")
// Verify original content is preserved // Verify original content is preserved
content, err := afero.ReadFile(fs, "/output.mf") content, err := afero.ReadFile(fs, testOutput)
require.NoError(t, err) require.NoError(t, err)
assert.Equal(t, "existing", string(content), "original file should be preserved") assert.Equal(t, "existing", string(content), "original file should be preserved")
} }
func TestGenerateAtomicWriteUsesTemp(t *testing.T) { func TestGenerateAtomicWriteUsesTemp(t *testing.T) {
t.Parallel()
// This test verifies that generate uses a temp file by checking // This test verifies that generate uses a temp file by checking
// that the output file doesn't exist until generation completes. // that the output file doesn't exist until generation completes.
// We do this by generating to a path and verifying the temp file // We do this by generating to a path and verifying the temp file
@@ -411,183 +526,239 @@ func TestGenerateAtomicWriteUsesTemp(t *testing.T) {
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create test file // Create test file
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("hello"), 0o644)) writeTestFile(t, fs, testFile1, "hello")
// Generate manifest // Generate manifest
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/output.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testOutput, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode) require.Equal(t, 0, exitCode)
// Both output file should exist and temp should not // Both output file should exist and temp should not
exists, _ := afero.Exists(fs, "/output.mf") exists, _ := afero.Exists(fs, testOutput)
assert.True(t, exists, "output file should exist") assert.True(t, exists, "output file should exist")
tmpExists, _ := afero.Exists(fs, "/output.mf.tmp") tmpExists, _ := afero.Exists(fs, testOutputTmp)
assert.False(t, tmpExists, "temp file should be cleaned up") assert.False(t, tmpExists, "temp file should be cleaned up")
// Verify manifest is valid (not empty) // Verify manifest is valid (not empty)
content, err := afero.ReadFile(fs, "/output.mf") content, err := afero.ReadFile(fs, testOutput)
require.NoError(t, err) require.NoError(t, err)
assert.True(t, len(content) > 0, "manifest should not be empty") assert.NotEmpty(t, content, "manifest should not be empty")
} }
// failingWriterFs wraps a filesystem and makes writes fail after N bytes // failingWriterFs wraps a filesystem and makes writes fail after N bytes
type failingWriterFs struct { type failingWriterFs struct {
afero.Fs afero.Fs
failAfter int64 failAfter int64
written int64 written int64
} }
type failingFile struct { type failingFile struct {
afero.File afero.File
fs *failingWriterFs fs *failingWriterFs
} }
func (f *failingFile) Write(p []byte) (int, error) { func (f *failingFile) Write(p []byte) (int, error) {
f.fs.written += int64(len(p)) f.fs.written += int64(len(p))
if f.fs.written > f.fs.failAfter { if f.fs.written > f.fs.failAfter {
return 0, fmt.Errorf("simulated write failure") return 0, errSimulatedWrite
} }
return f.File.Write(p) return f.File.Write(p)
} }
//nolint:ireturn // Create must return afero.File to satisfy afero.Fs.
func (fs *failingWriterFs) Create(name string) (afero.File, error) { func (fs *failingWriterFs) Create(name string) (afero.File, error) {
f, err := fs.Fs.Create(name) f, err := fs.Fs.Create(name)
if err != nil { if err != nil {
return nil, err return nil, err
} }
return &failingFile{File: f, fs: fs}, nil return &failingFile{File: f, fs: fs}, nil
} }
func TestGenerateAtomicWriteCleansUpOnError(t *testing.T) { func TestGenerateAtomicWriteCleansUpOnError(t *testing.T) {
t.Parallel()
baseFs := afero.NewMemMapFs() baseFs := afero.NewMemMapFs()
// Create test files - need enough content to trigger the write failure // Create test files - need enough content to trigger the write failure
require.NoError(t, baseFs.MkdirAll("/testdir", 0o755)) require.NoError(t, baseFs.MkdirAll(testDir, 0o755))
require.NoError(t, afero.WriteFile(baseFs, "/testdir/file1.txt", []byte("hello world this is a test file"), 0o644)) writeTestFile(t, baseFs, testFile1, "hello world this is a test file")
// Wrap with failing writer that fails after writing some bytes // Wrap with failing writer that fails after writing some bytes
fs := &failingWriterFs{Fs: baseFs, failAfter: 10} fs := &failingWriterFs{Fs: baseFs, failAfter: 10}
// Generate manifest - should fail during write // Generate manifest - should fail during write
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/output.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testOutput, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
assert.Equal(t, 1, exitCode, "should fail due to write error") assert.Equal(t, 1, exitCode, "should fail due to write error")
// With atomic writes: output.mf should NOT exist (temp was cleaned up) // With atomic writes: output.mf should NOT exist (temp was cleaned up)
// With non-atomic writes: output.mf WOULD exist (partial/empty) // With non-atomic writes: output.mf WOULD exist (partial/empty)
exists, _ := afero.Exists(baseFs, "/output.mf") exists, _ := afero.Exists(baseFs, testOutput)
assert.False(t, exists, "output file should not exist after failed generation (atomic write)") assert.False(t, exists,
"output file should not exist after failed generation (atomic write)")
// Temp file should also not exist // Temp file should also not exist
tmpExists, _ := afero.Exists(baseFs, "/output.mf.tmp") tmpExists, _ := afero.Exists(baseFs, testOutputTmp)
assert.False(t, tmpExists, "temp file should be cleaned up after failed generation") assert.False(t, tmpExists,
"temp file should be cleaned up after failed generation")
} }
func TestGenerateValidatesInputPaths(t *testing.T) { func TestGenerateValidatesInputPaths(t *testing.T) {
fs := afero.NewMemMapFs() t.Parallel()
// Create one valid directory seedValidDir := func(t *testing.T, fs afero.Fs) {
require.NoError(t, fs.MkdirAll("/validdir", 0o755)) t.Helper()
require.NoError(t, afero.WriteFile(fs, "/validdir/file.txt", []byte("content"), 0o644))
require.NoError(t, fs.MkdirAll("/validdir", 0o755))
writeTestFile(t, fs, "/validdir/file.txt", "content")
}
t.Run("nonexistent path fails fast", func(t *testing.T) { t.Run("nonexistent path fails fast", func(t *testing.T) {
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/output.mf", "/nonexistent"}, fs) t.Parallel()
exitCode := RunWithOptions(opts)
fs := afero.NewMemMapFs()
seedValidDir(t, fs)
opts := testOpts([]string{
testApp, cmdGenerate, "-q", "-o", testOutput, "/nonexistent",
}, fs)
exitCode := runCLI(opts)
assert.Equal(t, 1, exitCode) assert.Equal(t, 1, exitCode)
stderr := opts.Stderr.(*bytes.Buffer).String()
stderr := testStderr(t, opts)
assert.Contains(t, stderr, "path does not exist") assert.Contains(t, stderr, "path does not exist")
assert.Contains(t, stderr, "/nonexistent") assert.Contains(t, stderr, "/nonexistent")
}) })
t.Run("mix of valid and invalid paths fails fast", func(t *testing.T) { t.Run("mix of valid and invalid paths fails fast", func(t *testing.T) {
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/output.mf", "/validdir", "/alsononexistent"}, fs) t.Parallel()
exitCode := RunWithOptions(opts)
fs := afero.NewMemMapFs()
seedValidDir(t, fs)
opts := testOpts([]string{
testApp, cmdGenerate, "-q", "-o", testOutput,
"/validdir", "/alsononexistent",
}, fs)
exitCode := runCLI(opts)
assert.Equal(t, 1, exitCode) assert.Equal(t, 1, exitCode)
stderr := opts.Stderr.(*bytes.Buffer).String()
stderr := testStderr(t, opts)
assert.Contains(t, stderr, "path does not exist") assert.Contains(t, stderr, "path does not exist")
assert.Contains(t, stderr, "/alsononexistent") assert.Contains(t, stderr, "/alsononexistent")
// Output file should not have been created // Output file should not have been created
exists, _ := afero.Exists(fs, "/output.mf") exists, _ := afero.Exists(fs, testOutput)
assert.False(t, exists, "output file should not exist when path validation fails") assert.False(t, exists,
"output file should not exist when path validation fails")
}) })
t.Run("valid paths succeed", func(t *testing.T) { t.Run("valid paths succeed", func(t *testing.T) {
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/output.mf", "/validdir"}, fs) t.Parallel()
exitCode := RunWithOptions(opts)
fs := afero.NewMemMapFs()
seedValidDir(t, fs)
opts := testOpts([]string{
testApp, cmdGenerate, "-q", "-o", testOutput, "/validdir",
}, fs)
exitCode := runCLI(opts)
assert.Equal(t, 0, exitCode) assert.Equal(t, 0, exitCode)
}) })
} }
func TestCheckDetectsManifestCorruption(t *testing.T) { func TestCheckDetectsManifestCorruption(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
rng := rand.New(rand.NewSource(42)) rng := rand.New(rand.NewSource(42)) //nolint:gosec // deterministic test data
// Create many small files with random names to generate a ~1MB manifest // Create many small files with random names to generate a ~1MB manifest
// Each manifest entry is roughly 50-60 bytes, so we need ~20000 files // Each manifest entry is roughly 50-60 bytes, so we need ~20000 files
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll(testDir, 0o755))
numFiles := 20000 numFiles := 20000
for i := 0; i < numFiles; i++ { for range numFiles {
// Generate random filename // Generate random filename
filename := fmt.Sprintf("/testdir/%08x%08x%08x.dat", rng.Uint32(), rng.Uint32(), rng.Uint32()) filename := fmt.Sprintf("/testdir/%08x%08x%08x.dat",
rng.Uint32(), rng.Uint32(), rng.Uint32())
// Small random content // Small random content
content := make([]byte, 16+rng.Intn(48)) content := make([]byte, 16+rng.Intn(48))
rng.Read(content) _, _ = rng.Read(content)
require.NoError(t, afero.WriteFile(fs, filename, content, 0o644)) require.NoError(t, afero.WriteFile(fs, filename, content, 0o644))
} }
// Generate manifest outside of testdir // Generate manifest outside of testdir
opts := testOpts([]string{"mfer", "generate", "-q", "-o", "/manifest.mf", "/testdir"}, fs) opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", testManifest, testDir}, fs)
exitCode := RunWithOptions(opts) exitCode := runCLI(opts)
require.Equal(t, 0, exitCode, "generate should succeed") require.Equal(t, 0, exitCode, "generate should succeed")
// Read the valid manifest and verify it's approximately 1MB // Read the valid manifest and verify it's approximately 1MB
validManifest, err := afero.ReadFile(fs, "/manifest.mf") validManifest, err := afero.ReadFile(fs, testManifest)
require.NoError(t, err) require.NoError(t, err)
require.True(t, len(validManifest) >= 1024*1024, "manifest should be at least 1MB, got %d bytes", len(validManifest)) require.GreaterOrEqual(t, len(validManifest), 1024*1024,
"manifest should be at least 1MB, got %d bytes", len(validManifest))
t.Logf("manifest size: %d bytes (%d files)", len(validManifest), numFiles) t.Logf("manifest size: %d bytes (%d files)", len(validManifest), numFiles)
// First corruption: truncate the manifest // First corruption: truncate the manifest
require.NoError(t, afero.WriteFile(fs, "/manifest.mf", validManifest[:len(validManifest)/2], 0o644)) require.NoError(t, afero.WriteFile(fs, testManifest,
validManifest[:len(validManifest)/2], 0o644))
// Check should fail with truncated manifest // Check should fail with truncated manifest
opts = testOpts([]string{"mfer", "check", "-q", "--base", "/testdir", "/manifest.mf"}, fs) opts = testOpts([]string{
exitCode = RunWithOptions(opts) testApp, cmdCheck, "-q", testFlagBase, testDir, testManifest,
}, fs)
exitCode = runCLI(opts)
assert.Equal(t, 1, exitCode, "check should fail with truncated manifest") assert.Equal(t, 1, exitCode, "check should fail with truncated manifest")
// Verify check passes with valid manifest // Verify check passes with valid manifest
require.NoError(t, afero.WriteFile(fs, "/manifest.mf", validManifest, 0o644)) require.NoError(t, afero.WriteFile(fs, testManifest, validManifest, 0o644))
opts = testOpts([]string{"mfer", "check", "-q", "--base", "/testdir", "/manifest.mf"}, fs)
exitCode = RunWithOptions(opts) opts = testOpts([]string{
testApp, cmdCheck, "-q", testFlagBase, testDir, testManifest,
}, fs)
exitCode = runCLI(opts)
require.Equal(t, 0, exitCode, "check should pass with valid manifest") require.Equal(t, 0, exitCode, "check should pass with valid manifest")
// Now do 500 random corruption iterations // Now do 500 random corruption iterations
for i := 0; i < 500; i++ { for i := range 500 {
// Corrupt: write a random byte at a random offset // Corrupt: write a random byte at a random offset
corrupted := make([]byte, len(validManifest)) corrupted := make([]byte, len(validManifest))
copy(corrupted, validManifest) copy(corrupted, validManifest)
offset := rng.Intn(len(corrupted)) offset := rng.Intn(len(corrupted))
originalByte := corrupted[offset] originalByte := corrupted[offset]
// Make sure we actually change the byte // Make sure we actually change the byte
newByte := byte(rng.Intn(256)) buf := make([]byte, 1)
newByte := originalByte
for newByte == originalByte { for newByte == originalByte {
newByte = byte(rng.Intn(256)) _, _ = rng.Read(buf)
newByte = buf[0]
} }
corrupted[offset] = newByte corrupted[offset] = newByte
require.NoError(t, afero.WriteFile(fs, "/manifest.mf", corrupted, 0o644)) require.NoError(t, afero.WriteFile(fs, testManifest, corrupted, 0o644))
// Check should fail with corrupted manifest // Check should fail with corrupted manifest
opts = testOpts([]string{"mfer", "check", "-q", "--base", "/testdir", "/manifest.mf"}, fs) opts = testOpts([]string{
exitCode = RunWithOptions(opts) testApp, cmdCheck, "-q", testFlagBase, testDir, testManifest,
assert.Equal(t, 1, exitCode, "iteration %d: check should fail with corrupted manifest (offset %d, 0x%02x -> 0x%02x)", }, fs)
exitCode = runCLI(opts)
assert.Equal(t, 1, exitCode,
"iteration %d: check should fail with corrupted manifest "+
"(offset %d, 0x%02x -> 0x%02x)",
i, offset, originalByte, newByte) i, offset, originalByte, newByte)
// Restore valid manifest for next iteration // Restore valid manifest for next iteration
require.NoError(t, afero.WriteFile(fs, "/manifest.mf", validManifest, 0o644)) require.NoError(t, afero.WriteFile(fs, testManifest, validManifest, 0o644))
} }
} }

168
internal/cli/errmsg_test.go Normal file
View File

@@ -0,0 +1,168 @@
//nolint:testpackage // white-box tests exercise unexported internals
package cli
import (
"fmt"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// errMsgCase is one pinned user-visible error message.
type errMsgCase struct {
name string
err error
want string
}
const (
msgFpA = "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"
msgFpB = "BBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBB"
)
func checkErrMsgCases(t *testing.T, cases []errMsgCase) {
t.Helper()
for _, tc := range cases {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
assert.Equal(t, tc.want, tc.err.Error())
})
}
}
// TestErrorMessagesVerbatim pins the exact rendered text of the CLI's
// user-visible error messages.
//
// These strings are an interface: they are grepped for in CI pipelines
// and quoted in bug reports. The messages are assembled by wrapping
// static sentinels, and it is easy to change what a user sees while
// only meaning to make an error matchable with errors.Is - which is
// precisely what happened once already. Any change to a string below is
// therefore a deliberate, separately stated change, never a side effect
// of a refactor.
func TestErrorMessagesVerbatim(t *testing.T) {
t.Parallel()
checkErrMsgCases(t, []errMsgCase{
{
name: "check: no manifest found",
err: fmt.Errorf("%w in %s (looked for index.mf and .index.mf)",
errNoManifestFound, "/tmp/x"),
want: "no manifest found in /tmp/x " +
"(looked for index.mf and .index.mf)",
},
{
name: "check: invalid fingerprint length",
err: fmt.Errorf("%w, got %d", errInvalidFingerprint, 8),
want: "invalid fingerprint: must be exactly 40 hex characters, got 8",
},
{
name: "check: manifest not signed",
err: fmt.Errorf("%w, but signature from %s is required",
errManifestNotSigned, msgFpA),
want: "manifest is not signed, but signature from " + msgFpA +
" is required",
},
{
name: "check: signer mismatch",
err: fmt.Errorf("embedded signing key fingerprint %s %w %s",
msgFpA, errSignerMismatch, msgFpB),
want: "embedded signing key fingerprint " + msgFpA +
" does not match required " + msgFpB,
},
{
name: "gen: path does not exist",
err: fmt.Errorf("%w: %s", errPathNotExist, "nope"),
want: "path does not exist: nope",
},
{
name: "gen: output file exists",
err: fmt.Errorf("output file %s %w", "index.mf", errOutputExists),
want: "output file index.mf already exists " +
"(use --force to overwrite)",
},
{
name: "mfer: unknown command",
err: fmt.Errorf("%w %q", errUnknownCommand, "bogus"),
want: `unknown command "bogus"`,
},
})
}
// TestFetchErrorMessagesVerbatim pins the fetch and manifest-loader
// messages; see TestErrorMessagesVerbatim for why.
func TestFetchErrorMessagesVerbatim(t *testing.T) {
t.Parallel()
checkErrMsgCases(t, []errMsgCase{
{
name: "manifest_loader: http status",
err: fmt.Errorf("failed to fetch %s: %w %d",
"https://example.com/index.mf", errHTTPStatus, 404),
want: "failed to fetch https://example.com/index.mf: HTTP 404",
},
{
name: "fetch: manifest http status",
err: fmt.Errorf("failed to fetch manifest: %w %d",
errHTTPStatus, 404),
want: "failed to fetch manifest: HTTP 404",
},
{
name: "fetch: file http status",
err: fmt.Errorf("%w %d", errHTTPStatus, 500),
want: "HTTP 500",
},
{
name: "fetch: empty path",
err: errEmptyPath,
want: "empty path",
},
{
name: "fetch: absolute path",
err: fmt.Errorf("%w: %s", errAbsolutePath, "/etc/passwd"),
want: "absolute path not allowed: /etc/passwd",
},
{
name: "fetch: path traversal",
err: fmt.Errorf("%w: %s", errPathTraversal, "../x"),
want: "path traversal not allowed: ../x",
},
{
name: "fetch: size mismatch",
err: fmt.Errorf("%w: expected %d bytes, got %d",
errSizeMismatch, 10, 9),
want: "size mismatch: expected 10 bytes, got 9",
},
{
name: "fetch: url required",
err: errURLRequired,
want: "URL argument required",
},
{
name: "fetch: hash mismatch",
err: errHashMismatch,
want: "hash mismatch",
},
})
}
// TestSentinelsAreMatchable checks that the wrapped forms of the
// messages above remain matchable with errors.Is, which is the reason
// the sentinels exist at all.
func TestSentinelsAreMatchable(t *testing.T) {
t.Parallel()
wrapped := fmt.Errorf("embedded signing key fingerprint %s %w %s",
"a", errSignerMismatch, "b")
require.ErrorIs(t, wrapped, errSignerMismatch)
wrapped = fmt.Errorf("output file %s %w", "index.mf", errOutputExists)
require.ErrorIs(t, wrapped, errOutputExists)
wrapped = fmt.Errorf("failed to fetch manifest: %w %d", errHTTPStatus, 404)
require.ErrorIs(t, wrapped, errHTTPStatus)
assert.NotErrorIs(t, errHashMismatch, errSizeMismatch)
}

77
internal/cli/export.go Normal file
View File

@@ -0,0 +1,77 @@
package cli
import (
"encoding/hex"
"encoding/json"
"fmt"
"time"
"github.com/urfave/cli/v2"
"sneak.berlin/go/mfer/mfer"
)
// ExportEntry represents a single file entry in the exported JSON output.
type ExportEntry struct {
Path string `json:"path"`
Size int64 `json:"size"`
Hashes []string `json:"hashes"`
Mtime *string `json:"mtime,omitempty"`
Ctime *string `json:"ctime,omitempty"`
}
func (mfa *CLIApp) exportManifestOperation(ctx *cli.Context) error {
pathOrURL, err := mfa.resolveManifestArg(ctx)
if err != nil {
return fmt.Errorf("export: %w", err)
}
rc, err := mfa.openManifestReader(pathOrURL)
if err != nil {
return fmt.Errorf("export: %w", err)
}
defer func() { _ = rc.Close() }()
manifest, err := mfer.NewManifestFromReader(rc)
if err != nil {
return fmt.Errorf("export: failed to parse manifest: %w", err)
}
files := manifest.Files()
entries := make([]ExportEntry, 0, len(files))
for _, f := range files {
entry := ExportEntry{
Path: f.GetPath(),
Size: f.GetSize(),
Hashes: make([]string, 0, len(f.GetHashes())),
}
for _, h := range f.GetHashes() {
entry.Hashes = append(entry.Hashes, hex.EncodeToString(h.GetMultiHash()))
}
if mtime, ok := entryMtime(f); ok {
t := mtime.UTC().Format(time.RFC3339Nano)
entry.Mtime = &t
}
if f.GetCtime() != nil {
t := time.Unix(f.GetCtime().GetSeconds(), int64(f.GetCtime().GetNanos())).
UTC().Format(time.RFC3339Nano)
entry.Ctime = &t
}
entries = append(entries, entry)
}
enc := json.NewEncoder(mfa.Stdout)
enc.SetIndent("", " ")
err = enc.Encode(entries)
if err != nil {
return fmt.Errorf("export: failed to encode JSON: %w", err)
}
return nil
}

156
internal/cli/export_test.go Normal file
View File

@@ -0,0 +1,156 @@
package cli
import (
"bytes"
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/mfer/mfer"
)
const testCmdExport = "export"
// buildTestManifest creates a manifest from in-memory files and returns its bytes.
func buildTestManifest(t *testing.T, files map[string][]byte) []byte {
t.Helper()
sourceFs := afero.NewMemMapFs()
for path, content := range files {
require.NoError(t, sourceFs.MkdirAll("/", 0o755))
require.NoError(t, afero.WriteFile(sourceFs, "/"+path, content, 0o644))
}
opts := &mfer.ScannerOptions{Fs: sourceFs}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
var buf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &buf, nil))
return buf.Bytes()
}
func TestExportManifestOperation(t *testing.T) {
t.Parallel()
testFiles := map[string][]byte{
"hello.txt": []byte("Hello, World!"),
"sub/file.txt": []byte("nested content"),
}
manifestData := buildTestManifest(t, testFiles)
// Write manifest to memfs
fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/test.mf", manifestData, 0o644))
var stdout, stderr bytes.Buffer
exitCode := runCLI(&RunOptions{
Appname: testApp,
Args: []string{testApp, testCmdExport, "/test.mf"},
Stdin: &bytes.Buffer{},
Stdout: &stdout,
Stderr: &stderr,
Fs: fs,
})
require.Equal(t, 0, exitCode, "stderr: %s", stderr.String())
var entries []ExportEntry
require.NoError(t, json.Unmarshal(stdout.Bytes(), &entries))
assert.Len(t, entries, 2)
// Verify entries have expected fields
pathSet := make(map[string]bool)
for _, e := range entries {
pathSet[e.Path] = true
assert.NotEmpty(t, e.Hashes, "entry %s should have hashes", e.Path)
assert.Positive(t, e.Size, "entry %s should have positive size", e.Path)
}
assert.True(t, pathSet["hello.txt"])
assert.True(t, pathSet["sub/file.txt"])
}
func TestExportFromHTTPURL(t *testing.T) {
t.Parallel()
testFiles := map[string][]byte{
"a.txt": []byte("aaa"),
}
manifestData := buildTestManifest(t, testFiles)
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(manifestData)
}))
defer server.Close()
var stdout, stderr bytes.Buffer
exitCode := runCLI(&RunOptions{
Appname: testApp,
Args: []string{testApp, testCmdExport, server.URL + "/index.mf"},
Stdin: &bytes.Buffer{},
Stdout: &stdout,
Stderr: &stderr,
Fs: afero.NewMemMapFs(),
})
require.Equal(t, 0, exitCode, "stderr: %s", stderr.String())
var entries []ExportEntry
require.NoError(t, json.Unmarshal(stdout.Bytes(), &entries))
assert.Len(t, entries, 1)
assert.Equal(t, "a.txt", entries[0].Path)
}
func TestListFromHTTPURL(t *testing.T) {
t.Parallel()
testFiles := map[string][]byte{
"one.txt": []byte("1"),
"two.txt": []byte("22"),
}
manifestData := buildTestManifest(t, testFiles)
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
_, _ = w.Write(manifestData)
}))
defer server.Close()
var stdout, stderr bytes.Buffer
exitCode := runCLI(&RunOptions{
Appname: testApp,
Args: []string{testApp, "list", server.URL + "/index.mf"},
Stdin: &bytes.Buffer{},
Stdout: &stdout,
Stderr: &stderr,
Fs: afero.NewMemMapFs(),
})
require.Equal(t, 0, exitCode, "stderr: %s", stderr.String())
output := stdout.String()
assert.Contains(t, output, "one.txt")
assert.Contains(t, output, "two.txt")
}
func TestIsHTTPURL(t *testing.T) {
t.Parallel()
assert.True(t, isHTTPURL("http://example.com/manifest.mf"))
assert.True(t, isHTTPURL("https://example.com/manifest.mf"))
assert.False(t, isHTTPURL("/local/path.mf"))
assert.False(t, isHTTPURL("relative/path.mf"))
assert.False(t, isHTTPURL("ftp://example.com/file"))
}

View File

@@ -2,7 +2,9 @@ package cli
import ( import (
"bytes" "bytes"
"context"
"crypto/sha256" "crypto/sha256"
"errors"
"fmt" "fmt"
"io" "io"
"net/http" "net/http"
@@ -20,6 +22,45 @@ import (
"sneak.berlin/go/mfer/mfer" "sneak.berlin/go/mfer/mfer"
) )
const (
// progressChanBuffer is the buffer size of the download progress
// channel.
progressChanBuffer = 10
// bitsPerByte converts a bytes-per-second rate to bits per second.
bitsPerByte = 8
// dirPerms is the permission mode for directories created for
// downloaded files. Fetched trees are content that is normally
// published (served by a web server, read by another uid), so the
// traversal bit for group and other must stay set.
dirPerms os.FileMode = 0o755
// Bitrate unit thresholds in bits per second.
bpsPerGbps = 1e9
bpsPerMbps = 1e6
bpsPerKbps = 1e3
)
var (
// errURLRequired indicates the fetch command was run without a URL
// argument.
errURLRequired = errors.New("URL argument required")
// errEmptyPath indicates an empty file path in the manifest.
errEmptyPath = errors.New("empty path")
// errAbsolutePath indicates an absolute file path in the manifest.
errAbsolutePath = errors.New("absolute path not allowed")
// errPathTraversal indicates a manifest path escaping the target
// directory.
errPathTraversal = errors.New("path traversal not allowed")
// errSizeMismatch indicates a downloaded file with an unexpected
// size.
errSizeMismatch = errors.New("size mismatch")
// errHashMismatch indicates a downloaded file whose hash matches no
// manifest hash.
errHashMismatch = errors.New("hash mismatch")
)
// DownloadProgress reports the progress of a single file download. // DownloadProgress reports the progress of a single file download.
type DownloadProgress struct { type DownloadProgress struct {
Path string // File path being downloaded Path string // File path being downloaded
@@ -29,14 +70,98 @@ type DownloadProgress struct {
ETA time.Duration // Estimated time to completion ETA time.Duration // Estimated time to completion
} }
// httpGet issues a GET request for the given URL using the provided
// context and returns the response. The caller must close the body.
//
// Errors are returned unwrapped: this helper replaced direct http.Get
// calls, and each caller already supplies its own context string, so
// adding one here would change user-visible messages.
func httpGet(ctx context.Context, fileURL string) (*http.Response, error) {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, fileURL, nil)
if err != nil {
return nil, err
}
resp, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
return resp, nil
}
// reportDownloadProgress renders download progress until the channel
// closes, then closes done.
func reportDownloadProgress(progress <-chan DownloadProgress, done chan<- struct{}) {
defer close(done)
for p := range progress {
rate := formatBitrate(p.BytesPerSec * bitsPerByte)
if p.ETA > 0 {
log.Infof("%s: %s/%s, %s, ETA %s",
p.Path, humanize.IBytes(safeUint64(p.BytesRead)),
humanize.IBytes(safeUint64(p.TotalBytes)),
rate, p.ETA.Round(time.Second))
} else {
log.Infof("%s: %s/%s, %s",
p.Path, humanize.IBytes(safeUint64(p.BytesRead)),
humanize.IBytes(safeUint64(p.TotalBytes)), rate)
}
}
}
// manifestBaseURL returns the URL of the directory containing the
// manifest, with a trailing slash.
func manifestBaseURL(manifestURL string) (*url.URL, error) {
baseURL, err := url.Parse(manifestURL)
if err != nil {
return nil, fmt.Errorf("fetch: invalid manifest URL: %w", err)
}
baseURL.Path = path.Dir(baseURL.Path)
if !strings.HasSuffix(baseURL.Path, "/") {
baseURL.Path += "/"
}
return baseURL, nil
}
// downloadManifestFiles downloads every file in the manifest, reporting
// progress on the progress channel.
func downloadManifestFiles(
ctx context.Context,
baseURL *url.URL,
files []*mfer.MFFilePath,
progress chan<- DownloadProgress,
) error {
for _, f := range files {
// Sanitize the path to prevent path traversal attacks
localPath, err := sanitizePath(f.GetPath())
if err != nil {
return fmt.Errorf("invalid path in manifest: %w", err)
}
fileURL := baseURL.String() + encodeFilePath(f.GetPath())
log.Infof("fetching %s", f.GetPath())
err = downloadFile(ctx, fileURL, localPath, f, progress)
if err != nil {
return fmt.Errorf("failed to download %s: %w", f.GetPath(), err)
}
}
return nil
}
func (mfa *CLIApp) fetchManifestOperation(ctx *cli.Context) error { func (mfa *CLIApp) fetchManifestOperation(ctx *cli.Context) error {
log.Debug("fetchManifestOperation()") log.Debug("fetchManifestOperation()")
if ctx.Args().Len() == 0 { if ctx.Args().Len() == 0 {
return fmt.Errorf("URL argument required") return errURLRequired
} }
inputURL := ctx.Args().Get(0) inputURL := ctx.Args().Get(0)
manifestURL, err := resolveManifestURL(inputURL) manifestURL, err := resolveManifestURL(inputURL)
if err != nil { if err != nil {
return fmt.Errorf("invalid URL: %w", err) return fmt.Errorf("invalid URL: %w", err)
@@ -45,14 +170,16 @@ func (mfa *CLIApp) fetchManifestOperation(ctx *cli.Context) error {
log.Infof("fetching manifest from %s", manifestURL) log.Infof("fetching manifest from %s", manifestURL)
// Fetch manifest // Fetch manifest
resp, err := http.Get(manifestURL) resp, err := httpGet(ctx.Context, manifestURL)
if err != nil { if err != nil {
return fmt.Errorf("failed to fetch manifest: %w", err) return fmt.Errorf("failed to fetch manifest: %w", err)
} }
defer func() { _ = resp.Body.Close() }() defer func() { _ = resp.Body.Close() }()
if resp.StatusCode != http.StatusOK { if resp.StatusCode != http.StatusOK {
return fmt.Errorf("failed to fetch manifest: HTTP %d", resp.StatusCode) return fmt.Errorf("failed to fetch manifest: %w %d",
errHTTPStatus, resp.StatusCode)
} }
// Parse manifest // Parse manifest
@@ -65,74 +192,43 @@ func (mfa *CLIApp) fetchManifestOperation(ctx *cli.Context) error {
log.Infof("manifest contains %d files", len(files)) log.Infof("manifest contains %d files", len(files))
// Compute base URL (directory containing manifest) // Compute base URL (directory containing manifest)
baseURL, err := url.Parse(manifestURL) baseURL, err := manifestBaseURL(manifestURL)
if err != nil { if err != nil {
return err return err
} }
baseURL.Path = path.Dir(baseURL.Path)
if !strings.HasSuffix(baseURL.Path, "/") {
baseURL.Path += "/"
}
// Calculate total bytes to download // Calculate total bytes to download
var totalBytes int64 var totalBytes int64
for _, f := range files { for _, f := range files {
totalBytes += f.Size totalBytes += f.GetSize()
} }
// Create progress channel // Create progress channel and start progress reporter goroutine
progress := make(chan DownloadProgress, 10) progress := make(chan DownloadProgress, progressChanBuffer)
// Start progress reporter goroutine
done := make(chan struct{}) done := make(chan struct{})
go func() {
defer close(done) go reportDownloadProgress(progress, done)
for p := range progress {
rate := formatBitrate(p.BytesPerSec * 8)
if p.ETA > 0 {
log.Infof("%s: %s/%s, %s, ETA %s",
p.Path, humanize.IBytes(uint64(p.BytesRead)), humanize.IBytes(uint64(p.TotalBytes)),
rate, p.ETA.Round(time.Second))
} else {
log.Infof("%s: %s/%s, %s",
p.Path, humanize.IBytes(uint64(p.BytesRead)), humanize.IBytes(uint64(p.TotalBytes)), rate)
}
}
}()
// Track download start time // Track download start time
startTime := time.Now() startTime := time.Now()
// Download each file // Download each file
for _, f := range files { dlErr := downloadManifestFiles(ctx.Context, baseURL, files, progress)
// Sanitize the path to prevent path traversal attacks
localPath, err := sanitizePath(f.Path)
if err != nil {
close(progress)
<-done
return fmt.Errorf("invalid path in manifest: %w", err)
}
fileURL := baseURL.String() + encodeFilePath(f.Path)
log.Infof("fetching %s", f.Path)
if err := downloadFile(fileURL, localPath, f, progress); err != nil {
close(progress)
<-done
return fmt.Errorf("failed to download %s: %w", f.Path, err)
}
}
close(progress) close(progress)
<-done <-done
if dlErr != nil {
return dlErr
}
// Print summary // Print summary
elapsed := time.Since(startTime) elapsed := time.Since(startTime)
avgBytesPerSec := float64(totalBytes) / elapsed.Seconds() avgBytesPerSec := float64(totalBytes) / elapsed.Seconds()
avgRate := formatBitrate(avgBytesPerSec * 8) avgRate := formatBitrate(avgBytesPerSec * bitsPerByte)
log.Infof("downloaded %d files (%s) in %.1fs (%s avg)", log.Infof("downloaded %d files (%s) in %.1fs (%s avg)",
len(files), len(files),
humanize.IBytes(uint64(totalBytes)), humanize.IBytes(safeUint64(totalBytes)),
elapsed.Seconds(), elapsed.Seconds(),
avgRate) avgRate)
@@ -145,6 +241,7 @@ func encodeFilePath(p string) string {
for i, seg := range segments { for i, seg := range segments {
segments[i] = url.PathEscape(seg) segments[i] = url.PathEscape(seg)
} }
return strings.Join(segments, "/") return strings.Join(segments, "/")
} }
@@ -153,12 +250,12 @@ func encodeFilePath(p string) string {
func sanitizePath(p string) (string, error) { func sanitizePath(p string) (string, error) {
// Reject empty paths // Reject empty paths
if p == "" { if p == "" {
return "", fmt.Errorf("empty path") return "", errEmptyPath
} }
// Reject absolute paths // Reject absolute paths
if filepath.IsAbs(p) { if filepath.IsAbs(p) {
return "", fmt.Errorf("absolute path not allowed: %s", p) return "", fmt.Errorf("%w: %s", errAbsolutePath, p)
} }
// Clean the path to resolve . and .. // Clean the path to resolve . and ..
@@ -166,12 +263,12 @@ func sanitizePath(p string) (string, error) {
// Reject paths that escape the current directory // Reject paths that escape the current directory
if strings.HasPrefix(cleaned, ".."+string(filepath.Separator)) || cleaned == ".." { if strings.HasPrefix(cleaned, ".."+string(filepath.Separator)) || cleaned == ".." {
return "", fmt.Errorf("path traversal not allowed: %s", p) return "", fmt.Errorf("%w: %s", errPathTraversal, p)
} }
// Also check for absolute paths after cleaning (handles edge cases) // Also check for absolute paths after cleaning (handles edge cases)
if filepath.IsAbs(cleaned) { if filepath.IsAbs(cleaned) {
return "", fmt.Errorf("absolute path not allowed: %s", p) return "", fmt.Errorf("%w: %s", errAbsolutePath, p)
} }
return cleaned, nil return cleaned, nil
@@ -214,10 +311,14 @@ type progressWriter struct {
func (pw *progressWriter) Write(p []byte) (int, error) { func (pw *progressWriter) Write(p []byte) (int, error) {
n, err := pw.w.Write(p) n, err := pw.w.Write(p)
pw.written += int64(n) pw.written += int64(n)
if pw.progress != nil { if pw.progress != nil {
var bytesPerSec float64 var (
var eta time.Duration bytesPerSec float64
eta time.Duration
)
elapsed := time.Since(pw.startTime) elapsed := time.Since(pw.startTime)
if elapsed > 0 && pw.written > 0 { if elapsed > 0 && pw.written > 0 {
bytesPerSec = float64(pw.written) / elapsed.Seconds() bytesPerSec = float64(pw.written) / elapsed.Seconds()
@@ -226,6 +327,7 @@ func (pw *progressWriter) Write(p []byte) (int, error) {
eta = time.Duration(float64(remainingBytes)/bytesPerSec) * time.Second eta = time.Duration(float64(remainingBytes)/bytesPerSec) * time.Second
} }
} }
sendProgress(pw.progress, DownloadProgress{ sendProgress(pw.progress, DownloadProgress{
Path: pw.path, Path: pw.path,
BytesRead: pw.written, BytesRead: pw.written,
@@ -234,18 +336,19 @@ func (pw *progressWriter) Write(p []byte) (int, error) {
ETA: eta, ETA: eta,
}) })
} }
return n, err return n, err
} }
// formatBitrate formats a bits-per-second value with appropriate unit prefix. // formatBitrate formats a bits-per-second value with appropriate unit prefix.
func formatBitrate(bps float64) string { func formatBitrate(bps float64) string {
switch { switch {
case bps >= 1e9: case bps >= bpsPerGbps:
return fmt.Sprintf("%.1f Gbps", bps/1e9) return fmt.Sprintf("%.1f Gbps", bps/bpsPerGbps)
case bps >= 1e6: case bps >= bpsPerMbps:
return fmt.Sprintf("%.1f Mbps", bps/1e6) return fmt.Sprintf("%.1f Mbps", bps/bpsPerMbps)
case bps >= 1e3: case bps >= bpsPerKbps:
return fmt.Sprintf("%.1f Kbps", bps/1e3) return fmt.Sprintf("%.1f Kbps", bps/bpsPerKbps)
default: default:
return fmt.Sprintf("%.0f bps", bps) return fmt.Sprintf("%.0f bps", bps)
} }
@@ -259,55 +362,102 @@ func sendProgress(ch chan<- DownloadProgress, p DownloadProgress) {
} }
} }
// downloadFile downloads a URL to a local file path with hash verification. // tempPathFor computes the temporary download path for a local file.
// It downloads to a temporary file, verifies the hash, then renames to the final path. // For dotfiles, just append .tmp (they're already hidden); for regular
// Progress is reported via the progress channel. // files, prefix with . and append .tmp.
func downloadFile(fileURL, localPath string, entry *mfer.MFFilePath, progress chan<- DownloadProgress) error { func tempPathFor(localPath string) string {
// Create parent directories if needed
dir := filepath.Dir(localPath) dir := filepath.Dir(localPath)
if dir != "" && dir != "." {
if err := os.MkdirAll(dir, 0o755); err != nil {
return err
}
}
// Compute temp file path in the same directory
// For dotfiles, just append .tmp (they're already hidden)
// For regular files, prefix with . and append .tmp
base := filepath.Base(localPath) base := filepath.Base(localPath)
var tmpName string var tmpName string
if strings.HasPrefix(base, ".") { if strings.HasPrefix(base, ".") {
tmpName = base + ".tmp" tmpName = base + ".tmp"
} else { } else {
tmpName = "." + base + ".tmp" tmpName = "." + base + ".tmp"
} }
tmpPath := filepath.Join(dir, tmpName)
if dir == "" || dir == "." { if dir == "" || dir == "." {
tmpPath = tmpName return tmpName
} }
// Fetch file return filepath.Join(dir, tmpName)
resp, err := http.Get(fileURL) }
// verifyDownloadedHash checks the computed sha256 digest against the
// manifest entry's hashes; at least one must match.
func verifyDownloadedHash(digest []byte, entry *mfer.MFFilePath) error {
computed, err := multihash.Encode(digest, multihash.SHA2_256)
if err != nil { if err != nil {
return err return fmt.Errorf("failed to encode hash: %w", err)
} }
for _, hash := range entry.GetHashes() {
if bytes.Equal(computed, hash.GetMultiHash()) {
return nil
}
}
return errHashMismatch
}
// downloadFile downloads a URL to a local file path with hash verification.
// It downloads to a temporary file, verifies the hash, then renames to the final path.
// Progress is reported via the progress channel.
func downloadFile(
ctx context.Context,
fileURL, localPath string,
entry *mfer.MFFilePath,
progress chan<- DownloadProgress,
) error {
// Enforce the path invariant here rather than relying on the caller,
// so every entry point to downloadFile gets the same treatment.
localPath, err := sanitizePath(localPath)
if err != nil {
return fmt.Errorf("invalid path: %w", err)
}
// Create parent directories if needed
dir := filepath.Dir(localPath)
if dir != "" && dir != "." {
err := os.MkdirAll(dir, dirPerms)
if err != nil {
return fmt.Errorf("failed to create directory %s: %w", dir, err)
}
}
tmpPath := tempPathFor(localPath)
// Fetch file
resp, err := httpGet(ctx, fileURL)
if err != nil {
return fmt.Errorf("HTTP request failed: %w", err)
}
defer func() { _ = resp.Body.Close() }() defer func() { _ = resp.Body.Close() }()
if resp.StatusCode != http.StatusOK { if resp.StatusCode != http.StatusOK {
return fmt.Errorf("HTTP %d", resp.StatusCode) return fmt.Errorf("%w %d", errHTTPStatus, resp.StatusCode)
} }
// Determine expected size // Determine expected size
expectedSize := entry.Size expectedSize := entry.GetSize()
totalBytes := resp.ContentLength totalBytes := resp.ContentLength
if totalBytes < 0 { if totalBytes < 0 {
totalBytes = expectedSize totalBytes = expectedSize
} }
// Create temp file // Create temp file.
out, err := os.Create(tmpPath) //
// G304: tmpPath is derived from localPath, which sanitizePath above
// constrains lexically to a relative path that does not escape the
// destination directory. That is a purely lexical guarantee: it does
// not resolve symlinks, so a pre-existing symlink inside the
// destination tree can still redirect this write outside of it
// (tracked in issue #86).
out, err := os.Create(tmpPath) //nolint:gosec // G304: see comment above
if err != nil { if err != nil {
return err return fmt.Errorf("failed to create temp file: %w", err)
} }
// Set up hash computation // Set up hash computation
@@ -328,45 +478,50 @@ func downloadFile(fileURL, localPath string, entry *mfer.MFFilePath, progress ch
// Close file before checking errors (to flush writes) // Close file before checking errors (to flush writes)
closeErr := out.Close() closeErr := out.Close()
// If copy failed, clean up temp file and return error err = finishDownload(
if copyErr != nil { tmpPath, localPath, written, expectedSize, h.Sum(nil), entry,
copyErr, closeErr)
if err != nil {
_ = os.Remove(tmpPath) _ = os.Remove(tmpPath)
return err
}
return nil
}
// finishDownload validates the copy result, verifies size and hash, and
// moves the temp file into place. On error the caller removes tmpPath.
func finishDownload(
tmpPath, localPath string,
written, expectedSize int64,
digest []byte,
entry *mfer.MFFilePath,
copyErr, closeErr error,
) error {
if copyErr != nil {
return copyErr return copyErr
} }
if closeErr != nil { if closeErr != nil {
_ = os.Remove(tmpPath)
return closeErr return closeErr
} }
// Verify size // Verify size
if written != expectedSize { if written != expectedSize {
_ = os.Remove(tmpPath) return fmt.Errorf("%w: expected %d bytes, got %d",
return fmt.Errorf("size mismatch: expected %d bytes, got %d", expectedSize, written) errSizeMismatch, expectedSize, written)
}
// Encode computed hash as multihash
computed, err := multihash.Encode(h.Sum(nil), multihash.SHA2_256)
if err != nil {
_ = os.Remove(tmpPath)
return fmt.Errorf("failed to encode hash: %w", err)
} }
// Verify hash against manifest (at least one must match) // Verify hash against manifest (at least one must match)
hashMatch := false err := verifyDownloadedHash(digest, entry)
for _, hash := range entry.Hashes { if err != nil {
if bytes.Equal(computed, hash.MultiHash) { return err
hashMatch = true
break
}
}
if !hashMatch {
_ = os.Remove(tmpPath)
return fmt.Errorf("hash mismatch")
} }
// Rename temp file to final path // Rename temp file to final path
if err := os.Rename(tmpPath, localPath); err != nil { err = os.Rename(tmpPath, localPath)
_ = os.Remove(tmpPath) if err != nil {
return fmt.Errorf("failed to rename temp file: %w", err) return fmt.Errorf("failed to rename temp file: %w", err)
} }

View File

@@ -1,3 +1,4 @@
//nolint:testpackage // white-box tests exercise unexported internals
package cli package cli
import ( import (
@@ -16,13 +17,25 @@ import (
"sneak.berlin/go/mfer/mfer" "sneak.berlin/go/mfer/mfer"
) )
const (
testFileTxt = "file.txt"
testDirFile = "dir/file.txt"
testIndexMF = "https://example.com/path/index.mf"
// Exactly what url.Parse renders, with no wrapper of our own.
urlParseControlCharErr = `parse "http://example.com/\x7f": ` +
`net/url: invalid control character in URL`
)
func TestEncodeFilePath(t *testing.T) { func TestEncodeFilePath(t *testing.T) {
t.Parallel()
tests := []struct { tests := []struct {
input string input string
expected string expected string
}{ }{
{"file.txt", "file.txt"}, {testFileTxt, testFileTxt},
{"dir/file.txt", "dir/file.txt"}, {testDirFile, testDirFile},
{"my file.txt", "my%20file.txt"}, {"my file.txt", "my%20file.txt"},
{"dir/my file.txt", "dir/my%20file.txt"}, {"dir/my file.txt", "dir/my%20file.txt"},
{"file#1.txt", "file%231.txt"}, {"file#1.txt", "file%231.txt"},
@@ -33,6 +46,8 @@ func TestEncodeFilePath(t *testing.T) {
for _, tt := range tests { for _, tt := range tests {
t.Run(tt.input, func(t *testing.T) { t.Run(tt.input, func(t *testing.T) {
t.Parallel()
result := encodeFilePath(tt.input) result := encodeFilePath(tt.input)
assert.Equal(t, tt.expected, result) assert.Equal(t, tt.expected, result)
}) })
@@ -40,23 +55,27 @@ func TestEncodeFilePath(t *testing.T) {
} }
func TestSanitizePath(t *testing.T) { func TestSanitizePath(t *testing.T) {
t.Parallel()
// Valid paths that should be accepted // Valid paths that should be accepted
validTests := []struct { validTests := []struct {
input string input string
expected string expected string
}{ }{
{"file.txt", "file.txt"}, {testFileTxt, testFileTxt},
{"dir/file.txt", "dir/file.txt"}, {testDirFile, testDirFile},
{"dir/subdir/file.txt", "dir/subdir/file.txt"}, {"dir/subdir/file.txt", "dir/subdir/file.txt"},
{"./file.txt", "file.txt"}, {"./file.txt", testFileTxt},
{"./dir/file.txt", "dir/file.txt"}, {"./dir/file.txt", testDirFile},
{"dir/./file.txt", "dir/file.txt"}, {"dir/./file.txt", testDirFile},
} }
for _, tt := range validTests { for _, tt := range validTests {
t.Run("valid:"+tt.input, func(t *testing.T) { t.Run("valid:"+tt.input, func(t *testing.T) {
t.Parallel()
result, err := sanitizePath(tt.input) result, err := sanitizePath(tt.input)
assert.NoError(t, err) require.NoError(t, err)
assert.Equal(t, tt.expected, result) assert.Equal(t, tt.expected, result)
}) })
} }
@@ -78,6 +97,8 @@ func TestSanitizePath(t *testing.T) {
for _, tt := range invalidTests { for _, tt := range invalidTests {
t.Run("invalid:"+tt.desc, func(t *testing.T) { t.Run("invalid:"+tt.desc, func(t *testing.T) {
t.Parallel()
_, err := sanitizePath(tt.input) _, err := sanitizePath(tt.input)
assert.Error(t, err, "expected error for path: %s", tt.input) assert.Error(t, err, "expected error for path: %s", tt.input)
}) })
@@ -85,36 +106,115 @@ func TestSanitizePath(t *testing.T) {
} }
func TestResolveManifestURL(t *testing.T) { func TestResolveManifestURL(t *testing.T) {
t.Parallel()
tests := []struct { tests := []struct {
input string input string
expected string expected string
}{ }{
// Already ends with .mf - use as-is // Already ends with .mf - use as-is
{"https://example.com/path/index.mf", "https://example.com/path/index.mf"}, {testIndexMF, testIndexMF},
{"https://example.com/path/custom.mf", "https://example.com/path/custom.mf"}, {"https://example.com/path/custom.mf", "https://example.com/path/custom.mf"},
{"https://example.com/foo.mf", "https://example.com/foo.mf"}, {"https://example.com/foo.mf", "https://example.com/foo.mf"},
// Directory with trailing slash - append index.mf // Directory with trailing slash - append index.mf
{"https://example.com/path/", "https://example.com/path/index.mf"}, {"https://example.com/path/", testIndexMF},
{"https://example.com/", "https://example.com/index.mf"}, {"https://example.com/", "https://example.com/index.mf"},
// Directory without trailing slash - add slash and index.mf // Directory without trailing slash - add slash and index.mf
{"https://example.com/path", "https://example.com/path/index.mf"}, {"https://example.com/path", testIndexMF},
{"https://example.com", "https://example.com/index.mf"}, {"https://example.com", "https://example.com/index.mf"},
// With query strings // With query strings
{"https://example.com/path?foo=bar", "https://example.com/path/index.mf?foo=bar"}, {
"https://example.com/path?foo=bar",
"https://example.com/path/index.mf?foo=bar",
},
} }
for _, tt := range tests { for _, tt := range tests {
t.Run(tt.input, func(t *testing.T) { t.Run(tt.input, func(t *testing.T) {
t.Parallel()
result, err := resolveManifestURL(tt.input) result, err := resolveManifestURL(tt.input)
assert.NoError(t, err) require.NoError(t, err)
assert.Equal(t, tt.expected, result) assert.Equal(t, tt.expected, result)
}) })
} }
// The sole caller wraps this error as "invalid URL: %w", so
// resolveManifestURL must return url.Parse's error unadorned.
t.Run("invalid:control character", func(t *testing.T) {
t.Parallel()
_, err := resolveManifestURL("http://example.com/\x7f")
require.ErrorContains(t, err, urlParseControlCharErr)
assert.NotContains(t, err.Error(), "failed to parse URL")
})
} }
// scanToManifest scans sourceFs and returns the serialized manifest bytes.
func scanToManifest(t *testing.T, sourceFs afero.Fs) []byte {
t.Helper()
s := mfer.NewScannerWithOptions(&mfer.ScannerOptions{Fs: sourceFs})
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
return manifestBuf.Bytes()
}
// chdirTemp switches the working directory to a fresh temp dir for the
// duration of the test and returns its path.
func chdirTemp(t *testing.T) string {
t.Helper()
destDir := t.TempDir()
origDir, err := os.Getwd()
require.NoError(t, err)
require.NoError(t, os.Chdir(destDir))
t.Cleanup(func() { _ = os.Chdir(origDir) })
return destDir
}
// fetchTestHandler serves the manifest at /index.mf and the given files
// at their paths.
func fetchTestHandler(
manifestData []byte, testFiles map[string][]byte,
) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
path := r.URL.Path
if path == "/index.mf" {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(manifestData)
return
}
// Strip leading slash
if len(path) > 0 && path[0] == '/' {
path = path[1:]
}
content, exists := testFiles[path]
if !exists {
http.NotFound(w, r)
return
}
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(content)
}
}
//nolint:paralleltest // changes the process-global working directory
func TestFetchFromHTTP(t *testing.T) { func TestFetchFromHTTP(t *testing.T) {
// Create source filesystem with test files // Create source filesystem with test files
sourceFs := afero.NewMemMapFs() sourceFs := afero.NewMemMapFs()
@@ -134,51 +234,14 @@ func TestFetchFromHTTP(t *testing.T) {
} }
// Generate manifest using scanner // Generate manifest using scanner
opts := &mfer.ScannerOptions{ manifestData := scanToManifest(t, sourceFs)
Fs: sourceFs,
}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
manifestData := manifestBuf.Bytes()
// Create HTTP server that serves the source filesystem // Create HTTP server that serves the source filesystem
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { server := httptest.NewServer(fetchTestHandler(manifestData, testFiles))
path := r.URL.Path
if path == "/index.mf" {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(manifestData)
return
}
// Strip leading slash
if len(path) > 0 && path[0] == '/' {
path = path[1:]
}
content, exists := testFiles[path]
if !exists {
http.NotFound(w, r)
return
}
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(content)
}))
defer server.Close() defer server.Close()
// Create destination directory // Change to a fresh destination directory for the test
destDir, err := os.MkdirTemp("", "mfer-fetch-test-*") destDir := chdirTemp(t)
require.NoError(t, err)
defer func() { _ = os.RemoveAll(destDir) }()
// Change to dest directory for the test
origDir, err := os.Getwd()
require.NoError(t, err)
require.NoError(t, os.Chdir(destDir))
defer func() { _ = os.Chdir(origDir) }()
// Parse the manifest to get file entries // Parse the manifest to get file entries
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestData)) manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestData))
@@ -189,132 +252,125 @@ func TestFetchFromHTTP(t *testing.T) {
// Download each file using downloadFile // Download each file using downloadFile
progress := make(chan DownloadProgress, 10) progress := make(chan DownloadProgress, 10)
go func() { go func() {
for range progress { for p := range progress {
// Drain progress channel _ = p // drain progress channel
} }
}() }()
baseURL := server.URL + "/" baseURL := server.URL + "/"
for _, f := range files { for _, f := range files {
localPath, err := sanitizePath(f.Path) localPath, err := sanitizePath(f.GetPath())
require.NoError(t, err) require.NoError(t, err)
fileURL := baseURL + f.Path fileURL := baseURL + f.GetPath()
err = downloadFile(fileURL, localPath, f, progress) err = downloadFile(context.Background(), fileURL, localPath, f, progress)
require.NoError(t, err, "failed to download %s", f.Path) require.NoError(t, err, "failed to download %s", f.GetPath())
} }
close(progress) close(progress)
// Verify downloaded files match originals // Verify downloaded files match originals
for path, expectedContent := range testFiles { for path, expectedContent := range testFiles {
downloadedPath := filepath.Join(destDir, path) downloadedPath := filepath.Join(destDir, path)
//nolint:gosec // test-controlled path
downloadedContent, err := os.ReadFile(downloadedPath) downloadedContent, err := os.ReadFile(downloadedPath)
require.NoError(t, err, "failed to read downloaded file %s", path) require.NoError(t, err, "failed to read downloaded file %s", path)
assert.Equal(t, expectedContent, downloadedContent, "content mismatch for %s", path) assert.Equal(t, expectedContent, downloadedContent,
"content mismatch for %s", path)
} }
} }
//nolint:paralleltest // changes the process-global working directory
func TestFetchHashMismatch(t *testing.T) { func TestFetchHashMismatch(t *testing.T) {
// Create source filesystem with a test file // Create source filesystem with a test file
sourceFs := afero.NewMemMapFs() sourceFs := afero.NewMemMapFs()
originalContent := []byte("Original content") originalContent := []byte("Original content")
require.NoError(t, afero.WriteFile(sourceFs, "/file.txt", originalContent, 0o644)) require.NoError(t, afero.WriteFile(sourceFs, "/file.txt", originalContent, 0o644))
// Generate manifest // Generate and parse manifest
opts := &mfer.ScannerOptions{Fs: sourceFs} manifestData := scanToManifest(t, sourceFs)
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
var manifestBuf bytes.Buffer manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestData))
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
// Parse manifest
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestBuf.Bytes()))
require.NoError(t, err) require.NoError(t, err)
files := manifest.Files() files := manifest.Files()
require.Len(t, files, 1) require.Len(t, files, 1)
// Create server that serves DIFFERENT content (to trigger hash mismatch) // Create server that serves DIFFERENT content (to trigger hash mismatch)
tamperedContent := []byte("Tampered content!") tamperedContent := []byte("Tampered content!")
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream") server := httptest.NewServer(
_, _ = w.Write(tamperedContent) http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
})) w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(tamperedContent)
}))
defer server.Close() defer server.Close()
// Create temp directory // Work in a fresh temp directory
destDir, err := os.MkdirTemp("", "mfer-fetch-hash-test-*") chdirTemp(t)
require.NoError(t, err)
defer func() { _ = os.RemoveAll(destDir) }()
origDir, err := os.Getwd()
require.NoError(t, err)
require.NoError(t, os.Chdir(destDir))
defer func() { _ = os.Chdir(origDir) }()
// Try to download - should fail with hash mismatch // Try to download - should fail with hash mismatch
err = downloadFile(server.URL+"/file.txt", "file.txt", files[0], nil) err = downloadFile(context.Background(),
assert.Error(t, err) server.URL+"/file.txt", testFileTxt, files[0], nil)
require.Error(t, err)
assert.Contains(t, err.Error(), "mismatch") assert.Contains(t, err.Error(), "mismatch")
// Verify temp file was cleaned up // Verify temp file was cleaned up
_, err = os.Stat(".file.txt.tmp") _, err = os.Stat(".file.txt.tmp")
assert.True(t, os.IsNotExist(err), "temp file should be cleaned up on hash mismatch") assert.True(t, os.IsNotExist(err),
"temp file should be cleaned up on hash mismatch")
// Verify final file was not created // Verify final file was not created
_, err = os.Stat("file.txt") _, err = os.Stat(testFileTxt)
assert.True(t, os.IsNotExist(err), "final file should not exist on hash mismatch") assert.True(t, os.IsNotExist(err),
"final file should not exist on hash mismatch")
} }
//nolint:paralleltest // changes the process-global working directory
func TestFetchSizeMismatch(t *testing.T) { func TestFetchSizeMismatch(t *testing.T) {
// Create source filesystem with a test file // Create source filesystem with a test file
sourceFs := afero.NewMemMapFs() sourceFs := afero.NewMemMapFs()
originalContent := []byte("Original content with specific size") originalContent := []byte("Original content with specific size")
require.NoError(t, afero.WriteFile(sourceFs, "/file.txt", originalContent, 0o644)) require.NoError(t, afero.WriteFile(sourceFs, "/file.txt", originalContent, 0o644))
// Generate manifest // Generate and parse manifest
opts := &mfer.ScannerOptions{Fs: sourceFs} manifestData := scanToManifest(t, sourceFs)
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
var manifestBuf bytes.Buffer manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestData))
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
// Parse manifest
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestBuf.Bytes()))
require.NoError(t, err) require.NoError(t, err)
files := manifest.Files() files := manifest.Files()
require.Len(t, files, 1) require.Len(t, files, 1)
// Create server that serves content with wrong size // Create server that serves content with wrong size
wrongSizeContent := []byte("Short") wrongSizeContent := []byte("Short")
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream") server := httptest.NewServer(
_, _ = w.Write(wrongSizeContent) http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
})) w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(wrongSizeContent)
}))
defer server.Close() defer server.Close()
// Create temp directory // Work in a fresh temp directory
destDir, err := os.MkdirTemp("", "mfer-fetch-size-test-*") chdirTemp(t)
require.NoError(t, err)
defer func() { _ = os.RemoveAll(destDir) }()
origDir, err := os.Getwd()
require.NoError(t, err)
require.NoError(t, os.Chdir(destDir))
defer func() { _ = os.Chdir(origDir) }()
// Try to download - should fail with size mismatch // Try to download - should fail with size mismatch
err = downloadFile(server.URL+"/file.txt", "file.txt", files[0], nil) err = downloadFile(context.Background(),
assert.Error(t, err) server.URL+"/file.txt", testFileTxt, files[0], nil)
require.Error(t, err)
assert.Contains(t, err.Error(), "size mismatch") assert.Contains(t, err.Error(), "size mismatch")
// Verify temp file was cleaned up // Verify temp file was cleaned up
_, err = os.Stat(".file.txt.tmp") _, err = os.Stat(".file.txt.tmp")
assert.True(t, os.IsNotExist(err), "temp file should be cleaned up on size mismatch") assert.True(t, os.IsNotExist(err),
"temp file should be cleaned up on size mismatch")
} }
//nolint:paralleltest // changes the process-global working directory
func TestFetchProgress(t *testing.T) { func TestFetchProgress(t *testing.T) {
// Create source filesystem with a larger test file // Create source filesystem with a larger test file
sourceFs := afero.NewMemMapFs() sourceFs := afero.NewMemMapFs()
@@ -322,53 +378,47 @@ func TestFetchProgress(t *testing.T) {
content := bytes.Repeat([]byte("x"), 100*1024) // 100KB content := bytes.Repeat([]byte("x"), 100*1024) // 100KB
require.NoError(t, afero.WriteFile(sourceFs, "/large.txt", content, 0o644)) require.NoError(t, afero.WriteFile(sourceFs, "/large.txt", content, 0o644))
// Generate manifest // Generate and parse manifest
opts := &mfer.ScannerOptions{Fs: sourceFs} manifestData := scanToManifest(t, sourceFs)
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
var manifestBuf bytes.Buffer manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestData))
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
// Parse manifest
manifest, err := mfer.NewManifestFromReader(bytes.NewReader(manifestBuf.Bytes()))
require.NoError(t, err) require.NoError(t, err)
files := manifest.Files() files := manifest.Files()
require.Len(t, files, 1) require.Len(t, files, 1)
// Create server that serves the content // Create server that serves the content
server := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) { server := httptest.NewServer(
w.Header().Set("Content-Type", "application/octet-stream") http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Length", "102400") w.Header().Set("Content-Type", "application/octet-stream")
// Write in chunks to allow progress reporting w.Header().Set("Content-Length", "102400")
reader := bytes.NewReader(content) // Write in chunks to allow progress reporting
_, _ = io.Copy(w, reader) reader := bytes.NewReader(content)
})) _, _ = io.Copy(w, reader)
}))
defer server.Close() defer server.Close()
// Create temp directory // Work in a fresh temp directory
destDir, err := os.MkdirTemp("", "mfer-fetch-progress-test-*") chdirTemp(t)
require.NoError(t, err)
defer func() { _ = os.RemoveAll(destDir) }()
origDir, err := os.Getwd()
require.NoError(t, err)
require.NoError(t, os.Chdir(destDir))
defer func() { _ = os.Chdir(origDir) }()
// Set up progress channel and collect updates // Set up progress channel and collect updates
progress := make(chan DownloadProgress, 100) progress := make(chan DownloadProgress, 100)
var progressUpdates []DownloadProgress var progressUpdates []DownloadProgress
done := make(chan struct{}) done := make(chan struct{})
go func() { go func() {
for p := range progress { for p := range progress {
progressUpdates = append(progressUpdates, p) progressUpdates = append(progressUpdates, p)
} }
close(done) close(done)
}() }()
// Download // Download
err = downloadFile(server.URL+"/large.txt", "large.txt", files[0], progress) err = downloadFile(context.Background(),
server.URL+"/large.txt", "large.txt", files[0], progress)
close(progress) close(progress)
<-done <-done
@@ -380,7 +430,8 @@ func TestFetchProgress(t *testing.T) {
// Verify final progress shows complete // Verify final progress shows complete
if len(progressUpdates) > 0 { if len(progressUpdates) > 0 {
last := progressUpdates[len(progressUpdates)-1] last := progressUpdates[len(progressUpdates)-1]
assert.Equal(t, int64(len(content)), last.BytesRead, "final progress should show all bytes read") assert.Equal(t, int64(len(content)), last.BytesRead,
"final progress should show all bytes read")
assert.Equal(t, "large.txt", last.Path) assert.Equal(t, "large.txt", last.Path)
} }

View File

@@ -2,6 +2,7 @@ package cli
import ( import (
"crypto/sha256" "crypto/sha256"
"errors"
"fmt" "fmt"
"io" "io"
"io/fs" "io/fs"
@@ -16,6 +17,19 @@ import (
"sneak.berlin/go/mfer/mfer" "sneak.berlin/go/mfer/mfer"
) )
const (
// hashBufSize is the read buffer size used when hashing files.
hashBufSize = 64 * 1024
// scanProgressInterval is how many scanned files pass between
// progress updates.
scanProgressInterval = 100
)
// errEntryMissingMtime indicates a manifest entry that carries no
// modification time where one is required to carry it forward unchanged.
var errEntryMissingMtime = errors.New("manifest entry has no mtime")
// FreshenStatus contains progress information for the freshen operation. // FreshenStatus contains progress information for the freshen operation.
type FreshenStatus struct { type FreshenStatus struct {
Phase string // "scan" or "hash" Phase string // "scan" or "hash"
@@ -36,196 +50,296 @@ type freshenEntry struct {
existing *mfer.MFFilePath // existing manifest entry if unchanged existing *mfer.MFFilePath // existing manifest entry if unchanged
} }
func (mfa *CLIApp) freshenManifestOperation(ctx *cli.Context) error { // freshenScanner walks the filesystem and compares it against the
log.Debug("freshenManifestOperation()") // entries of an existing manifest.
type freshenScanner struct {
fs afero.Fs
absBase string
manifestBase string
includeDotfiles bool
followSymlinks bool
showProgress bool
existingByPath map[string]*mfer.MFFilePath
basePath := ctx.String("base") entries []*freshenEntry
showProgress := ctx.Bool("progress") scanCount int64
includeDotfiles := ctx.Bool("IncludeDotfiles") changed int64
followSymlinks := ctx.Bool("FollowSymLinks") added int64
unchanged int64
}
// Find manifest file // resolveSymlink resolves a symlink to its target's FileInfo. The
var manifestPath string // second return value is false when the entry should be skipped.
var err error func (s *freshenScanner) resolveSymlink(path string) (fs.FileInfo, bool) {
if !s.followSymlinks {
return nil, false
}
if ctx.Args().Len() > 0 { realPath, err := filepath.EvalSymlinks(path)
arg := ctx.Args().Get(0) if err != nil {
info, statErr := mfa.Fs.Stat(arg) return nil, false // Skip broken symlinks
if statErr == nil && info.IsDir() { }
manifestPath, err = findManifest(mfa.Fs, arg)
if err != nil { realInfo, err := s.fs.Stat(realPath)
return err if err != nil || realInfo.IsDir() {
} return nil, false
} else { }
manifestPath = arg
} return realInfo, true
}
// recordEntry classifies a scanned file as changed, unchanged, or added
// relative to the existing manifest.
func (s *freshenScanner) recordEntry(relPath string, info fs.FileInfo) {
existing, inManifest := s.existingByPath[relPath]
if !inManifest {
s.added++
log.Verbosef("A %s", relPath)
s.entries = append(s.entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
return
}
// Check if changed (size or mtime). An entry with no recorded mtime
// cannot be compared, so it counts as changed and gets re-hashed;
// silently treating the absent mtime as the Unix epoch would classify
// every such entry as changed without saying why.
existingMtime, haveMtime := entryMtime(existing)
if !haveMtime {
log.Debugf("%s: manifest entry has no mtime, treating as changed",
relPath)
}
if !haveMtime || existing.GetSize() != info.Size() ||
!existingMtime.Equal(info.ModTime()) {
s.changed++
log.Verbosef("M %s", relPath)
s.entries = append(s.entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
} else { } else {
manifestPath, err = findManifest(mfa.Fs, ".") s.unchanged++
if err != nil {
return err s.entries = append(s.entries, &freshenEntry{
} path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: false,
existing: existing,
})
}
// Mark as seen
delete(s.existingByPath, relPath)
}
// walk is the afero.Walk callback for the scan phase.
func (s *freshenScanner) walk(path string, info fs.FileInfo, walkErr error) error {
if walkErr != nil {
return walkErr
} }
log.Infof("loading manifest from %s", manifestPath) // Get relative path
relPath, err := filepath.Rel(s.absBase, path)
// Load existing manifest
manifest, err := mfer.NewManifestFromFile(mfa.Fs, manifestPath)
if err != nil { if err != nil {
return fmt.Errorf("failed to load manifest: %w", err) return fmt.Errorf(
"freshen: failed to compute relative path for %s: %w", path, err)
} }
existingFiles := manifest.Files() // Skip the manifest file itself
log.Infof("manifest contains %d files", len(existingFiles)) if relPath == s.manifestBase || relPath == "."+s.manifestBase {
return nil
// Build map of existing entries by path
existingByPath := make(map[string]*mfer.MFFilePath, len(existingFiles))
for _, f := range existingFiles {
existingByPath[f.Path] = f
} }
// Phase 1: Scan filesystem // Handle dotfiles
log.Infof("scanning filesystem...") if !s.includeDotfiles && mfer.IsHiddenPath(filepath.ToSlash(relPath)) {
startScan := time.Now()
var entries []*freshenEntry
var scanCount int64
var removed, changed, added, unchanged int64
absBase, err := filepath.Abs(basePath)
if err != nil {
return err
}
err = afero.Walk(mfa.Fs, absBase, func(path string, info fs.FileInfo, walkErr error) error {
if walkErr != nil {
return walkErr
}
// Get relative path
relPath, err := filepath.Rel(absBase, path)
if err != nil {
return err
}
// Skip the manifest file itself
if relPath == filepath.Base(manifestPath) || relPath == "."+filepath.Base(manifestPath) {
return nil
}
// Handle dotfiles
if !includeDotfiles && mfer.IsHiddenPath(filepath.ToSlash(relPath)) {
if info.IsDir() {
return filepath.SkipDir
}
return nil
}
// Skip directories
if info.IsDir() { if info.IsDir() {
return nil return filepath.SkipDir
}
// Handle symlinks
if info.Mode()&fs.ModeSymlink != 0 {
if !followSymlinks {
return nil
}
realPath, err := filepath.EvalSymlinks(path)
if err != nil {
return nil // Skip broken symlinks
}
realInfo, err := mfa.Fs.Stat(realPath)
if err != nil || realInfo.IsDir() {
return nil
}
info = realInfo
}
scanCount++
// Check against existing manifest
existing, inManifest := existingByPath[relPath]
if inManifest {
// Check if changed (size or mtime)
existingMtime := time.Unix(existing.Mtime.Seconds, int64(existing.Mtime.Nanos))
if existing.Size != info.Size() || !existingMtime.Equal(info.ModTime()) {
changed++
log.Verbosef("M %s", relPath)
entries = append(entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
} else {
unchanged++
entries = append(entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: false,
existing: existing,
})
}
// Mark as seen
delete(existingByPath, relPath)
} else {
added++
log.Verbosef("A %s", relPath)
entries = append(entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
}
// Report scan progress
if showProgress && scanCount%100 == 0 {
log.Progressf("Scanning: %d files found", scanCount)
} }
return nil return nil
})
if showProgress {
log.ProgressDone()
} }
if err != nil { // Skip directories
return fmt.Errorf("failed to scan filesystem: %w", err) if info.IsDir() {
return nil
} }
// Remaining entries in existingByPath are removed files // Handle symlinks
removed = int64(len(existingByPath)) if info.Mode()&fs.ModeSymlink != 0 {
for path := range existingByPath { realInfo, keep := s.resolveSymlink(path)
log.Verbosef("D %s", path) if !keep {
return nil
}
info = realInfo
} }
scanDuration := time.Since(startScan) s.scanCount++
log.Infof("scan complete in %s: %d unchanged, %d changed, %d added, %d removed",
scanDuration.Round(time.Millisecond), unchanged, changed, added, removed)
// Calculate total bytes to hash // Check against existing manifest
var totalHashBytes int64 s.recordEntry(relPath, info)
var filesToHash int64
for _, e := range entries { // Report scan progress
if e.needsHash { if s.showProgress && s.scanCount%scanProgressInterval == 0 {
totalHashBytes += e.size log.Progressf("Scanning: %d files found", s.scanCount)
filesToHash++ }
return nil
}
// resolveFreshenManifestPath determines the manifest path from the CLI
// arguments, searching directories for a manifest where needed.
func (mfa *CLIApp) resolveFreshenManifestPath(ctx *cli.Context) (string, error) {
if ctx.Args().Len() == 0 {
return findManifest(mfa.Fs, ".")
}
arg := ctx.Args().Get(0)
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
return findManifest(mfa.Fs, arg)
}
return arg, nil
}
// freshenHasher hashes changed and added files and feeds all entries to
// a manifest builder.
type freshenHasher struct {
fs afero.Fs
absBase string
showProgress bool
totalHashBytes int64
filesToHash int64
startHash time.Time
builder *mfer.Builder
hashedFiles int64
hashedBytes int64
}
// reportProgress renders hashing progress for the current byte count.
func (h *freshenHasher) reportProgress(n int64) {
if !h.showProgress {
return
}
currentBytes := h.hashedBytes + n
elapsed := time.Since(h.startHash)
var (
rate float64
eta time.Duration
)
if elapsed > 0 && currentBytes > 0 {
rate = float64(currentBytes) / elapsed.Seconds()
remaining := h.totalHashBytes - currentBytes
if rate > 0 {
eta = time.Duration(float64(remaining)/rate) * time.Second
} }
} }
// Phase 2: Hash changed and new files if eta > 0 {
if filesToHash > 0 { log.Progressf("Hashing: %d/%d files, %s/s, ETA %s",
log.Infof("hashing %d files (%s)...", filesToHash, humanize.IBytes(uint64(totalHashBytes))) h.hashedFiles, h.filesToHash, humanize.IBytes(safeRateUint64(rate)),
eta.Round(time.Second))
} else {
log.Progressf("Hashing: %d/%d files, %s/s",
h.hashedFiles, h.filesToHash, humanize.IBytes(safeRateUint64(rate)))
}
}
// processEntry hashes the entry if needed and adds it to the builder.
func (h *freshenHasher) processEntry(e *freshenEntry) error {
if !e.needsHash {
// Use existing entry
err := addExistingToBuilder(h.builder, e.existing)
if err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
}
return nil
} }
startHash := time.Now() // Need to read and hash the file
var hashedFiles int64 absPath := filepath.Join(h.absBase, e.path)
var hashedBytes int64
f, err := h.fs.Open(absPath)
if err != nil {
return fmt.Errorf("failed to open %s: %w", e.path, err)
}
hash, bytesRead, err := hashFile(f, h.reportProgress)
_ = f.Close()
if err != nil {
return fmt.Errorf("failed to hash %s: %w", e.path, err)
}
h.hashedBytes += bytesRead
h.hashedFiles++
// Add to builder with computed hash
err = addFileToBuilder(h.builder, e.path, e.size, e.mtime, hash)
if err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
}
return nil
}
// writeFreshenedManifest writes the manifest atomically (write to a
// temp file, then rename over the target).
func writeFreshenedManifest(
afs afero.Fs, builder *mfer.Builder, manifestPath string,
) error {
tmpPath := manifestPath + ".tmp"
outFile, err := afs.Create(tmpPath)
if err != nil {
return fmt.Errorf("failed to create temp file: %w", err)
}
err = builder.Build(outFile)
_ = outFile.Close()
if err != nil {
_ = afs.Remove(tmpPath)
return fmt.Errorf("failed to write manifest: %w", err)
}
// Rename temp to final
err = afs.Rename(tmpPath, manifestPath)
if err != nil {
_ = afs.Remove(tmpPath)
return fmt.Errorf("failed to rename manifest: %w", err)
}
return nil
}
// newFreshenBuilder constructs the manifest builder configured from CLI
// flags.
func newFreshenBuilder(ctx *cli.Context) *mfer.Builder {
builder := mfer.NewBuilder() builder := mfer.NewBuilder()
if ctx.Bool("include-timestamps") {
builder.SetIncludeTimestamps(true)
}
// Set up signing options if sign-key is provided // Set up signing options if sign-key is provided
if signKey := ctx.String("sign-key"); signKey != "" { if signKey := ctx.String("sign-key"); signKey != "" {
@@ -235,6 +349,77 @@ func (mfa *CLIApp) freshenManifestOperation(ctx *cli.Context) error {
log.Infof("signing manifest with GPG key: %s", signKey) log.Infof("signing manifest with GPG key: %s", signKey)
} }
return builder
}
// freshenScan runs the scan phase against the loaded manifest entries
// and returns the populated scanner and the count of removed files.
func (mfa *CLIApp) freshenScan(
ctx *cli.Context, manifestPath, absBase string,
existingByPath map[string]*mfer.MFFilePath,
) (*freshenScanner, int64, error) {
log.Infof("scanning filesystem...")
startScan := time.Now()
showProgress := ctx.Bool("progress")
scanner := &freshenScanner{
fs: mfa.Fs,
absBase: absBase,
manifestBase: filepath.Base(manifestPath),
includeDotfiles: ctx.Bool("include-dotfiles"),
followSymlinks: ctx.Bool("follow-symlinks"),
showProgress: showProgress,
existingByPath: existingByPath,
}
err := afero.Walk(mfa.Fs, absBase, scanner.walk)
if showProgress {
log.ProgressDone()
}
if err != nil {
return nil, 0, fmt.Errorf("failed to scan filesystem: %w", err)
}
// Remaining entries in existingByPath are removed files
removed := int64(len(existingByPath))
for path := range existingByPath {
log.Verbosef("D %s", path)
}
scanDuration := time.Since(startScan)
log.Infof("scan complete in %s: %d unchanged, %d changed, %d added, %d removed",
scanDuration.Round(time.Millisecond), scanner.unchanged, scanner.changed,
scanner.added, removed)
return scanner, removed, nil
}
// hashTotals returns the total byte count and file count of entries
// that need hashing.
func hashTotals(entries []*freshenEntry) (int64, int64) {
var (
totalHashBytes int64
filesToHash int64
)
for _, e := range entries {
if e.needsHash {
totalHashBytes += e.size
filesToHash++
}
}
return totalHashBytes, filesToHash
}
// runFreshenHash processes every entry through the hasher, aborting if
// the context is canceled.
func runFreshenHash(
ctx *cli.Context, hasher *freshenHasher, entries []*freshenEntry,
) error {
for _, e := range entries { for _, e := range entries {
select { select {
case <-ctx.Done(): case <-ctx.Done():
@@ -242,122 +427,154 @@ func (mfa *CLIApp) freshenManifestOperation(ctx *cli.Context) error {
default: default:
} }
if e.needsHash { err := hasher.processEntry(e)
// Need to read and hash the file if err != nil {
absPath := filepath.Join(absBase, e.path) return err
f, err := mfa.Fs.Open(absPath)
if err != nil {
return fmt.Errorf("failed to open %s: %w", e.path, err)
}
hash, bytesRead, err := hashFile(f, e.size, func(n int64) {
if showProgress {
currentBytes := hashedBytes + n
elapsed := time.Since(startHash)
var rate float64
var eta time.Duration
if elapsed > 0 && currentBytes > 0 {
rate = float64(currentBytes) / elapsed.Seconds()
remaining := totalHashBytes - currentBytes
if rate > 0 {
eta = time.Duration(float64(remaining)/rate) * time.Second
}
}
if eta > 0 {
log.Progressf("Hashing: %d/%d files, %s/s, ETA %s",
hashedFiles, filesToHash, humanize.IBytes(uint64(rate)), eta.Round(time.Second))
} else {
log.Progressf("Hashing: %d/%d files, %s/s",
hashedFiles, filesToHash, humanize.IBytes(uint64(rate)))
}
}
})
_ = f.Close()
if err != nil {
return fmt.Errorf("failed to hash %s: %w", e.path, err)
}
hashedBytes += bytesRead
hashedFiles++
// Add to builder with computed hash
if err := addFileToBuilder(builder, e.path, e.size, e.mtime, hash); err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
}
} else {
// Use existing entry
if err := addExistingToBuilder(builder, e.existing); err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
}
} }
} }
return nil
}
// loadExistingEntries loads the manifest and indexes its file entries
// by path.
func (mfa *CLIApp) loadExistingEntries(
manifestPath string,
) (map[string]*mfer.MFFilePath, error) {
log.Infof("loading manifest from %s", manifestPath)
// Load existing manifest
manifest, err := mfer.NewManifestFromFile(mfa.Fs, manifestPath)
if err != nil {
return nil, fmt.Errorf("failed to load manifest: %w", err)
}
existingFiles := manifest.Files()
log.Infof("manifest contains %d files", len(existingFiles))
// Build map of existing entries by path
existingByPath := make(map[string]*mfer.MFFilePath, len(existingFiles))
for _, f := range existingFiles {
existingByPath[f.GetPath()] = f
}
return existingByPath, nil
}
func (mfa *CLIApp) freshenManifestOperation(ctx *cli.Context) error {
log.Debug("freshenManifestOperation()")
basePath := ctx.String("base")
showProgress := ctx.Bool("progress")
// Find manifest file
manifestPath, err := mfa.resolveFreshenManifestPath(ctx)
if err != nil {
return fmt.Errorf("freshen: %w", err)
}
existingByPath, err := mfa.loadExistingEntries(manifestPath)
if err != nil {
return err
}
absBase, err := filepath.Abs(basePath)
if err != nil {
return fmt.Errorf("freshen: invalid base path: %w", err)
}
// Phase 1: Scan filesystem
scanner, removed, err := mfa.freshenScan(ctx, manifestPath, absBase,
existingByPath)
if err != nil {
return err
}
// Calculate total bytes to hash
totalHashBytes, filesToHash := hashTotals(scanner.entries)
// Phase 2: Hash changed and new files
if filesToHash > 0 {
log.Infof("hashing %d files (%s)...", filesToHash,
humanize.IBytes(safeUint64(totalHashBytes)))
}
hasher := &freshenHasher{
fs: mfa.Fs,
absBase: absBase,
showProgress: showProgress,
totalHashBytes: totalHashBytes,
filesToHash: filesToHash,
startHash: time.Now(),
builder: newFreshenBuilder(ctx),
}
err = runFreshenHash(ctx, hasher, scanner.entries)
if err != nil {
return err
}
if showProgress && filesToHash > 0 { if showProgress && filesToHash > 0 {
log.ProgressDone() log.ProgressDone()
} }
// Print summary // Print summary
log.Infof("freshen complete: %d unchanged, %d changed, %d added, %d removed", log.Infof("freshen complete: %d unchanged, %d changed, %d added, %d removed",
unchanged, changed, added, removed) scanner.unchanged, scanner.changed, scanner.added, removed)
// Skip writing if nothing changed // Skip writing if nothing changed
if changed == 0 && added == 0 && removed == 0 { if scanner.changed == 0 && scanner.added == 0 && removed == 0 {
log.Infof("manifest unchanged, skipping write") log.Infof("manifest unchanged, skipping write")
return nil return nil
} }
// Write updated manifest atomically (write to temp, then rename) // Write updated manifest atomically (write to temp, then rename)
tmpPath := manifestPath + ".tmp" err = writeFreshenedManifest(mfa.Fs, hasher.builder, manifestPath)
outFile, err := mfa.Fs.Create(tmpPath)
if err != nil { if err != nil {
return fmt.Errorf("failed to create temp file: %w", err) return err
}
err = builder.Build(outFile)
_ = outFile.Close()
if err != nil {
_ = mfa.Fs.Remove(tmpPath)
return fmt.Errorf("failed to write manifest: %w", err)
}
// Rename temp to final
if err := mfa.Fs.Rename(tmpPath, manifestPath); err != nil {
_ = mfa.Fs.Remove(tmpPath)
return fmt.Errorf("failed to rename manifest: %w", err)
} }
totalDuration := time.Since(mfa.startupTime) totalDuration := time.Since(mfa.startupTime)
if hashedBytes > 0 { if hasher.hashedBytes > 0 {
hashDuration := time.Since(startHash) hashDuration := time.Since(hasher.startHash)
hashRate := float64(hashedBytes) / hashDuration.Seconds() hashRate := float64(hasher.hashedBytes) / hashDuration.Seconds()
log.Infof("hashed %s in %.1fs (%s/s)", log.Infof("hashed %s in %.1fs (%s/s)",
humanize.IBytes(uint64(hashedBytes)), totalDuration.Seconds(), humanize.IBytes(uint64(hashRate))) humanize.IBytes(safeUint64(hasher.hashedBytes)),
totalDuration.Seconds(), humanize.IBytes(safeRateUint64(hashRate)))
} }
log.Infof("wrote %d files to %s", len(entries), manifestPath)
log.Infof("wrote %d files to %s", len(scanner.entries), manifestPath)
return nil return nil
} }
// hashFile reads a file and computes its SHA256 multihash. // hashFile reads a file and computes its SHA256 multihash.
// Progress callback is called with bytes read so far. // Progress callback is called with bytes read so far.
func hashFile(r io.Reader, size int64, progress func(int64)) ([]byte, int64, error) { func hashFile(r io.Reader, progress func(int64)) ([]byte, int64, error) {
h := sha256.New() h := sha256.New()
buf := make([]byte, 64*1024) buf := make([]byte, hashBufSize)
var total int64 var total int64
for { for {
n, err := r.Read(buf) n, err := r.Read(buf)
if n > 0 { if n > 0 {
h.Write(buf[:n]) h.Write(buf[:n])
total += int64(n) total += int64(n)
if progress != nil { if progress != nil {
progress(total) progress(total)
} }
} }
if err == io.EOF { if err == io.EOF {
break break
} }
// Returned unwrapped: the caller renders this as
// "failed to hash <path>: <err>" and adding a second layer here
// would change that message.
if err != nil { if err != nil {
return nil, total, err return nil, total, err
} }
@@ -372,15 +589,29 @@ func hashFile(r io.Reader, size int64, progress func(int64)) ([]byte, int64, err
} }
// addFileToBuilder adds a new file entry to the builder // addFileToBuilder adds a new file entry to the builder
func addFileToBuilder(b *mfer.Builder, path string, size int64, mtime time.Time, hash []byte) error { func addFileToBuilder(
return b.AddFileWithHash(mfer.RelFilePath(path), mfer.FileSize(size), mfer.ModTime(mtime), hash) b *mfer.Builder, path string, size int64, mtime time.Time, hash []byte,
) error {
return b.AddFileWithHash(
mfer.RelFilePath(path), mfer.FileSize(size), mfer.ModTime(mtime), hash)
} }
// addExistingToBuilder adds an existing manifest entry to the builder // addExistingToBuilder adds an existing manifest entry to the builder.
//
// Entries reach this path only when recordEntry classified them as
// unchanged, which requires a recorded mtime, so an absent mtime here is
// an error rather than something to paper over with the Unix epoch.
func addExistingToBuilder(b *mfer.Builder, entry *mfer.MFFilePath) error { func addExistingToBuilder(b *mfer.Builder, entry *mfer.MFFilePath) error {
mtime := time.Unix(entry.Mtime.Seconds, int64(entry.Mtime.Nanos)) mtime, ok := entryMtime(entry)
if len(entry.Hashes) == 0 { if !ok {
return fmt.Errorf("%w: %s", errEntryMissingMtime, entry.GetPath())
}
if len(entry.GetHashes()) == 0 {
return nil return nil
} }
return b.AddFileWithHash(mfer.RelFilePath(entry.Path), mfer.FileSize(entry.Size), mfer.ModTime(mtime), entry.Hashes[0].MultiHash)
return b.AddFileWithHash(mfer.RelFilePath(entry.GetPath()),
mfer.FileSize(entry.GetSize()), mfer.ModTime(mtime),
entry.GetHashes()[0].GetMultiHash())
} }

View File

@@ -1,9 +1,12 @@
//nolint:testpackage // white-box tests exercise unexported internals
package cli package cli
import ( import (
"bytes" "bytes"
"context" "context"
"os"
"testing" "testing"
"time"
"github.com/spf13/afero" "github.com/spf13/afero"
"github.com/stretchr/testify/assert" "github.com/stretchr/testify/assert"
@@ -11,24 +14,48 @@ import (
"sneak.berlin/go/mfer/mfer" "sneak.berlin/go/mfer/mfer"
) )
func TestFreshenUnchanged(t *testing.T) { // stubFileInfo is a minimal fs.FileInfo for exercising recordEntry
// Create filesystem with test files // without touching a filesystem.
fs := afero.NewMemMapFs() type stubFileInfo struct {
size int64
mtime time.Time
}
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) func (s stubFileInfo) Name() string { return "stub" }
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644)) func (s stubFileInfo) Size() int64 { return s.size }
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644)) func (s stubFileInfo) Mode() os.FileMode { return 0 }
func (s stubFileInfo) ModTime() time.Time { return s.mtime }
func (s stubFileInfo) IsDir() bool { return false }
func (s stubFileInfo) Sys() any { return nil }
// setupFreshenDir populates /testdir with two files, scans it, and
// writes the resulting manifest to /testdir/.index.mf.
func setupFreshenDir(t *testing.T, fs afero.Fs) {
t.Helper()
require.NoError(t, fs.MkdirAll(testDir, 0o755))
writeTestFile(t, fs, testFile1, "content1")
writeTestFile(t, fs, "/testdir/file2.txt", "content2")
// Generate initial manifest // Generate initial manifest
opts := &mfer.ScannerOptions{Fs: fs} opts := &mfer.ScannerOptions{Fs: fs}
s := mfer.NewScannerWithOptions(opts) s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumeratePath("/testdir", nil)) require.NoError(t, s.EnumeratePath(testDir, nil))
var manifestBuf bytes.Buffer var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil)) require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
// Write manifest to filesystem // Write manifest to filesystem
require.NoError(t, afero.WriteFile(fs, "/testdir/.index.mf", manifestBuf.Bytes(), 0o644)) require.NoError(t,
afero.WriteFile(fs, "/testdir/.index.mf", manifestBuf.Bytes(), 0o644))
}
func TestFreshenUnchanged(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
setupFreshenDir(t, fs)
// Parse manifest to verify // Parse manifest to verify
manifest, err := mfer.NewManifestFromFile(fs, "/testdir/.index.mf") manifest, err := mfer.NewManifestFromFile(fs, "/testdir/.index.mf")
@@ -37,23 +64,10 @@ func TestFreshenUnchanged(t *testing.T) {
} }
func TestFreshenWithChanges(t *testing.T) { func TestFreshenWithChanges(t *testing.T) {
// Create filesystem with test files t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
setupFreshenDir(t, fs)
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644))
// Generate initial manifest
opts := &mfer.ScannerOptions{Fs: fs}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumeratePath("/testdir", nil))
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
// Write manifest to filesystem
require.NoError(t, afero.WriteFile(fs, "/testdir/.index.mf", manifestBuf.Bytes(), 0o644))
// Verify initial manifest has 2 files // Verify initial manifest has 2 files
manifest, err := mfer.NewManifestFromFile(fs, "/testdir/.index.mf") manifest, err := mfer.NewManifestFromFile(fs, "/testdir/.index.mf")
@@ -61,17 +75,17 @@ func TestFreshenWithChanges(t *testing.T) {
assert.Len(t, manifest.Files(), 2) assert.Len(t, manifest.Files(), 2)
// Add a new file // Add a new file
require.NoError(t, afero.WriteFile(fs, "/testdir/file3.txt", []byte("content3"), 0o644)) writeTestFile(t, fs, "/testdir/file3.txt", "content3")
// Modify file2 (change content and size) // Modify file2 (change content and size)
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("modified content2"), 0o644)) writeTestFile(t, fs, "/testdir/file2.txt", "modified content2")
// Remove file1 // Remove file1
require.NoError(t, fs.Remove("/testdir/file1.txt")) require.NoError(t, fs.Remove(testFile1))
// Note: The freshen operation would need to be run here // Note: The freshen operation would need to be run here
// For now, we just verify the test setup is correct // For now, we just verify the test setup is correct
exists, _ := afero.Exists(fs, "/testdir/file1.txt") exists, _ := afero.Exists(fs, testFile1)
assert.False(t, exists) assert.False(t, exists)
exists, _ = afero.Exists(fs, "/testdir/file3.txt") exists, _ = afero.Exists(fs, "/testdir/file3.txt")
@@ -80,3 +94,104 @@ func TestFreshenWithChanges(t *testing.T) {
content, _ := afero.ReadFile(fs, "/testdir/file2.txt") content, _ := afero.ReadFile(fs, "/testdir/file2.txt")
assert.Equal(t, "modified content2", string(content)) assert.Equal(t, "modified content2", string(content))
} }
// TestFreshenRecordEntryMtimePresence pins the behavior of recordEntry
// with respect to MFFilePath.Mtime, which is a message pointer with
// proto3 field presence and may legitimately be absent.
//
// An absent mtime must never be read as time.Unix(0, 0): that value
// never equals a real modification time, so every entry would be
// classified as changed, re-hashed, and the manifest rewritten
// unconditionally - the exact inverse of what freshen is for, and
// silent. An entry with no mtime is therefore "changed" because it
// cannot be compared, not because it looks like it dates from 1970.
func TestFreshenRecordEntryMtimePresence(t *testing.T) {
t.Parallel()
const relPath = "file1.txt"
mtime := time.Unix(1_700_000_000, 0)
info := stubFileInfo{size: 8, mtime: mtime}
for _, tc := range []struct {
name string
entry *mfer.MFFilePath
needsHash bool
changed int64
unchanged int64
}{
{
name: "matching mtime and size is unchanged",
entry: &mfer.MFFilePath{
Path: relPath,
Size: 8,
Mtime: &mfer.Timestamp{Seconds: mtime.Unix()},
},
needsHash: false,
changed: 0,
unchanged: 1,
},
{
name: "absent mtime is changed, not epoch",
entry: &mfer.MFFilePath{
Path: relPath,
Size: 8,
Mtime: nil,
},
needsHash: true,
changed: 1,
unchanged: 0,
},
} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
s := &freshenScanner{
existingByPath: map[string]*mfer.MFFilePath{relPath: tc.entry},
}
s.recordEntry(relPath, info)
require.Len(t, s.entries, 1)
assert.Equal(t, tc.needsHash, s.entries[0].needsHash)
assert.Equal(t, tc.changed, s.changed)
assert.Equal(t, tc.unchanged, s.unchanged)
assert.Zero(t, s.added)
})
}
}
// TestFreshenAddExistingRejectsMissingMtime pins that an entry with no
// mtime is never carried forward into a rebuilt manifest with a
// fabricated epoch timestamp.
func TestFreshenAddExistingRejectsMissingMtime(t *testing.T) {
t.Parallel()
b := mfer.NewBuilder()
entry := &mfer.MFFilePath{
Path: "file1.txt",
Size: 8,
Mtime: nil,
Hashes: []*mfer.MFFileChecksum{
{MultiHash: []byte{0x12, 0x20}},
},
}
err := addExistingToBuilder(b, entry)
require.ErrorIs(t, err, errEntryMissingMtime)
assert.Contains(t, err.Error(), "file1.txt")
}
// TestEntryMtime pins the presence semantics the callers depend on.
func TestEntryMtime(t *testing.T) {
t.Parallel()
got, ok := entryMtime(&mfer.MFFilePath{Mtime: nil})
assert.False(t, ok)
assert.True(t, got.IsZero())
got, ok = entryMtime(&mfer.MFFilePath{
Mtime: &mfer.Timestamp{Seconds: 1_700_000_000, Nanos: 500},
})
assert.True(t, ok)
assert.Equal(t, time.Unix(1_700_000_000, 500), got)
}

View File

@@ -1,6 +1,7 @@
package cli package cli
import ( import (
"errors"
"fmt" "fmt"
"os" "os"
"os/signal" "os/signal"
@@ -16,13 +17,90 @@ import (
"sneak.berlin/go/mfer/mfer" "sneak.berlin/go/mfer/mfer"
) )
func (mfa *CLIApp) generateManifestOperation(ctx *cli.Context) error { var (
log.Debug("generateManifestOperation()") // errPathNotExist indicates an input path that does not exist.
errPathNotExist = errors.New("path does not exist")
// errOutputExists indicates the output file already exists and
// --force was not given. It is wrapped mid-sentence so that the
// rendered message stays exactly as mfer has always printed it.
errOutputExists = errors.New(
"already exists (use --force to overwrite)")
)
// reportEnumProgress renders enumeration progress until the channel
// closes.
func reportEnumProgress(progress <-chan mfer.EnumerateStatus, wg *sync.WaitGroup) {
defer wg.Done()
for status := range progress {
log.Progressf("Enumerating: %d files, %s",
status.FilesFound,
humanize.IBytes(safeUint64(int64(status.BytesFound))))
}
log.ProgressDone()
}
// reportScanProgress renders scan progress until the channel closes.
func reportScanProgress(progress <-chan mfer.ScanStatus, wg *sync.WaitGroup) {
defer wg.Done()
for status := range progress {
if status.ETA > 0 {
log.Progressf("Scanning: %d/%d files, %s/s, ETA %s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)),
status.ETA.Round(time.Second))
} else {
log.Progressf("Scanning: %d/%d files, %s/s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)))
}
}
log.ProgressDone()
}
// collectInputPaths validates the input path arguments and returns them
// as absolute paths.
func (mfa *CLIApp) collectInputPaths(args cli.Args) ([]string, error) {
paths := make([]string, 0, args.Len())
for i := range args.Len() {
inputPath := args.Get(i)
ap, err := filepath.Abs(inputPath)
if err != nil {
return nil, fmt.Errorf("generate: invalid path %q: %w", inputPath, err)
}
// Validate path exists before adding to list
if exists, _ := afero.Exists(mfa.Fs, ap); !exists {
return nil, fmt.Errorf("%w: %s", errPathNotExist, inputPath)
}
log.Debugf("enumerating path: %s", ap)
paths = append(paths, ap)
}
return paths, nil
}
// buildScannerOptions constructs scanner options from the CLI flags.
func (mfa *CLIApp) buildScannerOptions(ctx *cli.Context) *mfer.ScannerOptions {
opts := &mfer.ScannerOptions{ opts := &mfer.ScannerOptions{
IncludeDotfiles: ctx.Bool("IncludeDotfiles"), IncludeDotfiles: ctx.Bool("include-dotfiles"),
FollowSymLinks: ctx.Bool("FollowSymLinks"), FollowSymLinks: ctx.Bool("follow-symlinks"),
Fs: mfa.Fs, IncludeTimestamps: ctx.Bool("include-timestamps"),
Fs: mfa.Fs,
}
// Set seed for deterministic UUID if provided
if seed := ctx.String("seed"); seed != "" {
opts.Seed = seed
log.Infof("using deterministic seed for manifest UUID")
} }
// Set up signing options if sign-key is provided // Set up signing options if sign-key is provided
@@ -33,136 +111,167 @@ func (mfa *CLIApp) generateManifestOperation(ctx *cli.Context) error {
log.Infof("signing manifest with GPG key: %s", signKey) log.Infof("signing manifest with GPG key: %s", signKey)
} }
s := mfer.NewScannerWithOptions(opts) return opts
}
// Phase 1: Enumeration - collect paths and stat files
args := ctx.Args()
showProgress := ctx.Bool("progress")
// Set up enumeration progress reporting
var enumProgress chan mfer.EnumerateStatus
var enumWg sync.WaitGroup
if showProgress {
enumProgress = make(chan mfer.EnumerateStatus, 1)
enumWg.Add(1)
go func() {
defer enumWg.Done()
for status := range enumProgress {
log.Progressf("Enumerating: %d files, %s",
status.FilesFound,
humanize.IBytes(uint64(status.BytesFound)))
}
log.ProgressDone()
}()
}
// enumerateInputs runs the enumeration phase over the argument paths,
// or the current directory when no arguments are given.
func (mfa *CLIApp) enumerateInputs(
s *mfer.Scanner, args cli.Args, enumProgress chan mfer.EnumerateStatus,
) error {
if args.Len() == 0 { if args.Len() == 0 {
// Default to current directory // Default to current directory
if err := s.EnumeratePath(".", enumProgress); err != nil { err := s.EnumeratePath(".", enumProgress)
return err if err != nil {
} return fmt.Errorf(
} else { "generate: failed to enumerate current directory: %w", err)
// Collect and validate all paths first
paths := make([]string, 0, args.Len())
for i := 0; i < args.Len(); i++ {
inputPath := args.Get(i)
ap, err := filepath.Abs(inputPath)
if err != nil {
return err
}
// Validate path exists before adding to list
if exists, _ := afero.Exists(mfa.Fs, ap); !exists {
return fmt.Errorf("path does not exist: %s", inputPath)
}
log.Debugf("enumerating path: %s", ap)
paths = append(paths, ap)
}
if err := s.EnumeratePaths(enumProgress, paths...); err != nil {
return err
} }
return nil
} }
// Collect and validate all paths first
paths, err := mfa.collectInputPaths(args)
if err != nil {
return err
}
err = s.EnumeratePaths(enumProgress, paths...)
if err != nil {
return fmt.Errorf("generate: failed to enumerate paths: %w", err)
}
return nil
}
// cleanupOnSignal installs a handler that removes the temp output file
// and exits when the process is interrupted. It returns the signal
// channel so the caller can stop and close it when done.
func (mfa *CLIApp) cleanupOnSignal(outFile afero.File, tmpPath string) chan os.Signal {
sigChan := make(chan os.Signal, 1)
signal.Notify(sigChan, os.Interrupt, syscall.SIGTERM)
go func() {
sig, ok := <-sigChan
if !ok || sig == nil {
return // Channel closed normally, not a signal
}
_ = outFile.Close()
_ = mfa.Fs.Remove(tmpPath)
os.Exit(1)
}()
return sigChan
}
// runEnumeratePhase enumerates all input paths with optional progress
// reporting and logs the totals.
func (mfa *CLIApp) runEnumeratePhase(ctx *cli.Context, s *mfer.Scanner) error {
// Set up enumeration progress reporting
var (
enumProgress chan mfer.EnumerateStatus
enumWg sync.WaitGroup
)
if ctx.Bool("progress") {
enumProgress = make(chan mfer.EnumerateStatus, 1)
enumWg.Add(1)
go reportEnumProgress(enumProgress, &enumWg)
}
err := mfa.enumerateInputs(s, ctx.Args(), enumProgress)
if err != nil {
return err
}
enumWg.Wait() enumWg.Wait()
log.Infof("enumerated %d files, %s total", s.FileCount(), humanize.IBytes(uint64(s.TotalBytes()))) log.Infof("enumerated %d files, %s total", s.FileCount(),
humanize.IBytes(safeUint64(int64(s.TotalBytes()))))
return nil
}
func (mfa *CLIApp) generateManifestOperation(ctx *cli.Context) error {
log.Debug("generateManifestOperation()")
s := mfer.NewScannerWithOptions(mfa.buildScannerOptions(ctx))
// Phase 1: Enumeration - collect paths and stat files
err := mfa.runEnumeratePhase(ctx, s)
if err != nil {
return err
}
showProgress := ctx.Bool("progress")
// Check if output file exists // Check if output file exists
outputPath := ctx.String("output") outputPath := ctx.String("output")
if exists, _ := afero.Exists(mfa.Fs, outputPath); exists { if exists, _ := afero.Exists(mfa.Fs, outputPath); exists && !ctx.Bool("force") {
if !ctx.Bool("force") { return fmt.Errorf("output file %s %w", outputPath, errOutputExists)
return fmt.Errorf("output file %s already exists (use --force to overwrite)", outputPath)
}
} }
// Create temp file for atomic write // Create temp file for atomic write
tmpPath := outputPath + ".tmp" tmpPath := outputPath + ".tmp"
outFile, err := mfa.Fs.Create(tmpPath) outFile, err := mfa.Fs.Create(tmpPath)
if err != nil { if err != nil {
return fmt.Errorf("failed to create temp file: %w", err) return fmt.Errorf("failed to create temp file: %w", err)
} }
// Set up signal handler to clean up temp file on Ctrl-C // Set up signal handler to clean up temp file on Ctrl-C
sigChan := make(chan os.Signal, 1) sigChan := mfa.cleanupOnSignal(outFile, tmpPath)
signal.Notify(sigChan, os.Interrupt, syscall.SIGTERM)
go func() {
sig, ok := <-sigChan
if !ok || sig == nil {
return // Channel closed normally, not a signal
}
_ = outFile.Close()
_ = mfa.Fs.Remove(tmpPath)
os.Exit(1)
}()
// Clean up temp file on any error or interruption // Clean up temp file on any error or interruption
success := false success := false
defer func() { defer func() {
signal.Stop(sigChan) signal.Stop(sigChan)
close(sigChan) close(sigChan)
_ = outFile.Close() _ = outFile.Close()
if !success { if !success {
_ = mfa.Fs.Remove(tmpPath) _ = mfa.Fs.Remove(tmpPath)
} }
}() }()
// Phase 2: Scan - read file contents and generate manifest // Phase 2: Scan - read file contents and generate manifest
var scanProgress chan mfer.ScanStatus var (
var scanWg sync.WaitGroup scanProgress chan mfer.ScanStatus
scanWg sync.WaitGroup
)
if showProgress { if showProgress {
scanProgress = make(chan mfer.ScanStatus, 1) scanProgress = make(chan mfer.ScanStatus, 1)
scanWg.Add(1) scanWg.Add(1)
go func() {
defer scanWg.Done() go reportScanProgress(scanProgress, &scanWg)
for status := range scanProgress {
if status.ETA > 0 {
log.Progressf("Scanning: %d/%d files, %s/s, ETA %s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)),
status.ETA.Round(time.Second))
} else {
log.Progressf("Scanning: %d/%d files, %s/s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)))
}
}
log.ProgressDone()
}()
} }
err = s.ToManifest(ctx.Context, outFile, scanProgress) err = s.ToManifest(ctx.Context, outFile, scanProgress)
scanWg.Wait() scanWg.Wait()
if err != nil { if err != nil {
return fmt.Errorf("failed to generate manifest: %w", err) return fmt.Errorf("failed to generate manifest: %w", err)
} }
// Close file before rename to ensure all data is flushed // Close file before rename to ensure all data is flushed
if err := outFile.Close(); err != nil { err = outFile.Close()
if err != nil {
return fmt.Errorf("failed to close temp file: %w", err) return fmt.Errorf("failed to close temp file: %w", err)
} }
// Atomic rename // Atomic rename
if err := mfa.Fs.Rename(tmpPath, outputPath); err != nil { err = mfa.Fs.Rename(tmpPath, outputPath)
if err != nil {
return fmt.Errorf("failed to rename temp file: %w", err) return fmt.Errorf("failed to rename temp file: %w", err)
} }
@@ -170,7 +279,9 @@ func (mfa *CLIApp) generateManifestOperation(ctx *cli.Context) error {
elapsed := time.Since(mfa.startupTime).Seconds() elapsed := time.Since(mfa.startupTime).Seconds()
rate := float64(s.TotalBytes()) / elapsed rate := float64(s.TotalBytes()) / elapsed
log.Infof("wrote %d files (%s) to %s in %.1fs (%s/s)", s.FileCount(), humanize.IBytes(uint64(s.TotalBytes())), outputPath, elapsed, humanize.IBytes(uint64(rate))) log.Infof("wrote %d files (%s) to %s in %.1fs (%s/s)", s.FileCount(),
humanize.IBytes(safeUint64(int64(s.TotalBytes()))), outputPath, elapsed,
humanize.IBytes(safeRateUint64(rate)))
return nil return nil
} }

View File

@@ -16,32 +16,21 @@ func (mfa *CLIApp) listManifestOperation(ctx *cli.Context) error {
longFormat := ctx.Bool("long") longFormat := ctx.Bool("long")
print0 := ctx.Bool("print0") print0 := ctx.Bool("print0")
// Find manifest file pathOrURL, err := mfa.resolveManifestArg(ctx)
var manifestPath string if err != nil {
var err error return fmt.Errorf("list: %w", err)
if ctx.Args().Len() > 0 {
arg := ctx.Args().Get(0)
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
manifestPath, err = findManifest(mfa.Fs, arg)
if err != nil {
return err
}
} else {
manifestPath = arg
}
} else {
manifestPath, err = findManifest(mfa.Fs, ".")
if err != nil {
return err
}
} }
// Load manifest rc, err := mfa.openManifestReader(pathOrURL)
manifest, err := mfer.NewManifestFromFile(mfa.Fs, manifestPath)
if err != nil { if err != nil {
return fmt.Errorf("failed to load manifest: %w", err) return fmt.Errorf("list: %w", err)
}
defer func() { _ = rc.Close() }()
manifest, err := mfer.NewManifestFromReader(rc)
if err != nil {
return fmt.Errorf("list: failed to parse manifest: %w", err)
} }
files := manifest.Files() files := manifest.Files()
@@ -54,10 +43,17 @@ func (mfa *CLIApp) listManifestOperation(ctx *cli.Context) error {
for _, f := range files { for _, f := range files {
if longFormat { if longFormat {
mtime := time.Unix(f.Mtime.Seconds, int64(f.Mtime.Nanos)) // An entry may legitimately carry no mtime; render that as
_, _ = fmt.Fprintf(mfa.Stdout, "%d\t%s\t%s%s", f.Size, mtime.Format(time.RFC3339), f.Path, lineEnd) // mtimeAbsent rather than as the Unix epoch.
mtimeStr := mtimeAbsent
if mtime, ok := entryMtime(f); ok {
mtimeStr = mtime.Format(time.RFC3339)
}
_, _ = fmt.Fprintf(mfa.Stdout, "%d\t%s\t%s%s",
f.GetSize(), mtimeStr, f.GetPath(), lineEnd)
} else { } else {
_, _ = fmt.Fprintf(mfa.Stdout, "%s%s", f.Path, lineEnd) _, _ = fmt.Fprintf(mfa.Stdout, "%s%s", f.GetPath(), lineEnd)
} }
} }

View File

@@ -0,0 +1,86 @@
package cli
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"strings"
"time"
"github.com/urfave/cli/v2"
)
// manifestFetchTimeout bounds HTTP requests made to fetch a manifest.
const manifestFetchTimeout = 30 * time.Second
// errHTTPStatus indicates an HTTP response with a non-OK status code.
//
// Its text is the literal "HTTP" prefix of the rendered "HTTP <code>"
// message that mfer has always printed, so that wrapping it does not
// change any user-visible output. Match it with errors.Is; do not read
// its message.
var errHTTPStatus = errors.New("HTTP")
// isHTTPURL returns true if the string starts with http:// or https://.
func isHTTPURL(s string) bool {
return strings.HasPrefix(s, "http://") || strings.HasPrefix(s, "https://")
}
// openManifestReader opens a manifest from a path or URL and returns a ReadCloser.
// The caller must close the returned reader.
func (mfa *CLIApp) openManifestReader(pathOrURL string) (io.ReadCloser, error) {
if isHTTPURL(pathOrURL) {
client := &http.Client{Timeout: manifestFetchTimeout}
req, err := http.NewRequestWithContext(
context.Background(), http.MethodGet, pathOrURL, nil,
)
if err != nil {
return nil, fmt.Errorf("failed to fetch %s: %w", pathOrURL, err)
}
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("failed to fetch %s: %w", pathOrURL, err)
}
if resp.StatusCode != http.StatusOK {
_ = resp.Body.Close()
return nil, fmt.Errorf("failed to fetch %s: %w %d",
pathOrURL, errHTTPStatus, resp.StatusCode)
}
return resp.Body, nil
}
f, err := mfa.Fs.Open(pathOrURL)
if err != nil {
return nil, err
}
return f, nil
}
// resolveManifestArg resolves the manifest path from CLI arguments.
// HTTP(S) URLs are returned as-is. Directories are searched for index.mf/.index.mf.
// If no argument is given, the current directory is searched.
func (mfa *CLIApp) resolveManifestArg(ctx *cli.Context) (string, error) {
if ctx.Args().Len() > 0 {
arg := ctx.Args().Get(0)
if isHTTPURL(arg) {
return arg, nil
}
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
return findManifest(mfa.Fs, arg)
}
return arg, nil
}
return findManifest(mfa.Fs, ".")
}

View File

@@ -1,6 +1,7 @@
package cli package cli
import ( import (
"errors"
"fmt" "fmt"
"io" "io"
"os" "os"
@@ -12,8 +13,24 @@ import (
"sneak.berlin/go/mfer/mfer" "sneak.berlin/go/mfer/mfer"
) )
// Command and flag names shared across command definitions and tests.
const (
cmdGenerate = "generate"
cmdCheck = "check"
cmdExport = "export"
flagProgress = "progress"
manifestArgsUsage = "[manifest file]"
)
// errUnknownCommand indicates an unrecognized command argument.
var errUnknownCommand = errors.New("unknown command")
// CLIApp is the main CLI application container. It holds configuration, // CLIApp is the main CLI application container. It holds configuration,
// I/O streams, and filesystem abstraction to enable testing and flexibility. // I/O streams, and filesystem abstraction to enable testing and flexibility.
//
//nolint:revive // established name used throughout the codebase and tests
type CLIApp struct { type CLIApp struct {
appname string appname string
version string version string
@@ -41,29 +58,34 @@ const banner = `
\ \:\ \ \:\ \ \::/ \ \:\ \ \:\ \ \:\ \ \::/ \ \:\
\__\/ \__\/ \__\/ \__\/` \__\/ \__\/ \__\/ \__\/`
func (mfa *CLIApp) printBanner() {
if log.GetLevel() <= log.InfoLevel {
_, _ = fmt.Fprintln(mfa.Stdout, banner)
_, _ = fmt.Fprintf(mfa.Stdout, " mfer by @sneak: v%s released %s\n", mfer.Version, mfer.ReleaseDate)
_, _ = fmt.Fprintln(mfa.Stdout, " https://sneak.berlin/go/mfer")
}
}
// VersionString returns the version and git revision formatted for display. // VersionString returns the version and git revision formatted for display.
func (mfa *CLIApp) VersionString() string { func (mfa *CLIApp) VersionString() string {
if mfa.gitrev != "" { if mfa.gitrev != "" {
return fmt.Sprintf("%s (%s)", mfer.Version, mfa.gitrev) return fmt.Sprintf("%s (%s)", mfer.Version, mfa.gitrev)
} }
return mfer.Version return mfer.Version
} }
func (mfa *CLIApp) printBanner() {
if log.GetLevel() <= log.InfoLevel {
_, _ = fmt.Fprintln(mfa.Stdout, banner)
_, _ = fmt.Fprintf(mfa.Stdout,
" mfer by @sneak: v%s released %s\n",
mfer.Version, mfer.ReleaseDate)
_, _ = fmt.Fprintln(mfa.Stdout, " https://sneak.berlin/go/mfer")
}
}
func (mfa *CLIApp) setVerbosity(c *cli.Context) { func (mfa *CLIApp) setVerbosity(c *cli.Context) {
_, present := os.LookupEnv("MFER_DEBUG") _, present := os.LookupEnv("MFER_DEBUG")
if present {
switch {
case present:
log.EnableDebugLogging() log.EnableDebugLogging()
} else if c.Bool("quiet") { case c.Bool("quiet"):
log.SetLevel(log.ErrorLevel) log.SetLevel(log.ErrorLevel)
} else { default:
log.SetLevelFromVerbosity(c.Count("verbose")) log.SetLevelFromVerbosity(c.Count("verbose"))
} }
} }
@@ -85,10 +107,215 @@ func commonFlags() []cli.Flag {
} }
} }
func (mfa *CLIApp) generateCommand() *cli.Command {
return &cli.Command{
Name: cmdGenerate,
Aliases: []string{"gen"},
Usage: "Generate manifest file",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.generateManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.BoolFlag{
Name: "follow-symlinks",
Aliases: []string{"L"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "include-dotfiles",
Aliases: []string{"IncludeDotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.StringFlag{
Name: "output",
Value: "./.index.mf",
Aliases: []string{"o"},
Usage: "Specify output filename",
},
&cli.BoolFlag{
Name: "force",
Aliases: []string{"f"},
Usage: "Overwrite output file if it exists",
},
&cli.BoolFlag{
Name: flagProgress,
Aliases: []string{"P"},
Usage: "Show progress during enumeration and scanning",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
EnvVars: []string{"MFER_SIGN_KEY"},
},
&cli.StringFlag{
Name: "seed",
Usage: "Seed value for deterministic manifest UUID",
EnvVars: []string{"MFER_SEED"},
},
&cli.BoolFlag{
Name: "include-timestamps",
Usage: "Include createdAt timestamp in manifest " +
"(omitted by default for determinism)",
},
),
}
}
func (mfa *CLIApp) checkCommand() *cli.Command {
return &cli.Command{
Name: cmdCheck,
Usage: "Validate files using manifest file",
ArgsUsage: manifestArgsUsage,
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.checkManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths from manifest",
},
&cli.BoolFlag{
Name: flagProgress,
Aliases: []string{"P"},
Usage: "Show progress during checking",
},
&cli.BoolFlag{
Name: "no-extra-files",
Usage: "Fail if files exist in base directory that are not in manifest",
},
&cli.StringFlag{
Name: "require-signature",
Aliases: []string{"S"},
Usage: "Require manifest to be signed by the specified GPG key ID",
EnvVars: []string{"MFER_REQUIRE_SIGNATURE"},
},
),
}
}
func (mfa *CLIApp) freshenCommand() *cli.Command {
return &cli.Command{
Name: "freshen",
Usage: "Update manifest with changed, new, and removed files",
ArgsUsage: manifestArgsUsage,
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.freshenManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths",
},
&cli.BoolFlag{
Name: "follow-symlinks",
Aliases: []string{"L"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "include-dotfiles",
Aliases: []string{"IncludeDotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.BoolFlag{
Name: flagProgress,
Aliases: []string{"P"},
Usage: "Show progress during scanning and hashing",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
EnvVars: []string{"MFER_SIGN_KEY"},
},
&cli.BoolFlag{
Name: "include-timestamps",
Usage: "Include createdAt timestamp in manifest " +
"(omitted by default for determinism)",
},
),
}
}
func (mfa *CLIApp) exportCommand() *cli.Command {
return &cli.Command{
Name: cmdExport,
Usage: "Export manifest contents as JSON",
ArgsUsage: "[manifest file or URL]",
Action: func(c *cli.Context) error {
return mfa.exportManifestOperation(c)
},
}
}
func (mfa *CLIApp) versionCommand() *cli.Command {
return &cli.Command{
Name: "version",
Usage: "Show version",
Action: func(_ *cli.Context) error {
_, _ = fmt.Fprintln(mfa.Stdout, mfa.VersionString())
return nil
},
}
}
func (mfa *CLIApp) listCommand() *cli.Command {
return &cli.Command{
Name: "list",
Aliases: []string{"ls"},
Usage: "List files in manifest",
ArgsUsage: manifestArgsUsage,
Action: func(c *cli.Context) error {
return mfa.listManifestOperation(c)
},
Flags: []cli.Flag{
&cli.BoolFlag{
Name: "long",
Aliases: []string{"l"},
Usage: "Show size and mtime",
},
&cli.BoolFlag{
Name: "print0",
Usage: "Separate entries with NUL character (for xargs -0)",
},
},
}
}
func (mfa *CLIApp) fetchCommand() *cli.Command {
return &cli.Command{
Name: "fetch",
Usage: "fetch manifest and referenced files",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.fetchManifestOperation(c)
},
Flags: commonFlags(),
}
}
func (mfa *CLIApp) run(args []string) { func (mfa *CLIApp) run(args []string) {
mfa.startupTime = time.Now() mfa.startupTime = time.Now()
if NO_COLOR { if NoColor {
// shoutout to rob pike who thinks it's juvenile // shoutout to rob pike who thinks it's juvenile
log.DisableStyling() log.DisableStyling()
} }
@@ -106,173 +333,30 @@ func (mfa *CLIApp) run(args []string) {
ErrWriter: mfa.Stderr, ErrWriter: mfa.Stderr,
Action: func(c *cli.Context) error { Action: func(c *cli.Context) error {
if c.Args().Len() > 0 { if c.Args().Len() > 0 {
return fmt.Errorf("unknown command %q", c.Args().First()) return fmt.Errorf("%w %q", errUnknownCommand, c.Args().First())
} }
mfa.printBanner() mfa.printBanner()
return cli.ShowAppHelp(c) return cli.ShowAppHelp(c)
}, },
Commands: []*cli.Command{ Commands: []*cli.Command{
{ mfa.generateCommand(),
Name: "generate", mfa.checkCommand(),
Aliases: []string{"gen"}, mfa.freshenCommand(),
Usage: "Generate manifest file", mfa.exportCommand(),
Action: func(c *cli.Context) error { mfa.versionCommand(),
mfa.setVerbosity(c) mfa.listCommand(),
mfa.printBanner() mfa.fetchCommand(),
return mfa.generateManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.BoolFlag{
Name: "FollowSymLinks",
Aliases: []string{"follow-symlinks"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "IncludeDotfiles",
Aliases: []string{"include-dotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.StringFlag{
Name: "output",
Value: "./.index.mf",
Aliases: []string{"o"},
Usage: "Specify output filename",
},
&cli.BoolFlag{
Name: "force",
Aliases: []string{"f"},
Usage: "Overwrite output file if it exists",
},
&cli.BoolFlag{
Name: "progress",
Aliases: []string{"P"},
Usage: "Show progress during enumeration and scanning",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
EnvVars: []string{"MFER_SIGN_KEY"},
},
),
},
{
Name: "check",
Usage: "Validate files using manifest file",
ArgsUsage: "[manifest file]",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.checkManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths from manifest",
},
&cli.BoolFlag{
Name: "progress",
Aliases: []string{"P"},
Usage: "Show progress during checking",
},
&cli.BoolFlag{
Name: "no-extra-files",
Usage: "Fail if files exist in base directory that are not in manifest",
},
&cli.StringFlag{
Name: "require-signature",
Aliases: []string{"S"},
Usage: "Require manifest to be signed by the specified GPG key ID",
EnvVars: []string{"MFER_REQUIRE_SIGNATURE"},
},
),
},
{
Name: "freshen",
Usage: "Update manifest with changed, new, and removed files",
ArgsUsage: "[manifest file]",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.freshenManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths",
},
&cli.BoolFlag{
Name: "FollowSymLinks",
Aliases: []string{"follow-symlinks"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "IncludeDotfiles",
Aliases: []string{"include-dotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.BoolFlag{
Name: "progress",
Aliases: []string{"P"},
Usage: "Show progress during scanning and hashing",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
EnvVars: []string{"MFER_SIGN_KEY"},
},
),
},
{
Name: "version",
Usage: "Show version",
Action: func(c *cli.Context) error {
_, _ = fmt.Fprintln(mfa.Stdout, mfa.VersionString())
return nil
},
},
{
Name: "list",
Aliases: []string{"ls"},
Usage: "List files in manifest",
ArgsUsage: "[manifest file]",
Action: func(c *cli.Context) error {
return mfa.listManifestOperation(c)
},
Flags: []cli.Flag{
&cli.BoolFlag{
Name: "long",
Aliases: []string{"l"},
Usage: "Show size and mtime",
},
&cli.BoolFlag{
Name: "print0",
Usage: "Separate entries with NUL character (for xargs -0)",
},
},
},
{
Name: "fetch",
Usage: "fetch manifest and referenced files",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.fetchManifestOperation(c)
},
Flags: commonFlags(),
},
}, },
} }
mfa.app.HideVersion = true mfa.app.HideVersion = false
err := mfa.app.Run(args) err := mfa.app.Run(args)
if err != nil { if err != nil {
mfa.exitCode = 1 mfa.exitCode = 1
log.WithError(err).Debugf("exiting") log.WithError(err).Debugf("exiting")
} }
} }

28
internal/cli/mtime.go Normal file
View File

@@ -0,0 +1,28 @@
package cli
import (
"time"
"sneak.berlin/go/mfer/mfer"
)
// mtimeAbsent is printed in place of a modification time when a manifest
// entry does not carry one.
const mtimeAbsent = "-"
// entryMtime returns the modification time recorded for a manifest entry.
//
// MFFilePath.Mtime is a message pointer with proto3 field presence, so an
// absent mtime is a representable, on-the-wire-valid state. It must never
// be conflated with a recorded mtime of the Unix epoch: callers that
// compare mtimes have to treat "absent" as "unknown", not as
// 1970-01-01T00:00:00Z, or every entry compares as modified. ok reports
// whether an mtime was actually recorded.
func entryMtime(entry *mfer.MFFilePath) (time.Time, bool) {
ts := entry.GetMtime()
if ts == nil {
return time.Time{}, false
}
return time.Unix(ts.GetSeconds(), int64(ts.GetNanos())), true
}

View File

@@ -1,3 +1,5 @@
// Package log provides leveled logging with progress output helpers
// on top of apex/log and pterm.
package log package log
import ( import (
@@ -52,6 +54,11 @@ func (l Level) String() string {
} }
} }
// callerSkip is the runtime.Caller stack depth from the public Debug
// helpers to the caller of the log package.
const callerSkip = 2
//nolint:gochecknoglobals // package-level logger state by design
var ( var (
// mu protects the output writers and level // mu protects the output writers and level
mu sync.RWMutex mu sync.RWMutex
@@ -60,7 +67,7 @@ var (
// stderr is the writer for log output // stderr is the writer for log output
stderr io.Writer = os.Stderr stderr io.Writer = os.Stderr
// currentLevel is our log level (includes Verbose) // currentLevel is our log level (includes Verbose)
currentLevel Level = InfoLevel currentLevel = InfoLevel
) )
// SetOutput configures the output writers for the log package. // SetOutput configures the output writers for the log package.
@@ -68,8 +75,10 @@ var (
func SetOutput(out, err io.Writer) { func SetOutput(out, err io.Writer) {
mu.Lock() mu.Lock()
defer mu.Unlock() defer mu.Unlock()
stdout = out stdout = out
stderr = err stderr = err
pterm.SetDefaultOutput(out) pterm.SetDefaultOutput(out)
} }
@@ -77,6 +86,7 @@ func SetOutput(out, err io.Writer) {
func GetStdout() io.Writer { func GetStdout() io.Writer {
mu.RLock() mu.RLock()
defer mu.RUnlock() defer mu.RUnlock()
return stdout return stdout
} }
@@ -84,6 +94,7 @@ func GetStdout() io.Writer {
func GetStderr() io.Writer { func GetStderr() io.Writer {
mu.RLock() mu.RLock()
defer mu.RUnlock() defer mu.RUnlock()
return stderr return stderr
} }
@@ -91,6 +102,7 @@ func GetStderr() io.Writer {
func DisableStyling() { func DisableStyling() {
pterm.DisableColor() pterm.DisableColor()
pterm.DisableStyling() pterm.DisableStyling()
pterm.Debug.Prefix.Text = "" pterm.Debug.Prefix.Text = ""
pterm.Info.Prefix.Text = "" pterm.Info.Prefix.Text = ""
pterm.Success.Prefix.Text = "" pterm.Success.Prefix.Text = ""
@@ -102,7 +114,9 @@ func DisableStyling() {
// Init initializes the logger with the CLI handler and default log level. // Init initializes the logger with the CLI handler and default log level.
func Init() { func Init() {
mu.RLock() mu.RLock()
w := stderr w := stderr
mu.RUnlock() mu.RUnlock()
log.SetHandler(acli.New(w)) log.SetHandler(acli.New(w))
log.SetLevel(log.DebugLevel) // Let apex/log pass everything; we filter ourselves log.SetLevel(log.DebugLevel) // Let apex/log pass everything; we filter ourselves
@@ -112,11 +126,12 @@ func Init() {
func isEnabled(l Level) bool { func isEnabled(l Level) bool {
mu.RLock() mu.RLock()
defer mu.RUnlock() defer mu.RUnlock()
return l >= currentLevel return l >= currentLevel
} }
// Fatalf logs a formatted message at fatal level. // Fatalf logs a formatted message at fatal level.
func Fatalf(format string, args ...interface{}) { func Fatalf(format string, args ...any) {
if isEnabled(FatalLevel) { if isEnabled(FatalLevel) {
log.Fatalf(format, args...) log.Fatalf(format, args...)
} }
@@ -130,7 +145,7 @@ func Fatal(arg string) {
} }
// Errorf logs a formatted message at error level. // Errorf logs a formatted message at error level.
func Errorf(format string, args ...interface{}) { func Errorf(format string, args ...any) {
if isEnabled(ErrorLevel) { if isEnabled(ErrorLevel) {
log.Errorf(format, args...) log.Errorf(format, args...)
} }
@@ -144,7 +159,7 @@ func Error(arg string) {
} }
// Warnf logs a formatted message at warn level. // Warnf logs a formatted message at warn level.
func Warnf(format string, args ...interface{}) { func Warnf(format string, args ...any) {
if isEnabled(WarnLevel) { if isEnabled(WarnLevel) {
log.Warnf(format, args...) log.Warnf(format, args...)
} }
@@ -158,7 +173,7 @@ func Warn(arg string) {
} }
// Infof logs a formatted message at info level. // Infof logs a formatted message at info level.
func Infof(format string, args ...interface{}) { func Infof(format string, args ...any) {
if isEnabled(InfoLevel) { if isEnabled(InfoLevel) {
log.Infof(format, args...) log.Infof(format, args...)
} }
@@ -172,7 +187,7 @@ func Info(arg string) {
} }
// Verbosef logs a formatted message at verbose level. // Verbosef logs a formatted message at verbose level.
func Verbosef(format string, args ...interface{}) { func Verbosef(format string, args ...any) {
if isEnabled(VerboseLevel) { if isEnabled(VerboseLevel) {
log.Infof(format, args...) log.Infof(format, args...)
} }
@@ -186,16 +201,16 @@ func Verbose(arg string) {
} }
// Debugf logs a formatted message at debug level with caller location. // Debugf logs a formatted message at debug level with caller location.
func Debugf(format string, args ...interface{}) { func Debugf(format string, args ...any) {
if isEnabled(DebugLevel) { if isEnabled(DebugLevel) {
DebugReal(fmt.Sprintf(format, args...), 2) DebugReal(fmt.Sprintf(format, args...), callerSkip)
} }
} }
// Debug logs a message at debug level with caller location. // Debug logs a message at debug level with caller location.
func Debug(arg string) { func Debug(arg string) {
if isEnabled(DebugLevel) { if isEnabled(DebugLevel) {
DebugReal(arg, 2) DebugReal(arg, callerSkip)
} }
} }
@@ -204,18 +219,20 @@ func DebugReal(arg string, cs int) {
if !isEnabled(DebugLevel) { if !isEnabled(DebugLevel) {
return return
} }
_, callerFile, callerLine, ok := runtime.Caller(cs) _, callerFile, callerLine, ok := runtime.Caller(cs)
if !ok { if !ok {
return return
} }
tag := fmt.Sprintf("%s:%d: ", filepath.Base(callerFile), callerLine) tag := fmt.Sprintf("%s:%d: ", filepath.Base(callerFile), callerLine)
log.Debug(tag + arg) log.Debug(tag + arg)
} }
// Dump logs a spew dump of the arguments at debug level. // Dump logs a spew dump of the arguments at debug level.
func Dump(args ...interface{}) { func Dump(args ...any) {
if isEnabled(DebugLevel) { if isEnabled(DebugLevel) {
DebugReal(spew.Sdump(args...), 2) DebugReal(spew.Sdump(args...), callerSkip)
} }
} }
@@ -246,6 +263,7 @@ func SetLevelFromVerbosity(l int) {
func SetLevel(l Level) { func SetLevel(l Level) {
mu.Lock() mu.Lock()
defer mu.Unlock() defer mu.Unlock()
currentLevel = l currentLevel = l
} }
@@ -253,6 +271,7 @@ func SetLevel(l Level) {
func GetLevel() Level { func GetLevel() Level {
mu.RLock() mu.RLock()
defer mu.RUnlock() defer mu.RUnlock()
return currentLevel return currentLevel
} }
@@ -263,7 +282,7 @@ func WithError(e error) *log.Entry {
// Progressf prints a progress message that overwrites the current line. // Progressf prints a progress message that overwrites the current line.
// Use ProgressDone() when progress is complete to move to the next line. // Use ProgressDone() when progress is complete to move to the next line.
func Progressf(format string, args ...interface{}) { func Progressf(format string, args ...any) {
pterm.Printf("\r"+format, args...) pterm.Printf("\r"+format, args...)
} }

View File

@@ -1,12 +1,12 @@
package log package log_test
import ( import (
"testing" "testing"
"github.com/stretchr/testify/assert" "sneak.berlin/go/mfer/internal/log"
) )
func TestBuild(t *testing.T) { func TestBuild(t *testing.T) {
Init() t.Parallel()
assert.True(t, true) log.Init()
} }

View File

@@ -1,3 +1,5 @@
// Package mfer implements the mfer manifest file format: building,
// serializing, verifying, and checking manifests of file trees.
package mfer package mfer
import ( import (
@@ -5,6 +7,7 @@ import (
"errors" "errors"
"fmt" "fmt"
"io" "io"
"sort"
"strings" "strings"
"sync" "sync"
"time" "time"
@@ -13,6 +16,27 @@ import (
"github.com/multiformats/go-multihash" "github.com/multiformats/go-multihash"
) )
// readChunkSize is the buffer size used when reading file contents for
// hashing.
const readChunkSize = 64 * 1024
// The errPath* sentinels below are worded as the trailing fragment of the
// message ValidatePath renders, because the offending path is quoted
// before them (`path %q ...`). Wrapping them mid-sentence keeps the
// rendered text exactly as mfer has always printed it. Match them with
// errors.Is rather than by reading their messages.
var (
errPathEmpty = errors.New("path cannot be empty")
errPathNotUTF8 = errors.New("is not valid UTF-8")
errPathBackslash = errors.New("contains backslash; use forward slashes only")
errPathAbsolute = errors.New("is absolute; must be relative")
errPathEmptySegment = errors.New("contains empty segment")
errPathDotDot = errors.New("contains '..' segment")
errSizeMismatch = errors.New("size mismatch")
errNegativeSize = errors.New("size cannot be negative")
errEmptyHash = errors.New("hash cannot be nil or empty")
)
// ValidatePath checks that a file path conforms to manifest path invariants: // ValidatePath checks that a file path conforms to manifest path invariants:
// - Must be valid UTF-8 // - Must be valid UTF-8
// - Must use forward slashes only (no backslashes) // - Must use forward slashes only (no backslashes)
@@ -22,25 +46,31 @@ import (
// - Must not be empty // - Must not be empty
func ValidatePath(p string) error { func ValidatePath(p string) error {
if p == "" { if p == "" {
return errors.New("path cannot be empty") return errPathEmpty
} }
if !utf8.ValidString(p) { if !utf8.ValidString(p) {
return fmt.Errorf("path %q is not valid UTF-8", p) return fmt.Errorf("path %q %w", p, errPathNotUTF8)
} }
if strings.ContainsRune(p, '\\') { if strings.ContainsRune(p, '\\') {
return fmt.Errorf("path %q contains backslash; use forward slashes only", p) return fmt.Errorf("path %q %w", p, errPathBackslash)
} }
if strings.HasPrefix(p, "/") { if strings.HasPrefix(p, "/") {
return fmt.Errorf("path %q is absolute; must be relative", p) return fmt.Errorf("path %q %w", p, errPathAbsolute)
} }
for _, seg := range strings.Split(p, "/") { for _, seg := range strings.Split(p, "/") {
if seg == "" { if seg == "" {
return fmt.Errorf("path %q contains empty segment", p) return fmt.Errorf("path %q %w", p, errPathEmptySegment)
} }
if seg == ".." { if seg == ".." {
return fmt.Errorf("path %q contains '..' segment", p) return fmt.Errorf("path %q %w", p, errPathDotDot)
} }
} }
return nil return nil
} }
@@ -67,11 +97,7 @@ type UnixNanos int32
// Timestamp converts ModTime to a protobuf Timestamp. // Timestamp converts ModTime to a protobuf Timestamp.
func (m ModTime) Timestamp() *Timestamp { func (m ModTime) Timestamp() *Timestamp {
t := time.Time(m) return newTimestampFromTime(time.Time(m))
return &Timestamp{
Seconds: t.Unix(),
Nanos: int32(t.Nanosecond()),
}
} }
// Multihash represents a multihash-encoded file hash (typically SHA2-256). // Multihash represents a multihash-encoded file hash (typically SHA2-256).
@@ -84,10 +110,12 @@ type FileHashProgress struct {
// Builder constructs a manifest by adding files one at a time. // Builder constructs a manifest by adding files one at a time.
type Builder struct { type Builder struct {
mu sync.Mutex mu sync.Mutex
files []*MFFilePath files []*MFFilePath
createdAt time.Time createdAt time.Time
signingOptions *SigningOptions includeTimestamps bool
signingOptions *SigningOptions
fixedUUID []byte // if set, use this UUID instead of generating one
} }
// NewBuilder creates a new Builder. // NewBuilder creates a new Builder.
@@ -98,6 +126,14 @@ func NewBuilder() *Builder {
} }
} }
// SetSeed derives a deterministic UUID from the given seed string.
// The seed is hashed once with SHA-256 and the first 16 bytes are used
// as a fixed UUID for the manifest.
func (b *Builder) SetSeed(seed string) {
hash := sha256.Sum256([]byte(seed))
b.fixedUUID = hash[:uuidLength]
}
// AddFile reads file content from reader, computes hashes, and adds to manifest. // AddFile reads file content from reader, computes hashes, and adds to manifest.
// Progress updates are sent to the progress channel (if non-nil) without blocking. // Progress updates are sent to the progress channel (if non-nil) without blocking.
// Returns the number of bytes read. // Returns the number of bytes read.
@@ -108,7 +144,8 @@ func (b *Builder) AddFile(
reader io.Reader, reader io.Reader,
progress chan<- FileHashProgress, progress chan<- FileHashProgress,
) (FileSize, error) { ) (FileSize, error) {
if err := ValidatePath(string(path)); err != nil { err := ValidatePath(string(path))
if err != nil {
return 0, err return 0, err
} }
@@ -117,7 +154,8 @@ func (b *Builder) AddFile(
// Read file in chunks, updating hash and progress // Read file in chunks, updating hash and progress
var totalRead FileSize var totalRead FileSize
buf := make([]byte, 64*1024) // 64KB chunks
buf := make([]byte, readChunkSize)
for { for {
n, err := reader.Read(buf) n, err := reader.Read(buf)
@@ -126,9 +164,11 @@ func (b *Builder) AddFile(
totalRead += FileSize(n) totalRead += FileSize(n)
sendFileHashProgress(progress, FileHashProgress{BytesRead: totalRead}) sendFileHashProgress(progress, FileHashProgress{BytesRead: totalRead})
} }
if err == io.EOF { if err == io.EOF {
break break
} }
if err != nil { if err != nil {
return totalRead, err return totalRead, err
} }
@@ -136,7 +176,10 @@ func (b *Builder) AddFile(
// Verify actual bytes read matches declared size // Verify actual bytes read matches declared size
if totalRead != size { if totalRead != size {
return totalRead, fmt.Errorf("size mismatch for %q: declared %d bytes but read %d bytes", path, size, totalRead) return totalRead, fmt.Errorf(
"%w for %q: declared %d bytes but read %d bytes",
errSizeMismatch, path, size, totalRead,
)
} }
// Encode hash as multihash (SHA2-256) // Encode hash as multihash (SHA2-256)
@@ -167,6 +210,7 @@ func sendFileHashProgress(ch chan<- FileHashProgress, p FileHashProgress) {
if ch == nil { if ch == nil {
return return
} }
select { select {
case ch <- p: case ch <- p:
default: default:
@@ -177,21 +221,30 @@ func sendFileHashProgress(ch chan<- FileHashProgress, p FileHashProgress) {
func (b *Builder) FileCount() int { func (b *Builder) FileCount() int {
b.mu.Lock() b.mu.Lock()
defer b.mu.Unlock() defer b.mu.Unlock()
return len(b.files) return len(b.files)
} }
// AddFileWithHash adds a file entry with a pre-computed hash. // AddFileWithHash adds a file entry with a pre-computed hash.
// This is useful when the hash is already known (e.g., from an existing manifest). // This is useful when the hash is already known (e.g., from an existing manifest).
// Returns an error if path is empty, size is negative, or hash is nil/empty. // Returns an error if path is empty, size is negative, or hash is nil/empty.
func (b *Builder) AddFileWithHash(path RelFilePath, size FileSize, mtime ModTime, hash Multihash) error { func (b *Builder) AddFileWithHash(
if err := ValidatePath(string(path)); err != nil { path RelFilePath,
return err size FileSize,
mtime ModTime,
hash Multihash,
) error {
err := ValidatePath(string(path))
if err != nil {
return fmt.Errorf("add file: %w", err)
} }
if size < 0 { if size < 0 {
return errors.New("size cannot be negative") return errNegativeSize
} }
if len(hash) == 0 { if len(hash) == 0 {
return errors.New("hash cannot be nil or empty") return errEmptyHash
} }
entry := &MFFilePath{ entry := &MFFilePath{
@@ -206,14 +259,25 @@ func (b *Builder) AddFileWithHash(path RelFilePath, size FileSize, mtime ModTime
b.mu.Lock() b.mu.Lock()
b.files = append(b.files, entry) b.files = append(b.files, entry)
b.mu.Unlock() b.mu.Unlock()
return nil return nil
} }
// SetIncludeTimestamps controls whether the manifest includes a createdAt timestamp.
// By default timestamps are omitted for deterministic output.
func (b *Builder) SetIncludeTimestamps(include bool) {
b.mu.Lock()
defer b.mu.Unlock()
b.includeTimestamps = include
}
// SetSigningOptions sets the GPG signing options for the manifest. // SetSigningOptions sets the GPG signing options for the manifest.
// If opts is non-nil, the manifest will be signed when Build() is called. // If opts is non-nil, the manifest will be signed when Build() is called.
func (b *Builder) SetSigningOptions(opts *SigningOptions) { func (b *Builder) SetSigningOptions(opts *SigningOptions) {
b.mu.Lock() b.mu.Lock()
defer b.mu.Unlock() defer b.mu.Unlock()
b.signingOptions = opts b.signingOptions = opts
} }
@@ -222,30 +286,44 @@ func (b *Builder) Build(w io.Writer) error {
b.mu.Lock() b.mu.Lock()
defer b.mu.Unlock() defer b.mu.Unlock()
// Sort files by path for deterministic output
sort.Slice(b.files, func(i, j int) bool {
return b.files[i].GetPath() < b.files[j].GetPath()
})
// Create inner manifest // Create inner manifest
inner := &MFFile{ inner := &MFFile{
Version: MFFile_VERSION_ONE, Version: MFFile_VERSION_ONE,
CreatedAt: newTimestampFromTime(b.createdAt), Files: b.files,
Files: b.files, }
if b.includeTimestamps {
inner.CreatedAt = newTimestampFromTime(b.createdAt)
} }
// Create a temporary manifest to use existing serialization // Create a temporary manifest to use existing serialization
m := &manifest{ m := &manifest{
pbInner: inner, pbInner: inner,
signingOptions: b.signingOptions, signingOptions: b.signingOptions,
fixedUUID: b.fixedUUID,
} }
// Generate outer wrapper // Generate outer wrapper
if err := m.generateOuter(); err != nil { err := m.generateOuter()
return err if err != nil {
return fmt.Errorf("build: generate outer: %w", err)
} }
// Generate final output // Generate final output
if err := m.generate(); err != nil { err = m.generate()
return err if err != nil {
return fmt.Errorf("build: generate: %w", err)
} }
// Write to output // Write to output
_, err := w.Write(m.output.Bytes()) _, err = w.Write(m.output.Bytes())
return err if err != nil {
return fmt.Errorf("build: write output: %w", err)
}
return nil
} }

View File

@@ -1,3 +1,4 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer package mfer
import ( import (
@@ -10,24 +11,34 @@ import (
"github.com/stretchr/testify/require" "github.com/stretchr/testify/require"
) )
const testFileName = "file.txt"
func TestNewBuilder(t *testing.T) { func TestNewBuilder(t *testing.T) {
t.Parallel()
b := NewBuilder() b := NewBuilder()
assert.NotNil(t, b) assert.NotNil(t, b)
assert.Equal(t, 0, b.FileCount()) assert.Equal(t, 0, b.FileCount())
} }
func TestBuilderAddFile(t *testing.T) { func TestBuilderAddFile(t *testing.T) {
t.Parallel()
b := NewBuilder() b := NewBuilder()
content := []byte("test content") content := []byte("test content")
reader := bytes.NewReader(content) reader := bytes.NewReader(content)
bytesRead, err := b.AddFile("test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil) bytesRead, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
require.NoError(t, err) require.NoError(t, err)
assert.Equal(t, FileSize(len(content)), bytesRead) assert.Equal(t, FileSize(len(content)), bytesRead)
assert.Equal(t, 1, b.FileCount()) assert.Equal(t, 1, b.FileCount())
} }
func TestBuilderAddFileWithHash(t *testing.T) { func TestBuilderAddFileWithHash(t *testing.T) {
t.Parallel()
b := NewBuilder() b := NewBuilder()
hash := make([]byte, 34) // SHA256 multihash is 34 bytes hash := make([]byte, 34) // SHA256 multihash is 34 bytes
@@ -37,54 +48,71 @@ func TestBuilderAddFileWithHash(t *testing.T) {
} }
func TestBuilderAddFileWithHashValidation(t *testing.T) { func TestBuilderAddFileWithHashValidation(t *testing.T) {
t.Parallel()
t.Run("empty path", func(t *testing.T) { t.Run("empty path", func(t *testing.T) {
t.Parallel()
b := NewBuilder() b := NewBuilder()
hash := make([]byte, 34) hash := make([]byte, 34)
err := b.AddFileWithHash("", 100, ModTime(time.Now()), hash) err := b.AddFileWithHash("", 100, ModTime(time.Now()), hash)
assert.Error(t, err) require.Error(t, err)
assert.Contains(t, err.Error(), "path") assert.Contains(t, err.Error(), "path")
}) })
t.Run("negative size", func(t *testing.T) { t.Run("negative size", func(t *testing.T) {
t.Parallel()
b := NewBuilder() b := NewBuilder()
hash := make([]byte, 34) hash := make([]byte, 34)
err := b.AddFileWithHash("test.txt", -1, ModTime(time.Now()), hash) err := b.AddFileWithHash("test.txt", -1, ModTime(time.Now()), hash)
assert.Error(t, err) require.Error(t, err)
assert.Contains(t, err.Error(), "size") assert.Contains(t, err.Error(), "size")
}) })
t.Run("nil hash", func(t *testing.T) { t.Run("nil hash", func(t *testing.T) {
t.Parallel()
b := NewBuilder() b := NewBuilder()
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), nil) err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), nil)
assert.Error(t, err) require.Error(t, err)
assert.Contains(t, err.Error(), "hash") assert.Contains(t, err.Error(), "hash")
}) })
t.Run("empty hash", func(t *testing.T) { t.Run("empty hash", func(t *testing.T) {
t.Parallel()
b := NewBuilder() b := NewBuilder()
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), []byte{}) err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), []byte{})
assert.Error(t, err) require.Error(t, err)
assert.Contains(t, err.Error(), "hash") assert.Contains(t, err.Error(), "hash")
}) })
t.Run("valid inputs", func(t *testing.T) { t.Run("valid inputs", func(t *testing.T) {
t.Parallel()
b := NewBuilder() b := NewBuilder()
hash := make([]byte, 34) hash := make([]byte, 34)
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), hash) err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), hash)
assert.NoError(t, err) require.NoError(t, err)
assert.Equal(t, 1, b.FileCount()) assert.Equal(t, 1, b.FileCount())
}) })
} }
func TestBuilderBuild(t *testing.T) { func TestBuilderBuild(t *testing.T) {
t.Parallel()
b := NewBuilder() b := NewBuilder()
content := []byte("test content") content := []byte("test content")
reader := bytes.NewReader(content) reader := bytes.NewReader(content)
_, err := b.AddFile("test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil) _, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
require.NoError(t, err) require.NoError(t, err)
var buf bytes.Buffer var buf bytes.Buffer
err = b.Build(&buf) err = b.Build(&buf)
require.NoError(t, err) require.NoError(t, err)
@@ -93,6 +121,8 @@ func TestBuilderBuild(t *testing.T) {
} }
func TestNewTimestampFromTimeExtremeDate(t *testing.T) { func TestNewTimestampFromTimeExtremeDate(t *testing.T) {
t.Parallel()
// Regression test: newTimestampFromTime used UnixNano() which panics // Regression test: newTimestampFromTime used UnixNano() which panics
// for dates outside ~1678-2262. Now uses Nanosecond() which is safe. // for dates outside ~1678-2262. Now uses Nanosecond() which is safe.
tests := []struct { tests := []struct {
@@ -107,21 +137,351 @@ func TestNewTimestampFromTimeExtremeDate(t *testing.T) {
for _, tt := range tests { for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) { t.Run(tt.name, func(t *testing.T) {
t.Parallel()
// Should not panic // Should not panic
ts := newTimestampFromTime(tt.time) ts := newTimestampFromTime(tt.time)
assert.Equal(t, tt.time.Unix(), ts.Seconds) assert.Equal(t, tt.time.Unix(), ts.GetSeconds())
assert.Equal(t, int32(tt.time.Nanosecond()), ts.Nanos) assert.Equal(t, tt.time.Nanosecond(), int(ts.GetNanos()))
}) })
} }
} }
func TestBuilderDeterministicOutput(t *testing.T) {
t.Parallel()
buildManifest := func() []byte {
b := NewBuilder()
// Use a fixed createdAt and UUID so output is reproducible
b.createdAt = time.Date(2025, 1, 1, 0, 0, 0, 0, time.UTC)
b.fixedUUID = make([]byte, 16) // all zeros
mtime := ModTime(time.Date(2025, 6, 1, 0, 0, 0, 0, time.UTC))
// Add files in reverse order to test sorting
files := []struct {
path string
content string
}{
{"c/file.txt", "content c"},
{"a/file.txt", "content a"},
{"b/file.txt", "content b"},
}
for _, f := range files {
r := bytes.NewReader([]byte(f.content))
_, err := b.AddFile(
RelFilePath(f.path), FileSize(len(f.content)), mtime, r, nil,
)
require.NoError(t, err)
}
var buf bytes.Buffer
err := b.Build(&buf)
require.NoError(t, err)
return buf.Bytes()
}
out1 := buildManifest()
out2 := buildManifest()
assert.Equal(t, out1, out2,
"two builds with same input should produce byte-identical output")
}
func TestSetSeedDeterministic(t *testing.T) {
t.Parallel()
b1 := NewBuilder()
b1.SetSeed("test-seed-value")
b2 := NewBuilder()
b2.SetSeed("test-seed-value")
assert.Equal(t, b1.fixedUUID, b2.fixedUUID, "same seed should produce same UUID")
assert.Len(t, b1.fixedUUID, 16, "UUID should be 16 bytes")
b3 := NewBuilder()
b3.SetSeed("different-seed")
assert.NotEqual(t, b1.fixedUUID, b3.fixedUUID,
"different seeds should produce different UUIDs")
}
func TestValidatePath(t *testing.T) {
t.Parallel()
valid := []string{
testFileName,
"dir/file.txt",
"a/b/c/d.txt",
"file with spaces.txt",
"日本語.txt", //nolint:gosmopolitan // deliberately tests non-ASCII UTF-8 paths
}
for _, p := range valid {
t.Run("valid:"+p, func(t *testing.T) {
t.Parallel()
assert.NoError(t, ValidatePath(p))
})
}
invalid := []struct {
path string
desc string
}{
{"", "empty"},
{"/absolute", "absolute path"},
{"has\\backslash", "backslash"},
{"has/../traversal", "dot-dot segment"},
{"has//double", "empty segment"},
{"..", "just dot-dot"},
{string([]byte{0xff, 0xfe}), "invalid UTF-8"},
}
for _, tt := range invalid {
t.Run("invalid:"+tt.desc, func(t *testing.T) {
t.Parallel()
assert.Error(t, ValidatePath(tt.path))
})
}
}
func TestBuilderAddFileSizeMismatch(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("short")
reader := bytes.NewReader(content)
// Declare wrong size
_, err := b.AddFile("test.txt", FileSize(100), ModTime(time.Now()), reader, nil)
require.Error(t, err)
assert.Contains(t, err.Error(), "size mismatch")
}
func TestBuilderAddFileInvalidPath(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("data")
reader := bytes.NewReader(content)
_, err := b.AddFile("", FileSize(len(content)), ModTime(time.Now()), reader, nil)
require.Error(t, err)
reader.Reset(content)
_, err = b.AddFile(
"/absolute", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
assert.Error(t, err)
}
func TestBuilderAddFileWithProgress(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := bytes.Repeat([]byte("x"), 1000)
reader := bytes.NewReader(content)
progress := make(chan FileHashProgress, 100)
bytesRead, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, progress,
)
close(progress)
require.NoError(t, err)
assert.Equal(t, FileSize(1000), bytesRead)
var updates []FileHashProgress
for p := range progress {
updates = append(updates, p)
}
assert.NotEmpty(t, updates)
// Last update should show all bytes
assert.Equal(t, FileSize(1000), updates[len(updates)-1].BytesRead)
}
func TestBuilderBuildRoundTrip(t *testing.T) {
t.Parallel()
// Build a manifest, deserialize it, verify all fields survive round-trip
b := NewBuilder()
now := time.Date(2025, 6, 15, 12, 0, 0, 0, time.UTC)
files := []struct {
path string
content []byte
}{
{"alpha.txt", []byte("alpha content")},
{"beta/gamma.txt", []byte("gamma content")},
{"beta/delta.txt", []byte("delta content")},
}
for _, f := range files {
reader := bytes.NewReader(f.content)
_, err := b.AddFile(
RelFilePath(f.path), FileSize(len(f.content)), ModTime(now), reader, nil,
)
require.NoError(t, err)
}
var buf bytes.Buffer
require.NoError(t, b.Build(&buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
mfiles := m.Files()
require.Len(t, mfiles, 3)
// Verify sorted order
assert.Equal(t, "alpha.txt", mfiles[0].GetPath())
assert.Equal(t, "beta/delta.txt", mfiles[1].GetPath())
assert.Equal(t, "beta/gamma.txt", mfiles[2].GetPath())
// Verify sizes
assert.Equal(t, int64(len("alpha content")), mfiles[0].GetSize())
// Verify hashes are present
for _, f := range mfiles {
require.NotEmpty(t, f.GetHashes(), "file %s should have hashes", f.GetPath())
assert.NotEmpty(t, f.GetHashes()[0].GetMultiHash())
}
}
func TestNewManifestFromReaderInvalidMagic(t *testing.T) {
t.Parallel()
_, err := NewManifestFromReader(bytes.NewReader([]byte("NOT_VALID")))
require.Error(t, err)
assert.Contains(t, err.Error(), "invalid file format")
}
func TestNewManifestFromReaderEmpty(t *testing.T) {
t.Parallel()
_, err := NewManifestFromReader(bytes.NewReader([]byte{}))
assert.Error(t, err)
}
func TestNewManifestFromReaderTruncated(t *testing.T) {
t.Parallel()
// Just the magic with nothing after
_, err := NewManifestFromReader(bytes.NewReader([]byte(MAGIC)))
assert.Error(t, err)
}
func TestManifestString(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("test")
reader := bytes.NewReader(content)
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
require.NoError(t, err)
var buf bytes.Buffer
require.NoError(t, b.Build(&buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.Contains(t, m.String(), "count=1")
}
func TestBuilderBuildEmpty(t *testing.T) { func TestBuilderBuildEmpty(t *testing.T) {
t.Parallel()
b := NewBuilder() b := NewBuilder()
var buf bytes.Buffer var buf bytes.Buffer
err := b.Build(&buf) err := b.Build(&buf)
require.NoError(t, err) require.NoError(t, err)
// Should still produce valid manifest with 0 files // Should still produce valid manifest with 0 files
assert.True(t, strings.HasPrefix(buf.String(), MAGIC)) assert.True(t, strings.HasPrefix(buf.String(), MAGIC))
} }
func TestBuilderOmitsCreatedAtByDefault(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("hello")
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()),
bytes.NewReader(content), nil,
)
require.NoError(t, err)
var buf bytes.Buffer
require.NoError(t, b.Build(&buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.Nil(t, m.pbInner.GetCreatedAt(),
"createdAt should be nil by default for deterministic output")
}
func TestBuilderIncludesCreatedAtWhenRequested(t *testing.T) {
t.Parallel()
b := NewBuilder()
b.SetIncludeTimestamps(true)
content := []byte("hello")
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()),
bytes.NewReader(content), nil,
)
require.NoError(t, err)
var buf bytes.Buffer
require.NoError(t, b.Build(&buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.NotNil(t, m.pbInner.GetCreatedAt(),
"createdAt should be set when IncludeTimestamps is true")
}
func TestBuilderDeterministicFileOrder(t *testing.T) {
t.Parallel()
// Two builds with same files in different order should produce same file ordering.
// Note: UUIDs differ per build, so we compare parsed file lists, not raw bytes.
buildAndParse := func(order []string) []*MFFilePath {
b := NewBuilder()
for _, name := range order {
content := []byte("content of " + name)
_, err := b.AddFile(
RelFilePath(name), FileSize(len(content)),
ModTime(time.Unix(1000, 0)), bytes.NewReader(content), nil,
)
require.NoError(t, err)
}
var buf bytes.Buffer
require.NoError(t, b.Build(&buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
return m.Files()
}
files1 := buildAndParse([]string{"b.txt", "a.txt"})
files2 := buildAndParse([]string{"a.txt", "b.txt"})
require.Len(t, files1, 2)
require.Len(t, files2, 2)
for i := range files1 {
assert.Equal(t, files1[i].GetPath(), files2[i].GetPath())
assert.Equal(t, files1[i].GetSize(), files2[i].GetSize())
}
assert.Equal(t, "a.txt", files1[0].GetPath())
assert.Equal(t, "b.txt", files1[1].GetPath())
}

View File

@@ -14,6 +14,8 @@ import (
"github.com/spf13/afero" "github.com/spf13/afero"
) )
var errNoSigningPubKey = errors.New("manifest has no signing public key")
// Result represents the outcome of checking a single file. // Result represents the outcome of checking a single file.
type Result struct { type Result struct {
Path RelFilePath // Relative path from manifest Path RelFilePath // Relative path from manifest
@@ -24,6 +26,7 @@ type Result struct {
// Status represents the verification status of a file. // Status represents the verification status of a file.
type Status int type Status int
// Verification result statuses reported for each checked file.
const ( const (
StatusOK Status = iota // File matches manifest (size and hash verified) StatusOK Status = iota // File matches manifest (size and hash verified)
StatusMissing // File not found on disk StatusMissing // File not found on disk
@@ -70,6 +73,9 @@ type Checker struct {
fs afero.Fs fs afero.Fs
// manifestPaths is a set of paths in the manifest for quick lookup // manifestPaths is a set of paths in the manifest for quick lookup
manifestPaths map[RelFilePath]struct{} manifestPaths map[RelFilePath]struct{}
// manifestRelPath is the relative path of the manifest file from
// basePath (for exclusion)
manifestRelPath RelFilePath
// signature info from the manifest // signature info from the manifest
signature []byte signature []byte
signer []byte signer []byte
@@ -95,19 +101,32 @@ func NewChecker(manifestPath string, basePath string, fs afero.Fs) (*Checker, er
} }
files := m.Files() files := m.Files()
manifestPaths := make(map[RelFilePath]struct{}, len(files)) manifestPaths := make(map[RelFilePath]struct{}, len(files))
for _, f := range files { for _, f := range files {
manifestPaths[RelFilePath(f.Path)] = struct{}{} manifestPaths[RelFilePath(f.GetPath())] = struct{}{}
}
// Compute manifest's relative path from basePath for exclusion in FindExtraFiles
absManifest, err := filepath.Abs(manifestPath)
if err != nil {
return nil, err
}
manifestRel, err := filepath.Rel(abs, absManifest)
if err != nil {
manifestRel = ""
} }
return &Checker{ return &Checker{
basePath: AbsFilePath(abs), basePath: AbsFilePath(abs),
files: files, files: files,
fs: fs, fs: fs,
manifestPaths: manifestPaths, manifestPaths: manifestPaths,
signature: m.pbOuter.Signature, manifestRelPath: RelFilePath(manifestRel),
signer: m.pbOuter.Signer, signature: m.pbOuter.GetSignature(),
signingPubKey: m.pbOuter.SigningPubKey, signer: m.pbOuter.GetSigner(),
signingPubKey: m.pbOuter.GetSigningPubKey(),
}, nil }, nil
} }
@@ -120,8 +139,9 @@ func (c *Checker) FileCount() FileCount {
func (c *Checker) TotalBytes() FileSize { func (c *Checker) TotalBytes() FileSize {
var total FileSize var total FileSize
for _, f := range c.files { for _, f := range c.files {
total += FileSize(f.Size) total += FileSize(f.GetSize())
} }
return total return total
} }
@@ -135,7 +155,8 @@ func (c *Checker) Signer() []byte {
return c.signer return c.signer
} }
// SigningPubKey returns the signing public key if the manifest is signed, nil otherwise. // SigningPubKey returns the signing public key if the manifest is signed,
// nil otherwise.
func (c *Checker) SigningPubKey() []byte { func (c *Checker) SigningPubKey() []byte {
return c.signingPubKey return c.signingPubKey
} }
@@ -145,8 +166,9 @@ func (c *Checker) SigningPubKey() []byte {
// returns its actual fingerprint from the key material itself. // returns its actual fingerprint from the key material itself.
func (c *Checker) ExtractEmbeddedSigningKeyFP() (string, error) { func (c *Checker) ExtractEmbeddedSigningKeyFP() (string, error) {
if len(c.signingPubKey) == 0 { if len(c.signingPubKey) == 0 {
return "", errors.New("manifest has no signing public key") return "", errNoSigningPubKey
} }
return gpgExtractPubKeyFingerprint(c.signingPubKey) return gpgExtractPubKeyFingerprint(c.signingPubKey)
} }
@@ -154,10 +176,15 @@ func (c *Checker) ExtractEmbeddedSigningKeyFP() (string, error) {
// Results are sent to the results channel as files are checked. // Results are sent to the results channel as files are checked.
// Progress updates are sent to the progress channel approximately once per second. // Progress updates are sent to the progress channel approximately once per second.
// Both channels are closed when the method returns. // Both channels are closed when the method returns.
func (c *Checker) Check(ctx context.Context, results chan<- Result, progress chan<- CheckStatus) error { func (c *Checker) Check(
ctx context.Context,
results chan<- Result,
progress chan<- CheckStatus,
) error {
if results != nil { if results != nil {
defer close(results) defer close(results)
} }
if progress != nil { if progress != nil {
defer close(progress) defer close(progress)
} }
@@ -165,11 +192,14 @@ func (c *Checker) Check(ctx context.Context, results chan<- Result, progress cha
totalFiles := FileCount(len(c.files)) totalFiles := FileCount(len(c.files))
totalBytes := c.TotalBytes() totalBytes := c.TotalBytes()
var checkedFiles FileCount var (
var checkedBytes FileSize checkedFiles FileCount
var failures FileCount checkedBytes FileSize
failures FileCount
)
startTime := time.Now() startTime := time.Now()
lastProgressTime := time.Now()
for _, entry := range c.files { for _, entry := range c.files {
select { select {
@@ -182,94 +212,41 @@ func (c *Checker) Check(ctx context.Context, results chan<- Result, progress cha
if result.Status != StatusOK { if result.Status != StatusOK {
failures++ failures++
} }
checkedFiles++ checkedFiles++
if results != nil { if results != nil {
results <- result results <- result
} }
// Send progress with rate and ETA calculation // Send progress at most once per second (rate-limited)
if progress != nil { if progress != nil {
elapsed := time.Since(startTime) now := time.Now()
var bytesPerSec float64
var eta time.Duration
if elapsed > 0 && checkedBytes > 0 { isLast := checkedFiles == totalFiles
bytesPerSec = float64(checkedBytes) / elapsed.Seconds() if isLast || now.Sub(lastProgressTime) >= time.Second {
remainingBytes := totalBytes - checkedBytes bytesPerSec, eta := computeRateETA(
if bytesPerSec > 0 { time.Since(startTime), checkedBytes, totalBytes,
eta = time.Duration(float64(remainingBytes)/bytesPerSec) * time.Second )
}
sendCheckStatus(progress, CheckStatus{
TotalFiles: totalFiles,
CheckedFiles: checkedFiles,
TotalBytes: totalBytes,
CheckedBytes: checkedBytes,
BytesPerSec: bytesPerSec,
ETA: eta,
Failures: failures,
})
lastProgressTime = now
} }
sendCheckStatus(progress, CheckStatus{
TotalFiles: totalFiles,
CheckedFiles: checkedFiles,
TotalBytes: totalBytes,
CheckedBytes: checkedBytes,
BytesPerSec: bytesPerSec,
ETA: eta,
Failures: failures,
})
} }
} }
return nil return nil
} }
func (c *Checker) checkFile(entry *MFFilePath, checkedBytes *FileSize) Result {
absPath := filepath.Join(string(c.basePath), entry.Path)
relPath := RelFilePath(entry.Path)
// Check if file exists
info, err := c.fs.Stat(absPath)
if err != nil {
if errors.Is(err, os.ErrNotExist) || errors.Is(err, afero.ErrFileNotFound) {
return Result{Path: relPath, Status: StatusMissing, Message: "file not found"}
}
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check size
if info.Size() != entry.Size {
*checkedBytes += FileSize(info.Size())
return Result{
Path: relPath,
Status: StatusSizeMismatch,
Message: "size mismatch",
}
}
// Open and hash file
f, err := c.fs.Open(absPath)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
defer func() { _ = f.Close() }()
h := sha256.New()
n, err := io.Copy(h, f)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
*checkedBytes += FileSize(n)
// Encode as multihash and compare
computed, err := multihash.Encode(h.Sum(nil), multihash.SHA2_256)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check against all hashes in manifest (at least one must match)
for _, hash := range entry.Hashes {
if bytes.Equal(computed, hash.MultiHash) {
return Result{Path: relPath, Status: StatusOK}
}
}
return Result{Path: relPath, Status: StatusHashMismatch, Message: "hash mismatch"}
}
// FindExtraFiles walks the filesystem and reports files not in the manifest. // FindExtraFiles walks the filesystem and reports files not in the manifest.
// Results are sent to the results channel. The channel is closed when done. // Results are sent to the results channel. The channel is closed when done.
// Hidden files/directories (starting with .) are skipped, as they are excluded // Hidden files/directories (starting with .) are skipped, as they are excluded
@@ -279,7 +256,7 @@ func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) err
defer close(results) defer close(results)
} }
return afero.Walk(c.fs, string(c.basePath), func(walkPath string, info os.FileInfo, err error) error { walkFn := func(walkPath string, info os.FileInfo, err error) error {
if err != nil { if err != nil {
return err return err
} }
@@ -301,6 +278,7 @@ func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) err
if info.IsDir() { if info.IsDir() {
return filepath.SkipDir return filepath.SkipDir
} }
return nil return nil
} }
@@ -309,14 +287,13 @@ func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) err
return nil return nil
} }
// Skip manifest files relPath := RelFilePath(rel)
base := filepath.Base(rel)
if base == "index.mf" || base == ".index.mf" { // Skip the manifest file itself
if relPath == c.manifestRelPath {
return nil return nil
} }
relPath := RelFilePath(rel)
// Check if path is in manifest // Check if path is in manifest
if _, exists := c.manifestPaths[relPath]; !exists { if _, exists := c.manifestPaths[relPath]; !exists {
if results != nil { if results != nil {
@@ -329,7 +306,75 @@ func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) err
} }
return nil return nil
}) }
return afero.Walk(c.fs, string(c.basePath), walkFn)
}
func (c *Checker) checkFile(entry *MFFilePath, checkedBytes *FileSize) Result {
absPath := filepath.Join(string(c.basePath), entry.GetPath())
relPath := RelFilePath(entry.GetPath())
// Check if file exists
info, err := c.fs.Stat(absPath)
if err != nil {
if errors.Is(err, os.ErrNotExist) || errors.Is(err, afero.ErrFileNotFound) {
return Result{
Path: relPath,
Status: StatusMissing,
Message: "file not found",
}
}
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check size
if info.Size() != entry.GetSize() {
*checkedBytes += FileSize(info.Size())
return Result{
Path: relPath,
Status: StatusSizeMismatch,
Message: "size mismatch",
}
}
// Open and hash file
f, err := c.fs.Open(absPath)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
defer func() { _ = f.Close() }()
h := sha256.New()
n, err := io.Copy(h, f)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
*checkedBytes += FileSize(n)
// Encode as multihash and compare
computed, err := multihash.Encode(h.Sum(nil), multihash.SHA2_256)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check against all hashes in manifest (at least one must match)
for _, hash := range entry.GetHashes() {
if bytes.Equal(computed, hash.GetMultiHash()) {
return Result{Path: relPath, Status: StatusOK}
}
}
return Result{
Path: relPath,
Status: StatusHashMismatch,
Message: "hash mismatch",
}
} }
// sendCheckStatus sends a status update without blocking. // sendCheckStatus sends a status update without blocking.
@@ -337,6 +382,7 @@ func sendCheckStatus(ch chan<- CheckStatus, status CheckStatus) {
if ch == nil { if ch == nil {
return return
} }
select { select {
case ch <- status: case ch <- status:
default: default:

View File

@@ -1,8 +1,10 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer package mfer
import ( import (
"bytes" "bytes"
"context" "context"
"fmt"
"testing" "testing"
"time" "time"
@@ -11,7 +13,15 @@ import (
"github.com/stretchr/testify/require" "github.com/stretchr/testify/require"
) )
const (
testFile1 = "file1.txt"
testFile2 = "file2.txt"
testExistsFile = "exists.txt"
)
func TestStatusString(t *testing.T) { func TestStatusString(t *testing.T) {
t.Parallel()
tests := []struct { tests := []struct {
status Status status Status
expected string expected string
@@ -27,19 +37,26 @@ func TestStatusString(t *testing.T) {
for _, tt := range tests { for _, tt := range tests {
t.Run(tt.expected, func(t *testing.T) { t.Run(tt.expected, func(t *testing.T) {
t.Parallel()
assert.Equal(t, tt.expected, tt.status.String()) assert.Equal(t, tt.expected, tt.status.String())
}) })
} }
} }
// createTestManifest creates a manifest file in the filesystem with the given files. // createTestManifest creates a manifest file in the filesystem with the given files.
func createTestManifest(t *testing.T, fs afero.Fs, manifestPath string, files map[string][]byte) { func createTestManifest(
t *testing.T, fs afero.Fs, manifestPath string, files map[string][]byte,
) {
t.Helper() t.Helper()
builder := NewBuilder() builder := NewBuilder()
for path, content := range files { for path, content := range files {
reader := bytes.NewReader(content) reader := bytes.NewReader(content)
_, err := builder.AddFile(RelFilePath(path), FileSize(len(content)), ModTime(time.Now()), reader, nil) _, err := builder.AddFile(
RelFilePath(path), FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
require.NoError(t, err) require.NoError(t, err)
} }
@@ -48,10 +65,13 @@ func createTestManifest(t *testing.T, fs afero.Fs, manifestPath string, files ma
require.NoError(t, afero.WriteFile(fs, manifestPath, buf.Bytes(), 0o644)) require.NoError(t, afero.WriteFile(fs, manifestPath, buf.Bytes(), 0o644))
} }
// createFilesOnDisk creates the given files on the filesystem. // createFilesOnDisk creates the given files on the filesystem under
func createFilesOnDisk(t *testing.T, fs afero.Fs, basePath string, files map[string][]byte) { // /data.
func createFilesOnDisk(t *testing.T, fs afero.Fs, files map[string][]byte) {
t.Helper() t.Helper()
basePath := "/data"
for path, content := range files { for path, content := range files {
fullPath := basePath + "/" + path fullPath := basePath + "/" + path
require.NoError(t, fs.MkdirAll(basePath, 0o755)) require.NoError(t, fs.MkdirAll(basePath, 0o755))
@@ -60,11 +80,15 @@ func createFilesOnDisk(t *testing.T, fs afero.Fs, basePath string, files map[str
} }
func TestNewChecker(t *testing.T) { func TestNewChecker(t *testing.T) {
t.Parallel()
t.Run("valid manifest", func(t *testing.T) { t.Run("valid manifest", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{ files := map[string][]byte{
"file1.txt": []byte("hello"), testFile1: []byte("hello"),
"file2.txt": []byte("world"), testFile2: []byte("world"),
} }
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
@@ -75,12 +99,16 @@ func TestNewChecker(t *testing.T) {
}) })
t.Run("missing manifest", func(t *testing.T) { t.Run("missing manifest", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
_, err := NewChecker("/nonexistent.mf", "/", fs) _, err := NewChecker("/nonexistent.mf", "/", fs)
assert.Error(t, err) assert.Error(t, err)
}) })
t.Run("invalid manifest", func(t *testing.T) { t.Run("invalid manifest", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/bad.mf", []byte("not a manifest"), 0o644)) require.NoError(t, afero.WriteFile(fs, "/bad.mf", []byte("not a manifest"), 0o644))
_, err := NewChecker("/bad.mf", "/", fs) _, err := NewChecker("/bad.mf", "/", fs)
@@ -89,6 +117,8 @@ func TestNewChecker(t *testing.T) {
} }
func TestCheckerFileCountAndTotalBytes(t *testing.T) { func TestCheckerFileCountAndTotalBytes(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{ files := map[string][]byte{
"small.txt": []byte("hi"), "small.txt": []byte("hi"),
@@ -105,13 +135,15 @@ func TestCheckerFileCountAndTotalBytes(t *testing.T) {
} }
func TestCheckAllFilesOK(t *testing.T) { func TestCheckAllFilesOK(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{ files := map[string][]byte{
"file1.txt": []byte("content one"), testFile1: []byte("content one"),
"file2.txt": []byte("content two"), testFile2: []byte("content two"),
} }
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files) createFilesOnDisk(t, fs, files)
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err) require.NoError(t, err)
@@ -126,21 +158,24 @@ func TestCheckAllFilesOK(t *testing.T) {
} }
assert.Len(t, resultList, 2) assert.Len(t, resultList, 2)
for _, r := range resultList { for _, r := range resultList {
assert.Equal(t, StatusOK, r.Status, "file %s should be OK", r.Path) assert.Equal(t, StatusOK, r.Status, "file %s should be OK", r.Path)
} }
} }
func TestCheckMissingFile(t *testing.T) { func TestCheckMissingFile(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{ files := map[string][]byte{
"exists.txt": []byte("I exist"), testExistsFile: []byte("I exist"),
"missing.txt": []byte("I don't exist on disk"), "missing.txt": []byte("I don't exist on disk"),
} }
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
// Only create one file // Only create one file
createFilesOnDisk(t, fs, "/data", map[string][]byte{ createFilesOnDisk(t, fs, map[string][]byte{
"exists.txt": []byte("I exist"), testExistsFile: []byte("I exist"),
}) })
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
@@ -151,13 +186,17 @@ func TestCheckMissingFile(t *testing.T) {
require.NoError(t, err) require.NoError(t, err)
var okCount, missingCount int var okCount, missingCount int
for r := range results { for r := range results {
switch r.Status { switch r.Status {
case StatusOK: case StatusOK:
okCount++ okCount++
case StatusMissing: case StatusMissing:
missingCount++ missingCount++
assert.Equal(t, RelFilePath("missing.txt"), r.Path) assert.Equal(t, RelFilePath("missing.txt"), r.Path)
case StatusSizeMismatch, StatusHashMismatch, StatusExtra, StatusError:
// Not expected in this test; counted assertions below will fail.
} }
} }
@@ -166,14 +205,16 @@ func TestCheckMissingFile(t *testing.T) {
} }
func TestCheckSizeMismatch(t *testing.T) { func TestCheckSizeMismatch(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{ files := map[string][]byte{
"file.txt": []byte("original content"), testFileName: []byte("original content"),
} }
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
// Create file with different size // Create file with different size
createFilesOnDisk(t, fs, "/data", map[string][]byte{ createFilesOnDisk(t, fs, map[string][]byte{
"file.txt": []byte("short"), testFileName: []byte("short"),
}) })
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
@@ -185,21 +226,23 @@ func TestCheckSizeMismatch(t *testing.T) {
r := <-results r := <-results
assert.Equal(t, StatusSizeMismatch, r.Status) assert.Equal(t, StatusSizeMismatch, r.Status)
assert.Equal(t, RelFilePath("file.txt"), r.Path) assert.Equal(t, RelFilePath(testFileName), r.Path)
} }
func TestCheckHashMismatch(t *testing.T) { func TestCheckHashMismatch(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
originalContent := []byte("original content") originalContent := []byte("original content")
files := map[string][]byte{ files := map[string][]byte{
"file.txt": originalContent, testFileName: originalContent,
} }
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
// Create file with same size but different content // Create file with same size but different content
differentContent := []byte("different contnt") // same length (16 bytes) but different differentContent := []byte("different contnt") // same length (16 bytes) but different
require.Equal(t, len(originalContent), len(differentContent), "test requires same length") require.Len(t, differentContent, len(originalContent), "test requires same length")
createFilesOnDisk(t, fs, "/data", map[string][]byte{ createFilesOnDisk(t, fs, map[string][]byte{
"file.txt": differentContent, testFileName: differentContent,
}) })
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
@@ -211,17 +254,19 @@ func TestCheckHashMismatch(t *testing.T) {
r := <-results r := <-results
assert.Equal(t, StatusHashMismatch, r.Status) assert.Equal(t, StatusHashMismatch, r.Status)
assert.Equal(t, RelFilePath("file.txt"), r.Path) assert.Equal(t, RelFilePath(testFileName), r.Path)
} }
func TestCheckWithProgress(t *testing.T) { func TestCheckWithProgress(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{ files := map[string][]byte{
"file1.txt": bytes.Repeat([]byte("a"), 100), testFile1: bytes.Repeat([]byte("a"), 100),
"file2.txt": bytes.Repeat([]byte("b"), 200), testFile2: bytes.Repeat([]byte("b"), 200),
} }
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files) createFilesOnDisk(t, fs, files)
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err) require.NoError(t, err)
@@ -232,9 +277,7 @@ func TestCheckWithProgress(t *testing.T) {
err = chk.Check(context.Background(), results, progress) err = chk.Check(context.Background(), results, progress)
require.NoError(t, err) require.NoError(t, err)
// Drain results // results is fully buffered and closed; no draining needed
for range results {
}
// Check progress was sent // Check progress was sent
var progressUpdates []CheckStatus var progressUpdates []CheckStatus
@@ -253,14 +296,17 @@ func TestCheckWithProgress(t *testing.T) {
} }
func TestCheckContextCancellation(t *testing.T) { func TestCheckContextCancellation(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create many files to ensure we have time to cancel // Create many files to ensure we have time to cancel
files := make(map[string][]byte) files := make(map[string][]byte)
for i := 0; i < 100; i++ { for i := range 100 {
files[string(rune('a'+i%26))+".txt"] = bytes.Repeat([]byte("x"), 1000) files[string(rune('a'+i%26))+".txt"] = bytes.Repeat([]byte("x"), 1000)
} }
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files) createFilesOnDisk(t, fs, files)
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err) require.NoError(t, err)
@@ -274,17 +320,19 @@ func TestCheckContextCancellation(t *testing.T) {
} }
func TestFindExtraFiles(t *testing.T) { func TestFindExtraFiles(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Manifest only contains file1 // Manifest only contains file1
manifestFiles := map[string][]byte{ manifestFiles := map[string][]byte{
"file1.txt": []byte("in manifest"), testFile1: []byte("in manifest"),
} }
createTestManifest(t, fs, "/manifest.mf", manifestFiles) createTestManifest(t, fs, "/manifest.mf", manifestFiles)
// Disk has file1 and file2 // Disk has file1 and file2
createFilesOnDisk(t, fs, "/data", map[string][]byte{ createFilesOnDisk(t, fs, map[string][]byte{
"file1.txt": []byte("in manifest"), testFile1: []byte("in manifest"),
"file2.txt": []byte("extra file"), testFile2: []byte("extra file"),
}) })
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
@@ -300,19 +348,21 @@ func TestFindExtraFiles(t *testing.T) {
} }
assert.Len(t, extras, 1) assert.Len(t, extras, 1)
assert.Equal(t, RelFilePath("file2.txt"), extras[0].Path) assert.Equal(t, RelFilePath(testFile2), extras[0].Path)
assert.Equal(t, StatusExtra, extras[0].Status) assert.Equal(t, StatusExtra, extras[0].Status)
assert.Equal(t, "not in manifest", extras[0].Message) assert.Equal(t, "not in manifest", extras[0].Message)
} }
func TestFindExtraFilesSkipsManifestAndDotfiles(t *testing.T) { func TestFindExtraFilesSkipsManifestAndDotfiles(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
manifestFiles := map[string][]byte{ manifestFiles := map[string][]byte{
"file1.txt": []byte("in manifest"), testFile1: []byte("in manifest"),
} }
createTestManifest(t, fs, "/data/.index.mf", manifestFiles) createTestManifest(t, fs, "/data/.index.mf", manifestFiles)
createFilesOnDisk(t, fs, "/data", map[string][]byte{ createFilesOnDisk(t, fs, map[string][]byte{
"file1.txt": []byte("in manifest"), testFile1: []byte("in manifest"),
}) })
// Create dotfile and manifest that should be skipped // Create dotfile and manifest that should be skipped
require.NoError(t, afero.WriteFile(fs, "/data/.hidden", []byte("hidden"), 0o644)) require.NoError(t, afero.WriteFile(fs, "/data/.hidden", []byte("hidden"), 0o644))
@@ -337,17 +387,21 @@ func TestFindExtraFilesSkipsManifestAndDotfiles(t *testing.T) {
for _, e := range extras { for _, e := range extras {
t.Logf("extra: %s", e.Path) t.Logf("extra: %s", e.Path)
} }
assert.Len(t, extras, 1) assert.Len(t, extras, 1)
if len(extras) > 0 { if len(extras) > 0 {
assert.Equal(t, RelFilePath("extra.txt"), extras[0].Path) assert.Equal(t, RelFilePath("extra.txt"), extras[0].Path)
} }
} }
func TestFindExtraFilesContextCancellation(t *testing.T) { func TestFindExtraFilesContextCancellation(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{"file.txt": []byte("data")} files := map[string][]byte{testFileName: []byte("data")}
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files) createFilesOnDisk(t, fs, files)
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err) require.NoError(t, err)
@@ -361,10 +415,12 @@ func TestFindExtraFilesContextCancellation(t *testing.T) {
} }
func TestCheckNilChannels(t *testing.T) { func TestCheckNilChannels(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{"file.txt": []byte("data")} files := map[string][]byte{testFileName: []byte("data")}
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files) createFilesOnDisk(t, fs, files)
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err) require.NoError(t, err)
@@ -375,10 +431,12 @@ func TestCheckNilChannels(t *testing.T) {
} }
func TestFindExtraFilesNilChannel(t *testing.T) { func TestFindExtraFilesNilChannel(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{"file.txt": []byte("data")} files := map[string][]byte{testFileName: []byte("data")}
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files) createFilesOnDisk(t, fs, files)
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err) require.NoError(t, err)
@@ -389,6 +447,8 @@ func TestFindExtraFilesNilChannel(t *testing.T) {
} }
func TestCheckSubdirectories(t *testing.T) { func TestCheckSubdirectories(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{ files := map[string][]byte{
"dir1/file1.txt": []byte("content1"), "dir1/file1.txt": []byte("content1"),
@@ -400,6 +460,7 @@ func TestCheckSubdirectories(t *testing.T) {
// Create files with full directory structure // Create files with full directory structure
for path, content := range files { for path, content := range files {
fullPath := "/data/" + path fullPath := "/data/" + path
require.NoError(t, fs.MkdirAll("/data/dir1/dir2/dir3", 0o755)) require.NoError(t, fs.MkdirAll("/data/dir1/dir2/dir3", 0o755))
require.NoError(t, afero.WriteFile(fs, fullPath, content, 0o644)) require.NoError(t, afero.WriteFile(fs, fullPath, content, 0o644))
} }
@@ -412,25 +473,30 @@ func TestCheckSubdirectories(t *testing.T) {
require.NoError(t, err) require.NoError(t, err)
var okCount int var okCount int
for r := range results { for r := range results {
assert.Equal(t, StatusOK, r.Status, "file %s should be OK", r.Path) assert.Equal(t, StatusOK, r.Status, "file %s should be OK", r.Path)
okCount++ okCount++
} }
assert.Equal(t, 3, okCount) assert.Equal(t, 3, okCount)
} }
func TestCheckMissingFileDetectedWithoutFallback(t *testing.T) { func TestCheckMissingFileDetectedWithoutFallback(t *testing.T) {
t.Parallel()
// Regression test: errors.Is(err, errors.New("...")) never matches because // Regression test: errors.Is(err, errors.New("...")) never matches because
// errors.New creates a new value each time. The fix uses os.ErrNotExist instead. // errors.New creates a new value each time. The fix uses os.ErrNotExist instead.
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
files := map[string][]byte{ files := map[string][]byte{
"exists.txt": []byte("here"), testExistsFile: []byte("here"),
"missing.txt": []byte("not on disk"), "missing.txt": []byte("not on disk"),
} }
createTestManifest(t, fs, "/manifest.mf", files) createTestManifest(t, fs, "/manifest.mf", files)
// Only create one file on disk // Only create one file on disk
createFilesOnDisk(t, fs, "/data", map[string][]byte{ createFilesOnDisk(t, fs, map[string][]byte{
"exists.txt": []byte("here"), testExistsFile: []byte("here"),
}) })
chk, err := NewChecker("/manifest.mf", "/data", fs) chk, err := NewChecker("/manifest.mf", "/data", fs)
@@ -447,12 +513,78 @@ func TestCheckMissingFileDetectedWithoutFallback(t *testing.T) {
assert.Equal(t, RelFilePath("missing.txt"), r.Path) assert.Equal(t, RelFilePath("missing.txt"), r.Path)
} }
} }
assert.Equal(t, 1, statusCounts[StatusOK], "one file should be OK") assert.Equal(t, 1, statusCounts[StatusOK], "one file should be OK")
assert.Equal(t, 1, statusCounts[StatusMissing], "one file should be MISSING") assert.Equal(t, 1, statusCounts[StatusMissing], "one file should be MISSING")
assert.Equal(t, 0, statusCounts[StatusError], "no files should be ERROR") assert.Equal(t, 0, statusCounts[StatusError], "no files should be ERROR")
} }
func TestFindExtraFilesSkipsDotfiles(t *testing.T) {
t.Parallel()
// Regression test for #16: FindExtraFiles should not report dotfiles
// or the manifest file itself as extra files.
fs := afero.NewMemMapFs()
files := map[string][]byte{
testFile1: []byte("in manifest"),
}
createTestManifest(t, fs, "/data/.index.mf", files)
createFilesOnDisk(t, fs, files)
// Add dotfiles and manifest file on disk
require.NoError(t, afero.WriteFile(fs, "/data/.hidden", []byte("dotfile"), 0o644))
require.NoError(t, fs.MkdirAll("/data/.git", 0o755))
require.NoError(t,
afero.WriteFile(fs, "/data/.git/config", []byte("git config"), 0o644))
chk, err := NewChecker("/data/.index.mf", "/data", fs)
require.NoError(t, err)
results := make(chan Result, 10)
err = chk.FindExtraFiles(context.Background(), results)
require.NoError(t, err)
var extras []Result
for r := range results {
extras = append(extras, r)
}
// Should report NO extra files — dotfiles and manifest should be skipped
assert.Empty(t, extras,
"FindExtraFiles should not report dotfiles or manifest file as extra; got: %v",
extras)
}
func TestFindExtraFilesSkipsManifestFile(t *testing.T) {
t.Parallel()
// The manifest file itself should never be reported as extra
fs := afero.NewMemMapFs()
files := map[string][]byte{
testFile1: []byte("content"),
}
createTestManifest(t, fs, "/data/index.mf", files)
createFilesOnDisk(t, fs, files)
chk, err := NewChecker("/data/index.mf", "/data", fs)
require.NoError(t, err)
results := make(chan Result, 10)
err = chk.FindExtraFiles(context.Background(), results)
require.NoError(t, err)
var extras []Result
for r := range results {
extras = append(extras, r)
}
assert.Empty(t, extras,
"manifest file should not be reported as extra; got: %v", extras)
}
func TestCheckEmptyManifest(t *testing.T) { func TestCheckEmptyManifest(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
// Create manifest with no files // Create manifest with no files
createTestManifest(t, fs, "/manifest.mf", map[string][]byte{}) createTestManifest(t, fs, "/manifest.mf", map[string][]byte{})
@@ -471,5 +603,47 @@ func TestCheckEmptyManifest(t *testing.T) {
for range results { for range results {
count++ count++
} }
assert.Equal(t, 0, count) assert.Equal(t, 0, count)
} }
func TestCheckProgressRateLimited(t *testing.T) {
t.Parallel()
// Create many small files - progress should be rate-limited, not one per file.
// With rate-limiting to once per second, we should get far fewer progress
// updates than files (plus one final update).
fs := afero.NewMemMapFs()
files := make(map[string][]byte, 100)
for i := range 100 {
name := fmt.Sprintf("file%03d.txt", i)
files[name] = []byte("content")
}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
require.NoError(t, err)
results := make(chan Result, 200)
progress := make(chan CheckStatus, 200)
err = chk.Check(context.Background(), results, progress)
require.NoError(t, err)
// results is fully buffered and closed; no draining needed
// Count progress updates
var progressCount int
for range progress {
progressCount++
}
// Should be far fewer than 100 (rate-limited to once per second)
// At minimum we get the final update
assert.GreaterOrEqual(t, progressCount, 1,
"should get at least the final progress update")
assert.Less(t, progressCount, 100,
"progress should be rate-limited, not one per file")
}

View File

@@ -1,11 +1,17 @@
package mfer package mfer
const ( const (
Version = "0.1.0" // Version is the current mfer release version.
Version = "0.1.0"
// ReleaseDate is the date on which Version was released.
ReleaseDate = "2025-12-17" ReleaseDate = "2025-12-17"
// MaxDecompressedSize is the maximum allowed size of decompressed manifest // MaxDecompressedSize is the maximum allowed size of decompressed manifest
// data (256 MB). This prevents decompression bombs from consuming excessive // data (256 MB). This prevents decompression bombs from consuming excessive
// memory. // memory.
MaxDecompressedSize int64 = 256 * 1024 * 1024 MaxDecompressedSize int64 = 256 * 1024 * 1024
// uuidLength is the length in bytes of a binary UUID.
uuidLength = 16
) )

View File

@@ -15,105 +15,174 @@ import (
"sneak.berlin/go/mfer/internal/log" "sneak.berlin/go/mfer/internal/log"
) )
var (
errInvalidUUIDLength = errors.New("invalid UUID length")
errInvalidUUIDFormat = errors.New("invalid UUID format")
errUnknownVersion = errors.New("unknown version")
errUnknownCompression = errors.New("unknown compression type")
errCompressedHashWrong = errors.New("compressed data hash mismatch")
errSignatureNoPubKey = errors.New("signature present but no public key")
errDecompressedTooLarge = errors.New("decompressed data exceeds maximum allowed size")
errUUIDMismatch = errors.New("outer and inner UUID mismatch")
errInvalidFileFormat = errors.New("invalid file format")
)
// validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable). // validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable).
func validateUUID(data []byte) error { func validateUUID(data []byte) error {
if len(data) != 16 { if len(data) != uuidLength {
return errors.New("invalid UUID length") return errInvalidUUIDLength
} }
// Try to parse as UUID to validate format // Try to parse as UUID to validate format
_, err := uuid.FromBytes(data) _, err := uuid.FromBytes(data)
if err != nil { if err != nil {
return errors.New("invalid UUID format") return errInvalidUUIDFormat
} }
return nil return nil
} }
func (m *manifest) deserializeInner() error { // validateOuterHeader checks the outer message's version, compression
if m.pbOuter.Version != MFFileOuter_VERSION_ONE { // type, and UUID.
return errors.New("unknown version") func (m *manifest) validateOuterHeader() error {
if m.pbOuter.GetVersion() != MFFileOuter_VERSION_ONE {
return errUnknownVersion
} }
if m.pbOuter.CompressionType != MFFileOuter_COMPRESSION_ZSTD {
return errors.New("unknown compression type") if m.pbOuter.GetCompressionType() != MFFileOuter_COMPRESSION_ZSTD {
return errUnknownCompression
} }
// Validate outer UUID before any decompression // Validate outer UUID before any decompression
if err := validateUUID(m.pbOuter.Uuid); err != nil { err := validateUUID(m.pbOuter.GetUuid())
return errors.New("outer UUID invalid: " + err.Error()) if err != nil {
return fmt.Errorf("outer UUID invalid: %w", err)
} }
// Verify hash of compressed data before decompression return nil
}
// verifyOuterIntegrity checks the hash of the compressed payload and,
// if a signature is present, verifies it against the embedded public key.
func (m *manifest) verifyOuterIntegrity() error {
h := sha256.New() h := sha256.New()
if _, err := h.Write(m.pbOuter.InnerMessage); err != nil {
return err _, err := h.Write(m.pbOuter.GetInnerMessage())
if err != nil {
return fmt.Errorf("deserialize: hash write: %w", err)
} }
sha256Hash := h.Sum(nil) sha256Hash := h.Sum(nil)
if !bytes.Equal(sha256Hash, m.pbOuter.Sha256) { if !bytes.Equal(sha256Hash, m.pbOuter.GetSha256()) {
return errors.New("compressed data hash mismatch") return errCompressedHashWrong
} }
// Verify signature if present if len(m.pbOuter.GetSignature()) == 0 {
if len(m.pbOuter.Signature) > 0 { return nil
if len(m.pbOuter.SigningPubKey) == 0 {
return errors.New("signature present but no public key")
}
sigString, err := m.signatureString()
if err != nil {
return fmt.Errorf("failed to generate signature string for verification: %w", err)
}
if err := gpgVerify([]byte(sigString), m.pbOuter.Signature, m.pbOuter.SigningPubKey); err != nil {
return fmt.Errorf("signature verification failed: %w", err)
}
log.Infof("signature verified successfully")
} }
bb := bytes.NewBuffer(m.pbOuter.InnerMessage) if len(m.pbOuter.GetSigningPubKey()) == 0 {
return errSignatureNoPubKey
}
sigString, err := m.signatureString()
if err != nil {
return fmt.Errorf(
"failed to generate signature string for verification: %w", err,
)
}
err = gpgVerify(
[]byte(sigString),
m.pbOuter.GetSignature(),
m.pbOuter.GetSigningPubKey(),
)
if err != nil {
return fmt.Errorf("signature verification failed: %w", err)
}
log.Infof("signature verified successfully")
return nil
}
// decompressInner decompresses the inner payload, enforcing size limits
// to prevent decompression bombs.
func (m *manifest) decompressInner() ([]byte, error) {
bb := bytes.NewBuffer(m.pbOuter.GetInnerMessage())
zr, err := zstd.NewReader(bb) zr, err := zstd.NewReader(bb)
if err != nil { if err != nil {
return err return nil, fmt.Errorf("deserialize: zstd reader: %w", err)
} }
defer zr.Close() defer zr.Close()
// Limit decompressed size to prevent decompression bombs. // Limit decompressed size to prevent decompression bombs.
// Use declared size + 1 byte to detect overflow, capped at MaxDecompressedSize. // Use declared size + 1 byte to detect overflow, capped at MaxDecompressedSize.
maxSize := MaxDecompressedSize maxSize := MaxDecompressedSize
if m.pbOuter.Size > 0 && m.pbOuter.Size < int64(maxSize) { if m.pbOuter.GetSize() > 0 && m.pbOuter.GetSize() < maxSize {
maxSize = int64(m.pbOuter.Size) + 1 maxSize = m.pbOuter.GetSize() + 1
} }
limitedReader := io.LimitReader(zr, maxSize) limitedReader := io.LimitReader(zr, maxSize)
dat, err := io.ReadAll(limitedReader) dat, err := io.ReadAll(limitedReader)
if err != nil {
return nil, fmt.Errorf("deserialize: decompress: %w", err)
}
if int64(len(dat)) >= MaxDecompressedSize {
return nil, fmt.Errorf(
"%w of %d bytes", errDecompressedTooLarge, MaxDecompressedSize,
)
}
return dat, nil
}
func (m *manifest) deserializeInner() error {
err := m.validateOuterHeader()
if err != nil { if err != nil {
return err return err
} }
if int64(len(dat)) >= MaxDecompressedSize {
return fmt.Errorf("decompressed data exceeds maximum allowed size of %d bytes", MaxDecompressedSize) err = m.verifyOuterIntegrity()
if err != nil {
return err
}
dat, err := m.decompressInner()
if err != nil {
return err
} }
isize := len(dat) isize := len(dat)
if int64(isize) != m.pbOuter.Size { if int64(isize) != m.pbOuter.GetSize() {
log.Debugf("truncated data, got %d expected %d", isize, m.pbOuter.Size) log.Debugf("truncated data, got %d expected %d", isize, m.pbOuter.GetSize())
return bork.ErrFileTruncated return bork.ErrFileTruncated
} }
// Deserialize inner message // Deserialize inner message
m.pbInner = new(MFFile) m.pbInner = new(MFFile)
if err := proto.Unmarshal(dat, m.pbInner); err != nil {
return err err = proto.Unmarshal(dat, m.pbInner)
if err != nil {
return fmt.Errorf("deserialize: unmarshal inner: %w", err)
} }
// Validate inner UUID // Validate inner UUID
if err := validateUUID(m.pbInner.Uuid); err != nil { err = validateUUID(m.pbInner.GetUuid())
return errors.New("inner UUID invalid: " + err.Error()) if err != nil {
return fmt.Errorf("inner UUID invalid: %w", err)
} }
// Verify UUIDs match // Verify UUIDs match
if !bytes.Equal(m.pbOuter.Uuid, m.pbInner.Uuid) { if !bytes.Equal(m.pbOuter.GetUuid(), m.pbInner.GetUuid()) {
return errors.New("outer and inner UUID mismatch") return errUUIDMismatch
} }
log.Infof("loaded manifest with %d files", len(m.pbInner.Files)) log.Infof("loaded manifest with %d files", len(m.pbInner.GetFiles()))
return nil return nil
} }
@@ -122,20 +191,26 @@ func validateMagic(dat []byte) bool {
if len(dat) < ml { if len(dat) < ml {
return false return false
} }
got := dat[0:ml] got := dat[0:ml]
expected := []byte(MAGIC) expected := []byte(MAGIC)
return bytes.Equal(got, expected) return bytes.Equal(got, expected)
} }
// NewManifestFromReader reads a manifest from an io.Reader. // NewManifestFromReader reads a manifest from an io.Reader.
//
//nolint:revive // unexported-return: exporting manifest is owner question 13
func NewManifestFromReader(input io.Reader) (*manifest, error) { func NewManifestFromReader(input io.Reader) (*manifest, error) {
m := &manifest{} m := &manifest{}
dat, err := io.ReadAll(input) dat, err := io.ReadAll(input)
if err != nil { if err != nil {
return nil, err return nil, err
} }
if !validateMagic(dat) { if !validateMagic(dat) {
return nil, errors.New("invalid file format") return nil, errInvalidFileFormat
} }
// remove magic bytes prefix: // remove magic bytes prefix:
@@ -145,12 +220,15 @@ func NewManifestFromReader(input io.Reader) (*manifest, error) {
// deserialize outer: // deserialize outer:
m.pbOuter = new(MFFileOuter) m.pbOuter = new(MFFileOuter)
if err := proto.Unmarshal(dat, m.pbOuter); err != nil {
err = proto.Unmarshal(dat, m.pbOuter)
if err != nil {
return nil, err return nil, err
} }
// deserialize inner: // deserialize inner:
if err := m.deserializeInner(); err != nil { err = m.deserializeInner()
if err != nil {
return nil, err return nil, err
} }
@@ -159,14 +237,19 @@ func NewManifestFromReader(input io.Reader) (*manifest, error) {
// NewManifestFromFile reads a manifest from a file path using the given filesystem. // NewManifestFromFile reads a manifest from a file path using the given filesystem.
// If fs is nil, the real filesystem (OsFs) is used. // If fs is nil, the real filesystem (OsFs) is used.
//
//nolint:revive // unexported-return: exporting manifest is owner question 13
func NewManifestFromFile(fs afero.Fs, path string) (*manifest, error) { func NewManifestFromFile(fs afero.Fs, path string) (*manifest, error) {
if fs == nil { if fs == nil {
fs = afero.NewOsFs() fs = afero.NewOsFs()
} }
f, err := fs.Open(path) f, err := fs.Open(path)
if err != nil { if err != nil {
return nil, err return nil, err
} }
defer func() { _ = f.Close() }() defer func() { _ = f.Close() }()
return NewManifestFromReader(f) return NewManifestFromReader(f)
} }

85
mfer/errmsg_test.go Normal file
View File

@@ -0,0 +1,85 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// TestValidatePathMessagesVerbatim pins the exact rendered text of every
// ValidatePath rejection.
//
// These strings are user-visible and are assembled by wrapping static
// sentinels mid-sentence, which makes them easy to reword by accident
// while refactoring for errors.Is matchability. Changing one is a
// deliberate change, not a refactoring side effect.
func TestValidatePathMessagesVerbatim(t *testing.T) {
t.Parallel()
for _, tc := range []struct {
name string
path string
want string
is error
}{
{
name: "empty",
path: "",
want: "path cannot be empty",
is: errPathEmpty,
},
{
name: "not utf8",
path: "a\xffb",
want: `path "a\xffb" is not valid UTF-8`,
is: errPathNotUTF8,
},
{
name: "backslash",
path: `a\b`,
want: `path "a\\b" contains backslash; ` +
"use forward slashes only",
is: errPathBackslash,
},
{
name: "absolute",
path: "/a/b",
want: `path "/a/b" is absolute; must be relative`,
is: errPathAbsolute,
},
{
name: "empty segment",
path: "a//b",
want: `path "a//b" contains empty segment`,
is: errPathEmptySegment,
},
{
name: "dotdot segment",
path: "a/../b",
want: `path "a/../b" contains '..' segment`,
is: errPathDotDot,
},
} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
err := ValidatePath(tc.path)
require.Error(t, err)
assert.Equal(t, tc.want, err.Error())
require.ErrorIs(t, err, tc.is)
})
}
}
// TestSerializeInternalErrorMessagesVerbatim pins the two distinct
// "internal error" messages, which differ between generate and
// generateOuter and have always done so.
func TestSerializeInternalErrorMessagesVerbatim(t *testing.T) {
t.Parallel()
m := &manifest{}
require.EqualError(t, m.generate(), "internal error: pbInner not set")
require.EqualError(t, m.generateOuter(), "internal error")
}

View File

@@ -2,13 +2,45 @@ package mfer
import ( import (
"bytes" "bytes"
"context"
"errors"
"fmt" "fmt"
"io"
"os" "os"
"os/exec" "os/exec"
"path/filepath" "path/filepath"
"strings" "strings"
) )
const (
// privateDirPerms is the permission mode for temporary GPG home
// directories.
privateDirPerms os.FileMode = 0o700
// privateFilePerms is the permission mode for temporary key,
// signature, and data files.
privateFilePerms os.FileMode = 0o600
// gpgFingerprintField is the record type tag for fingerprint lines
// in gpg --with-colons output.
gpgFingerprintField = "fpr"
// gpgFingerprintMinFields is the minimum number of colon-separated
// fields in a gpg fingerprint record (the fingerprint is field 10).
gpgFingerprintMinFields = 10
// gpg option names used from more than one call site.
gpgOptArmor = "--armor"
gpgOptHomedir = "--homedir"
gpgOptVerify = "--verify"
)
var (
errGPGKeyNotFound = errors.New("gpg key not found")
errFingerprintNotFound = errors.New("fingerprint not found for key")
errImportedFPRNotFound = errors.New("fingerprint not found in imported key")
)
// GPGKeyID represents a GPG key identifier (fingerprint or key ID). // GPGKeyID represents a GPG key identifier (fingerprint or key ID).
type GPGKeyID string type GPGKeyID string
@@ -17,22 +49,69 @@ type SigningOptions struct {
KeyID GPGKeyID KeyID GPGKeyID
} }
// gpgSign creates a detached signature of the data using the specified key. // gpgArgs builds a gpg argument list from opts followed by positional
// Returns the armored detached signature. // arguments, separated by an explicit "--" end-of-options marker.
func gpgSign(data []byte, keyID GPGKeyID) ([]byte, error) { //
cmd := exec.Command("gpg", // This matters because key IDs reach gpg as bare positional arguments
"--detach-sign", // (from --sign-key / MFER_SIGN_KEY) and gpg would otherwise parse a value
"--armor", // beginning with "-" as one of its own options. Callers must route every
"--local-user", string(keyID), // non-option argument through here.
) func gpgArgs(opts []string, positional ...string) []string {
args := make([]string, 0, len(opts)+1+len(positional))
args = append(args, opts...)
args = append(args, "--")
args = append(args, positional...)
cmd.Stdin = bytes.NewReader(data) return args
}
// runGPG runs the gpg binary in batch mode with the given arguments and
// optional stdin, returning captured stdout and stderr.
func runGPG(stdin io.Reader, args ...string) (*bytes.Buffer, *bytes.Buffer, error) {
fullArgs := append([]string{"--batch", "--no-tty"}, args...)
// G204: the executable name is a compile-time constant. The arguments
// are not, so the guarantee that matters is placement: every
// caller-supplied value is passed either as the value of a named
// option or after the "--" end-of-options marker inserted by gpgArgs,
// and therefore cannot be reinterpreted by gpg as an option.
cmd := exec.CommandContext( //nolint:gosec // G204: see comment above
context.Background(), "gpg", fullArgs...)
cmd.Stdin = stdin
var stdout, stderr bytes.Buffer var stdout, stderr bytes.Buffer
cmd.Stdout = &stdout cmd.Stdout = &stdout
cmd.Stderr = &stderr cmd.Stderr = &stderr
if err := cmd.Run(); err != nil { err := cmd.Run()
return &stdout, &stderr, err
}
// parseFingerprint extracts the first fingerprint from gpg --with-colons
// output, or returns ok=false if none is present.
func parseFingerprint(colonOutput string) (string, bool) {
for _, line := range strings.Split(colonOutput, "\n") {
fields := strings.Split(line, ":")
if len(fields) >= gpgFingerprintMinFields &&
fields[0] == gpgFingerprintField {
return fields[9], true
}
}
return "", false
}
// gpgSign creates a detached signature of the data using the specified key.
// Returns the armored detached signature.
func gpgSign(data []byte, keyID GPGKeyID) ([]byte, error) {
stdout, stderr, err := runGPG(bytes.NewReader(data),
"--detach-sign",
gpgOptArmor,
"--local-user", string(keyID),
)
if err != nil {
return nil, fmt.Errorf("gpg sign failed: %w: %s", err, stderr.String()) return nil, fmt.Errorf("gpg sign failed: %w: %s", err, stderr.String())
} }
@@ -42,22 +121,15 @@ func gpgSign(data []byte, keyID GPGKeyID) ([]byte, error) {
// gpgExportPublicKey exports the public key for the specified key ID. // gpgExportPublicKey exports the public key for the specified key ID.
// Returns the armored public key. // Returns the armored public key.
func gpgExportPublicKey(keyID GPGKeyID) ([]byte, error) { func gpgExportPublicKey(keyID GPGKeyID) ([]byte, error) {
cmd := exec.Command("gpg", stdout, stderr, err := runGPG(nil,
"--export", gpgArgs([]string{"--export", gpgOptArmor}, string(keyID))...,
"--armor",
string(keyID),
) )
if err != nil {
var stdout, stderr bytes.Buffer
cmd.Stdout = &stdout
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
return nil, fmt.Errorf("gpg export failed: %w: %s", err, stderr.String()) return nil, fmt.Errorf("gpg export failed: %w: %s", err, stderr.String())
} }
if stdout.Len() == 0 { if stdout.Len() == 0 {
return nil, fmt.Errorf("gpg key not found: %s", keyID) return nil, fmt.Errorf("%w: %s", errGPGKeyNotFound, keyID)
} }
return stdout.Bytes(), nil return stdout.Bytes(), nil
@@ -65,30 +137,21 @@ func gpgExportPublicKey(keyID GPGKeyID) ([]byte, error) {
// gpgGetKeyFingerprint gets the full fingerprint for a key ID. // gpgGetKeyFingerprint gets the full fingerprint for a key ID.
func gpgGetKeyFingerprint(keyID GPGKeyID) ([]byte, error) { func gpgGetKeyFingerprint(keyID GPGKeyID) ([]byte, error) {
cmd := exec.Command("gpg", stdout, stderr, err := runGPG(nil,
"--with-colons", gpgArgs([]string{"--with-colons", "--fingerprint"}, string(keyID))...,
"--fingerprint",
string(keyID),
) )
if err != nil {
var stdout, stderr bytes.Buffer return nil, fmt.Errorf(
cmd.Stdout = &stdout "gpg fingerprint lookup failed: %w: %s", err, stderr.String(),
cmd.Stderr = &stderr )
if err := cmd.Run(); err != nil {
return nil, fmt.Errorf("gpg fingerprint lookup failed: %w: %s", err, stderr.String())
} }
// Parse the colon-delimited output to find the fingerprint fpr, ok := parseFingerprint(stdout.String())
lines := strings.Split(stdout.String(), "\n") if !ok {
for _, line := range lines { return nil, fmt.Errorf("%w: %s", errFingerprintNotFound, keyID)
fields := strings.Split(line, ":")
if len(fields) >= 10 && fields[0] == "fpr" {
return []byte(fields[9]), nil
}
} }
return nil, fmt.Errorf("fingerprint not found for key: %s", keyID) return []byte(fpr), nil
} }
// gpgExtractPubKeyFingerprint imports a public key into a temporary keyring // gpgExtractPubKeyFingerprint imports a public key into a temporary keyring
@@ -100,54 +163,51 @@ func gpgExtractPubKeyFingerprint(pubKey []byte) (string, error) {
if err != nil { if err != nil {
return "", fmt.Errorf("failed to create temp dir: %w", err) return "", fmt.Errorf("failed to create temp dir: %w", err)
} }
defer func() { _ = os.RemoveAll(tmpDir) }() defer func() { _ = os.RemoveAll(tmpDir) }()
// Set restrictive permissions // Set restrictive permissions
if err := os.Chmod(tmpDir, 0o700); err != nil { err = os.Chmod(tmpDir, privateDirPerms)
if err != nil {
return "", fmt.Errorf("failed to set temp dir permissions: %w", err) return "", fmt.Errorf("failed to set temp dir permissions: %w", err)
} }
// Write public key to temp file // Write public key to temp file
pubKeyFile := filepath.Join(tmpDir, "pubkey.asc") pubKeyFile := filepath.Join(tmpDir, "pubkey.asc")
if err := os.WriteFile(pubKeyFile, pubKey, 0o600); err != nil {
err = os.WriteFile(pubKeyFile, pubKey, privateFilePerms)
if err != nil {
return "", fmt.Errorf("failed to write public key: %w", err) return "", fmt.Errorf("failed to write public key: %w", err)
} }
// Import the public key into the temporary keyring // Import the public key into the temporary keyring
importCmd := exec.Command("gpg", _, importStderr, err := runGPG(nil,
"--homedir", tmpDir, gpgArgs([]string{gpgOptHomedir, tmpDir, "--import"}, pubKeyFile)...,
"--import",
pubKeyFile,
) )
var importStderr bytes.Buffer if err != nil {
importCmd.Stderr = &importStderr return "", fmt.Errorf(
if err := importCmd.Run(); err != nil { "failed to import public key: %w: %s", err, importStderr.String(),
return "", fmt.Errorf("failed to import public key: %w: %s", err, importStderr.String()) )
} }
// List keys to get fingerprint // List keys to get fingerprint
listCmd := exec.Command("gpg", listStdout, listStderr, err := runGPG(nil,
"--homedir", tmpDir, "--homedir", tmpDir,
"--with-colons", "--with-colons",
"--fingerprint", "--fingerprint",
) )
var listStdout, listStderr bytes.Buffer if err != nil {
listCmd.Stdout = &listStdout return "", fmt.Errorf(
listCmd.Stderr = &listStderr "failed to list keys: %w: %s", err, listStderr.String(),
if err := listCmd.Run(); err != nil { )
return "", fmt.Errorf("failed to list keys: %w: %s", err, listStderr.String())
} }
// Parse the colon-delimited output to find the fingerprint fpr, ok := parseFingerprint(listStdout.String())
lines := strings.Split(listStdout.String(), "\n") if !ok {
for _, line := range lines { return "", errImportedFPRNotFound
fields := strings.Split(line, ":")
if len(fields) >= 10 && fields[0] == "fpr" {
return fields[9], nil
}
} }
return "", fmt.Errorf("fingerprint not found in imported key") return fpr, nil
} }
// gpgVerify verifies a detached signature against data using the provided public key. // gpgVerify verifies a detached signature against data using the provided public key.
@@ -158,54 +218,58 @@ func gpgVerify(data, signature, pubKey []byte) error {
if err != nil { if err != nil {
return fmt.Errorf("failed to create temp dir: %w", err) return fmt.Errorf("failed to create temp dir: %w", err)
} }
defer func() { _ = os.RemoveAll(tmpDir) }() defer func() { _ = os.RemoveAll(tmpDir) }()
// Set restrictive permissions // Set restrictive permissions
if err := os.Chmod(tmpDir, 0o700); err != nil { err = os.Chmod(tmpDir, privateDirPerms)
if err != nil {
return fmt.Errorf("failed to set temp dir permissions: %w", err) return fmt.Errorf("failed to set temp dir permissions: %w", err)
} }
// Write public key to temp file // Write public key to temp file
pubKeyFile := filepath.Join(tmpDir, "pubkey.asc") pubKeyFile := filepath.Join(tmpDir, "pubkey.asc")
if err := os.WriteFile(pubKeyFile, pubKey, 0o600); err != nil {
err = os.WriteFile(pubKeyFile, pubKey, privateFilePerms)
if err != nil {
return fmt.Errorf("failed to write public key: %w", err) return fmt.Errorf("failed to write public key: %w", err)
} }
// Write signature to temp file // Write signature to temp file
sigFile := filepath.Join(tmpDir, "signature.asc") sigFile := filepath.Join(tmpDir, "signature.asc")
if err := os.WriteFile(sigFile, signature, 0o600); err != nil {
err = os.WriteFile(sigFile, signature, privateFilePerms)
if err != nil {
return fmt.Errorf("failed to write signature: %w", err) return fmt.Errorf("failed to write signature: %w", err)
} }
// Write data to temp file // Write data to temp file
dataFile := filepath.Join(tmpDir, "data") dataFile := filepath.Join(tmpDir, "data")
if err := os.WriteFile(dataFile, data, 0o600); err != nil {
err = os.WriteFile(dataFile, data, privateFilePerms)
if err != nil {
return fmt.Errorf("failed to write data: %w", err) return fmt.Errorf("failed to write data: %w", err)
} }
// Import the public key into the temporary keyring // Import the public key into the temporary keyring
importCmd := exec.Command("gpg", _, importStderr, err := runGPG(nil,
"--homedir", tmpDir, gpgArgs([]string{gpgOptHomedir, tmpDir, "--import"}, pubKeyFile)...,
"--import",
pubKeyFile,
) )
var importStderr bytes.Buffer if err != nil {
importCmd.Stderr = &importStderr return fmt.Errorf(
if err := importCmd.Run(); err != nil { "failed to import public key: %w: %s", err, importStderr.String(),
return fmt.Errorf("failed to import public key: %w: %s", err, importStderr.String()) )
} }
// Verify the signature // Verify the signature
verifyCmd := exec.Command("gpg", _, verifyStderr, err := runGPG(nil,
"--homedir", tmpDir, gpgArgs([]string{gpgOptHomedir, tmpDir, gpgOptVerify},
"--verify", sigFile, dataFile)...,
sigFile,
dataFile,
) )
var verifyStderr bytes.Buffer if err != nil {
verifyCmd.Stderr = &verifyStderr return fmt.Errorf(
if err := verifyCmd.Run(); err != nil { "signature verification failed: %w: %s", err, verifyStderr.String(),
return fmt.Errorf("signature verification failed: %w: %s", err, verifyStderr.String()) )
} }
return nil return nil

View File

@@ -1,3 +1,4 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer package mfer
import ( import (
@@ -15,35 +16,20 @@ import (
) )
// testGPGEnv sets up a temporary GPG home directory with a test key. // testGPGEnv sets up a temporary GPG home directory with a test key.
// Returns the key ID and a cleanup function. // Returns the key ID and the GPG home directory; callers must point
func testGPGEnv(t *testing.T) (GPGKeyID, func()) { // GNUPGHOME at the returned directory (via t.Setenv) before using the
// gpg helpers under test.
func testGPGEnv(t *testing.T) (GPGKeyID, string) {
t.Helper() t.Helper()
// Check if gpg is installed // Check if gpg is installed
if _, err := exec.LookPath("gpg"); err != nil { _, err := exec.LookPath("gpg")
if err != nil {
t.Skip("gpg not installed, skipping signing test") t.Skip("gpg not installed, skipping signing test")
return "", func() {}
} }
// Create temporary GPG home directory // Create temporary GPG home directory (0700 by default)
gpgHome, err := os.MkdirTemp("", "mfer-gpg-test-*") gpgHome := t.TempDir()
require.NoError(t, err)
// Set restrictive permissions on GPG home
require.NoError(t, os.Chmod(gpgHome, 0o700))
// Save original GNUPGHOME and set new one
origGPGHome := os.Getenv("GNUPGHOME")
require.NoError(t, os.Setenv("GNUPGHOME", gpgHome))
cleanup := func() {
if origGPGHome == "" {
_ = os.Unsetenv("GNUPGHOME")
} else {
_ = os.Setenv("GNUPGHOME", origGPGHome)
}
_ = os.RemoveAll(gpgHome)
}
// Generate a test key with no passphrase // Generate a test key with no passphrase
keyParams := `%no-protection keyParams := `%no-protection
@@ -57,45 +43,51 @@ Expire-Date: 0
paramsFile := filepath.Join(gpgHome, "key-params") paramsFile := filepath.Join(gpgHome, "key-params")
require.NoError(t, os.WriteFile(paramsFile, []byte(keyParams), 0o600)) require.NoError(t, os.WriteFile(paramsFile, []byte(keyParams), 0o600))
cmd := exec.Command("gpg", "--batch", "--gen-key", paramsFile) //nolint:gosec // paramsFile is a test-controlled path inside t.TempDir()
cmd := exec.CommandContext(context.Background(), "gpg",
"--batch", "--gen-key", paramsFile)
cmd.Env = append(os.Environ(), "GNUPGHOME="+gpgHome) cmd.Env = append(os.Environ(), "GNUPGHOME="+gpgHome)
output, err := cmd.CombinedOutput() output, err := cmd.CombinedOutput()
if err != nil { if err != nil {
cleanup()
t.Skipf("failed to generate test GPG key: %v: %s", err, output) t.Skipf("failed to generate test GPG key: %v: %s", err, output)
return "", func() {}
} }
// Get the key fingerprint // Get the key fingerprint
cmd = exec.Command("gpg", "--list-keys", "--with-colons", "test@mfer.test") cmd = exec.CommandContext(context.Background(), "gpg",
"--list-keys", "--with-colons", "test@mfer.test")
cmd.Env = append(os.Environ(), "GNUPGHOME="+gpgHome) cmd.Env = append(os.Environ(), "GNUPGHOME="+gpgHome)
output, err = cmd.Output() output, err = cmd.Output()
if err != nil { if err != nil {
cleanup()
t.Fatalf("failed to list test key: %v", err) t.Fatalf("failed to list test key: %v", err)
} }
// Parse fingerprint from output // Parse fingerprint from output
var keyID string var keyID string
for _, line := range strings.Split(string(output), "\n") { for _, line := range strings.Split(string(output), "\n") {
fields := strings.Split(line, ":") fields := strings.Split(line, ":")
if len(fields) >= 10 && fields[0] == "fpr" { if len(fields) >= gpgFingerprintMinFields &&
fields[0] == gpgFingerprintField {
keyID = fields[9] keyID = fields[9]
break break
} }
} }
if keyID == "" { if keyID == "" {
cleanup()
t.Fatal("failed to find test key fingerprint") t.Fatal("failed to find test key fingerprint")
} }
return GPGKeyID(keyID), cleanup return GPGKeyID(keyID), gpgHome
} }
func TestGPGSign(t *testing.T) { func TestGPGSign(t *testing.T) {
keyID, cleanup := testGPGEnv(t) keyID, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
data := []byte("test data to sign") data := []byte("test data to sign")
sig, err := gpgSign(data, keyID) sig, err := gpgSign(data, keyID)
@@ -106,8 +98,8 @@ func TestGPGSign(t *testing.T) {
} }
func TestGPGExportPublicKey(t *testing.T) { func TestGPGExportPublicKey(t *testing.T) {
keyID, cleanup := testGPGEnv(t) keyID, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
pubKey, err := gpgExportPublicKey(keyID) pubKey, err := gpgExportPublicKey(keyID)
require.NoError(t, err) require.NoError(t, err)
@@ -117,8 +109,8 @@ func TestGPGExportPublicKey(t *testing.T) {
} }
func TestGPGGetKeyFingerprint(t *testing.T) { func TestGPGGetKeyFingerprint(t *testing.T) {
keyID, cleanup := testGPGEnv(t) keyID, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
fingerprint, err := gpgGetKeyFingerprint(keyID) fingerprint, err := gpgGetKeyFingerprint(keyID)
require.NoError(t, err) require.NoError(t, err)
@@ -127,10 +119,47 @@ func TestGPGGetKeyFingerprint(t *testing.T) {
assert.Len(t, fingerprint, 40, "fingerprint should be 40 hex chars") assert.Len(t, fingerprint, 40, "fingerprint should be 40 hex chars")
} }
// TestGPGArgsSeparatesPositionals pins that caller-supplied values are
// placed after an end-of-options marker. Key IDs arrive from --sign-key
// and MFER_SIGN_KEY as bare positional arguments, so without the marker
// a value beginning with "-" would be parsed by gpg as one of its own
// options.
func TestGPGArgsSeparatesPositionals(t *testing.T) {
t.Parallel()
assert.Equal(t,
[]string{"--opt-a", "--opt-b", "--", "--version"},
gpgArgs([]string{"--opt-a", "--opt-b"}, "--version"))
assert.Equal(t,
[]string{"--opt-c", "--", "sig", "data"},
gpgArgs([]string{"--opt-c"}, "sig", "data"))
assert.Equal(t, []string{"--opt-d", "--"},
gpgArgs([]string{"--opt-d"}))
}
// TestGPGOptionLikeKeyIDIsNotAnOption drives real gpg with a key ID that
// looks like an option and asserts it is treated as a (nonexistent) key
// rather than executed as gpg's own --version.
func TestGPGOptionLikeKeyIDIsNotAnOption(t *testing.T) {
_, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
pubKey, err := gpgExportPublicKey(GPGKeyID("--version"))
require.Error(t, err)
require.ErrorIs(t, err, errGPGKeyNotFound)
assert.NotContains(t, string(pubKey), "gpg (GnuPG)")
fpr, err := gpgGetKeyFingerprint(GPGKeyID("--version"))
require.Error(t, err)
assert.NotContains(t, string(fpr), "gpg (GnuPG)")
}
func TestGPGSignInvalidKey(t *testing.T) { func TestGPGSignInvalidKey(t *testing.T) {
// Set up test environment (we need GNUPGHOME set) // Set up test environment (we need GNUPGHOME set)
_, cleanup := testGPGEnv(t) _, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
data := []byte("test data") data := []byte("test data")
_, err := gpgSign(data, GPGKeyID("NONEXISTENT_KEY_ID_12345")) _, err := gpgSign(data, GPGKeyID("NONEXISTENT_KEY_ID_12345"))
@@ -138,8 +167,8 @@ func TestGPGSignInvalidKey(t *testing.T) {
} }
func TestBuilderWithSigning(t *testing.T) { func TestBuilderWithSigning(t *testing.T) {
keyID, cleanup := testGPGEnv(t) keyID, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
// Create a builder with signing options // Create a builder with signing options
b := NewBuilder() b := NewBuilder()
@@ -155,6 +184,7 @@ func TestBuilderWithSigning(t *testing.T) {
// Build the manifest // Build the manifest
var buf bytes.Buffer var buf bytes.Buffer
err = b.Build(&buf) err = b.Build(&buf)
require.NoError(t, err) require.NoError(t, err)
@@ -163,26 +193,32 @@ func TestBuilderWithSigning(t *testing.T) {
require.NoError(t, err) require.NoError(t, err)
require.NotNil(t, manifest.pbOuter) require.NotNil(t, manifest.pbOuter)
assert.NotEmpty(t, manifest.pbOuter.Signature, "signature should be populated") assert.NotEmpty(t, manifest.pbOuter.GetSignature(),
assert.NotEmpty(t, manifest.pbOuter.Signer, "signer should be populated") "signature should be populated")
assert.NotEmpty(t, manifest.pbOuter.SigningPubKey, "signing public key should be populated") assert.NotEmpty(t, manifest.pbOuter.GetSigner(), "signer should be populated")
assert.NotEmpty(t, manifest.pbOuter.GetSigningPubKey(),
"signing public key should be populated")
// Verify signature is a valid PGP signature // Verify signature is a valid PGP signature
assert.Contains(t, string(manifest.pbOuter.Signature), "-----BEGIN PGP SIGNATURE-----") assert.Contains(t, string(manifest.pbOuter.GetSignature()),
"-----BEGIN PGP SIGNATURE-----")
// Verify public key is a valid PGP public key block // Verify public key is a valid PGP public key block
assert.Contains(t, string(manifest.pbOuter.SigningPubKey), "-----BEGIN PGP PUBLIC KEY BLOCK-----") assert.Contains(t, string(manifest.pbOuter.GetSigningPubKey()),
"-----BEGIN PGP PUBLIC KEY BLOCK-----")
} }
func TestScannerWithSigning(t *testing.T) { func TestScannerWithSigning(t *testing.T) {
keyID, cleanup := testGPGEnv(t) keyID, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
// Create in-memory filesystem with test files // Create in-memory filesystem with test files
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644)) require.NoError(t,
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644)) afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644))
// Create scanner with signing options // Create scanner with signing options
opts := &ScannerOptions{ opts := &ScannerOptions{
@@ -205,14 +241,14 @@ func TestScannerWithSigning(t *testing.T) {
manifest, err := NewManifestFromReader(&buf) manifest, err := NewManifestFromReader(&buf)
require.NoError(t, err) require.NoError(t, err)
assert.NotEmpty(t, manifest.pbOuter.Signature) assert.NotEmpty(t, manifest.pbOuter.GetSignature())
assert.NotEmpty(t, manifest.pbOuter.Signer) assert.NotEmpty(t, manifest.pbOuter.GetSigner())
assert.NotEmpty(t, manifest.pbOuter.SigningPubKey) assert.NotEmpty(t, manifest.pbOuter.GetSigningPubKey())
} }
func TestGPGVerify(t *testing.T) { func TestGPGVerify(t *testing.T) {
keyID, cleanup := testGPGEnv(t) keyID, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
data := []byte("test data to sign and verify") data := []byte("test data to sign and verify")
sig, err := gpgSign(data, keyID) sig, err := gpgSign(data, keyID)
@@ -227,8 +263,8 @@ func TestGPGVerify(t *testing.T) {
} }
func TestGPGVerifyInvalidSignature(t *testing.T) { func TestGPGVerifyInvalidSignature(t *testing.T) {
keyID, cleanup := testGPGEnv(t) keyID, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
data := []byte("test data to sign") data := []byte("test data to sign")
sig, err := gpgSign(data, keyID) sig, err := gpgSign(data, keyID)
@@ -244,8 +280,8 @@ func TestGPGVerifyInvalidSignature(t *testing.T) {
} }
func TestGPGVerifyBadPublicKey(t *testing.T) { func TestGPGVerifyBadPublicKey(t *testing.T) {
keyID, cleanup := testGPGEnv(t) keyID, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
data := []byte("test data") data := []byte("test data")
sig, err := gpgSign(data, keyID) sig, err := gpgSign(data, keyID)
@@ -258,8 +294,8 @@ func TestGPGVerifyBadPublicKey(t *testing.T) {
} }
func TestManifestSignatureVerification(t *testing.T) { func TestManifestSignatureVerification(t *testing.T) {
keyID, cleanup := testGPGEnv(t) keyID, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
// Create a builder with signing options // Create a builder with signing options
b := NewBuilder() b := NewBuilder()
@@ -275,6 +311,7 @@ func TestManifestSignatureVerification(t *testing.T) {
// Build the manifest // Build the manifest
var buf bytes.Buffer var buf bytes.Buffer
err = b.Build(&buf) err = b.Build(&buf)
require.NoError(t, err) require.NoError(t, err)
@@ -284,12 +321,12 @@ func TestManifestSignatureVerification(t *testing.T) {
require.NotNil(t, manifest) require.NotNil(t, manifest)
// Signature should be present and valid // Signature should be present and valid
assert.NotEmpty(t, manifest.pbOuter.Signature) assert.NotEmpty(t, manifest.pbOuter.GetSignature())
} }
func TestManifestTamperedSignatureFails(t *testing.T) { func TestManifestTamperedSignatureFails(t *testing.T) {
keyID, cleanup := testGPGEnv(t) keyID, gpgHome := testGPGEnv(t)
defer cleanup() t.Setenv("GNUPGHOME", gpgHome)
// Create a signed manifest // Create a signed manifest
b := NewBuilder() b := NewBuilder()
@@ -303,6 +340,7 @@ func TestManifestTamperedSignatureFails(t *testing.T) {
require.NoError(t, err) require.NoError(t, err)
var buf bytes.Buffer var buf bytes.Buffer
err = b.Build(&buf) err = b.Build(&buf)
require.NoError(t, err) require.NoError(t, err)
@@ -312,6 +350,7 @@ func TestManifestTamperedSignatureFails(t *testing.T) {
for i := range data { for i := range data {
if i > 100 && data[i] == 'A' { if i > 100 && data[i] == 'A' {
data[i] = 'B' data[i] = 'B'
break break
} }
} }
@@ -322,6 +361,8 @@ func TestManifestTamperedSignatureFails(t *testing.T) {
} }
func TestBuilderWithoutSigning(t *testing.T) { func TestBuilderWithoutSigning(t *testing.T) {
t.Parallel()
// Create a builder without signing options // Create a builder without signing options
b := NewBuilder() b := NewBuilder()
@@ -333,6 +374,7 @@ func TestBuilderWithoutSigning(t *testing.T) {
// Build the manifest // Build the manifest
var buf bytes.Buffer var buf bytes.Buffer
err = b.Build(&buf) err = b.Build(&buf)
require.NoError(t, err) require.NoError(t, err)
@@ -341,7 +383,10 @@ func TestBuilderWithoutSigning(t *testing.T) {
require.NoError(t, err) require.NoError(t, err)
require.NotNil(t, manifest.pbOuter) require.NotNil(t, manifest.pbOuter)
assert.Empty(t, manifest.pbOuter.Signature, "signature should be empty when not signing") assert.Empty(t, manifest.pbOuter.GetSignature(),
assert.Empty(t, manifest.pbOuter.Signer, "signer should be empty when not signing") "signature should be empty when not signing")
assert.Empty(t, manifest.pbOuter.SigningPubKey, "signing public key should be empty when not signing") assert.Empty(t, manifest.pbOuter.GetSigner(),
"signer should be empty when not signing")
assert.Empty(t, manifest.pbOuter.GetSigningPubKey(),
"signing public key should be empty when not signing")
} }

View File

@@ -9,21 +9,32 @@ import (
"github.com/multiformats/go-multihash" "github.com/multiformats/go-multihash"
) )
var (
errOuterNotSet = errors.New("pbOuter not set")
errUUIDNotSet = errors.New("UUID not set")
errSHA256NotSet = errors.New("SHA256 hash not set")
)
// manifest holds the internal representation of a manifest file. // manifest holds the internal representation of a manifest file.
// Use NewManifestFromFile or NewManifestFromReader to load an existing manifest, // Use NewManifestFromFile or NewManifestFromReader to load an existing
// or use Builder to create a new one. // manifest, or use Builder to create a new one.
//
// Whether this type should be exported is an open design question owned by
// the repository owner; see README design question 13.
type manifest struct { type manifest struct {
pbInner *MFFile pbInner *MFFile
pbOuter *MFFileOuter pbOuter *MFFileOuter
output *bytes.Buffer output *bytes.Buffer
signingOptions *SigningOptions signingOptions *SigningOptions
fixedUUID []byte // if set, use this UUID instead of generating one
} }
func (m *manifest) String() string { func (m *manifest) String() string {
count := 0 count := 0
if m.pbInner != nil { if m.pbInner != nil {
count = len(m.pbInner.Files) count = len(m.pbInner.GetFiles())
} }
return fmt.Sprintf("<Manifest count=%d>", count) return fmt.Sprintf("<Manifest count=%d>", count)
} }
@@ -32,7 +43,8 @@ func (m *manifest) Files() []*MFFilePath {
if m.pbInner == nil { if m.pbInner == nil {
return nil return nil
} }
return m.pbInner.Files
return m.pbInner.GetFiles()
} }
// signatureString generates the canonical string used for signing/verification. // signatureString generates the canonical string used for signing/verification.
@@ -40,20 +52,24 @@ func (m *manifest) Files() []*MFFilePath {
// Requires pbOuter to be set with Uuid and Sha256 fields. // Requires pbOuter to be set with Uuid and Sha256 fields.
func (m *manifest) signatureString() (string, error) { func (m *manifest) signatureString() (string, error) {
if m.pbOuter == nil { if m.pbOuter == nil {
return "", errors.New("pbOuter not set") return "", errOuterNotSet
}
if len(m.pbOuter.Uuid) == 0 {
return "", errors.New("UUID not set")
}
if len(m.pbOuter.Sha256) == 0 {
return "", errors.New("SHA256 hash not set")
} }
mh, err := multihash.Encode(m.pbOuter.Sha256, multihash.SHA2_256) if len(m.pbOuter.GetUuid()) == 0 {
return "", errUUIDNotSet
}
if len(m.pbOuter.GetSha256()) == 0 {
return "", errSHA256NotSet
}
mh, err := multihash.Encode(m.pbOuter.GetSha256(), multihash.SHA2_256)
if err != nil { if err != nil {
return "", fmt.Errorf("failed to encode multihash: %w", err) return "", fmt.Errorf("failed to encode multihash: %w", err)
} }
uuidStr := hex.EncodeToString(m.pbOuter.Uuid)
uuidStr := hex.EncodeToString(m.pbOuter.GetUuid())
mhStr := hex.EncodeToString(mh) mhStr := hex.EncodeToString(mh)
return fmt.Sprintf("%s-%s-%s", MAGIC, uuidStr, mhStr), nil return fmt.Sprintf("%s-%s-%s", MAGIC, uuidStr, mhStr), nil
} }

View File

@@ -1,7 +1,7 @@
// Code generated by protoc-gen-go. DO NOT EDIT. // Code generated by protoc-gen-go. DO NOT EDIT.
// versions: // versions:
// protoc-gen-go v1.36.11 // protoc-gen-go v1.36.11
// protoc v6.33.0 // protoc v6.33.4
// source: mf.proto // source: mf.proto
package mfer package mfer
@@ -329,6 +329,9 @@ func (x *MFFileOuter) GetSigningPubKey() []byte {
type MFFilePath struct { type MFFilePath struct {
state protoimpl.MessageState `protogen:"open.v1"` state protoimpl.MessageState `protogen:"open.v1"`
// required attributes: // required attributes:
// Path invariants: must be valid UTF-8, use forward slashes only,
// be relative (no leading /), contain no ".." segments, and no
// empty segments (no "//").
Path string `protobuf:"bytes,1,opt,name=path,proto3" json:"path,omitempty"` Path string `protobuf:"bytes,1,opt,name=path,proto3" json:"path,omitempty"`
Size int64 `protobuf:"varint,2,opt,name=size,proto3" json:"size,omitempty"` Size int64 `protobuf:"varint,2,opt,name=size,proto3" json:"size,omitempty"`
// gotta have at least one: // gotta have at least one:
@@ -337,7 +340,6 @@ type MFFilePath struct {
MimeType *string `protobuf:"bytes,301,opt,name=mimeType,proto3,oneof" json:"mimeType,omitempty"` MimeType *string `protobuf:"bytes,301,opt,name=mimeType,proto3,oneof" json:"mimeType,omitempty"`
Mtime *Timestamp `protobuf:"bytes,302,opt,name=mtime,proto3,oneof" json:"mtime,omitempty"` Mtime *Timestamp `protobuf:"bytes,302,opt,name=mtime,proto3,oneof" json:"mtime,omitempty"`
Ctime *Timestamp `protobuf:"bytes,303,opt,name=ctime,proto3,oneof" json:"ctime,omitempty"` Ctime *Timestamp `protobuf:"bytes,303,opt,name=ctime,proto3,oneof" json:"ctime,omitempty"`
Atime *Timestamp `protobuf:"bytes,304,opt,name=atime,proto3,oneof" json:"atime,omitempty"`
unknownFields protoimpl.UnknownFields unknownFields protoimpl.UnknownFields
sizeCache protoimpl.SizeCache sizeCache protoimpl.SizeCache
} }
@@ -414,13 +416,6 @@ func (x *MFFilePath) GetCtime() *Timestamp {
return nil return nil
} }
func (x *MFFilePath) GetAtime() *Timestamp {
if x != nil {
return x.Atime
}
return nil
}
type MFFileChecksum struct { type MFFileChecksum struct {
state protoimpl.MessageState `protogen:"open.v1"` state protoimpl.MessageState `protogen:"open.v1"`
// 1.0 golang implementation must write a multihash here // 1.0 golang implementation must write a multihash here
@@ -566,7 +561,7 @@ const file_mf_proto_rawDesc = "" +
"\n" + "\n" +
"_signatureB\t\n" + "_signatureB\t\n" +
"\a_signerB\x10\n" + "\a_signerB\x10\n" +
"\x0e_signingPubKey\"\xa2\x02\n" + "\x0e_signingPubKey\"\xf0\x01\n" +
"\n" + "\n" +
"MFFilePath\x12\x12\n" + "MFFilePath\x12\x12\n" +
"\x04path\x18\x01 \x01(\tR\x04path\x12\x12\n" + "\x04path\x18\x01 \x01(\tR\x04path\x12\x12\n" +
@@ -576,13 +571,10 @@ const file_mf_proto_rawDesc = "" +
"\x05mtime\x18\xae\x02 \x01(\v2\n" + "\x05mtime\x18\xae\x02 \x01(\v2\n" +
".TimestampH\x01R\x05mtime\x88\x01\x01\x12&\n" + ".TimestampH\x01R\x05mtime\x88\x01\x01\x12&\n" +
"\x05ctime\x18\xaf\x02 \x01(\v2\n" + "\x05ctime\x18\xaf\x02 \x01(\v2\n" +
".TimestampH\x02R\x05ctime\x88\x01\x01\x12&\n" + ".TimestampH\x02R\x05ctime\x88\x01\x01B\v\n" +
"\x05atime\x18\xb0\x02 \x01(\v2\n" +
".TimestampH\x03R\x05atime\x88\x01\x01B\v\n" +
"\t_mimeTypeB\b\n" + "\t_mimeTypeB\b\n" +
"\x06_mtimeB\b\n" + "\x06_mtimeB\b\n" +
"\x06_ctimeB\b\n" + "\x06_ctime\".\n" +
"\x06_atime\".\n" +
"\x0eMFFileChecksum\x12\x1c\n" + "\x0eMFFileChecksum\x12\x1c\n" +
"\tmultiHash\x18\x01 \x01(\fR\tmultiHash\"\xd6\x01\n" + "\tmultiHash\x18\x01 \x01(\fR\tmultiHash\"\xd6\x01\n" +
"\x06MFFile\x12)\n" + "\x06MFFile\x12)\n" +
@@ -627,15 +619,14 @@ var file_mf_proto_depIdxs = []int32{
6, // 2: MFFilePath.hashes:type_name -> MFFileChecksum 6, // 2: MFFilePath.hashes:type_name -> MFFileChecksum
3, // 3: MFFilePath.mtime:type_name -> Timestamp 3, // 3: MFFilePath.mtime:type_name -> Timestamp
3, // 4: MFFilePath.ctime:type_name -> Timestamp 3, // 4: MFFilePath.ctime:type_name -> Timestamp
3, // 5: MFFilePath.atime:type_name -> Timestamp 2, // 5: MFFile.version:type_name -> MFFile.Version
2, // 6: MFFile.version:type_name -> MFFile.Version 5, // 6: MFFile.files:type_name -> MFFilePath
5, // 7: MFFile.files:type_name -> MFFilePath 3, // 7: MFFile.createdAt:type_name -> Timestamp
3, // 8: MFFile.createdAt:type_name -> Timestamp 8, // [8:8] is the sub-list for method output_type
9, // [9:9] is the sub-list for method output_type 8, // [8:8] is the sub-list for method input_type
9, // [9:9] is the sub-list for method input_type 8, // [8:8] is the sub-list for extension type_name
9, // [9:9] is the sub-list for extension type_name 8, // [8:8] is the sub-list for extension extendee
9, // [9:9] is the sub-list for extension extendee 0, // [0:8] is the sub-list for field type_name
0, // [0:9] is the sub-list for field type_name
} }
func init() { file_mf_proto_init() } func init() { file_mf_proto_init() }

View File

@@ -59,7 +59,6 @@ message MFFilePath {
optional string mimeType = 301; optional string mimeType = 301;
optional Timestamp mtime = 302; optional Timestamp mtime = 302;
optional Timestamp ctime = 303; optional Timestamp ctime = 303;
optional Timestamp atime = 304;
} }
message MFFileChecksum { message MFFileChecksum {

View File

@@ -43,10 +43,20 @@ type ScanStatus struct {
// ScannerOptions configures scanner behavior. // ScannerOptions configures scanner behavior.
type ScannerOptions struct { type ScannerOptions struct {
IncludeDotfiles bool // Include files and directories starting with a dot (default: exclude) // IncludeDotfiles includes files and directories starting with a dot
FollowSymLinks bool // Resolve symlinks instead of skipping them // (default: exclude).
Fs afero.Fs // Filesystem to use, defaults to OsFs if nil IncludeDotfiles bool
SigningOptions *SigningOptions // GPG signing options (nil = no signing) // FollowSymLinks resolves symlinks instead of skipping them.
FollowSymLinks bool
// IncludeTimestamps includes a createdAt timestamp in the manifest
// (default: omit for determinism).
IncludeTimestamps bool
// Fs is the filesystem to use, defaults to OsFs if nil.
Fs afero.Fs
// SigningOptions holds GPG signing options (nil = no signing).
SigningOptions *SigningOptions
// Seed, if set, derives a deterministic UUID from this seed.
Seed string
} }
// FileEntry represents a file that has been enumerated. // FileEntry represents a file that has been enumerated.
@@ -77,10 +87,12 @@ func NewScannerWithOptions(opts *ScannerOptions) *Scanner {
if opts == nil { if opts == nil {
opts = &ScannerOptions{} opts = &ScannerOptions{}
} }
fs := opts.Fs fs := opts.Fs
if fs == nil { if fs == nil {
fs = afero.NewOsFs() fs = afero.NewOsFs()
} }
return &Scanner{ return &Scanner{
files: make([]*FileEntry, 0), files: make([]*FileEntry, 0),
options: opts, options: opts,
@@ -94,47 +106,63 @@ func (s *Scanner) EnumerateFile(filePath string) error {
if err != nil { if err != nil {
return err return err
} }
info, err := s.fs.Stat(abs) info, err := s.fs.Stat(abs)
if err != nil { if err != nil {
return err return err
} }
// For single files, use the filename as the relative path // For single files, use the filename as the relative path
basePath := filepath.Dir(abs) basePath := filepath.Dir(abs)
return s.enumerateFileWithInfo(filepath.Base(abs), basePath, info, nil) return s.enumerateFileWithInfo(filepath.Base(abs), basePath, info, nil)
} }
// EnumeratePath walks a directory path and adds all files to the scanner. // EnumeratePath walks a directory path and adds all files to the scanner.
// If progress is non-nil, status updates are sent as files are discovered. // If progress is non-nil, status updates are sent as files are discovered.
// The progress channel is closed when the method returns. // The progress channel is closed when the method returns.
func (s *Scanner) EnumeratePath(inputPath string, progress chan<- EnumerateStatus) error { func (s *Scanner) EnumeratePath(
inputPath string,
progress chan<- EnumerateStatus,
) error {
if progress != nil { if progress != nil {
defer close(progress) defer close(progress)
} }
abs, err := filepath.Abs(inputPath) abs, err := filepath.Abs(inputPath)
if err != nil { if err != nil {
return err return err
} }
afs := afero.NewReadOnlyFs(afero.NewBasePathFs(s.fs, abs)) afs := afero.NewReadOnlyFs(afero.NewBasePathFs(s.fs, abs))
return s.enumerateFS(afs, abs, progress) return s.enumerateFS(afs, abs, progress)
} }
// EnumeratePaths walks multiple directory paths and adds all files to the scanner. // EnumeratePaths walks multiple directory paths and adds all files to the scanner.
// If progress is non-nil, status updates are sent as files are discovered. // If progress is non-nil, status updates are sent as files are discovered.
// The progress channel is closed when the method returns. // The progress channel is closed when the method returns.
func (s *Scanner) EnumeratePaths(progress chan<- EnumerateStatus, inputPaths ...string) error { func (s *Scanner) EnumeratePaths(
progress chan<- EnumerateStatus,
inputPaths ...string,
) error {
if progress != nil { if progress != nil {
defer close(progress) defer close(progress)
} }
for _, p := range inputPaths { for _, p := range inputPaths {
abs, err := filepath.Abs(p) abs, err := filepath.Abs(p)
if err != nil { if err != nil {
return err return err
} }
afs := afero.NewReadOnlyFs(afero.NewBasePathFs(s.fs, abs)) afs := afero.NewReadOnlyFs(afero.NewBasePathFs(s.fs, abs))
if err := s.enumerateFS(afs, abs, progress); err != nil {
err = s.enumerateFS(afs, abs, progress)
if err != nil {
return err return err
} }
} }
return nil return nil
} }
@@ -142,31 +170,231 @@ func (s *Scanner) EnumeratePaths(progress chan<- EnumerateStatus, inputPaths ...
// If progress is non-nil, status updates are sent as files are discovered. // If progress is non-nil, status updates are sent as files are discovered.
// The progress channel is closed when the method returns. // The progress channel is closed when the method returns.
// basePath is used to compute absolute paths for file reading. // basePath is used to compute absolute paths for file reading.
func (s *Scanner) EnumerateFS(afs afero.Fs, basePath string, progress chan<- EnumerateStatus) error { func (s *Scanner) EnumerateFS(
afs afero.Fs,
basePath string,
progress chan<- EnumerateStatus,
) error {
if progress != nil { if progress != nil {
defer close(progress) defer close(progress)
} }
return s.enumerateFS(afs, basePath, progress) return s.enumerateFS(afs, basePath, progress)
} }
// enumerateFS is the internal implementation that doesn't close the progress channel. // Files returns a copy of all files added to the scanner.
func (s *Scanner) enumerateFS(afs afero.Fs, basePath string, progress chan<- EnumerateStatus) error { func (s *Scanner) Files() []*FileEntry {
s.mu.RLock()
defer s.mu.RUnlock()
out := make([]*FileEntry, len(s.files))
copy(out, s.files)
return out
}
// FileCount returns the number of files in the scanner.
func (s *Scanner) FileCount() FileCount {
s.mu.RLock()
defer s.mu.RUnlock()
return FileCount(len(s.files))
}
// TotalBytes returns the total size of all files in the scanner.
func (s *Scanner) TotalBytes() FileSize {
s.mu.RLock()
defer s.mu.RUnlock()
return s.totalBytes
}
// ToManifest reads all file contents, computes hashes, and generates a manifest.
// If progress is non-nil, status updates are sent approximately once per second.
// The progress channel is closed when the method returns.
// The manifest is written to the provided io.Writer.
func (s *Scanner) ToManifest(
ctx context.Context, w io.Writer, progress chan<- ScanStatus,
) error {
if progress != nil {
defer close(progress)
}
s.mu.RLock()
files := make([]*FileEntry, len(s.files))
copy(files, s.files)
totalFiles := FileCount(len(files))
var totalBytes FileSize
for _, f := range files {
totalBytes += f.Size
}
s.mu.RUnlock()
builder := s.configureBuilder()
var (
scannedFiles FileCount
scannedBytes FileSize
)
lastProgressTime := time.Now()
startTime := time.Now()
pt := &scanProgressTracker{
progress: progress,
totalFiles: totalFiles,
totalBytes: totalBytes,
startTime: startTime,
lastProgress: &lastProgressTime,
}
for _, entry := range files {
// Check for cancellation
select {
case <-ctx.Done():
return ctx.Err()
default:
}
bytesRead, err := s.scanFile(builder, pt, entry, scannedFiles, scannedBytes)
if err != nil {
return err
}
scannedFiles++
scannedBytes += bytesRead
}
// Send final progress (ETA is 0 at completion; remaining bytes are 0,
// so computeRateETA yields eta 0 and the same average rate as before)
if progress != nil {
rate, _ := computeRateETA(time.Since(startTime), scannedBytes, totalBytes)
sendScanStatus(progress, ScanStatus{
TotalFiles: totalFiles,
ScannedFiles: scannedFiles,
TotalBytes: totalBytes,
ScannedBytes: scannedBytes,
BytesPerSec: rate,
ETA: 0,
})
}
// Build and write manifest
//nolint:contextcheck // Build's GPG signing exec is not cancellable by design
return builder.Build(w)
}
// configureBuilder constructs a manifest builder configured from the
// scanner options.
func (s *Scanner) configureBuilder() *Builder {
builder := NewBuilder()
if s.options.IncludeTimestamps {
builder.SetIncludeTimestamps(true)
}
if s.options.SigningOptions != nil {
builder.SetSigningOptions(s.options.SigningOptions)
}
if s.options.Seed != "" {
builder.SetSeed(s.options.Seed)
}
return builder
}
// scanFile hashes a single file into the builder, forwarding per-file
// progress updates, and returns the number of bytes read.
func (s *Scanner) scanFile(
builder *Builder,
pt *scanProgressTracker,
entry *FileEntry,
scannedFiles FileCount,
scannedBytes FileSize,
) (FileSize, error) {
// Open file
f, err := s.fs.Open(string(entry.AbsPath))
if err != nil {
return 0, err
}
// Create progress channel for this file
var (
fileProgress chan FileHashProgress
wg sync.WaitGroup
)
if pt.progress != nil {
fileProgress = make(chan FileHashProgress, 1)
wg.Add(1)
go func(base FileSize, done FileCount) {
defer wg.Done()
pt.forward(fileProgress, done, base)
}(scannedBytes, scannedFiles)
}
// Add to manifest with progress channel
bytesRead, err := builder.AddFile(
entry.Path,
entry.Size,
entry.Mtime,
f,
fileProgress,
)
_ = f.Close()
// Close channel and wait for goroutine to finish
if fileProgress != nil {
close(fileProgress)
wg.Wait()
}
if err != nil {
return 0, err
}
log.Verbosef("+ %s (%s)", entry.Path, humanize.IBytes(sizeToUint64(bytesRead)))
return bytesRead, nil
}
// enumerateFS is the internal implementation that doesn't close the
// progress channel.
func (s *Scanner) enumerateFS(
afs afero.Fs,
basePath string,
progress chan<- EnumerateStatus,
) error {
return afero.Walk(afs, "/", func(p string, info fs.FileInfo, err error) error { return afero.Walk(afs, "/", func(p string, info fs.FileInfo, err error) error {
if err != nil { if err != nil {
return err return err
} }
if !s.options.IncludeDotfiles && IsHiddenPath(p) { if !s.options.IncludeDotfiles && IsHiddenPath(p) {
if info.IsDir() { if info.IsDir() {
return filepath.SkipDir return filepath.SkipDir
} }
return nil return nil
} }
return s.enumerateFileWithInfo(p, basePath, info, progress) return s.enumerateFileWithInfo(p, basePath, info, progress)
}) })
} }
// enumerateFileWithInfo adds a file with pre-existing fs.FileInfo. // enumerateFileWithInfo adds a file with pre-existing fs.FileInfo.
func (s *Scanner) enumerateFileWithInfo(filePath string, basePath string, info fs.FileInfo, progress chan<- EnumerateStatus) error { func (s *Scanner) enumerateFileWithInfo(
filePath string,
basePath string,
info fs.FileInfo,
progress chan<- EnumerateStatus,
) error {
if info.IsDir() { if info.IsDir() {
// Manifests contain only files, directories are implied // Manifests contain only files, directories are implied
return nil return nil
@@ -191,11 +419,13 @@ func (s *Scanner) enumerateFileWithInfo(filePath string, basePath string, info f
realPath, err := filepath.EvalSymlinks(absPath) realPath, err := filepath.EvalSymlinks(absPath)
if err != nil { if err != nil {
// Skip broken symlinks // Skip broken symlinks
return nil return nil //nolint:nilerr // broken symlinks are skipped by design
} }
realInfo, err := s.fs.Stat(realPath) realInfo, err := s.fs.Stat(realPath)
if err != nil { if err != nil {
return nil // Skip symlinks whose target cannot be stat'd
return nil //nolint:nilerr // unreadable targets are skipped by design
} }
// Skip if symlink points to a directory // Skip if symlink points to a directory
if realInfo.IsDir() { if realInfo.IsDir() {
@@ -230,154 +460,78 @@ func (s *Scanner) enumerateFileWithInfo(filePath string, basePath string, info f
return nil return nil
} }
// Files returns a copy of all files added to the scanner. // scanProgressTracker carries the shared state needed to report rate-limited
func (s *Scanner) Files() []*FileEntry { // scan progress updates.
s.mu.RLock() type scanProgressTracker struct {
defer s.mu.RUnlock() progress chan<- ScanStatus
out := make([]*FileEntry, len(s.files)) totalFiles FileCount
copy(out, s.files) totalBytes FileSize
return out startTime time.Time
lastProgress *time.Time
} }
// FileCount returns the number of files in the scanner. // forward relays per-file hash progress to the scan progress channel,
func (s *Scanner) FileCount() FileCount { // rate-limited to one update per second.
s.mu.RLock() func (pt *scanProgressTracker) forward(
defer s.mu.RUnlock() fileProgress <-chan FileHashProgress,
return FileCount(len(s.files)) scannedFiles FileCount,
} baseBytes FileSize,
) {
// TotalBytes returns the total size of all files in the scanner. for p := range fileProgress {
func (s *Scanner) TotalBytes() FileSize { // Send progress at most once per second
s.mu.RLock() now := time.Now()
defer s.mu.RUnlock() if now.Sub(*pt.lastProgress) < time.Second {
return s.totalBytes continue
}
// ToManifest reads all file contents, computes hashes, and generates a manifest.
// If progress is non-nil, status updates are sent approximately once per second.
// The progress channel is closed when the method returns.
// The manifest is written to the provided io.Writer.
func (s *Scanner) ToManifest(ctx context.Context, w io.Writer, progress chan<- ScanStatus) error {
if progress != nil {
defer close(progress)
}
s.mu.RLock()
files := make([]*FileEntry, len(s.files))
copy(files, s.files)
totalFiles := FileCount(len(files))
var totalBytes FileSize
for _, f := range files {
totalBytes += f.Size
}
s.mu.RUnlock()
builder := NewBuilder()
if s.options.SigningOptions != nil {
builder.SetSigningOptions(s.options.SigningOptions)
}
var scannedFiles FileCount
var scannedBytes FileSize
lastProgressTime := time.Now()
startTime := time.Now()
for _, entry := range files {
// Check for cancellation
select {
case <-ctx.Done():
return ctx.Err()
default:
} }
// Open file currentBytes := baseBytes + p.BytesRead
f, err := s.fs.Open(string(entry.AbsPath)) rate, eta := computeRateETA(now.Sub(pt.startTime), currentBytes, pt.totalBytes)
if err != nil {
return err
}
// Create progress channel for this file sendScanStatus(pt.progress, ScanStatus{
var fileProgress chan FileHashProgress TotalFiles: pt.totalFiles,
var wg sync.WaitGroup
if progress != nil {
fileProgress = make(chan FileHashProgress, 1)
wg.Add(1)
go func(baseScannedBytes FileSize) {
defer wg.Done()
for p := range fileProgress {
// Send progress at most once per second
now := time.Now()
if now.Sub(lastProgressTime) >= time.Second {
elapsed := now.Sub(startTime).Seconds()
currentBytes := baseScannedBytes + p.BytesRead
var rate float64
var eta time.Duration
if elapsed > 0 && currentBytes > 0 {
rate = float64(currentBytes) / elapsed
remainingBytes := totalBytes - currentBytes
if rate > 0 {
eta = time.Duration(float64(remainingBytes)/rate) * time.Second
}
}
sendScanStatus(progress, ScanStatus{
TotalFiles: totalFiles,
ScannedFiles: scannedFiles,
TotalBytes: totalBytes,
ScannedBytes: currentBytes,
BytesPerSec: rate,
ETA: eta,
})
lastProgressTime = now
}
}
}(scannedBytes)
}
// Add to manifest with progress channel
bytesRead, err := builder.AddFile(
entry.Path,
entry.Size,
entry.Mtime,
f,
fileProgress,
)
_ = f.Close()
// Close channel and wait for goroutine to finish
if fileProgress != nil {
close(fileProgress)
wg.Wait()
}
if err != nil {
return err
}
log.Verbosef("+ %s (%s)", entry.Path, humanize.IBytes(uint64(bytesRead)))
scannedFiles++
scannedBytes += bytesRead
}
// Send final progress (ETA is 0 at completion)
if progress != nil {
elapsed := time.Since(startTime).Seconds()
var rate float64
if elapsed > 0 {
rate = float64(scannedBytes) / elapsed
}
sendScanStatus(progress, ScanStatus{
TotalFiles: totalFiles,
ScannedFiles: scannedFiles, ScannedFiles: scannedFiles,
TotalBytes: totalBytes, TotalBytes: pt.totalBytes,
ScannedBytes: scannedBytes, ScannedBytes: currentBytes,
BytesPerSec: rate, BytesPerSec: rate,
ETA: 0, ETA: eta,
}) })
*pt.lastProgress = now
}
}
// computeRateETA returns the average throughput over elapsed time and the
// estimated time to process the remaining bytes at that rate.
func computeRateETA(
elapsed time.Duration,
done FileSize,
total FileSize,
) (float64, time.Duration) {
var (
rate float64
eta time.Duration
)
if elapsed > 0 && done > 0 {
rate = float64(done) / elapsed.Seconds()
remaining := total - done
if rate > 0 {
eta = time.Duration(float64(remaining)/rate) * time.Second
}
} }
// Build and write manifest return rate, eta
return builder.Build(w) }
// sizeToUint64 converts a FileSize to uint64 for display, clamping
// negative values to zero so the conversion cannot overflow.
func sizeToUint64(v FileSize) uint64 {
if v < 0 {
return 0
}
return uint64(v)
} }
// IsHiddenPath returns true if the path or any of its parent directories // IsHiddenPath returns true if the path or any of its parent directories
@@ -388,17 +542,21 @@ func IsHiddenPath(p string) bool {
if tp == "." || tp == "/" { if tp == "." || tp == "/" {
return false return false
} }
if strings.HasPrefix(tp, ".") { if strings.HasPrefix(tp, ".") {
return true return true
} }
for { for {
d, f := path.Split(tp) d, f := path.Split(tp)
if strings.HasPrefix(f, ".") { if strings.HasPrefix(f, ".") {
return true return true
} }
if d == "" { if d == "" {
return false return false
} }
tp = d[0 : len(d)-1] // trim trailing slash from dir tp = d[0 : len(d)-1] // trim trailing slash from dir
} }
} }
@@ -409,6 +567,7 @@ func sendEnumerateStatus(ch chan<- EnumerateStatus, status EnumerateStatus) {
if ch == nil { if ch == nil {
return return
} }
select { select {
case ch <- status: case ch <- status:
default: default:
@@ -422,6 +581,7 @@ func sendScanStatus(ch chan<- ScanStatus, status ScanStatus) {
if ch == nil { if ch == nil {
return return
} }
select { select {
case ch <- status: case ch <- status:
default: default:

View File

@@ -1,3 +1,4 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer package mfer
import ( import (
@@ -12,6 +13,8 @@ import (
) )
func TestNewScanner(t *testing.T) { func TestNewScanner(t *testing.T) {
t.Parallel()
s := NewScanner() s := NewScanner()
assert.NotNil(t, s) assert.NotNil(t, s)
assert.Equal(t, FileCount(0), s.FileCount()) assert.Equal(t, FileCount(0), s.FileCount())
@@ -19,12 +22,18 @@ func TestNewScanner(t *testing.T) {
} }
func TestNewScannerWithOptions(t *testing.T) { func TestNewScannerWithOptions(t *testing.T) {
t.Parallel()
t.Run("nil options", func(t *testing.T) { t.Run("nil options", func(t *testing.T) {
t.Parallel()
s := NewScannerWithOptions(nil) s := NewScannerWithOptions(nil)
assert.NotNil(t, s) assert.NotNil(t, s)
}) })
t.Run("with options", func(t *testing.T) { t.Run("with options", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
opts := &ScannerOptions{ opts := &ScannerOptions{
IncludeDotfiles: true, IncludeDotfiles: true,
@@ -37,6 +46,8 @@ func TestNewScannerWithOptions(t *testing.T) {
} }
func TestScannerEnumerateFile(t *testing.T) { func TestScannerEnumerateFile(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("hello world"), 0o644)) require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("hello world"), 0o644))
@@ -54,6 +65,8 @@ func TestScannerEnumerateFile(t *testing.T) {
} }
func TestScannerEnumerateFileMissing(t *testing.T) { func TestScannerEnumerateFileMissing(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs}) s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumerateFile("/nonexistent.txt") err := s.EnumerateFile("/nonexistent.txt")
@@ -61,11 +74,14 @@ func TestScannerEnumerateFileMissing(t *testing.T) {
} }
func TestScannerEnumeratePath(t *testing.T) { func TestScannerEnumeratePath(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir/subdir", 0o755)) require.NoError(t, fs.MkdirAll("/testdir/subdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("one"), 0o644)) require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("one"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("two"), 0o644)) require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("two"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/subdir/file3.txt", []byte("three"), 0o644)) require.NoError(t,
afero.WriteFile(fs, "/testdir/subdir/file3.txt", []byte("three"), 0o644))
s := NewScannerWithOptions(&ScannerOptions{Fs: fs}) s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumeratePath("/testdir", nil) err := s.EnumeratePath("/testdir", nil)
@@ -76,6 +92,8 @@ func TestScannerEnumeratePath(t *testing.T) {
} }
func TestScannerEnumeratePathWithProgress(t *testing.T) { func TestScannerEnumeratePathWithProgress(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("one"), 0o644)) require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("one"), 0o644))
@@ -100,6 +118,8 @@ func TestScannerEnumeratePathWithProgress(t *testing.T) {
} }
func TestScannerEnumeratePaths(t *testing.T) { func TestScannerEnumeratePaths(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/dir1", 0o755)) require.NoError(t, fs.MkdirAll("/dir1", 0o755))
require.NoError(t, fs.MkdirAll("/dir2", 0o755)) require.NoError(t, fs.MkdirAll("/dir2", 0o755))
@@ -114,13 +134,20 @@ func TestScannerEnumeratePaths(t *testing.T) {
} }
func TestScannerExcludeDotfiles(t *testing.T) { func TestScannerExcludeDotfiles(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir/.hidden", 0o755)) require.NoError(t, fs.MkdirAll("/testdir/.hidden", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/visible.txt", []byte("visible"), 0o644)) require.NoError(t,
require.NoError(t, afero.WriteFile(fs, "/testdir/.hidden.txt", []byte("hidden"), 0o644)) afero.WriteFile(fs, "/testdir/visible.txt", []byte("visible"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/.hidden/inside.txt", []byte("inside"), 0o644)) require.NoError(t,
afero.WriteFile(fs, "/testdir/.hidden.txt", []byte("hidden"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/.hidden/inside.txt", []byte("inside"), 0o644))
t.Run("exclude by default", func(t *testing.T) { t.Run("exclude by default", func(t *testing.T) {
t.Parallel()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs, IncludeDotfiles: false}) s := NewScannerWithOptions(&ScannerOptions{Fs: fs, IncludeDotfiles: false})
err := s.EnumeratePath("/testdir", nil) err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err) require.NoError(t, err)
@@ -131,6 +158,8 @@ func TestScannerExcludeDotfiles(t *testing.T) {
}) })
t.Run("include when enabled", func(t *testing.T) { t.Run("include when enabled", func(t *testing.T) {
t.Parallel()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs, IncludeDotfiles: true}) s := NewScannerWithOptions(&ScannerOptions{Fs: fs, IncludeDotfiles: true})
err := s.EnumeratePath("/testdir", nil) err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err) require.NoError(t, err)
@@ -140,34 +169,43 @@ func TestScannerExcludeDotfiles(t *testing.T) {
} }
func TestScannerToManifest(t *testing.T) { func TestScannerToManifest(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content one"), 0o644)) require.NoError(t,
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content two"), 0o644)) afero.WriteFile(fs, "/testdir/file1.txt", []byte("content one"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file2.txt", []byte("content two"), 0o644))
s := NewScannerWithOptions(&ScannerOptions{Fs: fs}) s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumeratePath("/testdir", nil) err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err) require.NoError(t, err)
var buf bytes.Buffer var buf bytes.Buffer
err = s.ToManifest(context.Background(), &buf, nil) err = s.ToManifest(context.Background(), &buf, nil)
require.NoError(t, err) require.NoError(t, err)
// Manifest should have magic bytes // Manifest should have magic bytes
assert.True(t, buf.Len() > 0) assert.Positive(t, buf.Len())
assert.Equal(t, MAGIC, string(buf.Bytes()[:8])) assert.Equal(t, MAGIC, string(buf.Bytes()[:8]))
} }
func TestScannerToManifestWithProgress(t *testing.T) { func TestScannerToManifestWithProgress(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file.txt", bytes.Repeat([]byte("x"), 1000), 0o644)) require.NoError(t,
afero.WriteFile(fs, "/testdir/file.txt", bytes.Repeat([]byte("x"), 1000), 0o644))
s := NewScannerWithOptions(&ScannerOptions{Fs: fs}) s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumeratePath("/testdir", nil) err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err) require.NoError(t, err)
var buf bytes.Buffer var buf bytes.Buffer
progress := make(chan ScanStatus, 10) progress := make(chan ScanStatus, 10)
err = s.ToManifest(context.Background(), &buf, progress) err = s.ToManifest(context.Background(), &buf, progress)
@@ -188,12 +226,15 @@ func TestScannerToManifestWithProgress(t *testing.T) {
} }
func TestScannerToManifestContextCancellation(t *testing.T) { func TestScannerToManifestContextCancellation(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll("/testdir", 0o755))
// Create many files to ensure we have time to cancel // Create many files to ensure we have time to cancel
for i := 0; i < 100; i++ { for i := range 100 {
name := string(rune('a'+i%26)) + string(rune('0'+i/26)) + ".txt" name := string(rune('a'+i%26)) + string(rune('0'+i/26)) + ".txt"
require.NoError(t, afero.WriteFile(fs, "/testdir/"+name, bytes.Repeat([]byte("x"), 100), 0o644)) require.NoError(t,
afero.WriteFile(fs, "/testdir/"+name, bytes.Repeat([]byte("x"), 100), 0o644))
} }
s := NewScannerWithOptions(&ScannerOptions{Fs: fs}) s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
@@ -204,24 +245,30 @@ func TestScannerToManifestContextCancellation(t *testing.T) {
cancel() // Cancel immediately cancel() // Cancel immediately
var buf bytes.Buffer var buf bytes.Buffer
err = s.ToManifest(ctx, &buf, nil) err = s.ToManifest(ctx, &buf, nil)
assert.ErrorIs(t, err, context.Canceled) assert.ErrorIs(t, err, context.Canceled)
} }
func TestScannerToManifestEmptyScanner(t *testing.T) { func TestScannerToManifestEmptyScanner(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs}) s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
var buf bytes.Buffer var buf bytes.Buffer
err := s.ToManifest(context.Background(), &buf, nil) err := s.ToManifest(context.Background(), &buf, nil)
require.NoError(t, err) require.NoError(t, err)
// Should still produce a valid manifest // Should still produce a valid manifest
assert.True(t, buf.Len() > 0) assert.Positive(t, buf.Len())
assert.Equal(t, MAGIC, string(buf.Bytes()[:8])) assert.Equal(t, MAGIC, string(buf.Bytes()[:8]))
} }
func TestScannerFilesCopiesSlice(t *testing.T) { func TestScannerFilesCopiesSlice(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("hello"), 0o644)) require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("hello"), 0o644))
@@ -236,10 +283,13 @@ func TestScannerFilesCopiesSlice(t *testing.T) {
} }
func TestScannerEnumerateFS(t *testing.T) { func TestScannerEnumerateFS(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir/sub", 0o755)) require.NoError(t, fs.MkdirAll("/testdir/sub", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file.txt", []byte("hello"), 0o644)) require.NoError(t, afero.WriteFile(fs, "/testdir/file.txt", []byte("hello"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/sub/nested.txt", []byte("world"), 0o644)) require.NoError(t,
afero.WriteFile(fs, "/testdir/sub/nested.txt", []byte("world"), 0o644))
// Create a basepath filesystem // Create a basepath filesystem
baseFs := afero.NewBasePathFs(fs, "/testdir") baseFs := afero.NewBasePathFs(fs, "/testdir")
@@ -252,13 +302,17 @@ func TestScannerEnumerateFS(t *testing.T) {
} }
func TestSendEnumerateStatusNonBlocking(t *testing.T) { func TestSendEnumerateStatusNonBlocking(t *testing.T) {
t.Parallel()
// Channel with no buffer - send should not block // Channel with no buffer - send should not block
ch := make(chan EnumerateStatus) ch := make(chan EnumerateStatus)
// This should not block // This should not block
done := make(chan bool) done := make(chan bool)
go func() { go func() {
sendEnumerateStatus(ch, EnumerateStatus{FilesFound: 1}) sendEnumerateStatus(ch, EnumerateStatus{FilesFound: 1})
done <- true done <- true
}() }()
@@ -271,12 +325,16 @@ func TestSendEnumerateStatusNonBlocking(t *testing.T) {
} }
func TestSendScanStatusNonBlocking(t *testing.T) { func TestSendScanStatusNonBlocking(t *testing.T) {
t.Parallel()
// Channel with no buffer - send should not block // Channel with no buffer - send should not block
ch := make(chan ScanStatus) ch := make(chan ScanStatus)
done := make(chan bool) done := make(chan bool)
go func() { go func() {
sendScanStatus(ch, ScanStatus{ScannedFiles: 1}) sendScanStatus(ch, ScanStatus{ScannedFiles: 1})
done <- true done <- true
}() }()
@@ -289,14 +347,19 @@ func TestSendScanStatusNonBlocking(t *testing.T) {
} }
func TestSendStatusNilChannel(t *testing.T) { func TestSendStatusNilChannel(t *testing.T) {
t.Parallel()
// Should not panic with nil channel // Should not panic with nil channel
sendEnumerateStatus(nil, EnumerateStatus{}) sendEnumerateStatus(nil, EnumerateStatus{})
sendScanStatus(nil, ScanStatus{}) sendScanStatus(nil, ScanStatus{})
} }
func TestScannerFileEntryFields(t *testing.T) { func TestScannerFileEntryFields(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
now := time.Now().Truncate(time.Second) now := time.Now().Truncate(time.Second)
require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("content"), 0o644)) require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("content"), 0o644))
require.NoError(t, fs.Chtimes("/test.txt", now, now)) require.NoError(t, fs.Chtimes("/test.txt", now, now))
@@ -315,11 +378,13 @@ func TestScannerFileEntryFields(t *testing.T) {
} }
func TestScannerLargeFileEnumeration(t *testing.T) { func TestScannerLargeFileEnumeration(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs() fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755)) require.NoError(t, fs.MkdirAll("/testdir", 0o755))
// Create 100 files // Create 100 files
for i := 0; i < 100; i++ { for i := range 100 {
name := "/testdir/" + string(rune('a'+i%26)) + string(rune('0'+i/26%10)) + ".txt" name := "/testdir/" + string(rune('a'+i%26)) + string(rune('0'+i/26%10)) + ".txt"
require.NoError(t, afero.WriteFile(fs, name, []byte("data"), 0o644)) require.NoError(t, afero.WriteFile(fs, name, []byte("data"), 0o644))
} }
@@ -330,20 +395,20 @@ func TestScannerLargeFileEnumeration(t *testing.T) {
err := s.EnumeratePath("/testdir", progress) err := s.EnumeratePath("/testdir", progress)
require.NoError(t, err) require.NoError(t, err)
// Drain channel // progress is fully buffered and closed; no draining needed
for range progress {
}
assert.Equal(t, FileCount(100), s.FileCount()) assert.Equal(t, FileCount(100), s.FileCount())
assert.Equal(t, FileSize(400), s.TotalBytes()) // 100 * 4 bytes assert.Equal(t, FileSize(400), s.TotalBytes()) // 100 * 4 bytes
} }
func TestIsHiddenPath(t *testing.T) { func TestIsHiddenPath(t *testing.T) {
t.Parallel()
tests := []struct { tests := []struct {
path string path string
hidden bool hidden bool
}{ }{
{"file.txt", false}, {testFileName, false},
{".hidden", true}, {".hidden", true},
{"dir/file.txt", false}, {"dir/file.txt", false},
{"dir/.hidden", true}, {"dir/.hidden", true},
@@ -352,12 +417,16 @@ func TestIsHiddenPath(t *testing.T) {
{"/absolute/.hidden", true}, {"/absolute/.hidden", true},
{"./relative", false}, // path.Clean removes leading ./ {"./relative", false}, // path.Clean removes leading ./
{"a/b/c/.d/e", true}, {"a/b/c/.d/e", true},
{".", false}, // current directory is not hidden {".", false}, // current directory is not hidden (#14)
{"/", false}, // root is not hidden {"/", false}, // root is not hidden
{"./", false}, // current directory with trailing slash
{"./file.txt", false}, // file in current directory
} }
for _, tt := range tests { for _, tt := range tests {
t.Run(tt.path, func(t *testing.T) { t.Run(tt.path, func(t *testing.T) {
t.Parallel()
assert.Equal(t, tt.hidden, IsHiddenPath(tt.path), "IsHiddenPath(%q)", tt.path) assert.Equal(t, tt.hidden, IsHiddenPath(tt.path), "IsHiddenPath(%q)", tt.path)
}) })
} }

View File

@@ -5,6 +5,7 @@ import (
"crypto/sha256" "crypto/sha256"
"errors" "errors"
"fmt" "fmt"
"math"
"time" "time"
"github.com/google/uuid" "github.com/google/uuid"
@@ -15,68 +16,112 @@ import (
// MAGIC is the file format magic bytes prefix (rot13 of "MANIFEST"). // MAGIC is the file format magic bytes prefix (rot13 of "MANIFEST").
const MAGIC string = "ZNAVSRFG" const MAGIC string = "ZNAVSRFG"
var (
// errInnerNotSet is returned by generate when the inner manifest is
// missing.
errInnerNotSet = errors.New("internal error: pbInner not set")
// errInternal is returned by generateOuter for the same condition.
// The two messages differ, and both are load-bearing for callers that
// match on text, so they are kept distinct.
errInternal = errors.New("internal error")
)
// nanosecondsInt32 converts t's nanosecond component to int32.
// time.Time.Nanosecond is documented to return a value in [0, 999999999],
// so the conversion cannot overflow. This sits directly in the manifest
// content path: silently substituting a default would zero every entry's
// mtime nanos and change the serialized bytes and their hash, so an
// out-of-contract value is a programming error and panics rather than
// being papered over.
func nanosecondsInt32(t time.Time) int32 {
n := t.Nanosecond()
if n < 0 || n > math.MaxInt32 {
panic(fmt.Sprintf(
"mfer: time.Time.Nanosecond out of contract: %d", n))
}
return int32(n)
}
func newTimestampFromTime(t time.Time) *Timestamp { func newTimestampFromTime(t time.Time) *Timestamp {
return &Timestamp{ return &Timestamp{
Seconds: t.Unix(), Seconds: t.Unix(),
Nanos: int32(t.Nanosecond()), Nanos: nanosecondsInt32(t),
} }
} }
func (m *manifest) generate() error { func (m *manifest) generate() error {
if m.pbInner == nil { if m.pbInner == nil {
return errors.New("internal error: pbInner not set") return errInnerNotSet
} }
if m.pbOuter == nil { if m.pbOuter == nil {
e := m.generateOuter() e := m.generateOuter()
if e != nil { if e != nil {
return e return e
} }
} }
dat, err := proto.MarshalOptions{Deterministic: true}.Marshal(m.pbOuter) dat, err := proto.MarshalOptions{Deterministic: true}.Marshal(m.pbOuter)
if err != nil { if err != nil {
return err return fmt.Errorf("serialize: marshal outer: %w", err)
} }
m.output = bytes.NewBuffer([]byte(MAGIC))
m.output = bytes.NewBufferString(MAGIC)
_, err = m.output.Write(dat) _, err = m.output.Write(dat)
if err != nil { if err != nil {
return err return fmt.Errorf("serialize: write output: %w", err)
} }
return nil return nil
} }
func (m *manifest) generateOuter() error { func (m *manifest) generateOuter() error {
if m.pbInner == nil { if m.pbInner == nil {
return errors.New("internal error") return errInternal
}
// Use fixed UUID if provided, otherwise generate a new one
var manifestUUID uuid.UUID
if len(m.fixedUUID) == uuidLength {
copy(manifestUUID[:], m.fixedUUID)
} else {
manifestUUID = uuid.New()
} }
// Generate UUID and set on inner message
manifestUUID := uuid.New()
m.pbInner.Uuid = manifestUUID[:] m.pbInner.Uuid = manifestUUID[:]
innerData, err := proto.MarshalOptions{Deterministic: true}.Marshal(m.pbInner) innerData, err := proto.MarshalOptions{Deterministic: true}.Marshal(m.pbInner)
if err != nil { if err != nil {
return err return fmt.Errorf("serialize: marshal inner: %w", err)
} }
// Compress the inner data // Compress the inner data
idc := new(bytes.Buffer) idc := new(bytes.Buffer)
zw, err := zstd.NewWriter(idc, zstd.WithEncoderLevel(zstd.SpeedBestCompression)) zw, err := zstd.NewWriter(idc, zstd.WithEncoderLevel(zstd.SpeedBestCompression))
if err != nil { if err != nil {
return err return fmt.Errorf("serialize: create compressor: %w", err)
} }
_, err = zw.Write(innerData) _, err = zw.Write(innerData)
if err != nil { if err != nil {
return err return fmt.Errorf("serialize: compress: %w", err)
} }
_ = zw.Close() _ = zw.Close()
compressedData := idc.Bytes() compressedData := idc.Bytes()
// Hash the compressed data for integrity verification before decompression // Hash the compressed data for integrity verification before decompression
h := sha256.New() h := sha256.New()
if _, err := h.Write(compressedData); err != nil {
return err _, err = h.Write(compressedData)
if err != nil {
return fmt.Errorf("serialize: hash write: %w", err)
} }
sha256Hash := h.Sum(nil) sha256Hash := h.Sum(nil)
m.pbOuter = &MFFileOuter{ m.pbOuter = &MFFileOuter{
@@ -90,29 +135,40 @@ func (m *manifest) generateOuter() error {
// Sign the manifest if signing options are provided // Sign the manifest if signing options are provided
if m.signingOptions != nil && m.signingOptions.KeyID != "" { if m.signingOptions != nil && m.signingOptions.KeyID != "" {
sigString, err := m.signatureString() return m.signOuter()
if err != nil {
return fmt.Errorf("failed to generate signature string: %w", err)
}
sig, err := gpgSign([]byte(sigString), m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to sign manifest: %w", err)
}
m.pbOuter.Signature = sig
fingerprint, err := gpgGetKeyFingerprint(m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to get key fingerprint: %w", err)
}
m.pbOuter.Signer = fingerprint
pubKey, err := gpgExportPublicKey(m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to export public key: %w", err)
}
m.pbOuter.SigningPubKey = pubKey
} }
return nil return nil
} }
// signOuter signs the outer message with the configured GPG key and
// embeds the signature, signer fingerprint, and public key.
func (m *manifest) signOuter() error {
sigString, err := m.signatureString()
if err != nil {
return fmt.Errorf("failed to generate signature string: %w", err)
}
sig, err := gpgSign([]byte(sigString), m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to sign manifest: %w", err)
}
m.pbOuter.Signature = sig
fingerprint, err := gpgGetKeyFingerprint(m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to get key fingerprint: %w", err)
}
m.pbOuter.Signer = fingerprint
pubKey, err := gpgExportPublicKey(m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to export public key: %w", err)
}
m.pbOuter.SigningPubKey = pubKey
return nil
}

View File

@@ -27,13 +27,19 @@ func (b BaseURL) JoinPath(path RelFilePath) (FileURL, error) {
base.Path += "/" base.Path += "/"
} }
// Parse and encode the relative path // Encode each path segment individually to preserve slashes
ref, err := url.Parse(url.PathEscape(string(path))) segments := strings.Split(string(path), "/")
for i, seg := range segments {
segments[i] = url.PathEscape(seg)
}
ref, err := url.Parse(strings.Join(segments, "/"))
if err != nil { if err != nil {
return "", err return "", err
} }
resolved := base.ResolveReference(ref) resolved := base.ResolveReference(ref)
return FileURL(resolved.String()), nil return FileURL(resolved.String()), nil
} }

59
mfer/url_test.go Normal file
View File

@@ -0,0 +1,59 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
func TestBaseURLJoinPath(t *testing.T) {
t.Parallel()
tests := []struct {
base BaseURL
path RelFilePath
expected string
}{
{"https://example.com/dir/", testFileName, "https://example.com/dir/file.txt"},
{"https://example.com/dir", testFileName, "https://example.com/dir/file.txt"},
{"https://example.com/", "sub/file.txt", "https://example.com/sub/file.txt"},
{
"https://example.com/dir/",
"file with spaces.txt",
"https://example.com/dir/file%20with%20spaces.txt",
},
}
for _, tt := range tests {
t.Run(string(tt.base)+"+"+string(tt.path), func(t *testing.T) {
t.Parallel()
result, err := tt.base.JoinPath(tt.path)
require.NoError(t, err)
assert.Equal(t, tt.expected, string(result))
})
}
}
func TestBaseURLString(t *testing.T) {
t.Parallel()
b := BaseURL("https://example.com/")
assert.Equal(t, "https://example.com/", b.String())
}
func TestFileURLString(t *testing.T) {
t.Parallel()
f := FileURL("https://example.com/file.txt")
assert.Equal(t, "https://example.com/file.txt", f.String())
}
func TestManifestURLString(t *testing.T) {
t.Parallel()
m := ManifestURL("https://example.com/index.mf")
assert.Equal(t, "https://example.com/index.mf", m.String())
}

Binary file not shown.

10
package.json Normal file
View File

@@ -0,0 +1,10 @@
{
"name": "mfer",
"version": "0.1.0",
"private": true,
"description": "Development tooling for the mfer repository: prettier, used by script/fmt and script/fmt-check to format and verify Markdown and JSON.",
"license": "WTFPL",
"devDependencies": {
"prettier": "3.9.6"
}
}

159
script/bootstrap Executable file
View File

@@ -0,0 +1,159 @@
#!/bin/sh
# script/bootstrap: install all dependencies needed to build and develop
# this repo. Idempotent: every install is guarded by a check so already
# installed tools are skipped. Base tooling comes from nix, apt, brew,
# or apk (detected in that order); assumes NOTHING is present (not git,
# make, node, yarn, go, or python). Node is used directly if installed;
# otherwise a pinned version is installed via nvm (installing nvm
# itself first, from a hash-verified release archive, never curl | sh).
#
# Uncomment the language sections in main() that apply to this repo.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Pinned versions, 2026-07-06. Never "latest" or "lts"; exact versions.
NODE_VERSION="22.17.0"
NVM_VERSION="0.40.3"
# sha256 of https://github.com/nvm-sh/nvm/archive/refs/tags/v0.40.3.tar.gz
NVM_SHA256="5f4d6aaa04a177dc93c985e31dbc411ab6b8c6e1e21d8015dbc1372625fcd1d0"
YARN_VERSION="1.22.22"
PKGMGR=""
SUDO=""
detect_pkgmgr() {
[ -n "$PKGMGR" ] && return 0
if command -v nix-env >/dev/null 2>&1; then
PKGMGR="nix"
elif command -v apt-get >/dev/null 2>&1; then
PKGMGR="apt"
elif command -v brew >/dev/null 2>&1; then
PKGMGR="brew"
elif command -v apk >/dev/null 2>&1; then
PKGMGR="apk"
else
echo "bootstrap: no supported package manager (nix, apt, brew, apk)" >&2
exit 1
fi
if [ "$PKGMGR" = "apt" ]; then
export DEBIAN_FRONTEND=noninteractive
if [ "$(id -u)" != "0" ]; then
SUDO="sudo"
fi
fi
}
# pkg_install <nix-attr> <apt-pkg> <brew-formula> <apk-pkg>
pkg_install() {
detect_pkgmgr
case "$PKGMGR" in
nix) nix-env -iA "nixpkgs.$1" ;;
apt) $SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2" ;;
brew) brew install "$3" ;;
apk) apk add --no-cache "$4" ;;
esac
}
missing() {
! command -v "$1" >/dev/null 2>&1
}
# verify_sha256 <file> <expected-hash>
verify_sha256() {
if command -v sha256sum >/dev/null 2>&1; then
actual="$(sha256sum "$1" | cut -d' ' -f1)"
else
actual="$(shasum -a 256 "$1" | cut -d' ' -f1)"
fi
if [ "$actual" != "$2" ]; then
echo "bootstrap: sha256 mismatch for $1" >&2
echo " expected: $2" >&2
echo " actual: $actual" >&2
exit 1
fi
}
# nvm is a bash script; run a command in a bash with nvm loaded
nvm_sh() {
bash -c ". \"\$HOME/.nvm/nvm.sh\" && $*"
}
ensure_nvm() {
[ -s "$HOME/.nvm/nvm.sh" ] && return 0
# nvm prerequisites; nvm itself requires bash, so install it too
if missing bash; then pkg_install bash bash bash bash; fi
if missing curl; then pkg_install curl curl curl curl; fi
if missing git; then pkg_install git git git git; fi
tmp="$(mktemp -d)"
curl -fsSL -o "$tmp/nvm.tar.gz" \
"https://github.com/nvm-sh/nvm/archive/refs/tags/v${NVM_VERSION}.tar.gz"
verify_sha256 "$tmp/nvm.tar.gz" "$NVM_SHA256"
mkdir -p "$HOME/.nvm"
tar -xzf "$tmp/nvm.tar.gz" -C "$HOME/.nvm" --strip-components=1
rm -rf "$tmp"
}
ensure_node() {
if ! missing node; then return 0; fi
ensure_nvm
nvm_sh "nvm install $NODE_VERSION"
}
ensure_yarn() {
if ! missing yarn; then return 0; fi
if ! missing corepack; then
corepack enable
corepack prepare "yarn@$YARN_VERSION" --activate
elif [ -s "$HOME/.nvm/nvm.sh" ]; then
nvm_sh "nvm use $NODE_VERSION >/dev/null && corepack enable && \
corepack prepare yarn@$YARN_VERSION --activate"
else
npm install -g "yarn@$YARN_VERSION"
fi
}
install_js_deps() {
if missing yarn && [ -s "$HOME/.nvm/nvm.sh" ]; then
nvm_sh "nvm use $NODE_VERSION >/dev/null && cd \"$ROOT\" && \
yarn install --frozen-lockfile"
else
yarn install --frozen-lockfile
fi
}
main() {
cd "$ROOT"
# Base tooling (every repo)
if missing git; then pkg_install git git git git; fi
if missing make; then pkg_install gnumake make make make; fi
# ---- JS / docs repos ----
# This is a Go repo, but node and yarn are required anyway: prettier
# formats the Markdown and JSON, and script/fmt-check verifies it.
# The version is pinned by package.json/yarn.lock, whose integrity
# hashes --frozen-lockfile enforces.
ensure_node
ensure_yarn
install_js_deps
# ---- Go repos ----
if missing go; then pkg_install go golang go go; fi
# golangci-lint: packaged in nix, brew, and apk. On apt there is no
# package: download a specific release archive from GitHub and
# verify its hash (verify_sha256), never curl | sh.
if missing golangci-lint; then
pkg_install golangci-lint golangci-lint golangci-lint golangci-lint
fi
go mod download
# ---- Python repos ----
# if missing python3; then pkg_install python3 python3 python3 python3; fi
# python3 -m venv .venv
# ./.venv/bin/pip install -e '.[dev]'
echo "bootstrap complete"
}
main "$@"

15
script/check Executable file
View File

@@ -0,0 +1,15 @@
#!/bin/sh
# script/check: run all checks (test, lint, fmt-check). Our own
# extension to scripts-to-rule-them-all. Must not modify any files.
# Generic: usually needs no adaptation.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
main() {
"$SCRIPT_DIR/test"
"$SCRIPT_DIR/lint"
"$SCRIPT_DIR/fmt-check"
}
main "$@"

14
script/cibuild Executable file
View File

@@ -0,0 +1,14 @@
#!/bin/sh
# script/cibuild: run the CI build. The Dockerfile runs script/check
# (via make check), so a successful build implies all checks pass.
# Generic: needs no adaptation. The Gitea workflow runs this on push.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
docker build .
}
main "$@"

15
script/docker Executable file
View File

@@ -0,0 +1,15 @@
#!/bin/sh
# script/docker: build the Docker image tagged with the project name.
# Identical in all repos; the tag comes from script/projectname.
# Generic: needs no adaptation.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
docker build -t "$("$SCRIPT_DIR/projectname")" .
}
main "$@"

27
script/fmt Executable file
View File

@@ -0,0 +1,27 @@
#!/bin/sh
# script/fmt: format all files (writes).
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
# Regenerate mfer/mf.pb.go from mfer/mf.proto if it is missing or stale
# (mirrors the old Makefile prerequisite; the generated file is
# committed, so this is normally a no-op).
ensure_pb() {
if [ ! -f mfer/mf.pb.go ] ||
[ -n "$(find mfer/mf.proto -newer mfer/mf.pb.go 2>/dev/null)" ]; then
(cd mfer && go generate .)
fi
}
main() {
cd "$ROOT"
ensure_pb
gofumpt -l -w mfer internal cmd
golangci-lint run --fix
# Markdown and JSON, over the same file set script/fmt-check verifies.
"$SCRIPT_DIR/prettier" --write
}
main "$@"

14
script/fmt-check Executable file
View File

@@ -0,0 +1,14 @@
#!/bin/sh
# script/fmt-check: check formatting (read-only). Same scope as
# script/fmt, but fails instead of writing: Go via script/fmt-check-go,
# Markdown and JSON via script/prettier.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
main() {
"$SCRIPT_DIR/fmt-check-go"
"$SCRIPT_DIR/prettier" --check
}
main "$@"

29
script/fmt-check-go Executable file
View File

@@ -0,0 +1,29 @@
#!/bin/sh
# script/fmt-check-go: check Go formatting (read-only). Split out from
# script/fmt-check so the Docker lint stage, whose image has no node and
# therefore no prettier, can run the Go half on its own.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Regenerate mfer/mf.pb.go from mfer/mf.proto if it is missing or stale
# (mirrors the old Makefile prerequisite; the generated file is
# committed, so this is normally a no-op).
ensure_pb() {
if [ ! -f mfer/mf.pb.go ] ||
[ -n "$(find mfer/mf.proto -newer mfer/mf.pb.go 2>/dev/null)" ]; then
(cd mfer && go generate .)
fi
}
main() {
cd "$ROOT"
ensure_pb
if [ -n "$(gofmt -l .)" ]; then
echo "gofmt: files need formatting:" >&2
gofmt -l . >&2
exit 1
fi
}
main "$@"

17
script/install-precommit Executable file
View File

@@ -0,0 +1,17 @@
#!/bin/sh
# script/install-precommit: install the git pre-commit hook that runs
# script/precommit. Our own extension to scripts-to-rule-them-all.
# Generic: needs no adaptation.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
hook=".git/hooks/pre-commit"
printf '#!/bin/sh\nset -e\nscript/precommit\n' > .git/hooks/pre-commit
chmod +x .git/hooks/pre-commit
echo "pre-commit hook installed: runs script/precommit"
}
main "$@"

17
script/lint Executable file
View File

@@ -0,0 +1,17 @@
#!/bin/sh
# script/lint: run the linter.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
golangci-lint run
if [ -n "$(gofmt -l .)" ]; then
echo "gofmt: files need formatting:" >&2
gofmt -l . >&2
exit 1
fi
}
main "$@"

20
script/precommit Executable file
View File

@@ -0,0 +1,20 @@
#!/bin/sh
# script/precommit: run by the git pre-commit hook; fails the commit if
# checks fail. Our own extension to scripts-to-rule-them-all. Go repo
# extras run first: go mod tidy and go fmt, failing the commit if they
# change go.mod or go.sum.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
go mod tidy
go fmt ./...
git diff --exit-code -- go.mod go.sum ||
{ echo "go mod tidy changed files; stage and retry" >&2; exit 1; }
"$SCRIPT_DIR/check"
}
main "$@"

69
script/prettier Executable file
View File

@@ -0,0 +1,69 @@
#!/bin/sh
# script/prettier: run prettier over this repo's canonical file set.
#
# Takes exactly one mode argument, --write or --check, and applies the
# same patterns in both modes. script/fmt and script/fmt-check both go
# through here, so the set of files that get formatted and the set that
# get verified cannot drift apart.
#
# Failures are never swallowed: a missing prettier is an error, not a
# silent skip. A formatter that quietly does nothing is worse than one
# that fails loudly.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
usage() {
echo "usage: script/prettier --write|--check" >&2
exit 2
}
# Prefer the version pinned by package.json/yarn.lock so that CI and
# developer machines format identically. Fall back to a prettier on PATH,
# but say so, because a different version formats differently.
find_prettier() {
if [ -x "$ROOT/node_modules/.bin/prettier" ]; then
printf '%s\n' "$ROOT/node_modules/.bin/prettier"
return 0
fi
if command -v prettier >/dev/null 2>&1; then
echo "prettier: node_modules/.bin/prettier is absent; using the" \
"prettier on PATH, which may be a different version than the" \
"one pinned in package.json. Run script/bootstrap to install" \
"the pinned version." >&2
command -v prettier
return 0
fi
return 1
}
main() {
[ "$#" -eq 1 ] || usage
case "$1" in
--write | --check) mode="$1" ;;
*) usage ;;
esac
cd "$ROOT"
if ! prettier_bin="$(find_prettier)"; then
echo "prettier: not found." >&2
echo " Install it with: script/bootstrap" >&2
echo " (installs the version pinned in package.json/yarn.lock)" >&2
exit 1
fi
# Markdown and JSON, repo-wide rather than root-only, so files in
# subdirectories (docs/, once it exists) are covered too. Exclusions
# live in .prettierignore; REPO_POLICIES.md is excluded there because
# it is a verbatim copy of an upstream document.
#
# --no-error-on-unmatched-pattern is deliberately NOT used: both
# patterns always match at least one tracked file (README.md,
# package.json), so an empty match means the glob broke, and prettier
# erroring out is exactly what we want rather than a vacuous pass.
"$prettier_bin" "$mode" "**/*.md"
"$prettier_bin" "$mode" "**/*.json"
}
main "$@"

12
script/projectname Executable file
View File

@@ -0,0 +1,12 @@
#!/bin/sh
# script/projectname: output the name of this project. Our own
# extension to scripts-to-rule-them-all. Other scripts that need the
# name (e.g. script/docker) call this, so they can stay identical
# across all repos.
set -eu
main() {
echo "mfer"
}
main "$@"

14
script/setup Executable file
View File

@@ -0,0 +1,14 @@
#!/bin/sh
# script/setup: set up the repo for development after a fresh clone:
# installs dependencies (script/bootstrap) and the git pre-commit hook.
# Add any repo-specific initialization (db init, .env template) here.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
main() {
"$SCRIPT_DIR/bootstrap"
"$SCRIPT_DIR/install-precommit"
}
main "$@"

23
script/test Executable file
View File

@@ -0,0 +1,23 @@
#!/bin/sh
# script/test: run the test suite.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Regenerate mfer/mf.pb.go from mfer/mf.proto if it is missing or stale
# (mirrors the old Makefile prerequisite; the generated file is
# committed, so this is normally a no-op).
ensure_pb() {
if [ ! -f mfer/mf.pb.go ] ||
[ -n "$(find mfer/mf.proto -newer mfer/mf.pb.go 2>/dev/null)" ]; then
(cd mfer && go generate .)
fi
}
main() {
cd "$ROOT"
ensure_pb
go test -v --timeout 10s ./...
}
main "$@"

Binary file not shown.

8
yarn.lock Normal file
View File

@@ -0,0 +1,8 @@
# THIS IS AN AUTOGENERATED FILE. DO NOT EDIT THIS FILE DIRECTLY.
# yarn lockfile v1
prettier@3.9.6:
version "3.9.6"
resolved "https://registry.yarnpkg.com/prettier/-/prettier-3.9.6.tgz#b3ea5146515d40fc53f18aa63f74dfab1e10dbf6"
integrity sha512-OpN0zzVdiaiAhxpuuj5efpIS4sY9j7bY6uR5mnj5yPzGkdkjNKSJeUThPb60Jw29QuAZgA4o+/iB49kFiaBX6g==