Author SHA1 Message Date
sneak 265e4f7d8a Move the CLI to urfave/cli v3 (closes #110)
check / check (push) Failing after 2s
Ported per the library's v2-to-v3 migration guide: the app is a root
cli.Command, actions take a context and the command, and flag
environment variables become value sources. -v, -q and the version flag
are local so, as before, only the commands defining them accept them.
Every command stops reading flags at its first argument, as v2 did.
ErrWriter is stdout so usage errors print with their help. The action's
context reaches the manifest download in check, export and list.
testify rises to v1.12.1, which v3 requires; the urfave_cli_no_docs
build tag, which v3 lacks, is dropped. v3 accepts a flag given under two
names, so -v --verbose gives debug output, and the test pinning the
refusal becomes a TestVerboseCount case.

Model: opus-5-5
2026-10-04 19:34:32 +00:00
clawbot ab72692439 Pin the remaining developer tool installs (closes #68)
check / check (push) Failing after 3s
gofumpt and protoc-gen-go are tools of a separate Go module in bin/tools,
so `go tool` builds them from source checked against bin/tools/go.sum and
mfer's own module gains no dependencies. script/bootstrap downloads that
module, and unpacks protoc 33.4 into bin/protoc from its release archive
after checking the sha256 it holds for the platform. script/generate runs
that protoc with the pinned plugin, so mfer/mf.go and its go:generate line
go. script/prettier runs only the node_modules prettier, fails when it
differs from the package.json pin, finds the node bootstrap installed
through nvm by the version in .nvmrc, and runs prettier once. Comment and
package.json fixes.

Model: opus-5-5
2026-10-04 20:48:51 +02:00
clawbot 9bb0ab3a03 fetch refuses a manifest that lists another file's temp name (closes #151)
check / check (push) Failing after 2s
fetch downloads each file to a temp name beside it and first removes
whatever is there. A manifest listing both a.txt and .a.txt.tmp had
fetch delete the second while fetching the first, then exit 0 with a
tree check rejects.

The refusal of a manifest that lists the saved manifest's own name or
temp name now covers this too: a listed file, or a directory a listed
file is in, may not sit at any name fetch writes besides the listed
files themselves. Temp names come from tempPathFor, names are compared
ignoring case as before, and the refusal still happens before the
destination is created or any file requested.

Model: opus-5-5
2026-10-04 20:02:17 +02:00
clawbot d3394bd2a2 Makefile: thin shims only, add make build (closes #73)
check / check (push) Failing after 2s
The Makefile is now one shim per script/ entrypoint, in the order of
the canonical model Makefile, plus build, generate and fuzz. The new
script/build writes bin/mfer with the urfave_cli_no_docs tag and the
git describe revision that script/docker stamps. script/generate now
adds the Go bin directory to PATH itself, and the Docker lint stage
calls script/gofumpt --check directly. Removed: the tarball-caching
rules and their ignore entries, godoc, ci, fixme, clean, run, default,
fmt-check-go, fmt-check-md, and bin/gitrev.sh, which only the
Makefile called.

Model: opus-5-5
2026-10-04 18:38:13 +02:00
clawbot a3e37d1ab8 fetch: destination directory, skip files already present, save the manifest, require a signer (closes #101)
check / check (push) Failing after 3s
fetch takes --dest (default .) and writes every file there through the
existing symlink and hard-link guards, which now work relative to that
directory. A file already there with the listed size and hash is
skipped; a leftover temp file is still replaced. Once every file
verifies, the manifest is saved as index.mf through the same temp file
and rename, so check runs on the result; a manifest that lists index.mf
or its temp name at the top of the tree is refused, since saving would
replace that file. --require-signature is shared with check and
enforced through verifyRequiredSigner. Both refusals come before any
file is downloaded or anything is written.

Model: opus-5-5
2026-10-04 18:31:51 +02:00
clawbot acff23d92b Check Go formatting with gofumpt, as make fmt writes it (closes #70)
check / check (push) Failing after 2s
The Go format check ran plain gofmt, which accepts code that gofumpt,
the formatter script/fmt runs, rewrites; and the two covered different
files. Both now go through script/gofumpt, which runs one pinned
gofumpt version over every Go file in --write or --check mode, as
script/prettier does for Markdown. It replaces script/fmt-check-go;
make fmt-check-go and the Docker lint stage call it. go run builds the
pinned version, so nothing has to install gofumpt, and the check fails
when gofumpt cannot run instead of passing. Generated mfer/mf.pb.go
passes: gofumpt holds generated files to gofmt's rules. The check is
not redundant: .golangci.yml enables no golangci-lint formatters.

Model: opus-5-5
2026-10-04 17:54:50 +02:00
clawbot 8fe0244291 check always warns about files the manifest does not list (closes #103)
check / check (push) Failing after 2s
check now always looks under the base directory for files the manifest
does not list, hidden files and directories included, and prints one
warning per file. The result still depends only on the listed files;
--no-extra-files turns each unlisted file into a failure, and --quiet
hides the warnings but not the failures. A directory that cannot be
listed is reported the same way, and the search goes on past it.

The manifest itself is left out by file identity, as gen and freshen
do; a symlink under the base is compared by what it points to. A base
directory named through a symlink is resolved before the search.

Model: opus-5-5
2026-10-04 17:48:51 +02:00
clawbot 400a2f8f63 Move FORMAT.md to docs/ and contrib/ to bin/, fix bash script style (closes #74)
check / check (push) Failing after 3s
FORMAT.md moves to docs/ and every reference follows; its two relative
links to mfer/mf.proto now point up one directory, its text is otherwise
unchanged. contrib/usage.sh moves to bin/, not docs/: it is a script you
run (it builds mfer, then generates and checks a manifest of the repo),
not reading material. Both scripts now use #!/usr/bin/env bash,
set -euo pipefail and a main function. bin/gitrev.sh still exits non-zero
outside a git checkout, so the Makefile still falls back to unknown.
.gitignore now ignores only the built bin/mfer rather than all of bin/,
which holds tracked scripts.

Model: opus-5-5
2026-10-04 17:09:20 +02:00
clawbot ba3d24ca42 End-to-end tests for freshen and fetch (closes #66)
check / check (push) Failing after 3s
The freshen tests built a fixture and never ran freshen. They now run
gen and freshen through the command entry point on a temp dir: an
unchanged tree leaves the manifest's entries as they were, and a
modified file (also one edited without changing its size), a new or a
deleted file gives a manifest that lists exactly the tree, with sizes,
hashes and mtimes from disk, on which check passes.

New fetch tests run the command against an httptest server: a nested
tree lands complete, a hash mismatch exits non-zero and leaves no file,
and a partly filled destination gets every listed file downloaded again
and replaced, while a file the manifest does not list is left alone.

Model: opus-5-5
2026-10-04 17:02:22 +02:00
clawbot 11041330a7 Move internal/log from apex/log to log/slog (closes #77)
check / check (push) Failing after 2s
internal/log now logs through log/slog. After Init, records go to a small
slog.Handler that prints the same lines as before; before Init they go to
slog.Default(). The helpers and their call sites are unchanged. With
NO_COLOR set on a terminal, log lines are no longer colored.

pterm is dropped: it only printed progress lines, which fmt now writes
directly, outside slog. apex/log, pterm and their indirect dependencies
leave go.mod; golang.org/x/term becomes direct.

simplelog is not imported: importing it replaces slog's default handler in
every program that imports package mfer. The logger stays process-global,
since injecting it would change package mfer's API.

Progress and log lines now share the logger's lock, so the check progress
test runs check ten times to catch a missing wait.

Model: opus-5-5
2026-10-04 16:31:54 +02:00
clawbot 5c7ef94799 Set go_package to the sneak.berlin/go/mfer module path (closes #79)
check / check (push) Failing after 1s
mf.proto named git.eeqj.de/sneak/mfer/mfer as its Go package, which
disagrees with the sneak.berlin/go/mfer module path go.mod declares.
go_package is now sneak.berlin/go/mfer/mfer, the import path of the
mfer/ package. mf.pb.go is regenerated with make generate, which
changes only the go_package bytes in its embedded descriptor, and
mf.proto.sha256 records the new mf.proto hash.

Model: opus-5-5
2026-10-04 16:02:21 +02:00
clawbot 82b9434444 fetch: client timeout, retry with backoff, url.JoinPath (closes #63)
check / check (push) Failing after 2s
fetch now makes every request through an http.Client with a time
limit, ten minutes by default and set with --timeout, which must be
greater than zero. A connection error, a timeout, or a 5xx or 429
response is retried, up to five tries in all, after a random wait
whose limit doubles from one second, or after the wait the server's
Retry-After asks for, up to one minute. Each try of a file starts a
new temp file, so a retry never keeps a partial file; the size and
hash checks are unchanged. Manifest and file URLs are built with
URL.JoinPath, so trailing slashes, query strings and names that need
escaping all work.

Model: opus-5-5
2026-10-04 15:48:55 +02:00
clawbot 6a2307a82b check waits for its progress output before the summary (closes #140)
check / check (push) Failing after 1s
With --progress, runCheck started the progress goroutine but waited
only for the results goroutine, so check could log its summary, or
return, before the last progress line was written and cleared. The
progress goroutine now signals a sync.WaitGroup when it finishes, and
runCheck waits on it as soon as Check returns, as generate does.

The new test sends stdout and stderr to one buffer and slows down
stdout writes, so without the wait the progress output lands after
the summary.

Model: opus-5-5
2026-10-04 15:31:57 +02:00
clawbot 4bf87d1ccf AddFileWithHash rejects hashes that are not multihashes (closes #129)
check / check (push) Failing after 2s
AddFileWithHash took any non-empty bytes as a hash, so the builder could
write a manifest that mfer refuses to load. It now decodes the hash with
go-multihash and also requires a digest of at least 32 bytes, the SHA-256
length the reader's decoding-cost limit assumes: a valid but shorter
multihash, such as an empty identity hash or SHA-1, still makes a
manifest of one-character paths too costly to load. When mfer freshen
meets such a hash in an existing manifest, its error names the manifest
entry and says to regenerate the manifest with mfer generate. Test
fixtures whose stand-in hashes were not valid multihashes now use a
SHA-256 multihash.

Model: opus-5-5
2026-10-04 14:48:49 +02:00
clawbot e412d20c20 Default the manifest to index.mf and keep it out of its own listing (closes #100)
check / check (push) Failing after 1s
mfer gen now writes index.mf instead of .index.mf, the name fetch
requests and the README calls the standard filename. Given a
directory, check, freshen, list and export look only for index.mf;
.index.mf is no longer recognized. Because index.mf is not hidden, gen
and freshen leave out of their own listing the manifest they write and
a temp file an interrupted run left beside it, matched by file
identity (os.SameFile) however the path is spelled. The scanner takes
these as a plain ExcludePaths list; manifestTempPath alone names the
temp file, for both writes, both skips and the tests.

Model: opus-5-5
2026-10-04 13:48:52 +02:00
clawbot 0501568203 Never regenerate mf.pb.go during checks; fail when it is stale (closes #71)
check / check (push) Failing after 1s
The test and format scripts regenerated mfer/mf.pb.go whenever
mfer/mf.proto looked newer by mtime, which a fresh checkout often causes,
so make check could rewrite a committed file and needed protoc. Nothing
regenerates it any more except make generate (script/generate), which
refuses to run unless protoc 33.4 and protoc-gen-go v1.36.11, the
versions that wrote the committed file, are on PATH, and records the
hash of mf.proto in mfer/mf.proto.sha256. A Go test compares that hash
with mf.proto and fails, naming make generate, when they differ. The
mtime rule, the unused protoc-gen-go v1.28.1 install rule, make clean's
deletion of mf.pb.go and the Dockerfile's touch workarounds are removed.

Model: opus-5-5
2026-10-04 13:31:57 +02:00
clawbot 0a9963002c NewChecker takes CheckerOptions instead of positional arguments (closes #78)
check / check (push) Failing after 1s
NewChecker now takes *CheckerOptions (ManifestPath, BasePath, Fs), named
like ScannerOptions. A nil Fs still means the OS filesystem, as before and
as in ScannerOptions; nil options or an empty path return an error naming
the missing path.

Audit of the other exported constructors in mfer: NewManifestFromFile took
a filesystem and a path positionally; it now takes
*ManifestFromFileOptions (Path, Fs) with the same nil and empty rules.
NewBuilder and NewScanner take no arguments, NewScannerWithOptions already
takes options, and NewManifestFromReader takes one reader, which the style
guide exempts; these are unchanged.

Model: opus-5-5
2026-10-04 12:19:31 +02:00
clawbot 76116005c8 Reject manifests whose file entries decode far larger than their bytes (closes #123)
check / check (push) Failing after 2s
Before decoding the manifest, the parser adds up what decoding sets
aside for each file entry, hash, timestamp and MIME type, however short
its encoding, and refuses the manifest once that passes 8 times the
decompressed size; manifests mfer writes come to at most about 7.15
times. Empty entries decoded to about 50 times their size: under 1 KB of
manifest allocated about 500 MB. Fields the decoder does not know are
dropped; kept, they took up to 5 times more. A test refuses entries
counted just over 8 times and loads them just under. The fuzz ceiling
rises from 16 to 20 times the input and decompressed data; seeds of
empty entries and of empty hashes fail it without the fix.

Model: opus-5-5
2026-10-04 12:02:27 +02:00
clawbot 7088857692 Count -v once and document -v -v for debug output (closes #125)
check / check (push) Failing after 2s
urfave/cli before v2.25.5 counted a flag given by its alias twice, so
one -v or --verbose already gave debug output. Bump it to v2.27.7,
which counts it once: one -v gives verbose output, two give debug.

The -v help text now says -v -v instead of -vv, which stays refused:
urfave/cli's option for combined short flags would let a flag that
takes a value read the next letter as its value.

-v and --verbose together stay refused: urfave/cli v2 refuses a flag
given under two of its names, and one flag with an alias keeps help and
parsing simple.

The bump changes some help output; generate and fetch now name their
arguments. Tests start each run at the default log level.

Model: opus-5-5
2026-10-04 11:48:52 +02:00
clawbot 588c1bae74 Give the image CA certificates so fetch works over HTTPS (closes #131)
check / check (push) Failing after 2s
The final stage is scratch, which has no CA certificates, so fetch from
an HTTPS URL failed to verify any server. Copy the CA bundle from the
pinned builder image into the final stage.

Model: opus-5-5
2026-10-04 10:31:54 +02:00
clawbot 64eb5cbd40 Scanner status tests no longer depend on a 100 ms timer (closes #124)
check / check (push) Failing after 3s
The two status-send tests now call the send directly on a channel nobody
receives from, so a blocking send hangs the test into its timeout instead
of racing a 100 ms timer that a loaded host can miss.

The gpg test for a child holding gpg's output had the same problem: its
100 ms deadline could fire before the fake gpg wrote the PID of sleep,
and the cleanup then failed. The fake gpg now writes that PID to a named
pipe, and the test cancels only after reading it. It then waits up to
10 s for the call, so a call that waits for sleep fails the test and the
cleanup still kills sleep.

Model: opus-5-5
2026-10-04 09:48:51 +02:00
clawbot 45eac1f6f8 Run all linting in Docker via the Dockerfile lint stage (closes #90)
check / check (push) Failing after 3s
script/lint now builds only the lint stage of the main Dockerfile
(docker build --no-cache --target lint), whose build runs the linter, so
a successful build is a clean lint. It is uncached because a cached
build runs no linter, and a trap removes the image it tagged; the tag
carries the process ID so concurrent runs do not collide. The lint stage
calls golangci-lint directly, since make lint now needs Docker. Nothing
installs or runs golangci-lint on the host any more: bootstrap and the
Makefile drop the install, and script/fmt drops golangci-lint run --fix.

Model: opus-5-5
2026-10-04 09:31:54 +02:00
clawbot b91e92b070 Build a static binary so the scratch image runs (closes #126)
check / check (push) Failing after 1s
The final stage is scratch, which has no C library, but the builder
compiled mfer with cgo on (the golang image's default). The binary
imports net (mfer's own HTTP code does, and so does google/uuid), so
with cgo on it came out dynamically linked and the image could not
start. The image's go build now sets CGO_ENABLED=0; nothing in mfer
needs cgo. A new builder step runs ldd on the binary and fails the
build unless it reports a static executable.

Model: opus-5-5
2026-10-04 08:48:52 +02:00
clawbot a2732cf8da Add the required README sections (closes #75)
check / check (push) Successful in 1m28s
The Description first line now names the project, purpose, category,
WTFPL license, and author. A new Getting Started section gives a
copy-pasteable build-from-source block and gen/check/fetch usage. Problem
Statement and Proposed Solution move under a new Rationale heading, the
prose kept. A new Design section documents the package layout: the mfer/
library with its committed protobuf code, the internal/cli commands,
internal/log and internal/bork, and the cmd/mfer entrypoint. Authors is
renamed Author with the canonical link.

Getting Started uses go build because the make target that builds the
binary runs protoc, which script/bootstrap does not install.

Model: opus-4-8 (implementation); opus-5-5 (rebase, rework)
2026-10-04 06:32:07 +02:00
clawbot a789eb3083 Give -v to verbose, move --version to -V (closes #64)
check / check (push) Successful in 1m45s
urfave/cli's built-in version flag claimed -v, so "mfer -v --version"
failed to parse, and -v meant version at the root but verbose on
generate, check, freshen and fetch. Verbose is the more common meaning,
so the root now takes -v and -q too, and the version flag takes -V. A
-v or -q before generate, check, freshen, fetch or export applies to
that subcommand; list keeps its fixed quiet logging.

User-visible changes: -v no longer prints the version; use -V or
--version. "mfer version" prints "mfer version 0.1.0 (...)", the same
line as --version, instead of the bare "0.1.0 (...)".

Tests: runCLI now points the logger at io.Discard after each run, so
other tests' log lines cannot land in a finished run's output.

Model: opus-4-8 (implementation); opus-5-5 (rework)
2026-10-04 06:02:24 +02:00
clawbot d00982b329 Fuzz NewManifestFromReader and cap the zstd decoder (closes #65)
check / check (push) Successful in 1m5s
FuzzNewManifestFromReader fails when the parser returns both or neither
of a manifest and an error, or allocates more than a fixed multiple of
its input and the decompressed data it may read, plus room for the
decoder's window buffers. make test runs the seed corpus; make fuzz
fuzzes for one minute.

Parser bug: MaxDecompressedSize did not bound decompression; the zstd
decoder decoded a payload under 128 KiB in full before the LimitReader
read any of it. It now decodes only what the LimitReader reads and
refuses windows over the 8 MiB mfer writes with. Seeds: a frame claiming
8 GiB, two frames together over the limit, empty frames with growing
windows.

Model: opus-5-5
2026-10-04 05:31:51 +02:00
clawbot 51f69c960d Enforce real timeouts on gpg subprocess calls (closes #62)
check / check (push) Successful in 2m2s
Every gpg run now has a one-minute deadline (gpgTimeout) on top of its
caller's context and is killed when either ends. A timeout is reported
as "gpg timed out" under the failing operation instead of "signal:
killed". Only gpg itself is killed; WaitDelay (one second) stops the run
from waiting on a process gpg left behind that still holds its output,
such as a wrapper script that does not exec the real gpg.
Builder.Build and Checker.ExtractEmbeddedSigningKeyFP take a context, so
ToManifest's context now reaches signing and the contextcheck
suppression calling signing non-cancellable is gone. Manifest loading
takes no context, so its signature check is bounded by the timeout
alone.

Model: opus-5-5
2026-10-04 04:31:53 +02:00
clawbot 1adad7d3bc Remove TODO.md; track open work in the issues (closes #76)
check / check (push) Successful in 1m29s
TODO.md is deleted and the README TODO section now only points to the
issue tracker, where open work and open design questions live from now
on. AGENTS.md says so too, and says this overrides the shared policy's
README todo list for this repo. The README Open Questions section
becomes Original Design Questions, each noting where it now stands.

Every open item in both lists was already done or already owned by an
issue, apart from one, now #121,
and the chunk checksum question, now on
#81.

Model: opus-5-5
2026-10-04 03:32:09 +02:00
clawbot c31796998f Pin CLI error messages by driving their real call sites (closes #87)
check / check (push) Successful in 58s
errmsg_test.go now calls the function that emits each user-visible CLI
error message and asserts on the full text it returns, instead of
rendering format strings copied from production, so rewording any pinned
message fails the suite. The unknown-command message is driven through
the root command's action.

The signer-mismatch case signs a manifest with a throwaway gpg key and
is skipped where gpg is absent, like the repo's other signing tests.

The freshen mtime-presence test gives the scanned file an epoch mtime,
so it fails if recordEntry reads an absent manifest mtime as the epoch.

Model: opus-4-8 (first implementation); opus-5-5 (completion, this message)
2026-10-03 18:24:30 +02:00
clawbot a3749e9e9b cibuild: build with --no-cache like script/docker (closes #89)
check / check (push) Successful in 1m1s
A bare `docker build .` served the Dockerfile's check steps from the
layer cache on an unchanged tree and exited 0 without running them.
script/cibuild now computes the version on its own line and runs the
same build command as script/docker and the shared policy, --no-cache
included, so every CI build runs the checks. README.md describes
script/cibuild accordingly.

Model: opus-5-5
2026-10-03 17:41:55 +02:00
clawbot fa97c4519c Never write fetch's temp file into an existing file (closes #115)
check / check (push) Successful in 53s
downloadFile opened the temp file with os.Create, which opens and
empties a file already at that name. If that file was a hard link to a
file outside the destination directory, fetch overwrote the outside
file. It now removes whatever is at the temp name, which removes only
that name, and creates the temp file with O_EXCL, so the create fails if
anything is still or again there. A leftover temp file from an
interrupted run is still replaced. The new test puts a hard link at the
temp name and checks that fetch succeeds and the outside file is
unchanged. The command name is now the constant cmdFetch, like the other
command names, because lint requires it once a third test uses it.

Model: opus-5-5
2026-10-03 16:58:35 +02:00
clawbot 7e601929c8 Refuse fetch writes through a symlink in the destination (closes #86)
check / check (push) Successful in 1m10s
sanitizePath checks manifest paths only as text, so a symlink already
inside the destination directory could send fetch's writes outside it.
checkNoSymlinks now looks at each existing part of a path with os.Lstat
and refuses the path if any part is a symlink, wherever it points.
fetch runs it immediately before each write: creating the parent
directories, creating the temp file, and renaming it into place. The new
test puts such a symlink at each of those three places, and once inside
a plain directory, and checks that the fetch fails and nothing outside
changes. The G304 comment now states what holds. A symlink swapped in
between a check and its write is not caught; os.Root closes that once
the Go version is raised.

Model: opus-5-5
2026-10-03 16:08:09 +02:00
clawbot c3b5fe651a Stamp the tag or short commit in a plain docker build (closes #112)
check / check (push) Successful in 48s
.dockerignore now sends .git but not .git/config, which can hold a
credential and which git describe does not need. The build stage stamps
main.Gitrev from the VERSION build argument when one is given, otherwise
from git describe --tags --always, and fails if .git is present and no
version comes out. git there trusts /src whoever owns it, since a
context sent as a tar archive keeps its files' owners. script/docker is
replaced by the canonical copy, which passes VERSION; bin/gitrev.sh uses
--tags too, so every entrypoint stamps the same value for a clean
commit.

Model: opus-5-5
2026-10-02 08:53:34 +02:00
clawbot 0deacfc7ed Validate manifest entry paths on deserialize (closes #61)
check / check (push) Failing after 1s
Untrusted .mf files were parsed with no path validation, so an entry like ../../etc/passwd reached filepath.Join against the checker's base path and mfer check could stat and read outside it. ValidatePath ran only when building a manifest. It now runs on every entry as the manifest loads, so every consumer is covered. The whole manifest is rejected on the first bad entry instead of dropping it, which could hide files from a check; the error wraps errInvalidManifestPath and names the path.

Disclosure: a path that is not valid UTF-8 is refused earlier, by the protobuf string decoder, so that error does not name the path.

Model: opus-4-8 (implementation); fable-5-1 (summary)
2026-09-22 00:47:26 +02:00
clawbot 7de4d6ec1c Rewrite script/test to the canonical race-enabled pattern (closes #67)
check / check (push) Successful in 51s
script/test now runs go test -timeout 30s -race -cover ./..., quiet on success and rerunning verbose on failure. -race exposed a data race on the process-global apex/log logger: Init reconfigured it on every CLI run while other goroutines logged through it. internal/log now mutates the global under the write lock and reads it under the read lock; WithError is dropped because its Entry logged outside that lock, and run() reports a failed command's error via log.Errorf instead (still shown under -q). The corruption fuzz test is scaled to 1500 files / 100 iterations to fit the budget under -race; it pins the same claims at reduced breadth. Independent review ran the Docker lint gate uncached.

Model: opus-4-8 (implementation, review)
model: claude-fable-5
2026-09-21 15:17:50 +02:00
clawbot 6de3f1d714 Add .editorconfig, broaden .gitignore, drop dead Drone refs (closes #72)
check / check (push) Failing after 1s
Add the canonical .editorconfig from sneak/prompts verbatim. Broaden .gitignore to cover secrets (.env, .env.*, *.key, *.pem), OS files, editor files, and Go artifacts while keeping every repo-specific entry; no tracked file becomes ignored, and bin/gitrev.sh stays tracked despite the /bin/ rule. Remove the stale .drone.yml ignore entry and the DRONE_COMMIT_SHA branch from bin/gitrev.sh, left from the Gitea Actions migration; the GITREV and git-describe paths still work.

A reader may trip over .editorconfig saying four spaces for every file while Go files are tab-indented: gofmt governs Go, editorconfig does not, and the file is taken verbatim by policy.

Disclosure: the worker's cold docker build hit the 10s test timeout in the gpg signature test; this change has no Go code, and that timeout is tracked by issues 67 and 62.

Model: opus-4-8 (implementation); fable-5-1 (merge)
2026-09-21 09:56:15 +02:00
clawbot 5683d0f4ff Configure prettier and make fmt-check cover markdown (closes #69)
check / check (push) Successful in 4s
Adds .prettierrc and .prettierignore, a single script/prettier entrypoint shared by fmt and fmt-check so the two cannot drift, and prettier 3.9.6 pinned by yarn.lock integrity hash.

Markdown formatting is now gated in the authoritative Docker build via a new mdfmt stage, since the golangci-lint image has no node. REPO_POLICIES.md is ignored so local tooling cannot drift it from upstream.

Removes the || true that made the previous prettier invocation unable to fail.
2026-08-10 16:16:42 +02:00
clawbot de476708e9 Update golangci-lint to v2.12.2 with canonical config (closes #60)
check / check (push) Has been cancelled
Adopts golangci-lint v2.12.2 and the canonical .golangci.yml (default: all), and fixes all resulting findings across the tree.

Two intended behavior changes: absent MFFilePath.Mtime is handled explicitly in freshen, list and export rather than dereferenced (main panicked); gpg positional key IDs now follow an explicit -- end-of-options marker.

All twelve reworded user-visible error messages restored to byte-identical parity with main and pinned by tests.
2026-08-10 16:06:12 +02:00
sneak 6d19de74e7 scripts-to-rule-them-all (#58)
check / check (push) Successful in 6s
Reviewed-on: #58
Co-authored-by: sneak <sneak@sneak.berlin>
Co-committed-by: sneak <sneak@sneak.berlin>
2026-07-07 02:13:35 +02:00
sneak 688ebd90a4 TODO (#57)
check / check (push) Successful in 5s
Reviewed-on: #57
Co-authored-by: sneak <sneak@sneak.berlin>
Co-committed-by: sneak <sneak@sneak.berlin>
2026-07-06 21:19:31 +02:00
sneak 5dc4433477 move to standardized repo policies (#56)
check / check (push) Successful in 4s
Reviewed-on: #56
2026-06-28 09:35:33 +02:00
sneak 3be8064865 Merge branch 'main' into next
check / check (push) Successful in 42s
2026-06-28 09:35:17 +02:00
sneak 958a5f8fc4 move to standardized repo policies 2026-06-28 09:30:11 +02:00
9ce16a83ad Add TODO section to README with 1.0 roadmap, remove TODO.md (#54)
check / check (push) Successful in 4s
## Summary

Performs a design and status review of the codebase and adds a comprehensive TODO section to `README.md` listing remaining work for a 1.0 release.

### What changed

- **README.md**: Added a `TODO: Remaining Work for 1.0` section covering:
  - **7 design questions** requiring @sneak's input before implementation (manifest type export, Go module path, GPG vs pure-Go crypto, format framing, etc.) — each with an answer field for inline decisions
  - **Implementation tasks** organized by category: repo infrastructure, format & correctness, library, CLI, testing & robustness, documentation, and release checklist
  - Updated build status section (removed stale Drone CI badge, replaced with description of current Docker-based CI)
- **TODO.md**: Removed — items integrated into README TODO section
- **AGENTS.md**: Updated reference from `TODO.md` to README TODO section

### Design review findings

**What works well:**
- Core library (Builder, Scanner, Checker) is solid with good test coverage
- Format specification is well-designed (protobuf + zstd, multihash, deterministic serialization)
- CLI covers all major operations (gen, check, list, export, freshen, fetch)
- Test suite is thorough — builder, scanner, checker, GPG, CLI integration, corruption detection
- afero abstraction enables clean testing without filesystem side effects

**Key gaps for 1.0:**
- Missing repo infrastructure (`.golangci.yml`, `.editorconfig`, CI workflow)
- `manifest` type is unexported — consumers can't use it in their own type declarations
- GPG signing shells out to `gpg` subprocess — fragile and may not be installed
- Go module path inconsistency between `go.mod` and proto `go_package`
- `fetch` command lacks retry logic and has no HTTP timeout
- Missing fuzz tests for untrusted input deserialization
- Freshen CLI command has incomplete integration test coverage

closes #47
closes #50

Co-authored-by: user <user@Mac.lan guest wan>
Co-authored-by: clawbot <clawbot@noreply.git.eeqj.de>
Co-authored-by: Jeffrey Paul <sneak@noreply.example.org>
Reviewed-on: #54
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-04-07 00:43:56 +02:00
clawbot 01124c10d9 Add Gitea Actions CI workflow (#53)
check / check (push) Successful in 18s
Follow-up to [PR #51](#51) — addresses sneak's review comment requesting Gitea Actions to run the checks.

## Changes

Adds `.gitea/workflows/check.yml` that runs `docker build .` on every push, matching the standard CI pattern used across all `sneak/*` repos (identical to [chat](https://git.eeqj.de/sneak/chat/src/branch/main/.gitea/workflows/check.yml) and [dnswatcher](https://git.eeqj.de/sneak/dnswatcher/src/branch/main/.gitea/workflows/check.yml)).

Since the Dockerfile already runs `make check` (fmt-check, lint, test), a successful build implies all checks pass.

The `actions/checkout` action is pinned by SHA (`@11bd71901bbe5b1630ceea73d27597364c9af683` = v4.2.2) per REPO_POLICIES.

## Verification

`docker build .` passes clean.

Reviewed-on: #53
Co-authored-by: clawbot <sneak+clawbot@sneak.cloud>
Co-committed-by: clawbot <sneak+clawbot@sneak.cloud>
2026-03-20 06:57:27 +01:00
clawbotandclawbot 6ba32f5b35 Add REPO_POLICIES.md, rename CLAUDE.md to AGENTS.md, deduplicate (#51)
Closes #48

## Changes

- **Added `REPO_POLICIES.md`** — copied from the standard template at [sneak/prompts](https://git.eeqj.de/sneak/prompts/src/branch/main/prompts/REPO_POLICIES.md) (last_modified: 2026-03-10). This is the authoritative cross-project policy document covering repository structure, tooling, Docker, formatting, testing, and workflow standards.

- **Renamed `CLAUDE.md` → `AGENTS.md`** — deduplicated content:
  - Rules already covered by `REPO_POLICIES.md` (e.g. `git add -A`, Makefile targets) are no longer repeated
  - `AGENTS.md` retains only agent-specific workflow instructions: test-first bug fixing, no AI attribution in commits, per-change make fmt/test/lint workflow, and repo-specific notes (proto files, FORMAT.md, TODO.md)

- **Updated `README.md`** — added a reference to `REPO_POLICIES.md` in the Participation section

- **Formatting** — `make fmt` (prettier) applied to all markdown files

## Verification

`docker build .` passes clean — lint, fmt-check, and all tests green.

Co-authored-by: clawbot <clawbot@noreply.git.eeqj.de>
Reviewed-on: #51
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-03-17 05:07:43 +01:00
clawbotanduser e62c709d42 Remove committed .index.mf and add to .gitignore (#52)
Closes #49

`.index.mf` is a generated manifest file produced at runtime by `mfer gen`. It was committed to the repo but should not be tracked in version control.

Changes:
- `git rm .index.mf` to remove it from tracking
- Added `.index.mf` to `.gitignore` to prevent accidental re-commits

Co-authored-by: user <user@Mac.lan guest wan>
Reviewed-on: #52
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-03-17 05:06:37 +01:00
clawbotandclawbot 89903fa1cd Split Dockerfile: pre-built golangci-lint stage for faster CI (#45)
Closes #39

Splits the Dockerfile into a dedicated lint stage using the `golangci/golangci-lint` image directly, rather than copying the binary into the builder stage.

## Changes

### Dockerfile

- **Lint stage** (`AS lint`): Uses the pre-built `golangci/golangci-lint` image (pinned by sha256) to run `make fmt-check` and `make lint`. This is a self-contained stage with its own `go mod download` and source copy.
- **Builder stage** (`AS builder`): Runs only `make test` and the final binary build. No longer needs golangci-lint installed.
- **Stage dependency**: `COPY --from=lint /src/go.sum /dev/null` forces BuildKit to always execute the lint stage (without this, unused stages are silently skipped).
- Both stages touch `mfer/mf.pb.go` to prevent make from trying to regenerate via protoc.

With BuildKit, the lint and builder stages run in parallel after their shared `go mod download` layers complete, so lint/formatting failures surface much faster without blocking on test execution.

### Makefile

- Added `lint` to the `check` target prereqs: `check: test lint fmt-check` (was `check: test fmt-check`), matching the [REPO_POLICIES](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/REPO_POLICIES.md) requirement.

Co-authored-by: clawbot <clawbot@noreply.git.eeqj.de>
Reviewed-on: #45
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-03-15 18:56:25 +01:00
sneak b3d10106e1 next (#44)
Reviewed-on: #44
2026-03-15 18:49:19 +01:00
clawbot 9712c10fe3 Merge main into next: resolve conflicts, rewrite Dockerfile for Go 1.23
- Resolve merge conflicts (README.md, TODO.md, go.mod) keeping next's versions
- Rewrite Dockerfile: replace sneak/builder:2022-12-08 (Go 1.19) with
  golang@sha256-pinned (Go 1.23), add golangci-lint for future use
- Remove references to deleted vendor.tzst, modcache.tzst
- Simplify to standard multi-stage build: check + build + scratch final image
- Keep module path sneak.berlin/go/mfer from next branch
- Add Makefile targets: check, fmt-check, hooks (per REPO_POLICIES)
- Pin golangci-lint@v2.0.2 in devprereqs
- Dockerfile version/date comment on pinned image hash
- make check runs test + fmt-check (lint deferred to follow-up issue)
2026-03-14 17:38:40 -07:00
clawbot 43916c7746 1.0 quality polish — code review, tests, bug fixes, documentation (#32)
Comprehensive quality pass targeting 1.0 release:

- Code review and refactoring
- Fix open bugs (#14, #16, #23)
- Expand test coverage
- Lint clean
- README update with build instructions (#9)
- Documentation improvements

Branched from `next` (active dev branch).

Reviewed-on: #32
Co-authored-by: clawbot <clawbot@noreply.example.org>
Co-committed-by: clawbot <clawbot@noreply.example.org>
2026-03-01 23:58:37 +01:00
sneak 1f7ee256ec fix: update go directive from 1.17 to 1.22 (fixes #33) (#34)
Reviewed-on: #34
2026-02-20 11:35:44 +01:00
clawbot 28c6fbd220 fix: update go directive from 1.17 to 1.22 (fixes #33)
go.mod specified go 1.17 but protobuf generated code (mf.pb.go) uses
go1.18+ features (any) and go1.20+ features (unsafe.StringData).
Updated to go 1.22 and ran go mod tidy.
2026-02-20 02:25:21 -08:00
sneak 5aae442156 add links to metalink format (#7)
Reviewed-on: #7
2026-02-09 02:15:58 +01:00
104 changed files with 10483 additions and 2731 deletions
+66 -3
View File
@@ -1,3 +1,66 @@
*.tmp
*.dockerimage
.git
# .dockerignore does NOT use .gitignore semantics. Docker matches with
# moby/patternmatcher: filepath.Match plus `**`, so `*` does not cross
# `/` and an unprefixed pattern is anchored at the context root. Every
# depth-independent pattern therefore needs `**/`, or `config/.env` and
# `certs/server.key` still ship while this file reads as solved. Only
# genuinely root-anchored entries go unprefixed. Never transplant these
# into .gitignore, where `**/` is wrong.
#
# Matching is case-sensitive, so secrets use character ranges rather
# than an ALL-CAPS twin, which would still miss `Server.Key`.
#
# Extend with this repo's own host-built artifacts, written anchored:
# `/myapp`, never `**/myapp`, which also matches `cmd/myapp/` and
# deletes the package directory from the context.
# .git is sent without its config. Without a VERSION build argument the
# stage that compiles runs `git describe --tags --always` on .git, which
# does not need .git/config; that file can hold a credential, such as a
# password in a remote URL or the token the CI checkout step stores there.
.git/config
# Agent scratch: one full checkout of the repo per in-flight agent.
# Anchored because it occurs once where agents run at the repo root.
# KNOWN GAP: a repo running agents in subdirectories still ships
# `services/api/.claude/` and must add its own anchored entry.
.claude
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Re-include a committed template with a negation if the
# build needs one: `!docs/example.env`.
**/*.[eE][nN][vV]
**/.[eE][nN][vV].*
**/.[eE][nN][vV][rR][cC]
# Private keys and the bundles carrying them. Public certificates
# (*.crt, *.cer) are deliberately absent: they are legitimate inputs.
**/*.[pP][eE][mM]
**/*.[kK][eE][yY]
**/*.[pP]12
**/*.[pP][fF][xX]
**/[iI][dD]_[rR][sS][aA]
**/[iI][dD]_[dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]
**/[iI][dD]_[eE][dD]25519
# Dependencies: restored inside the image, never copied in.
**/node_modules
# OS metadata.
**/.DS_Store
**/Thumbs.db
# Editor state: never a build input, and it churns COPY.
**/*.swp
**/*.swo
**/*~
**/*.bak
**/.idea
**/.vscode
**/*.sublime-*
# This repo's own host-built binary (make build).
/bin/mfer
# The protoc script/bootstrap unpacks for script/generate.
/bin/protoc
+12
View File
@@ -0,0 +1,12 @@
root = true
[*]
indent_style = space
indent_size = 4
end_of_line = lf
charset = utf-8
trim_trailing_whitespace = true
insert_final_newline = true
[Makefile]
indent_style = tab
+7 -22
View File
@@ -1,24 +1,9 @@
name: check
on:
push:
branches: [main, next]
pull_request:
branches: [main, next]
on: [push]
jobs:
check:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 # v4
- uses: actions/setup-go@40f1582b2485089dde7abd97c1529aa768e1baff # v5
with:
go-version-file: go.mod
- name: Install golangci-lint
run: go install github.com/golangci/golangci-lint/v2/cmd/golangci-lint@5d1e709b7be35cb2025444e19de266b056b7b7ee # v2.10.1
- name: Install protoc-gen-go
run: go install google.golang.org/protobuf/cmd/protoc-gen-go@4dfe9d308b477d29e8c2b6d75abf3e78aafe3cb8 # v1.28.1
- name: Generate protobuf
run: |
sudo apt-get update && sudo apt-get install -y protobuf-compiler
cd mfer && go generate .
- name: Run checks
run: make check
check:
runs-on: ubuntu-latest
steps:
# actions/checkout v4.2.2, 2026-03-16
- uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
- run: script/cibuild
+27 -8
View File
@@ -1,10 +1,29 @@
/bin/
/bin/mfer
/bin/protoc/
/tmp
*.tmp
*.dockerimage
/vendor
vendor.tzst
modcache.tzst
/node_modules/
# Stale files
.drone.yml
# Generated manifest files
/index.mf
# Secrets
.env
.env.*
*.key
*.pem
# OS files
.DS_Store
Thumbs.db
# Editor files
*.swp
*.swo
*~
.idea/
.vscode/
# Go build artifacts
*.log
*.out
*.test
+34
View File
@@ -0,0 +1,34 @@
version: "2"
# Config schema uses the golangci-lint v2 layout (settings live under
# linters.settings, not top-level linters-settings) so that the
# thresholds below are actually applied by golangci-lint >= v2.
run:
timeout: 5m
modules-download-mode: readonly
linters:
default: all
disable:
# Genuinely incompatible with project patterns
- exhaustruct # Requires all struct fields
- depguard # Dependency allow/block lists
- godot # Requires comments to end with periods
- wsl # Deprecated, replaced by wsl_v5
- wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go
settings:
lll:
line-length: 88
funlen:
lines: 80
statements: 50
cyclop:
max-complexity: 15
dupl:
threshold: 100
issues:
max-issues-per-linter: 0
max-same-issues: 0
BIN
View File
Binary file not shown.
+1
View File
@@ -0,0 +1 @@
22.17.0
+14
View File
@@ -0,0 +1,14 @@
# REPO_POLICIES.md is a verbatim copy of an authoritative upstream
# document (sneak/prompts). Local tooling must never rewrite it: any
# reformatting is silent drift from the source of truth.
REPO_POLICIES.md
# User-owned configuration, copied verbatim from upstream. Not matched by
# the current prettier file set, but listed so widening that set can never
# start rewriting it.
.golangci.yml
# Dependencies and build output
node_modules/
vendor/
bin/
+4
View File
@@ -0,0 +1,4 @@
{
"tabWidth": 4,
"proseWrap": "always"
}
+34
View File
@@ -0,0 +1,34 @@
# Agent Instructions
Read `REPO_POLICIES.md` before making any changes. It is the authoritative
source for coding standards, formatting, linting, and workflow rules.
## Workflow
- When fixing a bug, write a failing test FIRST. Only after the test fails,
write the code to fix the bug. Then ensure the test passes. Leave the test in
place and commit it with the bugfix. Don't run shell commands to test bugfixes
or reproduce bugs. Write tests!
- After each change, run `make fmt`, then `make test`, then `make lint`. Fix any
failures before committing.
- After each change, commit only the files you've changed. Push after
committing.
## Attribution
- Never mention Claude, Anthropic, or any AI/LLM tooling in commit messages. Do
not use attribution.
## Repository-Specific Notes
- This is a Go library + CLI tool for generating `.mf` manifest files.
- The proto definition is in `mfer/mf.proto`; generated `.pb.go` files are
committed (required for `go get` compatibility).
- The format specification is in `docs/FORMAT.md`.
- Open work, open design questions included, is tracked only in the repo's
issues: https://git.eeqj.de/sneak/mfer/issues. There is no `TODO.md` and no
TODO list in `README.md`; do not add either. For this repo this overrides the
`REPO_POLICIES.md` rule to put the todo list in the README, per sneak's
ruling: https://git.eeqj.de/sneak/mfer/issues/76#issuecomment-118130.
-20
View File
@@ -1,20 +0,0 @@
# Important Rules
- when fixing a bug, write a failing test FIRST. only after the test fails, write
the code to fix the bug. then ensure the test passes. leave the test in
place and commit it with the bugfix. don't run shell commands to test
bugfixes or reproduce bugs. write tests!
- never, ever mention claude or anthropic in commit messages. do not use attribution
- after each change, run "make fmt".
- after each change, run "make test" and ensure all tests pass.
- after each change, run "make lint" and ensure no linting errors. fix any
you find, one by one.
- after each change, commit the files you've changed. push after
committing.
- NEVER use `git add -A`. always add only individual files that you've changed.
+73 -34
View File
@@ -1,37 +1,76 @@
################################################################################
#2345678911234567892123456789312345678941234567895123456789612345678971234567898
################################################################################
FROM sneak/builder:2022-12-08 AS builder
ENV DEBIAN_FRONTEND noninteractive
WORKDIR /build
COPY ./Makefile ./.golangci.yml ./go.mod ./go.sum /build/
COPY ./vendor.tzst /build/vendor.tzst
COPY ./modcache.tzst /build/modcache.tzst
COPY ./internal ./internal
COPY ./bin/gitrev.sh ./bin/gitrev.sh
COPY ./mfer ./mfer
COPY ./cmd ./cmd
ARG GITREV unknown
ARG DRONE_COMMIT_SHA unknown
# Lint stage — fast feedback on formatting and lint issues
# golangci/golangci-lint:v2.12.2 (Debian-based), 2026-08-07
FROM golangci/golangci-lint:v2.12.2@sha256:5cceeef04e53efe1470638d4b4b4f5ceefd574955ab3941b2d9a68a8c9ad5240 AS lint
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Go half of fmt-check only: this image has no node, so no prettier. The
# markdown half runs in the mdfmt stage below. The image has no gofumpt
# either; script/gofumpt builds the version bin/tools/go.mod pins.
RUN script/gofumpt --check
# The linter directly, not `make lint`: script/lint builds this stage, and
# there is no docker inside this build.
RUN golangci-lint run --config .golangci.yml ./...
# Markdown/JSON format stage — prettier needs node, which the Go images
# do not have. node:22.17.0-bookworm-slim (2026-08-09); ships node
# 22.17.0 and yarn 1.22.22, the versions .nvmrc and script/bootstrap pin.
FROM node@sha256:b04ce4ae4e95b522112c2e5c52f781471a5cbc3b594527bcddedee9bc48c03a0 AS mdfmt
WORKDIR /src
COPY package.json yarn.lock ./
RUN yarn install --frozen-lockfile
COPY . .
# No make in this image; call the script entrypoint directly.
RUN script/prettier --check
# Build stage — tests and compilation
# golang:1.23.12, 2026-03-14
FROM golang@sha256:60deed95d3888cc5e4d9ff8a10c54e5edc008c6ae3fba6187be6fb592e19e8c0 AS builder
# Force BuildKit to run the lint and mdfmt stages by creating stage dependencies
COPY --from=lint /src/go.sum /dev/null
COPY --from=mdfmt /src/go.sum /dev/null
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make test
# A build context sent as a tar archive, as upaas sends it, keeps its files'
# owners, and git refuses to read a checkout owned by another user.
RUN git config --system --add safe.directory /src
# The revision `mfer version` prints, stamped into main.Gitrev: the VERSION
# build argument when one is given (script/docker passes one), otherwise
# `git describe --tags --always` of the .git the build context carries: the
# tag on a tagged commit, tag-N-gHASH on a commit after one, the short commit
# when no tag is reachable. git ships in this base image. A context that
# carries .git and still yields no version fails the build.
ARG VERSION
RUN version="${VERSION:-$(git describe --tags --always)}"; \
if [ -e .git ] && { [ -z "$version" ] || [ "$version" = dev ] || \
[ "$version" = unknown ]; }; then \
echo "no version could be derived although the build context carries .git" >&2; \
exit 1; \
fi; \
cd cmd/mfer && \
CGO_ENABLED=0 go build -ldflags "-X main.Gitrev=$version" -o /mfer .
# Fail unless /mfer is statically linked: scratch has no C library to run it.
RUN ldd /mfer 2>&1 | grep -q 'not a dynamic executable'
RUN mkdir -p "$(go env GOMODCACHE)" && cd "$(go env GOMODCACHE)" && \
zstdmt -d --stdout /build/modcache.tzst | tar xf - && \
rm /build/modcache.tzst && cd /build
RUN \
cd mfer && go generate . && cd .. && \
GOPACKAGESDEBUG=true golangci-lint run ./... && \
mkdir vendor && cd vendor && \
zstdmt -d --stdout /build/vendor.tzst | tar xf - && rm /build/vendor.tzst && \
cd .. && \
make mfer.cmd
RUN rm -rf /build/vendor && go mod vendor && tar -c . | zstdmt -19 > /src.tzst
################################################################################
#2345678911234567892123456789312345678941234567895123456789612345678971234567898
################################################################################
## final image
################################################################################
FROM scratch
# we put all the source into the final image for posterity, it's small
COPY --from=builder /src.tzst /src.tzst
COPY --from=builder /build/mfer.cmd /mfer
# scratch has no CA certificates; fetch needs them to verify HTTPS servers.
COPY --from=builder /etc/ssl/certs/ca-certificates.crt /etc/ssl/certs/
COPY --from=builder /mfer /mfer
ENTRYPOINT ["/mfer"]
+28 -75
View File
@@ -1,88 +1,41 @@
export DOCKER_BUILDKIT := 1
export PROGRESS_NO_TRUNC := 1
GOPATH := $(shell go env GOPATH)
export PATH := $(PATH):$(GOPATH)/bin
PROTOC_GEN_GO := $(GOPATH)/bin/protoc-gen-go
SOURCEFILES := mfer/*.go mfer/*.proto internal/*/*.go cmd/*/*.go go.mod go.sum
ARCH := $(shell uname -m)
GITREV_BUILD := $(shell bash $(PWD)/bin/gitrev.sh)
APPNAME := mfer
VERSION := 0.1.0
export DOCKER_IMAGE_CACHE_DIR := $(HOME)/Library/Caches/Docker/$(APPNAME)-$(ARCH)
GOLDFLAGS += -X main.Version=$(VERSION)
GOLDFLAGS += -X main.Gitrev=$(GITREV_BUILD)
GOFLAGS := -ldflags "$(GOLDFLAGS)"
.PHONY: bootstrap setup test lint fmt fmt-check check docker hooks build generate fuzz
.PHONY: docker default run ci test fixme check check-fmt lint
# Makefile targets are thin shims; the implementations live in script/
# per the scripts-to-rule-them-all pattern (see the Entrypoints section
# of README.md).
default: fmt test
bootstrap:
@script/bootstrap
run: ./bin/mfer
./$<
./$< gen
setup:
@script/setup
ci: test
test:
@script/test
test: $(SOURCEFILES) mfer/mf.pb.go
go test -v --timeout 10s ./...
lint:
@script/lint
$(PROTOC_GEN_GO):
test -e $(PROTOC_GEN_GO) || go install -v google.golang.org/protobuf/cmd/protoc-gen-go@v1.28.1
fmt:
@script/fmt
fixme:
@grep -nir fixme . | grep -v Makefile
fmt-check:
@script/fmt-check
devprereqs:
which golangci-lint || go install -v github.com/golangci/golangci-lint/cmd/golangci-lint@latest
check:
@script/check
mfer/mf.pb.go: mfer/mf.proto
cd mfer && go generate .
docker:
@script/docker
bin/mfer: $(SOURCEFILES) mfer/mf.pb.go
protoc --version
cd cmd/mfer && go build -tags urfave_cli_no_docs -o ../../bin/mfer $(GOFLAGS) .
hooks:
@script/install-precommit
clean:
rm -rfv mfer/*.pb.go bin/mfer cmd/mfer/mfer *.dockerimage
build:
@script/build
fmt: mfer/mf.pb.go
gofumpt -l -w mfer internal cmd
golangci-lint run --fix
-prettier -w *.json
-prettier -w *.md
generate:
@script/generate
lint: mfer/mf.pb.go
golangci-lint run ./...
docker: sneak-mfer.$(ARCH).tzst.dockerimage
sneak-mfer.$(ARCH).tzst.dockerimage: $(SOURCEFILES) vendor.tzst modcache.tzst
docker build --progress plain --build-arg GITREV=$(GITREV_BUILD) -t sneak/mfer .
docker save sneak/mfer | pv | zstdmt -19 > $@
du -sh $@
godoc:
open http://127.0.0.1:6060
godoc -http=:6060
vendor.tzst: go.mod go.sum
go mod tidy
go mod vendor
cd vendor && tar -c . | pv | zstdmt -19 > $(PWD)/$@.tmp
rm -rf vendor
mv $@.tmp $@
modcache.tzst: go.mod go.sum
go mod tidy
cd $(HOME)/go/pkg && chmod -R u+rw . && rm -rf mod sumdb
go mod download -x
cd $(shell go env GOMODCACHE) && tar -c . | pv | zstdmt -19 > $(PWD)/$@.tmp
mv $@.tmp $@
# Individual check targets
check-fmt:
@echo "==> Checking formatting..."
@test -z "$$(gofmt -l .)" || (echo "Files not formatted:" && gofmt -l . && exit 1)
# Run all checks (formatting, linting, tests) without modifying files
check: check-fmt lint test
fuzz:
@script/fuzz
+255 -133
View File
@@ -1,132 +1,222 @@
# mfer
[mfer](https://git.eeqj.de/sneak/mfer) is a reference implementation library
and thin wrapper command-line utility written in [Go](https://golang.org)
and first published in 2022 under the [WTFPL](https://wtfpl.net) (public
domain) license. It specifies and generates `.mf` manifest files over a
directory tree of files to encapsulate metadata about them (such as
cryptographic checksums or signatures over same) to aid in archiving,
downloading, and streaming, or mirroring. The manifest files' data is
serialized with Google's [protobuf serialization
format](https://developers.google.com/protocol-buffers). The structure of
these files can be found [in the format
specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto)
which is included in the [project
repository](https://git.eeqj.de/sneak/mfer).
[mfer](https://git.eeqj.de/sneak/mfer) is a [WTFPL](https://wtfpl.net)-licensed
(public domain) [Go](https://golang.org) library and command-line tool by
[@sneak](https://sneak.berlin) that specifies and generates `.mf` manifest files
over a directory tree to encapsulate metadata about the files — such as
cryptographic checksums and signatures over same — to aid in archiving,
downloading, streaming, and mirroring. It was first published in 2022. The
manifest files' data is serialized with Google's
[protobuf serialization format](https://developers.google.com/protocol-buffers).
The structure of these files can be found
[in the format specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto)
which is included in the [project repository](https://git.eeqj.de/sneak/mfer).
The current version is pre-1.0 and while the repo was published in 2022,
there has not yet been any versioned release. [SemVer](https://semver.org)
will be used for releases.
The current version is pre-1.0 and while the repo was published in 2022, there
has not yet been any versioned release. [SemVer](https://semver.org) will be
used for releases.
This project was started by [@sneak](https://sneak.berlin) to scratch an
itch in 2022 and is currently a one-person effort, though the goal is for
this to emerge as a de-facto standard and be incorporated into other
software. A compatible javascript library is planned.
This project was started by [@sneak](https://sneak.berlin) to scratch an itch in
2022 and is currently a one-person effort, though the goal is for this to emerge
as a de-facto standard and be incorporated into other software. A compatible
javascript library is planned.
# Phases
# Getting Started
Manifest generation happens in two distinct phases:
`mfer` builds from source with a Go 1.23+ toolchain. The generated protobuf code
is committed, so no `protoc` toolchain is required:
## Phase 1: Enumeration
```sh
git clone https://git.eeqj.de/sneak/mfer.git
cd mfer
make build
```
Walking directories and calling `stat()` on files to collect metadata (path, size, mtime, ctime). This builds the list of files to be scanned. Relatively fast as it only reads filesystem metadata, not file contents.
Generate a manifest for a directory tree, verify it later, and fetch a published
tree by URL:
**Progress:** `EnumerateStatus` with `FilesFound` and `BytesFound`
```sh
# Write index.mf, a manifest of the files under the current directory.
bin/mfer gen .
## Phase 2: Scan (ToManifest)
# Verify the files on disk against the manifest. Exits nonzero if any file
# it lists is missing or corrupted; warns about files it does not list.
bin/mfer check index.mf
Reading file contents and computing cryptographic hashes for manifest generation. This is the expensive phase that reads all file data from disk.
# Download and cryptographically verify a tree published over HTTP into
# ./mirror: mfer fetches <url>/index.mf, downloads every file it lists,
# skipping any already there with the right hash, then saves the manifest as
# mirror/index.mf.
bin/mfer fetch --dest mirror https://example.com/tree/
```
**Progress:** `ScanStatus` with `TotalFiles`, `ScannedFiles`, `TotalBytes`, `ScannedBytes`, `BytesPerSec`
# Code Conventions
- **Logging:** Never use `fmt.Printf` or write to stdout/stderr directly in normal code. Use the `internal/log` package for all output (`log.Info`, `log.Infof`, `log.Debug`, `log.Debugf`, `log.Progressf`, `log.ProgressDone`).
- **Filesystem abstraction:** Use `github.com/spf13/afero` for filesystem operations to enable testing and flexibility.
- **CLI framework:** Use `github.com/urfave/cli/v2` for command-line interface.
- **Serialization:** Use Protocol Buffers for manifest file format.
- **Internal packages:** Non-exported implementation details go in `internal/` subdirectories.
- **Concurrency:** Use `sync.RWMutex` for protecting shared state; prefer channels for progress reporting.
- **Progress channels:** Use buffered channels (size 1) with non-blocking sends to avoid blocking the main operation if the consumer is slow.
- **Context support:** Long-running operations should accept `context.Context` for cancellation.
- **NO_COLOR:** Respect the `NO_COLOR` environment variable for disabling colored output.
- **Options pattern:** Use `NewWithOptions(opts *Options)` constructor pattern for configurable types.
Run `bin/mfer help` for the full command list, or `bin/mfer <command> --help`
for a single command's options.
# Build Status
[![Build Status](https://drone.datavi.be/api/badges/sneak/mfer/status.svg)](https://drone.datavi.be/sneak/mfer)
CI runs `script/cibuild`, which builds the Docker image with `--no-cache`, so
the formatting, lint and test steps in the `Dockerfile` run on every build. The
`main` branch must always be green.
# Entrypoints
This repository adheres to the
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
standard: normalized scripts in `script/` are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call them. We
provide:
- `script/bootstrap` — install all dependencies, idempotently: Go and the
modules of both `go.mod` and `bin/tools/go.mod`; node (the version `.nvmrc`
names, through nvm when there is no node on `PATH`) and yarn, plus the
prettier version pinned in `package.json`/`yarn.lock`; and `protoc` 33.4,
unpacked into `bin/protoc` from its release archive once the archive matches
the sha256 the script holds for this platform. golangci-lint is not installed,
it runs only in Docker
- `script/setup` — make a fresh clone ready for development: runs
`script/bootstrap`, then `script/install-precommit`
- `script/projectname` — output the project name (`mfer`); used by other scripts
such as `script/docker`
- `script/test` — run the test suite (`go test`); one test fails when
`mfer/mf.proto` no longer matches the hash `script/generate` recorded
- `script/build` (`make build`) — build the `mfer` binary into `bin/mfer`,
stamped with the revision `mfer version` prints: the output of
`git describe --tags --always --dirty`, as `script/docker` passes it
- `script/generate` (`make generate`) — regenerate `mfer/mf.pb.go` from
`mfer/mf.proto` and record the hash of that `mfer/mf.proto` in
`mfer/mf.proto.sha256`; the only thing that regenerates the committed
`mfer/mf.pb.go`. It runs the `protoc` that `script/bootstrap` unpacks into
`bin/protoc`, refusing any version but 33.4, and the `protoc-gen-go` that
`bin/tools/go.mod` pins, which `go tool` builds from source checked against
the hashes in `bin/tools/go.sum`
- `script/fuzz` — fuzz the manifest parser for one minute; run by hand
(`make fuzz`), never by CI, while `script/test` runs its committed seed corpus
as ordinary tests
- `script/lint` — run `golangci-lint` in Docker: builds only the `lint` stage of
the `Dockerfile` (the Go format check, then the linter), uncached so it runs
every time, then removes the image
- `script/fmt` — format all code and docs (writes): `script/gofumpt --write` and
`script/prettier --write`
- `script/gofumpt` — run `gofumpt` over every Go file in the repository in the
given mode, `--write` or `--check`, at the version `bin/tools/go.mod` pins
(built on demand by `go tool` from source checked against the hashes in
`bin/tools/go.sum`, so nothing installs it); `script/fmt`, `script/fmt-check`
and the Docker lint stage all go through it, so they cannot disagree about Go
formatting
- `script/prettier` — run prettier over the repository's canonical file set
(Markdown and JSON, minus `.prettierignore`) in the given mode, `--write` or
`--check`, with the prettier version `yarn.lock` pins, run by the node on
`PATH` or else the one `script/bootstrap` installed through nvm; the single
definition of that file set, so `script/fmt` and `script/fmt-check` cannot
disagree about it
- `script/fmt-check` — check formatting without writing:
`script/gofumpt --check` plus `script/prettier --check`
- `script/check` — run `script/test`, `script/lint`, and `script/fmt-check`
- `script/docker` — build the Docker image tagged with the project name
- `script/cibuild` — CI entrypoint: builds the image with the same command as
`script/docker`, uncached, so the checks in the Dockerfile run every time
- `script/precommit` — pre-commit checks: `go mod tidy` verification, then
`script/check`
- `script/install-precommit` — install the git pre-commit hook that runs
`script/precommit`
# Participation
The community is as yet nonexistent so there are no defined policies or
norms yet. Primary development happens on a privately-run Gitea instance at
[https://git.eeqj.de/sneak/mfer](https://git.eeqj.de/sneak/mfer) and issues
are [tracked there](https://git.eeqj.de/sneak/mfer/issues).
The community is as yet nonexistent so there are no defined policies or norms
yet. Primary development happens on a privately-run Gitea instance at
[https://git.eeqj.de/sneak/mfer](https://git.eeqj.de/sneak/mfer) and issues are
[tracked there](https://git.eeqj.de/sneak/mfer/issues).
Changes must always be formatted with a standard `go fmt`, syntactically
valid, and must pass the linting defined in the repository (presently only
the `golangci-lint` defaults), which can be run with a `make lint`. The
`main` branch is protected and all changes must be made via [pull
requests](https://git.eeqj.de/sneak/mfer/pulls) and pass CI to be merged.
Any changes submitted to this project must also be
Changes must always be formatted with a standard `go fmt`, syntactically valid,
and must pass the linting defined in the repository's `.golangci.yml`, which
`make lint` runs in Docker. The `main` branch is protected and all changes must
be made via [pull requests](https://git.eeqj.de/sneak/mfer/pulls) and pass CI to
be merged. Any changes submitted to this project must also be
[WTFPL-licensed](https://wtfpl.net) to be considered.
# Problem Statement
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
tooling requirements, and workflow conventions.
# Rationale
## The problem
Given a plain URL, there is no standard way to safely and programmatically
download everything "under" that URL path. `wget -r` can traverse directory
listings if they're enabled, but every server has a different format, and
this does not verify cryptographic integrity of the files, or enable them to
be fetched using a different protocol other than HTTP/s.
listings if they're enabled, but every server has a different format, and this
does not verify cryptographic integrity of the files, or enable them to be
fetched using a different protocol other than HTTP/s.
Currently, the solution that people are using are sidecar files in the
format of `SHASUMS` checksum files, as well as a `SHASUMS.asc` PGP detached
signature. This is not checksum-algorithm-agnostic and the sidecar file is
not always consistently named.
Currently, the solution that people are using are sidecar files in the format of
`SHASUMS` checksum files, as well as a `SHASUMS.asc` PGP detached signature.
This is not checksum-algorithm-agnostic and the sidecar file is not always
consistently named.
Real issues I face:
- when I plug in an ExFAT hard drive, I don't know if any files on the
filesystem are corrupted or missing
- current ad-hoc solution are `SHASUMS`/`SHASUMS.asc` files
- current ad-hoc solution are `SHASUMS`/`SHASUMS.asc` files
- when I want to mirror an HTTP archive, I have to use special tools like
debmirror that understand the archive format
- the debian repository metadata structure is hot garbage
- when I download a large file via HTTP, I have no way of knowing if the
file content is what it's supposed to be
- the debian repository metadata structure is hot garbage
- when I download a large file via HTTP, I have no way of knowing if the file
content is what it's supposed to be
# Proposed Solution
## The solution
A standard, a manifest file format, and a tool for generating same.
The manifest file would be called `index.mf`, and the tool for generating such would be called `mfer`.
The manifest file would be called `index.mf`, and the tool for generating such
would be called `mfer`.
The manifest file would do several important things:
- have a standard filename, so if given
`https://example.com/downloadpackage/` one could fetch
`https://example.com/downloadpackage/index.mf` to enumerate the full
directory listing.
- have a standard filename, so if given `https://example.com/downloadpackage/`
one could fetch `https://example.com/downloadpackage/index.mf` to enumerate
the full directory listing.
- contain a version field for extensibility
- contain structured data (protobuf, json, or cbor)
- provide an inner signed container, so that the manifest file itself can
embed a signature and a public key alongside in a single file
- provide an inner signed container, so that the manifest file itself can embed
a signature and a public key alongside in a single file
- contain a list of files, each with a relative path to the manifest
- contain manifest timestamp
- contain ctime/mtime information for files so that file metadata can be
preserved
- contain cryptographic checksums in several different algorithms for each
file
- probably encoded with multihash to indicate algo + hash
- sha256 at the minimum
- would be nice to include an IPFS/IPLD CIDv1 root hash for each file,
which likely involves doing an ipfs file object chunking
- maybe even including the complete IPFS/IPLD directory tree objects and
chunklists?
- this is because generating an `index.mf` does not imply publishing on
ipfs at that time
- maybe a bittorrent chunklist for torrent client compatibility? perhaps a
top-level infohash for the whole manifest?
- contain cryptographic checksums in several different algorithms for each file
- probably encoded with multihash to indicate algo + hash
- sha256 at the minimum
- would be nice to include an IPFS/IPLD CIDv1 root hash for each file, which
likely involves doing an ipfs file object chunking
- maybe even including the complete IPFS/IPLD directory tree objects and
chunklists?
- this is because generating an `index.mf` does not imply publishing on
ipfs at that time
- maybe a bittorrent chunklist for torrent client compatibility? perhaps a
top-level infohash for the whole manifest?
# Design
The repository is split into a reusable library and a thin command-line wrapper
around it.
- `mfer/` is the reusable library and the heart of the project: it defines the
manifest format and implements building, scanning, checking, serialization,
and signing. The protobuf schema is `mfer/mf.proto`, and the generated code it
produces (`mfer/mf.pb.go`) is committed alongside it so the library builds
with `go get` and needs no `protoc` toolchain.
- `internal/cli/` holds the command implementations — `generate` (alias `gen`),
`check`, `freshen`, `export`, `list` (alias `ls`), `fetch`, and `version` —
that wire the library to the command-line interface.
- `internal/log/` provides the logging used by the commands and the library.
- `internal/bork/` provides the error the library returns when a manifest's
decompressed contents are not the size the manifest records.
- `cmd/mfer/` is the entrypoint: its `main` package runs `internal/cli` and
exits with the status it returns.
Everything under `internal/` is private to this repository; only the `mfer/`
package is intended for import by other software.
# Design Goals
@@ -140,41 +230,57 @@ The manifest file would do several important things:
# Non-Goals
- Manifest generation speed
- likely involves IPFS chunking, bittorrent chunking, and several
different cryptographic hash functions over the entirety of each and
every file
- likely involves IPFS chunking, bittorrent chunking, and several different
cryptographic hash functions over the entirety of each and every file
- Small manifest file size (within reason)
- 30MiB files are "small" these days, given modern storage/bandwidth
- metadata size should not be used as an excuse to sacrifice utility (such
as providing checksums over each chunk of a large file)
- 30MiB files are "small" these days, given modern storage/bandwidth
- metadata size should not be used as an excuse to sacrifice utility (such
as providing checksums over each chunk of a large file)
# Limitations
# Original Design Questions
- **Manifest size:** Manifests must fit entirely in system memory during reading and writing.
These were the open questions when the project started; open design questions
are now tracked only in the [issues](https://git.eeqj.de/sneak/mfer/issues).
# Open Questions
- Should the manifest file include checksums of individual file chunks, or just for the whole assembled file?
- If so, should the chunksize be fixed or dynamic?
- Should the manifest file include checksums of individual file chunks, or just
for the whole assembled file? If so, should the chunk size be fixed or
dynamic? Still open, on
[issue 81](https://git.eeqj.de/sneak/mfer/issues/81#issuecomment-118698).
- Should the manifest signature format be GnuPG signatures, or those from
OpenBSD's signify (of which there is a good [golang
implementation](https://github.com/frankbraun/gosignify)?
OpenBSD's signify (of which there is a good
[golang implementation](https://github.com/frankbraun/gosignify))? Still open,
as question 10 on [issue 82](https://git.eeqj.de/sneak/mfer/issues/82).
- Should the on-disk serialization format be proto3 or json?
- Should the on-disk serialization format be proto3 or json? Settled: it is
proto3, see `docs/FORMAT.md` and `mfer/mf.proto`.
# Tool Examples
- `mfer gen` / `mfer gen .`
- recurses under current directory and writes out an `index.mf`
- recurses under current directory and writes out an `index.mf`
- `mfer check` / `mfer check .`
- verifies checksums of all files in manifest, displaying error and
exiting nonzero if any files are missing or corrupted
- verifies checksums of all files in manifest, displaying error and exiting
nonzero if any files are missing or corrupted
- warns about each file under the base directory that the manifest does not
list, hidden files included; with `--no-extra-files` each one is a failure
instead
- `mfer fetch https://example.com/stuff/`
- fetches `/stuff/index.mf` and downloads all files listed in manifest,
optionally resuming any that already exist locally, and assures
cryptographic integrity of downloaded files.
- fetches `/stuff/index.mf` and downloads all files listed in manifest into
the current directory, or the one given with `--dest`, and assures
cryptographic integrity of downloaded files. A file already there with the
size and hash the manifest lists is skipped. Once every file is in place,
the manifest is saved there as `index.mf`, so `mfer check` can verify the
tree later. Each file is downloaded to a temp file beside it, such as
`.a.txt.tmp` for `a.txt`, then moved into place. A manifest is refused
before any file is downloaded if it lists a file where fetch writes
another: at the temp file of a listed file, or at `index.mf` or
`.index.mf.tmp` at the top of the tree. Names are compared in any letter
case.
- `mfer fetch --require-signature <fingerprint> https://example.com/stuff/`
- as above, but first refuses a manifest not signed by the key with that
fingerprint, as `mfer check --require-signature` does, before downloading
any file.
# Implementation Plan
@@ -191,30 +297,31 @@ The manifest file would do several important things:
# Hopes And Dreams
- `aria2c https://example.com/manifestdirectory/`
- (fetches `https://example.com/manifestdirectory/index.mf`, downloads and
checksums all files, resumes any that exist locally already)
- (fetches `https://example.com/manifestdirectory/index.mf`, downloads and
checksums all files, resumes any that exist locally already)
- `mfer fetch https://example.com/manifestdirectory/`
- a command line option to zero/omit mtime/ctime, as well as manifest
timestamp, and sort all directory listings so that manifest file
generation is deterministic/reproducible
- URL format `mfer fetch https://exmaple.com/manifestdirectory/?key=5539AD00DE4C42F3AFE11575052443F4DF2A55C2`
to assert in the URL which PGP signing key should be used in the manifest,
so that shared URLs have a cryptographic trust root
- a "well-known" key in the manifest that maps well known keys (could reuse
the http spec) to specific file paths in the manifest.
- example: a `berlin.sneak.app.slideshow` key that maps to a json
slideshow config listing what image paths to show, and for how long, and
in what order
- a command line option to zero/omit mtime/ctime, as well as manifest timestamp,
and sort all directory listings so that manifest file generation is
deterministic/reproducible
- URL format
`mfer fetch https://exmaple.com/manifestdirectory/?key=5539AD00DE4C42F3AFE11575052443F4DF2A55C2`
to assert in the URL which PGP signing key should be used in the manifest, so
that shared URLs have a cryptographic trust root
- a "well-known" key in the manifest that maps well known keys (could reuse the
http spec) to specific file paths in the manifest.
- example: a `berlin.sneak.app.slideshow` key that maps to a json slideshow
config listing what image paths to show, and for how long, and in what
order
# Use Cases
## Web Images
I'd like to be able to put a bunch of images into a directory, generate a
manifest, and then point a slideshow client (such as an ambient display, or
a react app with the target directory in a query string arg) at that
statically hosted directory, and have it discover the full list of images
available at that URL.
manifest, and then point a slideshow client (such as an ambient display, or a
react app with the target directory in a query string arg) at that statically
hosted directory, and have it discover the full list of images available at that
URL.
## Software Distribution
@@ -224,29 +331,44 @@ resumably by either HTTP or IPFS/BitTorrent without a .torrent file.
## Filesystem Archive Integrity
I use filesystems that don't include data checksums, and I would like a
cryptographically signed checksum file so that I can later verify that a set
of archive files have not been modified, none are missing, and that the
checksums have not been altered in storage by a second party.
cryptographically signed checksum file so that I can later verify that a set of
archive files have not been modified, none are missing, and that the checksums
have not been altered in storage by a second party.
## Filesystem-Independent Checksums
I would like to be able to plug in a hard drive or flash drive and, if there
is an `index.mf` in the root, automatically detect missing/corrupted files,
I would like to be able to plug in a hard drive or flash drive and, if there is
an `index.mf` in the root, automatically detect missing/corrupted files,
regardless of filesystem format.
# Collaboration
Please email [`sneak@sneak.berlin`](mailto:sneak@sneak.berlin) with your
desired username for an account on this Gitea instance.
Please email [`sneak@sneak.berlin`](mailto:sneak@sneak.berlin) with your desired
username for an account on this Gitea instance.
# TODO
Open work, open design questions included, is tracked in this repo's issues:
[https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues).
# See Also
## Prior Art: Metalink
- [Metalink - Mozilla Wiki](https://wiki.mozilla.org/Metalink)
- [Metalink - Wikipedia](https://en.wikipedia.org/wiki/Metalink)
- [RFC 5854 - The Metalink Download Description Format](https://datatracker.ietf.org/doc/html/rfc5854)
- [RFC 6249 - Metalink/HTTP: Mirrors and Hashes](https://www.rfc-editor.org/rfc/rfc6249.html)
## Links
- Repo: [https://git.eeqj.de/sneak/mfer](https://git.eeqj.de/sneak/mfer)
- Issues: [https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues)
- Issues:
[https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues)
# Authors
# Author
- [@sneak &lt;sneak@sneak.berlin&gt;](mailto:sneak@sneak.berlin)
- [@sneak](https://sneak.berlin)
# License
+408
View File
@@ -0,0 +1,408 @@
---
title: Repository Policies
last_modified: 2026-07-06
---
This document covers repository structure, tooling, and workflow standards. Code
style conventions are in separate documents:
- [Code Styleguide](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE.md)
(general, bash, Docker)
- [Go](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE_GO.md)
- [JavaScript](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE_JS.md)
- [Python](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/CODE_STYLEGUIDE_PYTHON.md)
- [Go HTTP Server Conventions](https://git.eeqj.de/sneak/prompts/raw/branch/main/prompts/GO_HTTP_SERVER_CONVENTIONS.md)
---
- Cross-project documentation (such as this file) must include
`last_modified: YYYY-MM-DD` in the YAML front matter so it can be kept in sync
with the authoritative source as policies evolve.
- **ALL external references must be pinned by cryptographic hash.** This
includes Docker base images, Go modules, npm packages, GitHub Actions, and
anything else fetched from a remote source. Version tags (`@v4`, `@latest`,
`:3.21`, etc.) are server-mutable and therefore remote code execution
vulnerabilities. The ONLY acceptable way to reference an external dependency
is by its content hash (Docker `@sha256:...`, Go module hash in `go.sum`, npm
integrity hash in lockfile, GitHub Actions `@<commit-sha>`). No exceptions.
This also means never `curl | bash` to install tools like pyenv, nvm, rustup,
etc. Instead, download a specific release archive from GitHub, verify its hash
(hardcoded in the Dockerfile or script), and only then install. Unverified
install scripts are arbitrary remote code execution. This is the single most
important rule in this document. Double-check every external reference in
every file before committing. There are zero exceptions to this rule.
- Every repo with software must have a root `Makefile` with these targets:
`make bootstrap`, `make setup`, `make test`, `make lint`, `make fmt` (writes),
`make fmt-check` (read-only), `make check` (runs `test`, `lint`, `fmt-check`),
`make docker`, and `make hooks` (installs pre-commit hook). A model Makefile
is at `https://git.eeqj.de/sneak/prompts/raw/branch/main/Makefile`.
- Repos follow the
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
pattern: the implementation of each Makefile target lives in an executable
script in `script/` (`script/bootstrap`, `script/setup`, `script/test`,
`script/lint`, `script/fmt`, `script/fmt-check`, `script/check`,
`script/docker`), and the Makefile targets are thin shims that call them. The
scripts must be POSIX sh (`#!/bin/sh`, `set -eu`, no bashisms) so they run in
minimal containers (e.g. alpine images have no bash); locate the repo root
with `$(cd "$(dirname "$0")/.." && pwd -P)` and `cd` there before acting. From
the standard's canonical set we use `bootstrap`, `setup` (make the repo ready
for development after a fresh clone: runs `bootstrap`, then
`install-precommit`, plus any repo-specific initialization), `test`, and
`cibuild`. `script/bootstrap` installs all dependencies idempotently and
assumes nothing is present: base tools come from nix, apt, brew, or apk
(detected in that order; apt runs noninteractive). For node it uses the
installed node if present; otherwise it installs a PINNED node version via
nvm, first installing nvm itself if missing — from a hash-verified GitHub
release archive (never `curl | sh`), with bash installed as an explicit
prerequisite since nvm requires bash. yarn is then pinned via
`corepack prepare yarn@<version> --activate`. Never install "latest" or "lts";
always exact versions. `script/cibuild` runs the CI build: it changes to the
repo root and runs `docker build .`; the Gitea workflow calls it. Four further
scripts are our own extensions to the standard: `script/check` runs
`script/test`, `script/lint`, and `script/fmt-check`; `script/precommit` is
what the git pre-commit hook runs, and it calls `script/check`;
`script/install-precommit` installs the git pre-commit hook (the `make hooks`
target shims to it); and `script/projectname` (literally that filename) simply
outputs the project's name. Scripts that need the name call
`script/projectname` — e.g. `script/docker` assembles its image tag from it —
so those scripts stay byte-identical across all repos. Repo-type-specific
pre-commit extras (e.g. `go mod tidy` verification in Go repos) belong in
`script/precommit`, not in the hook itself. Model scripts are at
`https://git.eeqj.de/sneak/prompts/raw/branch/main/script/<name>`. The README
must document the provided scripts in an **Entrypoints** section (see the
README requirements below).
- Always use Makefile targets (`make fmt`, `make test`, `make lint`, etc.)
instead of invoking the underlying tools directly. The Makefile is the single
source of truth for how these operations are run.
- The Makefile is authoritative documentation for how the repo is used. Beyond
the required targets above, it should have targets for every common operation:
running a local development server (`make run`, `make dev`), re-initializing
or migrating the database (`make db-reset`, `make migrate`), building
artifacts (`make build`), generating code, seeding data, or anything else a
developer would do regularly. If someone checks out the repo and types
`make<tab>`, they should see every meaningful operation available. A new
contributor should be able to understand the entire development workflow by
reading the Makefile.
- Every repo should have a `Dockerfile`. All Dockerfiles must run `make check`
as a build step so the build fails if the branch is not green. For non-server
repos, the Dockerfile should bring up a development environment and run
`make check`. For server repos, `make check` should run as an early build
stage before the final image is assembled. Dockerfiles install development
prerequisites by running `script/bootstrap` rather than duplicating installs
inline; COPY `script/` and the dependency manifests (`package.json` +
`yarn.lock`, `go.mod` + `go.sum`, etc.) before running it so the bootstrap
layer stays cached until dependencies change.
- **Dockerfiles must use a separate lint stage for fail-fast feedback.** Go
repos use a multistage build where linting runs in an independent stage based
on the `golangci/golangci-lint` image (pinned by hash). This stage runs
`make fmt-check` and `make lint` before the full build begins. The build stage
then declares an explicit dependency on the lint stage via
`COPY --from=lint /src/go.sum /dev/null`, which forces BuildKit to complete
linting before proceeding to compilation and tests. This ensures lint failures
surface in seconds rather than minutes, without blocking on dependency
download or compilation in the build stage.
The standard pattern for a Go repo Dockerfile is:
```dockerfile
# Lint stage — fast feedback on formatting and lint issues
# golangci/golangci-lint:v2.x.x, YYYY-MM-DD
FROM golangci/golangci-lint@sha256:... AS lint
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make fmt-check
RUN make lint
# Build stage
# golang:1.x-alpine, YYYY-MM-DD
FROM golang@sha256:... AS builder
WORKDIR /src
# Force BuildKit to run the lint stage before proceeding
COPY --from=lint /src/go.sum /dev/null
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN make test
ARG VERSION=dev
RUN CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X main.Version=${VERSION}" \
-o /app ./cmd/app/
# Runtime stage
FROM alpine@sha256:...
COPY --from=builder /app /usr/local/bin/app
ENTRYPOINT ["app"]
```
Key points:
- The lint stage uses the `golangci/golangci-lint` image directly (it
includes both Go and the linter), so there is no need to install the
linter separately.
- `COPY --from=lint /src/go.sum /dev/null` is a no-op file copy that creates
a stage dependency. BuildKit runs stages in parallel by default; without
this line, the build stage would not wait for lint to finish and a lint
failure might not fail the overall build.
- If the project uses `//go:embed` directives that reference build artifacts
(e.g. a web frontend compiled in a separate stage), the lint stage must
create placeholder files so the embed directives resolve. Example:
`RUN mkdir -p web/dist && touch web/dist/index.html web/dist/style.css`.
The lint stage should not depend on the actual build output — it exists to
fail fast.
- If the project requires CGO or system libraries for linting (e.g.
`vips-dev`), install them in the lint stage with `apk add`.
- The build stage runs `make test` after compilation setup. Tests run in the
build stage, not the lint stage, because they may require compiled
artifacts or heavier dependencies.
- Every repo should have a Gitea Actions workflow (`.gitea/workflows/`) that
runs `script/cibuild` (which runs `docker build .`) on push. Since the
Dockerfile already runs `make check`, a successful build implies all checks
pass.
- Use platform-standard formatters: `black` for Python, `prettier` for
JS/CSS/Markdown/HTML, `go fmt` for Go. Always use default configuration with
two exceptions: four-space indents (except Go), and `proseWrap: always` for
Markdown (hard-wrap at 80 columns). Documentation and writing repos (Markdown,
HTML, CSS) should also have `.prettierrc` and `.prettierignore`.
- Pre-commit hook: runs `script/precommit`, which calls `script/check`. If local
testing is not possible in the repo, `script/precommit` may skip `script/test`
and run only `script/lint` and `script/fmt-check`. The hook is installed by
`script/install-precommit`; the Makefile must provide a `make hooks` target
that shims to it.
- All repos with software must have tests that run via the platform-standard
test framework (`go test`, `pytest`, `jest`/`vitest`, etc.). If no meaningful
tests exist yet, add the most minimal test possible — e.g. importing the
module under test to verify it compiles/parses. There is no excuse for
`make test` to be a no-op.
- `make test` must complete in under 20 seconds. Add a 30-second timeout in the
Makefile.
- **`make test` should use the conditional verbose rerun pattern.** Run tests
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to
show full output. This keeps CI logs and `docker build` output clean on
success (just package/suite summaries) while providing full diagnostic detail
on failure (every test case, every assertion). The general shell pattern:
```makefile
test:
@<test-command> || \
{ echo "--- Rerunning with -v for details ---"; \
<test-command-with-v>; exit 1; }
```
Go example:
```makefile
test:
@go test -timeout 30s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 30s -race -v ./...; exit 1; }
```
Python example:
```makefile
test:
@python -m pytest || \
{ echo "--- Rerunning with -v for details ---"; \
python -m pytest -v; exit 1; }
```
The `exit 1` ensures the target always fails after a rerun — the first run
already proved the tests are broken, so the build must not pass even if a
flaky test happens to succeed on the second attempt. The rerun exists solely
for diagnostic output.
- Docker builds must complete in under 5 minutes.
- `make check` must not modify any files in the repo. Tests may use temporary
directories.
- `main` must always pass `make check`, no exceptions.
- Never commit secrets. `.env` files, credentials, API keys, and private keys
must be in `.gitignore`. No exceptions.
- `.gitignore` should be comprehensive from the start: OS files (`.DS_Store`),
editor files (`.swp`, `*~`), language build artifacts, and `node_modules/`.
Fetch the standard `.gitignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.gitignore` when setting up
a new repo.
- **No build artifacts in version control.** Code-derived data (compiled
bundles, minified output, generated assets) must never be committed to the
repository if it can be avoided. The build process (e.g. Dockerfile, Makefile)
should generate these at build time. Notable exception: Go protobuf generated
files (`.pb.go`) ARE committed because repos need to work with `go get`, which
downloads code but does not execute code generation.
- Never use `git add -A` or `git add .`. Always stage files explicitly by name.
- Never force-push to `main`.
- Make all changes on a feature branch. You can do whatever you want on a
feature branch.
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only
manually by the user. Fetch from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`.
- When pinning images or packages by hash, add a comment above the reference
with the version and date (YYYY-MM-DD).
- Use `yarn`, not `npm`.
- Write all dates as YYYY-MM-DD (ISO 8601).
- Simple projects should be configured with environment variables.
- Dockerized web services listen on port 8080 by default, overridable with
`PORT`.
- **HTTP/web services must be hardened for production internet exposure before
tagging 1.0.** This means full compliance with security best practices
including, without limitation, all of the following:
- **Security headers** on every response:
- `Strict-Transport-Security` (HSTS) with `max-age` of at least one year
and `includeSubDomains`.
- `Content-Security-Policy` (CSP) with a restrictive default policy
(`default-src 'self'` as a baseline, tightened per-resource as
needed). Never use `unsafe-inline` or `unsafe-eval` unless
unavoidable, and document the reason.
- `X-Frame-Options: DENY` (or `SAMEORIGIN` if framing is required).
Prefer the `frame-ancestors` CSP directive as the primary control.
- `X-Content-Type-Options: nosniff`.
- `Referrer-Policy: strict-origin-when-cross-origin` (or stricter).
- `Permissions-Policy` restricting access to browser features the
application does not use (camera, microphone, geolocation, etc.).
- **Request and response limits:**
- Maximum request body size enforced on all endpoints (e.g. Go
`http.MaxBytesReader`). Choose a sane default per-route; never accept
unbounded input.
- Maximum response body size where applicable (e.g. paginated APIs).
- `ReadTimeout` and `ReadHeaderTimeout` on the `http.Server` to defend
against slowloris attacks.
- `WriteTimeout` on the `http.Server`.
- `IdleTimeout` on the `http.Server`.
- Per-handler execution time limits via `context.WithTimeout` or
chi/stdlib `middleware.Timeout`.
- **Authentication and session security:**
- Rate limiting on password-based authentication endpoints. API keys are
high-entropy and not susceptible to brute force, so they are exempt.
- CSRF tokens on all state-mutating HTML forms. API endpoints
authenticated via `Authorization` header (Bearer token, API key) are
exempt because the browser does not attach these automatically.
- Passwords stored using bcrypt, scrypt, or argon2 — never plain-text,
MD5, or SHA.
- Session cookies set with `HttpOnly`, `Secure`, and `SameSite=Lax` (or
`Strict`) attributes.
- **Reverse proxy awareness:**
- True client IP detection when behind a reverse proxy
(`X-Forwarded-For`, `X-Real-IP`). The application must accept
forwarded headers only from a configured set of trusted proxy
addresses — never trust `X-Forwarded-For` unconditionally.
- **CORS:**
- Authenticated endpoints must restrict `Access-Control-Allow-Origin` to
an explicit allowlist of known origins. Wildcard (`*`) is acceptable
only for public, unauthenticated read-only APIs.
- **Error handling:**
- Internal errors must never leak stack traces, SQL queries, file paths,
or other implementation details to the client. Return generic error
messages in production; detailed errors only when `DEBUG` is enabled.
- **TLS:**
- Services never terminate TLS directly. They are always deployed behind
a TLS-terminating reverse proxy. The service itself listens on plain
HTTP. However, HSTS headers and `Secure` cookie flags must still be
set by the application so that the browser enforces HTTPS end-to-end.
This list is non-exhaustive. Apply defense-in-depth: if a standard security
hardening measure exists for HTTP services and is not listed here, it is
still expected. When in doubt, harden.
- `README.md` is the primary documentation. Required sections:
- **Description**: First line must include the project name, purpose,
category (web server, SPA, CLI tool, etc.), license, and author. Example:
"µPaaS is an MIT-licensed Go web application by @sneak that receives
git-frontend webhooks and deploys applications via Docker in realtime."
- **Getting Started**: Copy-pasteable install/usage code block.
- **Entrypoints**: Opens by stating that the repo adheres to the
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
standard (with that link), then documents each provided `script/`
entrypoint and its purpose.
- **Rationale**: Why does this exist?
- **Design**: How is the program structured?
- **TODO**: Update meticulously, even between commits. When planning, put
the todo list in the README so a new agent can pick up where the last one
left off.
- **License**: MIT, GPL, or WTFPL. Ask the user for new projects. Include a
`LICENSE` file in the repo root and a License section in the README.
- **Author**: [@sneak](https://sneak.berlin).
- First commit of a new repo should contain only `README.md`.
- Go module root: `sneak.berlin/go/<name>`. Always run `go mod tidy` before
committing.
- Use SemVer.
- Database migrations live in `internal/db/migrations/` and must be embedded in
the binary.
- `000_migration.sql` — contains ONLY the creation of the migrations
tracking table itself. Nothing else.
- `001_schema.sql` — the full application schema.
- **Pre-1.0.0:** never add additional migration files (002, 003, etc.).
There is no installed base to migrate. Edit `001_schema.sql` directly.
- **Post-1.0.0:** add new numbered migration files for each schema change.
Never edit existing migrations after release.
- All repos should have an `.editorconfig` enforcing the project's indentation
settings.
- Avoid putting files in the repo root unless necessary. Root should contain
only project-level config files (`README.md`, `Makefile`, `Dockerfile`,
`LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`, and
language-specific config). Everything else goes in a subdirectory. Canonical
subdirectory names:
- `bin/` — executable scripts and tools
- `cmd/` — Go command entrypoints
- `configs/` — configuration templates and examples
- `deploy/` — deployment manifests (k8s, compose, terraform)
- `docs/` — documentation and markdown (README.md stays in root)
- `internal/` — Go internal packages
- `internal/db/migrations/` — database migrations
- `pkg/` — Go library packages
- `share/` — systemd units, data files
- `static/` — static assets (images, fonts, etc.)
- `web/` — web frontend source
- When setting up a new repo, files from the `prompts` repo may be used as
templates. Fetch them from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/<path>`.
- New repos must contain at minimum:
- `README.md`, `.git`, `.gitignore`, `.editorconfig`
- `LICENSE`, `REPO_POLICIES.md` (copy from the `prompts` repo)
- `Makefile`
- `script/` entrypoints (`bootstrap`, `setup`, `projectname`, `test`,
`lint`, `fmt`, `fmt-check`, `check`, `docker`, `cibuild`, `precommit`,
`install-precommit`)
- `Dockerfile`, `.dockerignore`
- `.gitea/workflows/check.yml`
- Go: `go.mod`, `go.sum`, `.golangci.yml`
- JS: `package.json`, `yarn.lock`, `.prettierrc`, `.prettierignore`
- Python: `pyproject.toml`
-122
View File
@@ -1,122 +0,0 @@
# TODO: mfer 1.0
## Design Questions
*sneak: please answer inline below each question. These are preserved for posterity.*
### Format Design
**1. Should `MFFileChecksum` be simplified?**
Currently it's a separate message wrapping a single `bytes multiHash` field. Since multihash already self-describes the algorithm, `repeated bytes hashes` directly on `MFFilePath` would be simpler and reduce per-file protobuf overhead. Is the extra message layer intentional (e.g. planning to add per-hash metadata like `verified_at`)?
> *answer:*
**2. Should file permissions/mode be stored?**
The format stores mtime/ctime but not Unix file permissions. For archival use (ExFAT, filesystem-independent checksums) this may not matter, but for software distribution or filesystem restoration it's a gap. Should we reserve a field now (e.g. `optional uint32 mode = 305`) even if we don't populate it yet?
> *answer:*
**3. Should `atime` be removed from the schema?**
Access time is volatile, non-deterministic, and often disabled (`noatime`). Including it means two manifests of the same directory at different times will differ, which conflicts with the determinism goal. Remove it, or document it as "never set by default"?
> *answer:*
**4. What are the path normalization rules?**
The proto has `string path` with no specification about: always forward-slash? Must be relative? No `..` components allowed? UTF-8 NFC vs NFD normalization (macOS vs Linux)? Max path length? This is a security issue (path traversal) and a cross-platform compatibility issue. What rules should the spec mandate?
> *answer:*
**5. Should we add a version byte after the magic?**
Currently `ZNAVSRFG` is followed immediately by protobuf. Adding a version byte (`ZNAVSRFG\x01`) would allow future framing changes without requiring protobuf parsing to detect the version. `MFFileOuter.Version` serves this purpose but requires successful deserialization to read. Worth the extra byte?
> *answer:*
**6. Should we add a length-prefix after the magic?**
Protobuf is not self-delimiting. If we ever want to concatenate manifests or append data after the protobuf, the current framing is insufficient. Add a varint or fixed-width length-prefix?
> *answer:*
### Signature Design
**7. What does the outer SHA-256 hash cover — compressed or uncompressed data?**
The review notes it currently hashes compressed data (good for verifying before decompression), but this should be explicitly documented. Which is the intended behavior?
> *answer:*
**8. Should `signatureString()` sign raw bytes instead of a hex-encoded string?**
Currently the canonical string is `MAGIC-UUID-MULTIHASH` with hex encoding, which adds a transformation layer. Signing the raw `sha256` bytes (or compressed `innerMessage` directly) would be simpler. Keep the string format or switch to raw bytes?
> *answer:*
**9. Should we support detached signature files (`.mf.sig`)?**
Embedded signatures are better for single-file distribution. Detached `.mf.sig` files follow the familiar `SHASUMS`/`SHASUMS.asc` pattern and are simpler for HTTP serving. Support both modes?
> *answer:*
**10. GPG vs pure-Go crypto for signatures?**
Shelling out to `gpg` is fragile (may not be installed, version-dependent output). `github.com/ProtonMail/go-crypto` provides pure-Go OpenPGP, or we could go Ed25519/signify (simpler, no key management). Which direction?
> *answer:*
### Implementation Design
**11. Should manifests be deterministic by default?**
This means: sort file entries by path, omit `createdAt` timestamp (or make it opt-in), no `atime`. Should determinism be the default, with a `--include-timestamps` flag to opt in?
> *answer:*
**12. Should we consolidate or keep both scanner/checker implementations?**
There are two parallel implementations: `mfer/scanner.go` + `mfer/checker.go` (typed with `FileSize`, `RelFilePath`) and `internal/scanner/` + `internal/checker/` (raw `int64`, `string`). The `mfer/` versions are superior. Delete the `internal/` versions?
> *answer:*
**13. Should the `manifest` type be exported?**
Currently unexported with exported constructors (`New`, `NewFromPaths`, etc.). Consumers can't declare `var m *mfer.manifest`. Export the type, or define an interface?
> *answer:*
**14. What should the Go module path be for 1.0?**
Currently mixed between `sneak.berlin/go/mfer` and `git.eeqj.de/sneak/mfer`. Which is canonical?
> *answer:*
---
## Implementation Plan
### Phase 1: Foundation (format correctness)
- [ ] Delete `internal/scanner/` and `internal/checker/` — consolidate on `mfer/` package versions; update CLI code
- [ ] Add deterministic file ordering — sort entries by path (lexicographic, byte-order) in `Builder.Build()`; add test asserting byte-identical output from two runs
- [ ] Add decompression size limit — `io.LimitReader` in `deserializeInner()` with `m.pbOuter.Size` as bound
- [ ] Fix `errors.Is` dead code in checker — replace with `os.IsNotExist(err)` or `errors.Is(err, fs.ErrNotExist)`
- [ ] Fix `AddFile` to verify size — check `totalRead == size` after reading, return error on mismatch
- [ ] Specify path invariants — add proto comments (UTF-8, forward-slash, relative, no `..`, no leading `/`); validate in `Builder.AddFile` and `Builder.AddFileWithHash`
### Phase 2: CLI polish
- [ ] Fix flag naming — all CLI flags use kebab-case as primary (`--include-dotfiles`, `--follow-symlinks`)
- [ ] Fix URL construction in fetch — use `BaseURL.JoinPath()` or `url.JoinPath()` instead of string concatenation
- [ ] Add progress rate-limiting to Checker — throttle to once per second, matching Scanner
- [ ] Add `--deterministic` flag (or make it default) — omit `createdAt`, sort files
### Phase 3: Robustness
- [ ] Replace GPG subprocess with pure-Go crypto — `github.com/ProtonMail/go-crypto` or Ed25519/signify
- [ ] Add timeout to any remaining subprocess calls
- [ ] Add fuzzing tests for `NewManifestFromReader`
- [ ] Add retry logic to fetch — exponential backoff for transient HTTP errors
### Phase 4: Format finalization
- [ ] Remove or deprecate `atime` from proto (pending design question answer)
- [ ] Reserve `optional uint32 mode = 305` in `MFFilePath` for future file permissions
- [ ] Add version byte after magic — `ZNAVSRFG\x01` for format version 1
- [ ] Write format specification document — separate from README: magic, outer structure, compression, inner structure, path invariants, signature scheme, canonical serialization
### Phase 5: Release prep
- [ ] Finalize Go module path
- [ ] Audit all error messages for consistency and helpfulness
- [ ] Add `--version` output matching SemVer
- [ ] Tag v1.0.0
-12
View File
@@ -1,12 +0,0 @@
#!/bin/bash
#
if [[ ! -z "$DRONE_COMMIT_SHA" ]]; then
echo "${DRONE_COMMIT_SHA:0:7}"
exit 0
fi
if [[ ! -z "$GITREV" ]]; then
echo $GITREV
else
git describe --always --dirty=-dirty
fi
+22
View File
@@ -0,0 +1,22 @@
// The developer tools this repo runs with `go tool`: gofumpt for
// script/gofumpt and protoc-gen-go for script/generate. Kept out of the mfer
// module so they add nothing to what mfer's users download. `go tool` builds
// exactly the source whose hashes go.sum here records.
module sneak.berlin/go/mfer/bin/tools
go 1.26.0
tool (
google.golang.org/protobuf/cmd/protoc-gen-go
mvdan.cc/gofumpt
)
require (
golang.org/x/mod v0.40.0 // indirect
golang.org/x/sync v0.22.0 // indirect
golang.org/x/tools v0.49.0 // indirect
// protoc-gen-go v1.36.11, 2026-10-04
google.golang.org/protobuf v1.36.11 // indirect
// gofumpt v0.12.0, 2026-10-04
mvdan.cc/gofumpt v0.12.0 // indirect
)
+22
View File
@@ -0,0 +1,22 @@
github.com/go-quicktest/qt v1.102.0 h1:HSQxCeh5YZH3EL3W39ixjtyaEhcWSXQHtHnMBzSs474=
github.com/go-quicktest/qt v1.102.0/go.mod h1:p4lGIVX+8Wa6ZPNDvqcxq36XpUDLh42FLetFU7odllI=
github.com/google/go-cmp v0.7.0 h1:wk8382ETsv4JYUZwIsn6YpYiWiBsYLSJiTsyBybVuN8=
github.com/google/go-cmp v0.7.0/go.mod h1:pXiqmnSA92OHEEa9HXL2W4E7lf9JzCmGVUdgjX3N/iU=
github.com/kr/pretty v0.3.1 h1:flRD4NNwYAUpkphVc1HcthR4KEIFJ65n8Mw5qdRn3LE=
github.com/kr/pretty v0.3.1/go.mod h1:hoEshYVHaxMs3cyo3Yncou5ZscifuDolrwPKZanG3xk=
github.com/kr/text v0.2.0 h1:5Nx0Ya0ZqY2ygV366QzturHI13Jq95ApcVaJBhpS+AY=
github.com/kr/text v0.2.0/go.mod h1:eLer722TekiGuMkidMxC/pM04lWEeraHUUmBw8l2grE=
github.com/rogpeppe/go-internal v1.16.0 h1:O9DK+vNMDVGLr2BeZqmpLeMjiMNkuXfcqntWbZV6S5g=
github.com/rogpeppe/go-internal v1.16.0/go.mod h1:DrUVZyrJU+txYW5/1kwtXQSMFio52ZOxX7yM1VHvnxs=
golang.org/x/mod v0.40.0 h1:hUv+3cXcdRHz08UmSiOob7sadHig73uo5bkXxQ/tvUs=
golang.org/x/mod v0.40.0/go.mod h1:0/weTWkPWGBikyTWAX3dkjVztMmBA5hM0DH6BElSupE=
golang.org/x/sync v0.22.0 h1:SZjpbeLmrCk4xhRSZFNZW5gFUeCeFgjekvI/+gfScek=
golang.org/x/sync v0.22.0/go.mod h1:9xrNwdLfx4jkKbNva9FpL6vEN7evnE43NNNJQ2LF3+0=
golang.org/x/sys v0.47.0 h1:o7XGOvZQCADBQQ4Y7VNq2dRWQR7JmOUW8Kxx4ZsNgWs=
golang.org/x/sys v0.47.0/go.mod h1:4GL1E5IUh+htKOUEOaiffhrAeqysfVGipDYzABqnCmw=
golang.org/x/tools v0.49.0 h1:3NI7VXzL9+1WZD52Dx2ttoPwD5DWrFGpl9mFZDlmisI=
golang.org/x/tools v0.49.0/go.mod h1:SJNXV9DBKT0UbdttsQjbfJlAE/q+y36++zo3uL3N0Oo=
google.golang.org/protobuf v1.36.11 h1:fV6ZwhNocDyBLK0dj+fg8ektcVegBBuEolpbTQyBNVE=
google.golang.org/protobuf v1.36.11/go.mod h1:HTf+CrKn2C3g5S8VImy6tdcUvCska2kB7j23XfzDpco=
mvdan.cc/gofumpt v0.12.0 h1:1Lbudkz2kpM9Cjz2pL4M19u7q+GaEhCTNf7N9mfpcho=
mvdan.cc/gofumpt v0.12.0/go.mod h1:SmBHHrljiZu/uoypeKup3rFzP6eoC9UwCp2iH5E3jZA=
Executable
+23
View File
@@ -0,0 +1,23 @@
#!/usr/bin/env bash
set -euo pipefail
# usage.sh - Generate and check a manifest from the repo
# Run from repo root: bin/usage.sh
cleanup() {
rm -rf "$TMPDIR"
}
main() {
TMPDIR=$(mktemp -d)
MANIFEST="$TMPDIR/index.mf"
trap cleanup EXIT
echo "Building mfer..."
go build -o "$TMPDIR/mfer" ./cmd/mfer
"$TMPDIR/mfer" generate -o "$MANIFEST" .
"$TMPDIR/mfer" check --base . "$MANIFEST"
}
main "$@"
+7 -1
View File
@@ -1,3 +1,4 @@
// Command mfer generates and verifies file manifests.
package main
import (
@@ -6,8 +7,13 @@ import (
"sneak.berlin/go/mfer/internal/cli"
)
// Appname is the name of this program.
const Appname = "mfer"
// Version and Gitrev are injected at build time via -ldflags.
//
//nolint:gochecknoglobals // set via ldflags at build time
var (
Appname string = "mfer"
Version string
Gitrev string
)
+16 -2
View File
@@ -6,6 +6,20 @@ import (
"github.com/stretchr/testify/assert"
)
func TestBuild(t *testing.T) {
assert.True(t, true)
// TestAppname pins the program name that main passes to cli.Run; it is
// the name that appears in usage output and in the log prefix. It also
// keeps this package compiled under `go test`.
func TestAppname(t *testing.T) {
t.Parallel()
assert.Equal(t, "mfer", Appname)
}
// TestVersionDefaults documents that Version and Gitrev are empty unless
// injected at build time via -ldflags.
func TestVersionDefaults(t *testing.T) {
t.Parallel()
assert.Empty(t, Version)
assert.Empty(t, Gitrev)
}
-19
View File
@@ -1,19 +0,0 @@
#!/bin/bash
set -euo pipefail
# usage.sh - Generate and check a manifest from the repo
# Run from repo root: ./contrib/usage.sh
TMPDIR=$(mktemp -d)
MANIFEST="$TMPDIR/index.mf"
cleanup() {
rm -rf "$TMPDIR"
}
trap cleanup EXIT
echo "Building mfer..."
go build -o "$TMPDIR/mfer" ./cmd/mfer
"$TMPDIR/mfer" generate -o "$MANIFEST" .
"$TMPDIR/mfer" check --base . "$MANIFEST"
+148
View File
@@ -0,0 +1,148 @@
# .mf File Format Specification
Version 1.0
## Overview
An `.mf` file is a binary manifest that describes a directory tree of files,
including their paths, sizes, and cryptographic checksums. It supports optional
GPG signatures for integrity verification and optional timestamps for metadata
preservation.
## File Structure
An `.mf` file consists of two parts, concatenated:
1. **Magic bytes** (8 bytes): the ASCII string `ZNAVSRFG`
2. **Outer message**: a Protocol Buffers serialized `MFFileOuter` message
There is no length prefix or version byte between the magic and the protobuf
message. The protobuf message extends to the end of the file.
See [`mfer/mf.proto`](../mfer/mf.proto) for exact field numbers and types.
## Outer Message (`MFFileOuter`)
The outer message contains:
| Field | Number | Type | Description |
| ----------------- | ------ | ---------------- | ------------------------------------------------------------------------ |
| `version` | 101 | enum | Must be `VERSION_ONE` (1) |
| `compressionType` | 102 | enum | Compression of `innerMessage`; must be `COMPRESSION_ZSTD` (1) |
| `size` | 103 | int64 | Uncompressed size of `innerMessage` (corruption detection) |
| `sha256` | 104 | bytes | SHA-256 hash of the **compressed** `innerMessage` (corruption detection) |
| `uuid` | 105 | bytes | Random v4 UUID; must match the inner message UUID |
| `innerMessage` | 199 | bytes | Zstd-compressed serialized `MFFile` message |
| `signature` | 201 | bytes (optional) | GPG signature (ASCII-armored or binary) |
| `signer` | 202 | bytes (optional) | Full GPG key ID of the signer |
| `signingPubKey` | 203 | bytes (optional) | Full GPG signing public key |
### SHA-256 Hash
The `sha256` field (104) covers the **compressed** `innerMessage` bytes. This
allows verifying data integrity before decompression.
## Compression
The `innerMessage` field is compressed with
[Zstandard (zstd)](https://facebook.github.io/zstd/). Implementations must
enforce a decompression size limit to prevent decompression bombs. The reference
implementation limits decompressed size to 256 MB. It writes zstd frames with a
window of at most 8 MiB, the largest window the zstd format recommends decoders
support, and refuses frames that ask for a larger one. It also refuses an inner
message whose file entries, hashes, timestamps and MIME types, counted at 160,
112, 64 and 16 bytes each, add up to more than 8 times its size.
## Inner Message (`MFFile`)
After decompressing `innerMessage`, the result is a serialized `MFFile`
(referred to as the manifest):
| Field | Number | Type | Description |
| ----------- | ------ | --------------------- | ------------------------------------- |
| `version` | 100 | enum | Must be `VERSION_ONE` (1) |
| `files` | 101 | repeated `MFFilePath` | List of files in the manifest |
| `uuid` | 102 | bytes | Random v4 UUID; must match outer UUID |
| `createdAt` | 201 | Timestamp (optional) | When the manifest was created |
## File Entries (`MFFilePath`)
Each file entry contains:
| Field | Number | Type | Description |
| ---------- | ------ | ------------------------- | ----------------------------------- |
| `path` | 1 | string | Relative file path (see Path Rules) |
| `size` | 2 | int64 | File size in bytes |
| `hashes` | 3 | repeated `MFFileChecksum` | At least one hash required |
| `mimeType` | 301 | string (optional) | MIME type |
| `mtime` | 302 | Timestamp (optional) | Modification time |
| `ctime` | 303 | Timestamp (optional) | Change time (inode metadata change) |
Field 304 (`atime`) has been removed from the specification. Access time is
volatile and non-deterministic; it is not useful for integrity verification.
## Path Rules
All `path` values must satisfy these invariants:
- **UTF-8**: paths must be valid UTF-8
- **Forward slashes**: use `/` as the path separator (never `\`)
- **Relative only**: no leading `/`
- **No parent traversal**: no `..` path segments
- **No empty segments**: no `//` sequences
- **No trailing slash**: paths refer to files, not directories
Implementations must validate these invariants when reading and writing
manifests. Paths that violate these rules must be rejected.
## Hash Format (`MFFileChecksum`)
Each checksum is a single `bytes multiHash` field containing a
[multihash](https://multiformats.io/multihash/)-encoded value. Multihash is
self-describing: the encoded bytes include a varint algorithm identifier
followed by a varint digest length followed by the digest itself.
The 1.0 implementation writes SHA-256 multihashes (`0x12` algorithm code).
Implementations must be able to verify SHA-256 multihashes at minimum.
## Signature Scheme
Signing is optional. When present, the signature covers a canonical string
constructed as:
```
ZNAVSRFG-<UUID>-<SHA256>
```
Where:
- `ZNAVSRFG` is the magic bytes string (literal ASCII)
- `<UUID>` is the hex-encoded UUID from the outer message
- `<SHA256>` is the hex-encoded SHA-256 hash from the outer message (covering
compressed data)
Components are separated by hyphens. The signature is produced by GPG over this
canonical string and stored in the `signature` field of the outer message.
## Deterministic Serialization
By default, manifests are generated deterministically:
- File entries are sorted by `path` in **lexicographic byte order**
- `createdAt` is omitted unless explicitly requested
- `atime` is never included (field removed from schema)
This ensures that two independent runs over the same directory tree produce
byte-identical `.mf` files (assuming file contents and metadata have not
changed).
## MIME Type
The recommended MIME type for `.mf` files is `application/octet-stream`. The
`.mf` file extension is the canonical identifier.
## Reference
- Proto definition: [`mfer/mf.proto`](../mfer/mf.proto)
- Reference implementation:
[git.eeqj.de/sneak/mfer](https://git.eeqj.de/sneak/mfer)
+4 -19
View File
@@ -3,42 +3,27 @@ module sneak.berlin/go/mfer
go 1.23
require (
github.com/apex/log v1.9.0
github.com/davecgh/go-spew v1.1.1
github.com/dustin/go-humanize v1.0.1
github.com/google/uuid v1.1.2
github.com/klauspost/compress v1.18.2
github.com/multiformats/go-multihash v0.2.3
github.com/pterm/pterm v0.12.35
github.com/spf13/afero v1.8.0
github.com/stretchr/testify v1.8.1
github.com/urfave/cli/v2 v2.23.6
github.com/stretchr/testify v1.12.1
github.com/urfave/cli/v3 v3.14.0
golang.org/x/term v0.0.0-20210927222741-03fcf44c2211
google.golang.org/protobuf v1.28.1
)
require (
github.com/atomicgo/cursor v0.0.1 // indirect
github.com/cpuguy83/go-md2man/v2 v2.0.2 // indirect
github.com/fatih/color v1.7.0 // indirect
github.com/gookit/color v1.4.2 // indirect
github.com/klauspost/cpuid/v2 v2.0.9 // indirect
github.com/mattn/go-colorable v0.1.2 // indirect
github.com/mattn/go-isatty v0.0.8 // indirect
github.com/mattn/go-runewidth v0.0.13 // indirect
github.com/minio/sha256-simd v1.0.0 // indirect
github.com/mr-tron/base58 v1.2.0 // indirect
github.com/multiformats/go-varint v0.0.6 // indirect
github.com/pkg/errors v0.9.1 // indirect
github.com/pmezard/go-difflib v1.0.0 // indirect
github.com/rivo/uniseg v0.2.0 // indirect
github.com/russross/blackfriday/v2 v2.1.0 // indirect
github.com/spaolacci/murmur3 v1.1.0 // indirect
github.com/xo/terminfo v0.0.0-20210125001918-ca9a967f8778 // indirect
github.com/xrash/smetrics v0.0.0-20201216005158-039620a65673 // indirect
go.yaml.in/yaml/v3 v3.0.5 // indirect
golang.org/x/crypto v0.0.0-20220525230936-793ad666bf5e // indirect
golang.org/x/sys v0.1.0 // indirect
golang.org/x/term v0.0.0-20210927222741-03fcf44c2211 // indirect
golang.org/x/text v0.3.6 // indirect
gopkg.in/yaml.v3 v3.0.1 // indirect
lukechampine.com/blake3 v1.1.6 // indirect
)
+6 -99
View File
@@ -38,21 +38,6 @@ cloud.google.com/go/storage v1.14.0/go.mod h1:GrKmX003DSIwi9o29oFT7YDnHYwZoctc3f
dmitri.shuralyov.com/gpu/mtl v0.0.0-20190408044501-666a987793e9/go.mod h1:H6x//7gZCb22OMCxBHrMx7a5I7Hp++hsVxbQ4BYO7hU=
github.com/BurntSushi/toml v0.3.1/go.mod h1:xHWCNGjB5oqiDr8zfno3MHue2Ht5sIBksp03qcyfWMU=
github.com/BurntSushi/xgb v0.0.0-20160522181843-27f122750802/go.mod h1:IVnqGOEym/WlBOVXweHU+Q+/VP0lqqI8lqeDx9IjBqo=
github.com/MarvinJWendt/testza v0.1.0/go.mod h1:7AxNvlfeHP7Z/hDQ5JtE3OKYT3XFUeLCDE2DQninSqs=
github.com/MarvinJWendt/testza v0.2.1/go.mod h1:God7bhG8n6uQxwdScay+gjm9/LnO4D3kkcZX4hv9Rp8=
github.com/MarvinJWendt/testza v0.2.8/go.mod h1:nwIcjmr0Zz+Rcwfh3/4UhBp7ePKVhuBExvZqnKYWlII=
github.com/MarvinJWendt/testza v0.2.10/go.mod h1:pd+VWsoGUiFtq+hRKSU1Bktnn+DMCSrDrXDpX2bG66k=
github.com/MarvinJWendt/testza v0.2.12 h1:/PRp/BF+27t2ZxynTiqj0nyND5PbOtfJS0SuTuxmgeg=
github.com/MarvinJWendt/testza v0.2.12/go.mod h1:JOIegYyV7rX+7VZ9r77L/eH6CfJHHzXjB69adAhzZkI=
github.com/apex/log v1.9.0 h1:FHtw/xuaM8AgmvDDTI9fiwoAL25Sq2cxojnZICUU8l0=
github.com/apex/log v1.9.0/go.mod h1:m82fZlWIuiWzWP04XCTXmnX0xRkYYbCdYn8jbJeLBEA=
github.com/apex/logs v1.0.0/go.mod h1:XzxuLZ5myVHDy9SAmYpamKKRNApGj54PfYLcFrXqDwo=
github.com/aphistic/golf v0.0.0-20180712155816-02c07f170c5a/go.mod h1:3NqKYiepwy8kCu4PNA+aP7WUV72eXWJeP9/r3/K9aLE=
github.com/aphistic/sweet v0.2.0/go.mod h1:fWDlIh/isSE9n6EPsRmC0det+whmX6dJid3stzu0Xys=
github.com/atomicgo/cursor v0.0.1 h1:xdogsqa6YYlLfM+GyClC/Lchf7aiMerFiZQn7soTOoU=
github.com/atomicgo/cursor v0.0.1/go.mod h1:cBON2QmmrysudxNBFthvMtN32r3jxVRIvzkUiF/RuIk=
github.com/aws/aws-sdk-go v1.20.6/go.mod h1:KmX6BPdI08NWTb3/sm4ZGu5ShLoqVDhKgpiN924inxo=
github.com/aybabtme/rgbterm v0.0.0-20170906152045-cc83f3b3ce59/go.mod h1:q/89r3U2H7sSsE2t6Kca0lfwTK8JdoNGS/yzM/4iH5I=
github.com/census-instrumentation/opencensus-proto v0.2.1/go.mod h1:f6KPmirojxKA12rnyqOA5BBL4O983OfeGPqjHWSTneU=
github.com/chzyer/logex v1.1.10/go.mod h1:+Ywpsq7O8HXn0nuIou7OrIPyXbp3wmkHB+jjWRnGsAI=
github.com/chzyer/readline v0.0.0-20180603132655-2972be24d48e/go.mod h1:nSuG5e5PlCu98SY8svDHJxuZscDgtXS6KTTbou5AhLI=
@@ -61,8 +46,6 @@ github.com/client9/misspell v0.3.4/go.mod h1:qj6jICC3Q7zFZvVWo7KLAzC3yx5G7kyvSDk
github.com/cncf/udpa/go v0.0.0-20191209042840-269d4d468f6f/go.mod h1:M8M6+tZqaGXZJjfX53e64911xZQV5JYwmTeXPW+k8Sc=
github.com/cncf/udpa/go v0.0.0-20200629203442-efcf912fb354/go.mod h1:WmhPx2Nbnhtbo57+VJT5O0JRkEi1Wbu0z5j0R8u5Hbk=
github.com/cncf/udpa/go v0.0.0-20201120205902-5459f2c99403/go.mod h1:WmhPx2Nbnhtbo57+VJT5O0JRkEi1Wbu0z5j0R8u5Hbk=
github.com/cpuguy83/go-md2man/v2 v2.0.2 h1:p1EgwI/C7NhT0JmVkwCD2ZBK8j4aeHQX2pMHHBfMQ6w=
github.com/cpuguy83/go-md2man/v2 v2.0.2/go.mod h1:tgQtvFlXSQOSOSIRvRPT7W67SCa46tRHOmNcaadrF8o=
github.com/davecgh/go-spew v1.1.0/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
@@ -74,13 +57,9 @@ github.com/envoyproxy/go-control-plane v0.9.4/go.mod h1:6rpuAdCZL397s3pYoYcLgu1m
github.com/envoyproxy/go-control-plane v0.9.7/go.mod h1:cwu0lG7PUMfa9snN8LXBig5ynNVH9qI8YYLbd1fK2po=
github.com/envoyproxy/go-control-plane v0.9.9-0.20201210154907-fd9021fe5dad/go.mod h1:cXg6YxExXjJnVBQHBLXeUAgxn2UodCpnH306RInaBQk=
github.com/envoyproxy/protoc-gen-validate v0.1.0/go.mod h1:iSmxcyjqTsJpI2R4NaDN7+kN2VEUnK/pcBlmesArF7c=
github.com/fatih/color v1.7.0 h1:DkWD4oS2D8LGGgTQ6IvwJJXSL5Vp2ffcQg58nFV38Ys=
github.com/fatih/color v1.7.0/go.mod h1:Zm6kSWBoL9eyXnKyktHP6abPY2pDugNf5KwzbycvMj4=
github.com/fsnotify/fsnotify v1.4.7/go.mod h1:jwhsz4b93w/PPRr/qN1Yymfu8t87LnFCMoQvtojpjFo=
github.com/go-gl/glfw v0.0.0-20190409004039-e6da0acd62b1/go.mod h1:vR7hzQXu2zJy9AVAgeJqvqgH9Q5CA+iKCZ2gyEVpxRU=
github.com/go-gl/glfw/v3.3/glfw v0.0.0-20191125211704-12ad95a8df72/go.mod h1:tQ2UAYgL5IevRw8kRxooKSPJfGvJ9fJQFa0TUsXzTg8=
github.com/go-gl/glfw/v3.3/glfw v0.0.0-20200222043503-6f7a984d4dc4/go.mod h1:tQ2UAYgL5IevRw8kRxooKSPJfGvJ9fJQFa0TUsXzTg8=
github.com/go-logfmt/logfmt v0.4.0/go.mod h1:3RMwSq7FuexP4Kalkev3ejPJsZTpXXBr9+V4qmtdjCk=
github.com/golang/glog v0.0.0-20160126235308-23def4e6c14b/go.mod h1:SBH7ygxi8pfUlaOkMMuAQtPIUF8ecWP5IEl/CR7VP2Q=
github.com/golang/groupcache v0.0.0-20190702054246-869f871628b6/go.mod h1:cIg4eruTrX1D+g88fzRXU5OdNfaM+9IcxsU14FzY7Hc=
github.com/golang/groupcache v0.0.0-20191227052852-215e87163ea7/go.mod h1:cIg4eruTrX1D+g88fzRXU5OdNfaM+9IcxsU14FzY7Hc=
@@ -134,21 +113,15 @@ github.com/google/pprof v0.0.0-20201023163331-3e6fc7fc9c4c/go.mod h1:kpwsk12EmLe
github.com/google/pprof v0.0.0-20201203190320-1bf35d6f28c2/go.mod h1:kpwsk12EmLew5upagYY7GY0pfYCcupk39gWOCRROcvE=
github.com/google/pprof v0.0.0-20201218002935-b9804c9f04c2/go.mod h1:kpwsk12EmLew5upagYY7GY0pfYCcupk39gWOCRROcvE=
github.com/google/renameio v0.1.0/go.mod h1:KWCgfxg9yswjAJkECMjeO8J8rahYeXnNhOm40UhjYkI=
github.com/google/uuid v1.1.1/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
github.com/google/uuid v1.1.2 h1:EVhdT+1Kseyi1/pUmXKaFxYsDNy9RQYkMWRH68J/W7Y=
github.com/google/uuid v1.1.2/go.mod h1:TIyPZe4MgqvfeYDBFedMoGGpEw/LqOeaOT+nhxU+yHo=
github.com/googleapis/gax-go/v2 v2.0.4/go.mod h1:0Wqv26UfaUD9n4G6kQubkQ+KchISgw+vpHVxEJEs9eg=
github.com/googleapis/gax-go/v2 v2.0.5/go.mod h1:DWXyrwAJ9X0FpwwEdw+IPEYBICEFu5mhpdKc/us6bOk=
github.com/googleapis/google-cloud-go-testing v0.0.0-20200911160855-bcd43fbb19e8/go.mod h1:dvDLG8qkwmyD9a/MJJN3XJcT3xFxOKAvTZGvuZmac9g=
github.com/gookit/color v1.4.2 h1:tXy44JFSFkKnELV6WaMo/lLfu/meqITX3iAV52do7lk=
github.com/gookit/color v1.4.2/go.mod h1:fqRyamkC1W8uxl+lxCQxOT09l/vYfZ+QeiX3rKQHCoQ=
github.com/hashicorp/golang-lru v0.5.0/go.mod h1:/m3WP610KZHVQ1SGc6re/UDhFvYD7pJ4Ao+sR/qLZy8=
github.com/hashicorp/golang-lru v0.5.1/go.mod h1:/m3WP610KZHVQ1SGc6re/UDhFvYD7pJ4Ao+sR/qLZy8=
github.com/hpcloud/tail v1.0.0/go.mod h1:ab1qPbhIpdTxEkNHXyeSf5vhxWSCs/tWer42PpOxQnU=
github.com/ianlancetaylor/demangle v0.0.0-20181102032728-5e5cf60278f6/go.mod h1:aSSvb/t6k1mPoxDqO4vJh6VOCGPwU4O0C2/Eqndh1Sc=
github.com/ianlancetaylor/demangle v0.0.0-20200824232613-28f6c0f3b639/go.mod h1:aSSvb/t6k1mPoxDqO4vJh6VOCGPwU4O0C2/Eqndh1Sc=
github.com/jmespath/go-jmespath v0.0.0-20180206201540-c2b33e8439af/go.mod h1:Nht3zPeWKUH0NzdCt2Blrr5ys8VGpn0CEB0cQHVjt7k=
github.com/jpillora/backoff v0.0.0-20180909062703-3050d21c67d7/go.mod h1:2iMrUgbbvHEiQClaW2NsSzMyGHqN+rDFqY705q49KG0=
github.com/jstemmer/go-junit-report v0.0.0-20190106144839-af01ea7f8024/go.mod h1:6v2b51hI/fHJwM22ozAgKL4VKDeJcHhJFhtBdhmNjmU=
github.com/jstemmer/go-junit-report v0.9.1/go.mod h1:Brl9GWCQeLvo8nXZwPNNblvFj/XSXhF0NWZEnDohbsk=
github.com/kisielk/gotool v1.0.0/go.mod h1:XhKaO+MFFWcvkIS/tQcRk01m1F5IRFswLeQ+oQHNcck=
@@ -158,22 +131,9 @@ github.com/klauspost/cpuid/v2 v2.0.4/go.mod h1:FInQzS24/EEf25PyTYn52gqo7WaD8xa02
github.com/klauspost/cpuid/v2 v2.0.9 h1:lgaqFMSdTdQYdZ04uHyN2d/eKdOMyi2YLSvlQIBFYa4=
github.com/klauspost/cpuid/v2 v2.0.9/go.mod h1:FInQzS24/EEf25PyTYn52gqo7WaD8xa0213Md/qVLRg=
github.com/kr/fs v0.1.0/go.mod h1:FFnZGqtBN9Gxj7eW1uZ42v5BccTP0vu6NEaFoC2HwRg=
github.com/kr/logfmt v0.0.0-20140226030751-b84e30acd515/go.mod h1:+0opPa2QZZtGFBFZlji/RkVcI2GknAs/DXo4wKdlNEc=
github.com/kr/pretty v0.1.0/go.mod h1:dAy3ld7l9f0ibDNOQOHHMYYIIbhfbHSm3C4ZsoJORNo=
github.com/kr/pretty v0.2.0 h1:s5hAObm+yFO5uHYt5dYjxi2rXrsnmRpJx4OYvIWUaQs=
github.com/kr/pretty v0.2.0/go.mod h1:ipq/a2n7PKx3OHsz4KJII5eveXtPO4qwEXGdVfWzfnI=
github.com/kr/pty v1.1.1/go.mod h1:pFQYn66WHrOpPYNljwOMqo10TkYh1fy3cYio2l3bCsQ=
github.com/kr/text v0.1.0 h1:45sCR5RtlFHMR4UwH9sdQ5TC8v0qDQCHnXt+kaKSTVE=
github.com/kr/text v0.1.0/go.mod h1:4Jbv+DJW3UT/LiOwJeYQe1efqtUx/iVham/4vfdArNI=
github.com/mattn/go-colorable v0.1.1/go.mod h1:FuOcm+DKB9mbwrcAfNl7/TZVBZ6rcnceauSikq3lYCQ=
github.com/mattn/go-colorable v0.1.2 h1:/bC9yWikZXAL9uJdulbSfyVNIR3n3trXl+v8+1sx8mU=
github.com/mattn/go-colorable v0.1.2/go.mod h1:U0ppj6V5qS13XJ6of8GYAs25YV2eR4EVcfRqFIhoBtE=
github.com/mattn/go-isatty v0.0.5/go.mod h1:Iq45c/XA43vh69/j3iqttzPXn0bhXyGjM0Hdxcsrc5s=
github.com/mattn/go-isatty v0.0.8 h1:HLtExJ+uU2HOZ+wI0Tt5DtUDrx8yhUqDcp7fYERX4CE=
github.com/mattn/go-isatty v0.0.8/go.mod h1:Iq45c/XA43vh69/j3iqttzPXn0bhXyGjM0Hdxcsrc5s=
github.com/mattn/go-runewidth v0.0.13 h1:lTGmDsbAYt5DmK6OnoV7EuIF1wEIFAcxld6ypU4OSgU=
github.com/mattn/go-runewidth v0.0.13/go.mod h1:Jdepj2loyihRzMpdS35Xk/zdY8IAYHsh153qUoGf23w=
github.com/mgutz/ansi v0.0.0-20170206155736-9520e82c474b/go.mod h1:01TrycV0kFyexm33Z7vhZRXopbI8J3TDReVlkTgMUxE=
github.com/minio/sha256-simd v1.0.0 h1:v1ta+49hkWZyvaKwrQB8elexRqm6Y0aMLjCNsrYxo6g=
github.com/minio/sha256-simd v1.0.0/go.mod h1:OuYzVNI5vcoYIAmbIvHPl3N3jUzVedXbKy5RFepssQM=
github.com/mr-tron/base58 v1.2.0 h1:T/HDJBh4ZCPbU39/+c3rRvE0uKBQlU27+QI8LJ4t64o=
@@ -182,61 +142,23 @@ github.com/multiformats/go-multihash v0.2.3 h1:7Lyc8XfX/IY2jWb/gI7JP+o7JEq9hOa7B
github.com/multiformats/go-multihash v0.2.3/go.mod h1:dXgKXCXjBzdscBLk9JkjINiEsCKRVch90MdaGiKsvSM=
github.com/multiformats/go-varint v0.0.6 h1:gk85QWKxh3TazbLxED/NlDVv8+q+ReFJk7Y2W/KhfNY=
github.com/multiformats/go-varint v0.0.6/go.mod h1:3Ls8CIEsrijN6+B7PbrXRPxHRPuXSrVKRY101jdMZYE=
github.com/onsi/ginkgo v1.6.0/go.mod h1:lLunBs/Ym6LB5Z9jYTR76FiuTmxDTDusOGeTQH+WWjE=
github.com/onsi/gomega v1.5.0/go.mod h1:ex+gbHU/CVuBBDIJjb2X0qEXbFg53c61hWP/1CpauHY=
github.com/pkg/errors v0.8.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINEl0=
github.com/pkg/errors v0.9.1 h1:FEBLx1zS214owpjy7qsBeixbURkuhQAwrK5UwLGTwt4=
github.com/pkg/errors v0.9.1/go.mod h1:bwawxfHBFNV+L2hUp1rHADufV3IMtnDRdf1r5NINEl0=
github.com/pkg/sftp v1.13.1/go.mod h1:3HaPG6Dq1ILlpPZRO0HVMrsydcdLt6HRDccSgb87qRg=
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
github.com/prometheus/client_model v0.0.0-20190812154241-14fe0d1b01d4/go.mod h1:xMI15A0UPsDsEKsMN9yxemIoYk6Tm2C1GtYGdfGttqA=
github.com/pterm/pterm v0.12.27/go.mod h1:PhQ89w4i95rhgE+xedAoqous6K9X+r6aSOI2eFF7DZI=
github.com/pterm/pterm v0.12.29/go.mod h1:WI3qxgvoQFFGKGjGnJR849gU0TsEOvKn5Q8LlY1U7lg=
github.com/pterm/pterm v0.12.30/go.mod h1:MOqLIyMOgmTDz9yorcYbcw+HsgoZo3BQfg2wtl3HEFE=
github.com/pterm/pterm v0.12.31/go.mod h1:32ZAWZVXD7ZfG0s8qqHXePte42kdz8ECtRyEejaWgXU=
github.com/pterm/pterm v0.12.33/go.mod h1:x+h2uL+n7CP/rel9+bImHD5lF3nM9vJj80k9ybiiTTE=
github.com/pterm/pterm v0.12.35 h1:A/vHwDM+WByn0sTPlpL2L6kOTy12xqZuwNFMF/NlA+U=
github.com/pterm/pterm v0.12.35/go.mod h1:NjiL09hFhT/vWjQHSj1athJpx6H8cjpHXNAK5bUw8T8=
github.com/rivo/uniseg v0.2.0 h1:S1pD9weZBuJdFmowNwbpi7BJ8TNftyUImj/0WQi72jY=
github.com/rivo/uniseg v0.2.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJtxc=
github.com/rogpeppe/fastuuid v1.1.0/go.mod h1:jVj6XXZzXRy/MSR5jhDC/2q6DgLz+nrA6LYCDYWNEvQ=
github.com/rogpeppe/go-internal v1.3.0/go.mod h1:M8bDsm7K2OlrFYOpmOWEs/qY81heoFRclV5y23lUDJ4=
github.com/russross/blackfriday/v2 v2.1.0 h1:JIOH55/0cWyOuilr9/qlrm0BSXldqnqwMsf35Ld67mk=
github.com/russross/blackfriday/v2 v2.1.0/go.mod h1:+Rmxgy9KzJVeS9/2gXHxylqXiyQDYRxCVz55jmeOWTM=
github.com/sergi/go-diff v1.0.0/go.mod h1:0CfEIISq7TuYL3j771MWULgwwjU+GofnZX9QAmXWZgo=
github.com/smartystreets/assertions v1.0.0/go.mod h1:kHHU4qYBaI3q23Pp3VPrmWhuIUrLW/7eUrw0BU5VaoM=
github.com/smartystreets/go-aws-auth v0.0.0-20180515143844-0c1422d1fdb9/go.mod h1:SnhjPscd9TpLiy1LpzGSKh3bXCfxxXuqd9xmQJy3slM=
github.com/smartystreets/gunit v1.0.0/go.mod h1:qwPWnhz6pn0NnRBP++URONOVyNkPyr4SauJk4cUOwJs=
github.com/spaolacci/murmur3 v1.1.0 h1:7c1g84S4BPRrfL5Xrdp6fOJ206sU9y293DDHaoy0bLI=
github.com/spaolacci/murmur3 v1.1.0/go.mod h1:JwIasOWyU6f++ZhiEuf87xNszmSA2myDM2Kzu9HwQUA=
github.com/spf13/afero v1.8.0 h1:5MmtuhAgYeU6qpa7w7bP0dv6MBYuup0vekhSpSkoq60=
github.com/spf13/afero v1.8.0/go.mod h1:CtAatgMJh6bJEIs48Ay/FOnkljP3WeGUG0MC1RfAqwo=
github.com/stretchr/objx v0.1.0/go.mod h1:HFkY916IF+rwdDfMAkV7OtwuqBVzrE8GR6GFx+wExME=
github.com/stretchr/objx v0.4.0/go.mod h1:YvHI0jy2hoMjB+UWwv71VJQ9isScKT/TqJzVSSt89Yw=
github.com/stretchr/objx v0.5.0/go.mod h1:Yh+to48EsGEfYuaHDzXPcE3xhTkx73EhmCGUpEOglKo=
github.com/stretchr/testify v1.3.0/go.mod h1:M5WIy9Dh21IEIfnGCwXGc5bZfKNJtfHm1UVUgZn+9EI=
github.com/stretchr/testify v1.4.0/go.mod h1:j7eGeouHqKxXV5pUuKE4zz7dFj8WfuZ+81PSLYec5m4=
github.com/stretchr/testify v1.5.1/go.mod h1:5W2xD1RspED5o8YsWQXVCued0rvSQ+mT+I5cxcmMvtA=
github.com/stretchr/testify v1.6.1/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
github.com/stretchr/testify v1.7.0/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
github.com/stretchr/testify v1.7.1/go.mod h1:6Fq8oRcR53rry900zMqJjRRixrwX3KX962/h/Wwjteg=
github.com/stretchr/testify v1.8.0/go.mod h1:yNjHg4UonilssWZ8iaSj1OCr/vHnekPRkoO+kdMU+MU=
github.com/stretchr/testify v1.8.1 h1:w7B6lhMri9wdJUVmEZPGGhZzrYTPvgJArz7wNPgYKsk=
github.com/stretchr/testify v1.8.1/go.mod h1:w2LPCIKwWwSfY2zedu0+kehJoqGctiVI29o6fzry7u4=
github.com/tj/assert v0.0.0-20171129193455-018094318fb0/go.mod h1:mZ9/Rh9oLWpLLDRpvE+3b7gP/C2YyLFYxNmcLnPTMe0=
github.com/tj/assert v0.0.3 h1:Df/BlaZ20mq6kuai7f5z2TvPFiwC3xaWJSDQNiIS3Rk=
github.com/tj/assert v0.0.3/go.mod h1:Ne6X72Q+TB1AteidzQncjw9PabbMp4PBMZ1k+vd1Pvk=
github.com/tj/go-buffer v1.1.0/go.mod h1:iyiJpfFcR2B9sXu7KvjbT9fpM4mOelRSDTbntVj52Uc=
github.com/tj/go-elastic v0.0.0-20171221160941-36157cbbebc2/go.mod h1:WjeM0Oo1eNAjXGDx2yma7uG2XoyRZTq1uv3M/o7imD0=
github.com/tj/go-kinesis v0.0.0-20171128231115-08b17f58cb1b/go.mod h1:/yhzCV0xPfx6jb1bBgRFjl5lytqVqZXEaeqWP8lTEao=
github.com/tj/go-spin v1.1.0/go.mod h1:Mg1mzmePZm4dva8Qz60H2lHwmJ2loum4VIrLgVnKwh4=
github.com/urfave/cli/v2 v2.23.6 h1:iWmtKD+prGo1nKUtLO0Wg4z9esfBM4rAV4QRLQiEmJ4=
github.com/urfave/cli/v2 v2.23.6/go.mod h1:GHupkWPMM0M/sj1a2b4wUrWBPzazNrIjouW6fmdJLxc=
github.com/xo/terminfo v0.0.0-20210125001918-ca9a967f8778 h1:QldyIu/L63oPpyvQmHgvgickp1Yw510KJOqX7H24mg8=
github.com/xo/terminfo v0.0.0-20210125001918-ca9a967f8778/go.mod h1:2MuV+tbUrU1zIOPMxZ5EncGwgmMJsa+9ucAQZXxsObs=
github.com/xrash/smetrics v0.0.0-20201216005158-039620a65673 h1:bAn7/zixMGCfxrRTfdpNzjtPYqr8smhKouy9mxVdGPU=
github.com/xrash/smetrics v0.0.0-20201216005158-039620a65673/go.mod h1:N3UwUGtsrSj3ccvlPHLoLsHnpR27oXr4ZE984MbSER8=
github.com/stretchr/testify v1.12.1 h1:EuwCh5fleGS7H32xRwO3wRGT7DxrDhLAT6FF8MpWDWE=
github.com/stretchr/testify v1.12.1/go.mod h1:MDEgiDPPsNp5cuIrHPPCyornHKgEVbtFUmoNlxoYthg=
github.com/urfave/cli/v3 v3.14.0 h1:a8414NQlHJs0c/iBsulKLzlES0n/lEAskbL2LKpU4/s=
github.com/urfave/cli/v3 v3.14.0/go.mod h1:vXn6HxPNccJSzQr2QvwVncOKrgYGIHU0HY5h8B2nQj4=
github.com/yuin/goldmark v1.1.25/go.mod h1:3hX8gzYuyVAZsxl0MRgGTJEmQBFcNTphYh9decYSb74=
github.com/yuin/goldmark v1.1.27/go.mod h1:3hX8gzYuyVAZsxl0MRgGTJEmQBFcNTphYh9decYSb74=
github.com/yuin/goldmark v1.1.32/go.mod h1:3hX8gzYuyVAZsxl0MRgGTJEmQBFcNTphYh9decYSb74=
@@ -247,8 +169,9 @@ go.opencensus.io v0.22.2/go.mod h1:yxeiOL68Rb0Xd1ddK5vPZ/oVn4vY4Ynel7k9FzqtOIw=
go.opencensus.io v0.22.3/go.mod h1:yxeiOL68Rb0Xd1ddK5vPZ/oVn4vY4Ynel7k9FzqtOIw=
go.opencensus.io v0.22.4/go.mod h1:yxeiOL68Rb0Xd1ddK5vPZ/oVn4vY4Ynel7k9FzqtOIw=
go.opencensus.io v0.22.5/go.mod h1:5pWMHQbX5EPX2/62yrJeAkowc+lfs/XD7Uxpq3pI6kk=
go.yaml.in/yaml/v3 v3.0.5 h1:N6y/pJk8buWs9NY5ERU2HSMfm+IuD/OtfdAnq6kESPw=
go.yaml.in/yaml/v3 v3.0.5/go.mod h1:HVTZu1O7/Vkt2N+BFy8Zza+lnLsABggaTM2ZpNIGuKg=
golang.org/x/crypto v0.0.0-20190308221718-c2843e01d9a2/go.mod h1:djNgcEr1/C05ACkg1iLfiJU5Ep61QUkGW8qpdssI0+w=
golang.org/x/crypto v0.0.0-20190426145343-a29dc8fdc734/go.mod h1:yigFU9vqHzYiE8UmvKecakEJjdnWj3jj499lnFckfCI=
golang.org/x/crypto v0.0.0-20190510104115-cbcb75029529/go.mod h1:yigFU9vqHzYiE8UmvKecakEJjdnWj3jj499lnFckfCI=
golang.org/x/crypto v0.0.0-20190605123033-f99c8df09eb5/go.mod h1:yigFU9vqHzYiE8UmvKecakEJjdnWj3jj499lnFckfCI=
golang.org/x/crypto v0.0.0-20191011191535-87dc89f01550/go.mod h1:yigFU9vqHzYiE8UmvKecakEJjdnWj3jj499lnFckfCI=
@@ -292,7 +215,6 @@ golang.org/x/mod v0.4.0/go.mod h1:s0Qsj1ACt9ePp/hMypM3fl4fZqREWJwdYDEqhRiZZUA=
golang.org/x/mod v0.4.1/go.mod h1:s0Qsj1ACt9ePp/hMypM3fl4fZqREWJwdYDEqhRiZZUA=
golang.org/x/net v0.0.0-20180724234803-3673e40ba225/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20180826012351-8a410e7b638d/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20180906233101-161cd47e91fd/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20190108225652-1e06a53dbb7e/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20190213061140-3a22650c66bd/go.mod h1:mL1N/T3taQHkDXs73rZJwtUhF3w3ftmwwsq0BUmARs4=
golang.org/x/net v0.0.0-20190311183353-d8887717615a/go.mod h1:t9HGtf8HONx5eT2rtn7q6eTqICYqUVnKs3thJo3Qplg=
@@ -342,9 +264,7 @@ golang.org/x/sync v0.0.0-20200625203802-6e8e738ad208/go.mod h1:RxMgew5VJxzue5/jJ
golang.org/x/sync v0.0.0-20201020160332-67f06af15bc9/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sync v0.0.0-20201207232520-09787c993a3a/go.mod h1:RxMgew5VJxzue5/jJTE5uejpjVlOe/izrB70Jof72aM=
golang.org/x/sys v0.0.0-20180830151530-49385e6e1522/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20180909124046-d0be0721c37e/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20190215142949-d0b11bdaac8a/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20190222072716-a9d3bda3a223/go.mod h1:STP8DvDyc/dI5b8T5hshtkjS+E42TnysNCUPdjciGhY=
golang.org/x/sys v0.0.0-20190312061237-fead79001313/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20190412213103-97732733099d/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20190502145724-3ef323f4f1fd/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
@@ -375,15 +295,11 @@ golang.org/x/sys v0.0.0-20201201145000-ef89a241ccb3/go.mod h1:h1NjWce9XRLGQEsW7w
golang.org/x/sys v0.0.0-20210104204734-6f8348627aad/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20210119212857-b64e53b001e4/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20210225134936-a50acf3fe073/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20210330210617-4fbd30eecc44/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20210423185535-09eb48e85fd7/go.mod h1:h1NjWce9XRLGQEsW7wpKNCjG9DtNlClVuFLEZdDNbEs=
golang.org/x/sys v0.0.0-20210615035016-665e8c7367d1/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.0.0-20211013075003-97ac67df715c/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/sys v0.1.0 h1:kunALQeHf1/185U1i0GOB/fy1IPRDDpuoOOqRReG57U=
golang.org/x/sys v0.1.0/go.mod h1:oPkhp1MJrh7nUepCBck5+mAzfO9JrbApNNgaTdGDITg=
golang.org/x/term v0.0.0-20201126162022-7de9c90e9dd1/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210220032956-6a3ed077a48d/go.mod h1:bj7SfCRtBDWHUb9snDiAeCFNEtKQo2Wmx5Cou7ajbmo=
golang.org/x/term v0.0.0-20210615171337-6886f2dfbf5b/go.mod h1:jbD1KX2456YbFQfuXm/mYQcufACuNUgVhRMnK/tPxf8=
golang.org/x/term v0.0.0-20210927222741-03fcf44c2211 h1:JGgROgKl9N8DuW20oFS5gxc+lE67/N3FcwmBPMe7ArY=
golang.org/x/term v0.0.0-20210927222741-03fcf44c2211/go.mod h1:jbD1KX2456YbFQfuXm/mYQcufACuNUgVhRMnK/tPxf8=
golang.org/x/text v0.0.0-20170915032832-14c0d48ead0c/go.mod h1:NqM8EUOU14njkJ3fqMW+pc6Ldnwhi/IjpwHt7yyuwOQ=
@@ -542,18 +458,9 @@ google.golang.org/protobuf v1.28.1 h1:d0NfwRgPtno5B1Wa6L2DAG+KivqkdutMf1UhdNx175
google.golang.org/protobuf v1.28.1/go.mod h1:HV8QOd/L58Z+nl8r43ehVNZIU/HEI6OcFqwMG9pJV4I=
gopkg.in/check.v1 v0.0.0-20161208181325-20d25e280405/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20180628173108-788fd7840127/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15 h1:YR8cESwS4TdDjEe65xsg0ogRM/Nc3DYOhEAlW+xobZo=
gopkg.in/check.v1 v1.0.0-20190902080502-41f04d3bba15/go.mod h1:Co6ibVJAznAaIkqp8huTwlJQCZ016jof/cbN4VW5Yz0=
gopkg.in/errgo.v2 v2.1.0/go.mod h1:hNsd1EY+bozCKY1Ytp96fpM3vjJbqLJn88ws8XvfDNI=
gopkg.in/fsnotify.v1 v1.4.7/go.mod h1:Tz8NjZHkW78fSQdbUxIjBTcgA1z1m8ZHf0WmKUhAMys=
gopkg.in/tomb.v1 v1.0.0-20141024135613-dd632973f1e7/go.mod h1:dt/ZhP58zS4L8KSrWDmTeBkI65Dw0HsyUHuEVlX15mw=
gopkg.in/yaml.v2 v2.2.1/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI=
gopkg.in/yaml.v2 v2.2.2/go.mod h1:hI93XBmqTisBFMUTm0b8Fm+jr3Dg1NNxqwp+5A1VGuI=
gopkg.in/yaml.v3 v3.0.0-20200313102051-9f266ea9e77c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
gopkg.in/yaml.v3 v3.0.0-20200605160147-a5ece683394c/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
gopkg.in/yaml.v3 v3.0.0-20210107192922-496545a6307b/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
honnef.co/go/tools v0.0.0-20190102054323-c2f93a96b099/go.mod h1:rf3lG4BRIbNafJWhAfAdb/ePZxsR/4RtNHQocxwk9r4=
honnef.co/go/tools v0.0.0-20190106161140-3f1c8253044a/go.mod h1:rf3lG4BRIbNafJWhAfAdb/ePZxsR/4RtNHQocxwk9r4=
honnef.co/go/tools v0.0.0-20190418001031-e561f6794a2a/go.mod h1:rf3lG4BRIbNafJWhAfAdb/ePZxsR/4RtNHQocxwk9r4=
+5 -6
View File
@@ -1,15 +1,14 @@
// Package bork defines the sentinel errors used by the manifest
// reader and writer.
package bork
import (
"errors"
"fmt"
)
var (
ErrMissingMagic = errors.New("missing magic bytes in file")
// ErrMissingMagic indicates the input lacks the manifest magic bytes.
ErrMissingMagic = errors.New("missing magic bytes in file")
// ErrFileTruncated indicates the input ended before the expected length.
ErrFileTruncated = errors.New("file/stream is truncated abnormally")
)
func Newf(format string, args ...interface{}) error {
return fmt.Errorf(format, args...)
}
+5 -2
View File
@@ -1,11 +1,14 @@
package bork
package bork_test
import (
"testing"
"github.com/stretchr/testify/assert"
"sneak.berlin/go/mfer/internal/bork"
)
func TestBuild(t *testing.T) {
assert.NotNil(t, ErrMissingMagic)
t.Parallel()
assert.Error(t, bork.ErrMissingMagic)
}
+302 -125
View File
@@ -1,183 +1,360 @@
// Package cli implements the mfer command-line interface.
package cli
import (
"context"
"encoding/hex"
"errors"
"fmt"
"io"
"math"
"path/filepath"
"strconv"
"strings"
"sync"
"time"
"github.com/dustin/go-humanize"
"github.com/spf13/afero"
"github.com/urfave/cli/v2"
"github.com/urfave/cli/v3"
"sneak.berlin/go/mfer/internal/log"
"sneak.berlin/go/mfer/mfer"
)
// findManifest looks for a manifest file in the given directory.
// It checks for index.mf and .index.mf, returning the first one found.
func findManifest(fs afero.Fs, dir string) (string, error) {
candidates := []string{"index.mf", ".index.mf"}
for _, name := range candidates {
path := filepath.Join(dir, name)
exists, err := afero.Exists(fs, path)
if err != nil {
return "", err
}
if exists {
return path, nil
}
// fingerprintHexLen is the length of a full GPG key fingerprint in hex
// characters.
const fingerprintHexLen = 40
var (
// errNoManifestFound indicates no manifest file was found in the
// searched directory.
errNoManifestFound = errors.New("no manifest found")
// errInvalidFingerprint indicates a malformed --require-signature
// fingerprint argument. The length is spliced in from
// fingerprintHexLen so the two cannot drift apart.
errInvalidFingerprint = errors.New(
"invalid fingerprint: must be exactly " +
strconv.Itoa(fingerprintHexLen) + " hex characters")
// errManifestNotSigned indicates a signature was required but the
// manifest is unsigned. It is wrapped mid-sentence so that the
// rendered message stays exactly as mfer has always printed it.
errManifestNotSigned = errors.New("manifest is not signed")
// errSignerMismatch indicates the embedded signing key fingerprint
// does not match the required signer. Its text is the mid-sentence
// fragment of the rendered message, which users grep for in CI and
// which must therefore not change; match it with errors.Is rather
// than by reading it.
errSignerMismatch = errors.New("does not match required")
)
// safeUint64 converts a non-negative int64 to uint64, clamping negative
// values to zero.
func safeUint64(n int64) uint64 {
if n < 0 {
return 0
}
return "", fmt.Errorf("no manifest found in %s (looked for index.mf and .index.mf)", dir)
return uint64(n)
}
func (mfa *CLIApp) checkManifestOperation(ctx *cli.Context) error {
log.Debug("checkManifestOperation()")
var manifestPath string
var err error
if ctx.Args().Len() > 0 {
arg := ctx.Args().Get(0)
// Check if arg is a directory or a file
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
// It's a directory, look for manifest inside
manifestPath, err = findManifest(mfa.Fs, arg)
if err != nil {
return err
}
} else {
// Treat as a file path
manifestPath = arg
}
} else {
// No argument, look in current directory
manifestPath, err = findManifest(mfa.Fs, ".")
if err != nil {
return err
}
// safeRateUint64 converts a bytes-per-second rate to uint64 for display.
//
// A rate is computed as bytes/elapsed, so it is +Inf when the elapsed
// time rounds to zero and NaN when zero bytes were processed in zero
// time. Neither has a defined conversion to uint64, and on amd64 +Inf
// converts to a number that renders as "8.0 EiB/s"; both display as zero
// instead.
func safeRateUint64(rate float64) uint64 {
if math.IsNaN(rate) || math.IsInf(rate, 0) || rate <= 0 {
return 0
}
basePath := ctx.String("base")
showProgress := ctx.Bool("progress")
if rate >= math.MaxUint64 {
return math.MaxUint64
}
log.Infof("checking manifest %s with base %s", manifestPath, basePath)
return uint64(rate)
}
// Create checker
chk, err := mfer.NewChecker(manifestPath, basePath, mfa.Fs)
// findManifest returns the path of the manifest with the default name in
// dir, or an error if there is none.
func findManifest(fs afero.Fs, dir string) (string, error) {
path := filepath.Join(dir, defaultManifestName)
exists, err := afero.Exists(fs, path)
if err != nil {
return fmt.Errorf("failed to load manifest: %w", err)
return "", err
}
// Check signature requirement
requiredSigner := ctx.String("require-signature")
if requiredSigner != "" {
// Validate fingerprint format: must be exactly 40 hex characters
if len(requiredSigner) != 40 {
return fmt.Errorf("invalid fingerprint: must be exactly 40 hex characters, got %d", len(requiredSigner))
}
if _, err := hex.DecodeString(requiredSigner); err != nil {
return fmt.Errorf("invalid fingerprint: must be valid hex: %w", err)
}
if !chk.IsSigned() {
return fmt.Errorf("manifest is not signed, but signature from %s is required", requiredSigner)
}
// Extract fingerprint from the embedded public key (not from the signer field)
// This validates the key is importable and gets its actual fingerprint
embeddedFP, err := chk.ExtractEmbeddedSigningKeyFP()
if err != nil {
return fmt.Errorf("failed to extract fingerprint from embedded signing key: %w", err)
}
// Compare fingerprints - must be exact match (case-insensitive)
if !strings.EqualFold(embeddedFP, requiredSigner) {
return fmt.Errorf("embedded signing key fingerprint %s does not match required %s", embeddedFP, requiredSigner)
}
log.Infof("manifest signature verified (signer: %s)", embeddedFP)
if !exists {
return "", fmt.Errorf("%w in %s (looked for %s)",
errNoManifestFound, dir, defaultManifestName)
}
log.Infof("manifest contains %d files, %s", chk.FileCount(), humanize.IBytes(uint64(chk.TotalBytes())))
return path, nil
}
// fetchManifestToTemp downloads a manifest URL to a temporary file and
// returns the temp file path. The caller is responsible for removing it.
func (mfa *CLIApp) fetchManifestToTemp(
ctx context.Context, url string,
) (string, error) {
rc, fetchErr := mfa.openManifestReader(ctx, url)
if fetchErr != nil {
return "", fetchErr
}
tmpFile, tmpErr := afero.TempFile(mfa.Fs, "", "mfer-manifest-*.mf")
if tmpErr != nil {
_ = rc.Close()
return "", fmt.Errorf("failed to create temp file: %w", tmpErr)
}
tmpPath := tmpFile.Name()
_, cpErr := io.Copy(tmpFile, rc)
_ = rc.Close()
_ = tmpFile.Close()
if cpErr != nil {
_ = mfa.Fs.Remove(tmpPath)
return "", fmt.Errorf("failed to download manifest: %w", cpErr)
}
return tmpPath, nil
}
// verifyRequiredSigner enforces the --require-signature fingerprint
// against the manifest's embedded signing key.
func verifyRequiredSigner(
ctx context.Context, chk *mfer.Checker, requiredSigner string,
) error {
// Validate fingerprint format: must be exactly 40 hex characters
if len(requiredSigner) != fingerprintHexLen {
return fmt.Errorf("%w, got %d", errInvalidFingerprint, len(requiredSigner))
}
_, err := hex.DecodeString(requiredSigner)
if err != nil {
return fmt.Errorf("invalid fingerprint: must be valid hex: %w", err)
}
if !chk.IsSigned() {
return fmt.Errorf("%w, but signature from %s is required",
errManifestNotSigned, requiredSigner)
}
// Extract fingerprint from the embedded public key (not from the
// signer field). This validates the key is importable and gets its
// actual fingerprint.
embeddedFP, err := chk.ExtractEmbeddedSigningKeyFP(ctx)
if err != nil {
return fmt.Errorf(
"failed to extract fingerprint from embedded signing key: %w", err)
}
// Compare fingerprints - must be exact match (case-insensitive)
if !strings.EqualFold(embeddedFP, requiredSigner) {
return fmt.Errorf("embedded signing key fingerprint %s %w %s",
embeddedFP, errSignerMismatch, requiredSigner)
}
log.Infof("manifest signature verified (signer: %s)", embeddedFP)
return nil
}
// reportCheckProgress renders progress updates until the channel closes.
func reportCheckProgress(progress <-chan mfer.CheckStatus, wg *sync.WaitGroup) {
defer wg.Done()
for status := range progress {
if status.ETA > 0 {
log.Progressf("Checking: %d/%d files, %s/s, ETA %s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)),
status.ETA.Round(time.Second),
status.Failures)
} else {
log.Progressf("Checking: %d/%d files, %s/s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)),
status.Failures)
}
}
log.ProgressDone()
}
// countCheckFailures consumes check results, counting and logging
// failures, then closes done.
func countCheckFailures(
results <-chan mfer.Result, failures *int64, done chan<- struct{},
) {
for result := range results {
if result.Status != mfer.StatusOK {
*failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
} else {
log.Verbosef("%s: %s", result.Status, result.Path)
}
}
close(done)
}
// findExtraFiles reports files present on disk but absent from the
// manifest, and anything the search cannot read: each is a failure under
// --no-extra-files, otherwise a warning.
func findExtraFiles(
ctx context.Context, cmd *cli.Command, chk *mfer.Checker, failures *int64,
) error {
extraResults := make(chan mfer.Result, 1)
extraDone := make(chan struct{})
go func() {
for result := range extraResults {
if cmd.Bool("no-extra-files") {
*failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
} else {
log.Warnf("%s: %s (%s)", result.Status, result.Path, result.Message)
}
}
close(extraDone)
}()
err := chk.FindExtraFiles(ctx, extraResults)
if err != nil {
return fmt.Errorf("failed to check for extra files: %w", err)
}
<-extraDone
return nil
}
// runCheck runs the manifest check with progress and result reporting
// and returns the number of failures.
func runCheck(
ctx context.Context, cmd *cli.Command, chk *mfer.Checker, showProgress bool,
) (int64, error) {
// Set up results channel
results := make(chan mfer.Result, 1)
// Set up progress channel
var progress chan mfer.CheckStatus
var (
progress chan mfer.CheckStatus
progressWg sync.WaitGroup
)
if showProgress {
progress = make(chan mfer.CheckStatus, 1)
go func() {
for status := range progress {
if status.ETA > 0 {
log.Progressf("Checking: %d/%d files, %s/s, ETA %s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)),
status.ETA.Round(time.Second),
status.Failures)
} else {
log.Progressf("Checking: %d/%d files, %s/s, %d failures",
status.CheckedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)),
status.Failures)
}
}
log.ProgressDone()
}()
progressWg.Add(1)
go reportCheckProgress(progress, &progressWg)
}
// Process results in a goroutine
var failures int64
done := make(chan struct{})
go func() {
for result := range results {
if result.Status != mfer.StatusOK {
failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
} else {
log.Verbosef("%s: %s", result.Status, result.Path)
}
}
close(done)
}()
go countCheckFailures(results, &failures, done)
// Run check
err = chk.Check(ctx.Context, results, progress)
err := chk.Check(ctx, results, progress)
progressWg.Wait()
if err != nil {
return fmt.Errorf("check failed: %w", err)
return 0, fmt.Errorf("check failed: %w", err)
}
// Wait for results processing to complete
<-done
// Check for extra files if requested
if ctx.Bool("no-extra-files") {
extraResults := make(chan mfer.Result, 1)
extraDone := make(chan struct{})
go func() {
for result := range extraResults {
failures++
log.Infof("%s: %s (%s)", result.Status, result.Path, result.Message)
}
close(extraDone)
}()
err = findExtraFiles(ctx, cmd, chk, &failures)
if err != nil {
return 0, err
}
err = chk.FindExtraFiles(ctx.Context, extraResults)
if err != nil {
return fmt.Errorf("failed to check for extra files: %w", err)
return failures, nil
}
func (mfa *CLIApp) checkManifestOperation(
ctx context.Context, cmd *cli.Command,
) error {
log.Debug("checkManifestOperation()")
manifestPath, err := mfa.resolveManifestArg(cmd)
if err != nil {
return fmt.Errorf("check: %w", err)
}
// URL manifests need to be downloaded to a temp file for the checker
if isHTTPURL(manifestPath) {
tmpPath, tmpErr := mfa.fetchManifestToTemp(ctx, manifestPath)
if tmpErr != nil {
return fmt.Errorf("check: %w", tmpErr)
}
<-extraDone
defer func() { _ = mfa.Fs.Remove(tmpPath) }()
manifestPath = tmpPath
}
basePath := cmd.String("base")
showProgress := cmd.Bool("progress")
log.Infof("checking manifest %s with base %s", manifestPath, basePath)
// Create checker
//nolint:contextcheck // mfer loads a manifest without a context
chk, err := mfer.NewChecker(&mfer.CheckerOptions{
ManifestPath: manifestPath,
BasePath: basePath,
Fs: mfa.Fs,
})
if err != nil {
return fmt.Errorf("failed to load manifest: %w", err)
}
// Check signature requirement
requiredSigner := cmd.String(flagRequireSignature)
if requiredSigner != "" {
err = verifyRequiredSigner(ctx, chk, requiredSigner)
if err != nil {
return err
}
}
log.Infof("manifest contains %d files, %s", chk.FileCount(),
humanize.IBytes(safeUint64(int64(chk.TotalBytes()))))
failures, err := runCheck(ctx, cmd, chk, showProgress)
if err != nil {
return err
}
elapsed := time.Since(mfa.startupTime).Seconds()
rate := float64(chk.TotalBytes()) / elapsed
if failures == 0 {
log.Infof("checked %d files (%s) in %.1fs (%s/s): all OK", chk.FileCount(), humanize.IBytes(uint64(chk.TotalBytes())), elapsed, humanize.IBytes(uint64(rate)))
log.Infof("checked %d files (%s) in %.1fs (%s/s): all OK",
chk.FileCount(), humanize.IBytes(safeUint64(int64(chk.TotalBytes()))),
elapsed, humanize.IBytes(safeRateUint64(rate)))
} else {
log.Infof("checked %d files (%s) in %.1fs (%s/s): %d failed", chk.FileCount(), humanize.IBytes(uint64(chk.TotalBytes())), elapsed, humanize.IBytes(uint64(rate)), failures)
log.Infof("checked %d files (%s) in %.1fs (%s/s): %d failed",
chk.FileCount(), humanize.IBytes(safeUint64(int64(chk.TotalBytes()))),
elapsed, humanize.IBytes(safeRateUint64(rate)), failures)
}
if failures > 0 {
+11 -7
View File
@@ -7,15 +7,18 @@ import (
"github.com/spf13/afero"
)
// NO_COLOR disables colored output when set. Automatically true if the
// NoColor disables colored output when set. Automatically true if the
// NO_COLOR environment variable is present (per https://no-color.org/).
var NO_COLOR bool
//
//nolint:gochecknoglobals // process-wide setting derived from the environment
var NoColor = noColorEnvSet()
func init() {
NO_COLOR = false
if _, exists := os.LookupEnv("NO_COLOR"); exists {
NO_COLOR = true
}
// noColorEnvSet reports whether the NO_COLOR environment variable is
// present.
func noColorEnvSet() bool {
_, exists := os.LookupEnv("NO_COLOR")
return exists
}
// RunOptions contains all configuration for running the CLI application.
@@ -64,5 +67,6 @@ func RunWithOptions(opts *RunOptions) int {
}
m.run(opts.Args)
return m.exitCode
}
+901 -215
View File
File diff suppressed because it is too large Load Diff
+383
View File
@@ -0,0 +1,383 @@
//nolint:testpackage // white-box tests exercise unexported internals
package cli
import (
"bytes"
"context"
"net/http"
"net/http/httptest"
"os"
"os/exec"
"path/filepath"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
urfcli "github.com/urfave/cli/v3"
"sneak.berlin/go/mfer/mfer"
)
// These tests pin the exact rendered text of the CLI's user-visible error
// messages. The messages are grepped for in CI pipelines and quoted in bug
// reports, so a reword is a deliberate change, never a refactoring side
// effect.
//
// Every case drives the real function that emits the message and asserts on
// what it returns. No production format string is restated here: a test that
// only re-rendered a copied format string would keep passing after the real
// message changed, which is exactly the regression these tests exist to
// catch.
// Full 40-hex fingerprints used where a message embeds one.
const (
msgFpA = "AAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAAA"
msgFpB = "BBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBBB"
)
// runLocked runs fn while holding runMu, so operations that write to the
// process-global logger do not race the other CLI runs.
func runLocked(fn func() error) error {
runMu.Lock()
defer runMu.Unlock()
return fn()
}
// unsignedChecker builds a Checker over a freshly scanned, unsigned manifest.
func unsignedChecker(t *testing.T) *mfer.Checker {
t.Helper()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/d", 0o755))
require.NoError(t, afero.WriteFile(fs, "/d/f.txt", []byte("hi"), 0o644))
s := mfer.NewScannerWithOptions(&mfer.ScannerOptions{Fs: fs})
require.NoError(t, s.EnumeratePath("/d", nil))
var buf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &buf, nil))
require.NoError(t, afero.WriteFile(fs, "/d/index.mf", buf.Bytes(), 0o644))
chk, err := mfer.NewChecker(&mfer.CheckerOptions{
ManifestPath: "/d/index.mf",
BasePath: "/d",
Fs: fs,
})
require.NoError(t, err)
require.False(t, chk.IsSigned())
return chk
}
func TestNoManifestFoundMessage(t *testing.T) {
t.Parallel()
_, err := findManifest(afero.NewMemMapFs(), "/tmp/x")
require.ErrorIs(t, err, errNoManifestFound)
assert.EqualError(t, err,
"no manifest found in /tmp/x (looked for index.mf)")
}
func TestVerifyRequiredSignerMessages(t *testing.T) {
t.Parallel()
t.Run("invalid fingerprint length", func(t *testing.T) {
t.Parallel()
err := verifyRequiredSigner(context.Background(),
unsignedChecker(t), "12345678")
require.ErrorIs(t, err, errInvalidFingerprint)
assert.EqualError(t, err,
"invalid fingerprint: must be exactly 40 hex characters, got 8")
})
t.Run("manifest not signed", func(t *testing.T) {
t.Parallel()
err := verifyRequiredSigner(context.Background(),
unsignedChecker(t), msgFpA)
require.ErrorIs(t, err, errManifestNotSigned)
assert.EqualError(t, err,
"manifest is not signed, but signature from "+msgFpA+" is required")
})
}
// TestSignerMismatchMessage drives verifyRequiredSigner against a real signed
// manifest. The embedded fingerprint is whatever the generated key produced,
// so it is read back from the checker and substituted into the expected
// string; the required signer is a fixed value that cannot match it. Requires
// gpg and is skipped where it is absent, as the other signing tests are.
//
//nolint:paralleltest // signedManifest calls t.Setenv, which bars t.Parallel
func TestSignerMismatchMessage(t *testing.T) {
chk := signedChecker(t,
signedManifest(t, map[string][]byte{"f.txt": []byte("signed file")}))
embeddedFP, err := chk.ExtractEmbeddedSigningKeyFP(context.Background())
require.NoError(t, err)
err = verifyRequiredSigner(context.Background(), chk, msgFpB)
require.ErrorIs(t, err, errSignerMismatch)
assert.EqualError(t, err,
"embedded signing key fingerprint "+embeddedFP+
" does not match required "+msgFpB)
}
// signedManifest returns a manifest of files signed by a throwaway GPG key
// generated in a temporary GNUPGHOME, which it leaves set for the rest of
// the test.
func signedManifest(t *testing.T, files map[string][]byte) []byte {
t.Helper()
_, err := exec.LookPath("gpg")
if err != nil {
t.Skip("gpg not installed, skipping signing test")
}
gpgHome := t.TempDir()
params := "%no-protection\n" +
"Key-Type: RSA\nKey-Length: 2048\n" +
"Name-Real: MFER Test Key\nName-Email: test@mfer.test\n" +
"Expire-Date: 0\n%commit\n"
paramsFile := filepath.Join(gpgHome, "key-params")
require.NoError(t, os.WriteFile(paramsFile, []byte(params), 0o600))
//nolint:gosec // paramsFile is a test-controlled path inside t.TempDir()
cmd := exec.CommandContext(context.Background(), "gpg",
"--batch", "--gen-key", paramsFile)
cmd.Env = append(os.Environ(), "GNUPGHOME="+gpgHome)
out, err := cmd.CombinedOutput()
if err != nil {
t.Skipf("failed to generate test GPG key: %v: %s", err, out)
}
t.Setenv("GNUPGHOME", gpgHome)
b := mfer.NewBuilder()
b.SetSigningOptions(&mfer.SigningOptions{KeyID: mfer.GPGKeyID("test@mfer.test")})
for path, content := range files {
_, err = b.AddFile(mfer.RelFilePath(path), mfer.FileSize(len(content)),
mfer.ModTime{}, bytes.NewReader(content), nil)
require.NoError(t, err)
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
return buf.Bytes()
}
// signedChecker builds a Checker over manifest, a signed manifest.
func signedChecker(t *testing.T, manifest []byte) *mfer.Checker {
t.Helper()
fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/index.mf", manifest, 0o644))
chk, err := mfer.NewChecker(&mfer.CheckerOptions{
ManifestPath: "/index.mf",
BasePath: "/",
Fs: fs,
})
require.NoError(t, err)
require.True(t, chk.IsSigned())
return chk
}
func TestPathDoesNotExistMessage(t *testing.T) {
t.Parallel()
mfa := &CLIApp{Fs: afero.NewMemMapFs()}
cmd := &urfcli.Command{
Name: cmdGenerate,
Action: func(_ context.Context, c *urfcli.Command) error {
_, err := mfa.collectInputPaths(c.Args())
return err
},
}
err := cmd.Run(context.Background(), []string{cmdGenerate, "nope"})
require.ErrorIs(t, err, errPathNotExist)
assert.EqualError(t, err, "path does not exist: nope")
}
func TestOutputFileExistsMessage(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/d", 0o755))
require.NoError(t, afero.WriteFile(fs, "/d/f.txt", []byte("hi"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/out.mf", []byte("old"), 0o644))
mfa := &CLIApp{Fs: fs}
cmd := &urfcli.Command{
Name: cmdGenerate,
Flags: []urfcli.Flag{
&urfcli.StringFlag{Name: "output"},
&urfcli.BoolFlag{Name: "force"},
},
Action: mfa.generateManifestOperation,
}
// generateManifestOperation writes to the process-global logger during
// enumeration, so serialize with the other CLI runs.
err := runLocked(func() error {
return cmd.Run(context.Background(),
[]string{cmdGenerate, "--output", "/out.mf", "/d"})
})
require.ErrorIs(t, err, errOutputExists)
assert.EqualError(t, err,
"output file /out.mf already exists (use --force to overwrite)")
}
// TestUnknownCommandMessage drives the root command's action. run only logs
// the error that action returns, so the test lets run build the app with no
// command given and then runs that same app on an unknown command to get the
// error itself.
func TestUnknownCommandMessage(t *testing.T) {
t.Parallel()
mfa := &CLIApp{
appname: testApp,
Stdout: &bytes.Buffer{},
Stderr: &bytes.Buffer{},
Fs: afero.NewMemMapFs(),
}
// run points the process-global logger at this app's output, so
// serialize with the other CLI runs.
err := runLocked(func() error {
mfa.run([]string{testApp})
return mfa.app.Run(context.Background(), []string{testApp, "bogus"})
})
require.ErrorIs(t, err, errUnknownCommand)
assert.EqualError(t, err, `unknown command "bogus"`)
}
func TestManifestLoaderHTTPStatusMessage(t *testing.T) {
t.Parallel()
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusNotFound)
}))
defer server.Close()
mfa := &CLIApp{Fs: afero.NewMemMapFs()}
_, err := mfa.openManifestReader(context.Background(), server.URL+"/foo.mf")
require.ErrorIs(t, err, errHTTPStatus)
assert.EqualError(t, err,
"failed to fetch "+server.URL+"/foo.mf: HTTP 404")
}
func TestFetchManifestHTTPStatusMessage(t *testing.T) {
t.Parallel()
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusNotFound)
}))
defer server.Close()
mfa := &CLIApp{Fs: afero.NewMemMapFs()}
cmd := mfa.fetchCommand()
cmd.Action = mfa.fetchManifestOperation
// fetchManifestOperation logs to the process-global logger.
err := runLocked(func() error {
return cmd.Run(context.Background(), []string{cmdFetch, server.URL})
})
require.ErrorIs(t, err, errHTTPStatus)
assert.EqualError(t, err, "failed to fetch manifest: HTTP 404")
}
func TestFetchFileHTTPStatusMessage(t *testing.T) {
t.Parallel()
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.WriteHeader(http.StatusInternalServerError)
}))
defer server.Close()
// downloadFile logs each retry of the 500 to the process-global logger.
err := runLocked(func() error {
return downloadFile(context.Background(), testClient(), server.URL+"/x", ".", "x",
&mfer.MFFilePath{}, nil)
})
require.ErrorIs(t, err, errHTTPStatus)
assert.EqualError(t, err, "HTTP 500")
}
func TestURLRequiredMessage(t *testing.T) {
t.Parallel()
mfa := &CLIApp{Fs: afero.NewMemMapFs()}
cmd := &urfcli.Command{Name: cmdFetch, Action: mfa.fetchManifestOperation}
// fetchManifestOperation logs to the process-global logger.
err := runLocked(func() error {
return cmd.Run(context.Background(), []string{cmdFetch})
})
require.ErrorIs(t, err, errURLRequired)
assert.EqualError(t, err, "URL argument required")
}
func TestSanitizePathMessages(t *testing.T) {
t.Parallel()
t.Run("empty", func(t *testing.T) {
t.Parallel()
_, err := sanitizePath("")
require.ErrorIs(t, err, errEmptyPath)
assert.EqualError(t, err, "empty path")
})
t.Run("absolute", func(t *testing.T) {
t.Parallel()
_, err := sanitizePath("/etc/passwd")
require.ErrorIs(t, err, errAbsolutePath)
assert.EqualError(t, err, "absolute path not allowed: /etc/passwd")
})
t.Run("traversal", func(t *testing.T) {
t.Parallel()
_, err := sanitizePath("../x")
require.ErrorIs(t, err, errPathTraversal)
assert.EqualError(t, err, "path traversal not allowed: ../x")
})
}
func TestSizeMismatchMessage(t *testing.T) {
t.Parallel()
// finishDownload returns the size-mismatch error before it touches the
// paths, digest, or entry, so those can be zero here.
err := finishDownload("", "", "", 9, 10, nil, nil, nil, nil)
require.ErrorIs(t, err, errSizeMismatch)
assert.EqualError(t, err, "size mismatch: expected 10 bytes, got 9")
}
func TestHashMismatchMessage(t *testing.T) {
t.Parallel()
// A 32-byte digest that matches none of the (empty) manifest hashes.
err := verifyDownloadedHash(make([]byte, 32), &mfer.MFFilePath{})
require.ErrorIs(t, err, errHashMismatch)
require.NotErrorIs(t, err, errSizeMismatch)
assert.EqualError(t, err, "hash mismatch")
}
+81
View File
@@ -0,0 +1,81 @@
package cli
import (
"context"
"encoding/hex"
"encoding/json"
"fmt"
"time"
"github.com/urfave/cli/v3"
"sneak.berlin/go/mfer/mfer"
)
// ExportEntry represents a single file entry in the exported JSON output.
type ExportEntry struct {
Path string `json:"path"`
Size int64 `json:"size"`
Hashes []string `json:"hashes"`
Mtime *string `json:"mtime,omitempty"`
Ctime *string `json:"ctime,omitempty"`
}
func (mfa *CLIApp) exportManifestOperation(
ctx context.Context, cmd *cli.Command,
) error {
pathOrURL, err := mfa.resolveManifestArg(cmd)
if err != nil {
return fmt.Errorf("export: %w", err)
}
rc, err := mfa.openManifestReader(ctx, pathOrURL)
if err != nil {
return fmt.Errorf("export: %w", err)
}
defer func() { _ = rc.Close() }()
//nolint:contextcheck // mfer loads a manifest without a context
manifest, err := mfer.NewManifestFromReader(rc)
if err != nil {
return fmt.Errorf("export: failed to parse manifest: %w", err)
}
files := manifest.Files()
entries := make([]ExportEntry, 0, len(files))
for _, f := range files {
entry := ExportEntry{
Path: f.GetPath(),
Size: f.GetSize(),
Hashes: make([]string, 0, len(f.GetHashes())),
}
for _, h := range f.GetHashes() {
entry.Hashes = append(entry.Hashes, hex.EncodeToString(h.GetMultiHash()))
}
if mtime, ok := entryMtime(f); ok {
t := mtime.UTC().Format(time.RFC3339Nano)
entry.Mtime = &t
}
if f.GetCtime() != nil {
t := time.Unix(f.GetCtime().GetSeconds(), int64(f.GetCtime().GetNanos())).
UTC().Format(time.RFC3339Nano)
entry.Ctime = &t
}
entries = append(entries, entry)
}
enc := json.NewEncoder(mfa.Stdout)
enc.SetIndent("", " ")
err = enc.Encode(entries)
if err != nil {
return fmt.Errorf("export: failed to encode JSON: %w", err)
}
return nil
}
+156
View File
@@ -0,0 +1,156 @@
package cli
import (
"bytes"
"context"
"encoding/json"
"net/http"
"net/http/httptest"
"testing"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"sneak.berlin/go/mfer/mfer"
)
const testCmdExport = "export"
// buildTestManifest creates a manifest from in-memory files and returns its bytes.
func buildTestManifest(t *testing.T, files map[string][]byte) []byte {
t.Helper()
sourceFs := afero.NewMemMapFs()
for path, content := range files {
require.NoError(t, sourceFs.MkdirAll("/", 0o755))
require.NoError(t, afero.WriteFile(sourceFs, "/"+path, content, 0o644))
}
opts := &mfer.ScannerOptions{Fs: sourceFs}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumerateFS(sourceFs, "/", nil))
var buf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &buf, nil))
return buf.Bytes()
}
func TestExportManifestOperation(t *testing.T) {
t.Parallel()
testFiles := map[string][]byte{
"hello.txt": []byte("Hello, World!"),
"sub/file.txt": []byte("nested content"),
}
manifestData := buildTestManifest(t, testFiles)
// Write manifest to memfs
fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/test.mf", manifestData, 0o644))
var stdout, stderr bytes.Buffer
exitCode := runCLI(&RunOptions{
Appname: testApp,
Args: []string{testApp, testCmdExport, "/test.mf"},
Stdin: &bytes.Buffer{},
Stdout: &stdout,
Stderr: &stderr,
Fs: fs,
})
require.Equal(t, 0, exitCode, "stderr: %s", stderr.String())
var entries []ExportEntry
require.NoError(t, json.Unmarshal(stdout.Bytes(), &entries))
assert.Len(t, entries, 2)
// Verify entries have expected fields
pathSet := make(map[string]bool)
for _, e := range entries {
pathSet[e.Path] = true
assert.NotEmpty(t, e.Hashes, "entry %s should have hashes", e.Path)
assert.Positive(t, e.Size, "entry %s should have positive size", e.Path)
}
assert.True(t, pathSet["hello.txt"])
assert.True(t, pathSet["sub/file.txt"])
}
func TestExportFromHTTPURL(t *testing.T) {
t.Parallel()
testFiles := map[string][]byte{
"a.txt": []byte("aaa"),
}
manifestData := buildTestManifest(t, testFiles)
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
w.Header().Set("Content-Type", "application/octet-stream")
_, _ = w.Write(manifestData)
}))
defer server.Close()
var stdout, stderr bytes.Buffer
exitCode := runCLI(&RunOptions{
Appname: testApp,
Args: []string{testApp, testCmdExport, server.URL + "/index.mf"},
Stdin: &bytes.Buffer{},
Stdout: &stdout,
Stderr: &stderr,
Fs: afero.NewMemMapFs(),
})
require.Equal(t, 0, exitCode, "stderr: %s", stderr.String())
var entries []ExportEntry
require.NoError(t, json.Unmarshal(stdout.Bytes(), &entries))
assert.Len(t, entries, 1)
assert.Equal(t, "a.txt", entries[0].Path)
}
func TestListFromHTTPURL(t *testing.T) {
t.Parallel()
testFiles := map[string][]byte{
"one.txt": []byte("1"),
"two.txt": []byte("22"),
}
manifestData := buildTestManifest(t, testFiles)
server := httptest.NewServer(
http.HandlerFunc(func(w http.ResponseWriter, _ *http.Request) {
_, _ = w.Write(manifestData)
}))
defer server.Close()
var stdout, stderr bytes.Buffer
exitCode := runCLI(&RunOptions{
Appname: testApp,
Args: []string{testApp, "list", server.URL + "/index.mf"},
Stdin: &bytes.Buffer{},
Stdout: &stdout,
Stderr: &stderr,
Fs: afero.NewMemMapFs(),
})
require.Equal(t, 0, exitCode, "stderr: %s", stderr.String())
output := stdout.String()
assert.Contains(t, output, "one.txt")
assert.Contains(t, output, "two.txt")
}
func TestIsHTTPURL(t *testing.T) {
t.Parallel()
assert.True(t, isHTTPURL("http://example.com/manifest.mf"))
assert.True(t, isHTTPURL("https://example.com/manifest.mf"))
assert.False(t, isHTTPURL("/local/path.mf"))
assert.False(t, isHTTPURL("relative/path.mf"))
assert.False(t, isHTTPURL("ftp://example.com/file"))
}
+696 -154
View File
File diff suppressed because it is too large Load Diff
+1106 -140
View File
File diff suppressed because it is too large Load Diff
+505 -245
View File
@@ -1,21 +1,37 @@
package cli
import (
"context"
"crypto/sha256"
"errors"
"fmt"
"io"
"io/fs"
"os"
"path/filepath"
"time"
"github.com/dustin/go-humanize"
"github.com/multiformats/go-multihash"
"github.com/spf13/afero"
"github.com/urfave/cli/v2"
"github.com/urfave/cli/v3"
"sneak.berlin/go/mfer/internal/log"
"sneak.berlin/go/mfer/mfer"
)
const (
// hashBufSize is the read buffer size used when hashing files.
hashBufSize = 64 * 1024
// scanProgressInterval is how many scanned files pass between
// progress updates.
scanProgressInterval = 100
)
// errEntryMissingMtime indicates a manifest entry that carries no
// modification time where one is required to carry it forward unchanged.
var errEntryMissingMtime = errors.New("manifest entry has no mtime")
// FreshenStatus contains progress information for the freshen operation.
type FreshenStatus struct {
Phase string // "scan" or "hash"
@@ -36,42 +52,419 @@ type freshenEntry struct {
existing *mfer.MFFilePath // existing manifest entry if unchanged
}
func (mfa *CLIApp) freshenManifestOperation(ctx *cli.Context) error {
log.Debug("freshenManifestOperation()")
// freshenScanner walks the filesystem and compares it against the
// entries of an existing manifest.
type freshenScanner struct {
fs afero.Fs
absBase string
excluded []fs.FileInfo // files left out of the listing
includeDotfiles bool
followSymlinks bool
showProgress bool
existingByPath map[string]*mfer.MFFilePath
basePath := ctx.String("base")
showProgress := ctx.Bool("progress")
includeDotfiles := ctx.Bool("IncludeDotfiles")
followSymlinks := ctx.Bool("FollowSymLinks")
entries []*freshenEntry
scanCount int64
changed int64
added int64
unchanged int64
}
// Find manifest file
var manifestPath string
var err error
// resolveSymlink resolves a symlink to its target's FileInfo. The
// second return value is false when the entry should be skipped.
func (s *freshenScanner) resolveSymlink(path string) (fs.FileInfo, bool) {
if !s.followSymlinks {
return nil, false
}
if ctx.Args().Len() > 0 {
arg := ctx.Args().Get(0)
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
manifestPath, err = findManifest(mfa.Fs, arg)
if err != nil {
return err
}
} else {
manifestPath = arg
}
realPath, err := filepath.EvalSymlinks(path)
if err != nil {
return nil, false // Skip broken symlinks
}
realInfo, err := s.fs.Stat(realPath)
if err != nil || realInfo.IsDir() {
return nil, false
}
return realInfo, true
}
// recordEntry classifies a scanned file as changed, unchanged, or added
// relative to the existing manifest.
func (s *freshenScanner) recordEntry(relPath string, info fs.FileInfo) {
existing, inManifest := s.existingByPath[relPath]
if !inManifest {
s.added++
log.Verbosef("A %s", relPath)
s.entries = append(s.entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
return
}
// Check if changed (size or mtime). An entry with no recorded mtime
// cannot be compared, so it counts as changed and gets re-hashed;
// silently treating the absent mtime as the Unix epoch would classify
// every such entry as changed without saying why.
existingMtime, haveMtime := entryMtime(existing)
if !haveMtime {
log.Debugf("%s: manifest entry has no mtime, treating as changed",
relPath)
}
if !haveMtime || existing.GetSize() != info.Size() ||
!existingMtime.Equal(info.ModTime()) {
s.changed++
log.Verbosef("M %s", relPath)
s.entries = append(s.entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
} else {
manifestPath, err = findManifest(mfa.Fs, ".")
s.unchanged++
s.entries = append(s.entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: false,
existing: existing,
})
}
// Mark as seen
delete(s.existingByPath, relPath)
}
// walk is the afero.Walk callback for the scan phase.
func (s *freshenScanner) walk(path string, info fs.FileInfo, walkErr error) error {
if walkErr != nil {
return walkErr
}
// Get relative path
relPath, err := filepath.Rel(s.absBase, path)
if err != nil {
return fmt.Errorf(
"freshen: failed to compute relative path for %s: %w", path, err)
}
// Handle dotfiles
if !s.includeDotfiles && mfer.IsHiddenPath(filepath.ToSlash(relPath)) {
if info.IsDir() {
return filepath.SkipDir
}
return nil
}
// Skip directories
if info.IsDir() {
return nil
}
// Handle symlinks
if info.Mode()&fs.ModeSymlink != 0 {
realInfo, keep := s.resolveSymlink(path)
if !keep {
return nil
}
info = realInfo
}
for _, excluded := range s.excluded {
if os.SameFile(info, excluded) {
return nil
}
}
s.scanCount++
// Check against existing manifest
s.recordEntry(relPath, info)
// Report scan progress
if s.showProgress && s.scanCount%scanProgressInterval == 0 {
log.Progressf("Scanning: %d files found", s.scanCount)
}
return nil
}
// resolveFreshenManifestPath determines the manifest path from the CLI
// arguments, searching directories for a manifest where needed.
func (mfa *CLIApp) resolveFreshenManifestPath(cmd *cli.Command) (string, error) {
if cmd.Args().Len() == 0 {
return findManifest(mfa.Fs, ".")
}
arg := cmd.Args().Get(0)
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
return findManifest(mfa.Fs, arg)
}
return arg, nil
}
// freshenHasher hashes changed and added files and feeds all entries to
// a manifest builder.
type freshenHasher struct {
fs afero.Fs
absBase string
showProgress bool
totalHashBytes int64
filesToHash int64
startHash time.Time
builder *mfer.Builder
hashedFiles int64
hashedBytes int64
}
// reportProgress renders hashing progress for the current byte count.
func (h *freshenHasher) reportProgress(n int64) {
if !h.showProgress {
return
}
currentBytes := h.hashedBytes + n
elapsed := time.Since(h.startHash)
var (
rate float64
eta time.Duration
)
if elapsed > 0 && currentBytes > 0 {
rate = float64(currentBytes) / elapsed.Seconds()
remaining := h.totalHashBytes - currentBytes
if rate > 0 {
eta = time.Duration(float64(remaining)/rate) * time.Second
}
}
if eta > 0 {
log.Progressf("Hashing: %d/%d files, %s/s, ETA %s",
h.hashedFiles, h.filesToHash, humanize.IBytes(safeRateUint64(rate)),
eta.Round(time.Second))
} else {
log.Progressf("Hashing: %d/%d files, %s/s",
h.hashedFiles, h.filesToHash, humanize.IBytes(safeRateUint64(rate)))
}
}
// processEntry hashes the entry if needed and adds it to the builder.
func (h *freshenHasher) processEntry(e *freshenEntry) error {
if !e.needsHash {
// Use existing entry
err := addExistingToBuilder(h.builder, e.existing)
if err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
}
return nil
}
// Need to read and hash the file
absPath := filepath.Join(h.absBase, e.path)
f, err := h.fs.Open(absPath)
if err != nil {
return fmt.Errorf("failed to open %s: %w", e.path, err)
}
hash, bytesRead, err := hashFile(f, h.reportProgress)
_ = f.Close()
if err != nil {
return fmt.Errorf("failed to hash %s: %w", e.path, err)
}
h.hashedBytes += bytesRead
h.hashedFiles++
// Add to builder with computed hash
err = addFileToBuilder(h.builder, e.path, e.size, e.mtime, hash)
if err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
}
return nil
}
// writeFreshenedManifest writes the manifest atomically (write to a
// temp file, then rename over the target).
func writeFreshenedManifest(
ctx context.Context, afs afero.Fs, builder *mfer.Builder, manifestPath string,
) error {
tmpPath := manifestTempPath(manifestPath)
outFile, err := afs.Create(tmpPath)
if err != nil {
return fmt.Errorf("failed to create temp file: %w", err)
}
err = builder.Build(ctx, outFile)
_ = outFile.Close()
if err != nil {
_ = afs.Remove(tmpPath)
return fmt.Errorf("failed to write manifest: %w", err)
}
// Rename temp to final
err = afs.Rename(tmpPath, manifestPath)
if err != nil {
_ = afs.Remove(tmpPath)
return fmt.Errorf("failed to rename manifest: %w", err)
}
return nil
}
// newFreshenBuilder constructs the manifest builder configured from CLI
// flags.
func newFreshenBuilder(cmd *cli.Command) *mfer.Builder {
builder := mfer.NewBuilder()
if cmd.Bool("include-timestamps") {
builder.SetIncludeTimestamps(true)
}
// Set up signing options if sign-key is provided
if signKey := cmd.String("sign-key"); signKey != "" {
builder.SetSigningOptions(&mfer.SigningOptions{
KeyID: mfer.GPGKeyID(signKey),
})
log.Infof("signing manifest with GPG key: %s", signKey)
}
return builder
}
// freshenScan runs the scan phase against the loaded manifest entries
// and returns the populated scanner and the count of removed files.
func (mfa *CLIApp) freshenScan(
cmd *cli.Command, manifestPath, absBase string,
existingByPath map[string]*mfer.MFFilePath,
) (*freshenScanner, int64, error) {
log.Infof("scanning filesystem...")
startScan := time.Now()
showProgress := cmd.Bool("progress")
// Leave out the manifest and a temp file left by an interrupted run,
// as gen does. A path that cannot be stat'd, normally because no file
// is there, needs no leaving out.
var excluded []fs.FileInfo
for _, p := range []string{manifestPath, manifestTempPath(manifestPath)} {
info, err := mfa.Fs.Stat(p)
if err == nil {
excluded = append(excluded, info)
}
}
scanner := &freshenScanner{
fs: mfa.Fs,
absBase: absBase,
excluded: excluded,
includeDotfiles: cmd.Bool("include-dotfiles"),
followSymlinks: cmd.Bool("follow-symlinks"),
showProgress: showProgress,
existingByPath: existingByPath,
}
err := afero.Walk(mfa.Fs, absBase, scanner.walk)
if showProgress {
log.ProgressDone()
}
if err != nil {
return nil, 0, fmt.Errorf("failed to scan filesystem: %w", err)
}
// Remaining entries in existingByPath are removed files
removed := int64(len(existingByPath))
for path := range existingByPath {
log.Verbosef("D %s", path)
}
scanDuration := time.Since(startScan)
log.Infof("scan complete in %s: %d unchanged, %d changed, %d added, %d removed",
scanDuration.Round(time.Millisecond), scanner.unchanged, scanner.changed,
scanner.added, removed)
return scanner, removed, nil
}
// hashTotals returns the total byte count and file count of entries
// that need hashing.
func hashTotals(entries []*freshenEntry) (int64, int64) {
var (
totalHashBytes int64
filesToHash int64
)
for _, e := range entries {
if e.needsHash {
totalHashBytes += e.size
filesToHash++
}
}
return totalHashBytes, filesToHash
}
// runFreshenHash processes every entry through the hasher, aborting if
// the context is canceled.
func runFreshenHash(
ctx context.Context, hasher *freshenHasher, entries []*freshenEntry,
) error {
for _, e := range entries {
select {
case <-ctx.Done():
return ctx.Err()
default:
}
err := hasher.processEntry(e)
if err != nil {
return err
}
}
return nil
}
// loadExistingEntries loads the manifest and indexes its file entries
// by path.
func (mfa *CLIApp) loadExistingEntries(
manifestPath string,
) (map[string]*mfer.MFFilePath, error) {
log.Infof("loading manifest from %s", manifestPath)
// Load existing manifest
manifest, err := mfer.NewManifestFromFile(mfa.Fs, manifestPath)
manifest, err := mfer.NewManifestFromFile(&mfer.ManifestFromFileOptions{
Path: manifestPath,
Fs: mfa.Fs,
})
if err != nil {
return fmt.Errorf("failed to load manifest: %w", err)
return nil, fmt.Errorf("failed to load manifest: %w", err)
}
existingFiles := manifest.Files()
@@ -80,217 +473,66 @@ func (mfa *CLIApp) freshenManifestOperation(ctx *cli.Context) error {
// Build map of existing entries by path
existingByPath := make(map[string]*mfer.MFFilePath, len(existingFiles))
for _, f := range existingFiles {
existingByPath[f.Path] = f
existingByPath[f.GetPath()] = f
}
// Phase 1: Scan filesystem
log.Infof("scanning filesystem...")
startScan := time.Now()
return existingByPath, nil
}
var entries []*freshenEntry
var scanCount int64
var removed, changed, added, unchanged int64
func (mfa *CLIApp) freshenManifestOperation(
ctx context.Context, cmd *cli.Command,
) error {
log.Debug("freshenManifestOperation()")
absBase, err := filepath.Abs(basePath)
basePath := cmd.String("base")
showProgress := cmd.Bool("progress")
// Find manifest file
manifestPath, err := mfa.resolveFreshenManifestPath(cmd)
if err != nil {
return fmt.Errorf("freshen: %w", err)
}
//nolint:contextcheck // mfer loads a manifest without a context
existingByPath, err := mfa.loadExistingEntries(manifestPath)
if err != nil {
return err
}
err = afero.Walk(mfa.Fs, absBase, func(path string, info fs.FileInfo, walkErr error) error {
if walkErr != nil {
return walkErr
}
// Get relative path
relPath, err := filepath.Rel(absBase, path)
if err != nil {
return err
}
// Skip the manifest file itself
if relPath == filepath.Base(manifestPath) || relPath == "."+filepath.Base(manifestPath) {
return nil
}
// Handle dotfiles
if !includeDotfiles && mfer.IsHiddenPath(filepath.ToSlash(relPath)) {
if info.IsDir() {
return filepath.SkipDir
}
return nil
}
// Skip directories
if info.IsDir() {
return nil
}
// Handle symlinks
if info.Mode()&fs.ModeSymlink != 0 {
if !followSymlinks {
return nil
}
realPath, err := filepath.EvalSymlinks(path)
if err != nil {
return nil // Skip broken symlinks
}
realInfo, err := mfa.Fs.Stat(realPath)
if err != nil || realInfo.IsDir() {
return nil
}
info = realInfo
}
scanCount++
// Check against existing manifest
existing, inManifest := existingByPath[relPath]
if inManifest {
// Check if changed (size or mtime)
existingMtime := time.Unix(existing.Mtime.Seconds, int64(existing.Mtime.Nanos))
if existing.Size != info.Size() || !existingMtime.Equal(info.ModTime()) {
changed++
log.Verbosef("M %s", relPath)
entries = append(entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
} else {
unchanged++
entries = append(entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: false,
existing: existing,
})
}
// Mark as seen
delete(existingByPath, relPath)
} else {
added++
log.Verbosef("A %s", relPath)
entries = append(entries, &freshenEntry{
path: relPath,
size: info.Size(),
mtime: info.ModTime(),
needsHash: true,
})
}
// Report scan progress
if showProgress && scanCount%100 == 0 {
log.Progressf("Scanning: %d files found", scanCount)
}
return nil
})
if showProgress {
log.ProgressDone()
}
absBase, err := filepath.Abs(basePath)
if err != nil {
return fmt.Errorf("failed to scan filesystem: %w", err)
return fmt.Errorf("freshen: invalid base path: %w", err)
}
// Remaining entries in existingByPath are removed files
removed = int64(len(existingByPath))
for path := range existingByPath {
log.Verbosef("D %s", path)
// Phase 1: Scan filesystem
scanner, removed, err := mfa.freshenScan(cmd, manifestPath, absBase,
existingByPath)
if err != nil {
return err
}
scanDuration := time.Since(startScan)
log.Infof("scan complete in %s: %d unchanged, %d changed, %d added, %d removed",
scanDuration.Round(time.Millisecond), unchanged, changed, added, removed)
// Calculate total bytes to hash
var totalHashBytes int64
var filesToHash int64
for _, e := range entries {
if e.needsHash {
totalHashBytes += e.size
filesToHash++
}
}
totalHashBytes, filesToHash := hashTotals(scanner.entries)
// Phase 2: Hash changed and new files
if filesToHash > 0 {
log.Infof("hashing %d files (%s)...", filesToHash, humanize.IBytes(uint64(totalHashBytes)))
log.Infof("hashing %d files (%s)...", filesToHash,
humanize.IBytes(safeUint64(totalHashBytes)))
}
startHash := time.Now()
var hashedFiles int64
var hashedBytes int64
builder := mfer.NewBuilder()
// Set up signing options if sign-key is provided
if signKey := ctx.String("sign-key"); signKey != "" {
builder.SetSigningOptions(&mfer.SigningOptions{
KeyID: mfer.GPGKeyID(signKey),
})
log.Infof("signing manifest with GPG key: %s", signKey)
hasher := &freshenHasher{
fs: mfa.Fs,
absBase: absBase,
showProgress: showProgress,
totalHashBytes: totalHashBytes,
filesToHash: filesToHash,
startHash: time.Now(),
builder: newFreshenBuilder(cmd),
}
for _, e := range entries {
select {
case <-ctx.Done():
return ctx.Err()
default:
}
if e.needsHash {
// Need to read and hash the file
absPath := filepath.Join(absBase, e.path)
f, err := mfa.Fs.Open(absPath)
if err != nil {
return fmt.Errorf("failed to open %s: %w", e.path, err)
}
hash, bytesRead, err := hashFile(f, e.size, func(n int64) {
if showProgress {
currentBytes := hashedBytes + n
elapsed := time.Since(startHash)
var rate float64
var eta time.Duration
if elapsed > 0 && currentBytes > 0 {
rate = float64(currentBytes) / elapsed.Seconds()
remaining := totalHashBytes - currentBytes
if rate > 0 {
eta = time.Duration(float64(remaining)/rate) * time.Second
}
}
if eta > 0 {
log.Progressf("Hashing: %d/%d files, %s/s, ETA %s",
hashedFiles, filesToHash, humanize.IBytes(uint64(rate)), eta.Round(time.Second))
} else {
log.Progressf("Hashing: %d/%d files, %s/s",
hashedFiles, filesToHash, humanize.IBytes(uint64(rate)))
}
}
})
_ = f.Close()
if err != nil {
return fmt.Errorf("failed to hash %s: %w", e.path, err)
}
hashedBytes += bytesRead
hashedFiles++
// Add to builder with computed hash
if err := addFileToBuilder(builder, e.path, e.size, e.mtime, hash); err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
}
} else {
// Use existing entry
if err := addExistingToBuilder(builder, e.existing); err != nil {
return fmt.Errorf("failed to add %s: %w", e.path, err)
}
}
err = runFreshenHash(ctx, hasher, scanner.entries)
if err != nil {
return err
}
if showProgress && filesToHash > 0 {
@@ -299,65 +541,61 @@ func (mfa *CLIApp) freshenManifestOperation(ctx *cli.Context) error {
// Print summary
log.Infof("freshen complete: %d unchanged, %d changed, %d added, %d removed",
unchanged, changed, added, removed)
scanner.unchanged, scanner.changed, scanner.added, removed)
// Skip writing if nothing changed
if changed == 0 && added == 0 && removed == 0 {
if scanner.changed == 0 && scanner.added == 0 && removed == 0 {
log.Infof("manifest unchanged, skipping write")
return nil
}
// Write updated manifest atomically (write to temp, then rename)
tmpPath := manifestPath + ".tmp"
outFile, err := mfa.Fs.Create(tmpPath)
err = writeFreshenedManifest(ctx, mfa.Fs, hasher.builder, manifestPath)
if err != nil {
return fmt.Errorf("failed to create temp file: %w", err)
}
err = builder.Build(outFile)
_ = outFile.Close()
if err != nil {
_ = mfa.Fs.Remove(tmpPath)
return fmt.Errorf("failed to write manifest: %w", err)
}
// Rename temp to final
if err := mfa.Fs.Rename(tmpPath, manifestPath); err != nil {
_ = mfa.Fs.Remove(tmpPath)
return fmt.Errorf("failed to rename manifest: %w", err)
return err
}
totalDuration := time.Since(mfa.startupTime)
if hashedBytes > 0 {
hashDuration := time.Since(startHash)
hashRate := float64(hashedBytes) / hashDuration.Seconds()
if hasher.hashedBytes > 0 {
hashDuration := time.Since(hasher.startHash)
hashRate := float64(hasher.hashedBytes) / hashDuration.Seconds()
log.Infof("hashed %s in %.1fs (%s/s)",
humanize.IBytes(uint64(hashedBytes)), totalDuration.Seconds(), humanize.IBytes(uint64(hashRate)))
humanize.IBytes(safeUint64(hasher.hashedBytes)),
totalDuration.Seconds(), humanize.IBytes(safeRateUint64(hashRate)))
}
log.Infof("wrote %d files to %s", len(entries), manifestPath)
log.Infof("wrote %d files to %s", len(scanner.entries), manifestPath)
return nil
}
// hashFile reads a file and computes its SHA256 multihash.
// Progress callback is called with bytes read so far.
func hashFile(r io.Reader, size int64, progress func(int64)) ([]byte, int64, error) {
func hashFile(r io.Reader, progress func(int64)) ([]byte, int64, error) {
h := sha256.New()
buf := make([]byte, 64*1024)
buf := make([]byte, hashBufSize)
var total int64
for {
n, err := r.Read(buf)
if n > 0 {
h.Write(buf[:n])
total += int64(n)
if progress != nil {
progress(total)
}
}
if err == io.EOF {
break
}
// Returned unwrapped: the caller renders this as
// "failed to hash <path>: <err>" and adding a second layer here
// would change that message.
if err != nil {
return nil, total, err
}
@@ -372,15 +610,37 @@ func hashFile(r io.Reader, size int64, progress func(int64)) ([]byte, int64, err
}
// addFileToBuilder adds a new file entry to the builder
func addFileToBuilder(b *mfer.Builder, path string, size int64, mtime time.Time, hash []byte) error {
return b.AddFileWithHash(mfer.RelFilePath(path), mfer.FileSize(size), mfer.ModTime(mtime), hash)
func addFileToBuilder(
b *mfer.Builder, path string, size int64, mtime time.Time, hash []byte,
) error {
return b.AddFileWithHash(
mfer.RelFilePath(path), mfer.FileSize(size), mfer.ModTime(mtime), hash)
}
// addExistingToBuilder adds an existing manifest entry to the builder
// addExistingToBuilder adds an existing manifest entry to the builder.
//
// Entries reach this path only when recordEntry classified them as
// unchanged, which requires a recorded mtime, so an absent mtime here is
// an error rather than something to paper over with the Unix epoch.
func addExistingToBuilder(b *mfer.Builder, entry *mfer.MFFilePath) error {
mtime := time.Unix(entry.Mtime.Seconds, int64(entry.Mtime.Nanos))
if len(entry.Hashes) == 0 {
mtime, ok := entryMtime(entry)
if !ok {
return fmt.Errorf("%w: %s", errEntryMissingMtime, entry.GetPath())
}
if len(entry.GetHashes()) == 0 {
return nil
}
return b.AddFileWithHash(mfer.RelFilePath(entry.Path), mfer.FileSize(entry.Size), mfer.ModTime(mtime), entry.Hashes[0].MultiHash)
err := b.AddFileWithHash(mfer.RelFilePath(entry.GetPath()),
mfer.FileSize(entry.GetSize()), mfer.ModTime(mtime),
entry.GetHashes()[0].GetMultiHash())
if err != nil {
return fmt.Errorf(
"manifest entry %s: %w (regenerate the manifest with mfer generate)",
entry.GetPath(), err,
)
}
return nil
}
+367 -51
View File
@@ -1,82 +1,398 @@
//nolint:testpackage // white-box tests exercise unexported internals
package cli
import (
"bytes"
"context"
"crypto/sha256"
"maps"
"os"
"path/filepath"
"slices"
"testing"
"time"
"github.com/multiformats/go-multihash"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"google.golang.org/protobuf/proto"
"sneak.berlin/go/mfer/mfer"
)
// stubFileInfo is a minimal fs.FileInfo for exercising recordEntry
// without touching a filesystem.
type stubFileInfo struct {
size int64
mtime time.Time
}
func (s stubFileInfo) Name() string { return "stub" }
func (s stubFileInfo) Size() int64 { return s.size }
func (s stubFileInfo) Mode() os.FileMode { return 0 }
func (s stubFileInfo) ModTime() time.Time { return s.mtime }
func (s stubFileInfo) IsDir() bool { return false }
func (s stubFileInfo) Sys() any { return nil }
// setupFreshenDir writes files, by path and content, into a fresh temp
// dir and runs gen on it. It returns the temp dir and the path of the
// manifest gen wrote there.
func setupFreshenDir(
t *testing.T, fs afero.Fs, files map[string]string,
) (string, string) {
t.Helper()
root := t.TempDir()
manifestPath := filepath.Join(root, defaultManifestName)
for path, content := range files {
require.NoError(t, fs.MkdirAll(filepath.Dir(filepath.Join(root, path)), 0o750))
writeTestFile(t, fs, filepath.Join(root, path), content)
}
opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", manifestPath, root}, fs)
require.Equal(t, 0, runCLI(opts), "stderr: %s", testStderr(t, opts))
return root, manifestPath
}
// runFreshen runs freshen on the manifest at manifestPath for the tree
// at root and requires it to succeed.
func runFreshen(t *testing.T, fs afero.Fs, root, manifestPath string) {
t.Helper()
opts := testOpts([]string{
testApp, cmdFreshen, "-q", testFlagBase, root, manifestPath,
}, fs)
require.Equal(t, 0, runCLI(opts), "stderr: %s", testStderr(t, opts))
}
// manifestFiles returns the file entries of the manifest at path.
func manifestFiles(t *testing.T, fs afero.Fs, path string) []*mfer.MFFilePath {
t.Helper()
manifest, err := mfer.NewManifestFromFile(&mfer.ManifestFromFileOptions{
Path: path,
Fs: fs,
})
require.NoError(t, err)
return manifest.Files()
}
// TestFreshenUnchanged freshens a tree that has not changed since gen
// made its manifest: the manifest must list the same entries as before.
func TestFreshenUnchanged(t *testing.T) {
// Create filesystem with test files
fs := afero.NewMemMapFs()
t.Parallel()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644))
fs := afero.NewOsFs()
root, manifestPath := setupFreshenDir(t, fs,
map[string]string{testFileTxt: "content1", testDirFile: "content2"})
before := manifestFiles(t, fs, manifestPath)
// Generate initial manifest
opts := &mfer.ScannerOptions{Fs: fs}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumeratePath("/testdir", nil))
runFreshen(t, fs, root, manifestPath)
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
after := manifestFiles(t, fs, manifestPath)
require.Len(t, after, len(before))
// Write manifest to filesystem
require.NoError(t, afero.WriteFile(fs, "/testdir/.index.mf", manifestBuf.Bytes(), 0o644))
// Parse manifest to verify
manifest, err := mfer.NewManifestFromFile(fs, "/testdir/.index.mf")
require.NoError(t, err)
assert.Len(t, manifest.Files(), 2)
for i := range before {
assert.True(t, proto.Equal(before[i], after[i]),
"entry for %s changed", before[i].GetPath())
}
}
// assertManifestLists asserts that the manifest at manifestPath lists
// exactly the files in want, each with the size and SHA-256 hash of its
// content in want and the mtime of the file of that name under root.
func assertManifestLists(
t *testing.T, fs afero.Fs, root, manifestPath string, want map[string]string,
) {
t.Helper()
files := manifestFiles(t, fs, manifestPath)
listed := make([]string, 0, len(files))
for _, f := range files {
listed = append(listed, f.GetPath())
content, ok := want[f.GetPath()]
if !ok {
continue // reported by the ElementsMatch below
}
digest := sha256.Sum256([]byte(content))
hash, err := multihash.Encode(digest[:], multihash.SHA2_256)
require.NoError(t, err)
assert.Equal(t, int64(len(content)), f.GetSize(), f.GetPath())
require.NotEmpty(t, f.GetHashes(), f.GetPath())
assert.Equal(t, hash, f.GetHashes()[0].GetMultiHash(), f.GetPath())
info, err := fs.Stat(filepath.Join(root, f.GetPath()))
require.NoError(t, err)
mtime, ok := entryMtime(f)
assert.True(t, ok && mtime.Equal(info.ModTime()),
"%s: manifest has mtime %v, file has %v",
f.GetPath(), mtime, info.ModTime())
}
assert.ElementsMatch(t, slices.Collect(maps.Keys(want)), listed)
}
// TestFreshenWithChanges makes one change to a tree after gen made its
// manifest, then freshens the manifest. The rewritten manifest must list
// exactly the files now in the tree, and check must pass on the tree.
func TestFreshenWithChanges(t *testing.T) {
// Create filesystem with test files
fs := afero.NewMemMapFs()
t.Parallel()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644))
tree := map[string]string{testFileTxt: "content1", testDirFile: "content2"}
// Generate initial manifest
opts := &mfer.ScannerOptions{Fs: fs}
s := mfer.NewScannerWithOptions(opts)
require.NoError(t, s.EnumeratePath("/testdir", nil))
// Every file a case writes gets this mtime, which differs from the one
// gen recorded, so an edit that keeps the size is told apart by its
// mtime whatever the filesystem's clock resolution.
writtenMtime := time.Unix(1_700_000_000, 0)
var manifestBuf bytes.Buffer
require.NoError(t, s.ToManifest(context.Background(), &manifestBuf, nil))
for _, tc := range []struct {
name string
write map[string]string // files to write, by path and content
remove string // file to delete, if any
}{
{
name: "modified file",
write: map[string]string{testDirFile: "modified content2"},
},
{
name: "modified file, same size",
write: map[string]string{testDirFile: "CONTENT2"},
},
{
name: "new file",
write: map[string]string{"dir/new.txt": "content3"},
},
{
name: "deleted file",
remove: testFileTxt,
},
} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
// Write manifest to filesystem
require.NoError(t, afero.WriteFile(fs, "/testdir/.index.mf", manifestBuf.Bytes(), 0o644))
fs := afero.NewOsFs()
root, manifestPath := setupFreshenDir(t, fs, tree)
// Verify initial manifest has 2 files
manifest, err := mfer.NewManifestFromFile(fs, "/testdir/.index.mf")
require.NoError(t, err)
assert.Len(t, manifest.Files(), 2)
// want is the tree as it is after the change.
want := maps.Clone(tree)
// Add a new file
require.NoError(t, afero.WriteFile(fs, "/testdir/file3.txt", []byte("content3"), 0o644))
for path, content := range tc.write {
writeTestFile(t, fs, filepath.Join(root, path), content)
require.NoError(t, fs.Chtimes(
filepath.Join(root, path), writtenMtime, writtenMtime))
want[path] = content
}
// Modify file2 (change content and size)
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("modified content2"), 0o644))
if tc.remove != "" {
require.NoError(t, fs.Remove(filepath.Join(root, tc.remove)))
delete(want, tc.remove)
}
// Remove file1
require.NoError(t, fs.Remove("/testdir/file1.txt"))
runFreshen(t, fs, root, manifestPath)
// Note: The freshen operation would need to be run here
// For now, we just verify the test setup is correct
exists, _ := afero.Exists(fs, "/testdir/file1.txt")
assert.False(t, exists)
assertManifestLists(t, fs, root, manifestPath, want)
exists, _ = afero.Exists(fs, "/testdir/file3.txt")
assert.True(t, exists)
opts := testOpts([]string{
testApp, cmdCheck, "-q", testFlagNoExtra, testFlagBase, root, manifestPath,
}, fs)
assert.Equal(t, 0, runCLI(opts), "stderr: %s", testStderr(t, opts))
})
}
}
content, _ := afero.ReadFile(fs, "/testdir/file2.txt")
assert.Equal(t, "modified content2", string(content))
// TestFreshenLeavesManifestOutOfListing freshens a manifest kept in a
// subdirectory of the tree it lists: the manifest is not listed, while an
// ordinary file of the same name at the top of the tree is. The manifest
// is recognized by file identity, which needs the real filesystem.
func TestFreshenLeavesManifestOutOfListing(t *testing.T) {
t.Parallel()
root := t.TempDir()
manifestPath := filepath.Join(root, "sub", "listing.mf")
fs := afero.NewOsFs()
require.NoError(t, fs.MkdirAll(filepath.Join(root, "sub"), 0o750))
writeTestFile(t, fs, filepath.Join(root, testFileTxt), "hello")
writeTestFile(t, fs, filepath.Join(root, "listing.mf"), "an ordinary file")
opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", manifestPath, root}, fs)
require.Equal(t, 0, runCLI(opts), "stderr: %s", testStderr(t, opts))
// A new file gives freshen something to write.
writeTestFile(t, fs, filepath.Join(root, "added.txt"), "added")
opts = testOpts([]string{
testApp, "freshen", "-q", testFlagBase, root, manifestPath,
}, fs)
require.Equal(t, 0, runCLI(opts), "stderr: %s", testStderr(t, opts))
assert.ElementsMatch(t, []string{testFileTxt, "listing.mf", "added.txt"},
manifestPaths(t, fs, manifestPath))
}
// TestFreshenLeavesLeftoverTempFileOutOfListing freshens a manifest where
// an interrupted run left its temp file beside it: the leftover is not
// listed.
func TestFreshenLeavesLeftoverTempFileOutOfListing(t *testing.T) {
t.Parallel()
root := t.TempDir()
manifestPath := filepath.Join(root, "index.mf")
fs := afero.NewOsFs()
writeTestFile(t, fs, filepath.Join(root, testFileTxt), "hello")
opts := testOpts([]string{testApp, cmdGenerate, "-q", "-o", manifestPath, root}, fs)
require.Equal(t, 0, runCLI(opts), "stderr: %s", testStderr(t, opts))
writeTestFile(t, fs, manifestTempPath(manifestPath), "part of a manifest")
// A new file gives freshen something to write.
writeTestFile(t, fs, filepath.Join(root, "added.txt"), "added")
opts = testOpts([]string{
testApp, "freshen", "-q", testFlagBase, root, manifestPath,
}, fs)
require.Equal(t, 0, runCLI(opts), "stderr: %s", testStderr(t, opts))
assert.ElementsMatch(t, []string{testFileTxt, "added.txt"},
manifestPaths(t, fs, manifestPath))
}
// TestFreshenRecordEntryMtimePresence pins the behavior of recordEntry
// with respect to MFFilePath.Mtime, which is a message pointer with
// proto3 field presence and may legitimately be absent.
//
// An absent mtime must never be read as time.Unix(0, 0): that value
// never equals a real modification time, so every entry would be
// classified as changed, re-hashed, and the manifest rewritten
// unconditionally - the exact inverse of what freshen is for, and
// silent. An entry with no mtime is therefore "changed" because it
// cannot be compared, not because it looks like it dates from 1970.
func TestFreshenRecordEntryMtimePresence(t *testing.T) {
t.Parallel()
const relPath = "file1.txt"
// The scanned file's mtime is the Unix epoch. If recordEntry ever misreads
// an absent manifest mtime as the epoch, the "absent" case below would
// compare equal to this and be classified unchanged, so the test fails.
mtime := time.Unix(0, 0)
info := stubFileInfo{size: 8, mtime: mtime}
for _, tc := range []struct {
name string
entry *mfer.MFFilePath
needsHash bool
changed int64
unchanged int64
}{
{
name: "matching mtime and size is unchanged",
entry: &mfer.MFFilePath{
Path: relPath,
Size: 8,
Mtime: &mfer.Timestamp{Seconds: mtime.Unix()},
},
needsHash: false,
changed: 0,
unchanged: 1,
},
{
name: "absent mtime is changed, not epoch",
entry: &mfer.MFFilePath{
Path: relPath,
Size: 8,
Mtime: nil,
},
needsHash: true,
changed: 1,
unchanged: 0,
},
} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
s := &freshenScanner{
existingByPath: map[string]*mfer.MFFilePath{relPath: tc.entry},
}
s.recordEntry(relPath, info)
require.Len(t, s.entries, 1)
assert.Equal(t, tc.needsHash, s.entries[0].needsHash)
assert.Equal(t, tc.changed, s.changed)
assert.Equal(t, tc.unchanged, s.unchanged)
assert.Zero(t, s.added)
})
}
}
// TestFreshenAddExistingRejectsMissingMtime pins that an entry with no
// mtime is never carried forward into a rebuilt manifest with a
// fabricated epoch timestamp.
func TestFreshenAddExistingRejectsMissingMtime(t *testing.T) {
t.Parallel()
hash, err := multihash.Encode(make([]byte, sha256.Size), multihash.SHA2_256)
require.NoError(t, err)
b := mfer.NewBuilder()
entry := &mfer.MFFilePath{
Path: "file1.txt",
Size: 8,
Mtime: nil,
Hashes: []*mfer.MFFileChecksum{
{MultiHash: hash},
},
}
err = addExistingToBuilder(b, entry)
require.ErrorIs(t, err, errEntryMissingMtime)
assert.Contains(t, err.Error(), "file1.txt")
}
// TestFreshenAddExistingRejectsShortHash pins that an existing manifest
// entry whose hash the builder refuses is reported as a problem with the
// manifest, with the command that fixes it.
func TestFreshenAddExistingRejectsShortHash(t *testing.T) {
t.Parallel()
sha1Hash, err := multihash.Encode(make([]byte, 20), multihash.SHA1)
require.NoError(t, err)
b := mfer.NewBuilder()
entry := &mfer.MFFilePath{
Path: "old.txt",
Size: 8,
Mtime: &mfer.Timestamp{Seconds: 1_700_000_000},
Hashes: []*mfer.MFFileChecksum{
{MultiHash: sha1Hash},
},
}
err = addExistingToBuilder(b, entry)
require.Error(t, err)
assert.Contains(t, err.Error(), "manifest entry old.txt")
assert.Contains(t, err.Error(), "mfer generate")
assert.Zero(t, b.FileCount())
}
// TestEntryMtime pins the presence semantics the callers depend on.
func TestEntryMtime(t *testing.T) {
t.Parallel()
got, ok := entryMtime(&mfer.MFFilePath{Mtime: nil})
assert.False(t, ok)
assert.True(t, got.IsZero())
got, ok = entryMtime(&mfer.MFFilePath{
Mtime: &mfer.Timestamp{Seconds: 1_700_000_000, Nanos: 500},
})
assert.True(t, ok)
assert.Equal(t, time.Unix(1_700_000_000, 500), got)
}
+205 -93
View File
@@ -1,6 +1,8 @@
package cli
import (
"context"
"errors"
"fmt"
"os"
"os/signal"
@@ -11,164 +13,272 @@ import (
"github.com/dustin/go-humanize"
"github.com/spf13/afero"
"github.com/urfave/cli/v2"
"github.com/urfave/cli/v3"
"sneak.berlin/go/mfer/internal/log"
"sneak.berlin/go/mfer/mfer"
)
func (mfa *CLIApp) generateManifestOperation(ctx *cli.Context) error {
log.Debug("generateManifestOperation()")
var (
// errPathNotExist indicates an input path that does not exist.
errPathNotExist = errors.New("path does not exist")
// errOutputExists indicates the output file already exists and
// --force was not given. It is wrapped mid-sentence so that the
// rendered message stays exactly as mfer has always printed it.
errOutputExists = errors.New(
"already exists (use --force to overwrite)")
)
// reportEnumProgress renders enumeration progress until the channel
// closes.
func reportEnumProgress(progress <-chan mfer.EnumerateStatus, wg *sync.WaitGroup) {
defer wg.Done()
for status := range progress {
log.Progressf("Enumerating: %d files, %s",
status.FilesFound,
humanize.IBytes(safeUint64(int64(status.BytesFound))))
}
log.ProgressDone()
}
// reportScanProgress renders scan progress until the channel closes.
func reportScanProgress(progress <-chan mfer.ScanStatus, wg *sync.WaitGroup) {
defer wg.Done()
for status := range progress {
if status.ETA > 0 {
log.Progressf("Scanning: %d/%d files, %s/s, ETA %s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)),
status.ETA.Round(time.Second))
} else {
log.Progressf("Scanning: %d/%d files, %s/s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(safeRateUint64(status.BytesPerSec)))
}
}
log.ProgressDone()
}
// collectInputPaths validates the input path arguments and returns them
// as absolute paths.
func (mfa *CLIApp) collectInputPaths(args cli.Args) ([]string, error) {
paths := make([]string, 0, args.Len())
for i := range args.Len() {
inputPath := args.Get(i)
ap, err := filepath.Abs(inputPath)
if err != nil {
return nil, fmt.Errorf("generate: invalid path %q: %w", inputPath, err)
}
// Validate path exists before adding to list
if exists, _ := afero.Exists(mfa.Fs, ap); !exists {
return nil, fmt.Errorf("%w: %s", errPathNotExist, inputPath)
}
log.Debugf("enumerating path: %s", ap)
paths = append(paths, ap)
}
return paths, nil
}
// buildScannerOptions constructs scanner options from the CLI flags.
func (mfa *CLIApp) buildScannerOptions(cmd *cli.Command) *mfer.ScannerOptions {
output := cmd.String("output")
opts := &mfer.ScannerOptions{
IncludeDotfiles: ctx.Bool("IncludeDotfiles"),
FollowSymLinks: ctx.Bool("FollowSymLinks"),
Fs: mfa.Fs,
IncludeDotfiles: cmd.Bool("include-dotfiles"),
FollowSymLinks: cmd.Bool("follow-symlinks"),
IncludeTimestamps: cmd.Bool("include-timestamps"),
Fs: mfa.Fs,
// Neither a manifest being replaced nor a temp file left by an
// interrupted run belongs in the new manifest.
ExcludePaths: []string{output, manifestTempPath(output)},
}
// Set seed for deterministic UUID if provided
if seed := ctx.String("seed"); seed != "" {
if seed := cmd.String("seed"); seed != "" {
opts.Seed = seed
log.Infof("using deterministic seed for manifest UUID")
}
// Set up signing options if sign-key is provided
if signKey := ctx.String("sign-key"); signKey != "" {
if signKey := cmd.String("sign-key"); signKey != "" {
opts.SigningOptions = &mfer.SigningOptions{
KeyID: mfer.GPGKeyID(signKey),
}
log.Infof("signing manifest with GPG key: %s", signKey)
}
s := mfer.NewScannerWithOptions(opts)
// Phase 1: Enumeration - collect paths and stat files
args := ctx.Args()
showProgress := ctx.Bool("progress")
// Set up enumeration progress reporting
var enumProgress chan mfer.EnumerateStatus
var enumWg sync.WaitGroup
if showProgress {
enumProgress = make(chan mfer.EnumerateStatus, 1)
enumWg.Add(1)
go func() {
defer enumWg.Done()
for status := range enumProgress {
log.Progressf("Enumerating: %d files, %s",
status.FilesFound,
humanize.IBytes(uint64(status.BytesFound)))
}
log.ProgressDone()
}()
}
return opts
}
// enumerateInputs runs the enumeration phase over the argument paths,
// or the current directory when no arguments are given.
func (mfa *CLIApp) enumerateInputs(
s *mfer.Scanner, args cli.Args, enumProgress chan mfer.EnumerateStatus,
) error {
if args.Len() == 0 {
// Default to current directory
if err := s.EnumeratePath(".", enumProgress); err != nil {
return err
}
} else {
// Collect and validate all paths first
paths := make([]string, 0, args.Len())
for i := 0; i < args.Len(); i++ {
inputPath := args.Get(i)
ap, err := filepath.Abs(inputPath)
if err != nil {
return err
}
// Validate path exists before adding to list
if exists, _ := afero.Exists(mfa.Fs, ap); !exists {
return fmt.Errorf("path does not exist: %s", inputPath)
}
log.Debugf("enumerating path: %s", ap)
paths = append(paths, ap)
}
if err := s.EnumeratePaths(enumProgress, paths...); err != nil {
return err
err := s.EnumeratePath(".", enumProgress)
if err != nil {
return fmt.Errorf(
"generate: failed to enumerate current directory: %w", err)
}
return nil
}
// Collect and validate all paths first
paths, err := mfa.collectInputPaths(args)
if err != nil {
return err
}
err = s.EnumeratePaths(enumProgress, paths...)
if err != nil {
return fmt.Errorf("generate: failed to enumerate paths: %w", err)
}
return nil
}
// cleanupOnSignal installs a handler that removes the temp output file
// and exits when the process is interrupted. It returns the signal
// channel so the caller can stop and close it when done.
func (mfa *CLIApp) cleanupOnSignal(outFile afero.File, tmpPath string) chan os.Signal {
sigChan := make(chan os.Signal, 1)
signal.Notify(sigChan, os.Interrupt, syscall.SIGTERM)
go func() {
sig, ok := <-sigChan
if !ok || sig == nil {
return // Channel closed normally, not a signal
}
_ = outFile.Close()
_ = mfa.Fs.Remove(tmpPath)
os.Exit(1)
}()
return sigChan
}
// runEnumeratePhase enumerates all input paths with optional progress
// reporting and logs the totals.
func (mfa *CLIApp) runEnumeratePhase(cmd *cli.Command, s *mfer.Scanner) error {
// Set up enumeration progress reporting
var (
enumProgress chan mfer.EnumerateStatus
enumWg sync.WaitGroup
)
if cmd.Bool("progress") {
enumProgress = make(chan mfer.EnumerateStatus, 1)
enumWg.Add(1)
go reportEnumProgress(enumProgress, &enumWg)
}
err := mfa.enumerateInputs(s, cmd.Args(), enumProgress)
if err != nil {
return err
}
enumWg.Wait()
log.Infof("enumerated %d files, %s total", s.FileCount(), humanize.IBytes(uint64(s.TotalBytes())))
log.Infof("enumerated %d files, %s total", s.FileCount(),
humanize.IBytes(safeUint64(int64(s.TotalBytes()))))
return nil
}
func (mfa *CLIApp) generateManifestOperation(
ctx context.Context, cmd *cli.Command,
) error {
log.Debug("generateManifestOperation()")
s := mfer.NewScannerWithOptions(mfa.buildScannerOptions(cmd))
// Phase 1: Enumeration - collect paths and stat files
err := mfa.runEnumeratePhase(cmd, s)
if err != nil {
return err
}
showProgress := cmd.Bool("progress")
// Check if output file exists
outputPath := ctx.String("output")
if exists, _ := afero.Exists(mfa.Fs, outputPath); exists {
if !ctx.Bool("force") {
return fmt.Errorf("output file %s already exists (use --force to overwrite)", outputPath)
}
outputPath := cmd.String("output")
if exists, _ := afero.Exists(mfa.Fs, outputPath); exists && !cmd.Bool("force") {
return fmt.Errorf("output file %s %w", outputPath, errOutputExists)
}
// Create temp file for atomic write
tmpPath := outputPath + ".tmp"
tmpPath := manifestTempPath(outputPath)
outFile, err := mfa.Fs.Create(tmpPath)
if err != nil {
return fmt.Errorf("failed to create temp file: %w", err)
}
// Set up signal handler to clean up temp file on Ctrl-C
sigChan := make(chan os.Signal, 1)
signal.Notify(sigChan, os.Interrupt, syscall.SIGTERM)
go func() {
sig, ok := <-sigChan
if !ok || sig == nil {
return // Channel closed normally, not a signal
}
_ = outFile.Close()
_ = mfa.Fs.Remove(tmpPath)
os.Exit(1)
}()
sigChan := mfa.cleanupOnSignal(outFile, tmpPath)
// Clean up temp file on any error or interruption
success := false
defer func() {
signal.Stop(sigChan)
close(sigChan)
_ = outFile.Close()
if !success {
_ = mfa.Fs.Remove(tmpPath)
}
}()
// Phase 2: Scan - read file contents and generate manifest
var scanProgress chan mfer.ScanStatus
var scanWg sync.WaitGroup
var (
scanProgress chan mfer.ScanStatus
scanWg sync.WaitGroup
)
if showProgress {
scanProgress = make(chan mfer.ScanStatus, 1)
scanWg.Add(1)
go func() {
defer scanWg.Done()
for status := range scanProgress {
if status.ETA > 0 {
log.Progressf("Scanning: %d/%d files, %s/s, ETA %s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)),
status.ETA.Round(time.Second))
} else {
log.Progressf("Scanning: %d/%d files, %s/s",
status.ScannedFiles,
status.TotalFiles,
humanize.IBytes(uint64(status.BytesPerSec)))
}
}
log.ProgressDone()
}()
go reportScanProgress(scanProgress, &scanWg)
}
err = s.ToManifest(ctx.Context, outFile, scanProgress)
err = s.ToManifest(ctx, outFile, scanProgress)
scanWg.Wait()
if err != nil {
return fmt.Errorf("failed to generate manifest: %w", err)
}
// Close file before rename to ensure all data is flushed
if err := outFile.Close(); err != nil {
err = outFile.Close()
if err != nil {
return fmt.Errorf("failed to close temp file: %w", err)
}
// Atomic rename
if err := mfa.Fs.Rename(tmpPath, outputPath); err != nil {
err = mfa.Fs.Rename(tmpPath, outputPath)
if err != nil {
return fmt.Errorf("failed to rename temp file: %w", err)
}
@@ -176,7 +286,9 @@ func (mfa *CLIApp) generateManifestOperation(ctx *cli.Context) error {
elapsed := time.Since(mfa.startupTime).Seconds()
rate := float64(s.TotalBytes()) / elapsed
log.Infof("wrote %d files (%s) to %s in %.1fs (%s/s)", s.FileCount(), humanize.IBytes(uint64(s.TotalBytes())), outputPath, elapsed, humanize.IBytes(uint64(rate)))
log.Infof("wrote %d files (%s) to %s in %.1fs (%s/s)", s.FileCount(),
humanize.IBytes(safeUint64(int64(s.TotalBytes()))), outputPath, elapsed,
humanize.IBytes(safeRateUint64(rate)))
return nil
}
+28 -30
View File
@@ -1,47 +1,38 @@
package cli
import (
"context"
"fmt"
"time"
"github.com/urfave/cli/v2"
"github.com/urfave/cli/v3"
"sneak.berlin/go/mfer/internal/log"
"sneak.berlin/go/mfer/mfer"
)
func (mfa *CLIApp) listManifestOperation(ctx *cli.Context) error {
func (mfa *CLIApp) listManifestOperation(ctx context.Context, cmd *cli.Command) error {
// Default to ErrorLevel for clean output
log.SetLevel(log.ErrorLevel)
longFormat := ctx.Bool("long")
print0 := ctx.Bool("print0")
longFormat := cmd.Bool("long")
print0 := cmd.Bool("print0")
// Find manifest file
var manifestPath string
var err error
if ctx.Args().Len() > 0 {
arg := ctx.Args().Get(0)
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
manifestPath, err = findManifest(mfa.Fs, arg)
if err != nil {
return err
}
} else {
manifestPath = arg
}
} else {
manifestPath, err = findManifest(mfa.Fs, ".")
if err != nil {
return err
}
pathOrURL, err := mfa.resolveManifestArg(cmd)
if err != nil {
return fmt.Errorf("list: %w", err)
}
// Load manifest
manifest, err := mfer.NewManifestFromFile(mfa.Fs, manifestPath)
rc, err := mfa.openManifestReader(ctx, pathOrURL)
if err != nil {
return fmt.Errorf("failed to load manifest: %w", err)
return fmt.Errorf("list: %w", err)
}
defer func() { _ = rc.Close() }()
//nolint:contextcheck // mfer loads a manifest without a context
manifest, err := mfer.NewManifestFromReader(rc)
if err != nil {
return fmt.Errorf("list: failed to parse manifest: %w", err)
}
files := manifest.Files()
@@ -54,10 +45,17 @@ func (mfa *CLIApp) listManifestOperation(ctx *cli.Context) error {
for _, f := range files {
if longFormat {
mtime := time.Unix(f.Mtime.Seconds, int64(f.Mtime.Nanos))
_, _ = fmt.Fprintf(mfa.Stdout, "%d\t%s\t%s%s", f.Size, mtime.Format(time.RFC3339), f.Path, lineEnd)
// An entry may legitimately carry no mtime; render that as
// mtimeAbsent rather than as the Unix epoch.
mtimeStr := mtimeAbsent
if mtime, ok := entryMtime(f); ok {
mtimeStr = mtime.Format(time.RFC3339)
}
_, _ = fmt.Fprintf(mfa.Stdout, "%d\t%s\t%s%s",
f.GetSize(), mtimeStr, f.GetPath(), lineEnd)
} else {
_, _ = fmt.Fprintf(mfa.Stdout, "%s%s", f.Path, lineEnd)
_, _ = fmt.Fprintf(mfa.Stdout, "%s%s", f.GetPath(), lineEnd)
}
}
+86
View File
@@ -0,0 +1,86 @@
package cli
import (
"context"
"errors"
"fmt"
"io"
"net/http"
"strings"
"time"
"github.com/urfave/cli/v3"
)
// manifestFetchTimeout bounds HTTP requests made to fetch a manifest.
const manifestFetchTimeout = 30 * time.Second
// errHTTPStatus indicates an HTTP response with a non-OK status code.
//
// Its text is the literal "HTTP" prefix of the rendered "HTTP <code>"
// message that mfer has always printed, so that wrapping it does not
// change any user-visible output. Match it with errors.Is; do not read
// its message.
var errHTTPStatus = errors.New("HTTP")
// isHTTPURL returns true if the string starts with http:// or https://.
func isHTTPURL(s string) bool {
return strings.HasPrefix(s, "http://") || strings.HasPrefix(s, "https://")
}
// openManifestReader opens a manifest from a path or URL and returns a ReadCloser.
// The caller must close the returned reader.
func (mfa *CLIApp) openManifestReader(
ctx context.Context, pathOrURL string,
) (io.ReadCloser, error) {
if isHTTPURL(pathOrURL) {
client := &http.Client{Timeout: manifestFetchTimeout}
req, err := http.NewRequestWithContext(ctx, http.MethodGet, pathOrURL, nil)
if err != nil {
return nil, fmt.Errorf("failed to fetch %s: %w", pathOrURL, err)
}
resp, err := client.Do(req)
if err != nil {
return nil, fmt.Errorf("failed to fetch %s: %w", pathOrURL, err)
}
if resp.StatusCode != http.StatusOK {
_ = resp.Body.Close()
return nil, fmt.Errorf("failed to fetch %s: %w %d",
pathOrURL, errHTTPStatus, resp.StatusCode)
}
return resp.Body, nil
}
f, err := mfa.Fs.Open(pathOrURL)
if err != nil {
return nil, err
}
return f, nil
}
// resolveManifestArg resolves the manifest path from CLI arguments.
// HTTP(S) URLs are returned as-is. Directories are searched for index.mf.
// If no argument is given, the current directory is searched.
func (mfa *CLIApp) resolveManifestArg(cmd *cli.Command) (string, error) {
if cmd.Args().Len() > 0 {
arg := cmd.Args().Get(0)
if isHTTPURL(arg) {
return arg, nil
}
info, statErr := mfa.Fs.Stat(arg)
if statErr == nil && info.IsDir() {
return findManifest(mfa.Fs, arg)
}
return arg, nil
}
return findManifest(mfa.Fs, ".")
}
+371 -195
View File
@@ -1,26 +1,61 @@
package cli
import (
"context"
"errors"
"fmt"
"io"
"os"
"time"
"github.com/spf13/afero"
"github.com/urfave/cli/v2"
"github.com/urfave/cli/v3"
"sneak.berlin/go/mfer/internal/log"
"sneak.berlin/go/mfer/mfer"
)
// Command and flag names shared across command definitions and tests.
const (
cmdGenerate = "generate"
cmdCheck = "check"
cmdFreshen = "freshen"
cmdExport = "export"
cmdFetch = "fetch"
cmdVersion = "version"
flagProgress = "progress"
flagTimeout = "timeout"
flagDest = "dest"
flagRequireSignature = "require-signature"
manifestArgsUsage = "[manifest file]"
// defaultManifestName is the filename gen writes by default, the one
// looked for when a command is given a directory, and the one fetch
// appends to a directory URL.
defaultManifestName = "index.mf"
)
// manifestTempPath returns the temp file gen and freshen write a manifest
// to before renaming it to out.
func manifestTempPath(out string) string {
return out + ".tmp"
}
// errUnknownCommand indicates an unrecognized command argument.
var errUnknownCommand = errors.New("unknown command")
// CLIApp is the main CLI application container. It holds configuration,
// I/O streams, and filesystem abstraction to enable testing and flexibility.
//
//nolint:revive // established name used throughout the codebase and tests
type CLIApp struct {
appname string
version string
gitrev string
startupTime time.Time
exitCode int
app *cli.App
app *cli.Command
Stdin io.Reader // Standard input stream
Stdout io.Writer // Standard output stream for normal output
@@ -41,54 +76,324 @@ const banner = `
\ \:\ \ \:\ \ \::/ \ \:\
\__\/ \__\/ \__\/ \__\/`
func (mfa *CLIApp) printBanner() {
if log.GetLevel() <= log.InfoLevel {
_, _ = fmt.Fprintln(mfa.Stdout, banner)
_, _ = fmt.Fprintf(mfa.Stdout, " mfer by @sneak: v%s released %s\n", mfer.Version, mfer.ReleaseDate)
_, _ = fmt.Fprintln(mfa.Stdout, " https://sneak.berlin/go/mfer")
}
}
// VersionString returns the version and git revision formatted for display.
func (mfa *CLIApp) VersionString() string {
if mfa.gitrev != "" {
return fmt.Sprintf("%s (%s)", mfer.Version, mfa.gitrev)
}
return mfer.Version
}
func (mfa *CLIApp) setVerbosity(c *cli.Context) {
_, present := os.LookupEnv("MFER_DEBUG")
if present {
log.EnableDebugLogging()
} else if c.Bool("quiet") {
log.SetLevel(log.ErrorLevel)
} else {
log.SetLevelFromVerbosity(c.Count("verbose"))
// printVersion writes the version line shared by the --version flag and the
// version subcommand, so both produce identical output.
func (mfa *CLIApp) printVersion() {
_, _ = fmt.Fprintf(mfa.Stdout, "%s version %s\n", mfa.appname, mfa.VersionString())
}
func (mfa *CLIApp) printBanner() {
if log.GetLevel() <= log.InfoLevel {
_, _ = fmt.Fprintln(mfa.Stdout, banner)
_, _ = fmt.Fprintf(mfa.Stdout,
" mfer by @sneak: v%s released %s\n",
mfer.Version, mfer.ReleaseDate)
_, _ = fmt.Fprintln(mfa.Stdout, " https://sneak.berlin/go/mfer")
}
}
// commonFlags returns the flags shared by most commands (-v, -q)
// setVerbosity sets the log level from -v and -q, given before the subcommand
// name (the root's copies), after it (the subcommand's own copies), or both.
// urfave/cli reads a flag from the nearest command that defines it, so each
// command in the lineage is asked. The highest -v count wins rather than the
// sum, because a subcommand without its own copies reads the root's.
func (mfa *CLIApp) setVerbosity(cmd *cli.Command) {
_, present := os.LookupEnv("MFER_DEBUG")
verbosity := 0
quiet := false
for _, c := range cmd.Lineage() {
verbosity = max(verbosity, c.Count("verbose"))
quiet = quiet || c.Bool("quiet")
}
switch {
case present:
log.EnableDebugLogging()
case quiet:
log.SetLevel(log.ErrorLevel)
default:
log.SetLevelFromVerbosity(verbosity)
}
}
// commonFlags returns the -v and -q flags taken by the root and by the
// generate, check, freshen and fetch subcommands. They are local, so the
// root's copies are not inherited by the subcommands that do not take them.
func commonFlags() []cli.Flag {
return []cli.Flag{
&cli.BoolFlag{
Name: "verbose",
Aliases: []string{"v"},
Usage: "Increase verbosity (-v for verbose, -vv for debug)",
Count: new(int),
Usage: "Increase verbosity (-v for verbose, -v -v for debug)",
Local: true,
},
&cli.BoolFlag{
Name: "quiet",
Aliases: []string{"q"},
Usage: "Suppress output except errors",
Local: true,
},
}
}
// stopOnFirstArg returns the StopOnNthArg setting every command uses: flags
// are read only before the command's first argument, and everything after it
// is an argument, as with urfave/cli v2. So `mfer gen d -v` names a path "-v".
func stopOnFirstArg() *int {
n := 1
return &n
}
// requireSignatureFlag returns the --require-signature flag taken by the
// check and fetch subcommands.
func requireSignatureFlag() *cli.StringFlag {
return &cli.StringFlag{
Name: flagRequireSignature,
Aliases: []string{"S"},
Usage: "Require manifest to be signed by the specified GPG key ID",
Sources: cli.EnvVars("MFER_REQUIRE_SIGNATURE"),
}
}
func (mfa *CLIApp) generateCommand() *cli.Command {
return &cli.Command{
Name: cmdGenerate,
Aliases: []string{"gen"},
Usage: "Generate manifest file",
ArgsUsage: "[path ...]",
StopOnNthArg: stopOnFirstArg(),
Action: func(ctx context.Context, cmd *cli.Command) error {
mfa.setVerbosity(cmd)
mfa.printBanner()
return mfa.generateManifestOperation(ctx, cmd)
},
Flags: append(commonFlags(),
&cli.BoolFlag{
Name: "follow-symlinks",
Aliases: []string{"L"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "include-dotfiles",
Aliases: []string{"IncludeDotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.StringFlag{
Name: "output",
Value: defaultManifestName,
Aliases: []string{"o"},
Usage: "Specify output filename",
},
&cli.BoolFlag{
Name: "force",
Aliases: []string{"f"},
Usage: "Overwrite output file if it exists",
},
&cli.BoolFlag{
Name: flagProgress,
Aliases: []string{"P"},
Usage: "Show progress during enumeration and scanning",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
Sources: cli.EnvVars("MFER_SIGN_KEY"),
},
&cli.StringFlag{
Name: "seed",
Usage: "Seed value for deterministic manifest UUID",
Sources: cli.EnvVars("MFER_SEED"),
},
&cli.BoolFlag{
Name: "include-timestamps",
Usage: "Include createdAt timestamp in manifest " +
"(omitted by default for determinism)",
},
),
}
}
func (mfa *CLIApp) checkCommand() *cli.Command {
return &cli.Command{
Name: cmdCheck,
Usage: "Validate files using manifest file",
ArgsUsage: manifestArgsUsage,
StopOnNthArg: stopOnFirstArg(),
Action: func(ctx context.Context, cmd *cli.Command) error {
mfa.setVerbosity(cmd)
mfa.printBanner()
return mfa.checkManifestOperation(ctx, cmd)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths from manifest",
},
&cli.BoolFlag{
Name: flagProgress,
Aliases: []string{"P"},
Usage: "Show progress during checking",
},
&cli.BoolFlag{
Name: "no-extra-files",
Usage: "Fail, instead of warning, if files in base directory are not in manifest",
},
requireSignatureFlag(),
),
}
}
func (mfa *CLIApp) freshenCommand() *cli.Command {
return &cli.Command{
Name: cmdFreshen,
Usage: "Update manifest with changed, new, and removed files",
ArgsUsage: manifestArgsUsage,
StopOnNthArg: stopOnFirstArg(),
Action: func(ctx context.Context, cmd *cli.Command) error {
mfa.setVerbosity(cmd)
mfa.printBanner()
return mfa.freshenManifestOperation(ctx, cmd)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths",
},
&cli.BoolFlag{
Name: "follow-symlinks",
Aliases: []string{"L"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "include-dotfiles",
Aliases: []string{"IncludeDotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.BoolFlag{
Name: flagProgress,
Aliases: []string{"P"},
Usage: "Show progress during scanning and hashing",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
Sources: cli.EnvVars("MFER_SIGN_KEY"),
},
&cli.BoolFlag{
Name: "include-timestamps",
Usage: "Include createdAt timestamp in manifest " +
"(omitted by default for determinism)",
},
),
}
}
func (mfa *CLIApp) exportCommand() *cli.Command {
return &cli.Command{
Name: cmdExport,
Usage: "Export manifest contents as JSON",
ArgsUsage: "[manifest file or URL]",
StopOnNthArg: stopOnFirstArg(),
Action: func(ctx context.Context, cmd *cli.Command) error {
mfa.setVerbosity(cmd)
return mfa.exportManifestOperation(ctx, cmd)
},
}
}
func (mfa *CLIApp) versionCommand() *cli.Command {
return &cli.Command{
Name: cmdVersion,
Usage: "Show version",
StopOnNthArg: stopOnFirstArg(),
Action: func(context.Context, *cli.Command) error {
mfa.printVersion()
return nil
},
}
}
func (mfa *CLIApp) listCommand() *cli.Command {
return &cli.Command{
Name: "list",
Aliases: []string{"ls"},
Usage: "List files in manifest",
ArgsUsage: manifestArgsUsage,
StopOnNthArg: stopOnFirstArg(),
Action: mfa.listManifestOperation,
Flags: []cli.Flag{
&cli.BoolFlag{
Name: "long",
Aliases: []string{"l"},
Usage: "Show size and mtime",
},
&cli.BoolFlag{
Name: "print0",
Usage: "Separate entries with NUL character (for xargs -0)",
},
},
}
}
func (mfa *CLIApp) fetchCommand() *cli.Command {
return &cli.Command{
Name: cmdFetch,
Usage: "fetch manifest and referenced files",
ArgsUsage: "URL",
StopOnNthArg: stopOnFirstArg(),
Action: func(ctx context.Context, cmd *cli.Command) error {
mfa.setVerbosity(cmd)
mfa.printBanner()
return mfa.fetchManifestOperation(ctx, cmd)
},
Flags: append(commonFlags(),
&cli.DurationFlag{
Name: flagTimeout,
Value: httpTimeout,
Usage: "Time limit for each HTTP request, including the download " +
"of its body",
},
&cli.StringFlag{
Name: flagDest,
Aliases: []string{"d"},
Value: ".",
Usage: "Directory to download the files and the manifest into",
},
requireSignatureFlag(),
),
}
}
func (mfa *CLIApp) run(args []string) {
mfa.startupTime = time.Now()
if NO_COLOR {
if NoColor {
// shoutout to rob pike who thinks it's juvenile
log.DisableStyling()
}
@@ -97,187 +402,58 @@ func (mfa *CLIApp) run(args []string) {
log.SetOutput(mfa.Stdout, mfa.Stderr)
log.Init()
mfa.app = &cli.App{
Name: mfa.appname,
Usage: "Manifest generator",
Version: mfa.VersionString(),
EnableBashCompletion: true,
Writer: mfa.Stdout,
ErrWriter: mfa.Stderr,
Action: func(c *cli.Context) error {
if c.Args().Len() > 0 {
return fmt.Errorf("unknown command %q", c.Args().First())
// -v means verbose, not version, at the root as on the generate, check,
// freshen and fetch subcommands: verbose is the more common meaning of -v
// in tools that offer both. urfave/cli's built-in version flag claims -v
// by default, so the version flag takes the capital -V instead.
// VersionFlag and VersionPrinter are urfave/cli package globals; run() is
// serialized in tests, so assigning them here is safe.
cli.VersionFlag = &cli.BoolFlag{
Name: cmdVersion,
Aliases: []string{"V"},
Usage: "print the version",
Local: true,
}
cli.VersionPrinter = func(*cli.Command) { mfa.printVersion() }
mfa.app = &cli.Command{
Name: mfa.appname,
Usage: "Manifest generator",
Version: mfa.VersionString(),
EnableShellCompletion: true,
Writer: mfa.Stdout,
// v3 writes its "Incorrect Usage" line to ErrWriter; v2 wrote it to
// stdout, before the help, so it stays on stdout.
ErrWriter: mfa.Stdout,
Flags: commonFlags(),
StopOnNthArg: stopOnFirstArg(),
Action: func(_ context.Context, cmd *cli.Command) error {
if cmd.Args().Len() > 0 {
return fmt.Errorf("%w %q", errUnknownCommand, cmd.Args().First())
}
mfa.setVerbosity(cmd)
mfa.printBanner()
return cli.ShowAppHelp(c)
return cli.ShowRootCommandHelp(cmd)
},
Commands: []*cli.Command{
{
Name: "generate",
Aliases: []string{"gen"},
Usage: "Generate manifest file",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.generateManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.BoolFlag{
Name: "FollowSymLinks",
Aliases: []string{"follow-symlinks"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "IncludeDotfiles",
Aliases: []string{"include-dotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.StringFlag{
Name: "output",
Value: "./.index.mf",
Aliases: []string{"o"},
Usage: "Specify output filename",
},
&cli.BoolFlag{
Name: "force",
Aliases: []string{"f"},
Usage: "Overwrite output file if it exists",
},
&cli.BoolFlag{
Name: "progress",
Aliases: []string{"P"},
Usage: "Show progress during enumeration and scanning",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
EnvVars: []string{"MFER_SIGN_KEY"},
},
&cli.StringFlag{
Name: "seed",
Usage: "Seed value for deterministic manifest UUID",
EnvVars: []string{"MFER_SEED"},
},
),
},
{
Name: "check",
Usage: "Validate files using manifest file",
ArgsUsage: "[manifest file]",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.checkManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths from manifest",
},
&cli.BoolFlag{
Name: "progress",
Aliases: []string{"P"},
Usage: "Show progress during checking",
},
&cli.BoolFlag{
Name: "no-extra-files",
Usage: "Fail if files exist in base directory that are not in manifest",
},
&cli.StringFlag{
Name: "require-signature",
Aliases: []string{"S"},
Usage: "Require manifest to be signed by the specified GPG key ID",
EnvVars: []string{"MFER_REQUIRE_SIGNATURE"},
},
),
},
{
Name: "freshen",
Usage: "Update manifest with changed, new, and removed files",
ArgsUsage: "[manifest file]",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.freshenManifestOperation(c)
},
Flags: append(commonFlags(),
&cli.StringFlag{
Name: "base",
Aliases: []string{"b"},
Value: ".",
Usage: "Base directory for resolving relative paths",
},
&cli.BoolFlag{
Name: "FollowSymLinks",
Aliases: []string{"follow-symlinks"},
Usage: "Resolve encountered symlinks",
},
&cli.BoolFlag{
Name: "IncludeDotfiles",
Aliases: []string{"include-dotfiles"},
Usage: "Include dot (hidden) files (excluded by default)",
},
&cli.BoolFlag{
Name: "progress",
Aliases: []string{"P"},
Usage: "Show progress during scanning and hashing",
},
&cli.StringFlag{
Name: "sign-key",
Aliases: []string{"s"},
Usage: "GPG key ID to sign the manifest with",
EnvVars: []string{"MFER_SIGN_KEY"},
},
),
},
{
Name: "version",
Usage: "Show version",
Action: func(c *cli.Context) error {
_, _ = fmt.Fprintln(mfa.Stdout, mfa.VersionString())
return nil
},
},
{
Name: "list",
Aliases: []string{"ls"},
Usage: "List files in manifest",
ArgsUsage: "[manifest file]",
Action: func(c *cli.Context) error {
return mfa.listManifestOperation(c)
},
Flags: []cli.Flag{
&cli.BoolFlag{
Name: "long",
Aliases: []string{"l"},
Usage: "Show size and mtime",
},
&cli.BoolFlag{
Name: "print0",
Usage: "Separate entries with NUL character (for xargs -0)",
},
},
},
{
Name: "fetch",
Usage: "fetch manifest and referenced files",
Action: func(c *cli.Context) error {
mfa.setVerbosity(c)
mfa.printBanner()
return mfa.fetchManifestOperation(c)
},
Flags: commonFlags(),
},
mfa.generateCommand(),
mfa.checkCommand(),
mfa.freshenCommand(),
mfa.exportCommand(),
mfa.versionCommand(),
mfa.listCommand(),
mfa.fetchCommand(),
},
}
mfa.app.HideVersion = true
err := mfa.app.Run(args)
mfa.app.HideVersion = false
err := mfa.app.Run(context.Background(), args)
if err != nil {
mfa.exitCode = 1
log.WithError(err).Debugf("exiting")
log.Errorf("%s", err)
}
}
+28
View File
@@ -0,0 +1,28 @@
package cli
import (
"time"
"sneak.berlin/go/mfer/mfer"
)
// mtimeAbsent is printed in place of a modification time when a manifest
// entry does not carry one.
const mtimeAbsent = "-"
// entryMtime returns the modification time recorded for a manifest entry.
//
// MFFilePath.Mtime is a message pointer with proto3 field presence, so an
// absent mtime is a representable, on-the-wire-valid state. It must never
// be conflated with a recorded mtime of the Unix epoch: callers that
// compare mtimes have to treat "absent" as "unknown", not as
// 1970-01-01T00:00:00Z, or every entry compares as modified. ok reports
// whether an mtime was actually recorded.
func entryMtime(entry *mfer.MFFilePath) (time.Time, bool) {
ts := entry.GetMtime()
if ts == nil {
return time.Time{}, false
}
return time.Unix(ts.GetSeconds(), int64(ts.GetNanos())), true
}
+183 -77
View File
@@ -1,17 +1,25 @@
// Package log provides leveled logging on top of log/slog, and helpers that
// print progress lines which overwrite each other in place on a terminal.
//
// Until Init runs, log records go to slog.Default(), so a program that uses
// package mfer as a library gets them wherever it sends its own slog output.
// Init switches to the CLI's format on the stderr writer given to SetOutput.
// Progress lines are not log records: they are written straight to the stdout
// writer, never through slog.
package log
import (
"context"
"fmt"
"io"
"log/slog"
"os"
"path/filepath"
"runtime"
"sync"
"github.com/apex/log"
acli "github.com/apex/log/handlers/cli"
"github.com/davecgh/go-spew/spew"
"github.com/pterm/pterm"
"golang.org/x/term"
)
// Level represents log severity levels.
@@ -52,31 +60,109 @@ func (l Level) String() string {
}
}
// callerSkip is the runtime.Caller stack depth from the public Debug
// helpers to the caller of the log package.
const callerSkip = 2
// Escape sequences for colored log lines.
const (
ansiReset = "\x1b[0m"
ansiBold = "\x1b[1m"
ansiRed = "\x1b[31m"
ansiYellow = "\x1b[33m"
ansiBlue = "\x1b[34m"
ansiWhite = "\x1b[37m"
)
//nolint:gochecknoglobals // package-level logger state by design
var (
// mu protects the output writers and level
// mu protects the variables below
mu sync.RWMutex
// stdout is the writer for progress output
stdout io.Writer = os.Stdout
// stderr is the writer for log output
stderr io.Writer = os.Stderr
// styled is false once DisableStyling has been called
styled = true
// logger is the CLI logger Init builds; nil until Init runs
logger *slog.Logger
// currentLevel is our log level (includes Verbose)
currentLevel Level = InfoLevel
currentLevel = InfoLevel
)
// cliHandler is the slog.Handler the CLI logs through. Each record is one
// line: a symbol for its level right-aligned in four columns, then the
// message padded to 25 columns. In color, only the symbol is bold and in the
// level's color; the message is in the terminal's default color. Records are
// filtered by level before they are made, so the handler takes every record
// it is given.
type cliHandler struct {
mu sync.Mutex
w io.Writer
color bool
}
// Enabled reports true for every level.
func (h *cliHandler) Enabled(context.Context, slog.Level) bool {
return true
}
// Handle writes the record as one line.
func (h *cliHandler) Handle(_ context.Context, r slog.Record) error {
symbol, color := "•", ansiBlue
switch {
case r.Level >= slog.LevelError:
symbol, color = "⨯", ansiRed
case r.Level >= slog.LevelWarn:
color = ansiYellow
case r.Level < slog.LevelInfo:
color = ansiWhite
}
h.mu.Lock()
defer h.mu.Unlock()
if !h.color {
_, err := fmt.Fprintf(h.w, "%4s %-25s\n", symbol, r.Message)
return err
}
_, err := fmt.Fprintf(h.w, "%s%s%4s%s %-25s%s\n",
color, ansiBold, symbol, ansiReset, r.Message, ansiReset)
return err
}
// WithAttrs returns the handler unchanged: this package's helpers never
// attach attributes.
func (h *cliHandler) WithAttrs([]slog.Attr) slog.Handler {
return h
}
// WithGroup returns the handler unchanged: this package's helpers never open
// groups.
func (h *cliHandler) WithGroup(string) slog.Handler {
return h
}
// SetOutput configures the output writers for the log package.
// stdout is used for progress output, stderr is used for log messages.
// stdout is used for progress output, stderr is used for log messages
// from the next Init on.
func SetOutput(out, err io.Writer) {
mu.Lock()
defer mu.Unlock()
stdout = out
stderr = err
pterm.SetDefaultOutput(out)
}
// GetStdout returns the configured stdout writer.
func GetStdout() io.Writer {
mu.RLock()
defer mu.RUnlock()
return stdout
}
@@ -84,138 +170,154 @@ func GetStdout() io.Writer {
func GetStderr() io.Writer {
mu.RLock()
defer mu.RUnlock()
return stderr
}
// DisableStyling turns off colors and styling for terminal output.
// DisableStyling turns off colors in log lines from the next Init on.
func DisableStyling() {
pterm.DisableColor()
pterm.DisableStyling()
pterm.Debug.Prefix.Text = ""
pterm.Info.Prefix.Text = ""
pterm.Success.Prefix.Text = ""
pterm.Warning.Prefix.Text = ""
pterm.Error.Prefix.Text = ""
pterm.Fatal.Prefix.Text = ""
mu.Lock()
defer mu.Unlock()
styled = false
}
// Init initializes the logger with the CLI handler and default log level.
// Init sends log records to the CLI handler on the stderr writer given to
// SetOutput. Log lines are colored when stdout is a terminal whose TERM is
// not dumb, unless DisableStyling has been called.
//
// It replaces the logger under the write lock, and log calls hold the read
// lock while they write, so once Init returns nothing writes through the
// logger it replaced.
func Init() {
mu.RLock()
w := stderr
mu.RUnlock()
log.SetHandler(acli.New(w))
log.SetLevel(log.DebugLevel) // Let apex/log pass everything; we filter ourselves
mu.Lock()
defer mu.Unlock()
color := styled && os.Getenv("TERM") != "dumb" &&
term.IsTerminal(int(os.Stdout.Fd()))
logger = slog.New(&cliHandler{w: stderr, color: color})
}
// isEnabled returns true if messages at the given level should be logged.
func isEnabled(l Level) bool {
mu.RLock()
defer mu.RUnlock()
return l >= currentLevel
}
// Fatalf logs a formatted message at fatal level.
func Fatalf(format string, args ...interface{}) {
if isEnabled(FatalLevel) {
log.Fatalf(format, args...)
// logf logs a formatted message at level l if messages at l are enabled,
// holding the read lock while the record is written. slog has no verbose or
// fatal level: verbose messages are logged as info records, and fatal
// messages as error records.
func logf(l Level, format string, args ...any) {
mu.RLock()
defer mu.RUnlock()
if l < currentLevel {
return
}
lg := logger
if lg == nil {
lg = slog.Default()
}
msg := fmt.Sprintf(format, args...)
switch l {
case DebugLevel:
lg.Debug(msg)
case VerboseLevel, InfoLevel:
lg.Info(msg)
case WarnLevel:
lg.Warn(msg)
case ErrorLevel, FatalLevel:
lg.Error(msg)
}
}
// Fatal logs a message at fatal level.
// Fatalf logs a formatted message at fatal level, then exits with status 1.
func Fatalf(format string, args ...any) {
logf(FatalLevel, format, args...)
os.Exit(1)
}
// Fatal logs a message at fatal level, then exits with status 1.
func Fatal(arg string) {
if isEnabled(FatalLevel) {
log.Fatal(arg)
}
logf(FatalLevel, "%s", arg)
os.Exit(1)
}
// Errorf logs a formatted message at error level.
func Errorf(format string, args ...interface{}) {
if isEnabled(ErrorLevel) {
log.Errorf(format, args...)
}
func Errorf(format string, args ...any) {
logf(ErrorLevel, format, args...)
}
// Error logs a message at error level.
func Error(arg string) {
if isEnabled(ErrorLevel) {
log.Error(arg)
}
logf(ErrorLevel, "%s", arg)
}
// Warnf logs a formatted message at warn level.
func Warnf(format string, args ...interface{}) {
if isEnabled(WarnLevel) {
log.Warnf(format, args...)
}
func Warnf(format string, args ...any) {
logf(WarnLevel, format, args...)
}
// Warn logs a message at warn level.
func Warn(arg string) {
if isEnabled(WarnLevel) {
log.Warn(arg)
}
logf(WarnLevel, "%s", arg)
}
// Infof logs a formatted message at info level.
func Infof(format string, args ...interface{}) {
if isEnabled(InfoLevel) {
log.Infof(format, args...)
}
func Infof(format string, args ...any) {
logf(InfoLevel, format, args...)
}
// Info logs a message at info level.
func Info(arg string) {
if isEnabled(InfoLevel) {
log.Info(arg)
}
logf(InfoLevel, "%s", arg)
}
// Verbosef logs a formatted message at verbose level.
func Verbosef(format string, args ...interface{}) {
if isEnabled(VerboseLevel) {
log.Infof(format, args...)
}
func Verbosef(format string, args ...any) {
logf(VerboseLevel, format, args...)
}
// Verbose logs a message at verbose level.
func Verbose(arg string) {
if isEnabled(VerboseLevel) {
log.Info(arg)
}
logf(VerboseLevel, "%s", arg)
}
// Debugf logs a formatted message at debug level with caller location.
func Debugf(format string, args ...interface{}) {
func Debugf(format string, args ...any) {
if isEnabled(DebugLevel) {
DebugReal(fmt.Sprintf(format, args...), 2)
DebugReal(fmt.Sprintf(format, args...), callerSkip)
}
}
// Debug logs a message at debug level with caller location.
func Debug(arg string) {
if isEnabled(DebugLevel) {
DebugReal(arg, 2)
DebugReal(arg, callerSkip)
}
}
// DebugReal logs at debug level with caller info from the specified stack depth.
func DebugReal(arg string, cs int) {
if !isEnabled(DebugLevel) {
return
}
_, callerFile, callerLine, ok := runtime.Caller(cs)
if !ok {
return
}
tag := fmt.Sprintf("%s:%d: ", filepath.Base(callerFile), callerLine)
log.Debug(tag + arg)
logf(DebugLevel, "%s:%d: %s", filepath.Base(callerFile), callerLine, arg)
}
// Dump logs a spew dump of the arguments at debug level.
func Dump(args ...interface{}) {
func Dump(args ...any) {
if isEnabled(DebugLevel) {
DebugReal(spew.Sdump(args...), 2)
DebugReal(spew.Sdump(args...), callerSkip)
}
}
@@ -246,6 +348,7 @@ func SetLevelFromVerbosity(l int) {
func SetLevel(l Level) {
mu.Lock()
defer mu.Unlock()
currentLevel = l
}
@@ -253,22 +356,25 @@ func SetLevel(l Level) {
func GetLevel() Level {
mu.RLock()
defer mu.RUnlock()
return currentLevel
}
// WithError returns a log entry with the error attached.
func WithError(e error) *log.Entry {
return log.Log.WithError(e)
return currentLevel
}
// Progressf prints a progress message that overwrites the current line.
// Use ProgressDone() when progress is complete to move to the next line.
func Progressf(format string, args ...interface{}) {
pterm.Printf("\r"+format, args...)
// Progress goes to the stdout writer whatever the log level.
func Progressf(format string, args ...any) {
mu.Lock()
defer mu.Unlock()
_, _ = fmt.Fprintf(stdout, "\r"+format, args...)
}
// ProgressDone clears the progress line when progress is complete.
func ProgressDone() {
// Clear the line with spaces and return to beginning
pterm.Print("\r\033[K")
mu.Lock()
defer mu.Unlock()
// Return to the start of the line and erase it
_, _ = fmt.Fprint(stdout, "\r\033[K")
}
+218 -1
View File
@@ -1,12 +1,229 @@
//nolint:testpackage // white-box tests exercise unexported internals
package log
import (
"bytes"
stdlog "log"
"log/slog"
"os"
"testing"
"github.com/stretchr/testify/assert"
)
func TestBuild(t *testing.T) {
t.Parallel()
Init()
assert.True(t, true)
}
// capture points the package's writers at fresh buffers and sets its level,
// with colors off, then restores the default writers and level when the test
// ends. The state it changes is process-wide, so tests that use it do not
// run in parallel.
func capture(t *testing.T, l Level) (*bytes.Buffer, *bytes.Buffer) {
t.Helper()
var stdout, stderr bytes.Buffer
DisableStyling()
SetOutput(&stdout, &stderr)
SetLevel(l)
Init()
t.Cleanup(func() {
SetOutput(os.Stdout, os.Stderr)
SetLevel(InfoLevel)
Init()
})
return &stdout, &stderr
}
// TestLevelFiltering logs at every level, through both the plain and the
// formatting helpers, and checks that exactly the messages at or above the
// set level are written.
//
//nolint:paralleltest // changes the package's process-wide writers and level
func TestLevelFiltering(t *testing.T) {
messages := []struct {
level Level
text string
}{
{DebugLevel, "debug plain"},
{DebugLevel, "debug formatted"},
{VerboseLevel, "verbose plain"},
{VerboseLevel, "verbose formatted"},
{InfoLevel, "info plain"},
{InfoLevel, "info formatted"},
{WarnLevel, "warn plain"},
{WarnLevel, "warn formatted"},
{ErrorLevel, "error plain"},
{ErrorLevel, "error formatted"},
}
levels := []Level{DebugLevel, VerboseLevel, InfoLevel, WarnLevel, ErrorLevel}
for _, set := range levels {
t.Run(set.String(), func(t *testing.T) {
stdout, stderr := capture(t, set)
Debug("debug plain")
Debugf("debug %s", "formatted")
Verbose("verbose plain")
Verbosef("verbose %s", "formatted")
Info("info plain")
Infof("info %s", "formatted")
Warn("warn plain")
Warnf("warn %s", "formatted")
Error("error plain")
Errorf("error %s", "formatted")
for _, m := range messages {
if m.level >= set {
assert.Contains(t, stderr.String(), m.text)
} else {
assert.NotContains(t, stderr.String(), m.text)
}
}
assert.Empty(t, stdout.String())
})
}
}
// TestLineFormat checks the exact uncolored lines: a symbol per level
// right-aligned in four columns, then the message padded to 25 columns.
//
//nolint:paralleltest // changes the package's process-wide writers and level
func TestLineFormat(t *testing.T) {
_, stderr := capture(t, VerboseLevel)
Infof("scanning filesystem...")
Verbosef("+ %s (%s)", "a.txt", "1 B")
Warn("short")
Errorf("%s", "a message longer than twenty-five columns")
want := " • scanning filesystem... \n" +
" • + a.txt (1 B) \n" +
" • short \n" +
" ⨯ a message longer than twenty-five columns\n"
assert.Equal(t, want, stderr.String())
}
// TestDebugCallerTag checks that debug lines start with the file and line
// of the call that logged them.
//
//nolint:paralleltest // changes the package's process-wide writers and level
func TestDebugCallerTag(t *testing.T) {
_, stderr := capture(t, DebugLevel)
Debugf("enumerating path: %s", "/tmp")
assert.Regexp(t, `^ • log_test\.go:\d+: enumerating path: /tmp\s*\n$`,
stderr.String())
}
// TestColoredLine checks the colored lines: only the symbol is bold and in the
// level's color; the message is in the terminal's default color.
func TestColoredLine(t *testing.T) {
t.Parallel()
var buf bytes.Buffer
lg := slog.New(&cliHandler{w: &buf, color: true})
lg.Debug("debug")
lg.Info("info")
lg.Warn("warn")
lg.Error("error")
want := "\x1b[37m\x1b[1m •\x1b[0m debug \x1b[0m\n" +
"\x1b[34m\x1b[1m •\x1b[0m info \x1b[0m\n" +
"\x1b[33m\x1b[1m •\x1b[0m warn \x1b[0m\n" +
"\x1b[31m\x1b[1m ⨯\x1b[0m error \x1b[0m\n"
assert.Equal(t, want, buf.String())
}
// TestRecordLevels checks the slog level of the record each helper logs,
// through slog's text handler with the time left out. slog has no verbose
// level, so verbose messages are info records.
//
//nolint:paralleltest // changes the package's process-wide logger and level
func TestRecordLevels(t *testing.T) {
var buf bytes.Buffer
capture(t, DebugLevel)
opts := &slog.HandlerOptions{
Level: slog.LevelDebug,
ReplaceAttr: func(_ []string, a slog.Attr) slog.Attr {
if a.Key == slog.TimeKey {
return slog.Attr{}
}
return a
},
}
mu.Lock()
logger = slog.New(slog.NewTextHandler(&buf, opts))
mu.Unlock()
Debugf("debug")
Verbosef("verbose")
Infof("info")
Warnf("warn")
Errorf("error")
assert.Regexp(t, `^level=DEBUG msg="log_test\.go:\d+: debug"\n`+
"level=INFO msg=verbose\n"+
"level=INFO msg=info\n"+
"level=WARN msg=warn\n"+
"level=ERROR msg=error\n$", buf.String())
}
// TestProgress checks that progress lines go to the stdout writer whatever
// the log level, each starting with a carriage return so it overwrites the
// last, and that ProgressDone erases the line.
//
//nolint:paralleltest // changes the package's process-wide writers and level
func TestProgress(t *testing.T) {
stdout, stderr := capture(t, ErrorLevel)
Progressf("Scanning: %d files found", 1000)
Progressf("Scanning: %d files found", 2000)
ProgressDone()
assert.Equal(t,
"\rScanning: 1000 files found\rScanning: 2000 files found\r\x1b[K",
stdout.String())
assert.Empty(t, stderr.String())
}
// TestBeforeInit checks that records go to slog.Default() until Init runs,
// still filtered by the package's level. slog's default handler writes
// through the standard library's log package.
//
//nolint:paralleltest // changes the package's process-wide logger
func TestBeforeInit(t *testing.T) {
var buf bytes.Buffer
flags := stdlog.Flags()
stdlog.SetOutput(&buf)
stdlog.SetFlags(0)
mu.Lock()
logger = nil
mu.Unlock()
t.Cleanup(func() {
stdlog.SetOutput(os.Stderr)
stdlog.SetFlags(flags)
Init()
})
Infof("loaded manifest with %d files", 3)
Verbose("not shown at info level")
assert.Equal(t, "INFO loaded manifest with 3 files\n", buf.String())
}
+123 -46
View File
@@ -1,6 +1,9 @@
// Package mfer implements the mfer manifest file format: building,
// serializing, verifying, and checking manifests of file trees.
package mfer
import (
"context"
"crypto/sha256"
"errors"
"fmt"
@@ -14,6 +17,28 @@ import (
"github.com/multiformats/go-multihash"
)
// readChunkSize is the buffer size used when reading file contents for
// hashing.
const readChunkSize = 64 * 1024
// The errPath* sentinels below are worded as the trailing fragment of the
// message ValidatePath renders, because the offending path is quoted
// before them (`path %q ...`). Wrapping them mid-sentence keeps the
// rendered text exactly as mfer has always printed it. Match them with
// errors.Is rather than by reading their messages.
var (
errPathEmpty = errors.New("path cannot be empty")
errPathNotUTF8 = errors.New("is not valid UTF-8")
errPathBackslash = errors.New("contains backslash; use forward slashes only")
errPathAbsolute = errors.New("is absolute; must be relative")
errPathEmptySegment = errors.New("contains empty segment")
errPathDotDot = errors.New("contains '..' segment")
errSizeMismatch = errors.New("size mismatch")
errNegativeSize = errors.New("size cannot be negative")
errHashNotMultihash = errors.New("hash is not a valid multihash")
errHashTooShort = errors.New("hash digest is too short")
)
// ValidatePath checks that a file path conforms to manifest path invariants:
// - Must be valid UTF-8
// - Must use forward slashes only (no backslashes)
@@ -23,25 +48,31 @@ import (
// - Must not be empty
func ValidatePath(p string) error {
if p == "" {
return errors.New("path cannot be empty")
return errPathEmpty
}
if !utf8.ValidString(p) {
return fmt.Errorf("path %q is not valid UTF-8", p)
return fmt.Errorf("path %q %w", p, errPathNotUTF8)
}
if strings.ContainsRune(p, '\\') {
return fmt.Errorf("path %q contains backslash; use forward slashes only", p)
return fmt.Errorf("path %q %w", p, errPathBackslash)
}
if strings.HasPrefix(p, "/") {
return fmt.Errorf("path %q is absolute; must be relative", p)
return fmt.Errorf("path %q %w", p, errPathAbsolute)
}
for _, seg := range strings.Split(p, "/") {
if seg == "" {
return fmt.Errorf("path %q contains empty segment", p)
return fmt.Errorf("path %q %w", p, errPathEmptySegment)
}
if seg == ".." {
return fmt.Errorf("path %q contains '..' segment", p)
return fmt.Errorf("path %q %w", p, errPathDotDot)
}
}
return nil
}
@@ -68,11 +99,7 @@ type UnixNanos int32
// Timestamp converts ModTime to a protobuf Timestamp.
func (m ModTime) Timestamp() *Timestamp {
t := time.Time(m)
return &Timestamp{
Seconds: t.Unix(),
Nanos: int32(t.Nanosecond()),
}
return newTimestampFromTime(time.Time(m))
}
// Multihash represents a multihash-encoded file hash (typically SHA2-256).
@@ -85,19 +112,12 @@ type FileHashProgress struct {
// Builder constructs a manifest by adding files one at a time.
type Builder struct {
mu sync.Mutex
files []*MFFilePath
createdAt time.Time
signingOptions *SigningOptions
fixedUUID []byte // if set, use this UUID instead of generating one
}
// SetSeed derives a deterministic UUID from the given seed string.
// The seed is hashed once with SHA-256 and the first 16 bytes are used
// as a fixed UUID for the manifest.
func (b *Builder) SetSeed(seed string) {
hash := sha256.Sum256([]byte(seed))
b.fixedUUID = hash[:16]
mu sync.Mutex
files []*MFFilePath
createdAt time.Time
includeTimestamps bool
signingOptions *SigningOptions
fixedUUID []byte // if set, use this UUID instead of generating one
}
// NewBuilder creates a new Builder.
@@ -108,6 +128,14 @@ func NewBuilder() *Builder {
}
}
// SetSeed derives a deterministic UUID from the given seed string.
// The seed is hashed once with SHA-256 and the first 16 bytes are used
// as a fixed UUID for the manifest.
func (b *Builder) SetSeed(seed string) {
hash := sha256.Sum256([]byte(seed))
b.fixedUUID = hash[:uuidLength]
}
// AddFile reads file content from reader, computes hashes, and adds to manifest.
// Progress updates are sent to the progress channel (if non-nil) without blocking.
// Returns the number of bytes read.
@@ -118,7 +146,8 @@ func (b *Builder) AddFile(
reader io.Reader,
progress chan<- FileHashProgress,
) (FileSize, error) {
if err := ValidatePath(string(path)); err != nil {
err := ValidatePath(string(path))
if err != nil {
return 0, err
}
@@ -127,7 +156,8 @@ func (b *Builder) AddFile(
// Read file in chunks, updating hash and progress
var totalRead FileSize
buf := make([]byte, 64*1024) // 64KB chunks
buf := make([]byte, readChunkSize)
for {
n, err := reader.Read(buf)
@@ -136,9 +166,11 @@ func (b *Builder) AddFile(
totalRead += FileSize(n)
sendFileHashProgress(progress, FileHashProgress{BytesRead: totalRead})
}
if err == io.EOF {
break
}
if err != nil {
return totalRead, err
}
@@ -146,7 +178,10 @@ func (b *Builder) AddFile(
// Verify actual bytes read matches declared size
if totalRead != size {
return totalRead, fmt.Errorf("size mismatch for %q: declared %d bytes but read %d bytes", path, size, totalRead)
return totalRead, fmt.Errorf(
"%w for %q: declared %d bytes but read %d bytes",
errSizeMismatch, path, size, totalRead,
)
}
// Encode hash as multihash (SHA2-256)
@@ -177,6 +212,7 @@ func sendFileHashProgress(ch chan<- FileHashProgress, p FileHashProgress) {
if ch == nil {
return
}
select {
case ch <- p:
default:
@@ -187,21 +223,42 @@ func sendFileHashProgress(ch chan<- FileHashProgress, p FileHashProgress) {
func (b *Builder) FileCount() int {
b.mu.Lock()
defer b.mu.Unlock()
return len(b.files)
}
// AddFileWithHash adds a file entry with a pre-computed hash.
// This is useful when the hash is already known (e.g., from an existing manifest).
// Returns an error if path is empty, size is negative, or hash is nil/empty.
func (b *Builder) AddFileWithHash(path RelFilePath, size FileSize, mtime ModTime, hash Multihash) error {
if err := ValidatePath(string(path)); err != nil {
return err
// Returns an error if path is invalid, size is negative, or hash is not a
// multihash with a digest of at least 32 bytes, as long as SHA-256's.
func (b *Builder) AddFileWithHash(
path RelFilePath,
size FileSize,
mtime ModTime,
hash Multihash,
) error {
err := ValidatePath(string(path))
if err != nil {
return fmt.Errorf("add file: %w", err)
}
if size < 0 {
return errors.New("size cannot be negative")
return errNegativeSize
}
if len(hash) == 0 {
return errors.New("hash cannot be nil or empty")
decoded, err := multihash.Decode(hash)
if err != nil {
return fmt.Errorf("%w: %w", errHashNotMultihash, err)
}
// The reader's limit on decoding cost (maxDecodedGrowth) assumes every
// hash is at least as long as a SHA-256 multihash, so a manifest of
// shorter ones could fail to load.
if len(decoded.Digest) < sha256.Size {
return fmt.Errorf(
"%w: %d bytes, at least %d needed",
errHashTooShort, len(decoded.Digest), sha256.Size,
)
}
entry := &MFFilePath{
@@ -216,32 +273,46 @@ func (b *Builder) AddFileWithHash(path RelFilePath, size FileSize, mtime ModTime
b.mu.Lock()
b.files = append(b.files, entry)
b.mu.Unlock()
return nil
}
// SetIncludeTimestamps controls whether the manifest includes a createdAt timestamp.
// By default timestamps are omitted for deterministic output.
func (b *Builder) SetIncludeTimestamps(include bool) {
b.mu.Lock()
defer b.mu.Unlock()
b.includeTimestamps = include
}
// SetSigningOptions sets the GPG signing options for the manifest.
// If opts is non-nil, the manifest will be signed when Build() is called.
func (b *Builder) SetSigningOptions(opts *SigningOptions) {
b.mu.Lock()
defer b.mu.Unlock()
b.signingOptions = opts
}
// Build finalizes the manifest and writes it to the writer.
func (b *Builder) Build(w io.Writer) error {
// Build finalizes the manifest and writes it to the writer. ctx bounds the
// gpg runs that sign the manifest when signing options are set.
func (b *Builder) Build(ctx context.Context, w io.Writer) error {
b.mu.Lock()
defer b.mu.Unlock()
// Sort files by path for deterministic output
sort.Slice(b.files, func(i, j int) bool {
return b.files[i].Path < b.files[j].Path
return b.files[i].GetPath() < b.files[j].GetPath()
})
// Create inner manifest
inner := &MFFile{
Version: MFFile_VERSION_ONE,
CreatedAt: newTimestampFromTime(b.createdAt),
Files: b.files,
Version: MFFile_VERSION_ONE,
Files: b.files,
}
if b.includeTimestamps {
inner.CreatedAt = newTimestampFromTime(b.createdAt)
}
// Create a temporary manifest to use existing serialization
@@ -252,16 +323,22 @@ func (b *Builder) Build(w io.Writer) error {
}
// Generate outer wrapper
if err := m.generateOuter(); err != nil {
return err
err := m.generateOuter(ctx)
if err != nil {
return fmt.Errorf("build: generate outer: %w", err)
}
// Generate final output
if err := m.generate(); err != nil {
return err
err = m.generate(ctx)
if err != nil {
return fmt.Errorf("build: generate: %w", err)
}
// Write to output
_, err := w.Write(m.output.Bytes())
return err
_, err = w.Write(m.output.Bytes())
if err != nil {
return fmt.Errorf("build: write output: %w", err)
}
return nil
}
+422 -36
View File
@@ -1,91 +1,145 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"bytes"
"context"
"crypto/sha256"
"fmt"
"path/filepath"
"strings"
"testing"
"time"
"github.com/multiformats/go-multihash"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
const testFileName = "file.txt"
func TestNewBuilder(t *testing.T) {
t.Parallel()
b := NewBuilder()
assert.NotNil(t, b)
assert.Equal(t, 0, b.FileCount())
}
func TestBuilderAddFile(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("test content")
reader := bytes.NewReader(content)
bytesRead, err := b.AddFile("test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil)
bytesRead, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
require.NoError(t, err)
assert.Equal(t, FileSize(len(content)), bytesRead)
assert.Equal(t, 1, b.FileCount())
}
func TestBuilderAddFileWithHash(t *testing.T) {
b := NewBuilder()
hash := make([]byte, 34) // SHA256 multihash is 34 bytes
t.Parallel()
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), hash)
b := NewBuilder()
hash, err := multihash.Encode(make([]byte, sha256.Size), multihash.SHA2_256)
require.NoError(t, err)
err = b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), hash)
require.NoError(t, err)
assert.Equal(t, 1, b.FileCount())
}
func TestBuilderAddFileWithHashValidation(t *testing.T) {
t.Parallel()
sha256Hash, err := multihash.Encode(make([]byte, sha256.Size), multihash.SHA2_256)
require.NoError(t, err)
t.Run("empty path", func(t *testing.T) {
t.Parallel()
b := NewBuilder()
hash := make([]byte, 34)
err := b.AddFileWithHash("", 100, ModTime(time.Now()), hash)
assert.Error(t, err)
err := b.AddFileWithHash("", 100, ModTime(time.Now()), sha256Hash)
require.Error(t, err)
assert.Contains(t, err.Error(), "path")
})
t.Run("negative size", func(t *testing.T) {
t.Parallel()
b := NewBuilder()
hash := make([]byte, 34)
err := b.AddFileWithHash("test.txt", -1, ModTime(time.Now()), hash)
assert.Error(t, err)
err := b.AddFileWithHash("test.txt", -1, ModTime(time.Now()), sha256Hash)
require.Error(t, err)
assert.Contains(t, err.Error(), "size")
})
t.Run("nil hash", func(t *testing.T) {
b := NewBuilder()
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), nil)
assert.Error(t, err)
assert.Contains(t, err.Error(), "hash")
})
t.Run("empty hash", func(t *testing.T) {
b := NewBuilder()
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), []byte{})
assert.Error(t, err)
assert.Contains(t, err.Error(), "hash")
})
t.Run("valid inputs", func(t *testing.T) {
t.Parallel()
b := NewBuilder()
hash := make([]byte, 34)
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), hash)
assert.NoError(t, err)
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), sha256Hash)
require.NoError(t, err)
assert.Equal(t, 1, b.FileCount())
})
}
func TestBuilderAddFileWithHashRejectsBadHashes(t *testing.T) {
t.Parallel()
sha1Hash, err := multihash.Encode(make([]byte, 20), multihash.SHA1)
require.NoError(t, err)
truncatedHash, err := multihash.Encode(make([]byte, sha256.Size-1), multihash.SHA2_256)
require.NoError(t, err)
tests := []struct {
name string
hash Multihash
want error
}{
{"nil hash", nil, errHashNotMultihash},
{"empty hash", []byte{}, errHashNotMultihash},
{"one-byte hash", []byte{0x12}, errHashNotMultihash},
// A SHA-256 code and 32-byte length, then only two bytes of digest.
{"malformed multihash", []byte{0x12, 0x20, 0x01, 0x02}, errHashNotMultihash},
// A valid multihash, but its 20-byte SHA-1 digest is too short.
{"SHA-1 multihash", sha1Hash, errHashTooShort},
// A valid multihash whose 31-byte digest is one byte short.
{"31-byte digest", truncatedHash, errHashTooShort},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
b := NewBuilder()
err := b.AddFileWithHash("test.txt", 100, ModTime(time.Now()), tt.hash)
require.ErrorIs(t, err, tt.want)
assert.Equal(t, 0, b.FileCount())
})
}
}
func TestBuilderBuild(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("test content")
reader := bytes.NewReader(content)
_, err := b.AddFile("test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil)
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
require.NoError(t, err)
var buf bytes.Buffer
err = b.Build(&buf)
err = b.Build(context.Background(), &buf)
require.NoError(t, err)
// Should have magic bytes
@@ -93,6 +147,8 @@ func TestBuilderBuild(t *testing.T) {
}
func TestNewTimestampFromTimeExtremeDate(t *testing.T) {
t.Parallel()
// Regression test: newTimestampFromTime used UnixNano() which panics
// for dates outside ~1678-2262. Now uses Nanosecond() which is safe.
tests := []struct {
@@ -107,15 +163,19 @@ func TestNewTimestampFromTimeExtremeDate(t *testing.T) {
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
// Should not panic
ts := newTimestampFromTime(tt.time)
assert.Equal(t, tt.time.Unix(), ts.Seconds)
assert.Equal(t, int32(tt.time.Nanosecond()), ts.Nanos)
assert.Equal(t, tt.time.Unix(), ts.GetSeconds())
assert.Equal(t, tt.time.Nanosecond(), int(ts.GetNanos()))
})
}
}
func TestBuilderDeterministicOutput(t *testing.T) {
t.Parallel()
buildManifest := func() []byte {
b := NewBuilder()
// Use a fixed createdAt and UUID so output is reproducible
@@ -135,24 +195,32 @@ func TestBuilderDeterministicOutput(t *testing.T) {
}
for _, f := range files {
r := bytes.NewReader([]byte(f.content))
_, err := b.AddFile(RelFilePath(f.path), FileSize(len(f.content)), mtime, r, nil)
_, err := b.AddFile(
RelFilePath(f.path), FileSize(len(f.content)), mtime, r, nil,
)
require.NoError(t, err)
}
var buf bytes.Buffer
err := b.Build(&buf)
err := b.Build(context.Background(), &buf)
require.NoError(t, err)
return buf.Bytes()
}
out1 := buildManifest()
out2 := buildManifest()
assert.Equal(t, out1, out2, "two builds with same input should produce byte-identical output")
assert.Equal(t, out1, out2,
"two builds with same input should produce byte-identical output")
}
func TestSetSeedDeterministic(t *testing.T) {
t.Parallel()
b1 := NewBuilder()
b1.SetSeed("test-seed-value")
b2 := NewBuilder()
b2.SetSeed("test-seed-value")
assert.Equal(t, b1.fixedUUID, b2.fixedUUID, "same seed should produce same UUID")
@@ -160,16 +228,334 @@ func TestSetSeedDeterministic(t *testing.T) {
b3 := NewBuilder()
b3.SetSeed("different-seed")
assert.NotEqual(t, b1.fixedUUID, b3.fixedUUID, "different seeds should produce different UUIDs")
assert.NotEqual(t, b1.fixedUUID, b3.fixedUUID,
"different seeds should produce different UUIDs")
}
func TestValidatePath(t *testing.T) {
t.Parallel()
valid := []string{
testFileName,
"dir/file.txt",
"a/b/c/d.txt",
"file with spaces.txt",
"日本語.txt", //nolint:gosmopolitan // deliberately tests non-ASCII UTF-8 paths
}
for _, p := range valid {
t.Run("valid:"+p, func(t *testing.T) {
t.Parallel()
assert.NoError(t, ValidatePath(p))
})
}
invalid := []struct {
path string
desc string
}{
{"", "empty"},
{"/absolute", "absolute path"},
{"has\\backslash", "backslash"},
{"has/../traversal", "dot-dot segment"},
{"has//double", "empty segment"},
{"..", "just dot-dot"},
{string([]byte{0xff, 0xfe}), "invalid UTF-8"},
}
for _, tt := range invalid {
t.Run("invalid:"+tt.desc, func(t *testing.T) {
t.Parallel()
assert.Error(t, ValidatePath(tt.path))
})
}
}
func TestBuilderAddFileSizeMismatch(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("short")
reader := bytes.NewReader(content)
// Declare wrong size
_, err := b.AddFile("test.txt", FileSize(100), ModTime(time.Now()), reader, nil)
require.Error(t, err)
assert.Contains(t, err.Error(), "size mismatch")
}
func TestBuilderAddFileInvalidPath(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("data")
reader := bytes.NewReader(content)
_, err := b.AddFile("", FileSize(len(content)), ModTime(time.Now()), reader, nil)
require.Error(t, err)
reader.Reset(content)
_, err = b.AddFile(
"/absolute", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
assert.Error(t, err)
}
func TestBuilderAddFileWithProgress(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := bytes.Repeat([]byte("x"), 1000)
reader := bytes.NewReader(content)
progress := make(chan FileHashProgress, 100)
bytesRead, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, progress,
)
close(progress)
require.NoError(t, err)
assert.Equal(t, FileSize(1000), bytesRead)
var updates []FileHashProgress
for p := range progress {
updates = append(updates, p)
}
assert.NotEmpty(t, updates)
// Last update should show all bytes
assert.Equal(t, FileSize(1000), updates[len(updates)-1].BytesRead)
}
func TestBuilderBuildRoundTrip(t *testing.T) {
t.Parallel()
// Build a manifest, deserialize it, verify all fields survive round-trip
b := NewBuilder()
now := time.Date(2025, 6, 15, 12, 0, 0, 0, time.UTC)
files := []struct {
path string
content []byte
}{
{"alpha.txt", []byte("alpha content")},
{"beta/gamma.txt", []byte("gamma content")},
{"beta/delta.txt", []byte("delta content")},
}
for _, f := range files {
reader := bytes.NewReader(f.content)
_, err := b.AddFile(
RelFilePath(f.path), FileSize(len(f.content)), ModTime(now), reader, nil,
)
require.NoError(t, err)
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
mfiles := m.Files()
require.Len(t, mfiles, 3)
// Verify sorted order
assert.Equal(t, "alpha.txt", mfiles[0].GetPath())
assert.Equal(t, "beta/delta.txt", mfiles[1].GetPath())
assert.Equal(t, "beta/gamma.txt", mfiles[2].GetPath())
// Verify sizes
assert.Equal(t, int64(len("alpha content")), mfiles[0].GetSize())
// Verify hashes are present
for _, f := range mfiles {
require.NotEmpty(t, f.GetHashes(), "file %s should have hashes", f.GetPath())
assert.NotEmpty(t, f.GetHashes()[0].GetMultiHash())
}
}
// A manifest whose payload is larger than one zstd block (128 KiB) is
// written as a frame asking for the writer's whole window, zstdWindowSize,
// which is the most the parser accepts.
func TestBuilderBuildRoundTripLargeManifest(t *testing.T) {
t.Parallel()
hash, err := multihash.Encode(make([]byte, sha256.Size), multihash.SHA2_256)
require.NoError(t, err)
b := NewBuilder()
for i := range 4000 {
path := RelFilePath(fmt.Sprintf("dir/file-%05d.txt", i))
require.NoError(t, b.AddFileWithHash(path, FileSize(i), ModTime{}, hash))
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
require.Greater(t, m.pbOuter.GetSize(), int64(128<<10))
assert.Len(t, m.Files(), 4000)
}
func TestNewManifestFromReaderInvalidMagic(t *testing.T) {
t.Parallel()
_, err := NewManifestFromReader(bytes.NewReader([]byte("NOT_VALID")))
require.Error(t, err)
assert.Contains(t, err.Error(), "invalid file format")
}
func TestNewManifestFromReaderEmpty(t *testing.T) {
t.Parallel()
_, err := NewManifestFromReader(bytes.NewReader([]byte{}))
assert.Error(t, err)
}
func TestNewManifestFromReaderTruncated(t *testing.T) {
t.Parallel()
// Just the magic with nothing after
_, err := NewManifestFromReader(bytes.NewReader([]byte(MAGIC)))
assert.Error(t, err)
}
func TestNewManifestFromFileRequiresPath(t *testing.T) {
t.Parallel()
_, err := NewManifestFromFile(nil)
require.ErrorIs(t, err, errManifestPathEmpty)
_, err = NewManifestFromFile(&ManifestFromFileOptions{Fs: afero.NewMemMapFs()})
require.ErrorIs(t, err, errManifestPathEmpty)
}
func TestNewManifestFromFileNilFsUsesOsFs(t *testing.T) {
t.Parallel()
path := filepath.Join(t.TempDir(), "index.mf")
createTestManifest(t, afero.NewOsFs(), path, map[string][]byte{
testFileName: []byte("hello"),
})
m, err := NewManifestFromFile(&ManifestFromFileOptions{Path: path})
require.NoError(t, err)
assert.Len(t, m.Files(), 1)
}
func TestManifestString(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("test")
reader := bytes.NewReader(content)
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
require.NoError(t, err)
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.Contains(t, m.String(), "count=1")
}
func TestBuilderBuildEmpty(t *testing.T) {
t.Parallel()
b := NewBuilder()
var buf bytes.Buffer
err := b.Build(&buf)
err := b.Build(context.Background(), &buf)
require.NoError(t, err)
// Should still produce valid manifest with 0 files
assert.True(t, strings.HasPrefix(buf.String(), MAGIC))
}
func TestBuilderOmitsCreatedAtByDefault(t *testing.T) {
t.Parallel()
b := NewBuilder()
content := []byte("hello")
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()),
bytes.NewReader(content), nil,
)
require.NoError(t, err)
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.Nil(t, m.pbInner.GetCreatedAt(),
"createdAt should be nil by default for deterministic output")
}
func TestBuilderIncludesCreatedAtWhenRequested(t *testing.T) {
t.Parallel()
b := NewBuilder()
b.SetIncludeTimestamps(true)
content := []byte("hello")
_, err := b.AddFile(
"test.txt", FileSize(len(content)), ModTime(time.Now()),
bytes.NewReader(content), nil,
)
require.NoError(t, err)
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.NotNil(t, m.pbInner.GetCreatedAt(),
"createdAt should be set when IncludeTimestamps is true")
}
func TestBuilderDeterministicFileOrder(t *testing.T) {
t.Parallel()
// Two builds with same files in different order should produce same file ordering.
// Note: UUIDs differ per build, so we compare parsed file lists, not raw bytes.
buildAndParse := func(order []string) []*MFFilePath {
b := NewBuilder()
for _, name := range order {
content := []byte("content of " + name)
_, err := b.AddFile(
RelFilePath(name), FileSize(len(content)),
ModTime(time.Unix(1000, 0)), bytes.NewReader(content), nil,
)
require.NoError(t, err)
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
return m.Files()
}
files1 := buildAndParse([]string{"b.txt", "a.txt"})
files2 := buildAndParse([]string{"a.txt", "b.txt"})
require.Len(t, files1, 2)
require.Len(t, files2, 2)
for i := range files1 {
assert.Equal(t, files1[i].GetPath(), files2[i].GetPath())
assert.Equal(t, files1[i].GetSize(), files2[i].GetSize())
}
assert.Equal(t, "a.txt", files1[0].GetPath())
assert.Equal(t, "b.txt", files1[1].GetPath())
}
+202 -111
View File
@@ -14,6 +14,12 @@ import (
"github.com/spf13/afero"
)
var (
errNoSigningPubKey = errors.New("manifest has no signing public key")
errManifestPathEmpty = errors.New("manifest path cannot be empty")
errBasePathEmpty = errors.New("base path cannot be empty")
)
// Result represents the outcome of checking a single file.
type Result struct {
Path RelFilePath // Relative path from manifest
@@ -24,6 +30,7 @@ type Result struct {
// Status represents the verification status of a file.
type Status int
// Verification result statuses reported for each checked file.
const (
StatusOK Status = iota // File matches manifest (size and hash verified)
StatusMissing // File not found on disk
@@ -70,34 +77,65 @@ type Checker struct {
fs afero.Fs
// manifestPaths is a set of paths in the manifest for quick lookup
manifestPaths map[RelFilePath]struct{}
// manifestInfo is the manifest file, which FindExtraFiles leaves out,
// matched with os.SameFile.
manifestInfo os.FileInfo
// signature info from the manifest
signature []byte
signer []byte
signingPubKey []byte
}
// NewChecker creates a new Checker for the given manifest, base path, and filesystem.
// The basePath is the directory relative to which manifest paths are resolved.
// If fs is nil, the real filesystem (OsFs) is used.
func NewChecker(manifestPath string, basePath string, fs afero.Fs) (*Checker, error) {
// CheckerOptions configures a Checker.
type CheckerOptions struct {
// ManifestPath is the manifest file to check against (required).
ManifestPath string
// BasePath is the directory relative to which manifest paths are
// resolved (required).
BasePath string
// Fs is the filesystem to use, defaults to OsFs if nil.
Fs afero.Fs
}
// NewChecker creates a new Checker with the given options. It returns an
// error if opts is nil or either path is empty.
func NewChecker(opts *CheckerOptions) (*Checker, error) {
if opts == nil || opts.ManifestPath == "" {
return nil, errManifestPathEmpty
}
if opts.BasePath == "" {
return nil, errBasePathEmpty
}
fs := opts.Fs
if fs == nil {
fs = afero.NewOsFs()
}
m, err := NewManifestFromFile(fs, manifestPath)
m, err := NewManifestFromFile(&ManifestFromFileOptions{
Path: opts.ManifestPath,
Fs: fs,
})
if err != nil {
return nil, err
}
abs, err := filepath.Abs(basePath)
abs, err := filepath.Abs(opts.BasePath)
if err != nil {
return nil, err
}
files := m.Files()
manifestPaths := make(map[RelFilePath]struct{}, len(files))
for _, f := range files {
manifestPaths[RelFilePath(f.Path)] = struct{}{}
manifestPaths[RelFilePath(f.GetPath())] = struct{}{}
}
manifestInfo, err := fs.Stat(opts.ManifestPath)
if err != nil {
return nil, err
}
return &Checker{
@@ -105,9 +143,10 @@ func NewChecker(manifestPath string, basePath string, fs afero.Fs) (*Checker, er
files: files,
fs: fs,
manifestPaths: manifestPaths,
signature: m.pbOuter.Signature,
signer: m.pbOuter.Signer,
signingPubKey: m.pbOuter.SigningPubKey,
manifestInfo: manifestInfo,
signature: m.pbOuter.GetSignature(),
signer: m.pbOuter.GetSigner(),
signingPubKey: m.pbOuter.GetSigningPubKey(),
}, nil
}
@@ -120,8 +159,9 @@ func (c *Checker) FileCount() FileCount {
func (c *Checker) TotalBytes() FileSize {
var total FileSize
for _, f := range c.files {
total += FileSize(f.Size)
total += FileSize(f.GetSize())
}
return total
}
@@ -135,7 +175,8 @@ func (c *Checker) Signer() []byte {
return c.signer
}
// SigningPubKey returns the signing public key if the manifest is signed, nil otherwise.
// SigningPubKey returns the signing public key if the manifest is signed,
// nil otherwise.
func (c *Checker) SigningPubKey() []byte {
return c.signingPubKey
}
@@ -143,21 +184,27 @@ func (c *Checker) SigningPubKey() []byte {
// ExtractEmbeddedSigningKeyFP imports the manifest's embedded public key into a
// temporary keyring and extracts its fingerprint. This validates the key and
// returns its actual fingerprint from the key material itself.
func (c *Checker) ExtractEmbeddedSigningKeyFP() (string, error) {
func (c *Checker) ExtractEmbeddedSigningKeyFP(ctx context.Context) (string, error) {
if len(c.signingPubKey) == 0 {
return "", errors.New("manifest has no signing public key")
return "", errNoSigningPubKey
}
return gpgExtractPubKeyFingerprint(c.signingPubKey)
return gpgExtractPubKeyFingerprint(ctx, c.signingPubKey)
}
// Check verifies all files against the manifest.
// Results are sent to the results channel as files are checked.
// Progress updates are sent to the progress channel approximately once per second.
// Both channels are closed when the method returns.
func (c *Checker) Check(ctx context.Context, results chan<- Result, progress chan<- CheckStatus) error {
func (c *Checker) Check(
ctx context.Context,
results chan<- Result,
progress chan<- CheckStatus,
) error {
if results != nil {
defer close(results)
}
if progress != nil {
defer close(progress)
}
@@ -165,11 +212,14 @@ func (c *Checker) Check(ctx context.Context, results chan<- Result, progress cha
totalFiles := FileCount(len(c.files))
totalBytes := c.TotalBytes()
var checkedFiles FileCount
var checkedBytes FileSize
var failures FileCount
var (
checkedFiles FileCount
checkedBytes FileSize
failures FileCount
)
startTime := time.Now()
lastProgressTime := time.Now()
for _, entry := range c.files {
select {
@@ -182,108 +232,62 @@ func (c *Checker) Check(ctx context.Context, results chan<- Result, progress cha
if result.Status != StatusOK {
failures++
}
checkedFiles++
if results != nil {
results <- result
}
// Send progress with rate and ETA calculation
// Send progress at most once per second (rate-limited)
if progress != nil {
elapsed := time.Since(startTime)
var bytesPerSec float64
var eta time.Duration
now := time.Now()
if elapsed > 0 && checkedBytes > 0 {
bytesPerSec = float64(checkedBytes) / elapsed.Seconds()
remainingBytes := totalBytes - checkedBytes
if bytesPerSec > 0 {
eta = time.Duration(float64(remainingBytes)/bytesPerSec) * time.Second
}
isLast := checkedFiles == totalFiles
if isLast || now.Sub(lastProgressTime) >= time.Second {
bytesPerSec, eta := computeRateETA(
time.Since(startTime), checkedBytes, totalBytes,
)
sendCheckStatus(progress, CheckStatus{
TotalFiles: totalFiles,
CheckedFiles: checkedFiles,
TotalBytes: totalBytes,
CheckedBytes: checkedBytes,
BytesPerSec: bytesPerSec,
ETA: eta,
Failures: failures,
})
lastProgressTime = now
}
sendCheckStatus(progress, CheckStatus{
TotalFiles: totalFiles,
CheckedFiles: checkedFiles,
TotalBytes: totalBytes,
CheckedBytes: checkedBytes,
BytesPerSec: bytesPerSec,
ETA: eta,
Failures: failures,
})
}
}
return nil
}
func (c *Checker) checkFile(entry *MFFilePath, checkedBytes *FileSize) Result {
absPath := filepath.Join(string(c.basePath), entry.Path)
relPath := RelFilePath(entry.Path)
// Check if file exists
info, err := c.fs.Stat(absPath)
if err != nil {
if errors.Is(err, os.ErrNotExist) || errors.Is(err, afero.ErrFileNotFound) {
return Result{Path: relPath, Status: StatusMissing, Message: "file not found"}
}
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check size
if info.Size() != entry.Size {
*checkedBytes += FileSize(info.Size())
return Result{
Path: relPath,
Status: StatusSizeMismatch,
Message: "size mismatch",
}
}
// Open and hash file
f, err := c.fs.Open(absPath)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
defer func() { _ = f.Close() }()
h := sha256.New()
n, err := io.Copy(h, f)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
*checkedBytes += FileSize(n)
// Encode as multihash and compare
computed, err := multihash.Encode(h.Sum(nil), multihash.SHA2_256)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check against all hashes in manifest (at least one must match)
for _, hash := range entry.Hashes {
if bytes.Equal(computed, hash.MultiHash) {
return Result{Path: relPath, Status: StatusOK}
}
}
return Result{Path: relPath, Status: StatusHashMismatch, Message: "hash mismatch"}
}
// FindExtraFiles walks the filesystem and reports files not in the manifest.
// Results are sent to the results channel. The channel is closed when done.
// Hidden files/directories (starting with .) are skipped, as they are excluded
// from manifests by default. The manifest file itself is also skipped.
// FindExtraFiles walks the filesystem and reports files not in the manifest,
// hidden files and directories included. The manifest file itself is not
// reported. Anything the search cannot read, such as a directory that cannot
// be listed, is reported with StatusError and the search goes on. Results are
// sent to the results channel. The channel is closed when done.
func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) error {
if results != nil {
defer close(results)
}
return afero.Walk(c.fs, string(c.basePath), func(walkPath string, info os.FileInfo, err error) error {
if err != nil {
return err
}
// The search does not follow symlinks, so a base directory named
// through one is resolved first. If that fails, the base is searched as
// named and the search reports the problem.
root := string(c.basePath)
resolved, err := filepath.EvalSymlinks(root)
if err == nil {
root = resolved
}
walkFn := func(walkPath string, info os.FileInfo, walkErr error) error {
select {
case <-ctx.Done():
return ctx.Err()
@@ -291,16 +295,24 @@ func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) err
}
// Get relative path
rel, err := filepath.Rel(string(c.basePath), walkPath)
rel, err := filepath.Rel(root, walkPath)
if err != nil {
return err
}
// Skip hidden files and directories (dotfiles)
if IsHiddenPath(filepath.ToSlash(rel)) {
if info.IsDir() {
return filepath.SkipDir
relPath := RelFilePath(rel)
// Report what cannot be read, such as a directory that cannot be
// listed, and go on with the rest.
if walkErr != nil {
if results != nil {
results <- Result{
Path: relPath,
Status: StatusError,
Message: walkErr.Error(),
}
}
return nil
}
@@ -309,13 +321,19 @@ func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) err
return nil
}
// Skip manifest files
base := filepath.Base(rel)
if base == "index.mf" || base == ".index.mf" {
return nil
// A symlink is compared by what it points to, so a manifest reached
// through one is not reported either.
if info.Mode()&os.ModeSymlink != 0 {
target, statErr := c.fs.Stat(walkPath)
if statErr == nil {
info = target
}
}
relPath := RelFilePath(rel)
// Skip the manifest file itself, however its path is spelled
if os.SameFile(info, c.manifestInfo) {
return nil
}
// Check if path is in manifest
if _, exists := c.manifestPaths[relPath]; !exists {
@@ -329,7 +347,79 @@ func (c *Checker) FindExtraFiles(ctx context.Context, results chan<- Result) err
}
return nil
})
}
return afero.Walk(c.fs, root, walkFn)
}
func (c *Checker) checkFile(entry *MFFilePath, checkedBytes *FileSize) Result {
// entry.GetPath() is safe to join here: a manifest's entry paths are
// validated against the path invariants when it is loaded (see
// deserializeInner) or built (see Builder.AddFile), so a traversal or
// absolute path can never reach this point.
absPath := filepath.Join(string(c.basePath), entry.GetPath())
relPath := RelFilePath(entry.GetPath())
// Check if file exists
info, err := c.fs.Stat(absPath)
if err != nil {
if errors.Is(err, os.ErrNotExist) || errors.Is(err, afero.ErrFileNotFound) {
return Result{
Path: relPath,
Status: StatusMissing,
Message: "file not found",
}
}
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check size
if info.Size() != entry.GetSize() {
*checkedBytes += FileSize(info.Size())
return Result{
Path: relPath,
Status: StatusSizeMismatch,
Message: "size mismatch",
}
}
// Open and hash file
f, err := c.fs.Open(absPath)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
defer func() { _ = f.Close() }()
h := sha256.New()
n, err := io.Copy(h, f)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
*checkedBytes += FileSize(n)
// Encode as multihash and compare
computed, err := multihash.Encode(h.Sum(nil), multihash.SHA2_256)
if err != nil {
return Result{Path: relPath, Status: StatusError, Message: err.Error()}
}
// Check against all hashes in manifest (at least one must match)
for _, hash := range entry.GetHashes() {
if bytes.Equal(computed, hash.GetMultiHash()) {
return Result{Path: relPath, Status: StatusOK}
}
}
return Result{
Path: relPath,
Status: StatusHashMismatch,
Message: "hash mismatch",
}
}
// sendCheckStatus sends a status update without blocking.
@@ -337,6 +427,7 @@ func sendCheckStatus(ch chan<- CheckStatus, status CheckStatus) {
if ch == nil {
return
}
select {
case ch <- status:
default:
+425 -108
View File
@@ -1,8 +1,12 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"bytes"
"context"
"fmt"
"os"
"path/filepath"
"testing"
"time"
@@ -11,7 +15,17 @@ import (
"github.com/stretchr/testify/require"
)
const (
testFile1 = "file1.txt"
testFile2 = "file2.txt"
testExistsFile = "exists.txt"
testManifestPath = "/manifest.mf"
testDataDir = "/data"
)
func TestStatusString(t *testing.T) {
t.Parallel()
tests := []struct {
status Status
expected string
@@ -27,77 +41,169 @@ func TestStatusString(t *testing.T) {
for _, tt := range tests {
t.Run(tt.expected, func(t *testing.T) {
t.Parallel()
assert.Equal(t, tt.expected, tt.status.String())
})
}
}
// createTestManifest creates a manifest file in the filesystem with the given files.
func createTestManifest(t *testing.T, fs afero.Fs, manifestPath string, files map[string][]byte) {
func createTestManifest(
t *testing.T, fs afero.Fs, manifestPath string, files map[string][]byte,
) {
t.Helper()
builder := NewBuilder()
for path, content := range files {
reader := bytes.NewReader(content)
_, err := builder.AddFile(RelFilePath(path), FileSize(len(content)), ModTime(time.Now()), reader, nil)
_, err := builder.AddFile(
RelFilePath(path), FileSize(len(content)), ModTime(time.Now()), reader, nil,
)
require.NoError(t, err)
}
var buf bytes.Buffer
require.NoError(t, builder.Build(&buf))
require.NoError(t, builder.Build(context.Background(), &buf))
require.NoError(t, afero.WriteFile(fs, manifestPath, buf.Bytes(), 0o644))
}
// createFilesOnDisk creates the given files on the filesystem.
func createFilesOnDisk(t *testing.T, fs afero.Fs, basePath string, files map[string][]byte) {
// createFilesOnDisk creates the given files on the filesystem under
// testDataDir.
func createFilesOnDisk(t *testing.T, fs afero.Fs, files map[string][]byte) {
t.Helper()
for path, content := range files {
fullPath := basePath + "/" + path
require.NoError(t, fs.MkdirAll(basePath, 0o755))
fullPath := testDataDir + "/" + path
require.NoError(t, fs.MkdirAll(testDataDir, 0o755))
require.NoError(t, afero.WriteFile(fs, fullPath, content, 0o644))
}
}
func TestNewChecker(t *testing.T) {
t.Parallel()
t.Run("valid manifest", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
"file1.txt": []byte("hello"),
"file2.txt": []byte("world"),
testFile1: []byte("hello"),
testFile2: []byte("world"),
}
createTestManifest(t, fs, "/manifest.mf", files)
createTestManifest(t, fs, testManifestPath, files)
chk, err := NewChecker("/manifest.mf", "/", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: "/",
Fs: fs,
})
require.NoError(t, err)
assert.NotNil(t, chk)
assert.Equal(t, FileCount(2), chk.FileCount())
})
t.Run("missing manifest", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
_, err := NewChecker("/nonexistent.mf", "/", fs)
_, err := NewChecker(&CheckerOptions{
ManifestPath: "/nonexistent.mf",
BasePath: "/",
Fs: fs,
})
assert.Error(t, err)
})
t.Run("invalid manifest", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/bad.mf", []byte("not a manifest"), 0o644))
_, err := NewChecker("/bad.mf", "/", fs)
_, err := NewChecker(&CheckerOptions{
ManifestPath: "/bad.mf",
BasePath: "/",
Fs: fs,
})
assert.Error(t, err)
})
}
func TestNewCheckerRequiredPaths(t *testing.T) {
t.Parallel()
for _, tc := range []struct {
name string
opts *CheckerOptions
want string
is error
}{
{
name: "nil options",
opts: nil,
want: "manifest path cannot be empty",
is: errManifestPathEmpty,
},
{
name: "empty manifest path",
opts: &CheckerOptions{BasePath: testDataDir},
want: "manifest path cannot be empty",
is: errManifestPathEmpty,
},
{
name: "empty base path",
opts: &CheckerOptions{ManifestPath: testManifestPath},
want: "base path cannot be empty",
is: errBasePathEmpty,
},
} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
chk, err := NewChecker(tc.opts)
require.ErrorIs(t, err, tc.is)
require.EqualError(t, err, tc.want)
assert.Nil(t, chk)
})
}
}
func TestNewCheckerNilFsUsesOsFs(t *testing.T) {
t.Parallel()
dir := t.TempDir()
manifestPath := filepath.Join(dir, "index.mf")
content := []byte("hello")
createTestManifest(t, afero.NewOsFs(), manifestPath, map[string][]byte{
testFile1: content,
})
require.NoError(t, os.WriteFile(filepath.Join(dir, testFile1), content, 0o600))
chk, err := NewChecker(&CheckerOptions{ManifestPath: manifestPath, BasePath: dir})
require.NoError(t, err)
results := make(chan Result, 1)
require.NoError(t, chk.Check(context.Background(), results, nil))
assert.Equal(t, StatusOK, (<-results).Status)
}
func TestCheckerFileCountAndTotalBytes(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
"small.txt": []byte("hi"),
"medium.txt": []byte("hello world"),
"large.txt": bytes.Repeat([]byte("x"), 1000),
}
createTestManifest(t, fs, "/manifest.mf", files)
createTestManifest(t, fs, testManifestPath, files)
chk, err := NewChecker("/manifest.mf", "/", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: "/",
Fs: fs,
})
require.NoError(t, err)
assert.Equal(t, FileCount(3), chk.FileCount())
@@ -105,15 +211,21 @@ func TestCheckerFileCountAndTotalBytes(t *testing.T) {
}
func TestCheckAllFilesOK(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
"file1.txt": []byte("content one"),
"file2.txt": []byte("content two"),
testFile1: []byte("content one"),
testFile2: []byte("content two"),
}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files)
createTestManifest(t, fs, testManifestPath, files)
createFilesOnDisk(t, fs, files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 10)
@@ -126,24 +238,31 @@ func TestCheckAllFilesOK(t *testing.T) {
}
assert.Len(t, resultList, 2)
for _, r := range resultList {
assert.Equal(t, StatusOK, r.Status, "file %s should be OK", r.Path)
}
}
func TestCheckMissingFile(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
"exists.txt": []byte("I exist"),
"missing.txt": []byte("I don't exist on disk"),
testExistsFile: []byte("I exist"),
"missing.txt": []byte("I don't exist on disk"),
}
createTestManifest(t, fs, "/manifest.mf", files)
createTestManifest(t, fs, testManifestPath, files)
// Only create one file
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"exists.txt": []byte("I exist"),
createFilesOnDisk(t, fs, map[string][]byte{
testExistsFile: []byte("I exist"),
})
chk, err := NewChecker("/manifest.mf", "/data", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 10)
@@ -151,13 +270,17 @@ func TestCheckMissingFile(t *testing.T) {
require.NoError(t, err)
var okCount, missingCount int
for r := range results {
switch r.Status {
case StatusOK:
okCount++
case StatusMissing:
missingCount++
assert.Equal(t, RelFilePath("missing.txt"), r.Path)
case StatusSizeMismatch, StatusHashMismatch, StatusExtra, StatusError:
// Not expected in this test; counted assertions below will fail.
}
}
@@ -166,17 +289,23 @@ func TestCheckMissingFile(t *testing.T) {
}
func TestCheckSizeMismatch(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
"file.txt": []byte("original content"),
testFileName: []byte("original content"),
}
createTestManifest(t, fs, "/manifest.mf", files)
createTestManifest(t, fs, testManifestPath, files)
// Create file with different size
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"file.txt": []byte("short"),
createFilesOnDisk(t, fs, map[string][]byte{
testFileName: []byte("short"),
})
chk, err := NewChecker("/manifest.mf", "/data", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 10)
@@ -185,24 +314,30 @@ func TestCheckSizeMismatch(t *testing.T) {
r := <-results
assert.Equal(t, StatusSizeMismatch, r.Status)
assert.Equal(t, RelFilePath("file.txt"), r.Path)
assert.Equal(t, RelFilePath(testFileName), r.Path)
}
func TestCheckHashMismatch(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
originalContent := []byte("original content")
files := map[string][]byte{
"file.txt": originalContent,
testFileName: originalContent,
}
createTestManifest(t, fs, "/manifest.mf", files)
createTestManifest(t, fs, testManifestPath, files)
// Create file with same size but different content
differentContent := []byte("different contnt") // same length (16 bytes) but different
require.Equal(t, len(originalContent), len(differentContent), "test requires same length")
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"file.txt": differentContent,
require.Len(t, differentContent, len(originalContent), "test requires same length")
createFilesOnDisk(t, fs, map[string][]byte{
testFileName: differentContent,
})
chk, err := NewChecker("/manifest.mf", "/data", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 10)
@@ -211,19 +346,25 @@ func TestCheckHashMismatch(t *testing.T) {
r := <-results
assert.Equal(t, StatusHashMismatch, r.Status)
assert.Equal(t, RelFilePath("file.txt"), r.Path)
assert.Equal(t, RelFilePath(testFileName), r.Path)
}
func TestCheckWithProgress(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
"file1.txt": bytes.Repeat([]byte("a"), 100),
"file2.txt": bytes.Repeat([]byte("b"), 200),
testFile1: bytes.Repeat([]byte("a"), 100),
testFile2: bytes.Repeat([]byte("b"), 200),
}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files)
createTestManifest(t, fs, testManifestPath, files)
createFilesOnDisk(t, fs, files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 10)
@@ -232,9 +373,7 @@ func TestCheckWithProgress(t *testing.T) {
err = chk.Check(context.Background(), results, progress)
require.NoError(t, err)
// Drain results
for range results {
}
// results is fully buffered and closed; no draining needed
// Check progress was sent
var progressUpdates []CheckStatus
@@ -253,16 +392,23 @@ func TestCheckWithProgress(t *testing.T) {
}
func TestCheckContextCancellation(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
// Create many files to ensure we have time to cancel
files := make(map[string][]byte)
for i := 0; i < 100; i++ {
for i := range 100 {
files[string(rune('a'+i%26))+".txt"] = bytes.Repeat([]byte("x"), 1000)
}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files)
chk, err := NewChecker("/manifest.mf", "/data", fs)
createTestManifest(t, fs, testManifestPath, files)
createFilesOnDisk(t, fs, files)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
ctx, cancel := context.WithCancel(context.Background())
@@ -274,20 +420,26 @@ func TestCheckContextCancellation(t *testing.T) {
}
func TestFindExtraFiles(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
// Manifest only contains file1
manifestFiles := map[string][]byte{
"file1.txt": []byte("in manifest"),
testFile1: []byte("in manifest"),
}
createTestManifest(t, fs, "/manifest.mf", manifestFiles)
createTestManifest(t, fs, testManifestPath, manifestFiles)
// Disk has file1 and file2
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"file1.txt": []byte("in manifest"),
"file2.txt": []byte("extra file"),
createFilesOnDisk(t, fs, map[string][]byte{
testFile1: []byte("in manifest"),
testFile2: []byte("extra file"),
})
chk, err := NewChecker("/manifest.mf", "/data", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 10)
@@ -300,56 +452,140 @@ func TestFindExtraFiles(t *testing.T) {
}
assert.Len(t, extras, 1)
assert.Equal(t, RelFilePath("file2.txt"), extras[0].Path)
assert.Equal(t, RelFilePath(testFile2), extras[0].Path)
assert.Equal(t, StatusExtra, extras[0].Status)
assert.Equal(t, "not in manifest", extras[0].Message)
}
func TestFindExtraFilesSkipsManifestAndDotfiles(t *testing.T) {
fs := afero.NewMemMapFs()
manifestFiles := map[string][]byte{
"file1.txt": []byte("in manifest"),
}
createTestManifest(t, fs, "/data/.index.mf", manifestFiles)
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"file1.txt": []byte("in manifest"),
})
// Create dotfile and manifest that should be skipped
require.NoError(t, afero.WriteFile(fs, "/data/.hidden", []byte("hidden"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/data/.config/settings", []byte("cfg"), 0o644))
// Create a real extra file
require.NoError(t, fs.MkdirAll("/data", 0o755))
require.NoError(t, afero.WriteFile(fs, "/data/extra.txt", []byte("extra"), 0o644))
// TestFindExtraFilesReportsHiddenFilesButNotManifest keeps the manifest
// inside the checked tree: hidden files and directories are reported, the
// manifest is not. The manifest is recognized by file identity, which needs
// the real filesystem.
func TestFindExtraFilesReportsHiddenFilesButNotManifest(t *testing.T) {
t.Parallel()
chk, err := NewChecker("/data/.index.mf", "/data", fs)
dir := t.TempDir()
manifestPath := filepath.Join(dir, "index.mf")
fs := afero.NewOsFs()
createTestManifest(t, fs, manifestPath, map[string][]byte{
testFile1: []byte("in manifest"),
})
unlisted := []RelFilePath{"extra.txt", ".hidden", ".git/config"}
for _, p := range append([]RelFilePath{testFile1}, unlisted...) {
path := filepath.Join(dir, string(p))
require.NoError(t, fs.MkdirAll(filepath.Dir(path), 0o750))
require.NoError(t, afero.WriteFile(fs, path, []byte("x"), 0o600))
}
chk, err := NewChecker(&CheckerOptions{
ManifestPath: manifestPath,
BasePath: dir,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 10)
err = chk.FindExtraFiles(context.Background(), results)
require.NoError(t, chk.FindExtraFiles(context.Background(), results))
var extras []RelFilePath
for r := range results {
extras = append(extras, r.Path)
}
assert.ElementsMatch(t, unlisted, extras)
}
// TestFindExtraFilesSkipsManifestReachedThroughSymlink checks a tree whose
// index.mf is a symlink to the manifest kept outside the tree: the symlink is
// not reported.
func TestFindExtraFilesSkipsManifestReachedThroughSymlink(t *testing.T) {
t.Parallel()
dir := t.TempDir()
tree := filepath.Join(dir, "tree")
manifestPath := filepath.Join(dir, "real.mf")
linkPath := filepath.Join(tree, "index.mf")
fs := afero.NewOsFs()
createTestManifest(t, fs, manifestPath, map[string][]byte{testFile1: []byte("x")})
require.NoError(t, fs.MkdirAll(tree, 0o750))
require.NoError(t,
afero.WriteFile(fs, filepath.Join(tree, testFile1), []byte("x"), 0o600))
require.NoError(t, os.Symlink(manifestPath, linkPath))
chk, err := NewChecker(&CheckerOptions{
ManifestPath: linkPath,
BasePath: tree,
Fs: fs,
})
require.NoError(t, err)
var extras []Result
results := make(chan Result, 10)
require.NoError(t, chk.FindExtraFiles(context.Background(), results))
var extras []RelFilePath
for r := range results {
extras = append(extras, r)
extras = append(extras, r.Path)
}
// Should only report extra.txt, not .hidden, .config/settings, or .index.mf
for _, e := range extras {
t.Logf("extra: %s", e.Path)
assert.Empty(t, extras)
}
// TestFindExtraFilesSearchesBaseNamedThroughSymlink names the checked tree
// through a symlink to it: the files in the tree are searched, and the
// symlink itself is not reported.
func TestFindExtraFilesSearchesBaseNamedThroughSymlink(t *testing.T) {
t.Parallel()
dir := t.TempDir()
tree := filepath.Join(dir, "tree")
link := filepath.Join(dir, "link")
manifestPath := filepath.Join(dir, "index.mf")
fs := afero.NewOsFs()
createTestManifest(t, fs, manifestPath, map[string][]byte{testFile1: []byte("x")})
require.NoError(t, fs.MkdirAll(tree, 0o750))
for _, name := range []string{testFile1, testFile2} {
require.NoError(t,
afero.WriteFile(fs, filepath.Join(tree, name), []byte("x"), 0o600))
}
assert.Len(t, extras, 1)
if len(extras) > 0 {
assert.Equal(t, RelFilePath("extra.txt"), extras[0].Path)
require.NoError(t, os.Symlink(tree, link))
chk, err := NewChecker(&CheckerOptions{
ManifestPath: manifestPath,
BasePath: link,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 10)
require.NoError(t, chk.FindExtraFiles(context.Background(), results))
var extras []RelFilePath
for r := range results {
extras = append(extras, r.Path)
}
assert.Equal(t, []RelFilePath{testFile2}, extras)
}
func TestFindExtraFilesContextCancellation(t *testing.T) {
fs := afero.NewMemMapFs()
files := map[string][]byte{"file.txt": []byte("data")}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files)
t.Parallel()
chk, err := NewChecker("/manifest.mf", "/data", fs)
fs := afero.NewMemMapFs()
files := map[string][]byte{testFileName: []byte("data")}
createTestManifest(t, fs, testManifestPath, files)
createFilesOnDisk(t, fs, files)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
ctx, cancel := context.WithCancel(context.Background())
@@ -361,12 +597,18 @@ func TestFindExtraFilesContextCancellation(t *testing.T) {
}
func TestCheckNilChannels(t *testing.T) {
fs := afero.NewMemMapFs()
files := map[string][]byte{"file.txt": []byte("data")}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files)
t.Parallel()
chk, err := NewChecker("/manifest.mf", "/data", fs)
fs := afero.NewMemMapFs()
files := map[string][]byte{testFileName: []byte("data")}
createTestManifest(t, fs, testManifestPath, files)
createFilesOnDisk(t, fs, files)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
// Should not panic with nil channels
@@ -375,12 +617,18 @@ func TestCheckNilChannels(t *testing.T) {
}
func TestFindExtraFilesNilChannel(t *testing.T) {
fs := afero.NewMemMapFs()
files := map[string][]byte{"file.txt": []byte("data")}
createTestManifest(t, fs, "/manifest.mf", files)
createFilesOnDisk(t, fs, "/data", files)
t.Parallel()
chk, err := NewChecker("/manifest.mf", "/data", fs)
fs := afero.NewMemMapFs()
files := map[string][]byte{testFileName: []byte("data")}
createTestManifest(t, fs, testManifestPath, files)
createFilesOnDisk(t, fs, files)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
// Should not panic with nil channel
@@ -389,22 +637,29 @@ func TestFindExtraFilesNilChannel(t *testing.T) {
}
func TestCheckSubdirectories(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
files := map[string][]byte{
"dir1/file1.txt": []byte("content1"),
"dir1/dir2/file2.txt": []byte("content2"),
"dir1/dir2/dir3/deep.txt": []byte("deep content"),
}
createTestManifest(t, fs, "/manifest.mf", files)
createTestManifest(t, fs, testManifestPath, files)
// Create files with full directory structure
for path, content := range files {
fullPath := "/data/" + path
require.NoError(t, fs.MkdirAll("/data/dir1/dir2/dir3", 0o755))
require.NoError(t, afero.WriteFile(fs, fullPath, content, 0o644))
}
chk, err := NewChecker("/manifest.mf", "/data", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 10)
@@ -412,28 +667,37 @@ func TestCheckSubdirectories(t *testing.T) {
require.NoError(t, err)
var okCount int
for r := range results {
assert.Equal(t, StatusOK, r.Status, "file %s should be OK", r.Path)
okCount++
}
assert.Equal(t, 3, okCount)
}
func TestCheckMissingFileDetectedWithoutFallback(t *testing.T) {
t.Parallel()
// Regression test: errors.Is(err, errors.New("...")) never matches because
// errors.New creates a new value each time. The fix uses os.ErrNotExist instead.
fs := afero.NewMemMapFs()
files := map[string][]byte{
"exists.txt": []byte("here"),
"missing.txt": []byte("not on disk"),
testExistsFile: []byte("here"),
"missing.txt": []byte("not on disk"),
}
createTestManifest(t, fs, "/manifest.mf", files)
createTestManifest(t, fs, testManifestPath, files)
// Only create one file on disk
createFilesOnDisk(t, fs, "/data", map[string][]byte{
"exists.txt": []byte("here"),
createFilesOnDisk(t, fs, map[string][]byte{
testExistsFile: []byte("here"),
})
chk, err := NewChecker("/manifest.mf", "/data", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 10)
@@ -447,17 +711,24 @@ func TestCheckMissingFileDetectedWithoutFallback(t *testing.T) {
assert.Equal(t, RelFilePath("missing.txt"), r.Path)
}
}
assert.Equal(t, 1, statusCounts[StatusOK], "one file should be OK")
assert.Equal(t, 1, statusCounts[StatusMissing], "one file should be MISSING")
assert.Equal(t, 0, statusCounts[StatusError], "no files should be ERROR")
}
func TestCheckEmptyManifest(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
// Create manifest with no files
createTestManifest(t, fs, "/manifest.mf", map[string][]byte{})
createTestManifest(t, fs, testManifestPath, map[string][]byte{})
chk, err := NewChecker("/manifest.mf", "/data", fs)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
assert.Equal(t, FileCount(0), chk.FileCount())
@@ -471,5 +742,51 @@ func TestCheckEmptyManifest(t *testing.T) {
for range results {
count++
}
assert.Equal(t, 0, count)
}
func TestCheckProgressRateLimited(t *testing.T) {
t.Parallel()
// Create many small files - progress should be rate-limited, not one per file.
// With rate-limiting to once per second, we should get far fewer progress
// updates than files (plus one final update).
fs := afero.NewMemMapFs()
files := make(map[string][]byte, 100)
for i := range 100 {
name := fmt.Sprintf("file%03d.txt", i)
files[name] = []byte("content")
}
createTestManifest(t, fs, testManifestPath, files)
createFilesOnDisk(t, fs, files)
chk, err := NewChecker(&CheckerOptions{
ManifestPath: testManifestPath,
BasePath: testDataDir,
Fs: fs,
})
require.NoError(t, err)
results := make(chan Result, 200)
progress := make(chan CheckStatus, 200)
err = chk.Check(context.Background(), results, progress)
require.NoError(t, err)
// results is fully buffered and closed; no draining needed
// Count progress updates
var progressCount int
for range progress {
progressCount++
}
// Should be far fewer than 100 (rate-limited to once per second)
// At minimum we get the final update
assert.GreaterOrEqual(t, progressCount, 1,
"should get at least the final progress update")
assert.Less(t, progressCount, 100,
"progress should be rate-limited, not one per file")
}
+34 -1
View File
@@ -1,11 +1,44 @@
package mfer
const (
Version = "0.1.0"
// Version is the current mfer release version.
Version = "0.1.0"
// ReleaseDate is the date on which Version was released.
ReleaseDate = "2025-12-17"
// MaxDecompressedSize is the maximum allowed size of decompressed manifest
// data (256 MB). This prevents decompression bombs from consuming excessive
// memory.
MaxDecompressedSize int64 = 256 * 1024 * 1024
// zstdWindowSize is the zstd window zstd.SpeedBestCompression gives mfer's writer.
zstdWindowSize = 8 << 20
// uuidLength is the length in bytes of a binary UUID.
uuidLength = 16
// Numbers in mf.proto of MFFile.files and of the MFFilePath fields
// that decoding sets aside a fixed amount of memory for.
filesFieldNumber = 101
hashesFieldNumber = 3
mimeTypeFieldNumber = 301
mtimeFieldNumber = 302
ctimeFieldNumber = 303
// Bytes decoding sets aside for each file entry, hash, timestamp and
// MIME type, however short its encoding. checkDecodedSize refuses an
// inner message for which these add up to more than maxDecodedGrowth
// times its size.
decodedFileEntrySize = 160
decodedHashSize = 112
decodedTimestampSize = 64
decodedMIMETypeSize = 16
// Each file entry mfer writes holds a path of at least one byte, a
// multihash at least as long as SHA-256's 34 bytes (AddFileWithHash
// refuses shorter ones) and a modification time: at least 47 bytes,
// counted at 336. So its manifests add up to at most about 7.15 times
// their size, and this limit is about 12% above that.
maxDecodedGrowth = 8
)
+276 -63
View File
@@ -2,6 +2,7 @@ package mfer
import (
"bytes"
"context"
"crypto/sha256"
"errors"
"fmt"
@@ -10,110 +11,294 @@ import (
"github.com/google/uuid"
"github.com/klauspost/compress/zstd"
"github.com/spf13/afero"
"google.golang.org/protobuf/encoding/protowire"
"google.golang.org/protobuf/proto"
"sneak.berlin/go/mfer/internal/bork"
"sneak.berlin/go/mfer/internal/log"
)
var (
errInvalidUUIDLength = errors.New("invalid UUID length")
errInvalidUUIDFormat = errors.New("invalid UUID format")
errUnknownVersion = errors.New("unknown version")
errUnknownCompression = errors.New("unknown compression type")
errCompressedHashWrong = errors.New("compressed data hash mismatch")
errSignatureNoPubKey = errors.New("signature present but no public key")
errDecompressedTooLarge = errors.New("decompressed data exceeds maximum allowed size")
errUUIDMismatch = errors.New("outer and inner UUID mismatch")
errInvalidFileFormat = errors.New("invalid file format")
errInvalidManifestPath = errors.New("manifest contains invalid path")
errDecodedTooLarge = errors.New(
"manifest would take too much memory to decode")
)
// validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable).
func validateUUID(data []byte) error {
if len(data) != 16 {
return errors.New("invalid UUID length")
if len(data) != uuidLength {
return errInvalidUUIDLength
}
// Try to parse as UUID to validate format
_, err := uuid.FromBytes(data)
if err != nil {
return errors.New("invalid UUID format")
return errInvalidUUIDFormat
}
return nil
}
func (m *manifest) deserializeInner() error {
if m.pbOuter.Version != MFFileOuter_VERSION_ONE {
return errors.New("unknown version")
// validateOuterHeader checks the outer message's version, compression
// type, and UUID.
func (m *manifest) validateOuterHeader() error {
if m.pbOuter.GetVersion() != MFFileOuter_VERSION_ONE {
return errUnknownVersion
}
if m.pbOuter.CompressionType != MFFileOuter_COMPRESSION_ZSTD {
return errors.New("unknown compression type")
if m.pbOuter.GetCompressionType() != MFFileOuter_COMPRESSION_ZSTD {
return errUnknownCompression
}
// Validate outer UUID before any decompression
if err := validateUUID(m.pbOuter.Uuid); err != nil {
return errors.New("outer UUID invalid: " + err.Error())
}
// Verify hash of compressed data before decompression
h := sha256.New()
if _, err := h.Write(m.pbOuter.InnerMessage); err != nil {
return err
}
sha256Hash := h.Sum(nil)
if !bytes.Equal(sha256Hash, m.pbOuter.Sha256) {
return errors.New("compressed data hash mismatch")
}
// Verify signature if present
if len(m.pbOuter.Signature) > 0 {
if len(m.pbOuter.SigningPubKey) == 0 {
return errors.New("signature present but no public key")
}
sigString, err := m.signatureString()
if err != nil {
return fmt.Errorf("failed to generate signature string for verification: %w", err)
}
if err := gpgVerify([]byte(sigString), m.pbOuter.Signature, m.pbOuter.SigningPubKey); err != nil {
return fmt.Errorf("signature verification failed: %w", err)
}
log.Infof("signature verified successfully")
}
bb := bytes.NewBuffer(m.pbOuter.InnerMessage)
zr, err := zstd.NewReader(bb)
err := validateUUID(m.pbOuter.GetUuid())
if err != nil {
return err
return fmt.Errorf("outer UUID invalid: %w", err)
}
return nil
}
// verifyOuterIntegrity checks the hash of the compressed payload and,
// if a signature is present, verifies it against the embedded public key.
func (m *manifest) verifyOuterIntegrity() error {
h := sha256.New()
_, err := h.Write(m.pbOuter.GetInnerMessage())
if err != nil {
return fmt.Errorf("deserialize: hash write: %w", err)
}
sha256Hash := h.Sum(nil)
if !bytes.Equal(sha256Hash, m.pbOuter.GetSha256()) {
return errCompressedHashWrong
}
if len(m.pbOuter.GetSignature()) == 0 {
return nil
}
if len(m.pbOuter.GetSigningPubKey()) == 0 {
return errSignatureNoPubKey
}
sigString, err := m.signatureString()
if err != nil {
return fmt.Errorf(
"failed to generate signature string for verification: %w", err,
)
}
// Loading a manifest takes no context; gpgTimeout still bounds gpg.
err = gpgVerify(
context.Background(),
[]byte(sigString),
m.pbOuter.GetSignature(),
m.pbOuter.GetSigningPubKey(),
)
if err != nil {
return fmt.Errorf("signature verification failed: %w", err)
}
log.Infof("signature verified successfully")
return nil
}
// decompressInner decompresses the inner payload, enforcing size limits
// to prevent decompression bombs.
func (m *manifest) decompressInner() ([]byte, error) {
bb := bytes.NewBuffer(m.pbOuter.GetInnerMessage())
// By default the decoder decodes a payload under 128 KiB in full,
// each frame up to the decoder's limit, before the LimitReader below
// reads any of it. Decoding synchronously and never in full makes it
// decode only what the LimitReader asks for. It also sets aside a new
// buffer for each frame that asks for a larger window than the frames
// before it, even a frame holding no data. Refusing windows above
// zstdWindowSize keeps each buffer to a little over zstdWindowSize, and
// the buffers of frames holding no data to about 16 times it in total.
zr, err := zstd.NewReader(bb,
zstd.WithDecoderConcurrency(1),
zstd.WithDecodeBuffersBelow(0),
zstd.WithDecoderMaxWindow(zstdWindowSize))
if err != nil {
return nil, fmt.Errorf("deserialize: zstd reader: %w", err)
}
defer zr.Close()
// Limit decompressed size to prevent decompression bombs.
// Use declared size + 1 byte to detect overflow, capped at MaxDecompressedSize.
maxSize := MaxDecompressedSize
if m.pbOuter.Size > 0 && m.pbOuter.Size < int64(maxSize) {
maxSize = int64(m.pbOuter.Size) + 1
if m.pbOuter.GetSize() > 0 && m.pbOuter.GetSize() < maxSize {
maxSize = m.pbOuter.GetSize() + 1
}
limitedReader := io.LimitReader(zr, maxSize)
dat, err := io.ReadAll(limitedReader)
if err != nil {
return nil, fmt.Errorf("deserialize: decompress: %w", err)
}
if int64(len(dat)) >= MaxDecompressedSize {
return nil, fmt.Errorf(
"%w of %d bytes", errDecompressedTooLarge, MaxDecompressedSize,
)
}
return dat, nil
}
// checkDecodedSize refuses an encoded inner message whose file entries,
// hashes, timestamps and MIME types would take more than maxDecodedGrowth
// times its size to decode. Decoding sets aside a fixed amount for each,
// however short its encoding, so a message of empty ones would take about
// 50 times its size.
func checkDecodedSize(inner []byte) error {
limit := maxDecodedGrowth * int64(len(inner))
var decoded int64
add := func(size int64) error {
decoded += size
if decoded > limit {
return errDecodedTooLarge
}
return nil
}
return forEachBytesField(inner, func(num protowire.Number, entry []byte) error {
if num != filesFieldNumber {
return nil
}
err := add(decodedFileEntrySize)
if err != nil {
return err
}
return forEachBytesField(entry, func(num protowire.Number, _ []byte) error {
if num == hashesFieldNumber {
return add(decodedHashSize)
}
if num == mtimeFieldNumber || num == ctimeFieldNumber {
return add(decodedTimestampSize)
}
if num == mimeTypeFieldNumber {
return add(decodedMIMETypeSize)
}
return nil
})
})
}
// forEachBytesField calls fn with the number and value of each
// length-delimited field in the encoded message msg, and fails if msg is
// malformed.
func forEachBytesField(
msg []byte, fn func(num protowire.Number, value []byte) error,
) error {
for len(msg) > 0 {
num, wireType, tagLen := protowire.ConsumeTag(msg)
if tagLen < 0 {
return protowire.ParseError(tagLen)
}
valueLen := protowire.ConsumeFieldValue(num, wireType, msg[tagLen:])
if valueLen < 0 {
return protowire.ParseError(valueLen)
}
if wireType == protowire.BytesType {
value, _ := protowire.ConsumeBytes(msg[tagLen:])
err := fn(num, value)
if err != nil {
return err
}
}
msg = msg[tagLen+valueLen:]
}
return nil
}
func (m *manifest) deserializeInner() error {
err := m.validateOuterHeader()
if err != nil {
return err
}
if int64(len(dat)) >= MaxDecompressedSize {
return fmt.Errorf("decompressed data exceeds maximum allowed size of %d bytes", MaxDecompressedSize)
err = m.verifyOuterIntegrity()
if err != nil {
return err
}
dat, err := m.decompressInner()
if err != nil {
return err
}
isize := len(dat)
if int64(isize) != m.pbOuter.Size {
log.Debugf("truncated data, got %d expected %d", isize, m.pbOuter.Size)
if int64(isize) != m.pbOuter.GetSize() {
log.Debugf("truncated data, got %d expected %d", isize, m.pbOuter.GetSize())
return bork.ErrFileTruncated
}
err = checkDecodedSize(dat)
if err != nil {
return fmt.Errorf("deserialize: unmarshal inner: %w", err)
}
// Deserialize inner message
m.pbInner = new(MFFile)
if err := proto.Unmarshal(dat, m.pbInner); err != nil {
return err
// Unknown fields would cost memory; mfer never writes a loaded manifest out.
err = proto.UnmarshalOptions{DiscardUnknown: true}.Unmarshal(dat, m.pbInner)
if err != nil {
return fmt.Errorf("deserialize: unmarshal inner: %w", err)
}
// Validate inner UUID
if err := validateUUID(m.pbInner.Uuid); err != nil {
return errors.New("inner UUID invalid: " + err.Error())
err = validateUUID(m.pbInner.GetUuid())
if err != nil {
return fmt.Errorf("inner UUID invalid: %w", err)
}
// Verify UUIDs match
if !bytes.Equal(m.pbOuter.Uuid, m.pbInner.Uuid) {
return errors.New("outer and inner UUID mismatch")
if !bytes.Equal(m.pbOuter.GetUuid(), m.pbInner.GetUuid()) {
return errUUIDMismatch
}
log.Infof("loaded manifest with %d files", len(m.pbInner.Files))
// Enforce the manifest path invariants on every entry as it is loaded,
// so that no consumer of a manifest — Checker today, any restore or
// extract path tomorrow — acts on a traversal or absolute path from an
// untrusted .mf. Reject loudly on the first offender rather than
// dropping entries, which would let a hostile manifest hide files from a
// check.
for _, f := range m.pbInner.GetFiles() {
err = ValidatePath(f.GetPath())
if err != nil {
return fmt.Errorf("%w: %w", errInvalidManifestPath, err)
}
}
log.Infof("loaded manifest with %d files", len(m.pbInner.GetFiles()))
return nil
}
@@ -122,20 +307,26 @@ func validateMagic(dat []byte) bool {
if len(dat) < ml {
return false
}
got := dat[0:ml]
expected := []byte(MAGIC)
return bytes.Equal(got, expected)
}
// NewManifestFromReader reads a manifest from an io.Reader.
//
//nolint:revive // unexported-return: exporting manifest is owner question 13
func NewManifestFromReader(input io.Reader) (*manifest, error) {
m := &manifest{}
dat, err := io.ReadAll(input)
if err != nil {
return nil, err
}
if !validateMagic(dat) {
return nil, errors.New("invalid file format")
return nil, errInvalidFileFormat
}
// remove magic bytes prefix:
@@ -145,28 +336,50 @@ func NewManifestFromReader(input io.Reader) (*manifest, error) {
// deserialize outer:
m.pbOuter = new(MFFileOuter)
if err := proto.Unmarshal(dat, m.pbOuter); err != nil {
// Unknown fields would cost memory; mfer never writes a loaded manifest out.
err = proto.UnmarshalOptions{DiscardUnknown: true}.Unmarshal(dat, m.pbOuter)
if err != nil {
return nil, err
}
// deserialize inner:
if err := m.deserializeInner(); err != nil {
err = m.deserializeInner()
if err != nil {
return nil, err
}
return m, nil
}
// NewManifestFromFile reads a manifest from a file path using the given filesystem.
// If fs is nil, the real filesystem (OsFs) is used.
func NewManifestFromFile(fs afero.Fs, path string) (*manifest, error) {
// ManifestFromFileOptions configures NewManifestFromFile.
type ManifestFromFileOptions struct {
// Path is the manifest file to read (required).
Path string
// Fs is the filesystem to use, defaults to OsFs if nil.
Fs afero.Fs
}
// NewManifestFromFile reads a manifest from a file. It returns an error if
// opts is nil or its path is empty.
//
//nolint:revive // unexported-return: exporting manifest is owner question 13
func NewManifestFromFile(opts *ManifestFromFileOptions) (*manifest, error) {
if opts == nil || opts.Path == "" {
return nil, errManifestPathEmpty
}
fs := opts.Fs
if fs == nil {
fs = afero.NewOsFs()
}
f, err := fs.Open(path)
f, err := fs.Open(opts.Path)
if err != nil {
return nil, err
}
defer func() { _ = f.Close() }()
return NewManifestFromReader(f)
}
+90
View File
@@ -0,0 +1,90 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"bytes"
"runtime"
"testing"
"google.golang.org/protobuf/proto"
)
// FuzzNewManifestFromReader feeds arbitrary bytes to the manifest parser.
// `make test` runs it on the seed corpus in
// testdata/fuzz/FuzzNewManifestFromReader; `make fuzz` searches for new
// inputs.
//
// For every input the parser must return a manifest or an error, not both
// and not neither, and must not allocate more than a fixed multiple of its
// input and of the decompressed data it may read, plus room for the
// decoder's window buffers. A panic or a hang fails the test on its own.
func FuzzNewManifestFromReader(f *testing.F) {
// A signed manifest makes the parser write the key and signature to a
// temporary directory and run gpg on them. With gpg off the PATH and
// temporary files kept in the test's own directory, no process is
// started and nothing is written elsewhere; such input ends in an
// error instead.
f.Setenv("PATH", "")
f.Setenv("TMPDIR", f.TempDir())
f.Fuzz(func(t *testing.T, data []byte) {
var before, after runtime.MemStats
runtime.ReadMemStats(&before)
m, err := NewManifestFromReader(bytes.NewReader(data))
runtime.ReadMemStats(&after)
if (m == nil) == (err == nil) {
t.Fatalf("got manifest %p and error %v, want exactly one", m, err)
}
// The parser reads at most the declared size plus one byte of
// decompressed data, and never more than MaxDecompressedSize.
decompressed := uint64(MaxDecompressedSize)
outer := new(MFFileOuter)
if validateMagic(data) &&
proto.Unmarshal(data[len(MAGIC):], outer) == nil {
size := outer.GetSize()
if size > 0 && size < MaxDecompressedSize {
decompressed = uint64(size) + 1
}
}
// It also keeps a few copies of its input. Buffers grow by
// copying, so reaching those sizes allocates up to about six times
// them in total. Decoding the decompressed data takes up to
// maxDecodedGrowth times its size for file entries, hashes,
// timestamps and MIME types, and drops fields it does not know.
// The strings and bytes it copies out of it, such as many one-byte
// values in one hash, take up to about five times more under the
// race detector, which pads every small copy to 16 bytes, and about
// half that without it. Twenty times the input and the
// decompressed data leaves room for all of that.
//
// The decoder also sets aside a new buffer of one to two times the
// window for each frame that asks for a larger window than the
// frames before it, and refuses windows above zstdWindowSize. A
// frame that gives its content size instead of a window has that
// size as its window, refused above zstdWindowSize like any other.
// Frames asking for every window size up to that make it set aside
// about 16 times zstdWindowSize in total; 24 times leaves room.
//
// The seed whose frame claims 8 GiB fails if the decoder sets that
// size aside; the seed whose frames ask for ever larger windows
// fails if the decoder accepts windows of twice zstdWindowSize; the
// seed whose two frames together exceed MaxDecompressedSize fails
// if the decoder decodes them in full instead of stopping at the
// declared size; the seeds of empty file entries and of a file
// entry of empty hashes fail if the parser decodes them.
limit := 20*(uint64(len(data))+decompressed) + 24*zstdWindowSize
allocated := after.TotalAlloc - before.TotalAlloc
if allocated > limit {
t.Fatalf("allocated %d bytes for %d bytes of input, limit %d",
allocated, len(data), limit)
}
})
}
+259
View File
@@ -0,0 +1,259 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"bytes"
"context"
"crypto/sha256"
"fmt"
"strconv"
"strings"
"testing"
"time"
"github.com/google/uuid"
"github.com/klauspost/compress/zstd"
"github.com/multiformats/go-multihash"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
"google.golang.org/protobuf/encoding/protowire"
"google.golang.org/protobuf/proto"
)
// craftInnerBytes builds the wire bytes of an inner MFFile holding a single
// file entry whose path is exactly pathBytes. It writes the wire form by hand
// so a hostile path — including one that is not valid UTF-8 — can be embedded
// without proto.Marshal's own UTF-8 enforcement rejecting it first.
func craftInnerBytes(id uuid.UUID, pathBytes string) []byte {
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
entry = protowire.AppendString(entry, pathBytes)
inner := protowire.AppendTag(nil, 100, protowire.VarintType) // MFFile.version
inner = protowire.AppendVarint(inner, uint64(MFFile_VERSION_ONE))
inner = protowire.AppendTag(inner, 101, protowire.BytesType) // MFFile.files
inner = protowire.AppendBytes(inner, entry)
inner = protowire.AppendTag(inner, 102, protowire.BytesType) // MFFile.uuid
inner = protowire.AppendBytes(inner, id[:])
return inner
}
// wrapInner wraps inner MFFile wire bytes in a complete, well-formed .mf
// envelope (magic prefix, zstd-compressed payload, matching hash and UUID) so
// that deserialization reaches path validation rather than failing earlier on
// an integrity check.
func wrapInner(t *testing.T, id uuid.UUID, innerData []byte) []byte {
t.Helper()
var cbuf bytes.Buffer
zw, err := zstd.NewWriter(&cbuf, zstd.WithEncoderLevel(zstd.SpeedBestCompression))
require.NoError(t, err)
_, err = zw.Write(innerData)
require.NoError(t, err)
require.NoError(t, zw.Close())
compressed := cbuf.Bytes()
sum := sha256.Sum256(compressed)
outer := &MFFileOuter{
InnerMessage: compressed,
Size: int64(len(innerData)),
Sha256: sum[:],
Uuid: id[:],
Version: MFFileOuter_VERSION_ONE,
CompressionType: MFFileOuter_COMPRESSION_ZSTD,
}
ob, err := proto.Marshal(outer)
require.NoError(t, err)
return append([]byte(MAGIC), ob...)
}
func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
t.Parallel()
tests := []struct {
name string
path string
}{
{"parent traversal", "../escape"},
{"interior traversal", "a/../../escape"},
{"absolute path", "/etc/passwd"},
{"backslash path", `a\b`},
{"double slash", "a//b"},
{"empty path", ""},
{"invalid utf-8", "abc\xff"},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
id := uuid.New()
data := wrapInner(t, id, craftInnerBytes(id, tt.path))
_, err := NewManifestFromReader(bytes.NewReader(data))
require.Error(t, err)
if tt.path == "abc\xff" {
// A path that is not valid UTF-8 cannot survive the proto3
// string decoder, which rejects it before path validation
// runs; the manifest is still refused at load time.
return
}
require.ErrorIs(t, err, errInvalidManifestPath)
if tt.path != "" {
// ValidatePath quotes the path with %q; assert against the
// same rendering so escaped characters (e.g. a backslash)
// still match.
assert.Contains(t, err.Error(), fmt.Sprintf("%q", tt.path),
"error must name the offending path")
}
})
}
}
// Entries of a path, an empty hash, an empty MIME type and empty modification
// and change times are counted at 416 bytes each (160 + 112 + 16 + 64 + 64)
// and take 16 bytes plus the path to encode. A 35-character path makes that
// 51 bytes, about 8.2 times: refused, and leaving any one of the five
// uncounted, even the MIME type, brings it under 8. A 37-character path makes
// it 53 bytes, about 7.8 times: loaded.
func TestDeserializeRefusesEntriesThatDecodeTooLarge(t *testing.T) {
t.Parallel()
tests := []struct {
pathLen int
refused bool
}{
{35, true},
{37, false},
}
for _, tt := range tests {
t.Run(strconv.Itoa(tt.pathLen), func(t *testing.T) {
t.Parallel()
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
entry = protowire.AppendString(entry, strings.Repeat("a", tt.pathLen))
entry = protowire.AppendTag(entry, 3, protowire.BytesType) // MFFilePath.hashes
entry = protowire.AppendBytes(entry, nil)
entry = protowire.AppendTag(entry, 301, protowire.BytesType) // MFFilePath.mimeType
entry = protowire.AppendBytes(entry, nil)
entry = protowire.AppendTag(entry, 302, protowire.BytesType) // MFFilePath.mtime
entry = protowire.AppendBytes(entry, nil)
entry = protowire.AppendTag(entry, 303, protowire.BytesType) // MFFilePath.ctime
entry = protowire.AppendBytes(entry, nil)
id := uuid.New()
inner := protowire.AppendTag(nil, 102, protowire.BytesType) // MFFile.uuid
inner = protowire.AppendBytes(inner, id[:])
for range 1000 {
inner = protowire.AppendTag(inner, 101, protowire.BytesType) // MFFile.files
inner = protowire.AppendBytes(inner, entry)
}
_, err := NewManifestFromReader(bytes.NewReader(wrapInner(t, id, inner)))
if tt.refused {
require.ErrorIs(t, err, errDecodedTooLarge)
} else {
require.NoError(t, err)
}
})
}
}
// Fields the decoder does not know are dropped, in the outer message, the inner
// message and a file entry, so that they take no memory once loaded.
func TestDeserializeDropsUnknownFields(t *testing.T) {
t.Parallel()
unknown := protowire.AppendTag(nil, 99, protowire.BytesType) // in no message
unknown = protowire.AppendBytes(unknown, []byte("not known"))
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
entry = protowire.AppendString(entry, "a")
entry = append(entry, unknown...)
id := uuid.New()
inner := protowire.AppendTag(nil, 101, protowire.BytesType) // MFFile.files
inner = protowire.AppendBytes(inner, entry)
inner = protowire.AppendTag(inner, 102, protowire.BytesType) // MFFile.uuid
inner = protowire.AppendBytes(inner, id[:])
inner = append(inner, unknown...)
data := wrapInner(t, id, inner)
data = append(data, unknown...) // the outer message ends the file
m, err := NewManifestFromReader(bytes.NewReader(data))
require.NoError(t, err)
require.Len(t, m.Files(), 1)
assert.Empty(t, m.pbOuter.ProtoReflect().GetUnknown())
assert.Empty(t, m.pbInner.ProtoReflect().GetUnknown())
assert.Empty(t, m.Files()[0].ProtoReflect().GetUnknown())
}
// Many empty files with names of at most three characters and modification
// times at the epoch make about the densest manifest mfer writes: it takes
// about 7 times its size to decode, and still loads. A signature would not
// change the inner message, so none is added.
func TestDeserializeLoadsDensestManifest(t *testing.T) {
t.Parallel()
hash, err := multihash.Encode(make([]byte, sha256.Size), multihash.SHA2_256)
require.NoError(t, err)
b := NewBuilder()
b.SetIncludeTimestamps(true)
const files = 10000
for i := range files {
name := RelFilePath(strconv.FormatInt(int64(i), 36))
require.NoError(t, b.AddFileWithHash(name, 0, ModTime(time.Unix(0, 0)), hash))
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.Len(t, m.Files(), files)
}
func TestDeserializeValidManifestRoundTrips(t *testing.T) {
t.Parallel()
hash, err := multihash.Encode(make([]byte, sha256.Size), multihash.SHA2_256)
require.NoError(t, err)
b := NewBuilder()
require.NoError(t, b.AddFileWithHash("dir/file.txt", 123, ModTime{}, hash))
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(bytes.NewReader(buf.Bytes()))
require.NoError(t, err)
files := m.Files()
require.Len(t, files, 1)
assert.Equal(t, "dir/file.txt", files[0].GetPath())
assert.Equal(t, int64(123), files[0].GetSize())
}
// TestValidatePathRejectsInvalidUTF8 pins the ValidatePath rule that a manifest
// path must be valid UTF-8, independent of the proto decoder that also enforces
// it on the wire.
func TestValidatePathRejectsInvalidUTF8(t *testing.T) {
t.Parallel()
err := ValidatePath("abc\xff")
require.ErrorIs(t, err, errPathNotUTF8)
assert.Contains(t, err.Error(), "UTF-8")
}
+87
View File
@@ -0,0 +1,87 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"context"
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
// TestValidatePathMessagesVerbatim pins the exact rendered text of every
// ValidatePath rejection.
//
// These strings are user-visible and are assembled by wrapping static
// sentinels mid-sentence, which makes them easy to reword by accident
// while refactoring for errors.Is matchability. Changing one is a
// deliberate change, not a refactoring side effect.
func TestValidatePathMessagesVerbatim(t *testing.T) {
t.Parallel()
for _, tc := range []struct {
name string
path string
want string
is error
}{
{
name: "empty",
path: "",
want: "path cannot be empty",
is: errPathEmpty,
},
{
name: "not utf8",
path: "a\xffb",
want: `path "a\xffb" is not valid UTF-8`,
is: errPathNotUTF8,
},
{
name: "backslash",
path: `a\b`,
want: `path "a\\b" contains backslash; ` +
"use forward slashes only",
is: errPathBackslash,
},
{
name: "absolute",
path: "/a/b",
want: `path "/a/b" is absolute; must be relative`,
is: errPathAbsolute,
},
{
name: "empty segment",
path: "a//b",
want: `path "a//b" contains empty segment`,
is: errPathEmptySegment,
},
{
name: "dotdot segment",
path: "a/../b",
want: `path "a/../b" contains '..' segment`,
is: errPathDotDot,
},
} {
t.Run(tc.name, func(t *testing.T) {
t.Parallel()
err := ValidatePath(tc.path)
require.Error(t, err)
assert.Equal(t, tc.want, err.Error())
require.ErrorIs(t, err, tc.is)
})
}
}
// TestSerializeInternalErrorMessagesVerbatim pins the two distinct
// "internal error" messages, which differ between generate and
// generateOuter and have always done so.
func TestSerializeInternalErrorMessagesVerbatim(t *testing.T) {
t.Parallel()
m := &manifest{}
require.EqualError(t, m.generate(context.Background()),
"internal error: pbInner not set")
require.EqualError(t, m.generateOuter(context.Background()), "internal error")
}
+186 -89
View File
@@ -2,11 +2,55 @@ package mfer
import (
"bytes"
"context"
"errors"
"fmt"
"io"
"os"
"os/exec"
"path/filepath"
"strings"
"time"
)
const (
// gpgTimeout bounds every gpg run, which can otherwise wait forever on
// a passphrase prompt or a stalled gpg-agent. A minute leaves a person
// time to type a passphrase or touch a smartcard.
gpgTimeout = time.Minute
// gpgWaitDelay is how long a gpg run keeps waiting for gpg's stdout
// and stderr to close once gpg has been killed or has exited. Reading
// what gpg itself wrote takes far less; only a process gpg left behind
// holds them open longer.
gpgWaitDelay = time.Second
// privateDirPerms is the permission mode for temporary GPG home
// directories.
privateDirPerms os.FileMode = 0o700
// privateFilePerms is the permission mode for temporary key,
// signature, and data files.
privateFilePerms os.FileMode = 0o600
// gpgFingerprintField is the record type tag for fingerprint lines
// in gpg --with-colons output.
gpgFingerprintField = "fpr"
// gpgFingerprintMinFields is the minimum number of colon-separated
// fields in a gpg fingerprint record (the fingerprint is field 10).
gpgFingerprintMinFields = 10
// gpg option names used from more than one call site.
gpgOptArmor = "--armor"
gpgOptHomedir = "--homedir"
gpgOptVerify = "--verify"
)
var (
errGPGKeyNotFound = errors.New("gpg key not found")
errFingerprintNotFound = errors.New("fingerprint not found for key")
errImportedFPRNotFound = errors.New("fingerprint not found in imported key")
)
// GPGKeyID represents a GPG key identifier (fingerprint or key ID).
@@ -17,22 +61,90 @@ type SigningOptions struct {
KeyID GPGKeyID
}
// gpgSign creates a detached signature of the data using the specified key.
// Returns the armored detached signature.
func gpgSign(data []byte, keyID GPGKeyID) ([]byte, error) {
cmd := exec.Command("gpg",
"--detach-sign",
"--armor",
"--local-user", string(keyID),
)
// gpgArgs builds a gpg argument list from opts followed by positional
// arguments, separated by an explicit "--" end-of-options marker.
//
// This matters because key IDs reach gpg as bare positional arguments
// (from --sign-key / MFER_SIGN_KEY) and gpg would otherwise parse a value
// beginning with "-" as one of its own options. Callers must route every
// non-option argument through here.
func gpgArgs(opts []string, positional ...string) []string {
args := make([]string, 0, len(opts)+1+len(positional))
args = append(args, opts...)
args = append(args, "--")
args = append(args, positional...)
cmd.Stdin = bytes.NewReader(data)
return args
}
// runGPG runs the gpg binary in batch mode with the given arguments and
// optional stdin, returning captured stdout and stderr. gpg is killed when
// ctx ends or gpgTimeout passes, whichever comes first.
func runGPG(
ctx context.Context, stdin io.Reader, args ...string,
) (*bytes.Buffer, *bytes.Buffer, error) {
// exec.CommandContext kills only gpg itself. A gpg-agent that gpg
// starts runs detached and holds none of gpg's output, but another
// process gpg leaves behind (a wrapper script that runs the real gpg
// without exec, for example) can keep gpg's stdout or stderr open, and
// Run would wait for it to exit. WaitDelay stops that wait
// gpgWaitDelay after the kill; that process is left running.
ctx, cancel := context.WithTimeout(ctx, gpgTimeout)
defer cancel()
fullArgs := append([]string{"--batch", "--no-tty"}, args...)
// G204: the executable name is a compile-time constant. The arguments
// are not, so the guarantee that matters is placement: every
// caller-supplied value is passed either as the value of a named
// option or after the "--" end-of-options marker inserted by gpgArgs,
// and therefore cannot be reinterpreted by gpg as an option.
cmd := exec.CommandContext( //nolint:gosec // G204: see comment above
ctx, "gpg", fullArgs...)
cmd.WaitDelay = gpgWaitDelay
cmd.Stdin = stdin
var stdout, stderr bytes.Buffer
cmd.Stdout = &stdout
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
err := cmd.Run()
if err != nil && ctx.Err() != nil {
// gpg was killed because ctx ended, which Run reports only as
// "signal: killed"; return the reason instead.
err = ctx.Err()
if errors.Is(err, context.DeadlineExceeded) {
err = fmt.Errorf("gpg timed out: %w", err)
}
}
return &stdout, &stderr, err
}
// parseFingerprint extracts the first fingerprint from gpg --with-colons
// output, or returns ok=false if none is present.
func parseFingerprint(colonOutput string) (string, bool) {
for _, line := range strings.Split(colonOutput, "\n") {
fields := strings.Split(line, ":")
if len(fields) >= gpgFingerprintMinFields &&
fields[0] == gpgFingerprintField {
return fields[9], true
}
}
return "", false
}
// gpgSign creates a detached signature of the data using the specified key.
// Returns the armored detached signature.
func gpgSign(ctx context.Context, data []byte, keyID GPGKeyID) ([]byte, error) {
stdout, stderr, err := runGPG(ctx, bytes.NewReader(data),
"--detach-sign",
gpgOptArmor,
"--local-user", string(keyID),
)
if err != nil {
return nil, fmt.Errorf("gpg sign failed: %w: %s", err, stderr.String())
}
@@ -41,171 +153,156 @@ func gpgSign(data []byte, keyID GPGKeyID) ([]byte, error) {
// gpgExportPublicKey exports the public key for the specified key ID.
// Returns the armored public key.
func gpgExportPublicKey(keyID GPGKeyID) ([]byte, error) {
cmd := exec.Command("gpg",
"--export",
"--armor",
string(keyID),
func gpgExportPublicKey(ctx context.Context, keyID GPGKeyID) ([]byte, error) {
stdout, stderr, err := runGPG(ctx, nil,
gpgArgs([]string{"--export", gpgOptArmor}, string(keyID))...,
)
var stdout, stderr bytes.Buffer
cmd.Stdout = &stdout
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
if err != nil {
return nil, fmt.Errorf("gpg export failed: %w: %s", err, stderr.String())
}
if stdout.Len() == 0 {
return nil, fmt.Errorf("gpg key not found: %s", keyID)
return nil, fmt.Errorf("%w: %s", errGPGKeyNotFound, keyID)
}
return stdout.Bytes(), nil
}
// gpgGetKeyFingerprint gets the full fingerprint for a key ID.
func gpgGetKeyFingerprint(keyID GPGKeyID) ([]byte, error) {
cmd := exec.Command("gpg",
"--with-colons",
"--fingerprint",
string(keyID),
func gpgGetKeyFingerprint(ctx context.Context, keyID GPGKeyID) ([]byte, error) {
stdout, stderr, err := runGPG(ctx, nil,
gpgArgs([]string{"--with-colons", "--fingerprint"}, string(keyID))...,
)
var stdout, stderr bytes.Buffer
cmd.Stdout = &stdout
cmd.Stderr = &stderr
if err := cmd.Run(); err != nil {
return nil, fmt.Errorf("gpg fingerprint lookup failed: %w: %s", err, stderr.String())
if err != nil {
return nil, fmt.Errorf(
"gpg fingerprint lookup failed: %w: %s", err, stderr.String(),
)
}
// Parse the colon-delimited output to find the fingerprint
lines := strings.Split(stdout.String(), "\n")
for _, line := range lines {
fields := strings.Split(line, ":")
if len(fields) >= 10 && fields[0] == "fpr" {
return []byte(fields[9]), nil
}
fpr, ok := parseFingerprint(stdout.String())
if !ok {
return nil, fmt.Errorf("%w: %s", errFingerprintNotFound, keyID)
}
return nil, fmt.Errorf("fingerprint not found for key: %s", keyID)
return []byte(fpr), nil
}
// gpgExtractPubKeyFingerprint imports a public key into a temporary keyring
// and extracts its fingerprint. This verifies the key is valid and returns
// the actual fingerprint from the key material.
func gpgExtractPubKeyFingerprint(pubKey []byte) (string, error) {
func gpgExtractPubKeyFingerprint(ctx context.Context, pubKey []byte) (string, error) {
// Create temporary directory for GPG operations
tmpDir, err := os.MkdirTemp("", "mfer-gpg-fingerprint-*")
if err != nil {
return "", fmt.Errorf("failed to create temp dir: %w", err)
}
defer func() { _ = os.RemoveAll(tmpDir) }()
// Set restrictive permissions
if err := os.Chmod(tmpDir, 0o700); err != nil {
err = os.Chmod(tmpDir, privateDirPerms)
if err != nil {
return "", fmt.Errorf("failed to set temp dir permissions: %w", err)
}
// Write public key to temp file
pubKeyFile := filepath.Join(tmpDir, "pubkey.asc")
if err := os.WriteFile(pubKeyFile, pubKey, 0o600); err != nil {
err = os.WriteFile(pubKeyFile, pubKey, privateFilePerms)
if err != nil {
return "", fmt.Errorf("failed to write public key: %w", err)
}
// Import the public key into the temporary keyring
importCmd := exec.Command("gpg",
"--homedir", tmpDir,
"--import",
pubKeyFile,
_, importStderr, err := runGPG(ctx, nil,
gpgArgs([]string{gpgOptHomedir, tmpDir, "--import"}, pubKeyFile)...,
)
var importStderr bytes.Buffer
importCmd.Stderr = &importStderr
if err := importCmd.Run(); err != nil {
return "", fmt.Errorf("failed to import public key: %w: %s", err, importStderr.String())
if err != nil {
return "", fmt.Errorf(
"failed to import public key: %w: %s", err, importStderr.String(),
)
}
// List keys to get fingerprint
listCmd := exec.Command("gpg",
listStdout, listStderr, err := runGPG(ctx, nil,
"--homedir", tmpDir,
"--with-colons",
"--fingerprint",
)
var listStdout, listStderr bytes.Buffer
listCmd.Stdout = &listStdout
listCmd.Stderr = &listStderr
if err := listCmd.Run(); err != nil {
return "", fmt.Errorf("failed to list keys: %w: %s", err, listStderr.String())
if err != nil {
return "", fmt.Errorf(
"failed to list keys: %w: %s", err, listStderr.String(),
)
}
// Parse the colon-delimited output to find the fingerprint
lines := strings.Split(listStdout.String(), "\n")
for _, line := range lines {
fields := strings.Split(line, ":")
if len(fields) >= 10 && fields[0] == "fpr" {
return fields[9], nil
}
fpr, ok := parseFingerprint(listStdout.String())
if !ok {
return "", errImportedFPRNotFound
}
return "", fmt.Errorf("fingerprint not found in imported key")
return fpr, nil
}
// gpgVerify verifies a detached signature against data using the provided public key.
// It creates a temporary keyring to import the public key for verification.
func gpgVerify(data, signature, pubKey []byte) error {
func gpgVerify(ctx context.Context, data, signature, pubKey []byte) error {
// Create temporary directory for GPG operations
tmpDir, err := os.MkdirTemp("", "mfer-gpg-verify-*")
if err != nil {
return fmt.Errorf("failed to create temp dir: %w", err)
}
defer func() { _ = os.RemoveAll(tmpDir) }()
// Set restrictive permissions
if err := os.Chmod(tmpDir, 0o700); err != nil {
err = os.Chmod(tmpDir, privateDirPerms)
if err != nil {
return fmt.Errorf("failed to set temp dir permissions: %w", err)
}
// Write public key to temp file
pubKeyFile := filepath.Join(tmpDir, "pubkey.asc")
if err := os.WriteFile(pubKeyFile, pubKey, 0o600); err != nil {
err = os.WriteFile(pubKeyFile, pubKey, privateFilePerms)
if err != nil {
return fmt.Errorf("failed to write public key: %w", err)
}
// Write signature to temp file
sigFile := filepath.Join(tmpDir, "signature.asc")
if err := os.WriteFile(sigFile, signature, 0o600); err != nil {
err = os.WriteFile(sigFile, signature, privateFilePerms)
if err != nil {
return fmt.Errorf("failed to write signature: %w", err)
}
// Write data to temp file
dataFile := filepath.Join(tmpDir, "data")
if err := os.WriteFile(dataFile, data, 0o600); err != nil {
err = os.WriteFile(dataFile, data, privateFilePerms)
if err != nil {
return fmt.Errorf("failed to write data: %w", err)
}
// Import the public key into the temporary keyring
importCmd := exec.Command("gpg",
"--homedir", tmpDir,
"--import",
pubKeyFile,
_, importStderr, err := runGPG(ctx, nil,
gpgArgs([]string{gpgOptHomedir, tmpDir, "--import"}, pubKeyFile)...,
)
var importStderr bytes.Buffer
importCmd.Stderr = &importStderr
if err := importCmd.Run(); err != nil {
return fmt.Errorf("failed to import public key: %w: %s", err, importStderr.String())
if err != nil {
return fmt.Errorf(
"failed to import public key: %w: %s", err, importStderr.String(),
)
}
// Verify the signature
verifyCmd := exec.Command("gpg",
"--homedir", tmpDir,
"--verify",
sigFile,
dataFile,
_, verifyStderr, err := runGPG(ctx, nil,
gpgArgs([]string{gpgOptHomedir, tmpDir, gpgOptVerify},
sigFile, dataFile)...,
)
var verifyStderr bytes.Buffer
verifyCmd.Stderr = &verifyStderr
if err := verifyCmd.Run(); err != nil {
return fmt.Errorf("signature verification failed: %w: %s", err, verifyStderr.String())
if err != nil {
return fmt.Errorf(
"signature verification failed: %w: %s", err, verifyStderr.String(),
)
}
return nil
+224 -83
View File
@@ -1,13 +1,18 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"bytes"
"context"
"io"
"os"
"os/exec"
"path/filepath"
"strconv"
"strings"
"syscall"
"testing"
"time"
"github.com/spf13/afero"
"github.com/stretchr/testify/assert"
@@ -15,35 +20,20 @@ import (
)
// testGPGEnv sets up a temporary GPG home directory with a test key.
// Returns the key ID and a cleanup function.
func testGPGEnv(t *testing.T) (GPGKeyID, func()) {
// Returns the key ID and the GPG home directory; callers must point
// GNUPGHOME at the returned directory (via t.Setenv) before using the
// gpg helpers under test.
func testGPGEnv(t *testing.T) (GPGKeyID, string) {
t.Helper()
// Check if gpg is installed
if _, err := exec.LookPath("gpg"); err != nil {
_, err := exec.LookPath("gpg")
if err != nil {
t.Skip("gpg not installed, skipping signing test")
return "", func() {}
}
// Create temporary GPG home directory
gpgHome, err := os.MkdirTemp("", "mfer-gpg-test-*")
require.NoError(t, err)
// Set restrictive permissions on GPG home
require.NoError(t, os.Chmod(gpgHome, 0o700))
// Save original GNUPGHOME and set new one
origGPGHome := os.Getenv("GNUPGHOME")
require.NoError(t, os.Setenv("GNUPGHOME", gpgHome))
cleanup := func() {
if origGPGHome == "" {
_ = os.Unsetenv("GNUPGHOME")
} else {
_ = os.Setenv("GNUPGHOME", origGPGHome)
}
_ = os.RemoveAll(gpgHome)
}
// Create temporary GPG home directory (0700 by default)
gpgHome := t.TempDir()
// Generate a test key with no passphrase
keyParams := `%no-protection
@@ -57,48 +47,57 @@ Expire-Date: 0
paramsFile := filepath.Join(gpgHome, "key-params")
require.NoError(t, os.WriteFile(paramsFile, []byte(keyParams), 0o600))
cmd := exec.Command("gpg", "--batch", "--gen-key", paramsFile)
ctx, cancel := context.WithTimeout(context.Background(), gpgTimeout)
defer cancel()
//nolint:gosec // paramsFile is a test-controlled path inside t.TempDir()
cmd := exec.CommandContext(ctx, "gpg",
"--batch", "--gen-key", paramsFile)
cmd.Env = append(os.Environ(), "GNUPGHOME="+gpgHome)
output, err := cmd.CombinedOutput()
if err != nil {
cleanup()
t.Skipf("failed to generate test GPG key: %v: %s", err, output)
return "", func() {}
}
// Get the key fingerprint
cmd = exec.Command("gpg", "--list-keys", "--with-colons", "test@mfer.test")
cmd = exec.CommandContext(ctx, "gpg",
"--list-keys", "--with-colons", "test@mfer.test")
cmd.Env = append(os.Environ(), "GNUPGHOME="+gpgHome)
output, err = cmd.Output()
if err != nil {
cleanup()
t.Fatalf("failed to list test key: %v", err)
}
// Parse fingerprint from output
var keyID string
for _, line := range strings.Split(string(output), "\n") {
fields := strings.Split(line, ":")
if len(fields) >= 10 && fields[0] == "fpr" {
if len(fields) >= gpgFingerprintMinFields &&
fields[0] == gpgFingerprintField {
keyID = fields[9]
break
}
}
if keyID == "" {
cleanup()
t.Fatal("failed to find test key fingerprint")
}
return GPGKeyID(keyID), cleanup
return GPGKeyID(keyID), gpgHome
}
func TestGPGSign(t *testing.T) {
keyID, cleanup := testGPGEnv(t)
defer cleanup()
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
data := []byte("test data to sign")
sig, err := gpgSign(data, keyID)
sig, err := gpgSign(context.Background(), data, keyID)
require.NoError(t, err)
assert.NotEmpty(t, sig)
assert.Contains(t, string(sig), "-----BEGIN PGP SIGNATURE-----")
@@ -106,10 +105,10 @@ func TestGPGSign(t *testing.T) {
}
func TestGPGExportPublicKey(t *testing.T) {
keyID, cleanup := testGPGEnv(t)
defer cleanup()
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
pubKey, err := gpgExportPublicKey(keyID)
pubKey, err := gpgExportPublicKey(context.Background(), keyID)
require.NoError(t, err)
assert.NotEmpty(t, pubKey)
assert.Contains(t, string(pubKey), "-----BEGIN PGP PUBLIC KEY BLOCK-----")
@@ -117,29 +116,67 @@ func TestGPGExportPublicKey(t *testing.T) {
}
func TestGPGGetKeyFingerprint(t *testing.T) {
keyID, cleanup := testGPGEnv(t)
defer cleanup()
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
fingerprint, err := gpgGetKeyFingerprint(keyID)
fingerprint, err := gpgGetKeyFingerprint(context.Background(), keyID)
require.NoError(t, err)
assert.NotEmpty(t, fingerprint)
// The fingerprint should be 40 hex chars
assert.Len(t, fingerprint, 40, "fingerprint should be 40 hex chars")
}
// TestGPGArgsSeparatesPositionals pins that caller-supplied values are
// placed after an end-of-options marker. Key IDs arrive from --sign-key
// and MFER_SIGN_KEY as bare positional arguments, so without the marker
// a value beginning with "-" would be parsed by gpg as one of its own
// options.
func TestGPGArgsSeparatesPositionals(t *testing.T) {
t.Parallel()
assert.Equal(t,
[]string{"--opt-a", "--opt-b", "--", "--version"},
gpgArgs([]string{"--opt-a", "--opt-b"}, "--version"))
assert.Equal(t,
[]string{"--opt-c", "--", "sig", "data"},
gpgArgs([]string{"--opt-c"}, "sig", "data"))
assert.Equal(t, []string{"--opt-d", "--"},
gpgArgs([]string{"--opt-d"}))
}
// TestGPGOptionLikeKeyIDIsNotAnOption drives real gpg with a key ID that
// looks like an option and asserts it is treated as a (nonexistent) key
// rather than executed as gpg's own --version.
func TestGPGOptionLikeKeyIDIsNotAnOption(t *testing.T) {
_, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
pubKey, err := gpgExportPublicKey(context.Background(), GPGKeyID("--version"))
require.Error(t, err)
require.ErrorIs(t, err, errGPGKeyNotFound)
assert.NotContains(t, string(pubKey), "gpg (GnuPG)")
fpr, err := gpgGetKeyFingerprint(context.Background(), GPGKeyID("--version"))
require.Error(t, err)
assert.NotContains(t, string(fpr), "gpg (GnuPG)")
}
func TestGPGSignInvalidKey(t *testing.T) {
// Set up test environment (we need GNUPGHOME set)
_, cleanup := testGPGEnv(t)
defer cleanup()
_, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
data := []byte("test data")
_, err := gpgSign(data, GPGKeyID("NONEXISTENT_KEY_ID_12345"))
_, err := gpgSign(context.Background(), data,
GPGKeyID("NONEXISTENT_KEY_ID_12345"))
assert.Error(t, err)
}
func TestBuilderWithSigning(t *testing.T) {
keyID, cleanup := testGPGEnv(t)
defer cleanup()
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
// Create a builder with signing options
b := NewBuilder()
@@ -155,7 +192,8 @@ func TestBuilderWithSigning(t *testing.T) {
// Build the manifest
var buf bytes.Buffer
err = b.Build(&buf)
err = b.Build(context.Background(), &buf)
require.NoError(t, err)
// Parse the manifest and verify signature fields are populated
@@ -163,26 +201,32 @@ func TestBuilderWithSigning(t *testing.T) {
require.NoError(t, err)
require.NotNil(t, manifest.pbOuter)
assert.NotEmpty(t, manifest.pbOuter.Signature, "signature should be populated")
assert.NotEmpty(t, manifest.pbOuter.Signer, "signer should be populated")
assert.NotEmpty(t, manifest.pbOuter.SigningPubKey, "signing public key should be populated")
assert.NotEmpty(t, manifest.pbOuter.GetSignature(),
"signature should be populated")
assert.NotEmpty(t, manifest.pbOuter.GetSigner(), "signer should be populated")
assert.NotEmpty(t, manifest.pbOuter.GetSigningPubKey(),
"signing public key should be populated")
// Verify signature is a valid PGP signature
assert.Contains(t, string(manifest.pbOuter.Signature), "-----BEGIN PGP SIGNATURE-----")
assert.Contains(t, string(manifest.pbOuter.GetSignature()),
"-----BEGIN PGP SIGNATURE-----")
// Verify public key is a valid PGP public key block
assert.Contains(t, string(manifest.pbOuter.SigningPubKey), "-----BEGIN PGP PUBLIC KEY BLOCK-----")
assert.Contains(t, string(manifest.pbOuter.GetSigningPubKey()),
"-----BEGIN PGP PUBLIC KEY BLOCK-----")
}
func TestScannerWithSigning(t *testing.T) {
keyID, cleanup := testGPGEnv(t)
defer cleanup()
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
// Create in-memory filesystem with test files
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file1.txt", []byte("content1"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file2.txt", []byte("content2"), 0o644))
// Create scanner with signing options
opts := &ScannerOptions{
@@ -205,61 +249,61 @@ func TestScannerWithSigning(t *testing.T) {
manifest, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.NotEmpty(t, manifest.pbOuter.Signature)
assert.NotEmpty(t, manifest.pbOuter.Signer)
assert.NotEmpty(t, manifest.pbOuter.SigningPubKey)
assert.NotEmpty(t, manifest.pbOuter.GetSignature())
assert.NotEmpty(t, manifest.pbOuter.GetSigner())
assert.NotEmpty(t, manifest.pbOuter.GetSigningPubKey())
}
func TestGPGVerify(t *testing.T) {
keyID, cleanup := testGPGEnv(t)
defer cleanup()
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
data := []byte("test data to sign and verify")
sig, err := gpgSign(data, keyID)
sig, err := gpgSign(context.Background(), data, keyID)
require.NoError(t, err)
pubKey, err := gpgExportPublicKey(keyID)
pubKey, err := gpgExportPublicKey(context.Background(), keyID)
require.NoError(t, err)
// Verify the signature
err = gpgVerify(data, sig, pubKey)
err = gpgVerify(context.Background(), data, sig, pubKey)
require.NoError(t, err)
}
func TestGPGVerifyInvalidSignature(t *testing.T) {
keyID, cleanup := testGPGEnv(t)
defer cleanup()
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
data := []byte("test data to sign")
sig, err := gpgSign(data, keyID)
sig, err := gpgSign(context.Background(), data, keyID)
require.NoError(t, err)
pubKey, err := gpgExportPublicKey(keyID)
pubKey, err := gpgExportPublicKey(context.Background(), keyID)
require.NoError(t, err)
// Try to verify with different data - should fail
wrongData := []byte("different data")
err = gpgVerify(wrongData, sig, pubKey)
err = gpgVerify(context.Background(), wrongData, sig, pubKey)
assert.Error(t, err)
}
func TestGPGVerifyBadPublicKey(t *testing.T) {
keyID, cleanup := testGPGEnv(t)
defer cleanup()
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
data := []byte("test data")
sig, err := gpgSign(data, keyID)
sig, err := gpgSign(context.Background(), data, keyID)
require.NoError(t, err)
// Try to verify with invalid public key - should fail
badPubKey := []byte("not a valid public key")
err = gpgVerify(data, sig, badPubKey)
err = gpgVerify(context.Background(), data, sig, badPubKey)
assert.Error(t, err)
}
func TestManifestSignatureVerification(t *testing.T) {
keyID, cleanup := testGPGEnv(t)
defer cleanup()
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
// Create a builder with signing options
b := NewBuilder()
@@ -275,7 +319,8 @@ func TestManifestSignatureVerification(t *testing.T) {
// Build the manifest
var buf bytes.Buffer
err = b.Build(&buf)
err = b.Build(context.Background(), &buf)
require.NoError(t, err)
// Parse the manifest - signature should be verified during load
@@ -284,12 +329,12 @@ func TestManifestSignatureVerification(t *testing.T) {
require.NotNil(t, manifest)
// Signature should be present and valid
assert.NotEmpty(t, manifest.pbOuter.Signature)
assert.NotEmpty(t, manifest.pbOuter.GetSignature())
}
func TestManifestTamperedSignatureFails(t *testing.T) {
keyID, cleanup := testGPGEnv(t)
defer cleanup()
keyID, gpgHome := testGPGEnv(t)
t.Setenv("GNUPGHOME", gpgHome)
// Create a signed manifest
b := NewBuilder()
@@ -303,7 +348,8 @@ func TestManifestTamperedSignatureFails(t *testing.T) {
require.NoError(t, err)
var buf bytes.Buffer
err = b.Build(&buf)
err = b.Build(context.Background(), &buf)
require.NoError(t, err)
// Tamper with the signature by replacing some bytes
@@ -312,6 +358,7 @@ func TestManifestTamperedSignatureFails(t *testing.T) {
for i := range data {
if i > 100 && data[i] == 'A' {
data[i] = 'B'
break
}
}
@@ -322,6 +369,8 @@ func TestManifestTamperedSignatureFails(t *testing.T) {
}
func TestBuilderWithoutSigning(t *testing.T) {
t.Parallel()
// Create a builder without signing options
b := NewBuilder()
@@ -333,7 +382,8 @@ func TestBuilderWithoutSigning(t *testing.T) {
// Build the manifest
var buf bytes.Buffer
err = b.Build(&buf)
err = b.Build(context.Background(), &buf)
require.NoError(t, err)
// Parse the manifest and verify signature fields are empty
@@ -341,7 +391,98 @@ func TestBuilderWithoutSigning(t *testing.T) {
require.NoError(t, err)
require.NotNil(t, manifest.pbOuter)
assert.Empty(t, manifest.pbOuter.Signature, "signature should be empty when not signing")
assert.Empty(t, manifest.pbOuter.Signer, "signer should be empty when not signing")
assert.Empty(t, manifest.pbOuter.SigningPubKey, "signing public key should be empty when not signing")
assert.Empty(t, manifest.pbOuter.GetSignature(),
"signature should be empty when not signing")
assert.Empty(t, manifest.pbOuter.GetSigner(),
"signer should be empty when not signing")
assert.Empty(t, manifest.pbOuter.GetSigningPubKey(),
"signing public key should be empty when not signing")
}
// fakeGPGPath writes script as an executable named gpg into a temporary
// directory and returns a PATH value with that directory first.
func fakeGPGPath(t *testing.T, script string) string {
t.Helper()
binDir := t.TempDir()
//nolint:gosec // G306: the fake gpg has to be executable
require.NoError(t, os.WriteFile(filepath.Join(binDir, "gpg"),
[]byte(script), 0o700))
return binDir + string(os.PathListSeparator) + os.Getenv("PATH")
}
// TestGPGTimeoutKillsGPG puts a fake gpg that never finishes first on
// PATH and checks that a run past its deadline is killed and reported as
// a timeout of the named operation, instead of hanging.
func TestGPGTimeoutKillsGPG(t *testing.T) {
t.Setenv("PATH", fakeGPGPath(t, "#!/bin/sh\nexec sleep 10\n"))
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
_, err := gpgSign(ctx, []byte("data"), GPGKeyID("any"))
require.ErrorIs(t, err, context.DeadlineExceeded)
assert.Contains(t, err.Error(), "gpg sign failed: gpg timed out")
}
// TestGPGCancelWhenChildHoldsOutput uses a fake gpg that runs sleep as a
// child instead of exec-ing it, the way a wrapper script around the real
// gpg might. Killing the fake gpg leaves sleep holding its stdout and
// stderr open; the call must still return once ctx ends instead of waiting
// for sleep to exit. The fake gpg writes the process ID of sleep to a named
// pipe; the test ends ctx only after reading it, so sleep is running by
// then, and kills sleep before returning.
func TestGPGCancelWhenChildHoldsOutput(t *testing.T) {
pidPipe := filepath.Join(t.TempDir(), "sleep.pid")
require.NoError(t, syscall.Mkfifo(pidPipe, 0o600))
// sleep outlasts the 10 s wait below, so a call that waits for it fails.
t.Setenv("PATH", fakeGPGPath(t,
"#!/bin/sh\nsleep 60 &\necho $! >'"+pidPipe+"'\nwait\n"))
ctx, cancel := context.WithCancel(context.Background())
defer cancel()
signErr := make(chan error, 1)
go func() {
_, err := gpgSign(ctx, []byte("data"), GPGKeyID("any"))
signErr <- err
}()
pid, err := os.ReadFile(pidPipe) //nolint:gosec // G304: path inside t.TempDir()
require.NoError(t, err)
n, err := strconv.Atoi(strings.TrimSpace(string(pid)))
require.NoError(t, err)
sleep, err := os.FindProcess(n)
require.NoError(t, err)
t.Cleanup(func() { require.NoError(t, sleep.Kill()) })
cancel()
// The call should return about gpgWaitDelay (one second) after the
// cancel. 10 s is far above that and well under the 30 s test timeout,
// which would abort the whole package before the cleanup kills sleep.
select {
case err := <-signErr:
require.ErrorIs(t, err, context.Canceled)
case <-time.After(10 * time.Second):
t.Fatal("the call waited for the child holding gpg's output to exit")
}
}
// TestBuildPassesContextToSigning checks that a caller can cancel the gpg
// runs that sign a manifest through the context given to Build.
func TestBuildPassesContextToSigning(t *testing.T) {
t.Parallel()
b := NewBuilder()
b.SetSigningOptions(&SigningOptions{KeyID: "any"})
ctx, cancel := context.WithCancel(context.Background())
cancel()
require.ErrorIs(t, b.Build(ctx, io.Discard), context.Canceled)
}
+28 -13
View File
@@ -9,9 +9,18 @@ import (
"github.com/multiformats/go-multihash"
)
var (
errOuterNotSet = errors.New("pbOuter not set")
errUUIDNotSet = errors.New("UUID not set")
errSHA256NotSet = errors.New("SHA256 hash not set")
)
// manifest holds the internal representation of a manifest file.
// Use NewManifestFromFile or NewManifestFromReader to load an existing manifest,
// or use Builder to create a new one.
// Use NewManifestFromFile or NewManifestFromReader to load an existing
// manifest, or use Builder to create a new one.
//
// Whether this type should be exported is an open design question owned by
// the repository owner; see README design question 13.
type manifest struct {
pbInner *MFFile
pbOuter *MFFileOuter
@@ -23,8 +32,9 @@ type manifest struct {
func (m *manifest) String() string {
count := 0
if m.pbInner != nil {
count = len(m.pbInner.Files)
count = len(m.pbInner.GetFiles())
}
return fmt.Sprintf("<Manifest count=%d>", count)
}
@@ -33,7 +43,8 @@ func (m *manifest) Files() []*MFFilePath {
if m.pbInner == nil {
return nil
}
return m.pbInner.Files
return m.pbInner.GetFiles()
}
// signatureString generates the canonical string used for signing/verification.
@@ -41,20 +52,24 @@ func (m *manifest) Files() []*MFFilePath {
// Requires pbOuter to be set with Uuid and Sha256 fields.
func (m *manifest) signatureString() (string, error) {
if m.pbOuter == nil {
return "", errors.New("pbOuter not set")
}
if len(m.pbOuter.Uuid) == 0 {
return "", errors.New("UUID not set")
}
if len(m.pbOuter.Sha256) == 0 {
return "", errors.New("SHA256 hash not set")
return "", errOuterNotSet
}
mh, err := multihash.Encode(m.pbOuter.Sha256, multihash.SHA2_256)
if len(m.pbOuter.GetUuid()) == 0 {
return "", errUUIDNotSet
}
if len(m.pbOuter.GetSha256()) == 0 {
return "", errSHA256NotSet
}
mh, err := multihash.Encode(m.pbOuter.GetSha256(), multihash.SHA2_256)
if err != nil {
return "", fmt.Errorf("failed to encode multihash: %w", err)
}
uuidStr := hex.EncodeToString(m.pbOuter.Uuid)
uuidStr := hex.EncodeToString(m.pbOuter.GetUuid())
mhStr := hex.EncodeToString(mh)
return fmt.Sprintf("%s-%s-%s", MAGIC, uuidStr, mhStr), nil
}
-3
View File
@@ -1,3 +0,0 @@
package mfer
//go:generate protoc ./mf.proto --go_out=paths=source_relative:.
+16 -25
View File
@@ -1,7 +1,7 @@
// Code generated by protoc-gen-go. DO NOT EDIT.
// versions:
// protoc-gen-go v1.36.11
// protoc v6.33.0
// protoc v6.33.4
// source: mf.proto
package mfer
@@ -329,6 +329,9 @@ func (x *MFFileOuter) GetSigningPubKey() []byte {
type MFFilePath struct {
state protoimpl.MessageState `protogen:"open.v1"`
// required attributes:
// Path invariants: must be valid UTF-8, use forward slashes only,
// be relative (no leading /), contain no ".." segments, and no
// empty segments (no "//").
Path string `protobuf:"bytes,1,opt,name=path,proto3" json:"path,omitempty"`
Size int64 `protobuf:"varint,2,opt,name=size,proto3" json:"size,omitempty"`
// gotta have at least one:
@@ -337,7 +340,6 @@ type MFFilePath struct {
MimeType *string `protobuf:"bytes,301,opt,name=mimeType,proto3,oneof" json:"mimeType,omitempty"`
Mtime *Timestamp `protobuf:"bytes,302,opt,name=mtime,proto3,oneof" json:"mtime,omitempty"`
Ctime *Timestamp `protobuf:"bytes,303,opt,name=ctime,proto3,oneof" json:"ctime,omitempty"`
Atime *Timestamp `protobuf:"bytes,304,opt,name=atime,proto3,oneof" json:"atime,omitempty"`
unknownFields protoimpl.UnknownFields
sizeCache protoimpl.SizeCache
}
@@ -414,13 +416,6 @@ func (x *MFFilePath) GetCtime() *Timestamp {
return nil
}
func (x *MFFilePath) GetAtime() *Timestamp {
if x != nil {
return x.Atime
}
return nil
}
type MFFileChecksum struct {
state protoimpl.MessageState `protogen:"open.v1"`
// 1.0 golang implementation must write a multihash here
@@ -566,7 +561,7 @@ const file_mf_proto_rawDesc = "" +
"\n" +
"_signatureB\t\n" +
"\a_signerB\x10\n" +
"\x0e_signingPubKey\"\xa2\x02\n" +
"\x0e_signingPubKey\"\xf0\x01\n" +
"\n" +
"MFFilePath\x12\x12\n" +
"\x04path\x18\x01 \x01(\tR\x04path\x12\x12\n" +
@@ -576,13 +571,10 @@ const file_mf_proto_rawDesc = "" +
"\x05mtime\x18\xae\x02 \x01(\v2\n" +
".TimestampH\x01R\x05mtime\x88\x01\x01\x12&\n" +
"\x05ctime\x18\xaf\x02 \x01(\v2\n" +
".TimestampH\x02R\x05ctime\x88\x01\x01\x12&\n" +
"\x05atime\x18\xb0\x02 \x01(\v2\n" +
".TimestampH\x03R\x05atime\x88\x01\x01B\v\n" +
".TimestampH\x02R\x05ctime\x88\x01\x01B\v\n" +
"\t_mimeTypeB\b\n" +
"\x06_mtimeB\b\n" +
"\x06_ctimeB\b\n" +
"\x06_atime\".\n" +
"\x06_ctime\".\n" +
"\x0eMFFileChecksum\x12\x1c\n" +
"\tmultiHash\x18\x01 \x01(\fR\tmultiHash\"\xd6\x01\n" +
"\x06MFFile\x12)\n" +
@@ -595,7 +587,7 @@ const file_mf_proto_rawDesc = "" +
"\fVERSION_NONE\x10\x00\x12\x0f\n" +
"\vVERSION_ONE\x10\x01B\f\n" +
"\n" +
"_createdAtB\x1dZ\x1bgit.eeqj.de/sneak/mfer/mferb\x06proto3"
"_createdAtB\x1bZ\x19sneak.berlin/go/mfer/mferb\x06proto3"
var (
file_mf_proto_rawDescOnce sync.Once
@@ -627,15 +619,14 @@ var file_mf_proto_depIdxs = []int32{
6, // 2: MFFilePath.hashes:type_name -> MFFileChecksum
3, // 3: MFFilePath.mtime:type_name -> Timestamp
3, // 4: MFFilePath.ctime:type_name -> Timestamp
3, // 5: MFFilePath.atime:type_name -> Timestamp
2, // 6: MFFile.version:type_name -> MFFile.Version
5, // 7: MFFile.files:type_name -> MFFilePath
3, // 8: MFFile.createdAt:type_name -> Timestamp
9, // [9:9] is the sub-list for method output_type
9, // [9:9] is the sub-list for method input_type
9, // [9:9] is the sub-list for extension type_name
9, // [9:9] is the sub-list for extension extendee
0, // [0:9] is the sub-list for field type_name
2, // 5: MFFile.version:type_name -> MFFile.Version
5, // 6: MFFile.files:type_name -> MFFilePath
3, // 7: MFFile.createdAt:type_name -> Timestamp
8, // [8:8] is the sub-list for method output_type
8, // [8:8] is the sub-list for method input_type
8, // [8:8] is the sub-list for extension type_name
8, // [8:8] is the sub-list for extension extendee
0, // [0:8] is the sub-list for field type_name
}
func init() { file_mf_proto_init() }
+1 -2
View File
@@ -1,6 +1,6 @@
syntax = "proto3";
option go_package = "git.eeqj.de/sneak/mfer/mfer";
option go_package = "sneak.berlin/go/mfer/mfer";
message Timestamp {
int64 seconds = 1;
@@ -59,7 +59,6 @@ message MFFilePath {
optional string mimeType = 301;
optional Timestamp mtime = 302;
optional Timestamp ctime = 303;
optional Timestamp atime = 304;
}
message MFFileChecksum {
+1
View File
@@ -0,0 +1 @@
fa6fceaba5553c8667c631994535fe5c307c8adb19f3ad289babf79a0f438650 mf.proto
+29
View File
@@ -0,0 +1,29 @@
package mfer_test
import (
"crypto/sha256"
"encoding/hex"
"os"
"strings"
"testing"
"github.com/stretchr/testify/require"
)
// mf.pb.go is generated from mf.proto and committed. `make generate`
// records the hash of the mf.proto it generated from in mf.proto.sha256.
func TestGeneratedCodeMatchesProto(t *testing.T) {
t.Parallel()
proto, err := os.ReadFile("mf.proto")
require.NoError(t, err)
recorded, err := os.ReadFile("mf.proto.sha256")
require.NoError(t, err)
recordedHash, _, _ := strings.Cut(string(recorded), " ")
sum := sha256.Sum256(proto)
require.Equal(t, recordedHash, hex.EncodeToString(sum[:]),
"mfer/mf.proto has changed since mfer/mf.pb.go was generated "+
"from it: run `make generate` and commit the result")
}
+332 -156
View File
@@ -4,6 +4,7 @@ import (
"context"
"io"
"io/fs"
"os"
"path"
"path/filepath"
"strings"
@@ -43,11 +44,22 @@ type ScanStatus struct {
// ScannerOptions configures scanner behavior.
type ScannerOptions struct {
IncludeDotfiles bool // Include files and directories starting with a dot (default: exclude)
FollowSymLinks bool // Resolve symlinks instead of skipping them
Fs afero.Fs // Filesystem to use, defaults to OsFs if nil
SigningOptions *SigningOptions // GPG signing options (nil = no signing)
Seed string // If set, derive a deterministic UUID from this seed
// IncludeDotfiles includes files and directories starting with a dot
// (default: exclude).
IncludeDotfiles bool
// FollowSymLinks resolves symlinks instead of skipping them.
FollowSymLinks bool
// IncludeTimestamps includes a createdAt timestamp in the manifest
// (default: omit for determinism).
IncludeTimestamps bool
// Fs is the filesystem to use, defaults to OsFs if nil.
Fs afero.Fs
// SigningOptions holds GPG signing options (nil = no signing).
SigningOptions *SigningOptions
// Seed, if set, derives a deterministic UUID from this seed.
Seed string
// ExcludePaths lists files to leave out of the listing, matched with os.SameFile.
ExcludePaths []string
}
// FileEntry represents a file that has been enumerated.
@@ -66,6 +78,7 @@ type Scanner struct {
totalBytes FileSize // cached sum of all file sizes
options *ScannerOptions
fs afero.Fs
excluded []fs.FileInfo // the files named in ExcludePaths that exist
}
// NewScanner creates a new Scanner with default options.
@@ -78,15 +91,28 @@ func NewScannerWithOptions(opts *ScannerOptions) *Scanner {
if opts == nil {
opts = &ScannerOptions{}
}
fs := opts.Fs
if fs == nil {
fs = afero.NewOsFs()
}
return &Scanner{
s := &Scanner{
files: make([]*FileEntry, 0),
options: opts,
fs: fs,
}
// A path that cannot be stat'd, normally because no file is there,
// needs no leaving out.
for _, p := range opts.ExcludePaths {
info, err := s.fs.Stat(p)
if err == nil {
s.excluded = append(s.excluded, info)
}
}
return s
}
// EnumerateFile adds a single file to the scanner, calling stat() to get metadata.
@@ -95,47 +121,63 @@ func (s *Scanner) EnumerateFile(filePath string) error {
if err != nil {
return err
}
info, err := s.fs.Stat(abs)
if err != nil {
return err
}
// For single files, use the filename as the relative path
basePath := filepath.Dir(abs)
return s.enumerateFileWithInfo(filepath.Base(abs), basePath, info, nil)
}
// EnumeratePath walks a directory path and adds all files to the scanner.
// If progress is non-nil, status updates are sent as files are discovered.
// The progress channel is closed when the method returns.
func (s *Scanner) EnumeratePath(inputPath string, progress chan<- EnumerateStatus) error {
func (s *Scanner) EnumeratePath(
inputPath string,
progress chan<- EnumerateStatus,
) error {
if progress != nil {
defer close(progress)
}
abs, err := filepath.Abs(inputPath)
if err != nil {
return err
}
afs := afero.NewReadOnlyFs(afero.NewBasePathFs(s.fs, abs))
return s.enumerateFS(afs, abs, progress)
}
// EnumeratePaths walks multiple directory paths and adds all files to the scanner.
// If progress is non-nil, status updates are sent as files are discovered.
// The progress channel is closed when the method returns.
func (s *Scanner) EnumeratePaths(progress chan<- EnumerateStatus, inputPaths ...string) error {
func (s *Scanner) EnumeratePaths(
progress chan<- EnumerateStatus,
inputPaths ...string,
) error {
if progress != nil {
defer close(progress)
}
for _, p := range inputPaths {
abs, err := filepath.Abs(p)
if err != nil {
return err
}
afs := afero.NewReadOnlyFs(afero.NewBasePathFs(s.fs, abs))
if err := s.enumerateFS(afs, abs, progress); err != nil {
err = s.enumerateFS(afs, abs, progress)
if err != nil {
return err
}
}
return nil
}
@@ -143,31 +185,230 @@ func (s *Scanner) EnumeratePaths(progress chan<- EnumerateStatus, inputPaths ...
// If progress is non-nil, status updates are sent as files are discovered.
// The progress channel is closed when the method returns.
// basePath is used to compute absolute paths for file reading.
func (s *Scanner) EnumerateFS(afs afero.Fs, basePath string, progress chan<- EnumerateStatus) error {
func (s *Scanner) EnumerateFS(
afs afero.Fs,
basePath string,
progress chan<- EnumerateStatus,
) error {
if progress != nil {
defer close(progress)
}
return s.enumerateFS(afs, basePath, progress)
}
// enumerateFS is the internal implementation that doesn't close the progress channel.
func (s *Scanner) enumerateFS(afs afero.Fs, basePath string, progress chan<- EnumerateStatus) error {
// Files returns a copy of all files added to the scanner.
func (s *Scanner) Files() []*FileEntry {
s.mu.RLock()
defer s.mu.RUnlock()
out := make([]*FileEntry, len(s.files))
copy(out, s.files)
return out
}
// FileCount returns the number of files in the scanner.
func (s *Scanner) FileCount() FileCount {
s.mu.RLock()
defer s.mu.RUnlock()
return FileCount(len(s.files))
}
// TotalBytes returns the total size of all files in the scanner.
func (s *Scanner) TotalBytes() FileSize {
s.mu.RLock()
defer s.mu.RUnlock()
return s.totalBytes
}
// ToManifest reads all file contents, computes hashes, and generates a manifest.
// If progress is non-nil, status updates are sent approximately once per second.
// The progress channel is closed when the method returns.
// The manifest is written to the provided io.Writer.
func (s *Scanner) ToManifest(
ctx context.Context, w io.Writer, progress chan<- ScanStatus,
) error {
if progress != nil {
defer close(progress)
}
s.mu.RLock()
files := make([]*FileEntry, len(s.files))
copy(files, s.files)
totalFiles := FileCount(len(files))
var totalBytes FileSize
for _, f := range files {
totalBytes += f.Size
}
s.mu.RUnlock()
builder := s.configureBuilder()
var (
scannedFiles FileCount
scannedBytes FileSize
)
lastProgressTime := time.Now()
startTime := time.Now()
pt := &scanProgressTracker{
progress: progress,
totalFiles: totalFiles,
totalBytes: totalBytes,
startTime: startTime,
lastProgress: &lastProgressTime,
}
for _, entry := range files {
// Check for cancellation
select {
case <-ctx.Done():
return ctx.Err()
default:
}
bytesRead, err := s.scanFile(builder, pt, entry, scannedFiles, scannedBytes)
if err != nil {
return err
}
scannedFiles++
scannedBytes += bytesRead
}
// Send final progress (ETA is 0 at completion; remaining bytes are 0,
// so computeRateETA yields eta 0 and the same average rate as before)
if progress != nil {
rate, _ := computeRateETA(time.Since(startTime), scannedBytes, totalBytes)
sendScanStatus(progress, ScanStatus{
TotalFiles: totalFiles,
ScannedFiles: scannedFiles,
TotalBytes: totalBytes,
ScannedBytes: scannedBytes,
BytesPerSec: rate,
ETA: 0,
})
}
// Build and write manifest
return builder.Build(ctx, w)
}
// configureBuilder constructs a manifest builder configured from the
// scanner options.
func (s *Scanner) configureBuilder() *Builder {
builder := NewBuilder()
if s.options.IncludeTimestamps {
builder.SetIncludeTimestamps(true)
}
if s.options.SigningOptions != nil {
builder.SetSigningOptions(s.options.SigningOptions)
}
if s.options.Seed != "" {
builder.SetSeed(s.options.Seed)
}
return builder
}
// scanFile hashes a single file into the builder, forwarding per-file
// progress updates, and returns the number of bytes read.
func (s *Scanner) scanFile(
builder *Builder,
pt *scanProgressTracker,
entry *FileEntry,
scannedFiles FileCount,
scannedBytes FileSize,
) (FileSize, error) {
// Open file
f, err := s.fs.Open(string(entry.AbsPath))
if err != nil {
return 0, err
}
// Create progress channel for this file
var (
fileProgress chan FileHashProgress
wg sync.WaitGroup
)
if pt.progress != nil {
fileProgress = make(chan FileHashProgress, 1)
wg.Add(1)
go func(base FileSize, done FileCount) {
defer wg.Done()
pt.forward(fileProgress, done, base)
}(scannedBytes, scannedFiles)
}
// Add to manifest with progress channel
bytesRead, err := builder.AddFile(
entry.Path,
entry.Size,
entry.Mtime,
f,
fileProgress,
)
_ = f.Close()
// Close channel and wait for goroutine to finish
if fileProgress != nil {
close(fileProgress)
wg.Wait()
}
if err != nil {
return 0, err
}
log.Verbosef("+ %s (%s)", entry.Path, humanize.IBytes(sizeToUint64(bytesRead)))
return bytesRead, nil
}
// enumerateFS is the internal implementation that doesn't close the
// progress channel.
func (s *Scanner) enumerateFS(
afs afero.Fs,
basePath string,
progress chan<- EnumerateStatus,
) error {
return afero.Walk(afs, "/", func(p string, info fs.FileInfo, err error) error {
if err != nil {
return err
}
if !s.options.IncludeDotfiles && IsHiddenPath(p) {
if info.IsDir() {
return filepath.SkipDir
}
return nil
}
return s.enumerateFileWithInfo(p, basePath, info, progress)
})
}
// enumerateFileWithInfo adds a file with pre-existing fs.FileInfo.
func (s *Scanner) enumerateFileWithInfo(filePath string, basePath string, info fs.FileInfo, progress chan<- EnumerateStatus) error {
func (s *Scanner) enumerateFileWithInfo(
filePath string,
basePath string,
info fs.FileInfo,
progress chan<- EnumerateStatus,
) error {
if info.IsDir() {
// Manifests contain only files, directories are implied
return nil
@@ -192,11 +433,13 @@ func (s *Scanner) enumerateFileWithInfo(filePath string, basePath string, info f
realPath, err := filepath.EvalSymlinks(absPath)
if err != nil {
// Skip broken symlinks
return nil
return nil //nolint:nilerr // broken symlinks are skipped by design
}
realInfo, err := s.fs.Stat(realPath)
if err != nil {
return nil
// Skip symlinks whose target cannot be stat'd
return nil //nolint:nilerr // unreadable targets are skipped by design
}
// Skip if symlink points to a directory
if realInfo.IsDir() {
@@ -207,6 +450,12 @@ func (s *Scanner) enumerateFileWithInfo(filePath string, basePath string, info f
info = realInfo
}
for _, excluded := range s.excluded {
if os.SameFile(info, excluded) {
return nil
}
}
entry := &FileEntry{
Path: RelFilePath(cleanPath),
AbsPath: AbsFilePath(absPath),
@@ -231,157 +480,78 @@ func (s *Scanner) enumerateFileWithInfo(filePath string, basePath string, info f
return nil
}
// Files returns a copy of all files added to the scanner.
func (s *Scanner) Files() []*FileEntry {
s.mu.RLock()
defer s.mu.RUnlock()
out := make([]*FileEntry, len(s.files))
copy(out, s.files)
return out
// scanProgressTracker carries the shared state needed to report rate-limited
// scan progress updates.
type scanProgressTracker struct {
progress chan<- ScanStatus
totalFiles FileCount
totalBytes FileSize
startTime time.Time
lastProgress *time.Time
}
// FileCount returns the number of files in the scanner.
func (s *Scanner) FileCount() FileCount {
s.mu.RLock()
defer s.mu.RUnlock()
return FileCount(len(s.files))
}
// TotalBytes returns the total size of all files in the scanner.
func (s *Scanner) TotalBytes() FileSize {
s.mu.RLock()
defer s.mu.RUnlock()
return s.totalBytes
}
// ToManifest reads all file contents, computes hashes, and generates a manifest.
// If progress is non-nil, status updates are sent approximately once per second.
// The progress channel is closed when the method returns.
// The manifest is written to the provided io.Writer.
func (s *Scanner) ToManifest(ctx context.Context, w io.Writer, progress chan<- ScanStatus) error {
if progress != nil {
defer close(progress)
}
s.mu.RLock()
files := make([]*FileEntry, len(s.files))
copy(files, s.files)
totalFiles := FileCount(len(files))
var totalBytes FileSize
for _, f := range files {
totalBytes += f.Size
}
s.mu.RUnlock()
builder := NewBuilder()
if s.options.SigningOptions != nil {
builder.SetSigningOptions(s.options.SigningOptions)
}
if s.options.Seed != "" {
builder.SetSeed(s.options.Seed)
}
var scannedFiles FileCount
var scannedBytes FileSize
lastProgressTime := time.Now()
startTime := time.Now()
for _, entry := range files {
// Check for cancellation
select {
case <-ctx.Done():
return ctx.Err()
default:
// forward relays per-file hash progress to the scan progress channel,
// rate-limited to one update per second.
func (pt *scanProgressTracker) forward(
fileProgress <-chan FileHashProgress,
scannedFiles FileCount,
baseBytes FileSize,
) {
for p := range fileProgress {
// Send progress at most once per second
now := time.Now()
if now.Sub(*pt.lastProgress) < time.Second {
continue
}
// Open file
f, err := s.fs.Open(string(entry.AbsPath))
if err != nil {
return err
}
currentBytes := baseBytes + p.BytesRead
rate, eta := computeRateETA(now.Sub(pt.startTime), currentBytes, pt.totalBytes)
// Create progress channel for this file
var fileProgress chan FileHashProgress
var wg sync.WaitGroup
if progress != nil {
fileProgress = make(chan FileHashProgress, 1)
wg.Add(1)
go func(baseScannedBytes FileSize) {
defer wg.Done()
for p := range fileProgress {
// Send progress at most once per second
now := time.Now()
if now.Sub(lastProgressTime) >= time.Second {
elapsed := now.Sub(startTime).Seconds()
currentBytes := baseScannedBytes + p.BytesRead
var rate float64
var eta time.Duration
if elapsed > 0 && currentBytes > 0 {
rate = float64(currentBytes) / elapsed
remainingBytes := totalBytes - currentBytes
if rate > 0 {
eta = time.Duration(float64(remainingBytes)/rate) * time.Second
}
}
sendScanStatus(progress, ScanStatus{
TotalFiles: totalFiles,
ScannedFiles: scannedFiles,
TotalBytes: totalBytes,
ScannedBytes: currentBytes,
BytesPerSec: rate,
ETA: eta,
})
lastProgressTime = now
}
}
}(scannedBytes)
}
// Add to manifest with progress channel
bytesRead, err := builder.AddFile(
entry.Path,
entry.Size,
entry.Mtime,
f,
fileProgress,
)
_ = f.Close()
// Close channel and wait for goroutine to finish
if fileProgress != nil {
close(fileProgress)
wg.Wait()
}
if err != nil {
return err
}
log.Verbosef("+ %s (%s)", entry.Path, humanize.IBytes(uint64(bytesRead)))
scannedFiles++
scannedBytes += bytesRead
}
// Send final progress (ETA is 0 at completion)
if progress != nil {
elapsed := time.Since(startTime).Seconds()
var rate float64
if elapsed > 0 {
rate = float64(scannedBytes) / elapsed
}
sendScanStatus(progress, ScanStatus{
TotalFiles: totalFiles,
sendScanStatus(pt.progress, ScanStatus{
TotalFiles: pt.totalFiles,
ScannedFiles: scannedFiles,
TotalBytes: totalBytes,
ScannedBytes: scannedBytes,
TotalBytes: pt.totalBytes,
ScannedBytes: currentBytes,
BytesPerSec: rate,
ETA: 0,
ETA: eta,
})
*pt.lastProgress = now
}
}
// computeRateETA returns the average throughput over elapsed time and the
// estimated time to process the remaining bytes at that rate.
func computeRateETA(
elapsed time.Duration,
done FileSize,
total FileSize,
) (float64, time.Duration) {
var (
rate float64
eta time.Duration
)
if elapsed > 0 && done > 0 {
rate = float64(done) / elapsed.Seconds()
remaining := total - done
if rate > 0 {
eta = time.Duration(float64(remaining)/rate) * time.Second
}
}
// Build and write manifest
return builder.Build(w)
return rate, eta
}
// sizeToUint64 converts a FileSize to uint64 for display, clamping
// negative values to zero so the conversion cannot overflow.
func sizeToUint64(v FileSize) uint64 {
if v < 0 {
return 0
}
return uint64(v)
}
// IsHiddenPath returns true if the path or any of its parent directories
@@ -392,17 +562,21 @@ func IsHiddenPath(p string) bool {
if tp == "." || tp == "/" {
return false
}
if strings.HasPrefix(tp, ".") {
return true
}
for {
d, f := path.Split(tp)
if strings.HasPrefix(f, ".") {
return true
}
if d == "" {
return false
}
tp = d[0 : len(d)-1] // trim trailing slash from dir
}
}
@@ -413,6 +587,7 @@ func sendEnumerateStatus(ch chan<- EnumerateStatus, status EnumerateStatus) {
if ch == nil {
return
}
select {
case ch <- status:
default:
@@ -426,6 +601,7 @@ func sendScanStatus(ch chan<- ScanStatus, status ScanStatus) {
if ch == nil {
return
}
select {
case ch <- status:
default:
+88 -46
View File
@@ -1,3 +1,4 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
@@ -12,6 +13,8 @@ import (
)
func TestNewScanner(t *testing.T) {
t.Parallel()
s := NewScanner()
assert.NotNil(t, s)
assert.Equal(t, FileCount(0), s.FileCount())
@@ -19,12 +22,18 @@ func TestNewScanner(t *testing.T) {
}
func TestNewScannerWithOptions(t *testing.T) {
t.Parallel()
t.Run("nil options", func(t *testing.T) {
t.Parallel()
s := NewScannerWithOptions(nil)
assert.NotNil(t, s)
})
t.Run("with options", func(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
opts := &ScannerOptions{
IncludeDotfiles: true,
@@ -37,6 +46,8 @@ func TestNewScannerWithOptions(t *testing.T) {
}
func TestScannerEnumerateFile(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("hello world"), 0o644))
@@ -54,6 +65,8 @@ func TestScannerEnumerateFile(t *testing.T) {
}
func TestScannerEnumerateFileMissing(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumerateFile("/nonexistent.txt")
@@ -61,11 +74,14 @@ func TestScannerEnumerateFileMissing(t *testing.T) {
}
func TestScannerEnumeratePath(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir/subdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("one"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("two"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/subdir/file3.txt", []byte("three"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/subdir/file3.txt", []byte("three"), 0o644))
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumeratePath("/testdir", nil)
@@ -76,6 +92,8 @@ func TestScannerEnumeratePath(t *testing.T) {
}
func TestScannerEnumeratePathWithProgress(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("one"), 0o644))
@@ -100,6 +118,8 @@ func TestScannerEnumeratePathWithProgress(t *testing.T) {
}
func TestScannerEnumeratePaths(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/dir1", 0o755))
require.NoError(t, fs.MkdirAll("/dir2", 0o755))
@@ -114,13 +134,20 @@ func TestScannerEnumeratePaths(t *testing.T) {
}
func TestScannerExcludeDotfiles(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir/.hidden", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/visible.txt", []byte("visible"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/.hidden.txt", []byte("hidden"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/.hidden/inside.txt", []byte("inside"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/visible.txt", []byte("visible"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/.hidden.txt", []byte("hidden"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/.hidden/inside.txt", []byte("inside"), 0o644))
t.Run("exclude by default", func(t *testing.T) {
t.Parallel()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs, IncludeDotfiles: false})
err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err)
@@ -131,6 +158,8 @@ func TestScannerExcludeDotfiles(t *testing.T) {
})
t.Run("include when enabled", func(t *testing.T) {
t.Parallel()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs, IncludeDotfiles: true})
err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err)
@@ -140,34 +169,43 @@ func TestScannerExcludeDotfiles(t *testing.T) {
}
func TestScannerToManifest(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file1.txt", []byte("content one"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/file2.txt", []byte("content two"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file1.txt", []byte("content one"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file2.txt", []byte("content two"), 0o644))
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err)
var buf bytes.Buffer
err = s.ToManifest(context.Background(), &buf, nil)
require.NoError(t, err)
// Manifest should have magic bytes
assert.True(t, buf.Len() > 0)
assert.Positive(t, buf.Len())
assert.Equal(t, MAGIC, string(buf.Bytes()[:8]))
}
func TestScannerToManifestWithProgress(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file.txt", bytes.Repeat([]byte("x"), 1000), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/file.txt", bytes.Repeat([]byte("x"), 1000), 0o644))
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
err := s.EnumeratePath("/testdir", nil)
require.NoError(t, err)
var buf bytes.Buffer
progress := make(chan ScanStatus, 10)
err = s.ToManifest(context.Background(), &buf, progress)
@@ -188,12 +226,15 @@ func TestScannerToManifestWithProgress(t *testing.T) {
}
func TestScannerToManifestContextCancellation(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
// Create many files to ensure we have time to cancel
for i := 0; i < 100; i++ {
for i := range 100 {
name := string(rune('a'+i%26)) + string(rune('0'+i/26)) + ".txt"
require.NoError(t, afero.WriteFile(fs, "/testdir/"+name, bytes.Repeat([]byte("x"), 100), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/"+name, bytes.Repeat([]byte("x"), 100), 0o644))
}
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
@@ -204,24 +245,30 @@ func TestScannerToManifestContextCancellation(t *testing.T) {
cancel() // Cancel immediately
var buf bytes.Buffer
err = s.ToManifest(ctx, &buf, nil)
assert.ErrorIs(t, err, context.Canceled)
}
func TestScannerToManifestEmptyScanner(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
s := NewScannerWithOptions(&ScannerOptions{Fs: fs})
var buf bytes.Buffer
err := s.ToManifest(context.Background(), &buf, nil)
require.NoError(t, err)
// Should still produce a valid manifest
assert.True(t, buf.Len() > 0)
assert.Positive(t, buf.Len())
assert.Equal(t, MAGIC, string(buf.Bytes()[:8]))
}
func TestScannerFilesCopiesSlice(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("hello"), 0o644))
@@ -236,10 +283,13 @@ func TestScannerFilesCopiesSlice(t *testing.T) {
}
func TestScannerEnumerateFS(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir/sub", 0o755))
require.NoError(t, afero.WriteFile(fs, "/testdir/file.txt", []byte("hello"), 0o644))
require.NoError(t, afero.WriteFile(fs, "/testdir/sub/nested.txt", []byte("world"), 0o644))
require.NoError(t,
afero.WriteFile(fs, "/testdir/sub/nested.txt", []byte("world"), 0o644))
// Create a basepath filesystem
baseFs := afero.NewBasePathFs(fs, "/testdir")
@@ -252,51 +302,37 @@ func TestScannerEnumerateFS(t *testing.T) {
}
func TestSendEnumerateStatusNonBlocking(t *testing.T) {
// Channel with no buffer - send should not block
t.Parallel()
// Nobody receives, so a blocking send would hang the test into its timeout.
ch := make(chan EnumerateStatus)
// This should not block
done := make(chan bool)
go func() {
sendEnumerateStatus(ch, EnumerateStatus{FilesFound: 1})
done <- true
}()
select {
case <-done:
// Success - did not block
case <-time.After(100 * time.Millisecond):
t.Fatal("sendEnumerateStatus blocked on full channel")
}
sendEnumerateStatus(ch, EnumerateStatus{FilesFound: 1})
}
func TestSendScanStatusNonBlocking(t *testing.T) {
// Channel with no buffer - send should not block
t.Parallel()
// Nobody receives, so a blocking send would hang the test into its timeout.
ch := make(chan ScanStatus)
done := make(chan bool)
go func() {
sendScanStatus(ch, ScanStatus{ScannedFiles: 1})
done <- true
}()
select {
case <-done:
// Success - did not block
case <-time.After(100 * time.Millisecond):
t.Fatal("sendScanStatus blocked on full channel")
}
sendScanStatus(ch, ScanStatus{ScannedFiles: 1})
}
func TestSendStatusNilChannel(t *testing.T) {
t.Parallel()
// Should not panic with nil channel
sendEnumerateStatus(nil, EnumerateStatus{})
sendScanStatus(nil, ScanStatus{})
}
func TestScannerFileEntryFields(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
now := time.Now().Truncate(time.Second)
require.NoError(t, afero.WriteFile(fs, "/test.txt", []byte("content"), 0o644))
require.NoError(t, fs.Chtimes("/test.txt", now, now))
@@ -315,11 +351,13 @@ func TestScannerFileEntryFields(t *testing.T) {
}
func TestScannerLargeFileEnumeration(t *testing.T) {
t.Parallel()
fs := afero.NewMemMapFs()
require.NoError(t, fs.MkdirAll("/testdir", 0o755))
// Create 100 files
for i := 0; i < 100; i++ {
for i := range 100 {
name := "/testdir/" + string(rune('a'+i%26)) + string(rune('0'+i/26%10)) + ".txt"
require.NoError(t, afero.WriteFile(fs, name, []byte("data"), 0o644))
}
@@ -330,20 +368,20 @@ func TestScannerLargeFileEnumeration(t *testing.T) {
err := s.EnumeratePath("/testdir", progress)
require.NoError(t, err)
// Drain channel
for range progress {
}
// progress is fully buffered and closed; no draining needed
assert.Equal(t, FileCount(100), s.FileCount())
assert.Equal(t, FileSize(400), s.TotalBytes()) // 100 * 4 bytes
}
func TestIsHiddenPath(t *testing.T) {
t.Parallel()
tests := []struct {
path string
hidden bool
}{
{"file.txt", false},
{testFileName, false},
{".hidden", true},
{"dir/file.txt", false},
{"dir/.hidden", true},
@@ -352,12 +390,16 @@ func TestIsHiddenPath(t *testing.T) {
{"/absolute/.hidden", true},
{"./relative", false}, // path.Clean removes leading ./
{"a/b/c/.d/e", true},
{".", false}, // current directory is not hidden
{"/", false}, // root is not hidden
{".", false}, // current directory is not hidden (#14)
{"/", false}, // root is not hidden
{"./", false}, // current directory with trailing slash
{"./file.txt", false}, // file in current directory
}
for _, tt := range tests {
t.Run(tt.path, func(t *testing.T) {
t.Parallel()
assert.Equal(t, tt.hidden, IsHiddenPath(tt.path), "IsHiddenPath(%q)", tt.path)
})
}
+89 -37
View File
@@ -2,9 +2,11 @@ package mfer
import (
"bytes"
"context"
"crypto/sha256"
"errors"
"fmt"
"math"
"time"
"github.com/google/uuid"
@@ -15,73 +17,112 @@ import (
// MAGIC is the file format magic bytes prefix (rot13 of "MANIFEST").
const MAGIC string = "ZNAVSRFG"
var (
// errInnerNotSet is returned by generate when the inner manifest is
// missing.
errInnerNotSet = errors.New("internal error: pbInner not set")
// errInternal is returned by generateOuter for the same condition.
// The two messages differ, and both are load-bearing for callers that
// match on text, so they are kept distinct.
errInternal = errors.New("internal error")
)
// nanosecondsInt32 converts t's nanosecond component to int32.
// time.Time.Nanosecond is documented to return a value in [0, 999999999],
// so the conversion cannot overflow. This sits directly in the manifest
// content path: silently substituting a default would zero every entry's
// mtime nanos and change the serialized bytes and their hash, so an
// out-of-contract value is a programming error and panics rather than
// being papered over.
func nanosecondsInt32(t time.Time) int32 {
n := t.Nanosecond()
if n < 0 || n > math.MaxInt32 {
panic(fmt.Sprintf(
"mfer: time.Time.Nanosecond out of contract: %d", n))
}
return int32(n)
}
func newTimestampFromTime(t time.Time) *Timestamp {
return &Timestamp{
Seconds: t.Unix(),
Nanos: int32(t.Nanosecond()),
Nanos: nanosecondsInt32(t),
}
}
func (m *manifest) generate() error {
func (m *manifest) generate(ctx context.Context) error {
if m.pbInner == nil {
return errors.New("internal error: pbInner not set")
return errInnerNotSet
}
if m.pbOuter == nil {
e := m.generateOuter()
e := m.generateOuter(ctx)
if e != nil {
return e
}
}
dat, err := proto.MarshalOptions{Deterministic: true}.Marshal(m.pbOuter)
if err != nil {
return err
return fmt.Errorf("serialize: marshal outer: %w", err)
}
m.output = bytes.NewBuffer([]byte(MAGIC))
m.output = bytes.NewBufferString(MAGIC)
_, err = m.output.Write(dat)
if err != nil {
return err
return fmt.Errorf("serialize: write output: %w", err)
}
return nil
}
func (m *manifest) generateOuter() error {
func (m *manifest) generateOuter(ctx context.Context) error {
if m.pbInner == nil {
return errors.New("internal error")
return errInternal
}
// Use fixed UUID if provided, otherwise generate a new one
var manifestUUID uuid.UUID
if len(m.fixedUUID) == 16 {
if len(m.fixedUUID) == uuidLength {
copy(manifestUUID[:], m.fixedUUID)
} else {
manifestUUID = uuid.New()
}
m.pbInner.Uuid = manifestUUID[:]
innerData, err := proto.MarshalOptions{Deterministic: true}.Marshal(m.pbInner)
if err != nil {
return err
return fmt.Errorf("serialize: marshal inner: %w", err)
}
// Compress the inner data
idc := new(bytes.Buffer)
zw, err := zstd.NewWriter(idc, zstd.WithEncoderLevel(zstd.SpeedBestCompression))
if err != nil {
return err
return fmt.Errorf("serialize: create compressor: %w", err)
}
_, err = zw.Write(innerData)
if err != nil {
return err
return fmt.Errorf("serialize: compress: %w", err)
}
_ = zw.Close()
compressedData := idc.Bytes()
// Hash the compressed data for integrity verification before decompression
h := sha256.New()
if _, err := h.Write(compressedData); err != nil {
return err
_, err = h.Write(compressedData)
if err != nil {
return fmt.Errorf("serialize: hash write: %w", err)
}
sha256Hash := h.Sum(nil)
m.pbOuter = &MFFileOuter{
@@ -95,29 +136,40 @@ func (m *manifest) generateOuter() error {
// Sign the manifest if signing options are provided
if m.signingOptions != nil && m.signingOptions.KeyID != "" {
sigString, err := m.signatureString()
if err != nil {
return fmt.Errorf("failed to generate signature string: %w", err)
}
sig, err := gpgSign([]byte(sigString), m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to sign manifest: %w", err)
}
m.pbOuter.Signature = sig
fingerprint, err := gpgGetKeyFingerprint(m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to get key fingerprint: %w", err)
}
m.pbOuter.Signer = fingerprint
pubKey, err := gpgExportPublicKey(m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to export public key: %w", err)
}
m.pbOuter.SigningPubKey = pubKey
return m.signOuter(ctx)
}
return nil
}
// signOuter signs the outer message with the configured GPG key and
// embeds the signature, signer fingerprint, and public key.
func (m *manifest) signOuter(ctx context.Context) error {
sigString, err := m.signatureString()
if err != nil {
return fmt.Errorf("failed to generate signature string: %w", err)
}
sig, err := gpgSign(ctx, []byte(sigString), m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to sign manifest: %w", err)
}
m.pbOuter.Signature = sig
fingerprint, err := gpgGetKeyFingerprint(ctx, m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to get key fingerprint: %w", err)
}
m.pbOuter.Signer = fingerprint
pubKey, err := gpgExportPublicKey(ctx, m.signingOptions.KeyID)
if err != nil {
return fmt.Errorf("failed to export public key: %w", err)
}
m.pbOuter.SigningPubKey = pubKey
return nil
}
+2
View File
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06\xff\xff\xff\x03\xc2\x06 {\x16\xbdu\xa0\xa2\x11\xfcH\xef*\x1b7\r\x99\xefb\x04\x02g\n\xa9\xf3B5\xe5p\x96\x8c\x8c\xac\x0e\xca\x06\x10\x03Q\xb2\xd0\x19`F\xc1\xb1\xc0Z\xf4x\xf4g^\xba\f\xa1\x06(\xb5/\xfd\x04h\x04\x01\x00d\x01\xb2\x06\x10\x03Q\xb2\xd0\x19`F\xc1\xb1\xc0Z\xf4x\xf4g^\xaa\x06\x00\x01T\x13\x024\xce\xff\rL\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15M\x00\x00\x00\x01T\x00\x044\xfc\xff\x153\xea\a\xb4")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06\xff\xff\xff\x03\xc2\x06 .\xcd\x11|0\xfcP\xe5\x1b\xe3\xc6Ӡ\xcdڤx\xcd\x169t\x1a9~ǽB\xc9\xe8G`\x05\xca\x06\x10\x11*!\x0e\x95EF\xb8\xbd\x9f\xde\x12MF\r\x99\xba\f\xa6\x06(\xb5/\xfd\x04h,\x01\x00\xb4\x01\xb2\x06\x10\x11*!\x0e\x95EF\xb8\xbd\x9f\xde\x12MF\r\x99\xaa\x06\xe6\xff\xff\x03\x1a\x00\x01T\x14\x024\x8b\xff\x17L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15L\x00\x00\x00\x01T\x00\x044\xfd\xff\x15M\x00\x00\x00\x01T\x00\x044\xfc\xff\x15\x02\xd1.\xe3")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFGy[i\x04\xe5O\x82A\x1d\xf4\xb0\xe2z7:U\xee\xa3\xf9\xd6m\xacZ\x9b\xce\x1d\xd9/{@\x1d\xa5y[i\x04\xe5O\x82A\x1d\xf4\xb0\xe2z7:U\xee\xa3\xf9\xd6m\xacZ\x9b\xce\x1d\xd9/{@\x1d\xa5")
+2
View File
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 \xa3\xf7\x97\xa7\xf3\x87:\x90)\\ӊj\xb9\xf7\xfaTJ\x1b\xe7:\xee\xbe\"V\xe0:\x8d(Z\xd6%\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\fc(\xb5/\xfd\x04\x00\xb1\x02\x00\xa0\x06\x01\xaa\x06=\n\x05a.txt\x10\x01\x1a$\n\"\x12 ʗ\x81\x12\xca\x1b\xbd\xca\xfa\xc21\xb3\x9a#\xdcM\xa7\x86\xef\xf8\x14|Nr\xb9\x80w\x85\xaf\xeeH\xbb\xf2\x12\v\b\x80\x92\xb8Ø\xfe\xff\xff\xff\x01\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3s\xeeG\x80")
+2
View File
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 \xa3\xf7\x97\xa7\xf3\x87:\x90)\\ӊj\xb9\xf7\xfaTJ\x1b\xe7:\xee\xbe\"V\xe0:\x8d(Z\xd6%\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\fc(\xb5/\xfd\x04\x00\xb1\x02\x00\xa0\x06\x01\xaa\x06=\n\x05a.txt\x10\x01\x1a$\n\"\x12 ʗ\x81\x12\xca\x1b\xbd\xca\xfa\xc21\xb3\x9a#\xdcM\xa7\x86\xef\xf8\x14|Nr\xb9\x80w\x85\xaf\xeeH\xbb\xf2\x12\v\b\x80\x92\xb8Ø\xfe\xff\xff\xff\x01\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3s\xeeG\x80\xca\f\xe8\x03-----BEGIN PGP SIGNATURE-----\n\niQEzBAABCgAdFiEET1Yr+4Y/3GtRtO6IhypRF2zvI64FAmrBIXAACgkQhypRF2zv\nI67BQAf/QrpX2MjY15YGMGkjR5oIhnx/YV96aGYZyZThzb+l/R/N75iVFVkhX21d\nZhQqdCsORrodTPAXic2g2UGVXP9PhNMh7n6Wm3LsvQYjrRQGrQnqtCkut+3tUt8K\n7pt4OAnnwRSieaVImA1COmzxIrQQKNOs6UkgmAstGuPV0XZoeDiSG8TUYJ/vieCn\np5hC0FFXtzfw4NtkxSmkewE0xBxIwFCA/RfSHCGH3m5K+tRz41vMEgGbL1iEp6+V\nuBaoEc4hqCgEt+Af2pA8VHfqeu2vKiwggOpYpaILXZKVqH9+tWHL1EBv9t0vTsYE\n9D57euuR9+kOdngYNPieP1yn5dOSHg==\n=kvgL\n-----END PGP SIGNATURE-----\n\xd2\f(4F562BFB863FDC6B51B4EE88872A51176CEF23AE\xda\f\xb5\a-----BEGIN PGP PUBLIC KEY BLOCK-----\n\nmQENBGrBIW8BCADESetN5EdxIe7Fafgxl99Yoo5cOexf7wJyYT0wfUYlRaxt3neR\nhir7LOfH4PZWWoDx7qghxCS4+vs7yGypl6JOm7jnJlhn4HneDa2zeIlgGW2TamyE\nua9KPWBQqkFOYmKPmzp+KnL6ncnBLR5mDkNKFyON812KVvteu6Dp/DNk4Meufe44\nWWr49LSFZa9gEbmRCoQGKby9F0H0yIi4FAc74VdQudy0+fMKcfkKjEvByMzlbBEK\n92Hq3sRFzWd3kvPliNjZTmlh5n5m9aBhMpoy3GkKy8gpDdFc6NLA9iAJe7oNMriR\nkVoa5EjQL1xCXAiAWTYA9NScFfU/574sCTxZABEBAAG0Hk1GRVIgVGVzdCBLZXkg\nPHRlc3RAbWZlci50ZXN0PokBTwQTAQoAORYhBE9WK/uGP9xrUbTuiIcqURds7yOu\nBQJqwSFvAxsvBAULCQgHAgYVCgkICwIEFgIDAQIeAQIXgAAKCRCHKlEXbO8jrjml\nCAC8wUK9wmvxq0+NZUpFyP+P29klLZYzBDaBrLPJFs0GjnG4kvfUAktWx0Ro80F7\ncjTJ4f44XjDj4glvSjbe2VaDnZl9FTfzUfG+xjD4462NgntQ4fHk/uG4F6d1ikWx\nkEoMpIn1PlSMas1jTQSGlxUr+zFwWuUbGq4n6hRxEnwLlwJlwQt/Aw1vPDYuPDE3\nOYDhJIAJyP+6e9W8ToaAG9byg/22KA1u1qxnNQqsx5Tped2VltAzdYub+yeCuNc8\nIUo5ILo/fQq3GM5sUEaHjPolv88WlDm3vcdbSbDoh5m2inDtg5zuUKJwu32UGu8l\nKoTjp4nxgQGy5WmeBzs4Hdm/\n=TCB8\n-----END PGP PUBLIC KEY BLOCK-----\n")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06!\xc2\x06 \x91\x90*\xa5>\fݐ \x87\xbeaL\xc1\x05?\x0eR\xc18\xa4eՕ\xa95\xb9KʺoZ\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\f-(\xb5/\xfd\x04\x00\x01\x01\x00\xa0\x06\x01\xaa\x06\a\n\x05a.txt\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3a[k'")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06\x1f\xc2\x06 \x91\x90*\xa5>\fݐ \x87\xbeaL\xc1\x05?\x0eR\xc18\xa4eՕ\xa95\xb9KʺoZ\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\f-(\xb5/\xfd\x04\x00\x01\x01\x00\xa0\x06\x01\xaa\x06\a\n\x05a.txt\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3a[k'")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 \xa3\xf7\x97\xa7\xf3\x87:\x90)\\ӊj\xb9\xf7\xfaTJ\x1b\xe7:\xee\xbe\"V\xe0:\x8d(Z\xd6%\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\fc(\xb5/\xfd\x04\x00\xb1\x02\x00\xa0\x06\x01\xaa\x06=\n\x05a.txt\x10\x01\x1a$\n\"\x12 ʗ\x81\x12\xca\x1b\xbd\xca\xfa\xc21\xb3\x9a#\xdcM\xa7\x86\xef\xf8\x14|Nr\xb9\x80w\x85\xaf\xeeH\xbb\xf2\x12\v\b\x80\x92\xb8Ø\xfe\xff\xff\xff\x01\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3s\xeeG")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 ")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAV")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 \xa3\xf7\x97\xa7\xf3\x87:\x90)\\ӊj\xb9\xf7\xfaTJ\x1b\xe7:\xee\xbe\"V\xe0:\x8d(Z\xd6%\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\fc(\xb5/\xfd\x04\x00\xb1\x02\x00\xa0\x06\x01")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFX\xa8\x06\x01\xb0\x06\x01\xb8\x06V\xc2\x06 \xa3\xf7\x97\xa7\xf3\x87:\x90)\\ӊj\xb9\xf7\xfaTJ\x1b\xe7:\xee\xbe\"V\xe0:\x8d(Z\xd6%\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\fc(\xb5/\xfd\x04\x00\xb1\x02\x00\xa0\x06\x01\xaa\x06=\n\x05a.txt\x10\x01\x1a$\n\"\x12 ʗ\x81\x12\xca\x1b\xbd\xca\xfa\xc21\xb3\x9a#\xdcM\xa7\x86\xef\xf8\x14|Nr\xb9\x80w\x85\xaf\xeeH\xbb\xf2\x12\v\b\x80\x92\xb8Ø\xfe\xff\xff\xff\x01\xb2\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3s\xeeG\x80")
File diff suppressed because one or more lines are too long
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06\x01\xc2\x06 \xd6aQ\xa4\x85\xc0+\xbb\xca\x11\x11\a<\x019\x97\xb3\xbb3\xd0 \xd5U\xfa!\xaeAf<N@\x9d\xca\x06\x10\x93\x85\vpu\x85\xe4\x04\xe4\x95\x1a=\xdc\x1f\x05\xa3\xba\f\x12(\xb5/\xfd\xc0\x00\x00\x00\x00\x00\x02\x00\x00\x00\v\x00\x00\x00")
@@ -0,0 +1,2 @@
go test fuzz v1
[]byte("ZNAVSRFG\xa8\x06\x01\xb0\x06\x01\xb8\x06\x01\xc2\x06 \xc8g?\xa0\xc9\xc5]\xb57M\xe5O#\x04\xb2\x12\xc6)@=\xd2\xf3f/͉~\x10\x15\xfbkD\xca\x06\x10o\x1c*N;]L~\x9a\x8b\x1d.?@Qb\xba\f\x89\t(\xb5/\xfd\x00\x00\x01\x00\x00(\xb5/\xfd\x00\x01\x01\x00\x00(\xb5/\xfd\x00\x02\x01\x00\x00(\xb5/\xfd\x00\x03\x01\x00\x00(\xb5/\xfd\x00\x04\x01\x00\x00(\xb5/\xfd\x00\x05\x01\x00\x00(\xb5/\xfd\x00\x06\x01\x00\x00(\xb5/\xfd\x00\a\x01\x00\x00(\xb5/\xfd\x00\b\x01\x00\x00(\xb5/\xfd\x00\t\x01\x00\x00(\xb5/\xfd\x00\n\x01\x00\x00(\xb5/\xfd\x00\v\x01\x00\x00(\xb5/\xfd\x00\f\x01\x00\x00(\xb5/\xfd\x00\r\x01\x00\x00(\xb5/\xfd\x00\x0e\x01\x00\x00(\xb5/\xfd\x00\x0f\x01\x00\x00(\xb5/\xfd\x00\x10\x01\x00\x00(\xb5/\xfd\x00\x11\x01\x00\x00(\xb5/\xfd\x00\x12\x01\x00\x00(\xb5/\xfd\x00\x13\x01\x00\x00(\xb5/\xfd\x00\x14\x01\x00\x00(\xb5/\xfd\x00\x15\x01\x00\x00(\xb5/\xfd\x00\x16\x01\x00\x00(\xb5/\xfd\x00\x17\x01\x00\x00(\xb5/\xfd\x00\x18\x01\x00\x00(\xb5/\xfd\x00\x19\x01\x00\x00(\xb5/\xfd\x00\x1a\x01\x00\x00(\xb5/\xfd\x00\x1b\x01\x00\x00(\xb5/\xfd\x00\x1c\x01\x00\x00(\xb5/\xfd\x00\x1d\x01\x00\x00(\xb5/\xfd\x00\x1e\x01\x00\x00(\xb5/\xfd\x00\x1f\x01\x00\x00(\xb5/\xfd\x00 \x01\x00\x00(\xb5/\xfd\x00!\x01\x00\x00(\xb5/\xfd\x00\"\x01\x00\x00(\xb5/\xfd\x00#\x01\x00\x00(\xb5/\xfd\x00$\x01\x00\x00(\xb5/\xfd\x00%\x01\x00\x00(\xb5/\xfd\x00&\x01\x00\x00(\xb5/\xfd\x00'\x01\x00\x00(\xb5/\xfd\x00(\x01\x00\x00(\xb5/\xfd\x00)\x01\x00\x00(\xb5/\xfd\x00*\x01\x00\x00(\xb5/\xfd\x00+\x01\x00\x00(\xb5/\xfd\x00,\x01\x00\x00(\xb5/\xfd\x00-\x01\x00\x00(\xb5/\xfd\x00.\x01\x00\x00(\xb5/\xfd\x00/\x01\x00\x00(\xb5/\xfd\x000\x01\x00\x00(\xb5/\xfd\x001\x01\x00\x00(\xb5/\xfd\x002\x01\x00\x00(\xb5/\xfd\x003\x01\x00\x00(\xb5/\xfd\x004\x01\x00\x00(\xb5/\xfd\x005\x01\x00\x00(\xb5/\xfd\x006\x01\x00\x00(\xb5/\xfd\x007\x01\x00\x00(\xb5/\xfd\x008\x01\x00\x00(\xb5/\xfd\x009\x01\x00\x00(\xb5/\xfd\x00:\x01\x00\x00(\xb5/\xfd\x00;\x01\x00\x00(\xb5/\xfd\x00<\x01\x00\x00(\xb5/\xfd\x00=\x01\x00\x00(\xb5/\xfd\x00>\x01\x00\x00(\xb5/\xfd\x00?\x01\x00\x00(\xb5/\xfd\x00@\x01\x00\x00(\xb5/\xfd\x00A\x01\x00\x00(\xb5/\xfd\x00B\x01\x00\x00(\xb5/\xfd\x00C\x01\x00\x00(\xb5/\xfd\x00D\x01\x00\x00(\xb5/\xfd\x00E\x01\x00\x00(\xb5/\xfd\x00F\x01\x00\x00(\xb5/\xfd\x00G\x01\x00\x00(\xb5/\xfd\x00H\x01\x00\x00(\xb5/\xfd\x00I\x01\x00\x00(\xb5/\xfd\x00J\x01\x00\x00(\xb5/\xfd\x00K\x01\x00\x00(\xb5/\xfd\x00L\x01\x00\x00(\xb5/\xfd\x00M\x01\x00\x00(\xb5/\xfd\x00N\x01\x00\x00(\xb5/\xfd\x00O\x01\x00\x00(\xb5/\xfd\x00P\x01\x00\x00(\xb5/\xfd\x00Q\x01\x00\x00(\xb5/\xfd\x00R\x01\x00\x00(\xb5/\xfd\x00S\x01\x00\x00(\xb5/\xfd\x00T\x01\x00\x00(\xb5/\xfd\x00U\x01\x00\x00(\xb5/\xfd\x00V\x01\x00\x00(\xb5/\xfd\x00W\x01\x00\x00(\xb5/\xfd\x00X\x01\x00\x00(\xb5/\xfd\x00Y\x01\x00\x00(\xb5/\xfd\x00Z\x01\x00\x00(\xb5/\xfd\x00[\x01\x00\x00(\xb5/\xfd\x00\\\x01\x00\x00(\xb5/\xfd\x00]\x01\x00\x00(\xb5/\xfd\x00^\x01\x00\x00(\xb5/\xfd\x00_\x01\x00\x00(\xb5/\xfd\x00`\x01\x00\x00(\xb5/\xfd\x00a\x01\x00\x00(\xb5/\xfd\x00b\x01\x00\x00(\xb5/\xfd\x00c\x01\x00\x00(\xb5/\xfd\x00d\x01\x00\x00(\xb5/\xfd\x00e\x01\x00\x00(\xb5/\xfd\x00f\x01\x00\x00(\xb5/\xfd\x00g\x01\x00\x00(\xb5/\xfd\x00h\x01\x00\x00(\xb5/\xfd\x00i\x01\x00\x00(\xb5/\xfd\x00j\x01\x00\x00(\xb5/\xfd\x00k\x01\x00\x00(\xb5/\xfd\x00l\x01\x00\x00(\xb5/\xfd\x00m\x01\x00\x00(\xb5/\xfd\x00n\x01\x00\x00(\xb5/\xfd\x00o\x01\x00\x00(\xb5/\xfd\x00p\x01\x00\x00(\xb5/\xfd\x00q\x01\x00\x00(\xb5/\xfd\x00r\x01\x00\x00(\xb5/\xfd\x00s\x01\x00\x00(\xb5/\xfd\x00t\x01\x00\x00(\xb5/\xfd\x00u\x01\x00\x00(\xb5/\xfd\x00v\x01\x00\x00(\xb5/\xfd\x00w\x01\x00\x00(\xb5/\xfd\x00x\x01\x00\x00(\xb5/\xfd\x00y\x01\x00\x00(\xb5/\xfd\x00z\x01\x00\x00(\xb5/\xfd\x00{\x01\x00\x00(\xb5/\xfd\x00|\x01\x00\x00(\xb5/\xfd\x00}\x01\x00\x00(\xb5/\xfd\x00~\x01\x00\x00(\xb5/\xfd\x00\x7f\x01\x00\x00(\xb5/\xfd\x00\x80\x01\x00\x00")
File diff suppressed because one or more lines are too long
+8 -2
View File
@@ -27,13 +27,19 @@ func (b BaseURL) JoinPath(path RelFilePath) (FileURL, error) {
base.Path += "/"
}
// Parse and encode the relative path
ref, err := url.Parse(url.PathEscape(string(path)))
// Encode each path segment individually to preserve slashes
segments := strings.Split(string(path), "/")
for i, seg := range segments {
segments[i] = url.PathEscape(seg)
}
ref, err := url.Parse(strings.Join(segments, "/"))
if err != nil {
return "", err
}
resolved := base.ResolveReference(ref)
return FileURL(resolved.String()), nil
}
+59
View File
@@ -0,0 +1,59 @@
//nolint:testpackage // white-box tests exercise unexported internals
package mfer
import (
"testing"
"github.com/stretchr/testify/assert"
"github.com/stretchr/testify/require"
)
func TestBaseURLJoinPath(t *testing.T) {
t.Parallel()
tests := []struct {
base BaseURL
path RelFilePath
expected string
}{
{"https://example.com/dir/", testFileName, "https://example.com/dir/file.txt"},
{"https://example.com/dir", testFileName, "https://example.com/dir/file.txt"},
{"https://example.com/", "sub/file.txt", "https://example.com/sub/file.txt"},
{
"https://example.com/dir/",
"file with spaces.txt",
"https://example.com/dir/file%20with%20spaces.txt",
},
}
for _, tt := range tests {
t.Run(string(tt.base)+"+"+string(tt.path), func(t *testing.T) {
t.Parallel()
result, err := tt.base.JoinPath(tt.path)
require.NoError(t, err)
assert.Equal(t, tt.expected, string(result))
})
}
}
func TestBaseURLString(t *testing.T) {
t.Parallel()
b := BaseURL("https://example.com/")
assert.Equal(t, "https://example.com/", b.String())
}
func TestFileURLString(t *testing.T) {
t.Parallel()
f := FileURL("https://example.com/file.txt")
assert.Equal(t, "https://example.com/file.txt", f.String())
}
func TestManifestURLString(t *testing.T) {
t.Parallel()
m := ManifestURL("https://example.com/index.mf")
assert.Equal(t, "https://example.com/index.mf", m.String())
}
+9
View File
@@ -0,0 +1,9 @@
{
"name": "mfer",
"private": true,
"description": "Development tooling for the mfer repository: prettier, used by script/fmt and script/fmt-check to format and verify Markdown and JSON.",
"license": "WTFPL",
"devDependencies": {
"prettier": "3.9.6"
}
}
+208
View File
@@ -0,0 +1,208 @@
#!/bin/sh
# script/bootstrap: install all dependencies needed to build and develop
# this repo. Idempotent: every install is guarded by a check so already
# installed tools are skipped. Base tooling comes from nix, apt, brew,
# or apk (detected in that order); assumes NOTHING is present (not git,
# make, node, yarn, go, or python). Node is used directly if installed;
# otherwise a pinned version is installed via nvm (installing nvm
# itself first, from a hash-verified release archive, never curl | sh).
#
# Uncomment the language sections in main() that apply to this repo.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Pinned versions, 2026-07-06. Never "latest" or "lts"; exact versions.
# The node version is in .nvmrc, where script/prettier reads it too.
NODE_VERSION="$(cat "$ROOT/.nvmrc")"
NVM_VERSION="0.40.3"
# sha256 of https://github.com/nvm-sh/nvm/archive/refs/tags/v0.40.3.tar.gz
NVM_SHA256="5f4d6aaa04a177dc93c985e31dbc411ab6b8c6e1e21d8015dbc1372625fcd1d0"
YARN_VERSION="1.22.22"
# protoc v33.4, 2026-10-04, for script/generate. The sha256 of each
# platform's release archive is in ensure_protoc.
PROTOC_VERSION="33.4"
PKGMGR=""
SUDO=""
detect_pkgmgr() {
[ -n "$PKGMGR" ] && return 0
if command -v nix-env >/dev/null 2>&1; then
PKGMGR="nix"
elif command -v apt-get >/dev/null 2>&1; then
PKGMGR="apt"
elif command -v brew >/dev/null 2>&1; then
PKGMGR="brew"
elif command -v apk >/dev/null 2>&1; then
PKGMGR="apk"
else
echo "bootstrap: no supported package manager (nix, apt, brew, apk)" >&2
exit 1
fi
if [ "$PKGMGR" = "apt" ]; then
export DEBIAN_FRONTEND=noninteractive
if [ "$(id -u)" != "0" ]; then
SUDO="sudo"
fi
fi
}
# pkg_install <nix-attr> <apt-pkg> <brew-formula> <apk-pkg>
pkg_install() {
detect_pkgmgr
case "$PKGMGR" in
nix) nix-env -iA "nixpkgs.$1" ;;
apt) $SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2" ;;
brew) brew install "$3" ;;
apk) apk add --no-cache "$4" ;;
esac
}
missing() {
! command -v "$1" >/dev/null 2>&1
}
# verify_sha256 <file> <expected-hash>
verify_sha256() {
if command -v sha256sum >/dev/null 2>&1; then
actual="$(sha256sum "$1" | cut -d' ' -f1)"
else
actual="$(shasum -a 256 "$1" | cut -d' ' -f1)"
fi
if [ "$actual" != "$2" ]; then
echo "bootstrap: sha256 mismatch for $1" >&2
echo " expected: $2" >&2
echo " actual: $actual" >&2
exit 1
fi
}
# nvm is a bash script; run a command in a bash with nvm loaded.
# --no-use: otherwise loading nvm here switches to the version .nvmrc
# names, and fails silently while that version is not installed yet.
nvm_sh() {
bash -c ". \"\$HOME/.nvm/nvm.sh\" --no-use && $*"
}
ensure_nvm() {
[ -s "$HOME/.nvm/nvm.sh" ] && return 0
# nvm prerequisites; nvm itself requires bash, so install it too
if missing bash; then pkg_install bash bash bash bash; fi
if missing curl; then pkg_install curl curl curl curl; fi
if missing git; then pkg_install git git git git; fi
tmp="$(mktemp -d)"
curl -fsSL -o "$tmp/nvm.tar.gz" \
"https://github.com/nvm-sh/nvm/archive/refs/tags/v${NVM_VERSION}.tar.gz"
verify_sha256 "$tmp/nvm.tar.gz" "$NVM_SHA256"
mkdir -p "$HOME/.nvm"
tar -xzf "$tmp/nvm.tar.gz" -C "$HOME/.nvm" --strip-components=1
rm -rf "$tmp"
}
ensure_node() {
if ! missing node; then return 0; fi
ensure_nvm
nvm_sh "nvm install $NODE_VERSION"
}
ensure_yarn() {
if ! missing yarn; then return 0; fi
if ! missing corepack; then
corepack enable
corepack prepare "yarn@$YARN_VERSION" --activate
elif [ -s "$HOME/.nvm/nvm.sh" ]; then
nvm_sh "nvm use $NODE_VERSION >/dev/null && corepack enable && \
corepack prepare yarn@$YARN_VERSION --activate"
else
npm install -g "yarn@$YARN_VERSION"
fi
}
install_js_deps() {
if missing yarn && [ -s "$HOME/.nvm/nvm.sh" ]; then
nvm_sh "nvm use $NODE_VERSION >/dev/null && cd \"$ROOT\" && \
yarn install --frozen-lockfile"
else
yarn install --frozen-lockfile
fi
}
# Unpack protoc's release archive for this platform into bin/protoc, after
# checking the archive's sha256, unless bin/protoc already holds the pinned
# version.
ensure_protoc() {
dir="$ROOT/bin/protoc"
if [ "$("$dir/bin/protoc" --version 2>/dev/null)" = \
"libprotoc $PROTOC_VERSION" ]; then
return 0
fi
case "$(uname -s) $(uname -m)" in
"Linux x86_64")
platform="linux-x86_64"
sha256="c0040ea9aef08fdeb2c74ca609b18d5fdbfc44ea0042fcfbfb38860d35f7dd66"
;;
"Linux aarch64" | "Linux arm64")
platform="linux-aarch_64"
sha256="15aa988f4a6090636525ec236a8e4b3aab41eef402751bd5bb2df6afd9b7b5a5"
;;
"Darwin x86_64")
platform="osx-x86_64"
sha256="a49bec10d039e902d3b43e49938c42526f90011467609864fa6386ac4014da58"
;;
"Darwin arm64")
platform="osx-aarch_64"
sha256="726297dcfed58592fd35620a5a6246ae020c39e88f3fd4cb1827df7bcf3dfcf1"
;;
*)
echo "bootstrap: no protoc archive pinned for $(uname -s) $(uname -m)" >&2
exit 1
;;
esac
if missing curl; then pkg_install curl curl curl curl; fi
if missing unzip; then pkg_install unzip unzip unzip unzip; fi
tmp="$(mktemp -d)"
curl -fsSL -o "$tmp/protoc.zip" \
"https://github.com/protocolbuffers/protobuf/releases/download/v${PROTOC_VERSION}/protoc-${PROTOC_VERSION}-${platform}.zip"
verify_sha256 "$tmp/protoc.zip" "$sha256"
rm -rf "$dir"
unzip -q "$tmp/protoc.zip" -d "$dir"
rm -rf "$tmp"
}
main() {
cd "$ROOT"
# Base tooling (every repo)
if missing git; then pkg_install git git git git; fi
if missing make; then pkg_install gnumake make make make; fi
# ---- JS / docs repos ----
# This is a Go repo, but node and yarn are required anyway: prettier
# formats the Markdown and JSON, and script/fmt-check verifies it.
# The version is pinned by package.json/yarn.lock: yarn checks every
# package it fetches against its yarn.lock integrity hash, and
# --frozen-lockfile fails instead of rewriting a yarn.lock that no
# longer matches package.json.
ensure_node
ensure_yarn
install_js_deps
# ---- Go repos ----
if missing go; then pkg_install go golang go go; fi
# No golangci-lint: script/lint runs it in Docker only.
go mod download
# gofumpt and protoc-gen-go: bin/tools/go.mod pins them, and
# script/gofumpt and script/generate build them from there.
(cd "$ROOT/bin/tools" && go mod download)
ensure_protoc
# ---- Python repos ----
# if missing python3; then pkg_install python3 python3 python3 python3; fi
# python3 -m venv .venv
# ./.venv/bin/pip install -e '.[dev]'
echo "bootstrap complete"
}
main "$@"
Executable
+18
View File
@@ -0,0 +1,18 @@
#!/bin/sh
# script/build: build the mfer binary into bin/mfer, stamped with the
# revision `mfer version` prints, derived as script/docker derives it.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
# Own line: a failing command substitution inside an argument does
# not trip `set -e`, so the inline form degrades silently to an
# empty constant.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
go build -ldflags "-X main.Gitrev=$version" -o bin/mfer ./cmd/mfer
}
main "$@"
Executable
+15
View File
@@ -0,0 +1,15 @@
#!/bin/sh
# script/check: run all checks (test, lint, fmt-check). Our own
# extension to scripts-to-rule-them-all. Must not modify any files.
# Generic: usually needs no adaptation.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
main() {
"$SCRIPT_DIR/test"
"$SCRIPT_DIR/lint"
"$SCRIPT_DIR/fmt-check"
}
main "$@"
Executable
+24
View File
@@ -0,0 +1,24 @@
#!/bin/sh
# script/cibuild: run the CI build; the Gitea workflow runs this on push.
# It builds the image with the same command as script/docker. --no-cache
# because the checks the final stage depends on are RUN steps, and a
# cached one is a check that did not run.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
# Own line: a failing command substitution inside an argument does
# not trip `set -e`, so the inline form degrades silently to an
# empty constant. The VERSION build argument takes precedence over
# the version a build stage derives from the .git in the context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$("$SCRIPT_DIR/projectname")" .
}
main "$@"
Executable
+24
View File
@@ -0,0 +1,24 @@
#!/bin/sh
# script/docker: build the Docker image tagged with the project name.
# Identical in all repos; the tag comes from script/projectname.
# --no-cache because the gate phases the final stage depends on are RUN
# steps, and a cached one is a check that did not run.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
# Own line: a failing command substitution inside an argument does
# not trip `set -e`, so the inline form degrades silently to an
# empty constant. The VERSION build argument takes precedence over
# the version a build stage derives from the .git in the context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$("$SCRIPT_DIR/projectname")" .
}
main "$@"
Executable
+16
View File
@@ -0,0 +1,16 @@
#!/bin/sh
# script/fmt: format all files (writes).
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
# Go, then Markdown and JSON, through the same scripts script/fmt-check
# uses, so both see the same files and the same tool versions.
"$SCRIPT_DIR/gofumpt" --write
"$SCRIPT_DIR/prettier" --write
}
main "$@"
+14
View File
@@ -0,0 +1,14 @@
#!/bin/sh
# script/fmt-check: check formatting (read-only). Same scope as
# script/fmt, but fails instead of writing: Go via script/gofumpt,
# Markdown and JSON via script/prettier.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
main() {
"$SCRIPT_DIR/gofumpt" --check
"$SCRIPT_DIR/prettier" --check
}
main "$@"
Executable
+17
View File
@@ -0,0 +1,17 @@
#!/bin/sh
# script/fuzz: fuzz the manifest parser for one minute. Run by hand only:
# script/test already runs the committed seed corpus as ordinary tests,
# and CI never fuzzes. An input that fails is written to
# mfer/testdata/fuzz/FuzzNewManifestFromReader/; once the parser is fixed,
# commit it there as a regression seed.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
go test -run '^$' -fuzz '^FuzzNewManifestFromReader$' \
-fuzztime 1m -parallel 2 ./mfer
}
main "$@"
+54
View File
@@ -0,0 +1,54 @@
#!/bin/sh
# script/generate: regenerate mfer/mf.pb.go from mfer/mf.proto, and record
# the hash of that mf.proto in mfer/mf.proto.sha256. Nothing else
# regenerates mf.pb.go: it is committed, so building and checking need no
# protoc. A test fails while mf.proto no longer matches the recorded hash.
#
# Runs the protoc that script/bootstrap unpacks into bin/protoc, and the
# protoc-gen-go that bin/tools/go.mod pins. Another version of either
# writes a different mf.pb.go.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# The protoc version script/bootstrap installs. protoc 33.4 names itself
# v6.33.4 in the mf.pb.go header.
PROTOC_VERSION="33.4"
PROTOC="$ROOT/bin/protoc/bin/protoc"
# sha256 <file>: print "<hash> <file>", with sha256sum, or with shasum
# where there is no sha256sum.
sha256() {
if command -v sha256sum >/dev/null 2>&1; then
sha256sum "$1"
elif command -v shasum >/dev/null 2>&1; then
shasum -a 256 "$1"
else
echo "generate: needs sha256sum or shasum on PATH" >&2
exit 1
fi
}
main() {
# A bin/protoc left from before the pin moved fails here, until
# script/bootstrap replaces it.
actual="$("$PROTOC" --version 2>/dev/null || true)"
if [ "$actual" != "libprotoc $PROTOC_VERSION" ]; then
echo "generate: needs protoc $PROTOC_VERSION in bin/protoc," \
"found: ${actual:-none}; run script/bootstrap" >&2
exit 1
fi
# `go tool -n` builds protoc-gen-go from bin/tools and prints where the
# binary is, without running it.
plugin="$(cd "$ROOT/bin/tools" && go tool -n protoc-gen-go)"
cd "$ROOT/mfer"
# Hashed before regenerating, so a missing hash tool stops the script
# before it changes anything. Regenerating leaves mf.proto as it is.
proto_hash="$(sha256 mf.proto)"
"$PROTOC" --plugin=protoc-gen-go="$plugin" \
--go_out=paths=source_relative:. ./mf.proto
echo "$proto_hash" >mf.proto.sha256
}
main "$@"
Executable
+44
View File
@@ -0,0 +1,44 @@
#!/bin/sh
# script/gofumpt: run gofumpt over this repo's Go files.
#
# Takes exactly one mode argument, --write or --check, and runs the same
# gofumpt version over the same files in both modes. script/fmt and
# script/fmt-check both go through here, and so does the Docker lint
# stage, so what gets formatted and what gets verified cannot drift
# apart.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# The gofumpt version is the one bin/tools/go.mod pins. `go tool` run in
# bin/tools builds it from source checked against the hashes in
# bin/tools/go.sum, so neither a developer machine nor the lint image needs
# it installed.
usage() {
echo "usage: script/gofumpt --write|--check" >&2
exit 2
}
main() {
[ "$#" -eq 1 ] || usage
cd "$ROOT/bin/tools"
# Every Go file in the repo, from $ROOT down. gofumpt holds generated
# files, such as mfer/mf.pb.go, to gofmt's rules only.
case "$1" in
--write) go tool gofumpt -l -w "$ROOT" ;;
--check)
# Own line: a failing command inside `[ -n "$(...)" ]` does
# not trip `set -e`, so a gofumpt that never ran would pass.
unformatted="$(go tool gofumpt -l "$ROOT")"
if [ -n "$unformatted" ]; then
echo "gofumpt: files need formatting (run make fmt):" >&2
echo "$unformatted" >&2
exit 1
fi
;;
*) usage ;;
esac
}
main "$@"
+17
View File
@@ -0,0 +1,17 @@
#!/bin/sh
# script/install-precommit: install the git pre-commit hook that runs
# script/precommit. Our own extension to scripts-to-rule-them-all.
# Generic: needs no adaptation.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
hook=".git/hooks/pre-commit"
printf '#!/bin/sh\nset -e\nscript/precommit\n' > .git/hooks/pre-commit
chmod +x .git/hooks/pre-commit
echo "pre-commit hook installed: runs script/precommit"
}
main "$@"
Executable
+20
View File
@@ -0,0 +1,20 @@
#!/bin/sh
# script/lint: run golangci-lint, in Docker only. Builds the lint stage of
# the Dockerfile, whose build runs the linter, so a successful build is a
# clean lint. --no-cache because a cached build runs no linter. The image
# is removed afterwards, whatever the outcome.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
# Tagged per run, so concurrent runs never remove each other's image.
image="$("$SCRIPT_DIR/projectname")-lint:$$"
# A failed build leaves no image, so there is nothing to remove then.
trap 'docker image rm "$image" >/dev/null 2>&1 || true' EXIT INT TERM
docker build --no-cache --target lint -t "$image" .
}
main "$@"
+20
View File
@@ -0,0 +1,20 @@
#!/bin/sh
# script/precommit: run by the git pre-commit hook; fails the commit if
# checks fail. Our own extension to scripts-to-rule-them-all. Go repo
# extras run first: go mod tidy and go fmt, failing the commit if they
# change go.mod or go.sum.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() {
cd "$ROOT"
go mod tidy
go fmt ./...
git diff --exit-code -- go.mod go.sum ||
{ echo "go mod tidy changed files; stage and retry" >&2; exit 1; }
"$SCRIPT_DIR/check"
}
main "$@"
+69
View File
@@ -0,0 +1,69 @@
#!/bin/sh
# script/prettier: run prettier over this repo's canonical file set.
#
# Takes exactly one mode argument, --write or --check, and applies the
# same patterns in both modes. script/fmt and script/fmt-check both go
# through here, so the set of files that get formatted and the set that
# get verified cannot drift apart.
#
# Failures are never swallowed: a missing prettier is an error, not a
# silent skip. A formatter that quietly does nothing is worse than one
# that fails loudly.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
usage() {
echo "usage: script/prettier --write|--check" >&2
exit 2
}
# Only the prettier yarn installed from yarn.lock, never one on PATH: a
# different version formats differently.
PRETTIER="$ROOT/node_modules/.bin/prettier"
main() {
[ "$#" -eq 1 ] || usage
case "$1" in
--write | --check) mode="$1" ;;
*) usage ;;
esac
cd "$ROOT"
# Where there is no node on PATH, script/bootstrap installs the version
# .nvmrc names through nvm, which keeps it in this directory.
if ! command -v node >/dev/null 2>&1; then
PATH="$HOME/.nvm/versions/node/v$(cat .nvmrc)/bin:$PATH"
fi
if ! command -v node >/dev/null 2>&1; then
echo "prettier: node is missing; run script/bootstrap" >&2
exit 1
fi
# node_modules keeps the old prettier after package.json moves to a new
# one, until script/bootstrap runs again, so compare the two.
if ! installed="$("$PRETTIER" --version 2>/dev/null)"; then
echo "prettier: not installed; run script/bootstrap" >&2
exit 1
fi
pinned="$(node -p 'require("./package.json").devDependencies.prettier')"
if [ "$installed" != "$pinned" ]; then
echo "prettier: package.json pins $pinned but $installed is" \
"installed; run script/bootstrap" >&2
exit 1
fi
# Markdown and JSON, repo-wide rather than root-only, so files in
# subdirectories (docs/, once it exists) are covered too. Exclusions
# live in .prettierignore; REPO_POLICIES.md is excluded there because
# it is a verbatim copy of an upstream document.
#
# --no-error-on-unmatched-pattern is deliberately NOT used: both
# patterns always match at least one tracked file (README.md,
# package.json), so an empty match means the glob broke, and prettier
# erroring out is exactly what we want rather than a vacuous pass.
"$PRETTIER" "$mode" "**/*.md" "**/*.json"
}
main "$@"

Some files were not shown because too many files have changed in this diff Show More