Author SHA1 Message Date
sneak 09802fde10 Reject manifests whose file entries decode far larger than their bytes (closes #123)
check / check (push) Successful in 1m23s
Parser fix: before decoding the manifest, the parser now walks its file
entries and their hashes and rejects any entry shorter than 9 bytes or
hash shorter than 4, the smallest the format allows. Decoding allocates
a fixed amount per entry and per hash, so empty ones decoded to about 50
times their size: a 1.6 KB manifest allocated nearly 1 GB. What the
parser accepts now decodes to at most about 25 times its size, so the
fuzz target's ceiling rises from 16 to 36 times the input and
decompressed data. New seeds of empty entries and of empty hashes fail
the fuzz target without the fix.

Model: opus-5-5
2026-10-04 04:06:52 +00:00
15 changed files with 212 additions and 150 deletions
-1
View File
@@ -16,7 +16,6 @@ linters:
- depguard # Dependency allow/block lists
- godot # Requires comments to end with periods
- wsl # Deprecated, replaced by wsl_v5
- gomodguard # Deprecated, replaced by gomodguard_v2
- wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go
settings:
+2 -7
View File
@@ -14,9 +14,7 @@ RUN touch mfer/mf.pb.go
# Go half of fmt-check only: this image has no node, so no prettier. The
# markdown half runs in the mdfmt stage below.
RUN make fmt-check-go
# The linter directly, not `make lint`: script/lint builds this stage, and
# there is no docker inside this build.
RUN golangci-lint run --config .golangci.yml ./...
RUN make lint
# Markdown/JSON format stage — prettier needs node, which the Go images
# do not have. node:22.17.0-bookworm-slim (2026-08-09); ships node
@@ -69,10 +67,7 @@ RUN version="${VERSION:-$(git describe --tags --always)}"; \
exit 1; \
fi; \
cd cmd/mfer && \
CGO_ENABLED=0 go build -tags urfave_cli_no_docs -ldflags "-X main.Gitrev=$version" -o /mfer .
# Fail unless /mfer is statically linked: scratch has no C library to run it.
RUN ldd /mfer 2>&1 | grep -q 'not a dynamic executable'
go build -tags urfave_cli_no_docs -ldflags "-X main.Gitrev=$version" -o /mfer .
FROM scratch
COPY --from=builder /mfer /mfer
+3
View File
@@ -58,6 +58,9 @@ fmt-check-md:
hooks:
@script/install-precommit
devprereqs:
which golangci-lint || go install -v github.com/golangci/golangci-lint/v2/cmd/golangci-lint@v2.12.2
mfer/mf.pb.go: mfer/mf.proto
cd mfer && go generate .
+22 -77
View File
@@ -1,12 +1,12 @@
# mfer
[mfer](https://git.eeqj.de/sneak/mfer) is a [WTFPL](https://wtfpl.net)-licensed
(public domain) [Go](https://golang.org) library and command-line tool by
[@sneak](https://sneak.berlin) that specifies and generates `.mf` manifest files
over a directory tree to encapsulate metadata about the files — such as
cryptographic checksums and signatures over same — to aid in archiving,
downloading, streaming, and mirroring. It was first published in 2022. The
manifest files' data is serialized with Google's
[mfer](https://git.eeqj.de/sneak/mfer) is a reference implementation library and
thin wrapper command-line utility written in [Go](https://golang.org) and first
published in 2022 under the [WTFPL](https://wtfpl.net) (public domain) license.
It specifies and generates `.mf` manifest files over a directory tree of files
to encapsulate metadata about them (such as cryptographic checksums or
signatures over same) to aid in archiving, downloading, and streaming, or
mirroring. The manifest files' data is serialized with Google's
[protobuf serialization format](https://developers.google.com/protocol-buffers).
The structure of these files can be found
[in the format specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto)
@@ -21,36 +21,6 @@ This project was started by [@sneak](https://sneak.berlin) to scratch an itch in
as a de-facto standard and be incorporated into other software. A compatible
javascript library is planned.
# Getting Started
`mfer` builds from source with a Go 1.23+ toolchain. The generated protobuf code
is committed, so no `protoc` toolchain is required:
```sh
git clone https://git.eeqj.de/sneak/mfer.git
cd mfer
go build -o bin/mfer ./cmd/mfer
```
Generate a manifest for a directory tree, verify it later, and fetch a published
tree by URL:
```sh
# Write .index.mf, a manifest of the files under the current directory.
bin/mfer gen .
# Verify the files on disk against the manifest. Exits nonzero if any file
# is missing or corrupted.
bin/mfer check .index.mf
# Download and cryptographically verify a tree published over HTTP: mfer
# fetches <url>/index.mf, then downloads every file it lists.
bin/mfer fetch https://example.com/tree/
```
Run `bin/mfer help` for the full command list, or `bin/mfer <command> --help`
for a single command's options.
# Build Status
CI runs `script/cibuild`, which builds the Docker image with `--no-cache`, so
@@ -65,9 +35,9 @@ standard: normalized scripts in `script/` are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call them. We
provide:
- `script/bootstrap` — install all dependencies (Go, Go module download, and
node/yarn plus the prettier version pinned in `package.json`/`yarn.lock`),
idempotently; golangci-lint is not installed, it runs only in Docker
- `script/bootstrap` — install all dependencies (Go, golangci-lint, Go module
download, and node/yarn plus the prettier version pinned in
`package.json`/`yarn.lock`), idempotently
- `script/setup` — make a fresh clone ready for development: runs
`script/bootstrap`, then `script/install-precommit`
- `script/projectname` — output the project name (`mfer`); used by other scripts
@@ -77,11 +47,9 @@ provide:
- `script/fuzz` — fuzz the manifest parser for one minute; run by hand
(`make fuzz`), never by CI, while `script/test` runs its committed seed corpus
as ordinary tests
- `script/lint` — run `golangci-lint` in Docker: builds only the `lint` stage of
the `Dockerfile` (the Go format check, then the linter), uncached so it runs
every time, then removes the image
- `script/fmt` — format all code and docs (writes): `gofumpt` and
`script/prettier --write`
- `script/lint` — run `golangci-lint` and verify `gofmt` cleanliness
- `script/fmt` — format all code and docs (writes): `gofumpt`,
`golangci-lint run --fix`, and `script/prettier --write`
- `script/prettier` — run prettier over the repository's canonical file set
(Markdown and JSON, minus `.prettierignore`) in the given mode, `--write` or
`--check`; the single definition of that file set, so `script/fmt` and
@@ -107,18 +75,17 @@ yet. Primary development happens on a privately-run Gitea instance at
[tracked there](https://git.eeqj.de/sneak/mfer/issues).
Changes must always be formatted with a standard `go fmt`, syntactically valid,
and must pass the linting defined in the repository's `.golangci.yml`, which
`make lint` runs in Docker. The `main` branch is protected and all changes must
be made via [pull requests](https://git.eeqj.de/sneak/mfer/pulls) and pass CI to
be merged. Any changes submitted to this project must also be
and must pass the linting defined in the repository (presently only the
`golangci-lint` defaults), which can be run with a `make lint`. The `main`
branch is protected and all changes must be made via
[pull requests](https://git.eeqj.de/sneak/mfer/pulls) and pass CI to be merged.
Any changes submitted to this project must also be
[WTFPL-licensed](https://wtfpl.net) to be considered.
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
tooling requirements, and workflow conventions.
# Rationale
## The problem
# Problem Statement
Given a plain URL, there is no standard way to safely and programmatically
download everything "under" that URL path. `wget -r` can traverse directory
@@ -142,7 +109,7 @@ Real issues I face:
- when I download a large file via HTTP, I have no way of knowing if the file
content is what it's supposed to be
## The solution
# Proposed Solution
A standard, a manifest file format, and a tool for generating same.
@@ -174,28 +141,6 @@ The manifest file would do several important things:
- maybe a bittorrent chunklist for torrent client compatibility? perhaps a
top-level infohash for the whole manifest?
# Design
The repository is split into a reusable library and a thin command-line wrapper
around it.
- `mfer/` is the reusable library and the heart of the project: it defines the
manifest format and implements building, scanning, checking, serialization,
and signing. The protobuf schema is `mfer/mf.proto`, and the generated code it
produces (`mfer/mf.pb.go`) is committed alongside it so the library builds
with `go get` and needs no `protoc` toolchain.
- `internal/cli/` holds the command implementations — `generate` (alias `gen`),
`check`, `freshen`, `export`, `list` (alias `ls`), `fetch`, and `version` —
that wire the library to the command-line interface.
- `internal/log/` provides the logging used by the commands and the library.
- `internal/bork/` provides the error the library returns when a manifest's
decompressed contents are not the size the manifest records.
- `cmd/mfer/` is the entrypoint: its `main` package runs `internal/cli` and
exits with the status it returns.
Everything under `internal/` is private to this repository; only the `mfer/`
package is intended for import by other software.
# Design Goals
- Replace SHASUMS/SHASUMS.asc files
@@ -329,9 +274,9 @@ Open work, open design questions included, is tracked in this repo's issues:
- Issues:
[https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues)
# Author
# Authors
- [@sneak](https://sneak.berlin)
- [@sneak &lt;sneak@sneak.berlin&gt;](mailto:sneak@sneak.berlin)
# License
+13
View File
@@ -17,4 +17,17 @@ const (
// uuidLength is the length in bytes of a binary UUID.
uuidLength = 16
// filesFieldNumber and hashesFieldNumber are the numbers of
// MFFile.files and MFFilePath.hashes in mf.proto.
filesFieldNumber = 101
hashesFieldNumber = 3
// minHashSize is the encoded size of the smallest hash: a two-byte
// multihash (algorithm code, zero digest length) after its tag and length.
minHashSize = 2 + 2
// minFileEntrySize is the encoded size of the smallest file entry: a
// one-byte path and one hash, each after its tag and length.
minFileEntrySize = 2 + 1 + 2 + minHashSize
)
+59
View File
@@ -11,6 +11,7 @@ import (
"github.com/google/uuid"
"github.com/klauspost/compress/zstd"
"github.com/spf13/afero"
"google.golang.org/protobuf/encoding/protowire"
"google.golang.org/protobuf/proto"
"sneak.berlin/go/mfer/internal/bork"
"sneak.berlin/go/mfer/internal/log"
@@ -27,6 +28,8 @@ var (
errUUIDMismatch = errors.New("outer and inner UUID mismatch")
errInvalidFileFormat = errors.New("invalid file format")
errInvalidManifestPath = errors.New("manifest contains invalid path")
errEntryTooShort = errors.New(
"manifest contains a file entry or hash shorter than the format allows")
)
// validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable).
@@ -154,6 +157,57 @@ func (m *manifest) decompressInner() ([]byte, error) {
return dat, nil
}
// checkEntrySizes rejects an encoded inner message holding a file entry or
// a hash shorter than the format allows. Decoding allocates a fixed amount
// for each entry and each hash, however short, so a payload of empty ones
// would decode to about 50 times its size.
func checkEntrySizes(inner []byte) error {
return forEachBytesField(inner, filesFieldNumber, func(entry []byte) error {
if len(entry) < minFileEntrySize {
return errEntryTooShort
}
return forEachBytesField(entry, hashesFieldNumber, func(hash []byte) error {
if len(hash) < minHashSize {
return errEntryTooShort
}
return nil
})
})
}
// forEachBytesField calls fn with the value of each length-delimited field
// numbered num in the encoded message msg, and fails if msg is malformed.
func forEachBytesField(
msg []byte, num protowire.Number, fn func(value []byte) error,
) error {
for len(msg) > 0 {
fieldNum, wireType, tagLen := protowire.ConsumeTag(msg)
if tagLen < 0 {
return protowire.ParseError(tagLen)
}
valueLen := protowire.ConsumeFieldValue(fieldNum, wireType, msg[tagLen:])
if valueLen < 0 {
return protowire.ParseError(valueLen)
}
if fieldNum == num && wireType == protowire.BytesType {
value, _ := protowire.ConsumeBytes(msg[tagLen:])
err := fn(value)
if err != nil {
return err
}
}
msg = msg[tagLen+valueLen:]
}
return nil
}
func (m *manifest) deserializeInner() error {
err := m.validateOuterHeader()
if err != nil {
@@ -177,6 +231,11 @@ func (m *manifest) deserializeInner() error {
return bork.ErrFileTruncated
}
err = checkEntrySizes(dat)
if err != nil {
return fmt.Errorf("deserialize: unmarshal inner: %w", err)
}
// Deserialize inner message
m.pbInner = new(MFFile)
+8 -4
View File
@@ -55,8 +55,10 @@ func FuzzNewManifestFromReader(f *testing.F) {
// It also keeps a few copies of its input. Buffers grow by
// copying, so reaching those sizes allocates a few times them in
// total: sixteen times the input and the decompressed data leaves
// room for that.
// total. Decoding the decompressed data takes up to about 25 times
// its size, when every file entry and hash is as short as the
// parser accepts. Thirty-six times the input and the decompressed
// data leaves room for both.
//
// The decoder also sets aside a new buffer of one to two times the
// window for each frame that asks for a larger window than the
@@ -71,8 +73,10 @@ func FuzzNewManifestFromReader(f *testing.F) {
// fails if the decoder accepts windows of twice zstdWindowSize; the
// seed whose two frames together exceed MaxDecompressedSize fails
// if the decoder decodes them in full instead of stopping at the
// declared size.
limit := 16*(uint64(len(data))+decompressed) + 24*zstdWindowSize
// declared size; the seeds of empty file entries and of a file
// entry of empty hashes fail if the parser decodes entries or
// hashes shorter than the format allows.
limit := 36*(uint64(len(data))+decompressed) + 24*zstdWindowSize
allocated := after.TotalAlloc - before.TotalAlloc
if allocated > limit {
+31 -5
View File
@@ -17,12 +17,18 @@ import (
)
// craftInnerBytes builds the wire bytes of an inner MFFile holding a single
// file entry whose path is exactly pathBytes. It writes the wire form by hand
// so a hostile path — including one that is not valid UTF-8 — can be embedded
// without proto.Marshal's own UTF-8 enforcement rejecting it first.
func craftInnerBytes(id uuid.UUID, pathBytes string) []byte {
// file entry whose path is exactly pathBytes and whose one hash is multihash.
// It writes the wire form by hand so a hostile path — including one that is
// not valid UTF-8 — can be embedded without proto.Marshal's own UTF-8
// enforcement rejecting it first.
func craftInnerBytes(id uuid.UUID, pathBytes string, multihash []byte) []byte {
hash := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFileChecksum.multiHash
hash = protowire.AppendBytes(hash, multihash)
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
entry = protowire.AppendString(entry, pathBytes)
entry = protowire.AppendTag(entry, 3, protowire.BytesType) // MFFilePath.hashes
entry = protowire.AppendBytes(entry, hash)
inner := protowire.AppendTag(nil, 100, protowire.VarintType) // MFFile.version
inner = protowire.AppendVarint(inner, uint64(MFFile_VERSION_ONE))
@@ -89,7 +95,8 @@ func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
t.Parallel()
id := uuid.New()
data := wrapInner(t, id, craftInnerBytes(id, tt.path))
hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256
data := wrapInner(t, id, craftInnerBytes(id, tt.path, hash))
_, err := NewManifestFromReader(bytes.NewReader(data))
require.Error(t, err)
@@ -114,6 +121,25 @@ func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
}
}
// A one-byte path and a multihash of algorithm code and zero digest length
// make the smallest file entry and hash the format allows: they load, and an
// entry or hash one byte shorter is refused before decoding.
func TestDeserializeRejectsEntriesShorterThanFormatAllows(t *testing.T) {
t.Parallel()
load := func(path string, multihash []byte) error {
id := uuid.New()
data := wrapInner(t, id, craftInnerBytes(id, path, multihash))
_, err := NewManifestFromReader(bytes.NewReader(data))
return err
}
require.NoError(t, load("a", []byte{0, 0}))
require.ErrorIs(t, load("", []byte{0, 0}), errEntryTooShort)
require.ErrorIs(t, load("ab", []byte{0}), errEntryTooShort)
}
func TestDeserializeValidManifestRoundTrips(t *testing.T) {
t.Parallel()
+24 -40
View File
@@ -10,7 +10,6 @@ import (
"path/filepath"
"strconv"
"strings"
"syscall"
"testing"
"time"
@@ -426,51 +425,36 @@ func TestGPGTimeoutKillsGPG(t *testing.T) {
assert.Contains(t, err.Error(), "gpg sign failed: gpg timed out")
}
// TestGPGCancelWhenChildHoldsOutput uses a fake gpg that runs sleep as a
// TestGPGTimeoutWhenChildHoldsOutput uses a fake gpg that runs sleep as a
// child instead of exec-ing it, the way a wrapper script around the real
// gpg might. Killing the fake gpg leaves sleep holding its stdout and
// stderr open; the call must still return once ctx ends instead of waiting
// for sleep to exit. The fake gpg writes the process ID of sleep to a named
// pipe; the test ends ctx only after reading it, so sleep is running by
// then, and kills sleep before returning.
func TestGPGCancelWhenChildHoldsOutput(t *testing.T) {
pidPipe := filepath.Join(t.TempDir(), "sleep.pid")
require.NoError(t, syscall.Mkfifo(pidPipe, 0o600))
// sleep outlasts the 10 s wait below, so a call that waits for it fails.
// stderr open; the call must still return shortly after the deadline
// instead of waiting for sleep to exit. The fake gpg writes the process ID
// of sleep to a file so that the test can kill it before returning.
func TestGPGTimeoutWhenChildHoldsOutput(t *testing.T) {
pidFile := filepath.Join(t.TempDir(), "sleep.pid")
t.Setenv("PATH", fakeGPGPath(t,
"#!/bin/sh\nsleep 60 &\necho $! >'"+pidPipe+"'\nwait\n"))
"#!/bin/sh\nsleep 3 &\necho $! >'"+pidFile+"'\nwait\n"))
t.Cleanup(func() {
pid, err := os.ReadFile(pidFile) //nolint:gosec // G304: path inside t.TempDir()
require.NoError(t, err)
ctx, cancel := context.WithCancel(context.Background())
n, err := strconv.Atoi(strings.TrimSpace(string(pid)))
require.NoError(t, err)
sleep, err := os.FindProcess(n)
require.NoError(t, err)
require.NoError(t, sleep.Kill())
})
ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
signErr := make(chan error, 1)
go func() {
_, err := gpgSign(ctx, []byte("data"), GPGKeyID("any"))
signErr <- err
}()
pid, err := os.ReadFile(pidPipe) //nolint:gosec // G304: path inside t.TempDir()
require.NoError(t, err)
n, err := strconv.Atoi(strings.TrimSpace(string(pid)))
require.NoError(t, err)
sleep, err := os.FindProcess(n)
require.NoError(t, err)
t.Cleanup(func() { require.NoError(t, sleep.Kill()) })
cancel()
// The call should return about gpgWaitDelay (one second) after the
// cancel. 10 s is far above that and well under the 30 s test timeout,
// which would abort the whole package before the cleanup kills sleep.
select {
case err := <-signErr:
require.ErrorIs(t, err, context.Canceled)
case <-time.After(10 * time.Second):
t.Fatal("the call waited for the child holding gpg's output to exit")
}
start := time.Now()
_, err := gpgSign(ctx, []byte("data"), GPGKeyID("any"))
require.ErrorIs(t, err, context.DeadlineExceeded)
assert.Less(t, time.Since(start), 3*time.Second,
"the call waited for the child holding gpg's output to exit")
}
// TestBuildPassesContextToSigning checks that a caller can cancel the gpg
+31 -4
View File
@@ -304,19 +304,46 @@ func TestScannerEnumerateFS(t *testing.T) {
func TestSendEnumerateStatusNonBlocking(t *testing.T) {
t.Parallel()
// Nobody receives, so a blocking send would hang the test into its timeout.
// Channel with no buffer - send should not block
ch := make(chan EnumerateStatus)
sendEnumerateStatus(ch, EnumerateStatus{FilesFound: 1})
// This should not block
done := make(chan bool)
go func() {
sendEnumerateStatus(ch, EnumerateStatus{FilesFound: 1})
done <- true
}()
select {
case <-done:
// Success - did not block
case <-time.After(100 * time.Millisecond):
t.Fatal("sendEnumerateStatus blocked on full channel")
}
}
func TestSendScanStatusNonBlocking(t *testing.T) {
t.Parallel()
// Nobody receives, so a blocking send would hang the test into its timeout.
// Channel with no buffer - send should not block
ch := make(chan ScanStatus)
sendScanStatus(ch, ScanStatus{ScannedFiles: 1})
done := make(chan bool)
go func() {
sendScanStatus(ch, ScanStatus{ScannedFiles: 1})
done <- true
}()
select {
case <-done:
// Success - did not block
case <-time.After(100 * time.Millisecond):
t.Fatal("sendScanStatus blocked on full channel")
}
}
func TestSendStatusNilChannel(t *testing.T) {
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+6 -1
View File
@@ -140,7 +140,12 @@ main() {
# ---- Go repos ----
if missing go; then pkg_install go golang go go; fi
# No golangci-lint: script/lint runs it in Docker only.
# golangci-lint: packaged in nix, brew, and apk. On apt there is no
# package: download a specific release archive from GitHub and
# verify its hash (verify_sha256), never curl | sh.
if missing golangci-lint; then
pkg_install golangci-lint golangci-lint golangci-lint golangci-lint
fi
go mod download
# ---- Python repos ----
+1
View File
@@ -19,6 +19,7 @@ main() {
cd "$ROOT"
ensure_pb
gofumpt -l -w mfer internal cmd
golangci-lint run --fix
# Markdown and JSON, over the same file set script/fmt-check verifies.
"$SCRIPT_DIR/prettier" --write
}
+8 -11
View File
@@ -1,20 +1,17 @@
#!/bin/sh
# script/lint: run golangci-lint, in Docker only. Builds the lint stage of
# the Dockerfile, whose build runs the linter, so a successful build is a
# clean lint. --no-cache because a cached build runs no linter. The image
# is removed afterwards, whatever the outcome.
# script/lint: run the linter.
set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
main() {
cd "$ROOT"
# Tagged per run, so concurrent runs never remove each other's image.
image="$("$SCRIPT_DIR/projectname")-lint:$$"
# A failed build leaves no image, so there is nothing to remove then.
trap 'docker image rm "$image" >/dev/null 2>&1 || true' EXIT INT TERM
docker build --no-cache --target lint -t "$image" .
golangci-lint run
if [ -n "$(gofmt -l .)" ]; then
echo "gofmt: files need formatting:" >&2
gofmt -l . >&2
exit 1
fi
}
main "$@"