2 Commits
Author SHA1 Message Date
sneak e7331e8d11 Reject manifests whose file entries decode far larger than their bytes (closes #123)
check / check (push) Waiting to run
Parser fix: before decoding the manifest, the parser walks its file
entries and adds up what decoding sets aside for each entry, hash,
timestamp and MIME type, however short its encoding. It refuses the
manifest once that sum passes 8 times the decompressed size; the
densest manifests mfer writes come to about 7 times. Empty entries
decoded to about 50 times their size, so a 1.6 KB manifest allocated
nearly 1 GB. The fuzz target's ceiling falls to 20 times the input and
decompressed data, and a new seed of entries holding only an empty MIME
type and empty times fails it without the fix.

Model: opus-5-5
2026-10-04 05:56:21 +00:00
clawbot a2732cf8da Add the required README sections (closes #75)
check / check (push) Successful in 1m28s
The Description first line now names the project, purpose, category,
WTFPL license, and author. A new Getting Started section gives a
copy-pasteable build-from-source block and gen/check/fetch usage. Problem
Statement and Proposed Solution move under a new Rationale heading, the
prose kept. A new Design section documents the package layout: the mfer/
library with its committed protobuf code, the internal/cli commands,
internal/log and internal/bork, and the cmd/mfer entrypoint. Authors is
renamed Author with the canonical link.

Getting Started uses go build because the make target that builds the
binary runs protoc, which script/bootstrap does not install.

Model: opus-4-8 (implementation); opus-5-5 (rebase, rework)
2026-10-04 06:32:07 +02:00
6 changed files with 196 additions and 75 deletions
+65 -11
View File
@@ -1,12 +1,12 @@
# mfer # mfer
[mfer](https://git.eeqj.de/sneak/mfer) is a reference implementation library and [mfer](https://git.eeqj.de/sneak/mfer) is a [WTFPL](https://wtfpl.net)-licensed
thin wrapper command-line utility written in [Go](https://golang.org) and first (public domain) [Go](https://golang.org) library and command-line tool by
published in 2022 under the [WTFPL](https://wtfpl.net) (public domain) license. [@sneak](https://sneak.berlin) that specifies and generates `.mf` manifest files
It specifies and generates `.mf` manifest files over a directory tree of files over a directory tree to encapsulate metadata about the files — such as
to encapsulate metadata about them (such as cryptographic checksums or cryptographic checksums and signatures over same — to aid in archiving,
signatures over same) to aid in archiving, downloading, and streaming, or downloading, streaming, and mirroring. It was first published in 2022. The
mirroring. The manifest files' data is serialized with Google's manifest files' data is serialized with Google's
[protobuf serialization format](https://developers.google.com/protocol-buffers). [protobuf serialization format](https://developers.google.com/protocol-buffers).
The structure of these files can be found The structure of these files can be found
[in the format specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto) [in the format specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto)
@@ -21,6 +21,36 @@ This project was started by [@sneak](https://sneak.berlin) to scratch an itch in
as a de-facto standard and be incorporated into other software. A compatible as a de-facto standard and be incorporated into other software. A compatible
javascript library is planned. javascript library is planned.
# Getting Started
`mfer` builds from source with a Go 1.23+ toolchain. The generated protobuf code
is committed, so no `protoc` toolchain is required:
```sh
git clone https://git.eeqj.de/sneak/mfer.git
cd mfer
go build -o bin/mfer ./cmd/mfer
```
Generate a manifest for a directory tree, verify it later, and fetch a published
tree by URL:
```sh
# Write .index.mf, a manifest of the files under the current directory.
bin/mfer gen .
# Verify the files on disk against the manifest. Exits nonzero if any file
# is missing or corrupted.
bin/mfer check .index.mf
# Download and cryptographically verify a tree published over HTTP: mfer
# fetches <url>/index.mf, then downloads every file it lists.
bin/mfer fetch https://example.com/tree/
```
Run `bin/mfer help` for the full command list, or `bin/mfer <command> --help`
for a single command's options.
# Build Status # Build Status
CI runs `script/cibuild`, which builds the Docker image with `--no-cache`, so CI runs `script/cibuild`, which builds the Docker image with `--no-cache`, so
@@ -85,7 +115,9 @@ Any changes submitted to this project must also be
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards, See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
tooling requirements, and workflow conventions. tooling requirements, and workflow conventions.
# Problem Statement # Rationale
## The problem
Given a plain URL, there is no standard way to safely and programmatically Given a plain URL, there is no standard way to safely and programmatically
download everything "under" that URL path. `wget -r` can traverse directory download everything "under" that URL path. `wget -r` can traverse directory
@@ -109,7 +141,7 @@ Real issues I face:
- when I download a large file via HTTP, I have no way of knowing if the file - when I download a large file via HTTP, I have no way of knowing if the file
content is what it's supposed to be content is what it's supposed to be
# Proposed Solution ## The solution
A standard, a manifest file format, and a tool for generating same. A standard, a manifest file format, and a tool for generating same.
@@ -141,6 +173,28 @@ The manifest file would do several important things:
- maybe a bittorrent chunklist for torrent client compatibility? perhaps a - maybe a bittorrent chunklist for torrent client compatibility? perhaps a
top-level infohash for the whole manifest? top-level infohash for the whole manifest?
# Design
The repository is split into a reusable library and a thin command-line wrapper
around it.
- `mfer/` is the reusable library and the heart of the project: it defines the
manifest format and implements building, scanning, checking, serialization,
and signing. The protobuf schema is `mfer/mf.proto`, and the generated code it
produces (`mfer/mf.pb.go`) is committed alongside it so the library builds
with `go get` and needs no `protoc` toolchain.
- `internal/cli/` holds the command implementations — `generate` (alias `gen`),
`check`, `freshen`, `export`, `list` (alias `ls`), `fetch`, and `version` —
that wire the library to the command-line interface.
- `internal/log/` provides the logging used by the commands and the library.
- `internal/bork/` provides the error the library returns when a manifest's
decompressed contents are not the size the manifest records.
- `cmd/mfer/` is the entrypoint: its `main` package runs `internal/cli` and
exits with the status it returns.
Everything under `internal/` is private to this repository; only the `mfer/`
package is intended for import by other software.
# Design Goals # Design Goals
- Replace SHASUMS/SHASUMS.asc files - Replace SHASUMS/SHASUMS.asc files
@@ -274,9 +328,9 @@ Open work, open design questions included, is tracked in this repo's issues:
- Issues: - Issues:
[https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues) [https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues)
# Authors # Author
- [@sneak &lt;sneak@sneak.berlin&gt;](mailto:sneak@sneak.berlin) - [@sneak](https://sneak.berlin)
# License # License
+17 -10
View File
@@ -18,16 +18,23 @@ const (
// uuidLength is the length in bytes of a binary UUID. // uuidLength is the length in bytes of a binary UUID.
uuidLength = 16 uuidLength = 16
// filesFieldNumber and hashesFieldNumber are the numbers of // Numbers in mf.proto of MFFile.files and of the MFFilePath fields
// MFFile.files and MFFilePath.hashes in mf.proto. // that decoding sets aside a fixed amount of memory for.
filesFieldNumber = 101 filesFieldNumber = 101
hashesFieldNumber = 3 hashesFieldNumber = 3
mimeTypeFieldNumber = 301
mtimeFieldNumber = 302
ctimeFieldNumber = 303
// minHashSize is the encoded size of the smallest hash: a two-byte // Bytes decoding sets aside for each file entry, hash, timestamp and
// multihash (algorithm code, zero digest length) after its tag and length. // MIME type, however short its encoding. checkDecodedSize refuses an
minHashSize = 2 + 2 // inner message for which these add up to more than maxDecodedGrowth
// times its size.
decodedFileEntrySize = 160
decodedHashSize = 112
decodedTimestampSize = 64
decodedMIMETypeSize = 16
// minFileEntrySize is the encoded size of the smallest file entry: a // The densest manifests mfer writes add up to about 7 times their size.
// one-byte path and one hash, each after its tag and length. maxDecodedGrowth = 8
minFileEntrySize = 2 + 1 + 2 + minHashSize
) )
+49 -21
View File
@@ -28,8 +28,8 @@ var (
errUUIDMismatch = errors.New("outer and inner UUID mismatch") errUUIDMismatch = errors.New("outer and inner UUID mismatch")
errInvalidFileFormat = errors.New("invalid file format") errInvalidFileFormat = errors.New("invalid file format")
errInvalidManifestPath = errors.New("manifest contains invalid path") errInvalidManifestPath = errors.New("manifest contains invalid path")
errEntryTooShort = errors.New( errDecodedTooLarge = errors.New(
"manifest contains a file entry or hash shorter than the format allows") "manifest would take too much memory to decode")
) )
// validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable). // validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable).
@@ -157,19 +157,46 @@ func (m *manifest) decompressInner() ([]byte, error) {
return dat, nil return dat, nil
} }
// checkEntrySizes rejects an encoded inner message holding a file entry or // checkDecodedSize refuses an encoded inner message whose file entries,
// a hash shorter than the format allows. Decoding allocates a fixed amount // hashes, timestamps and MIME types would take more than maxDecodedGrowth
// for each entry and each hash, however short, so a payload of empty ones // times its size to decode. Decoding sets aside a fixed amount for each,
// would decode to about 50 times its size. // however short its encoding, so a message of empty ones would take about
func checkEntrySizes(inner []byte) error { // 50 times its size.
return forEachBytesField(inner, filesFieldNumber, func(entry []byte) error { func checkDecodedSize(inner []byte) error {
if len(entry) < minFileEntrySize { limit := maxDecodedGrowth * int64(len(inner))
return errEntryTooShort
var decoded int64
add := func(size int64) error {
decoded += size
if decoded > limit {
return errDecodedTooLarge
} }
return forEachBytesField(entry, hashesFieldNumber, func(hash []byte) error { return nil
if len(hash) < minHashSize { }
return errEntryTooShort
return forEachBytesField(inner, func(num protowire.Number, entry []byte) error {
if num != filesFieldNumber {
return nil
}
err := add(decodedFileEntrySize)
if err != nil {
return err
}
return forEachBytesField(entry, func(num protowire.Number, _ []byte) error {
if num == hashesFieldNumber {
return add(decodedHashSize)
}
if num == mtimeFieldNumber || num == ctimeFieldNumber {
return add(decodedTimestampSize)
}
if num == mimeTypeFieldNumber {
return add(decodedMIMETypeSize)
} }
return nil return nil
@@ -177,26 +204,27 @@ func checkEntrySizes(inner []byte) error {
}) })
} }
// forEachBytesField calls fn with the value of each length-delimited field // forEachBytesField calls fn with the number and value of each
// numbered num in the encoded message msg, and fails if msg is malformed. // length-delimited field in the encoded message msg, and fails if msg is
// malformed.
func forEachBytesField( func forEachBytesField(
msg []byte, num protowire.Number, fn func(value []byte) error, msg []byte, fn func(num protowire.Number, value []byte) error,
) error { ) error {
for len(msg) > 0 { for len(msg) > 0 {
fieldNum, wireType, tagLen := protowire.ConsumeTag(msg) num, wireType, tagLen := protowire.ConsumeTag(msg)
if tagLen < 0 { if tagLen < 0 {
return protowire.ParseError(tagLen) return protowire.ParseError(tagLen)
} }
valueLen := protowire.ConsumeFieldValue(fieldNum, wireType, msg[tagLen:]) valueLen := protowire.ConsumeFieldValue(num, wireType, msg[tagLen:])
if valueLen < 0 { if valueLen < 0 {
return protowire.ParseError(valueLen) return protowire.ParseError(valueLen)
} }
if fieldNum == num && wireType == protowire.BytesType { if wireType == protowire.BytesType {
value, _ := protowire.ConsumeBytes(msg[tagLen:]) value, _ := protowire.ConsumeBytes(msg[tagLen:])
err := fn(value) err := fn(num, value)
if err != nil { if err != nil {
return err return err
} }
@@ -231,7 +259,7 @@ func (m *manifest) deserializeInner() error {
return bork.ErrFileTruncated return bork.ErrFileTruncated
} }
err = checkEntrySizes(dat) err = checkDecodedSize(dat)
if err != nil { if err != nil {
return fmt.Errorf("deserialize: unmarshal inner: %w", err) return fmt.Errorf("deserialize: unmarshal inner: %w", err)
} }
+12 -9
View File
@@ -54,11 +54,14 @@ func FuzzNewManifestFromReader(f *testing.F) {
} }
// It also keeps a few copies of its input. Buffers grow by // It also keeps a few copies of its input. Buffers grow by
// copying, so reaching those sizes allocates a few times them in // copying, so reaching those sizes allocates up to about six times
// total. Decoding the decompressed data takes up to about 25 times // them in total. Decoding the decompressed data takes up to
// its size, when every file entry and hash is as short as the // maxDecodedGrowth times its size for file entries, hashes,
// parser accepts. Thirty-six times the input and the decompressed // timestamps and MIME types, and up to about five times more for
// data leaves room for both. // the bytes it copies out of it, such as fields it does not know,
// which it keeps in buffers that also grow by copying. Twenty
// times the input and the decompressed data leaves room for all of
// that.
// //
// The decoder also sets aside a new buffer of one to two times the // The decoder also sets aside a new buffer of one to two times the
// window for each frame that asks for a larger window than the // window for each frame that asks for a larger window than the
@@ -73,10 +76,10 @@ func FuzzNewManifestFromReader(f *testing.F) {
// fails if the decoder accepts windows of twice zstdWindowSize; the // fails if the decoder accepts windows of twice zstdWindowSize; the
// seed whose two frames together exceed MaxDecompressedSize fails // seed whose two frames together exceed MaxDecompressedSize fails
// if the decoder decodes them in full instead of stopping at the // if the decoder decodes them in full instead of stopping at the
// declared size; the seeds of empty file entries and of a file // declared size; the seeds of empty file entries, of a file entry
// entry of empty hashes fail if the parser decodes entries or // of empty hashes, and of file entries of only an empty MIME type
// hashes shorter than the format allows. // and empty times fail if the parser decodes them.
limit := 36*(uint64(len(data))+decompressed) + 24*zstdWindowSize limit := 20*(uint64(len(data))+decompressed) + 24*zstdWindowSize
allocated := after.TotalAlloc - before.TotalAlloc allocated := after.TotalAlloc - before.TotalAlloc
if allocated > limit { if allocated > limit {
+51 -24
View File
@@ -6,7 +6,9 @@ import (
"context" "context"
"crypto/sha256" "crypto/sha256"
"fmt" "fmt"
"strconv"
"testing" "testing"
"time"
"github.com/google/uuid" "github.com/google/uuid"
"github.com/klauspost/compress/zstd" "github.com/klauspost/compress/zstd"
@@ -17,18 +19,12 @@ import (
) )
// craftInnerBytes builds the wire bytes of an inner MFFile holding a single // craftInnerBytes builds the wire bytes of an inner MFFile holding a single
// file entry whose path is exactly pathBytes and whose one hash is multihash. // file entry whose path is exactly pathBytes. It writes the wire form by hand
// It writes the wire form by hand so a hostile path — including one that is // so a hostile path — including one that is not valid UTF-8 — can be embedded
// not valid UTF-8 — can be embedded without proto.Marshal's own UTF-8 // without proto.Marshal's own UTF-8 enforcement rejecting it first.
// enforcement rejecting it first. func craftInnerBytes(id uuid.UUID, pathBytes string) []byte {
func craftInnerBytes(id uuid.UUID, pathBytes string, multihash []byte) []byte {
hash := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFileChecksum.multiHash
hash = protowire.AppendBytes(hash, multihash)
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
entry = protowire.AppendString(entry, pathBytes) entry = protowire.AppendString(entry, pathBytes)
entry = protowire.AppendTag(entry, 3, protowire.BytesType) // MFFilePath.hashes
entry = protowire.AppendBytes(entry, hash)
inner := protowire.AppendTag(nil, 100, protowire.VarintType) // MFFile.version inner := protowire.AppendTag(nil, 100, protowire.VarintType) // MFFile.version
inner = protowire.AppendVarint(inner, uint64(MFFile_VERSION_ONE)) inner = protowire.AppendVarint(inner, uint64(MFFile_VERSION_ONE))
@@ -95,8 +91,7 @@ func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
t.Parallel() t.Parallel()
id := uuid.New() id := uuid.New()
hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256 data := wrapInner(t, id, craftInnerBytes(id, tt.path))
data := wrapInner(t, id, craftInnerBytes(id, tt.path, hash))
_, err := NewManifestFromReader(bytes.NewReader(data)) _, err := NewManifestFromReader(bytes.NewReader(data))
require.Error(t, err) require.Error(t, err)
@@ -121,23 +116,55 @@ func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
} }
} }
// A one-byte path and a multihash of algorithm code and zero digest length // Entries of a one-character path and empty modification and change times
// make the smallest file entry and hash the format allows: they load, and an // pass every other check, but would take about 23 times their size to decode.
// entry or hash one byte shorter is refused before decoding. func TestDeserializeRefusesEntriesThatDecodeTooLarge(t *testing.T) {
func TestDeserializeRejectsEntriesShorterThanFormatAllows(t *testing.T) {
t.Parallel() t.Parallel()
load := func(path string, multihash []byte) error { entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
id := uuid.New() entry = protowire.AppendString(entry, "a")
data := wrapInner(t, id, craftInnerBytes(id, path, multihash)) entry = protowire.AppendTag(entry, 302, protowire.BytesType) // MFFilePath.mtime
_, err := NewManifestFromReader(bytes.NewReader(data)) entry = protowire.AppendBytes(entry, nil)
entry = protowire.AppendTag(entry, 303, protowire.BytesType) // MFFilePath.ctime
entry = protowire.AppendBytes(entry, nil)
return err id := uuid.New()
inner := protowire.AppendTag(nil, 102, protowire.BytesType) // MFFile.uuid
inner = protowire.AppendBytes(inner, id[:])
for range 1000 {
inner = protowire.AppendTag(inner, 101, protowire.BytesType) // MFFile.files
inner = protowire.AppendBytes(inner, entry)
} }
require.NoError(t, load("a", []byte{0, 0})) _, err := NewManifestFromReader(bytes.NewReader(wrapInner(t, id, inner)))
require.ErrorIs(t, load("", []byte{0, 0}), errEntryTooShort) require.ErrorIs(t, err, errDecodedTooLarge)
require.ErrorIs(t, load("ab", []byte{0}), errEntryTooShort) }
// Many empty files with names of at most three characters and modification
// times at the epoch make about the densest manifest mfer writes: it takes
// about 7 times its size to decode, and still loads. A signature would not
// change the inner message, so none is added.
func TestDeserializeLoadsDensestManifest(t *testing.T) {
t.Parallel()
hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256
b := NewBuilder()
b.SetIncludeTimestamps(true)
const files = 10000
for i := range files {
name := RelFilePath(strconv.FormatInt(int64(i), 36))
require.NoError(t, b.AddFileWithHash(name, 0, ModTime(time.Unix(0, 0)), hash))
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.Len(t, m.Files(), files)
} }
func TestDeserializeValidManifestRoundTrips(t *testing.T) { func TestDeserializeValidManifestRoundTrips(t *testing.T) {
File diff suppressed because one or more lines are too long