2 Commits
Author SHA1 Message Date
sneak e7331e8d11 Reject manifests whose file entries decode far larger than their bytes (closes #123)
check / check (push) Waiting to run
Parser fix: before decoding the manifest, the parser walks its file
entries and adds up what decoding sets aside for each entry, hash,
timestamp and MIME type, however short its encoding. It refuses the
manifest once that sum passes 8 times the decompressed size; the
densest manifests mfer writes come to about 7 times. Empty entries
decoded to about 50 times their size, so a 1.6 KB manifest allocated
nearly 1 GB. The fuzz target's ceiling falls to 20 times the input and
decompressed data, and a new seed of entries holding only an empty MIME
type and empty times fails it without the fix.

Model: opus-5-5
2026-10-04 05:56:21 +00:00
clawbot a2732cf8da Add the required README sections (closes #75)
check / check (push) Waiting to run
The Description first line now names the project, purpose, category,
WTFPL license, and author. A new Getting Started section gives a
copy-pasteable build-from-source block and gen/check/fetch usage. Problem
Statement and Proposed Solution move under a new Rationale heading, the
prose kept. A new Design section documents the package layout: the mfer/
library with its committed protobuf code, the internal/cli commands,
internal/log and internal/bork, and the cmd/mfer entrypoint. Authors is
renamed Author with the canonical link.

Getting Started uses go build because the make target that builds the
binary runs protoc, which script/bootstrap does not install.

Model: opus-4-8 (implementation); opus-5-5 (rebase, rework)
2026-10-04 06:32:07 +02:00
8 changed files with 243 additions and 16 deletions
+65 -11
View File
@@ -1,12 +1,12 @@
# mfer
[mfer](https://git.eeqj.de/sneak/mfer) is a reference implementation library and
thin wrapper command-line utility written in [Go](https://golang.org) and first
published in 2022 under the [WTFPL](https://wtfpl.net) (public domain) license.
It specifies and generates `.mf` manifest files over a directory tree of files
to encapsulate metadata about them (such as cryptographic checksums or
signatures over same) to aid in archiving, downloading, and streaming, or
mirroring. The manifest files' data is serialized with Google's
[mfer](https://git.eeqj.de/sneak/mfer) is a [WTFPL](https://wtfpl.net)-licensed
(public domain) [Go](https://golang.org) library and command-line tool by
[@sneak](https://sneak.berlin) that specifies and generates `.mf` manifest files
over a directory tree to encapsulate metadata about the files — such as
cryptographic checksums and signatures over same — to aid in archiving,
downloading, streaming, and mirroring. It was first published in 2022. The
manifest files' data is serialized with Google's
[protobuf serialization format](https://developers.google.com/protocol-buffers).
The structure of these files can be found
[in the format specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto)
@@ -21,6 +21,36 @@ This project was started by [@sneak](https://sneak.berlin) to scratch an itch in
as a de-facto standard and be incorporated into other software. A compatible
javascript library is planned.
# Getting Started
`mfer` builds from source with a Go 1.23+ toolchain. The generated protobuf code
is committed, so no `protoc` toolchain is required:
```sh
git clone https://git.eeqj.de/sneak/mfer.git
cd mfer
go build -o bin/mfer ./cmd/mfer
```
Generate a manifest for a directory tree, verify it later, and fetch a published
tree by URL:
```sh
# Write .index.mf, a manifest of the files under the current directory.
bin/mfer gen .
# Verify the files on disk against the manifest. Exits nonzero if any file
# is missing or corrupted.
bin/mfer check .index.mf
# Download and cryptographically verify a tree published over HTTP: mfer
# fetches <url>/index.mf, then downloads every file it lists.
bin/mfer fetch https://example.com/tree/
```
Run `bin/mfer help` for the full command list, or `bin/mfer <command> --help`
for a single command's options.
# Build Status
CI runs `script/cibuild`, which builds the Docker image with `--no-cache`, so
@@ -85,7 +115,9 @@ Any changes submitted to this project must also be
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
tooling requirements, and workflow conventions.
# Problem Statement
# Rationale
## The problem
Given a plain URL, there is no standard way to safely and programmatically
download everything "under" that URL path. `wget -r` can traverse directory
@@ -109,7 +141,7 @@ Real issues I face:
- when I download a large file via HTTP, I have no way of knowing if the file
content is what it's supposed to be
# Proposed Solution
## The solution
A standard, a manifest file format, and a tool for generating same.
@@ -141,6 +173,28 @@ The manifest file would do several important things:
- maybe a bittorrent chunklist for torrent client compatibility? perhaps a
top-level infohash for the whole manifest?
# Design
The repository is split into a reusable library and a thin command-line wrapper
around it.
- `mfer/` is the reusable library and the heart of the project: it defines the
manifest format and implements building, scanning, checking, serialization,
and signing. The protobuf schema is `mfer/mf.proto`, and the generated code it
produces (`mfer/mf.pb.go`) is committed alongside it so the library builds
with `go get` and needs no `protoc` toolchain.
- `internal/cli/` holds the command implementations — `generate` (alias `gen`),
`check`, `freshen`, `export`, `list` (alias `ls`), `fetch`, and `version` —
that wire the library to the command-line interface.
- `internal/log/` provides the logging used by the commands and the library.
- `internal/bork/` provides the error the library returns when a manifest's
decompressed contents are not the size the manifest records.
- `cmd/mfer/` is the entrypoint: its `main` package runs `internal/cli` and
exits with the status it returns.
Everything under `internal/` is private to this repository; only the `mfer/`
package is intended for import by other software.
# Design Goals
- Replace SHASUMS/SHASUMS.asc files
@@ -274,9 +328,9 @@ Open work, open design questions included, is tracked in this repo's issues:
- Issues:
[https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues)
# Authors
# Author
- [@sneak &lt;sneak@sneak.berlin&gt;](mailto:sneak@sneak.berlin)
- [@sneak](https://sneak.berlin)
# License
+20
View File
@@ -17,4 +17,24 @@ const (
// uuidLength is the length in bytes of a binary UUID.
uuidLength = 16
// Numbers in mf.proto of MFFile.files and of the MFFilePath fields
// that decoding sets aside a fixed amount of memory for.
filesFieldNumber = 101
hashesFieldNumber = 3
mimeTypeFieldNumber = 301
mtimeFieldNumber = 302
ctimeFieldNumber = 303
// Bytes decoding sets aside for each file entry, hash, timestamp and
// MIME type, however short its encoding. checkDecodedSize refuses an
// inner message for which these add up to more than maxDecodedGrowth
// times its size.
decodedFileEntrySize = 160
decodedHashSize = 112
decodedTimestampSize = 64
decodedMIMETypeSize = 16
// The densest manifests mfer writes add up to about 7 times their size.
maxDecodedGrowth = 8
)
+87
View File
@@ -11,6 +11,7 @@ import (
"github.com/google/uuid"
"github.com/klauspost/compress/zstd"
"github.com/spf13/afero"
"google.golang.org/protobuf/encoding/protowire"
"google.golang.org/protobuf/proto"
"sneak.berlin/go/mfer/internal/bork"
"sneak.berlin/go/mfer/internal/log"
@@ -27,6 +28,8 @@ var (
errUUIDMismatch = errors.New("outer and inner UUID mismatch")
errInvalidFileFormat = errors.New("invalid file format")
errInvalidManifestPath = errors.New("manifest contains invalid path")
errDecodedTooLarge = errors.New(
"manifest would take too much memory to decode")
)
// validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable).
@@ -154,6 +157,85 @@ func (m *manifest) decompressInner() ([]byte, error) {
return dat, nil
}
// checkDecodedSize refuses an encoded inner message whose file entries,
// hashes, timestamps and MIME types would take more than maxDecodedGrowth
// times its size to decode. Decoding sets aside a fixed amount for each,
// however short its encoding, so a message of empty ones would take about
// 50 times its size.
func checkDecodedSize(inner []byte) error {
limit := maxDecodedGrowth * int64(len(inner))
var decoded int64
add := func(size int64) error {
decoded += size
if decoded > limit {
return errDecodedTooLarge
}
return nil
}
return forEachBytesField(inner, func(num protowire.Number, entry []byte) error {
if num != filesFieldNumber {
return nil
}
err := add(decodedFileEntrySize)
if err != nil {
return err
}
return forEachBytesField(entry, func(num protowire.Number, _ []byte) error {
if num == hashesFieldNumber {
return add(decodedHashSize)
}
if num == mtimeFieldNumber || num == ctimeFieldNumber {
return add(decodedTimestampSize)
}
if num == mimeTypeFieldNumber {
return add(decodedMIMETypeSize)
}
return nil
})
})
}
// forEachBytesField calls fn with the number and value of each
// length-delimited field in the encoded message msg, and fails if msg is
// malformed.
func forEachBytesField(
msg []byte, fn func(num protowire.Number, value []byte) error,
) error {
for len(msg) > 0 {
num, wireType, tagLen := protowire.ConsumeTag(msg)
if tagLen < 0 {
return protowire.ParseError(tagLen)
}
valueLen := protowire.ConsumeFieldValue(num, wireType, msg[tagLen:])
if valueLen < 0 {
return protowire.ParseError(valueLen)
}
if wireType == protowire.BytesType {
value, _ := protowire.ConsumeBytes(msg[tagLen:])
err := fn(num, value)
if err != nil {
return err
}
}
msg = msg[tagLen+valueLen:]
}
return nil
}
func (m *manifest) deserializeInner() error {
err := m.validateOuterHeader()
if err != nil {
@@ -177,6 +259,11 @@ func (m *manifest) deserializeInner() error {
return bork.ErrFileTruncated
}
err = checkDecodedSize(dat)
if err != nil {
return fmt.Errorf("deserialize: unmarshal inner: %w", err)
}
// Deserialize inner message
m.pbInner = new(MFFile)
+12 -5
View File
@@ -54,9 +54,14 @@ func FuzzNewManifestFromReader(f *testing.F) {
}
// It also keeps a few copies of its input. Buffers grow by
// copying, so reaching those sizes allocates a few times them in
// total: sixteen times the input and the decompressed data leaves
// room for that.
// copying, so reaching those sizes allocates up to about six times
// them in total. Decoding the decompressed data takes up to
// maxDecodedGrowth times its size for file entries, hashes,
// timestamps and MIME types, and up to about five times more for
// the bytes it copies out of it, such as fields it does not know,
// which it keeps in buffers that also grow by copying. Twenty
// times the input and the decompressed data leaves room for all of
// that.
//
// The decoder also sets aside a new buffer of one to two times the
// window for each frame that asks for a larger window than the
@@ -71,8 +76,10 @@ func FuzzNewManifestFromReader(f *testing.F) {
// fails if the decoder accepts windows of twice zstdWindowSize; the
// seed whose two frames together exceed MaxDecompressedSize fails
// if the decoder decodes them in full instead of stopping at the
// declared size.
limit := 16*(uint64(len(data))+decompressed) + 24*zstdWindowSize
// declared size; the seeds of empty file entries, of a file entry
// of empty hashes, and of file entries of only an empty MIME type
// and empty times fail if the parser decodes them.
limit := 20*(uint64(len(data))+decompressed) + 24*zstdWindowSize
allocated := after.TotalAlloc - before.TotalAlloc
if allocated > limit {
+53
View File
@@ -6,7 +6,9 @@ import (
"context"
"crypto/sha256"
"fmt"
"strconv"
"testing"
"time"
"github.com/google/uuid"
"github.com/klauspost/compress/zstd"
@@ -114,6 +116,57 @@ func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
}
}
// Entries of a one-character path and empty modification and change times
// pass every other check, but would take about 23 times their size to decode.
func TestDeserializeRefusesEntriesThatDecodeTooLarge(t *testing.T) {
t.Parallel()
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
entry = protowire.AppendString(entry, "a")
entry = protowire.AppendTag(entry, 302, protowire.BytesType) // MFFilePath.mtime
entry = protowire.AppendBytes(entry, nil)
entry = protowire.AppendTag(entry, 303, protowire.BytesType) // MFFilePath.ctime
entry = protowire.AppendBytes(entry, nil)
id := uuid.New()
inner := protowire.AppendTag(nil, 102, protowire.BytesType) // MFFile.uuid
inner = protowire.AppendBytes(inner, id[:])
for range 1000 {
inner = protowire.AppendTag(inner, 101, protowire.BytesType) // MFFile.files
inner = protowire.AppendBytes(inner, entry)
}
_, err := NewManifestFromReader(bytes.NewReader(wrapInner(t, id, inner)))
require.ErrorIs(t, err, errDecodedTooLarge)
}
// Many empty files with names of at most three characters and modification
// times at the epoch make about the densest manifest mfer writes: it takes
// about 7 times its size to decode, and still loads. A signature would not
// change the inner message, so none is added.
func TestDeserializeLoadsDensestManifest(t *testing.T) {
t.Parallel()
hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256
b := NewBuilder()
b.SetIncludeTimestamps(true)
const files = 10000
for i := range files {
name := RelFilePath(strconv.FormatInt(int64(i), 36))
require.NoError(t, b.AddFileWithHash(name, 0, ModTime(time.Unix(0, 0)), hash))
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.Len(t, m.Files(), files)
}
func TestDeserializeValidManifestRoundTrips(t *testing.T) {
t.Parallel()
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long