Compare commits

..
1 Commits
Author SHA1 Message Date
sneak 09802fde10 Reject manifests whose file entries decode far larger than their bytes (closes #123)
check / check (push) Successful in 1m23s
Parser fix: before decoding the manifest, the parser now walks its file
entries and their hashes and rejects any entry shorter than 9 bytes or
hash shorter than 4, the smallest the format allows. Decoding allocates
a fixed amount per entry and per hash, so empty ones decoded to about 50
times their size: a 1.6 KB manifest allocated nearly 1 GB. What the
parser accepts now decodes to at most about 25 times its size, so the
fuzz target's ceiling rises from 16 to 36 times the input and
decompressed data. New seeds of empty entries and of empty hashes fail
the fuzz target without the fix.

Model: opus-5-5
2026-10-04 04:06:52 +00:00
6 changed files with 75 additions and 196 deletions
+11 -65
View File
@@ -1,12 +1,12 @@
# mfer # mfer
[mfer](https://git.eeqj.de/sneak/mfer) is a [WTFPL](https://wtfpl.net)-licensed [mfer](https://git.eeqj.de/sneak/mfer) is a reference implementation library and
(public domain) [Go](https://golang.org) library and command-line tool by thin wrapper command-line utility written in [Go](https://golang.org) and first
[@sneak](https://sneak.berlin) that specifies and generates `.mf` manifest files published in 2022 under the [WTFPL](https://wtfpl.net) (public domain) license.
over a directory tree to encapsulate metadata about the files — such as It specifies and generates `.mf` manifest files over a directory tree of files
cryptographic checksums and signatures over same — to aid in archiving, to encapsulate metadata about them (such as cryptographic checksums or
downloading, streaming, and mirroring. It was first published in 2022. The signatures over same) to aid in archiving, downloading, and streaming, or
manifest files' data is serialized with Google's mirroring. The manifest files' data is serialized with Google's
[protobuf serialization format](https://developers.google.com/protocol-buffers). [protobuf serialization format](https://developers.google.com/protocol-buffers).
The structure of these files can be found The structure of these files can be found
[in the format specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto) [in the format specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto)
@@ -21,36 +21,6 @@ This project was started by [@sneak](https://sneak.berlin) to scratch an itch in
as a de-facto standard and be incorporated into other software. A compatible as a de-facto standard and be incorporated into other software. A compatible
javascript library is planned. javascript library is planned.
# Getting Started
`mfer` builds from source with a Go 1.23+ toolchain. The generated protobuf code
is committed, so no `protoc` toolchain is required:
```sh
git clone https://git.eeqj.de/sneak/mfer.git
cd mfer
go build -o bin/mfer ./cmd/mfer
```
Generate a manifest for a directory tree, verify it later, and fetch a published
tree by URL:
```sh
# Write .index.mf, a manifest of the files under the current directory.
bin/mfer gen .
# Verify the files on disk against the manifest. Exits nonzero if any file
# is missing or corrupted.
bin/mfer check .index.mf
# Download and cryptographically verify a tree published over HTTP: mfer
# fetches <url>/index.mf, then downloads every file it lists.
bin/mfer fetch https://example.com/tree/
```
Run `bin/mfer help` for the full command list, or `bin/mfer <command> --help`
for a single command's options.
# Build Status # Build Status
CI runs `script/cibuild`, which builds the Docker image with `--no-cache`, so CI runs `script/cibuild`, which builds the Docker image with `--no-cache`, so
@@ -115,9 +85,7 @@ Any changes submitted to this project must also be
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards, See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
tooling requirements, and workflow conventions. tooling requirements, and workflow conventions.
# Rationale # Problem Statement
## The problem
Given a plain URL, there is no standard way to safely and programmatically Given a plain URL, there is no standard way to safely and programmatically
download everything "under" that URL path. `wget -r` can traverse directory download everything "under" that URL path. `wget -r` can traverse directory
@@ -141,7 +109,7 @@ Real issues I face:
- when I download a large file via HTTP, I have no way of knowing if the file - when I download a large file via HTTP, I have no way of knowing if the file
content is what it's supposed to be content is what it's supposed to be
## The solution # Proposed Solution
A standard, a manifest file format, and a tool for generating same. A standard, a manifest file format, and a tool for generating same.
@@ -173,28 +141,6 @@ The manifest file would do several important things:
- maybe a bittorrent chunklist for torrent client compatibility? perhaps a - maybe a bittorrent chunklist for torrent client compatibility? perhaps a
top-level infohash for the whole manifest? top-level infohash for the whole manifest?
# Design
The repository is split into a reusable library and a thin command-line wrapper
around it.
- `mfer/` is the reusable library and the heart of the project: it defines the
manifest format and implements building, scanning, checking, serialization,
and signing. The protobuf schema is `mfer/mf.proto`, and the generated code it
produces (`mfer/mf.pb.go`) is committed alongside it so the library builds
with `go get` and needs no `protoc` toolchain.
- `internal/cli/` holds the command implementations — `generate` (alias `gen`),
`check`, `freshen`, `export`, `list` (alias `ls`), `fetch`, and `version` —
that wire the library to the command-line interface.
- `internal/log/` provides the logging used by the commands and the library.
- `internal/bork/` provides the error the library returns when a manifest's
decompressed contents are not the size the manifest records.
- `cmd/mfer/` is the entrypoint: its `main` package runs `internal/cli` and
exits with the status it returns.
Everything under `internal/` is private to this repository; only the `mfer/`
package is intended for import by other software.
# Design Goals # Design Goals
- Replace SHASUMS/SHASUMS.asc files - Replace SHASUMS/SHASUMS.asc files
@@ -328,9 +274,9 @@ Open work, open design questions included, is tracked in this repo's issues:
- Issues: - Issues:
[https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues) [https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues)
# Author # Authors
- [@sneak](https://sneak.berlin) - [@sneak &lt;sneak@sneak.berlin&gt;](mailto:sneak@sneak.berlin)
# License # License
+10 -17
View File
@@ -18,23 +18,16 @@ const (
// uuidLength is the length in bytes of a binary UUID. // uuidLength is the length in bytes of a binary UUID.
uuidLength = 16 uuidLength = 16
// Numbers in mf.proto of MFFile.files and of the MFFilePath fields // filesFieldNumber and hashesFieldNumber are the numbers of
// that decoding sets aside a fixed amount of memory for. // MFFile.files and MFFilePath.hashes in mf.proto.
filesFieldNumber = 101 filesFieldNumber = 101
hashesFieldNumber = 3 hashesFieldNumber = 3
mimeTypeFieldNumber = 301
mtimeFieldNumber = 302
ctimeFieldNumber = 303
// Bytes decoding sets aside for each file entry, hash, timestamp and // minHashSize is the encoded size of the smallest hash: a two-byte
// MIME type, however short its encoding. checkDecodedSize refuses an // multihash (algorithm code, zero digest length) after its tag and length.
// inner message for which these add up to more than maxDecodedGrowth minHashSize = 2 + 2
// times its size.
decodedFileEntrySize = 160
decodedHashSize = 112
decodedTimestampSize = 64
decodedMIMETypeSize = 16
// The densest manifests mfer writes add up to about 7 times their size. // minFileEntrySize is the encoded size of the smallest file entry: a
maxDecodedGrowth = 8 // one-byte path and one hash, each after its tag and length.
minFileEntrySize = 2 + 1 + 2 + minHashSize
) )
+21 -49
View File
@@ -28,8 +28,8 @@ var (
errUUIDMismatch = errors.New("outer and inner UUID mismatch") errUUIDMismatch = errors.New("outer and inner UUID mismatch")
errInvalidFileFormat = errors.New("invalid file format") errInvalidFileFormat = errors.New("invalid file format")
errInvalidManifestPath = errors.New("manifest contains invalid path") errInvalidManifestPath = errors.New("manifest contains invalid path")
errDecodedTooLarge = errors.New( errEntryTooShort = errors.New(
"manifest would take too much memory to decode") "manifest contains a file entry or hash shorter than the format allows")
) )
// validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable). // validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable).
@@ -157,46 +157,19 @@ func (m *manifest) decompressInner() ([]byte, error) {
return dat, nil return dat, nil
} }
// checkDecodedSize refuses an encoded inner message whose file entries, // checkEntrySizes rejects an encoded inner message holding a file entry or
// hashes, timestamps and MIME types would take more than maxDecodedGrowth // a hash shorter than the format allows. Decoding allocates a fixed amount
// times its size to decode. Decoding sets aside a fixed amount for each, // for each entry and each hash, however short, so a payload of empty ones
// however short its encoding, so a message of empty ones would take about // would decode to about 50 times its size.
// 50 times its size. func checkEntrySizes(inner []byte) error {
func checkDecodedSize(inner []byte) error { return forEachBytesField(inner, filesFieldNumber, func(entry []byte) error {
limit := maxDecodedGrowth * int64(len(inner)) if len(entry) < minFileEntrySize {
return errEntryTooShort
var decoded int64
add := func(size int64) error {
decoded += size
if decoded > limit {
return errDecodedTooLarge
} }
return nil return forEachBytesField(entry, hashesFieldNumber, func(hash []byte) error {
} if len(hash) < minHashSize {
return errEntryTooShort
return forEachBytesField(inner, func(num protowire.Number, entry []byte) error {
if num != filesFieldNumber {
return nil
}
err := add(decodedFileEntrySize)
if err != nil {
return err
}
return forEachBytesField(entry, func(num protowire.Number, _ []byte) error {
if num == hashesFieldNumber {
return add(decodedHashSize)
}
if num == mtimeFieldNumber || num == ctimeFieldNumber {
return add(decodedTimestampSize)
}
if num == mimeTypeFieldNumber {
return add(decodedMIMETypeSize)
} }
return nil return nil
@@ -204,27 +177,26 @@ func checkDecodedSize(inner []byte) error {
}) })
} }
// forEachBytesField calls fn with the number and value of each // forEachBytesField calls fn with the value of each length-delimited field
// length-delimited field in the encoded message msg, and fails if msg is // numbered num in the encoded message msg, and fails if msg is malformed.
// malformed.
func forEachBytesField( func forEachBytesField(
msg []byte, fn func(num protowire.Number, value []byte) error, msg []byte, num protowire.Number, fn func(value []byte) error,
) error { ) error {
for len(msg) > 0 { for len(msg) > 0 {
num, wireType, tagLen := protowire.ConsumeTag(msg) fieldNum, wireType, tagLen := protowire.ConsumeTag(msg)
if tagLen < 0 { if tagLen < 0 {
return protowire.ParseError(tagLen) return protowire.ParseError(tagLen)
} }
valueLen := protowire.ConsumeFieldValue(num, wireType, msg[tagLen:]) valueLen := protowire.ConsumeFieldValue(fieldNum, wireType, msg[tagLen:])
if valueLen < 0 { if valueLen < 0 {
return protowire.ParseError(valueLen) return protowire.ParseError(valueLen)
} }
if wireType == protowire.BytesType { if fieldNum == num && wireType == protowire.BytesType {
value, _ := protowire.ConsumeBytes(msg[tagLen:]) value, _ := protowire.ConsumeBytes(msg[tagLen:])
err := fn(num, value) err := fn(value)
if err != nil { if err != nil {
return err return err
} }
@@ -259,7 +231,7 @@ func (m *manifest) deserializeInner() error {
return bork.ErrFileTruncated return bork.ErrFileTruncated
} }
err = checkDecodedSize(dat) err = checkEntrySizes(dat)
if err != nil { if err != nil {
return fmt.Errorf("deserialize: unmarshal inner: %w", err) return fmt.Errorf("deserialize: unmarshal inner: %w", err)
} }
+9 -12
View File
@@ -54,14 +54,11 @@ func FuzzNewManifestFromReader(f *testing.F) {
} }
// It also keeps a few copies of its input. Buffers grow by // It also keeps a few copies of its input. Buffers grow by
// copying, so reaching those sizes allocates up to about six times // copying, so reaching those sizes allocates a few times them in
// them in total. Decoding the decompressed data takes up to // total. Decoding the decompressed data takes up to about 25 times
// maxDecodedGrowth times its size for file entries, hashes, // its size, when every file entry and hash is as short as the
// timestamps and MIME types, and up to about five times more for // parser accepts. Thirty-six times the input and the decompressed
// the bytes it copies out of it, such as fields it does not know, // data leaves room for both.
// which it keeps in buffers that also grow by copying. Twenty
// times the input and the decompressed data leaves room for all of
// that.
// //
// The decoder also sets aside a new buffer of one to two times the // The decoder also sets aside a new buffer of one to two times the
// window for each frame that asks for a larger window than the // window for each frame that asks for a larger window than the
@@ -76,10 +73,10 @@ func FuzzNewManifestFromReader(f *testing.F) {
// fails if the decoder accepts windows of twice zstdWindowSize; the // fails if the decoder accepts windows of twice zstdWindowSize; the
// seed whose two frames together exceed MaxDecompressedSize fails // seed whose two frames together exceed MaxDecompressedSize fails
// if the decoder decodes them in full instead of stopping at the // if the decoder decodes them in full instead of stopping at the
// declared size; the seeds of empty file entries, of a file entry // declared size; the seeds of empty file entries and of a file
// of empty hashes, and of file entries of only an empty MIME type // entry of empty hashes fail if the parser decodes entries or
// and empty times fail if the parser decodes them. // hashes shorter than the format allows.
limit := 20*(uint64(len(data))+decompressed) + 24*zstdWindowSize limit := 36*(uint64(len(data))+decompressed) + 24*zstdWindowSize
allocated := after.TotalAlloc - before.TotalAlloc allocated := after.TotalAlloc - before.TotalAlloc
if allocated > limit { if allocated > limit {
+24 -51
View File
@@ -6,9 +6,7 @@ import (
"context" "context"
"crypto/sha256" "crypto/sha256"
"fmt" "fmt"
"strconv"
"testing" "testing"
"time"
"github.com/google/uuid" "github.com/google/uuid"
"github.com/klauspost/compress/zstd" "github.com/klauspost/compress/zstd"
@@ -19,12 +17,18 @@ import (
) )
// craftInnerBytes builds the wire bytes of an inner MFFile holding a single // craftInnerBytes builds the wire bytes of an inner MFFile holding a single
// file entry whose path is exactly pathBytes. It writes the wire form by hand // file entry whose path is exactly pathBytes and whose one hash is multihash.
// so a hostile path — including one that is not valid UTF-8 — can be embedded // It writes the wire form by hand so a hostile path — including one that is
// without proto.Marshal's own UTF-8 enforcement rejecting it first. // not valid UTF-8 — can be embedded without proto.Marshal's own UTF-8
func craftInnerBytes(id uuid.UUID, pathBytes string) []byte { // enforcement rejecting it first.
func craftInnerBytes(id uuid.UUID, pathBytes string, multihash []byte) []byte {
hash := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFileChecksum.multiHash
hash = protowire.AppendBytes(hash, multihash)
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
entry = protowire.AppendString(entry, pathBytes) entry = protowire.AppendString(entry, pathBytes)
entry = protowire.AppendTag(entry, 3, protowire.BytesType) // MFFilePath.hashes
entry = protowire.AppendBytes(entry, hash)
inner := protowire.AppendTag(nil, 100, protowire.VarintType) // MFFile.version inner := protowire.AppendTag(nil, 100, protowire.VarintType) // MFFile.version
inner = protowire.AppendVarint(inner, uint64(MFFile_VERSION_ONE)) inner = protowire.AppendVarint(inner, uint64(MFFile_VERSION_ONE))
@@ -91,7 +95,8 @@ func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
t.Parallel() t.Parallel()
id := uuid.New() id := uuid.New()
data := wrapInner(t, id, craftInnerBytes(id, tt.path)) hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256
data := wrapInner(t, id, craftInnerBytes(id, tt.path, hash))
_, err := NewManifestFromReader(bytes.NewReader(data)) _, err := NewManifestFromReader(bytes.NewReader(data))
require.Error(t, err) require.Error(t, err)
@@ -116,55 +121,23 @@ func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
} }
} }
// Entries of a one-character path and empty modification and change times // A one-byte path and a multihash of algorithm code and zero digest length
// pass every other check, but would take about 23 times their size to decode. // make the smallest file entry and hash the format allows: they load, and an
func TestDeserializeRefusesEntriesThatDecodeTooLarge(t *testing.T) { // entry or hash one byte shorter is refused before decoding.
func TestDeserializeRejectsEntriesShorterThanFormatAllows(t *testing.T) {
t.Parallel() t.Parallel()
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path load := func(path string, multihash []byte) error {
entry = protowire.AppendString(entry, "a") id := uuid.New()
entry = protowire.AppendTag(entry, 302, protowire.BytesType) // MFFilePath.mtime data := wrapInner(t, id, craftInnerBytes(id, path, multihash))
entry = protowire.AppendBytes(entry, nil) _, err := NewManifestFromReader(bytes.NewReader(data))
entry = protowire.AppendTag(entry, 303, protowire.BytesType) // MFFilePath.ctime
entry = protowire.AppendBytes(entry, nil)
id := uuid.New() return err
inner := protowire.AppendTag(nil, 102, protowire.BytesType) // MFFile.uuid
inner = protowire.AppendBytes(inner, id[:])
for range 1000 {
inner = protowire.AppendTag(inner, 101, protowire.BytesType) // MFFile.files
inner = protowire.AppendBytes(inner, entry)
} }
_, err := NewManifestFromReader(bytes.NewReader(wrapInner(t, id, inner))) require.NoError(t, load("a", []byte{0, 0}))
require.ErrorIs(t, err, errDecodedTooLarge) require.ErrorIs(t, load("", []byte{0, 0}), errEntryTooShort)
} require.ErrorIs(t, load("ab", []byte{0}), errEntryTooShort)
// Many empty files with names of at most three characters and modification
// times at the epoch make about the densest manifest mfer writes: it takes
// about 7 times its size to decode, and still loads. A signature would not
// change the inner message, so none is added.
func TestDeserializeLoadsDensestManifest(t *testing.T) {
t.Parallel()
hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256
b := NewBuilder()
b.SetIncludeTimestamps(true)
const files = 10000
for i := range files {
name := RelFilePath(strconv.FormatInt(int64(i), 36))
require.NoError(t, b.AddFileWithHash(name, 0, ModTime(time.Unix(0, 0)), hash))
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.Len(t, m.Files(), files)
} }
func TestDeserializeValidManifestRoundTrips(t *testing.T) { func TestDeserializeValidManifestRoundTrips(t *testing.T) {
File diff suppressed because one or more lines are too long