Compare commits
1
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
09802fde10 |
@@ -1,12 +1,12 @@
|
||||
# mfer
|
||||
|
||||
[mfer](https://git.eeqj.de/sneak/mfer) is a [WTFPL](https://wtfpl.net)-licensed
|
||||
(public domain) [Go](https://golang.org) library and command-line tool by
|
||||
[@sneak](https://sneak.berlin) that specifies and generates `.mf` manifest files
|
||||
over a directory tree to encapsulate metadata about the files — such as
|
||||
cryptographic checksums and signatures over same — to aid in archiving,
|
||||
downloading, streaming, and mirroring. It was first published in 2022. The
|
||||
manifest files' data is serialized with Google's
|
||||
[mfer](https://git.eeqj.de/sneak/mfer) is a reference implementation library and
|
||||
thin wrapper command-line utility written in [Go](https://golang.org) and first
|
||||
published in 2022 under the [WTFPL](https://wtfpl.net) (public domain) license.
|
||||
It specifies and generates `.mf` manifest files over a directory tree of files
|
||||
to encapsulate metadata about them (such as cryptographic checksums or
|
||||
signatures over same) to aid in archiving, downloading, and streaming, or
|
||||
mirroring. The manifest files' data is serialized with Google's
|
||||
[protobuf serialization format](https://developers.google.com/protocol-buffers).
|
||||
The structure of these files can be found
|
||||
[in the format specification](https://git.eeqj.de/sneak/mfer/src/branch/main/mfer/mf.proto)
|
||||
@@ -21,36 +21,6 @@ This project was started by [@sneak](https://sneak.berlin) to scratch an itch in
|
||||
as a de-facto standard and be incorporated into other software. A compatible
|
||||
javascript library is planned.
|
||||
|
||||
# Getting Started
|
||||
|
||||
`mfer` builds from source with a Go 1.23+ toolchain. The generated protobuf code
|
||||
is committed, so no `protoc` toolchain is required:
|
||||
|
||||
```sh
|
||||
git clone https://git.eeqj.de/sneak/mfer.git
|
||||
cd mfer
|
||||
go build -o bin/mfer ./cmd/mfer
|
||||
```
|
||||
|
||||
Generate a manifest for a directory tree, verify it later, and fetch a published
|
||||
tree by URL:
|
||||
|
||||
```sh
|
||||
# Write .index.mf, a manifest of the files under the current directory.
|
||||
bin/mfer gen .
|
||||
|
||||
# Verify the files on disk against the manifest. Exits nonzero if any file
|
||||
# is missing or corrupted.
|
||||
bin/mfer check .index.mf
|
||||
|
||||
# Download and cryptographically verify a tree published over HTTP: mfer
|
||||
# fetches <url>/index.mf, then downloads every file it lists.
|
||||
bin/mfer fetch https://example.com/tree/
|
||||
```
|
||||
|
||||
Run `bin/mfer help` for the full command list, or `bin/mfer <command> --help`
|
||||
for a single command's options.
|
||||
|
||||
# Build Status
|
||||
|
||||
CI runs `script/cibuild`, which builds the Docker image with `--no-cache`, so
|
||||
@@ -115,9 +85,7 @@ Any changes submitted to this project must also be
|
||||
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
|
||||
tooling requirements, and workflow conventions.
|
||||
|
||||
# Rationale
|
||||
|
||||
## The problem
|
||||
# Problem Statement
|
||||
|
||||
Given a plain URL, there is no standard way to safely and programmatically
|
||||
download everything "under" that URL path. `wget -r` can traverse directory
|
||||
@@ -141,7 +109,7 @@ Real issues I face:
|
||||
- when I download a large file via HTTP, I have no way of knowing if the file
|
||||
content is what it's supposed to be
|
||||
|
||||
## The solution
|
||||
# Proposed Solution
|
||||
|
||||
A standard, a manifest file format, and a tool for generating same.
|
||||
|
||||
@@ -173,28 +141,6 @@ The manifest file would do several important things:
|
||||
- maybe a bittorrent chunklist for torrent client compatibility? perhaps a
|
||||
top-level infohash for the whole manifest?
|
||||
|
||||
# Design
|
||||
|
||||
The repository is split into a reusable library and a thin command-line wrapper
|
||||
around it.
|
||||
|
||||
- `mfer/` is the reusable library and the heart of the project: it defines the
|
||||
manifest format and implements building, scanning, checking, serialization,
|
||||
and signing. The protobuf schema is `mfer/mf.proto`, and the generated code it
|
||||
produces (`mfer/mf.pb.go`) is committed alongside it so the library builds
|
||||
with `go get` and needs no `protoc` toolchain.
|
||||
- `internal/cli/` holds the command implementations — `generate` (alias `gen`),
|
||||
`check`, `freshen`, `export`, `list` (alias `ls`), `fetch`, and `version` —
|
||||
that wire the library to the command-line interface.
|
||||
- `internal/log/` provides the logging used by the commands and the library.
|
||||
- `internal/bork/` provides the error the library returns when a manifest's
|
||||
decompressed contents are not the size the manifest records.
|
||||
- `cmd/mfer/` is the entrypoint: its `main` package runs `internal/cli` and
|
||||
exits with the status it returns.
|
||||
|
||||
Everything under `internal/` is private to this repository; only the `mfer/`
|
||||
package is intended for import by other software.
|
||||
|
||||
# Design Goals
|
||||
|
||||
- Replace SHASUMS/SHASUMS.asc files
|
||||
@@ -328,9 +274,9 @@ Open work, open design questions included, is tracked in this repo's issues:
|
||||
- Issues:
|
||||
[https://git.eeqj.de/sneak/mfer/issues](https://git.eeqj.de/sneak/mfer/issues)
|
||||
|
||||
# Author
|
||||
# Authors
|
||||
|
||||
- [@sneak](https://sneak.berlin)
|
||||
- [@sneak <sneak@sneak.berlin>](mailto:sneak@sneak.berlin)
|
||||
|
||||
# License
|
||||
|
||||
|
||||
+8
-15
@@ -18,23 +18,16 @@ const (
|
||||
// uuidLength is the length in bytes of a binary UUID.
|
||||
uuidLength = 16
|
||||
|
||||
// Numbers in mf.proto of MFFile.files and of the MFFilePath fields
|
||||
// that decoding sets aside a fixed amount of memory for.
|
||||
// filesFieldNumber and hashesFieldNumber are the numbers of
|
||||
// MFFile.files and MFFilePath.hashes in mf.proto.
|
||||
filesFieldNumber = 101
|
||||
hashesFieldNumber = 3
|
||||
mimeTypeFieldNumber = 301
|
||||
mtimeFieldNumber = 302
|
||||
ctimeFieldNumber = 303
|
||||
|
||||
// Bytes decoding sets aside for each file entry, hash, timestamp and
|
||||
// MIME type, however short its encoding. checkDecodedSize refuses an
|
||||
// inner message for which these add up to more than maxDecodedGrowth
|
||||
// times its size.
|
||||
decodedFileEntrySize = 160
|
||||
decodedHashSize = 112
|
||||
decodedTimestampSize = 64
|
||||
decodedMIMETypeSize = 16
|
||||
// minHashSize is the encoded size of the smallest hash: a two-byte
|
||||
// multihash (algorithm code, zero digest length) after its tag and length.
|
||||
minHashSize = 2 + 2
|
||||
|
||||
// The densest manifests mfer writes add up to about 7 times their size.
|
||||
maxDecodedGrowth = 8
|
||||
// minFileEntrySize is the encoded size of the smallest file entry: a
|
||||
// one-byte path and one hash, each after its tag and length.
|
||||
minFileEntrySize = 2 + 1 + 2 + minHashSize
|
||||
)
|
||||
|
||||
+21
-49
@@ -28,8 +28,8 @@ var (
|
||||
errUUIDMismatch = errors.New("outer and inner UUID mismatch")
|
||||
errInvalidFileFormat = errors.New("invalid file format")
|
||||
errInvalidManifestPath = errors.New("manifest contains invalid path")
|
||||
errDecodedTooLarge = errors.New(
|
||||
"manifest would take too much memory to decode")
|
||||
errEntryTooShort = errors.New(
|
||||
"manifest contains a file entry or hash shorter than the format allows")
|
||||
)
|
||||
|
||||
// validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable).
|
||||
@@ -157,46 +157,19 @@ func (m *manifest) decompressInner() ([]byte, error) {
|
||||
return dat, nil
|
||||
}
|
||||
|
||||
// checkDecodedSize refuses an encoded inner message whose file entries,
|
||||
// hashes, timestamps and MIME types would take more than maxDecodedGrowth
|
||||
// times its size to decode. Decoding sets aside a fixed amount for each,
|
||||
// however short its encoding, so a message of empty ones would take about
|
||||
// 50 times its size.
|
||||
func checkDecodedSize(inner []byte) error {
|
||||
limit := maxDecodedGrowth * int64(len(inner))
|
||||
|
||||
var decoded int64
|
||||
|
||||
add := func(size int64) error {
|
||||
decoded += size
|
||||
if decoded > limit {
|
||||
return errDecodedTooLarge
|
||||
// checkEntrySizes rejects an encoded inner message holding a file entry or
|
||||
// a hash shorter than the format allows. Decoding allocates a fixed amount
|
||||
// for each entry and each hash, however short, so a payload of empty ones
|
||||
// would decode to about 50 times its size.
|
||||
func checkEntrySizes(inner []byte) error {
|
||||
return forEachBytesField(inner, filesFieldNumber, func(entry []byte) error {
|
||||
if len(entry) < minFileEntrySize {
|
||||
return errEntryTooShort
|
||||
}
|
||||
|
||||
return nil
|
||||
}
|
||||
|
||||
return forEachBytesField(inner, func(num protowire.Number, entry []byte) error {
|
||||
if num != filesFieldNumber {
|
||||
return nil
|
||||
}
|
||||
|
||||
err := add(decodedFileEntrySize)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
|
||||
return forEachBytesField(entry, func(num protowire.Number, _ []byte) error {
|
||||
if num == hashesFieldNumber {
|
||||
return add(decodedHashSize)
|
||||
}
|
||||
|
||||
if num == mtimeFieldNumber || num == ctimeFieldNumber {
|
||||
return add(decodedTimestampSize)
|
||||
}
|
||||
|
||||
if num == mimeTypeFieldNumber {
|
||||
return add(decodedMIMETypeSize)
|
||||
return forEachBytesField(entry, hashesFieldNumber, func(hash []byte) error {
|
||||
if len(hash) < minHashSize {
|
||||
return errEntryTooShort
|
||||
}
|
||||
|
||||
return nil
|
||||
@@ -204,27 +177,26 @@ func checkDecodedSize(inner []byte) error {
|
||||
})
|
||||
}
|
||||
|
||||
// forEachBytesField calls fn with the number and value of each
|
||||
// length-delimited field in the encoded message msg, and fails if msg is
|
||||
// malformed.
|
||||
// forEachBytesField calls fn with the value of each length-delimited field
|
||||
// numbered num in the encoded message msg, and fails if msg is malformed.
|
||||
func forEachBytesField(
|
||||
msg []byte, fn func(num protowire.Number, value []byte) error,
|
||||
msg []byte, num protowire.Number, fn func(value []byte) error,
|
||||
) error {
|
||||
for len(msg) > 0 {
|
||||
num, wireType, tagLen := protowire.ConsumeTag(msg)
|
||||
fieldNum, wireType, tagLen := protowire.ConsumeTag(msg)
|
||||
if tagLen < 0 {
|
||||
return protowire.ParseError(tagLen)
|
||||
}
|
||||
|
||||
valueLen := protowire.ConsumeFieldValue(num, wireType, msg[tagLen:])
|
||||
valueLen := protowire.ConsumeFieldValue(fieldNum, wireType, msg[tagLen:])
|
||||
if valueLen < 0 {
|
||||
return protowire.ParseError(valueLen)
|
||||
}
|
||||
|
||||
if wireType == protowire.BytesType {
|
||||
if fieldNum == num && wireType == protowire.BytesType {
|
||||
value, _ := protowire.ConsumeBytes(msg[tagLen:])
|
||||
|
||||
err := fn(num, value)
|
||||
err := fn(value)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
@@ -259,7 +231,7 @@ func (m *manifest) deserializeInner() error {
|
||||
return bork.ErrFileTruncated
|
||||
}
|
||||
|
||||
err = checkDecodedSize(dat)
|
||||
err = checkEntrySizes(dat)
|
||||
if err != nil {
|
||||
return fmt.Errorf("deserialize: unmarshal inner: %w", err)
|
||||
}
|
||||
|
||||
@@ -54,14 +54,11 @@ func FuzzNewManifestFromReader(f *testing.F) {
|
||||
}
|
||||
|
||||
// It also keeps a few copies of its input. Buffers grow by
|
||||
// copying, so reaching those sizes allocates up to about six times
|
||||
// them in total. Decoding the decompressed data takes up to
|
||||
// maxDecodedGrowth times its size for file entries, hashes,
|
||||
// timestamps and MIME types, and up to about five times more for
|
||||
// the bytes it copies out of it, such as fields it does not know,
|
||||
// which it keeps in buffers that also grow by copying. Twenty
|
||||
// times the input and the decompressed data leaves room for all of
|
||||
// that.
|
||||
// copying, so reaching those sizes allocates a few times them in
|
||||
// total. Decoding the decompressed data takes up to about 25 times
|
||||
// its size, when every file entry and hash is as short as the
|
||||
// parser accepts. Thirty-six times the input and the decompressed
|
||||
// data leaves room for both.
|
||||
//
|
||||
// The decoder also sets aside a new buffer of one to two times the
|
||||
// window for each frame that asks for a larger window than the
|
||||
@@ -76,10 +73,10 @@ func FuzzNewManifestFromReader(f *testing.F) {
|
||||
// fails if the decoder accepts windows of twice zstdWindowSize; the
|
||||
// seed whose two frames together exceed MaxDecompressedSize fails
|
||||
// if the decoder decodes them in full instead of stopping at the
|
||||
// declared size; the seeds of empty file entries, of a file entry
|
||||
// of empty hashes, and of file entries of only an empty MIME type
|
||||
// and empty times fail if the parser decodes them.
|
||||
limit := 20*(uint64(len(data))+decompressed) + 24*zstdWindowSize
|
||||
// declared size; the seeds of empty file entries and of a file
|
||||
// entry of empty hashes fail if the parser decodes entries or
|
||||
// hashes shorter than the format allows.
|
||||
limit := 36*(uint64(len(data))+decompressed) + 24*zstdWindowSize
|
||||
|
||||
allocated := after.TotalAlloc - before.TotalAlloc
|
||||
if allocated > limit {
|
||||
|
||||
@@ -6,9 +6,7 @@ import (
|
||||
"context"
|
||||
"crypto/sha256"
|
||||
"fmt"
|
||||
"strconv"
|
||||
"testing"
|
||||
"time"
|
||||
|
||||
"github.com/google/uuid"
|
||||
"github.com/klauspost/compress/zstd"
|
||||
@@ -19,12 +17,18 @@ import (
|
||||
)
|
||||
|
||||
// craftInnerBytes builds the wire bytes of an inner MFFile holding a single
|
||||
// file entry whose path is exactly pathBytes. It writes the wire form by hand
|
||||
// so a hostile path — including one that is not valid UTF-8 — can be embedded
|
||||
// without proto.Marshal's own UTF-8 enforcement rejecting it first.
|
||||
func craftInnerBytes(id uuid.UUID, pathBytes string) []byte {
|
||||
// file entry whose path is exactly pathBytes and whose one hash is multihash.
|
||||
// It writes the wire form by hand so a hostile path — including one that is
|
||||
// not valid UTF-8 — can be embedded without proto.Marshal's own UTF-8
|
||||
// enforcement rejecting it first.
|
||||
func craftInnerBytes(id uuid.UUID, pathBytes string, multihash []byte) []byte {
|
||||
hash := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFileChecksum.multiHash
|
||||
hash = protowire.AppendBytes(hash, multihash)
|
||||
|
||||
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
|
||||
entry = protowire.AppendString(entry, pathBytes)
|
||||
entry = protowire.AppendTag(entry, 3, protowire.BytesType) // MFFilePath.hashes
|
||||
entry = protowire.AppendBytes(entry, hash)
|
||||
|
||||
inner := protowire.AppendTag(nil, 100, protowire.VarintType) // MFFile.version
|
||||
inner = protowire.AppendVarint(inner, uint64(MFFile_VERSION_ONE))
|
||||
@@ -91,7 +95,8 @@ func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
id := uuid.New()
|
||||
data := wrapInner(t, id, craftInnerBytes(id, tt.path))
|
||||
hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256
|
||||
data := wrapInner(t, id, craftInnerBytes(id, tt.path, hash))
|
||||
|
||||
_, err := NewManifestFromReader(bytes.NewReader(data))
|
||||
require.Error(t, err)
|
||||
@@ -116,55 +121,23 @@ func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
|
||||
}
|
||||
}
|
||||
|
||||
// Entries of a one-character path and empty modification and change times
|
||||
// pass every other check, but would take about 23 times their size to decode.
|
||||
func TestDeserializeRefusesEntriesThatDecodeTooLarge(t *testing.T) {
|
||||
// A one-byte path and a multihash of algorithm code and zero digest length
|
||||
// make the smallest file entry and hash the format allows: they load, and an
|
||||
// entry or hash one byte shorter is refused before decoding.
|
||||
func TestDeserializeRejectsEntriesShorterThanFormatAllows(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
|
||||
entry = protowire.AppendString(entry, "a")
|
||||
entry = protowire.AppendTag(entry, 302, protowire.BytesType) // MFFilePath.mtime
|
||||
entry = protowire.AppendBytes(entry, nil)
|
||||
entry = protowire.AppendTag(entry, 303, protowire.BytesType) // MFFilePath.ctime
|
||||
entry = protowire.AppendBytes(entry, nil)
|
||||
|
||||
load := func(path string, multihash []byte) error {
|
||||
id := uuid.New()
|
||||
inner := protowire.AppendTag(nil, 102, protowire.BytesType) // MFFile.uuid
|
||||
inner = protowire.AppendBytes(inner, id[:])
|
||||
data := wrapInner(t, id, craftInnerBytes(id, path, multihash))
|
||||
_, err := NewManifestFromReader(bytes.NewReader(data))
|
||||
|
||||
for range 1000 {
|
||||
inner = protowire.AppendTag(inner, 101, protowire.BytesType) // MFFile.files
|
||||
inner = protowire.AppendBytes(inner, entry)
|
||||
return err
|
||||
}
|
||||
|
||||
_, err := NewManifestFromReader(bytes.NewReader(wrapInner(t, id, inner)))
|
||||
require.ErrorIs(t, err, errDecodedTooLarge)
|
||||
}
|
||||
|
||||
// Many empty files with names of at most three characters and modification
|
||||
// times at the epoch make about the densest manifest mfer writes: it takes
|
||||
// about 7 times its size to decode, and still loads. A signature would not
|
||||
// change the inner message, so none is added.
|
||||
func TestDeserializeLoadsDensestManifest(t *testing.T) {
|
||||
t.Parallel()
|
||||
|
||||
hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256
|
||||
|
||||
b := NewBuilder()
|
||||
b.SetIncludeTimestamps(true)
|
||||
|
||||
const files = 10000
|
||||
for i := range files {
|
||||
name := RelFilePath(strconv.FormatInt(int64(i), 36))
|
||||
require.NoError(t, b.AddFileWithHash(name, 0, ModTime(time.Unix(0, 0)), hash))
|
||||
}
|
||||
|
||||
var buf bytes.Buffer
|
||||
require.NoError(t, b.Build(context.Background(), &buf))
|
||||
|
||||
m, err := NewManifestFromReader(&buf)
|
||||
require.NoError(t, err)
|
||||
assert.Len(t, m.Files(), files)
|
||||
require.NoError(t, load("a", []byte{0, 0}))
|
||||
require.ErrorIs(t, load("", []byte{0, 0}), errEntryTooShort)
|
||||
require.ErrorIs(t, load("ab", []byte{0}), errEntryTooShort)
|
||||
}
|
||||
|
||||
func TestDeserializeValidManifestRoundTrips(t *testing.T) {
|
||||
|
||||
-2
File diff suppressed because one or more lines are too long
Reference in New Issue
Block a user