Author SHA1 Message Date
sneak e7331e8d11 Reject manifests whose file entries decode far larger than their bytes (closes #123)
check / check (push) Waiting to run
Parser fix: before decoding the manifest, the parser walks its file
entries and adds up what decoding sets aside for each entry, hash,
timestamp and MIME type, however short its encoding. It refuses the
manifest once that sum passes 8 times the decompressed size; the
densest manifests mfer writes come to about 7 times. Empty entries
decoded to about 50 times their size, so a 1.6 KB manifest allocated
nearly 1 GB. The fuzz target's ceiling falls to 20 times the input and
decompressed data, and a new seed of entries holding only an empty MIME
type and empty times fails it without the fix.

Model: opus-5-5
2026-10-04 05:56:21 +00:00
16 changed files with 264 additions and 81 deletions
-1
View File
@@ -16,7 +16,6 @@ linters:
- depguard # Dependency allow/block lists - depguard # Dependency allow/block lists
- godot # Requires comments to end with periods - godot # Requires comments to end with periods
- wsl # Deprecated, replaced by wsl_v5 - wsl # Deprecated, replaced by wsl_v5
- gomodguard # Deprecated, replaced by gomodguard_v2
- wrapcheck # Too verbose for internal packages - wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go - varnamelen # Short names like db, id are idiomatic Go
settings: settings:
+2 -7
View File
@@ -14,9 +14,7 @@ RUN touch mfer/mf.pb.go
# Go half of fmt-check only: this image has no node, so no prettier. The # Go half of fmt-check only: this image has no node, so no prettier. The
# markdown half runs in the mdfmt stage below. # markdown half runs in the mdfmt stage below.
RUN make fmt-check-go RUN make fmt-check-go
# The linter directly, not `make lint`: script/lint builds this stage, and RUN make lint
# there is no docker inside this build.
RUN golangci-lint run --config .golangci.yml ./...
# Markdown/JSON format stage — prettier needs node, which the Go images # Markdown/JSON format stage — prettier needs node, which the Go images
# do not have. node:22.17.0-bookworm-slim (2026-08-09); ships node # do not have. node:22.17.0-bookworm-slim (2026-08-09); ships node
@@ -69,10 +67,7 @@ RUN version="${VERSION:-$(git describe --tags --always)}"; \
exit 1; \ exit 1; \
fi; \ fi; \
cd cmd/mfer && \ cd cmd/mfer && \
CGO_ENABLED=0 go build -tags urfave_cli_no_docs -ldflags "-X main.Gitrev=$version" -o /mfer . go build -tags urfave_cli_no_docs -ldflags "-X main.Gitrev=$version" -o /mfer .
# Fail unless /mfer is statically linked: scratch has no C library to run it.
RUN ldd /mfer 2>&1 | grep -q 'not a dynamic executable'
FROM scratch FROM scratch
COPY --from=builder /mfer /mfer COPY --from=builder /mfer /mfer
+3
View File
@@ -58,6 +58,9 @@ fmt-check-md:
hooks: hooks:
@script/install-precommit @script/install-precommit
devprereqs:
which golangci-lint || go install -v github.com/golangci/golangci-lint/v2/cmd/golangci-lint@v2.12.2
mfer/mf.pb.go: mfer/mf.proto mfer/mf.pb.go: mfer/mf.proto
cd mfer && go generate . cd mfer && go generate .
+11 -12
View File
@@ -65,9 +65,9 @@ standard: normalized scripts in `script/` are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call them. We development workflow, and the Makefile targets are thin shims that call them. We
provide: provide:
- `script/bootstrap` — install all dependencies (Go, Go module download, and - `script/bootstrap` — install all dependencies (Go, golangci-lint, Go module
node/yarn plus the prettier version pinned in `package.json`/`yarn.lock`), download, and node/yarn plus the prettier version pinned in
idempotently; golangci-lint is not installed, it runs only in Docker `package.json`/`yarn.lock`), idempotently
- `script/setup` — make a fresh clone ready for development: runs - `script/setup` — make a fresh clone ready for development: runs
`script/bootstrap`, then `script/install-precommit` `script/bootstrap`, then `script/install-precommit`
- `script/projectname` — output the project name (`mfer`); used by other scripts - `script/projectname` — output the project name (`mfer`); used by other scripts
@@ -77,11 +77,9 @@ provide:
- `script/fuzz` — fuzz the manifest parser for one minute; run by hand - `script/fuzz` — fuzz the manifest parser for one minute; run by hand
(`make fuzz`), never by CI, while `script/test` runs its committed seed corpus (`make fuzz`), never by CI, while `script/test` runs its committed seed corpus
as ordinary tests as ordinary tests
- `script/lint` — run `golangci-lint` in Docker: builds only the `lint` stage of - `script/lint` — run `golangci-lint` and verify `gofmt` cleanliness
the `Dockerfile` (the Go format check, then the linter), uncached so it runs - `script/fmt` — format all code and docs (writes): `gofumpt`,
every time, then removes the image `golangci-lint run --fix`, and `script/prettier --write`
- `script/fmt` — format all code and docs (writes): `gofumpt` and
`script/prettier --write`
- `script/prettier` — run prettier over the repository's canonical file set - `script/prettier` — run prettier over the repository's canonical file set
(Markdown and JSON, minus `.prettierignore`) in the given mode, `--write` or (Markdown and JSON, minus `.prettierignore`) in the given mode, `--write` or
`--check`; the single definition of that file set, so `script/fmt` and `--check`; the single definition of that file set, so `script/fmt` and
@@ -107,10 +105,11 @@ yet. Primary development happens on a privately-run Gitea instance at
[tracked there](https://git.eeqj.de/sneak/mfer/issues). [tracked there](https://git.eeqj.de/sneak/mfer/issues).
Changes must always be formatted with a standard `go fmt`, syntactically valid, Changes must always be formatted with a standard `go fmt`, syntactically valid,
and must pass the linting defined in the repository's `.golangci.yml`, which and must pass the linting defined in the repository (presently only the
`make lint` runs in Docker. The `main` branch is protected and all changes must `golangci-lint` defaults), which can be run with a `make lint`. The `main`
be made via [pull requests](https://git.eeqj.de/sneak/mfer/pulls) and pass CI to branch is protected and all changes must be made via
be merged. Any changes submitted to this project must also be [pull requests](https://git.eeqj.de/sneak/mfer/pulls) and pass CI to be merged.
Any changes submitted to this project must also be
[WTFPL-licensed](https://wtfpl.net) to be considered. [WTFPL-licensed](https://wtfpl.net) to be considered.
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards, See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
+20
View File
@@ -17,4 +17,24 @@ const (
// uuidLength is the length in bytes of a binary UUID. // uuidLength is the length in bytes of a binary UUID.
uuidLength = 16 uuidLength = 16
// Numbers in mf.proto of MFFile.files and of the MFFilePath fields
// that decoding sets aside a fixed amount of memory for.
filesFieldNumber = 101
hashesFieldNumber = 3
mimeTypeFieldNumber = 301
mtimeFieldNumber = 302
ctimeFieldNumber = 303
// Bytes decoding sets aside for each file entry, hash, timestamp and
// MIME type, however short its encoding. checkDecodedSize refuses an
// inner message for which these add up to more than maxDecodedGrowth
// times its size.
decodedFileEntrySize = 160
decodedHashSize = 112
decodedTimestampSize = 64
decodedMIMETypeSize = 16
// The densest manifests mfer writes add up to about 7 times their size.
maxDecodedGrowth = 8
) )
+87
View File
@@ -11,6 +11,7 @@ import (
"github.com/google/uuid" "github.com/google/uuid"
"github.com/klauspost/compress/zstd" "github.com/klauspost/compress/zstd"
"github.com/spf13/afero" "github.com/spf13/afero"
"google.golang.org/protobuf/encoding/protowire"
"google.golang.org/protobuf/proto" "google.golang.org/protobuf/proto"
"sneak.berlin/go/mfer/internal/bork" "sneak.berlin/go/mfer/internal/bork"
"sneak.berlin/go/mfer/internal/log" "sneak.berlin/go/mfer/internal/log"
@@ -27,6 +28,8 @@ var (
errUUIDMismatch = errors.New("outer and inner UUID mismatch") errUUIDMismatch = errors.New("outer and inner UUID mismatch")
errInvalidFileFormat = errors.New("invalid file format") errInvalidFileFormat = errors.New("invalid file format")
errInvalidManifestPath = errors.New("manifest contains invalid path") errInvalidManifestPath = errors.New("manifest contains invalid path")
errDecodedTooLarge = errors.New(
"manifest would take too much memory to decode")
) )
// validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable). // validateUUID checks that the byte slice is a valid UUID (16 bytes, parseable).
@@ -154,6 +157,85 @@ func (m *manifest) decompressInner() ([]byte, error) {
return dat, nil return dat, nil
} }
// checkDecodedSize refuses an encoded inner message whose file entries,
// hashes, timestamps and MIME types would take more than maxDecodedGrowth
// times its size to decode. Decoding sets aside a fixed amount for each,
// however short its encoding, so a message of empty ones would take about
// 50 times its size.
func checkDecodedSize(inner []byte) error {
limit := maxDecodedGrowth * int64(len(inner))
var decoded int64
add := func(size int64) error {
decoded += size
if decoded > limit {
return errDecodedTooLarge
}
return nil
}
return forEachBytesField(inner, func(num protowire.Number, entry []byte) error {
if num != filesFieldNumber {
return nil
}
err := add(decodedFileEntrySize)
if err != nil {
return err
}
return forEachBytesField(entry, func(num protowire.Number, _ []byte) error {
if num == hashesFieldNumber {
return add(decodedHashSize)
}
if num == mtimeFieldNumber || num == ctimeFieldNumber {
return add(decodedTimestampSize)
}
if num == mimeTypeFieldNumber {
return add(decodedMIMETypeSize)
}
return nil
})
})
}
// forEachBytesField calls fn with the number and value of each
// length-delimited field in the encoded message msg, and fails if msg is
// malformed.
func forEachBytesField(
msg []byte, fn func(num protowire.Number, value []byte) error,
) error {
for len(msg) > 0 {
num, wireType, tagLen := protowire.ConsumeTag(msg)
if tagLen < 0 {
return protowire.ParseError(tagLen)
}
valueLen := protowire.ConsumeFieldValue(num, wireType, msg[tagLen:])
if valueLen < 0 {
return protowire.ParseError(valueLen)
}
if wireType == protowire.BytesType {
value, _ := protowire.ConsumeBytes(msg[tagLen:])
err := fn(num, value)
if err != nil {
return err
}
}
msg = msg[tagLen+valueLen:]
}
return nil
}
func (m *manifest) deserializeInner() error { func (m *manifest) deserializeInner() error {
err := m.validateOuterHeader() err := m.validateOuterHeader()
if err != nil { if err != nil {
@@ -177,6 +259,11 @@ func (m *manifest) deserializeInner() error {
return bork.ErrFileTruncated return bork.ErrFileTruncated
} }
err = checkDecodedSize(dat)
if err != nil {
return fmt.Errorf("deserialize: unmarshal inner: %w", err)
}
// Deserialize inner message // Deserialize inner message
m.pbInner = new(MFFile) m.pbInner = new(MFFile)
+12 -5
View File
@@ -54,9 +54,14 @@ func FuzzNewManifestFromReader(f *testing.F) {
} }
// It also keeps a few copies of its input. Buffers grow by // It also keeps a few copies of its input. Buffers grow by
// copying, so reaching those sizes allocates a few times them in // copying, so reaching those sizes allocates up to about six times
// total: sixteen times the input and the decompressed data leaves // them in total. Decoding the decompressed data takes up to
// room for that. // maxDecodedGrowth times its size for file entries, hashes,
// timestamps and MIME types, and up to about five times more for
// the bytes it copies out of it, such as fields it does not know,
// which it keeps in buffers that also grow by copying. Twenty
// times the input and the decompressed data leaves room for all of
// that.
// //
// The decoder also sets aside a new buffer of one to two times the // The decoder also sets aside a new buffer of one to two times the
// window for each frame that asks for a larger window than the // window for each frame that asks for a larger window than the
@@ -71,8 +76,10 @@ func FuzzNewManifestFromReader(f *testing.F) {
// fails if the decoder accepts windows of twice zstdWindowSize; the // fails if the decoder accepts windows of twice zstdWindowSize; the
// seed whose two frames together exceed MaxDecompressedSize fails // seed whose two frames together exceed MaxDecompressedSize fails
// if the decoder decodes them in full instead of stopping at the // if the decoder decodes them in full instead of stopping at the
// declared size. // declared size; the seeds of empty file entries, of a file entry
limit := 16*(uint64(len(data))+decompressed) + 24*zstdWindowSize // of empty hashes, and of file entries of only an empty MIME type
// and empty times fail if the parser decodes them.
limit := 20*(uint64(len(data))+decompressed) + 24*zstdWindowSize
allocated := after.TotalAlloc - before.TotalAlloc allocated := after.TotalAlloc - before.TotalAlloc
if allocated > limit { if allocated > limit {
+53
View File
@@ -6,7 +6,9 @@ import (
"context" "context"
"crypto/sha256" "crypto/sha256"
"fmt" "fmt"
"strconv"
"testing" "testing"
"time"
"github.com/google/uuid" "github.com/google/uuid"
"github.com/klauspost/compress/zstd" "github.com/klauspost/compress/zstd"
@@ -114,6 +116,57 @@ func TestDeserializeRejectsInvalidEntryPaths(t *testing.T) {
} }
} }
// Entries of a one-character path and empty modification and change times
// pass every other check, but would take about 23 times their size to decode.
func TestDeserializeRefusesEntriesThatDecodeTooLarge(t *testing.T) {
t.Parallel()
entry := protowire.AppendTag(nil, 1, protowire.BytesType) // MFFilePath.path
entry = protowire.AppendString(entry, "a")
entry = protowire.AppendTag(entry, 302, protowire.BytesType) // MFFilePath.mtime
entry = protowire.AppendBytes(entry, nil)
entry = protowire.AppendTag(entry, 303, protowire.BytesType) // MFFilePath.ctime
entry = protowire.AppendBytes(entry, nil)
id := uuid.New()
inner := protowire.AppendTag(nil, 102, protowire.BytesType) // MFFile.uuid
inner = protowire.AppendBytes(inner, id[:])
for range 1000 {
inner = protowire.AppendTag(inner, 101, protowire.BytesType) // MFFile.files
inner = protowire.AppendBytes(inner, entry)
}
_, err := NewManifestFromReader(bytes.NewReader(wrapInner(t, id, inner)))
require.ErrorIs(t, err, errDecodedTooLarge)
}
// Many empty files with names of at most three characters and modification
// times at the epoch make about the densest manifest mfer writes: it takes
// about 7 times its size to decode, and still loads. A signature would not
// change the inner message, so none is added.
func TestDeserializeLoadsDensestManifest(t *testing.T) {
t.Parallel()
hash := make([]byte, 34) // multihash: 2-byte prefix + 32-byte SHA-256
b := NewBuilder()
b.SetIncludeTimestamps(true)
const files = 10000
for i := range files {
name := RelFilePath(strconv.FormatInt(int64(i), 36))
require.NoError(t, b.AddFileWithHash(name, 0, ModTime(time.Unix(0, 0)), hash))
}
var buf bytes.Buffer
require.NoError(t, b.Build(context.Background(), &buf))
m, err := NewManifestFromReader(&buf)
require.NoError(t, err)
assert.Len(t, m.Files(), files)
}
func TestDeserializeValidManifestRoundTrips(t *testing.T) { func TestDeserializeValidManifestRoundTrips(t *testing.T) {
t.Parallel() t.Parallel()
+18 -34
View File
@@ -10,7 +10,6 @@ import (
"path/filepath" "path/filepath"
"strconv" "strconv"
"strings" "strings"
"syscall"
"testing" "testing"
"time" "time"
@@ -426,31 +425,18 @@ func TestGPGTimeoutKillsGPG(t *testing.T) {
assert.Contains(t, err.Error(), "gpg sign failed: gpg timed out") assert.Contains(t, err.Error(), "gpg sign failed: gpg timed out")
} }
// TestGPGCancelWhenChildHoldsOutput uses a fake gpg that runs sleep as a // TestGPGTimeoutWhenChildHoldsOutput uses a fake gpg that runs sleep as a
// child instead of exec-ing it, the way a wrapper script around the real // child instead of exec-ing it, the way a wrapper script around the real
// gpg might. Killing the fake gpg leaves sleep holding its stdout and // gpg might. Killing the fake gpg leaves sleep holding its stdout and
// stderr open; the call must still return once ctx ends instead of waiting // stderr open; the call must still return shortly after the deadline
// for sleep to exit. The fake gpg writes the process ID of sleep to a named // instead of waiting for sleep to exit. The fake gpg writes the process ID
// pipe; the test ends ctx only after reading it, so sleep is running by // of sleep to a file so that the test can kill it before returning.
// then, and kills sleep before returning. func TestGPGTimeoutWhenChildHoldsOutput(t *testing.T) {
func TestGPGCancelWhenChildHoldsOutput(t *testing.T) { pidFile := filepath.Join(t.TempDir(), "sleep.pid")
pidPipe := filepath.Join(t.TempDir(), "sleep.pid")
require.NoError(t, syscall.Mkfifo(pidPipe, 0o600))
// sleep outlasts the 10 s wait below, so a call that waits for it fails.
t.Setenv("PATH", fakeGPGPath(t, t.Setenv("PATH", fakeGPGPath(t,
"#!/bin/sh\nsleep 60 &\necho $! >'"+pidPipe+"'\nwait\n")) "#!/bin/sh\nsleep 3 &\necho $! >'"+pidFile+"'\nwait\n"))
t.Cleanup(func() {
ctx, cancel := context.WithCancel(context.Background()) pid, err := os.ReadFile(pidFile) //nolint:gosec // G304: path inside t.TempDir()
defer cancel()
signErr := make(chan error, 1)
go func() {
_, err := gpgSign(ctx, []byte("data"), GPGKeyID("any"))
signErr <- err
}()
pid, err := os.ReadFile(pidPipe) //nolint:gosec // G304: path inside t.TempDir()
require.NoError(t, err) require.NoError(t, err)
n, err := strconv.Atoi(strings.TrimSpace(string(pid))) n, err := strconv.Atoi(strings.TrimSpace(string(pid)))
@@ -458,19 +444,17 @@ func TestGPGCancelWhenChildHoldsOutput(t *testing.T) {
sleep, err := os.FindProcess(n) sleep, err := os.FindProcess(n)
require.NoError(t, err) require.NoError(t, err)
t.Cleanup(func() { require.NoError(t, sleep.Kill()) }) require.NoError(t, sleep.Kill())
})
cancel() ctx, cancel := context.WithTimeout(context.Background(), 100*time.Millisecond)
defer cancel()
// The call should return about gpgWaitDelay (one second) after the start := time.Now()
// cancel. 10 s is far above that and well under the 30 s test timeout, _, err := gpgSign(ctx, []byte("data"), GPGKeyID("any"))
// which would abort the whole package before the cleanup kills sleep. require.ErrorIs(t, err, context.DeadlineExceeded)
select { assert.Less(t, time.Since(start), 3*time.Second,
case err := <-signErr: "the call waited for the child holding gpg's output to exit")
require.ErrorIs(t, err, context.Canceled)
case <-time.After(10 * time.Second):
t.Fatal("the call waited for the child holding gpg's output to exit")
}
} }
// TestBuildPassesContextToSigning checks that a caller can cancel the gpg // TestBuildPassesContextToSigning checks that a caller can cancel the gpg
+29 -2
View File
@@ -304,19 +304,46 @@ func TestScannerEnumerateFS(t *testing.T) {
func TestSendEnumerateStatusNonBlocking(t *testing.T) { func TestSendEnumerateStatusNonBlocking(t *testing.T) {
t.Parallel() t.Parallel()
// Nobody receives, so a blocking send would hang the test into its timeout. // Channel with no buffer - send should not block
ch := make(chan EnumerateStatus) ch := make(chan EnumerateStatus)
// This should not block
done := make(chan bool)
go func() {
sendEnumerateStatus(ch, EnumerateStatus{FilesFound: 1}) sendEnumerateStatus(ch, EnumerateStatus{FilesFound: 1})
done <- true
}()
select {
case <-done:
// Success - did not block
case <-time.After(100 * time.Millisecond):
t.Fatal("sendEnumerateStatus blocked on full channel")
}
} }
func TestSendScanStatusNonBlocking(t *testing.T) { func TestSendScanStatusNonBlocking(t *testing.T) {
t.Parallel() t.Parallel()
// Nobody receives, so a blocking send would hang the test into its timeout. // Channel with no buffer - send should not block
ch := make(chan ScanStatus) ch := make(chan ScanStatus)
done := make(chan bool)
go func() {
sendScanStatus(ch, ScanStatus{ScannedFiles: 1}) sendScanStatus(ch, ScanStatus{ScannedFiles: 1})
done <- true
}()
select {
case <-done:
// Success - did not block
case <-time.After(100 * time.Millisecond):
t.Fatal("sendScanStatus blocked on full channel")
}
} }
func TestSendStatusNilChannel(t *testing.T) { func TestSendStatusNilChannel(t *testing.T) {
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
File diff suppressed because one or more lines are too long
+6 -1
View File
@@ -140,7 +140,12 @@ main() {
# ---- Go repos ---- # ---- Go repos ----
if missing go; then pkg_install go golang go go; fi if missing go; then pkg_install go golang go go; fi
# No golangci-lint: script/lint runs it in Docker only. # golangci-lint: packaged in nix, brew, and apk. On apt there is no
# package: download a specific release archive from GitHub and
# verify its hash (verify_sha256), never curl | sh.
if missing golangci-lint; then
pkg_install golangci-lint golangci-lint golangci-lint golangci-lint
fi
go mod download go mod download
# ---- Python repos ---- # ---- Python repos ----
+1
View File
@@ -19,6 +19,7 @@ main() {
cd "$ROOT" cd "$ROOT"
ensure_pb ensure_pb
gofumpt -l -w mfer internal cmd gofumpt -l -w mfer internal cmd
golangci-lint run --fix
# Markdown and JSON, over the same file set script/fmt-check verifies. # Markdown and JSON, over the same file set script/fmt-check verifies.
"$SCRIPT_DIR/prettier" --write "$SCRIPT_DIR/prettier" --write
} }
+8 -11
View File
@@ -1,20 +1,17 @@
#!/bin/sh #!/bin/sh
# script/lint: run golangci-lint, in Docker only. Builds the lint stage of # script/lint: run the linter.
# the Dockerfile, whose build runs the linter, so a successful build is a
# clean lint. --no-cache because a cached build runs no linter. The image
# is removed afterwards, whatever the outcome.
set -eu set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)" ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
# Tagged per run, so concurrent runs never remove each other's image. golangci-lint run
image="$("$SCRIPT_DIR/projectname")-lint:$$" if [ -n "$(gofmt -l .)" ]; then
# A failed build leaves no image, so there is nothing to remove then. echo "gofmt: files need formatting:" >&2
trap 'docker image rm "$image" >/dev/null 2>&1 || true' EXIT INT TERM gofmt -l . >&2
docker build --no-cache --target lint -t "$image" . exit 1
fi
} }
main "$@" main "$@"