All checks were successful
check / check (push) Successful in 2m52s
MaxBodySize logged r.URL.Path untruncated at WARN, and routes.go registers it ahead of RequireAuth, so an unauthenticated POST /source/<8 KB>/edit with an oversize declared Content-Length wrote attacker-chosen text of attacker-chosen length into the operator's log, for the cost of a request with no body. The 2,560-byte per-line budget #146 established did not reach it: that budget lives in the access-log field capping and this is a separate slog call. The capping mechanism moves out of internal/middleware into internal/logfield so there is one budget and one implementation rather than a second ad-hoc truncation. Truncate and EncodedBytes are unchanged; the access log now spends logfield.MaxBytes where it spent maxLogFieldBytes. The sweep the issue asked for found five more call sites of the same shape, all reachable unauthenticated, all now capped: the CSRF 403 (also registered ahead of RequireAuth), the rate limiters' 429 (the per-entrypoint receiver limiter is unauthenticated), RequireAuth's own DEBUG line, the unknown-entrypoint DEBUG line on the receiver, and the failed-login DEBUG lines. DEBUG being off by default is not a bound: an operator turning it on to diagnose a flood must not thereby hand the flood an unbounded write. Every other slog call in the tree was read and judged; the PR body lists all of them, including the ones left alone and why. MaxBodySize stays ahead of RequireAuth. An oversize body should be refused before the request buys a cookie decrypt and a session load, and rejecting first is what keeps an unauthenticated flood from choosing how much session work the process does. The ordering and what it costs are now written at the registration, on maxFormBodySize. MaxAccessLogLineBytes is restated as the ceiling on every slog line carrying a client-supplied value, not just the access log's: each of these lines carries strictly fewer client-supplied fields than the access log does, so none can be wider. That is asserted per line under both handlers rather than argued. Two writers are called out as NOT covered, so the figure is not read as more than it is: the log delivery target, which exists to emit the whole event and is deliberate, and GORM's default logger, which prints the interpolated SQL to stdout on a record-not-found and is unbounded on the receiver and login lookups. That second one is a real defect this audit turned up and is filed separately as #178, not fixed here. Tests drive 8 KB of client-chosen text at all six sites, through both handlers internal/logger can install and through each character they escape — including a bare C0 control, which costs six bytes on the line against the one it cost to send and is the case a raw-byte budget breaks on first. Each holds the encoded line to the ceiling, holds the whole flood's output to what that ceiling allows, and asserts the markers at the far end of the input are absent, so a value that merely happened to be short cannot pass. internal/logfield gains a test that measures the per-rune charge against what the handlers really emit over roughly 3,000 code points on each, so an undercharged rune fails a test instead of quietly falsifying the ceiling. Verified by mutation: reverting the MaxBodySize cap alone fails 12 subtests with a 16,583-byte line against the 2,560 ceiling; reverting the other five fails 70; budgeting raw bytes instead of encoded ones fails 23 across three packages.
144 lines
4.9 KiB
Go
144 lines
4.9 KiB
Go
// Package logfield bounds the client-supplied values this service
|
|
// writes into its logs.
|
|
//
|
|
// Any log field whose content a client picks is spent against a budget
|
|
// here, in ENCODED bytes rather than in the bytes the client sent, so
|
|
// that escaping cannot multiply a field past its nominal size. One
|
|
// budget and one implementation serves the access log in
|
|
// internal/middleware and every other slog call that reaches a
|
|
// client-chosen path, header or form value; a second, ad-hoc
|
|
// truncation somewhere else in the tree is the thing this package
|
|
// exists to prevent.
|
|
package logfield
|
|
|
|
import (
|
|
"strings"
|
|
"unicode"
|
|
"unicode/utf8"
|
|
)
|
|
|
|
const (
|
|
// MaxBytes is the default budget for a log field whose value the
|
|
// client supplies outright: a URL, a path, a header, a form value.
|
|
// The budget is spent in ENCODED bytes (see Truncate), so 512 still
|
|
// holds a real browser's User-Agent whole — those are plain ASCII,
|
|
// which encodes one byte for one — while a value built from
|
|
// characters the encoder escapes keeps a shorter prefix. That is
|
|
// the intended trade: 500 quotation marks are not a debugging
|
|
// asset.
|
|
MaxBytes = 512
|
|
|
|
// TruncationMarker is appended to any field that was cut, so a
|
|
// short value and a truncated one cannot be confused. It is charged
|
|
// on top of the budget, not inside it.
|
|
TruncationMarker = "[truncated]"
|
|
)
|
|
|
|
// EncodedBytes is what r costs on the line once the log handler has
|
|
// escaped it, taking the worse of the two handlers internal/logger
|
|
// configures.
|
|
//
|
|
// slog's JSON handler escapes quote, backslash, newline, carriage
|
|
// return and tab to two bytes each, and every other C0 control plus
|
|
// LINE SEPARATOR and PARAGRAPH SEPARATOR to a six-byte \u escape; it
|
|
// passes every other rune through as its own UTF-8. Its text handler
|
|
// quotes with strconv.Quote, which spells a non-printable rune below
|
|
// U+10000 as \uXXXX but one at or above U+10000 as \UXXXXXXXX — ten
|
|
// bytes, not six. The text handler is therefore the worse of the two
|
|
// for every non-printable rune, and by four bytes apiece for the
|
|
// 955,086 unassigned, private-use and format code points on planes 1
|
|
// to 16.
|
|
//
|
|
// Charging ten there is what makes the stated line ceilings hold for
|
|
// the tty handler as well: U+1000C encodes as F0 90 80 8C, every byte
|
|
// >= 0x80, which httpguts.ValidHeaderFieldValue accepts and
|
|
// net/textproto does not strip, so a header can be filled with them.
|
|
//
|
|
// Both handlers pass printable runes through as their own UTF-8, so
|
|
// unicode.IsPrint separates the escaped cases from the plain ones for
|
|
// either handler.
|
|
func EncodedBytes(r rune) int {
|
|
const (
|
|
// A backslash and the character itself.
|
|
shortEscapeBytes = 2
|
|
// \uXXXX, which is also the width of \u00XX.
|
|
escapedRuneBytes = 6
|
|
// \UXXXXXXXX, strconv.Quote's spelling of a non-printable
|
|
// rune outside the basic multilingual plane.
|
|
escapedAstralRuneBytes = 10
|
|
// The first code point strconv.Quote spells with \U.
|
|
firstAstralRune = 0x10000
|
|
)
|
|
|
|
switch {
|
|
case r == '"' || r == '\\' || r == '\n' || r == '\r' || r == '\t':
|
|
return shortEscapeBytes
|
|
case !unicode.IsPrint(r) && r >= firstAstralRune:
|
|
return escapedAstralRuneBytes
|
|
case !unicode.IsPrint(r):
|
|
return escapedRuneBytes
|
|
default:
|
|
return utf8.RuneLen(r)
|
|
}
|
|
}
|
|
|
|
// Truncate caps s at maxBytes of ENCODED output, marking the value
|
|
// when it cuts.
|
|
//
|
|
// Budgeting raw bytes would not bound the line. Escaping only ever
|
|
// grows a value, so a raw budget spent on characters the encoder
|
|
// escapes buys a field several times its nominal size — and the line
|
|
// is the thing an operator is told to multiply by their request rate.
|
|
// Charging each rune what it will actually cost is what makes a stated
|
|
// ceiling true rather than merely larger. The visible consequence is
|
|
// that an escape-heavy value keeps a shorter prefix than a plain one,
|
|
// which is the correct trade.
|
|
//
|
|
// The result is always valid UTF-8. A cut on a byte boundary can split
|
|
// a multi-byte rune, and a header can carry bytes that were never
|
|
// valid UTF-8 to begin with; both are dropped rather than kept, since
|
|
// an encoder would otherwise spend six bytes replacing each one.
|
|
func Truncate(s string, maxBytes int) string {
|
|
// No rune encodes to fewer bytes than it occupies, so nothing past
|
|
// maxBytes raw can fit the budget. Slicing first bounds the scan
|
|
// below to the budget rather than to the size of the header the
|
|
// client sent.
|
|
window, cut := s, false
|
|
if len(window) > maxBytes {
|
|
window, cut = window[:maxBytes], true
|
|
}
|
|
|
|
var (
|
|
kept strings.Builder
|
|
spent int
|
|
)
|
|
|
|
for i := 0; i < len(window); {
|
|
r, size := utf8.DecodeRuneInString(window[i:])
|
|
if r == utf8.RuneError && size == 1 {
|
|
i += size
|
|
|
|
continue
|
|
}
|
|
|
|
cost := EncodedBytes(r)
|
|
if spent+cost > maxBytes {
|
|
cut = true
|
|
|
|
break
|
|
}
|
|
|
|
spent += cost
|
|
|
|
kept.WriteString(window[i : i+size])
|
|
|
|
i += size
|
|
}
|
|
|
|
if !cut {
|
|
return kept.String()
|
|
}
|
|
|
|
return kept.String() + TruncationMarker
|
|
}
|