Author SHA1 Message Date
clawbot 29922dba51 Move the command line from cmd/bsdaily into internal/cli (closes #8)
check / check (push) Waiting to run
cmd/bsdaily/main.go now only passes Version to cli.Main and exits with
the status it returns, as REPO_POLICIES.md requires of cmd/. The cobra
command, the flag rules and the date parsing moved unchanged into
internal/cli, so flags, help text, error messages, exit status and the
first log line stay as they were. The fallback to dev for an empty
version moved with them into cli.Main.

parseTargetDates is exported as ParseTargetDates so that a table test
outside the package can reach it; the test never touches the filesystem
or runs the extraction. README describes the new layout and TODO.md
records the step.

Model: opus-5-5
2026-10-06 19:42:41 +00:00
clawbot f5c7614768 Stamp the git tag or short commit into the binary (closes #4)
check / check (push) Waiting to run
A plain `docker build .` now stamps bsdaily's version: the VERSION
build argument when one is given, otherwise `git describe --tags
--always` on the .git in the build context, as the canonical
Dockerfile does. The build fails if .git is there and the version
still comes out empty, dev or unknown. A host `make` build stamps the
same `git describe`, or dev. bsdaily logs the version on the first
line of every run and prints it with --version; a build that stamps
nothing, or an empty value, reports dev.

Judgement call: the build line also takes the canonical -trimpath and
-s -w.
One //nolint (gochecknoglobals): -X can only set a package-level
variable.

Model: opus-5-5
Co-authored-by: clawbot <sneak+clawbot@sneak.cloud>
2026-10-06 18:41:51 +02:00
clawbot c16f177575 Adopt the shared .golangci.yml and fix the code to it (closes #6)
check / check (push) Successful in 5m14s
Vendor .golangci.yml byte-identical from sneak/prompts at cc440118 and
move the Dockerfile lint phase to golangci-lint v2.14.0 by the digest
REPO_POLICIES.md names. Fix the code to that config with flags, help
text, output files, SQL and the order of steps unchanged; long
functions are split into named steps.

Judgement call: the extraction transaction is now rolled back on every
early return; the old deferred rollback missed most failures and could
dereference a nil transaction.
Thirteen //nolint directives (gosec, unconvert, mnd, unqueryvet), each
with its reason.

Model: opus-5-5
2026-10-06 14:24:48 +02:00
clawbot 925a3896f7 Format Markdown with prettier in make fmt and make fmt-check (closes #7)
check / check (push) Successful in 5m9s
script/fmt now also runs prettier over every Markdown file, and
script/fmt-check checks them without writing. prettier is pinned in
package.json and yarn.lock. script/bootstrap keeps its Go and apt
handling and gains the canonical node and yarn install (nvm from a
hash-checked archive, yarn through corepack), at the pinned versions
when they are absent; a node or yarn already installed is used as is.
Both fmt scripts find yarn the canonical way. The new files come from
sneak/prompts at cc440118c876. README.md and TODO.md are rewrapped by
make fmt, with no wording changes.

Deviation: package.json drops the canonical "license": "MIT" line; this
repo is WTFPL.

Model: opus-5-5
Co-authored-by: clawbot <sneak+clawbot@sneak.cloud>
2026-10-06 12:24:51 +02:00
clawbot 17b6b84322 Bring the repo up to the standard layout (closes #1)
check / check (push) Successful in 5m6s
Vendors .gitignore, .dockerignore, .editorconfig, REPO_POLICIES.md,
script/cibuild, script/docker, script/lint, script/test and the CI
workflow byte-identical from sneak/prompts commit cc440118c876; the
two ignore files then add this repo's own build output, root-anchored
so cmd/bsdaily/ stays in. The Dockerfile gains a lint phase on the
golangci-lint image already pinned and a test phase on the Debian Go
image with sqlite3 and zstd; the build stage copies a file from each,
so a plain docker build cannot skip them. Nothing installs
golangci-lint on the host. On apt, script/bootstrap refreshes the
package lists once before its first install.

.golangci.yml and the lint cleanup: #6

Model: opus-5-5
2026-10-06 08:25:03 +02:00
35 changed files with 2033 additions and 676 deletions
+79 -6
View File
@@ -1,6 +1,79 @@
.git/ # .dockerignore does NOT use .gitignore semantics. Docker matches with
bin/ # moby/patternmatcher: filepath.Match plus `**`, so `*` does not cross
*.md # `/` and an unprefixed pattern is anchored at the context root. Every
LICENSE # depth-independent pattern therefore needs `**/`, or `config/.env` and
.editorconfig # `certs/server.key` still ship while this file reads as solved. Only
.gitignore # genuinely root-anchored entries go unprefixed. Never transplant these
# into .gitignore, where `**/` is wrong.
#
# Matching is case-sensitive, so secrets use character ranges rather
# than an ALL-CAPS twin, which would still miss `Server.Key`.
#
# Extend with this repo's own host-built artifacts, written anchored:
# `/myapp`, never `**/myapp`, which also matches `cmd/myapp/` and
# deletes the package directory from the context.
# .git is sent without its config. Without a VERSION build argument the
# stage that compiles runs `git describe --tags --always` on .git, which
# does not need .git/config; that file can hold a credential, such as a
# password in a remote URL or the token the CI checkout step stores there.
# Each submodule keeps a config with the same exposure in its git directory
# under .git/modules/, nested again for a submodule's own submodules, or in
# its own .git directory when it keeps one.
# KNOWN GAP: a submodule whose name has a `config` segment (`config`,
# `deploy/config`, `config/lib`) loses its whole git directory, because
# `**/.git/modules/**/config` also matches that segment's directory
# under .git/modules/. Go's version stamping then fails the build;
# nothing leaks. Name such a submodule without that segment:
# `git submodule add --name`.
**/.git/config
**/.git/modules/**/config
# Agent scratch: one full checkout of the repo per in-flight agent.
# Anchored because it occurs once where agents run at the repo root.
# KNOWN GAP: a repo running agents in subdirectories still ships
# `services/api/.claude/` and must add its own anchored entry.
.claude
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Re-include a committed template with a negation if the
# build needs one: `!docs/example.env`.
**/*.[eE][nN][vV]
**/.[eE][nN][vV].*
**/.[eE][nN][vV][rR][cC]
# Private keys and the bundles carrying them. Public certificates
# (*.crt, *.cer) are deliberately absent: they are legitimate inputs.
**/*.[pP][eE][mM]
**/*.[kK][eE][yY]
**/*.[pP]12
**/*.[pP][fF][xX]
**/[iI][dD]_[rR][sS][aA]
**/[iI][dD]_[dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]
**/[iI][dD]_[eE][cC][dD][sS][aA]_[sS][kK]
**/[iI][dD]_[eE][dD]25519
**/[iI][dD]_[eE][dD]25519_[sS][kK]
# Dependencies: restored inside the image, never copied in.
**/node_modules
# OS metadata.
**/.DS_Store
**/Thumbs.db
# Editor state: never a build input, and it churns COPY.
**/*.swp
**/*.swo
**/*~
**/*.bak
**/.idea
**/.vscode
**/*.sublime-*
# This repo's host-built artifacts: `make`, `go test -c` and
# `make test-coverage`.
/bsdaily
/*.test
/coverage.out
/coverage.html
+6
View File
@@ -10,3 +10,9 @@ insert_final_newline = true
[Makefile] [Makefile]
indent_style = tab indent_style = tab
[*.go]
indent_style = tab
# This repository's own sections, such as one for another language it
# uses, go below this comment, and a re-vendor keeps them.
+7 -12
View File
@@ -1,14 +1,9 @@
name: check name: check
on: on: [push]
push:
branches: [main]
pull_request:
branches: [main]
jobs: jobs:
check: check:
runs-on: ubuntu-latest runs-on: ubuntu-latest
steps: steps:
# actions/checkout v4, 2024-09-16 # actions/checkout v4.2.2, 2026-02-22
- uses: actions/checkout@34e114876b0b11c390a56381ad16ebd13914f8d5 - uses: actions/checkout@11bd71901bbe5b1630ceea73d27597364c9af683
- name: Build and check - run: script/cibuild
run: script/cibuild
+57 -5
View File
@@ -1,6 +1,58 @@
bin/ # OS
vendor/ .DS_Store
data/ Thumbs.db
.env
*.exe # Editors
*.swp
*.swo
*~
*.bak
.idea/
.vscode/
*.sublime-*
# Agent scratch (worktrees of this repo, created and destroyed by
# in-flight tooling). Unanchored: .gitignore patterns already match at
# every depth, so no prefix is wanted here. This is not a .dockerignore
# entry and must not be given a `**/` prefix on the way into one.
.claude/
# Node
node_modules/
# Secrets. Unanchored like every entry above, so each matches at every
# depth. Matching is case-sensitive on Linux, so names use character
# ranges rather than a lowercase form that misses `Server.Key`.
# Environment files. `*.env` covers bare `.env` and the `prod.env`
# convention. Only the templates `example.env` and `sample.env` are
# re-included below. A repository that commits any other template adds
# its own negation at the end of this file, for example `!.env.example`.
*.[eE][nN][vV]
.[eE][nN][vV].*
.[eE][nN][vV][rR][cC]
!example.env
!sample.env
# Private keys and the bundles carrying them.
*.[pP][eE][mM]
*.[kK][eE][yY]
*.[pP]12
*.[pP][fF][xX]
[iI][dD]_[rR][sS][aA]
[iI][dD]_[dD][sS][aA]
[iI][dD]_[eE][cC][dD][sS][aA]
[iI][dD]_[eE][cC][dD][sS][aA]_[sS][kK]
[iI][dD]_[eE][dD]25519
[iI][dD]_[eE][dD]25519_[sS][kK]
# This repository's own entries, such as its build outputs, go below
# this comment, and a re-vendor keeps them. Anchor a binary built at the
# root: `/myapp`, never `myapp`, which also ignores `cmd/myapp/`.
# Go: logs, test binaries, coverage output, and the binary `make` builds.
*.log
*.test
*.out
/coverage.html
/bsdaily /bsdaily
+99
View File
@@ -0,0 +1,99 @@
version: "2"
# Config schema uses the golangci-lint v2 layout (settings live under
# linters.settings, not top-level linters-settings) so that the
# thresholds below are actually applied by golangci-lint >= v2.
run:
timeout: 5m
modules-download-mode: readonly
linters:
default: all
enable:
# Successor to the deprecated gomodguard. Named explicitly, rather than
# left to `default: all`, because it carries the module policy below.
- gomodguard_v2
disable:
# Genuinely incompatible with project patterns
- exhaustruct # Requires all struct fields
- exhaustruct_v5 # Requires all struct fields (successor to exhaustruct)
- godot # Requires comments to end with periods
- wrapcheck # Too verbose for internal packages
- varnamelen # Short names like db, id are idiomatic Go
# Deprecated: the warning is attached to the old name, so it is
# silenced by disabling that name, not by enabling the successor.
- wsl # Deprecated, replaced by wsl_v5
- gomodguard # Deprecated, replaced by gomodguard_v2
settings:
lll:
line-length: 88
funlen:
lines: 80
statements: 50
cyclop:
max-complexity: 15
dupl:
threshold: 100
depguard:
# Test-support code must not be compiled into the shipped binary. A
# test-support package exists to hand a test privileges the program
# itself must never have, so a file that is not a test must not import
# one. Test files, and the files inside a package whose directory name
# ends in `test`, are where that code belongs, and are exempt.
#
# The deny list below is the one part of this file a repository is
# expected to extend, and the only part it may. depguard matches an
# import path against a list of prefixes, so it cannot be told "any path
# whose last segment ends in test"; a repository's own test-support
# packages have to be named here one at a time, by full import path,
# under a module path that differs from repository to repository. Add
# them; change nothing else.
rules:
test-support:
list-mode: lax
files:
- "$all"
- "!$test"
- "!**/*test/**"
deny:
- pkg: net/http/httptest
desc: >-
Test-support code belongs in test files and in packages whose
directory name ends in test, not in the shipped binary.
# Only decisions already recorded in the Go package defaults are
# listed here. Every entry matches the module path exactly.
gomodguard_v2:
blocked:
- module: github.com/rs/zerolog
recommendations:
- log/slog
reason: "Structured logging is stdlib log/slog."
# One entry per pre-fork module path, because the later releases
# are separate paths. A prefix match would be shorter but would
# also reach github.com/go-redis/redismock, the test double for
# the successor these entries recommend.
- module: github.com/go-redis/redis
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/go-redis/redis/v7
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/go-redis/redis/v8
recommendations:
- github.com/redis/go-redis/v9
reason: "Pre-fork module; use the maintained go-redis v9."
- module: github.com/sergi/go-diff
recommendations:
- github.com/aymanbagabas/go-udiff
reason: "No unified diff output; use go-udiff."
- module: github.com/hexops/gotextdiff
recommendations:
- github.com/aymanbagabas/go-udiff
reason: "Unmaintained fork; use go-udiff."
issues:
max-issues-per-linter: 0
max-same-issues: 0
+2
View File
@@ -0,0 +1,2 @@
node_modules/
yarn.lock
+4
View File
@@ -0,0 +1,4 @@
{
"tabWidth": 4,
"proseWrap": "always"
}
+54 -21
View File
@@ -1,8 +1,8 @@
# Lint stage # Lint phase. The linter is invoked directly rather than through `make
# golangci/golangci-lint:v2.12.2-alpine # lint` or `script/lint`, which are themselves a docker build and would
FROM golangci/golangci-lint:v2.12.2-alpine@sha256:91b27804074a0bacea298707f016911e60cf0cdbc6c7bf5ccacb5f0606d18d60 AS lint # recurse into a daemon that does not exist in a build step.
# golangci/golangci-lint:v2.14.0 (Debian-based), 2026-10-06
RUN apk add --no-cache make build-base FROM golangci/golangci-lint@sha256:ad862ba6b3798cbe0fd9fd7408d498fd74fbd2623a92406b2fd3898faf0bf98f AS lint
WORKDIR /src WORKDIR /src
@@ -13,21 +13,43 @@ RUN go mod download
# Copy source code # Copy source code
COPY . . COPY . .
# Run formatting check and linter RUN golangci-lint run ./...
RUN make fmt-check
RUN make lint
# Build stage # Test phase, invoking `go test` directly for the same reason. -race
# golang:1.26.4-alpine # needs cgo and so a C compiler, which the Debian Go image ships and the
# alpine one does not. The tool shells out to sqlite3, zstdmt and
# zstdcat, so tests that run it need them installed.
# golang:1.26.4-trixie, 2026-10-05
FROM golang:1.26.4-trixie@sha256:68b7145ec43d1820b9a56704554b53d1520aa2a15cb5233e374188a31b2a1bce AS test
RUN apt-get update \
&& apt-get install -y --no-install-recommends sqlite3 zstd \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
RUN go test -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 90s -race -v ./...; exit 1; }
# Build stage. Nothing is wanted from either phase above; the copies
# are what make BuildKit build them first, so this stage cannot run
# unless lint and test passed. A plain `docker build .` builds only the
# last stage and what it depends on, so the runtime stage stays last.
# golang:1.26.4-alpine, 2026-06-28
FROM golang:1.26.4-alpine@sha256:3ad57304ad93bbec8548a0437ad9e06a455660655d9af011d58b993f6f615648 AS builder FROM golang:1.26.4-alpine@sha256:3ad57304ad93bbec8548a0437ad9e06a455660655d9af011d58b993f6f615648 AS builder
# Depend on lint stage passing
COPY --from=lint /src/go.sum /dev/null COPY --from=lint /src/go.sum /dev/null
COPY --from=test /src/go.sum /dev/null
ARG VERSION=dev RUN apk add --no-cache git
# A tar-stream context keeps the sender's file owners, which git refuses.
# Install build deps plus the sqlite3 and zstd CLIs the tests/tool shell out to RUN git config --system --add safe.directory /src
RUN apk add --no-cache make build-base sqlite zstd
WORKDIR /src WORKDIR /src
@@ -38,14 +60,25 @@ RUN go mod download
# Copy source code # Copy source code
COPY . . COPY . .
# Run tests # Build (pure Go, no CGO required since we use modernc.org/sqlite).
RUN make test # The VERSION build arg when one is given, otherwise
# `git describe --tags --always` on the .git in the build context. With
# Build (pure Go, no CGO required since we use modernc.org/sqlite) # .git present, a version that is still empty, dev or unknown fails the
RUN CGO_ENABLED=0 go build -o /bsdaily ./cmd/bsdaily # build: git is missing or could not read the checkout.
ARG VERSION
RUN VERSION="${VERSION:-$(git describe --tags --always)}"; \
if [ -e .git ]; then \
case "$VERSION" in ""|dev|unknown) \
echo "version is '$VERSION' although .git is present" >&2; \
exit 1 ;; \
esac; \
fi; \
CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X main.Version=${VERSION}" \
-o /bsdaily ./cmd/bsdaily
# Runtime stage # Runtime stage
# alpine:3.21 # alpine:3.21, 2026-06-28
FROM alpine:3.21@sha256:48b0309ca019d89d40f670aa1bc06e426dc0931948452e8491e3d65087abc07d FROM alpine:3.21@sha256:48b0309ca019d89d40f670aa1bc06e426dc0931948452e8491e3d65087abc07d
# bsdaily shells out to sqlite3, zstdmt, zstdcat at runtime # bsdaily shells out to sqlite3, zstdmt, zstdcat at runtime
+5 -4
View File
@@ -1,7 +1,9 @@
.PHONY: all bootstrap setup check test lint fmt fmt-check build clean deps test-coverage test-integration install release release-snapshot docker hooks .PHONY: all bootstrap setup check test lint fmt fmt-check build clean deps test-coverage test-integration install release release-snapshot docker hooks
# Version number # Stamped into the binary: the same `git describe` a plain `docker build .`
VERSION := 0.1.0-dev # runs, or dev when it prints nothing (outside a git checkout, or where git is
# missing). ?= so that a VERSION already in the environment takes precedence.
VERSION ?= $(or $(shell git describe --tags --always 2>/dev/null),dev)
# Default target # Default target
all: bsdaily all: bsdaily
@@ -36,7 +38,7 @@ lint:
# Build binary (pure Go; no CGO required since we use modernc.org/sqlite). # Build binary (pure Go; no CGO required since we use modernc.org/sqlite).
bsdaily: internal/*/*.go cmd/bsdaily/*.go bsdaily: internal/*/*.go cmd/bsdaily/*.go
CGO_ENABLED=0 go build -o $@ ./cmd/bsdaily CGO_ENABLED=0 go build -ldflags "-X main.Version=$(VERSION)" -o $@ ./cmd/bsdaily
# Clean build artifacts. # Clean build artifacts.
clean: clean:
@@ -46,7 +48,6 @@ clean:
# Install dependencies. # Install dependencies.
deps: deps:
go mod download go mod download
go install github.com/golangci/golangci-lint/cmd/golangci-lint@latest
# Run tests with coverage. # Run tests with coverage.
test-coverage: test-coverage:
+146 -119
View File
@@ -1,28 +1,26 @@
# bsdaily # bsdaily
[bsdaily](https://git.eeqj.de/sneak/bsdaily) is a command-line utility [bsdaily](https://git.eeqj.de/sneak/bsdaily) is a command-line utility written
written in [Go](https://golang.org) that carves a single day (or a range of in [Go](https://golang.org) that carves a single day (or a range of days) of
days) of [Bluesky](https://bsky.app) firehose data out of a large, [Bluesky](https://bsky.app) firehose data out of a large, continuously-growing
continuously-growing SQLite database and writes it out as a self-contained, SQLite database and writes it out as a self-contained,
[zstd](https://facebook.github.io/zstd/)-compressed SQL dump. The dumps are [zstd](https://facebook.github.io/zstd/)-compressed SQL dump. The dumps are
named by date (e.g. `2026-06-27.sql.zst`), organized into per-month named by date (e.g. `2026-06-27.sql.zst`), organized into per-month directories,
directories, and are designed to be published, archived, mirrored, and later and are designed to be published, archived, mirrored, and later re-merged back
re-merged back into a single database. into a single database.
The source database is read from a read-only [ZFS](https://openzfs.org) The source database is read from a read-only [ZFS](https://openzfs.org)
snapshot, so extraction never contends with the live firehose ingester that snapshot, so extraction never contends with the live firehose ingester that is
is writing to the original database. The tool is operationally writing to the original database. The tool is operationally conservative: it
conservative: it checks free disk space before starting, copies the snapshot checks free disk space before starting, copies the snapshot to fast scratch
to fast scratch storage, processes one day at a time to avoid SQLite lock storage, processes one day at a time to avoid SQLite lock contention, verifies
contention, verifies every compressed output before publishing it, and every compressed output before publishing it, and writes output atomically via a
writes output atomically via a temp-file-and-rename so a partial run never temp-file-and-rename so a partial run never leaves a corrupt `.sql.zst` behind.
leaves a corrupt `.sql.zst` behind.
This project was written by [@sneak](https://sneak.berlin) to produce a This project was written by [@sneak](https://sneak.berlin) to produce a daily,
daily, mergeable, publicly-mirrorable archive of the Bluesky firehose. It is mergeable, publicly-mirrorable archive of the Bluesky firehose. It is currently
currently a one-person effort. The current version is pre-1.0 and there has a one-person effort. The current version is pre-1.0 and there has not yet been a
not yet been a versioned release; [SemVer](https://semver.org) will be used versioned release; [SemVer](https://semver.org) will be used for releases.
for releases.
# Build Status # Build Status
@@ -33,14 +31,14 @@ branch must always be green.
Primary development happens on a privately-run Gitea instance at Primary development happens on a privately-run Gitea instance at
[https://git.eeqj.de/sneak/bsdaily](https://git.eeqj.de/sneak/bsdaily) and [https://git.eeqj.de/sneak/bsdaily](https://git.eeqj.de/sneak/bsdaily) and
issues are [tracked issues are [tracked there](https://git.eeqj.de/sneak/bsdaily/issues).
there](https://git.eeqj.de/sneak/bsdaily/issues).
Changes must always be formatted with a standard `go fmt`, syntactically Changes must always be formatted with `make fmt` (`go fmt` for Go, prettier for
valid, and must pass the linting defined in the repository (presently the Markdown), syntactically valid, and must pass the linting defined in the
`golangci-lint` defaults), which can be run with a `make lint`. The `main` repository (the shared `.golangci.yml` from `sneak/prompts`), which can be run
branch is protected and all changes must be made via [pull with a `make lint`. The `main` branch is protected and all changes must be made
requests](https://git.eeqj.de/sneak/bsdaily/pulls) and pass CI to be merged. via [pull requests](https://git.eeqj.de/sneak/bsdaily/pulls) and pass CI to be
merged.
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards, See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
tooling requirements, and workflow conventions. tooling requirements, and workflow conventions.
@@ -50,48 +48,55 @@ tooling requirements, and workflow conventions.
This repository adheres to the This repository adheres to the
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all) [Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
standard: normalized scripts in `script/` are the entrypoints for the standard: normalized scripts in `script/` are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call development workflow, and the Makefile targets are thin shims that call them. We
them. We provide: provide:
- `script/bootstrap` — install all development dependencies (go, - `script/bootstrap` — install all development dependencies (go, Go module
golangci-lint, Go module download) download, node and yarn at pinned versions when absent, and the prettier
pinned in `package.json` and `yarn.lock`); a node or yarn already installed is
used whatever its version; the linter is not installed on the host
- `script/setup` — make a fresh clone ready for development: runs - `script/setup` — make a fresh clone ready for development: runs
`script/bootstrap`, then `script/install-precommit` `script/bootstrap`, then `script/install-precommit`
- `script/projectname` — print the project name (used for the Docker - `script/projectname` — print the project name (used for the Docker image tags)
image tag) - `script/test` — build the Dockerfile's `test` phase, which runs the test suite
- `script/test` — run the test suite (verbose rerun on failure) with `-race` (verbose rerun on failure)
- `script/lint` — run `golangci-lint run ./...` - `script/lint` — build the Dockerfile's `lint` phase, which runs
- `script/fmt` — format all code (writes) `golangci-lint run ./...`
- `script/fmt-check` — check formatting (read-only) - `script/fmt` — format the Go code with `go fmt` and every Markdown file with
- `script/check` — run `script/test`, `script/lint`, and prettier (writes)
`script/fmt-check` - `script/fmt-check` — check the Go formatting with `gofmt` and the Markdown
- `script/docker` — build the Docker image tagged via formatting with prettier (read-only)
`script/projectname` - `script/check` — run `script/test`, `script/lint`, and `script/fmt-check`
- `script/cibuild` — CI entrypoint: `docker build .` (the Dockerfile - `script/docker` — build the Docker image tagged via `script/projectname`; the
runs the checks) build runs the `lint` and `test` phases first
- `script/precommit` — pre-commit gate: `go mod tidy` + `go fmt` (must - `script/cibuild` — CI entrypoint: runs `script/bootstrap` and `script/check`,
not change files), then `script/check` then builds the Docker image tagged via `script/projectname`
- `script/install-precommit` — install the git pre-commit hook that - `script/precommit` — pre-commit gate: `go mod tidy` (must not change `go.mod`
runs `script/precommit` or `go.sum`) and `go fmt`, then `script/check`
- `script/install-precommit` — install the git pre-commit hook that runs
`script/precommit`
Every Docker build in `script/` is uncached, so the `lint` and `test` phases
always run rather than being served from the build cache.
# Problem Statement # Problem Statement
A Bluesky firehose ingester writes every observed post (and associated A Bluesky firehose ingester writes every observed post (and associated users,
users, hashtags, URLs, and media references) into a single ever-growing hashtags, URLs, and media references) into a single ever-growing SQLite
SQLite database, `firehose.db`. This database has several properties that database, `firehose.db`. This database has several properties that make it
make it awkward to publish or archive directly: awkward to publish or archive directly:
- It is **large and always growing**, so re-publishing the whole thing every - It is **large and always growing**, so re-publishing the whole thing every day
day is wasteful. is wasteful.
- It is **continuously written**, so reading from it directly risks lock - It is **continuously written**, so reading from it directly risks lock
contention with the live ingester and inconsistent reads. contention with the live ingester and inconsistent reads.
- It is **monolithic**, so there is no natural unit at which to mirror, - It is **monolithic**, so there is no natural unit at which to mirror, share,
share, or distribute "just yesterday's posts". or distribute "just yesterday's posts".
What is wanted instead is a stable, immutable, per-day artifact: a small What is wanted instead is a stable, immutable, per-day artifact: a small file
file containing exactly one calendar day of firehose data, cheap to publish, containing exactly one calendar day of firehose data, cheap to publish, cheap to
cheap to mirror, and trivially re-mergeable into a full database by anyone mirror, and trivially re-mergeable into a full database by anyone who collects a
who collects a set of them. set of them.
# Proposed Solution # Proposed Solution
@@ -101,81 +106,86 @@ A tool, `bsdaily`, that:
filesystem, so it reads from a consistent point-in-time copy that the live filesystem, so it reads from a consistent point-in-time copy that the live
ingester cannot be writing to; ingester cannot be writing to;
- copies the snapshot's database files to fast scratch storage; - copies the snapshot's database files to fast scratch storage;
- **extracts** a single day's `posts` (and all rows reachable from them) into - **extracts** a single day's `posts` (and all rows reachable from them) into a
a fresh, minimal per-day SQLite database; fresh, minimal per-day SQLite database;
- **dumps** that per-day database to SQL and pipes it through multithreaded - **dumps** that per-day database to SQL and pipes it through multithreaded zstd
zstd compression; compression;
- **verifies** the compressed output (zstd integrity check plus a sanity - **verifies** the compressed output (zstd integrity check plus a sanity check
check that the decompressed stream actually looks like SQL); that the decompressed stream actually looks like SQL);
- **publishes** the result atomically as - **publishes** the result atomically as
`DailiesBase/YYYY-MM/YYYY-MM-DD.sql.zst`. `DailiesBase/YYYY-MM/YYYY-MM-DD.sql.zst`.
Each daily dump is emitted with `INSERT` statements over the full schema Each daily dump is emitted with `INSERT` statements over the full schema
(including the deduplicated `users`, `hashtags`, and `urls` lookup tables), (including the deduplicated `users`, `hashtags`, and `urls` lookup tables), so
so any collection of daily dumps can be merged into a single database by any collection of daily dumps can be merged into a single database by rewriting
rewriting `INSERT INTO` to `INSERT OR IGNORE INTO` and replaying them in `INSERT INTO` to `INSERT OR IGNORE INTO` and replaying them in sequence. Two
sequence. Two helper scripts ([`merge_daily_dumps.sh`](merge_daily_dumps.sh) helper scripts ([`merge_daily_dumps.sh`](merge_daily_dumps.sh) and
and [`regenerate_auxiliary_tables.sql`](regenerate_auxiliary_tables.sql)) are [`regenerate_auxiliary_tables.sql`](regenerate_auxiliary_tables.sql)) are
included to do exactly this and to rebuild the aggregate statistics included to do exactly this and to rebuild the aggregate statistics
(`use_count`, `first_seen`, user `resolved_at`/`updated_at`) afterward. (`use_count`, `first_seen`, user `resolved_at`/`updated_at`) afterward.
# Design Goals # Design Goals
- **Never disturb the live ingester.** All reads come from a ZFS snapshot, - **Never disturb the live ingester.** All reads come from a ZFS snapshot, never
never the live database. the live database.
- **Crash-safe, idempotent runs.** Output is written to a temp file and - **Crash-safe, idempotent runs.** Output is written to a temp file and
atomically renamed; a day whose final output already exists is skipped, so atomically renamed; a day whose final output already exists is skipped, so
re-running a range is safe and resumable. re-running a range is safe and resumable.
- **Mergeable output.** Daily dumps re-combine losslessly into a full - **Mergeable output.** Daily dumps re-combine losslessly into a full database
database via `INSERT OR IGNORE`. via `INSERT OR IGNORE`.
- **Operationally cautious.** Free-space preflight checks on both scratch and - **Operationally cautious.** Free-space preflight checks on both scratch and
output filesystems; explicit verification of every artifact before it is output filesystems; explicit verification of every artifact before it is
published. published.
- **Fast where it's free.** Large snapshot copies use a 256MiB buffer, - **Fast where it's free.** Large snapshot copies use a 256MiB buffer,
pre-allocate the destination, and (on Linux) issue `posix_fadvise` pre-allocate the destination, and (on Linux) issue `posix_fadvise`
sequential/willneed hints; extraction uses aggressive, sequential/willneed hints; extraction uses aggressive, crash-unsafe-by-design
crash-unsafe-by-design SQLite pragmas because the working data lives only SQLite pragmas because the working data lives only in disposable scratch
in disposable scratch space. space.
# Non-Goals # Non-Goals
- **Real-time export.** `bsdaily` operates on daily snapshots; the freshest - **Real-time export.** `bsdaily` operates on daily snapshots; the freshest day
day it can produce is the snapshot date minus one. it can produce is the snapshot date minus one.
- **Schema ownership.** The schema is defined by the upstream firehose - **Schema ownership.** The schema is defined by the upstream firehose ingester;
ingester; [`schema.sql`](schema.sql) is included for reference only. [`schema.sql`](schema.sql) is included for reference only. `bsdaily` copies
`bsdaily` copies whatever table and index DDL it finds in the source. whatever table and index DDL it finds in the source.
- **Cross-platform deployment.** It is built and run on Linux (the - **Cross-platform deployment.** It is built and run on Linux (the free-space
free-space check and fadvise hints use `golang.org/x/sys/unix`; a non-Linux check and fadvise hints use `golang.org/x/sys/unix`; a non-Linux build
build compiles but is a no-op for the fadvise hints). The hard-coded paths compiles but is a no-op for the fadvise hints). The hard-coded paths assume
assume the production host's ZFS layout. the production host's ZFS layout.
# How It Works # How It Works
`cmd/bsdaily/main.go` only passes the version to `internal/cli` and exits with
the status it returns. `internal/cli` holds the command line: the flags, the
rules for combining them and the parsing of the dates they name.
`internal/bsdaily` does the extraction.
A single run proceeds as follows: A single run proceeds as follows:
1. **Find the snapshot.** Scan `SnapshotBase` for directories matching 1. **Find the snapshot.** Scan `SnapshotBase` for directories matching
`zfs-auto-snap_daily-YYYY-MM-DD-NNNN`, pick the most recent, and confirm `zfs-auto-snap_daily-YYYY-MM-DD-NNNN`, pick the most recent, and confirm it
it contains `firehose.db`. contains `firehose.db`.
2. **Determine target days.** Default to the snapshot date minus one day; 2. **Determine target days.** Default to the snapshot date minus one day; or use
or use `--date`, or every day in the inclusive `--from`/`--to` range. `--date`, or every day in the inclusive `--from`/`--to` range.
3. **Preflight disk space.** Require at least 500GiB free on the scratch 3. **Preflight disk space.** Require at least 500GiB free on the scratch
filesystem and 20GiB free on the output filesystem. filesystem and 20GiB free on the output filesystem.
4. **Copy the database to scratch.** Copy `firehose.db`, its `-wal`, and (if 4. **Copy the database to scratch.** Copy `firehose.db`, its `-wal`, and (if
present) its `-shm` from the snapshot into a fresh temp directory under present) its `-shm` from the snapshot into a fresh temp directory under
`TmpBase`. `TmpBase`.
5. **Per day**, processed strictly one at a time to avoid SQLite contention: 5. **Per day**, processed strictly one at a time to avoid SQLite contention:
- skip the day if its final output file already exists; - skip the day if its final output file already exists;
- `ATTACH` the copied source DB to a new empty per-day DB, recreate the - `ATTACH` the copied source DB to a new empty per-day DB, recreate the
table DDL, and `INSERT ... SELECT` the target day's `posts` plus all table DDL, and `INSERT ... SELECT` the target day's `posts` plus all rows
rows reachable from them (`posts_hashtags`, `posts_urls`, `hashtags`, reachable from them (`posts_hashtags`, `posts_urls`, `hashtags`, `urls`,
`urls`, `users`, and `media` if that table exists); `users`, and `media` if that table exists);
- abort the day cleanly if there are zero posts (`ErrNoPosts`), rather - abort the day cleanly if there are zero posts (`ErrNoPosts`), rather than
than emitting an empty dump; emitting an empty dump;
- recreate indexes, detach the source, and verify the inserted row count; - recreate indexes, detach the source, and verify the inserted row count;
- `sqlite3 .dump | zstdmt` into a hidden temp file; - `sqlite3 .dump | zstdmt` into a hidden temp file;
- run a zstd integrity check and confirm the decompressed head looks like - run a zstd integrity check and confirm the decompressed head looks like
SQL; SQL;
- atomically rename into place and delete the per-day scratch DB. - atomically rename into place and delete the per-day scratch DB.
6. **Clean up** the temp directory and log a processed/skipped/total summary. 6. **Clean up** the temp directory and log a processed/skipped/total summary.
# Usage # Usage
@@ -184,6 +194,7 @@ A single run proceeds as follows:
bsdaily # extract the snapshot date minus one day bsdaily # extract the snapshot date minus one day
bsdaily --date 2026-06-27 # extract a single specific day bsdaily --date 2026-06-27 # extract a single specific day
bsdaily --from 2026-06-01 --to 2026-06-27 # extract an inclusive range bsdaily --from 2026-06-01 --to 2026-06-27 # extract an inclusive range
bsdaily --version # print the version and exit
``` ```
Flags: Flags:
@@ -192,9 +203,24 @@ Flags:
`--from`/`--to`. `--from`/`--to`.
- `--from YYYY-MM-DD` — start of an inclusive range (requires `--to`). - `--from YYYY-MM-DD` — start of an inclusive range (requires `--to`).
- `--to YYYY-MM-DD` — end of an inclusive range (requires `--from`). - `--to YYYY-MM-DD` — end of an inclusive range (requires `--from`).
- `-v`, `--version` — print the version and exit.
With no flags, the tool extracts the day before the latest snapshot. All With no flags, the tool extracts the day before the latest snapshot. All
progress is logged as structured `slog` text to stderr. progress is logged as structured `slog` text to stderr; the first line of every
run carries the version.
The version is set at link time and depends on how the binary was built:
- `docker build .` takes it from the `VERSION` build argument when one is given,
otherwise from `git describe --tags --always` on the `.git` in the build
context. The build fails if `.git` is there and the version still comes out
empty, `dev` or `unknown`. With neither `.git` nor `VERSION`, the binary
reports `dev`.
- `script/docker`, `script/cibuild` and `make docker` pass the host's
`git describe --tags --always --dirty` as `VERSION`, so on a modified tree the
version ends in `-dirty`. When that prints nothing, they pass `unknown`.
- `make` stamps the host's `git describe --tags --always`, without `-dirty`, or
`dev` when that prints nothing.
## Merging dumps back into a database ## Merging dumps back into a database
@@ -212,9 +238,9 @@ sqlite3 merged.db < regenerate_auxiliary_tables.sql
- **Go** (see [`go.mod`](go.mod) for the toolchain version) to build. - **Go** (see [`go.mod`](go.mod) for the toolchain version) to build.
- **Linux** for production use (ZFS snapshots, `statfs` free-space checks, - **Linux** for production use (ZFS snapshots, `statfs` free-space checks,
`posix_fadvise` hints). `posix_fadvise` hints).
- The **`sqlite3`** and **`zstdmt`** (multithreaded zstd) binaries on - The **`sqlite3`** and **`zstdmt`** (multithreaded zstd) binaries on `PATH`;
`PATH`; `zstdcat` is used for verification. SQLite reads/writes during `zstdcat` is used for verification. SQLite reads/writes during extraction use
extraction use the pure-Go [`modernc.org/sqlite`](https://pkg.go.dev/modernc.org/sqlite) the pure-Go [`modernc.org/sqlite`](https://pkg.go.dev/modernc.org/sqlite)
driver, so no cgo is required for that part. driver, so no cgo is required for that part.
# Configuration # Configuration
@@ -234,21 +260,21 @@ elsewhere.
# Data Model # Data Model
The firehose schema (reference copy in [`schema.sql`](schema.sql)) centers The firehose schema (reference copy in [`schema.sql`](schema.sql)) centers on a
on a `posts` table, with `users` keyed by DID and many-to-many junction `posts` table, with `users` keyed by DID and many-to-many junction tables
tables linking posts to deduplicated `hashtags` and `urls`. An optional linking posts to deduplicated `hashtags` and `urls`. An optional `media` table
`media` table tracks downloaded blobs by content hash. `bsdaily` does not tracks downloaded blobs by content hash. `bsdaily` does not own this schema; it
own this schema; it reflects whatever DDL exists in the source snapshot and reflects whatever DDL exists in the source snapshot and selects forward from
selects forward from `posts` along the foreign-key relationships to produce a `posts` along the foreign-key relationships to produce a referentially-complete
referentially-complete per-day slice. per-day slice.
# Use Cases # Use Cases
## Daily public archive ## Daily public archive
Publish one small, immutable file per day to static HTTP (or IPFS, or a Publish one small, immutable file per day to static HTTP (or IPFS, or a mirror
mirror network) so that anyone can fetch exactly the day(s) they want and network) so that anyone can fetch exactly the day(s) they want and re-merge them
re-merge them locally. locally.
## Backfilling a range ## Backfilling a range
@@ -258,16 +284,17 @@ safe to re-run.
## Reconstituting a full database ## Reconstituting a full database
Collect any set of daily dumps and merge them with `INSERT OR IGNORE` to Collect any set of daily dumps and merge them with `INSERT OR IGNORE` to rebuild
rebuild a complete, queryable SQLite database, then regenerate the aggregate a complete, queryable SQLite database, then regenerate the aggregate statistics
statistics tables. tables.
# See Also # See Also
## Links ## Links
- Repo: [https://git.eeqj.de/sneak/bsdaily](https://git.eeqj.de/sneak/bsdaily) - Repo: [https://git.eeqj.de/sneak/bsdaily](https://git.eeqj.de/sneak/bsdaily)
- Issues: [https://git.eeqj.de/sneak/bsdaily/issues](https://git.eeqj.de/sneak/bsdaily/issues) - Issues:
[https://git.eeqj.de/sneak/bsdaily/issues](https://git.eeqj.de/sneak/bsdaily/issues)
- Bluesky: [https://bsky.app](https://bsky.app) - Bluesky: [https://bsky.app](https://bsky.app)
- zstd: [https://facebook.github.io/zstd/](https://facebook.github.io/zstd/) - zstd: [https://facebook.github.io/zstd/](https://facebook.github.io/zstd/)
+369 -86
View File
@@ -1,6 +1,6 @@
--- ---
title: Repository Policies title: Repository Policies
last_modified: 2026-07-06 last_modified: 2026-10-06
--- ---
This document covers repository structure, tooling, and workflow standards. Code This document covers repository structure, tooling, and workflow standards. Code
@@ -60,17 +60,28 @@ style conventions are in separate documents:
prerequisite since nvm requires bash. yarn is then pinned via prerequisite since nvm requires bash. yarn is then pinned via
`corepack prepare yarn@<version> --activate`. Never install "latest" or "lts"; `corepack prepare yarn@<version> --activate`. Never install "latest" or "lts";
always exact versions. `script/cibuild` runs the CI build: it changes to the always exact versions. `script/cibuild` runs the CI build: it changes to the
repo root and runs `docker build .`; the Gitea workflow calls it. Four further repo root, runs `script/bootstrap`, runs `script/check`, and builds the image
scripts are our own extensions to the standard: `script/check` runs with the version; the Gitea workflow calls it. **`script/cibuild` runs
`script/test`, `script/lint`, and `script/fmt-check`; `script/precommit` is `script/bootstrap` first**, because the workflow checks out the repo and runs
what the git pre-commit hook runs, and it calls `script/check`; nothing else, while `script/fmt-check` runs the formatter on the host: on a
`script/install-precommit` installs the git pre-commit hook (the `make hooks` pristine checkout with nothing installed the run dies there, after the
target shims to it); and `script/projectname` (literally that filename) simply containerised gates have passed. **The bootstrap alone is not enough**:
outputs the project's name. Scripts that need the name call `script/bootstrap` installs node and yarn under nvm and leaves neither on the
`script/projectname` — e.g. `script/docker` assembles its image tag from it — `PATH` of the shell that called it, so a bare `yarn` still exits 127. The host
so those scripts stay byte-identical across all repos. Repo-type-specific entrypoints that need yarn — `script/fmt` and `script/fmt-check` — therefore
pre-commit extras (e.g. `go mod tidy` verification in Go repos) belong in source nvm for the pinned node version before invoking it, exactly as
`script/precommit`, not in the hook itself. Model scripts are at `script/bootstrap`'s own install step does. A runner carrying nothing but
docker and git then gets through `script/check`. Four further scripts are our
own extensions to the standard: `script/check` runs `script/test`,
`script/lint` and `script/fmt-check`; `script/precommit` is what the git
pre-commit hook runs, and it calls `script/check`; `script/install-precommit`
installs the git pre-commit hook (the `make hooks` target shims to it); and
`script/projectname` (literally that filename) simply outputs the project's
name. Scripts that need the name call `script/projectname` — e.g.
`script/docker` assembles its image tag from it — so those scripts stay
byte-identical across all repos. Repo-type-specific pre-commit extras (e.g.
`go mod tidy` verification in Go repos) belong in `script/precommit`, not in
the hook itself. Model scripts are at
`https://git.eeqj.de/sneak/prompts/raw/branch/main/script/<name>`. The README `https://git.eeqj.de/sneak/prompts/raw/branch/main/script/<name>`. The README
must document the provided scripts in an **Entrypoints** section (see the must document the provided scripts in an **Entrypoints** section (see the
README requirements below). README requirements below).
@@ -89,87 +100,201 @@ style conventions are in separate documents:
contributor should be able to understand the entire development workflow by contributor should be able to understand the entire development workflow by
reading the Makefile. reading the Makefile.
- Every repo should have a `Dockerfile`. All Dockerfiles must run `make check` - Every repo should have a `Dockerfile`, and it carries the repo's gates: a
as a build step so the build fails if the branch is not green. For non-server `lint` phase and a `test` phase, with the final stage depending on both so the
repos, the Dockerfile should bring up a development environment and run image cannot be built unless they pass. For non-server repos the final stage
`make check`. For server repos, `make check` should run as an early build brings up a development environment; for server repos it is the runtime image.
stage before the final image is assembled. Dockerfiles install development The gate phases and the build stage start from their pinned base images and
prerequisites by running `script/bootstrap` rather than duplicating installs install what those images lack either inline, as the canonical Go `Dockerfile`
inline; COPY `script/` and the dependency manifests (`package.json` + below does for `git`, or by running `script/bootstrap`, as the `prompts`
`yarn.lock`, `go.mod` + `go.sum`, etc.) before running it so the bootstrap repo's own `Dockerfile` does for its yarn packages. The development
layer stays cached until dependencies change. environment stage installs development prerequisites by running
`script/bootstrap` rather than duplicating its installs inline. A stage that
runs `script/bootstrap` COPYs `script/` and the dependency manifests
(`package.json` + `yarn.lock`, `go.mod` + `go.sum`, etc.) before running it.
- **Dockerfiles must use a separate lint stage for fail-fast feedback.** Go - **Linting and testing run in Docker, as phases of the `Dockerfile`.** There is
repos use a multistage build where linting runs in an independent stage based no separate lint file. `script/lint` and `script/test` each build one phase
on the `golangci/golangci-lint` image (pinned by hash). This stage runs and nothing else:
`make fmt-check` and `make lint` before the full build begins. The build stage
then declares an explicit dependency on the lint stage via
`COPY --from=lint /src/go.sum /dev/null`, which forces BuildKit to complete
linting before proceeding to compilation and tests. This ensures lint failures
surface in seconds rather than minutes, without blocking on dependency
download or compilation in the build stage.
The standard pattern for a Go repo Dockerfile is: ```sh
tag="$(script/projectname)"
docker build --no-cache --target lint -t "$tag-lint" .
docker build --no-cache --target test -t "$tag-test" .
```
**A stage that is not the last one in the file is built only when the final
stage's chain depends on it, or when `--target` names it.** That is why the
two gates are always invoked by name here, and why the final stage carries a
`COPY --from=` of a harmless file from each of them: without that edge a
plain `docker build .` builds the last stage alone and exits 0 having linted
and tested nothing.
**Every `docker build` in `script/` is tagged**, here and in
`script/cibuild` and `script/docker`. An untagged build leaves a dangling
image behind on every invocation, on every developer host and every CI
runner; a tagged one replaces the previous image. Each script assigns the
tag on its own line before the build, so `set -e` stops it where
`script/projectname` fails.
Inside a phase the tool is invoked directly — `golangci-lint`, `go test`,
`eslint`, `prettier` — never through `make lint` or `script/test`, which are
themselves a `docker build` and would recurse into a daemon that does not
exist in a build step. Formatting is the exception and stays on the host:
`script/fmt` writes the working tree, and `script/fmt-check` is its
read-only twin.
**No lint verdict may come from a host invocation of the linter.** On a
shared host golangci-lint reads a result cache keyed on file content rather
than location, so a second checkout of the same content is served the first
one's findings, and a host-global lock in `$TMPDIR` makes concurrent runs
exit non-zero with `parallel golangci-lint is running` — a status a caller
cannot tell from real findings. Both have produced wrong verdicts in this
org, in both directions. A container has its own cache, its own `TMPDIR` and
a digest-pinned binary, so neither is reachable.
- **Any build that runs checks is built with `--no-cache`.** Docker invalidates
a `COPY` layer only when the copied content changes, so on an unchanged tree
the check `RUN` is served from cache, nothing executes, and the build still
exits 0. Every `docker build` in `script/` therefore passes `--no-cache`:
`script/lint`, `script/test`, `script/cibuild` and `script/docker` are the
four, and there is no fifth — `script/check` runs the two gate phases and
`script/fmt-check`, and builds no image of its own. A bare `docker build .` is
not evidence that anything ran: a sub-second build reporting success is a
cache hit, not a result. Never invalidate by pruning — `docker builder prune`
and friends destroy a build cache shared with every other build on the host.
When a check is added or changed, prove it works by planting a defect it must
catch and watching the run fail on it, then revert the defect. A green run
alone shows neither that the check ran nor that it covers what it should.
- **The gate phases are separate stages, and the build stage depends on both.**
The lint phase is based on the `golangci/golangci-lint` image (pinned by
hash), so lint failures surface in seconds rather than after a full compile,
and the test phase is based on the Debian Go image. The canonical Go repo
`Dockerfile`:
```dockerfile ```dockerfile
# Lint stage — fast feedback on formatting and lint issues # Lint phase
# golangci/golangci-lint:v2.x.x, YYYY-MM-DD # golangci/golangci-lint:v2.x.x, YYYY-MM-DD
FROM golangci/golangci-lint@sha256:... AS lint FROM golangci/golangci-lint@sha256:... AS lint
WORKDIR /src WORKDIR /src
COPY go.mod go.sum ./ COPY go.mod go.sum ./
RUN go mod download RUN go mod download
COPY . . COPY . .
RUN make fmt-check RUN golangci-lint run --config .golangci.yml ./...
RUN make lint
# Build stage # Test phase. -race needs cgo and so a C compiler, which the Debian Go
# golang:1.x-alpine, YYYY-MM-DD # image ships and the alpine one does not.
FROM golang@sha256:... AS builder # golang:1.x, YYYY-MM-DD
FROM golang@sha256:... AS test
WORKDIR /src WORKDIR /src
# Force BuildKit to run the lint stage before proceeding
COPY --from=lint /src/go.sum /dev/null
COPY go.mod go.sum ./ COPY go.mod go.sum ./
RUN go mod download RUN go mod download
COPY . . COPY . .
RUN make test RUN go test -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \
go test -timeout 90s -race -v ./...; exit 1; }
ARG VERSION=dev # Build stage. Nothing is wanted from either phase above; the copies
RUN CGO_ENABLED=0 go build -trimpath \ # are what make BuildKit build them first, so this stage cannot run
-ldflags="-s -w -X main.Version=${VERSION}" \ # unless lint and test passed.
-o /app ./cmd/app/ # golang:1.x-alpine, YYYY-MM-DD
FROM golang@sha256:... AS builder
COPY --from=lint /src/go.sum /dev/null
COPY --from=test /src/go.sum /dev/null
RUN apk add --no-cache git
# A tar-stream context keeps the sender's file owners, which git refuses.
RUN git config --system --add safe.directory /src
WORKDIR /src
COPY go.mod go.sum ./
RUN go mod download
COPY . .
# Runtime stage # The VERSION build arg when one is given, otherwise
# `git describe --tags --always` on the .git in the build context. With
# .git present, a version that is still empty, dev or unknown fails the
# build: git is missing or could not read the checkout.
ARG VERSION
RUN VERSION="${VERSION:-$(git describe --tags --always)}"; \
if [ -e .git ]; then \
case "$VERSION" in ""|dev|unknown) \
echo "version is '$VERSION' although .git is present" >&2; \
exit 1 ;; \
esac; \
fi; \
CGO_ENABLED=0 go build -trimpath \
-ldflags="-s -w -X main.Version=${VERSION}" \
-o /app ./cmd/app/
# Runtime stage, and the last one
FROM alpine@sha256:... FROM alpine@sha256:...
COPY --from=builder /app /usr/local/bin/app COPY --from=builder /app /usr/local/bin/app
ENTRYPOINT ["app"] ENTRYPOINT ["app"]
``` ```
Key points: Key points:
- The lint stage uses the `golangci/golangci-lint` image directly (it - The lint phase uses the `golangci/golangci-lint` image directly (it has
includes both Go and the linter), so there is no need to install the both Go and the linter), so nothing needs installing.
linter separately. - `COPY --from=<phase> /src/go.sum /dev/null` is a no-op copy whose only
- `COPY --from=lint /src/go.sum /dev/null` is a no-op file copy that creates purpose is the ordering edge. BuildKit runs stages in parallel by default,
a stage dependency. BuildKit runs stages in parallel by default; without and a stage nothing depends on is not built at all, so without these two
this line, the build stage would not wait for lint to finish and a lint lines a red gate would not fail the build.
failure might not fail the overall build. - Keep the runtime stage last, and if you add a stage after it, give it the
same two copies. A plain `docker build .` builds the last stage's chain
and nothing else.
- If the project uses `//go:embed` directives that reference build artifacts - If the project uses `//go:embed` directives that reference build artifacts
(e.g. a web frontend compiled in a separate stage), the lint stage must (e.g. a web frontend compiled in a separate stage), the lint phase must
create placeholder files so the embed directives resolve. Example: create placeholder files so the embed directives resolve. Example:
`RUN mkdir -p web/dist && touch web/dist/index.html web/dist/style.css`. `RUN mkdir -p web/dist && touch web/dist/index.html web/dist/style.css`.
The lint stage should not depend on the actual build output — it exists to - If the project requires CGO or system libraries for linting, install them
fail fast. in the lint phase. The `golangci/golangci-lint` image is Debian-based and
- If the project requires CGO or system libraries for linting (e.g. has no `apk`, so install with `apt-get` under the Debian package name
`vips-dev`), install them in the lint stage with `apk add`. (`libvips-dev`, where alpine says `vips-dev`), and delete the package
- The build stage runs `make test` after compilation setup. Tests run in the lists in the same `RUN`, so the layer does not keep them:
build stage, not the lint stage, because they may require compiled
artifacts or heavier dependencies. ```dockerfile
RUN apt-get update \
&& apt-get install -y --no-install-recommends libvips-dev \
&& rm -rf /var/lib/apt/lists/*
```
- `.dockerignore` lets `.git` into the build context. It keeps out every git
`config` at any depth (`**/.git/config`, `**/.git/modules/**/config`): the
repository's own, each submodule's under `.git/modules/`, and that of a
submodule keeping its own `.git` directory. `git describe` does not need
them, and each can hold a credential: a password in a remote URL, or the
token the CI checkout step stores there. A submodule whose name has a
`config` segment (`config`, `deploy/config`, `config/lib`) loses its whole
git directory to `**/.git/modules/**/config`, and Go's version stamping
then fails the build: give it a name without that segment
(`git submodule add --name`). The stage that compiles has `git` (the
Debian Go image has it; an alpine one needs `apk add --no-cache git`) and
takes the version from the `VERSION` build argument when one is given,
otherwise from `git describe --tags --always`. That gives the tag on a
tagged commit; on a later commit, the tag, the number of commits since it
and the short commit (`v1.2.3-4-gabc1234`); and the short commit when no
tag is reachable. The stage that compiles also marks its working directory
safe for git (`git config --system --add safe.directory /src`): a context
sent as a tar stream keeps the sender's file owners, and git refuses a
checkout owned by another user, so the version would come out empty.
`ARG VERSION` has no default, and the build fails if the context carries
`.git` and the version still comes out empty, `dev` or `unknown`. A plain
`docker build .` with no build arguments must succeed; a Dockerfile that
refuses an empty build argument drops that refusal and keeps the argument.
- Every repo should have a Gitea Actions workflow (`.gitea/workflows/`) that - Every repo should have a Gitea Actions workflow (`.gitea/workflows/`) that
runs `script/cibuild` (which runs `docker build .`) on push. Since the runs `script/cibuild` on push, and checks out the repo as its only other step.
Dockerfile already runs `make check`, a successful build implies all checks That script bootstraps, runs the gate phases, and then builds the image, so a
pass. successful run means every check passed; a bare `docker build .` does not
carry the same guarantee, because its gate phases may come from the cache. The
image build is uncached and so runs the gate phases a second time. That is the
price of the rule above, and it is worth paying: the image that ships is built
from a run of its own gates rather than from a cache entry. A separate
workflow limited to `main` by a `branches` list under `on: push` cannot be
checked by review: to try a change to it, add the feature branch to that list
and push, then remove the branch from the list again before merging. Keep any
job in it that publishes behind `if: github.ref_name == 'main'`, so the run
from the feature branch publishes nothing.
- Use platform-standard formatters: `black` for Python, `prettier` for - Use platform-standard formatters: `black` for Python, `prettier` for
JS/CSS/Markdown/HTML, `go fmt` for Go. Always use default configuration with JS/CSS/Markdown/HTML, `go fmt` for Go. Always use default configuration with
@@ -189,14 +314,21 @@ style conventions are in separate documents:
module under test to verify it compiles/parses. There is no excuse for module under test to verify it compiles/parses. There is no excuse for
`make test` to be a no-op. `make test` to be a no-op.
- `make test` must complete in under 20 seconds. Add a 30-second timeout in the - `make test` must complete in under 60 seconds. That is the hard cap, and a
Makefile. suite that exceeds it fails. Under 20 seconds is the target. A suite between
20 and 60 seconds is still green, but the overage must be filed as an
improvement bug against that repo. Add a 90-second timeout to the test
invocation (`go test -timeout 90s`). The backstop deliberately sits above the
hard cap so that it catches a genuinely hung test rather than a merely slow
one.
- **`make test` should use the conditional verbose rerun pattern.** Run tests - **The test command should use the conditional verbose rerun pattern.** Run
without `-v` (verbose) first. If tests fail, automatically rerun with `-v` to tests without `-v` (verbose) first. If tests fail, automatically rerun with
show full output. This keeps CI logs and `docker build` output clean on `-v` to show full output. This keeps CI logs and `docker build` output clean
success (just package/suite summaries) while providing full diagnostic detail on success (just package/suite summaries) while providing full diagnostic
on failure (every test case, every assertion). The general shell pattern: detail on failure (every test case, every assertion). The command lives in the
`test` phase of the `Dockerfile`, since `script/test` builds that phase; the
Makefile form below is the same pattern for any repo-local invocation:
```makefile ```makefile
test: test:
@@ -209,11 +341,26 @@ style conventions are in separate documents:
```makefile ```makefile
test: test:
@go test -timeout 30s -race -cover ./... || \ @go test -count=1 -timeout 90s -race -cover ./... || \
{ echo "--- Rerunning with -v for details ---"; \ { echo "--- Rerunning with -v for details ---"; \
go test -timeout 30s -race -v ./...; exit 1; } go test -count=1 -timeout 90s -race -v ./...; exit 1; }
``` ```
`-count=1` is required on both invocations: it defeats Go's test _result_
cache, so neither run can report a stored pass in place of running the
tests. It leaves the build cache alone, so it costs the runtime of the suite
and no recompilation.
That cache is Go's own, separate from Docker's layer cache. Go stores a
passing result in its cache directory (`GOCACHE`), and when the same tests
run again on unchanged code it prints that result, marked `(cached)`,
without running them. That matters on a developer's machine, where this
target runs and the directory lasts from one run to the next. The `test`
phase of the `Dockerfile` needs no `-count=1`: its base image holds no
result for this repo's tests and nothing before its `go test` step runs a
test, so there is nothing to replay. `--no-cache` (above) is what makes that
step run on an unchanged tree.
Python example: Python example:
```makefile ```makefile
@@ -239,10 +386,89 @@ style conventions are in separate documents:
must be in `.gitignore`. No exceptions. must be in `.gitignore`. No exceptions.
- `.gitignore` should be comprehensive from the start: OS files (`.DS_Store`), - `.gitignore` should be comprehensive from the start: OS files (`.DS_Store`),
editor files (`.swp`, `*~`), language build artifacts, and `node_modules/`. editor files (`.swp`, `*~`), in-repo agent scratch directories (`.claude/`),
Fetch the standard `.gitignore` from `node_modules/`, and the repo's own build outputs. Fetch the standard
`.gitignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.gitignore` when setting up `https://git.eeqj.de/sneak/prompts/raw/branch/main/.gitignore` when setting up
a new repo. a new repo. A repo's `.gitignore` is the standard file followed by the repo's
own entries, such as its binaries; a re-vendor replaces the standard part and
keeps those entries. These patterns are written to `.gitignore`'s own
semantics, in which an unanchored pattern already matches at every depth; they
are not a `.dockerignore` and must not be transplanted into one unmodified.
- **`.dockerignore` does not use `.gitignore` semantics, and copying patterns
across unmodified leaves secrets in the build context.** Docker matches with
`moby/patternmatcher`: `filepath.Match` semantics plus a `**` extension, so
`*` does not cross `/` and a pattern without a leading `**/` is anchored at
the build-context root. A `.dockerignore` listing `.env`, `*.pem` and `*.key`
therefore excludes only the copies at the repository root, while `config/.env`
and `certs/server.key` still reach the context and can land in an image layer
— which is more dangerous than a short file with no secret patterns at all,
because it reads as solved and stops anyone looking. Give every
depth-independent pattern the `**/` prefix and leave only genuinely
root-anchored entries unprefixed: `.claude`, and the repo's own host-built
binary, written `/myapp` and never `**/myapp`, which would also match
`cmd/myapp/` and delete the package directory from the context. Matching is
case-sensitive, and an ALL-CAPS twin per pattern still misses `Server.Key`, so
secret names use character ranges — `**/*.[kK][eE][yY]`, `**/*.[pP][eE][mM]`,
and likewise for `.envrc` and the extensionless SSH keys. Where such a pattern
also catches something the build needs, re-include it with a negation
(`!docs/example.env`); deleting the pattern reopens the exposure for every
other file it covers. Fetch the standard `.dockerignore` from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.dockerignore` and extend
it with the repo's own artifacts.
- **In-repo agent scratch belongs in both files, written to each file's own
semantics.** `.claude/` holds one worktree per in-flight agent — an entire
additional checkout of the repo — so under `COPY . .` the build context
inflates by a multiple of the repo and another session's unreviewed work can
be copied into an image layer. In `.gitignore` the entry is `.claude/`,
unanchored. In `.dockerignore` it is `.claude`, anchored and with **no** `**/`
prefix, because the prefixed form would also delete any nested directory of
that name from the build. Anchoring carries a known gap that the canonical
`.dockerignore` states in its own comment, since consuming repos receive the
file and not the tracker: the directory is created in the agent's working
directory, so a repo running agents in subdirectories still ships
`services/api/.claude/` and must add its own anchored entry there.
- **A plain `docker build .` of a clone stamps the version that
`git describe --tags --always` gives**, derived from the `.git` in the build
context as the canonical `Dockerfile` above shows. Without its failure check,
a missing `git` or an unreadable checkout would leave `-X main.Version=` empty
and the build would still exit 0. `script/docker` and `script/cibuild` pass
the version they compute on the host; it takes precedence. They do this
byte-identically across repos:
```sh
# The version and the tag each get their own line: a failing command
# substitution inside an argument does not trip `set -e`, so the inline
# form degrades to an empty constant.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
tag="$(script/projectname)"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$tag" .
```
`--always` makes an untagged repo yield an abbreviated commit hash rather
than failing, and the `[ -n "$version" ]` line is the single place the
fallback is applied — a live check that fires on a build from an export with
no `.git` and on a repository with no commits yet. Do not fold it into the
substitution as `|| echo unknown`, which makes the guard unreachable. The
Dockerfile's side is `ARG VERSION` in the stage that compiles, declared
there because `ARG` is stage-scoped; passing `VERSION` to a repo whose
Dockerfile declares no such `ARG` is ignored and costs nothing, which is why
the scripts stay byte-identical. One consequence for CI: the standard
checkout action clones shallow and fetches no tags, so a repo that embeds a
tag-derived version must set `fetch-depth: 0` on its checkout step.
- **Verify `.dockerignore` by enumerating the image, not by reading the
patterns.** Plant files at the root _and_ at least two directories deep, build
a probe image that does `COPY . .`, and list what actually landed
(`docker run --rm --entrypoint find IMAGE /app`). The `transferring context`
size is not a substitute: a nested secret is a few bytes, and BuildKit
transfers only the delta from the previous build.
- **No build artifacts in version control.** Code-derived data (compiled - **No build artifacts in version control.** Code-derived data (compiled
bundles, minified output, generated assets) must never be committed to the bundles, minified output, generated assets) must never be committed to the
@@ -258,9 +484,56 @@ style conventions are in separate documents:
- Make all changes on a feature branch. You can do whatever you want on a - Make all changes on a feature branch. You can do whatever you want on a
feature branch. feature branch.
- `.golangci.yml` is standardized and must _NEVER_ be modified by an agent, only - `.golangci.yml` is standardized. The vendored copy in a consuming repo must
manually by the user. Fetch from _NEVER_ be modified by an agent: fetch it from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml`. `https://git.eeqj.de/sneak/prompts/raw/branch/main/.golangci.yml` and keep it
byte-identical, so that no repo can quietly loosen its own linting. Linter
configuration changes are made to the canonical copy in the `prompts` repo and
reach consuming repos by re-vendoring; an agent may open a PR against
canonical, which only the user merges. One list is exempt from byte-identity,
because it cannot be written once for every repo: the `deny` list of the
`test-support` depguard rule, where a repo names its own test-support packages
by full import path. A repo adds entries there and changes nothing else, and a
re-vendor carries its entries forward. The canonical golangci-lint version is
v2.14.0 (released 2026-09-24), pinned as the digest of the lint phase's base
image
(`golangci/golangci-lint@sha256:ad862ba6b3798cbe0fd9fd7408d498fd74fbd2623a92406b2fd3898faf0bf98f`,
which reports `2.14.0 built with go1.27.0 from 114493f9`). A module's `go`
directive must not name a newer Go minor version than the one golangci-lint
was built with, or golangci-lint refuses to lint it: this release lints
`go 1.27.1` but not `go 1.28`. That digest is the only pin, since no repo
installs golangci-lint on the host. A repo sets the lint phase digest to the
one named here and re-vendors `.golangci.yml` in the same commit, whichever of
the two prompted the change: the canonical copy can name linters that an older
golangci-lint rejects, and a newer golangci-lint can add linters that
`default: all` switches on until the canonical copy disables them.
- **`script/bootstrap` installs a pinned tool by comparing versions, never by
testing presence.** An `if ! command -v <tool>; then install; fi` guard tests
`PATH` only, so on an already-provisioned machine the pin is inert and a
version bump is a silent no-op — while the Dockerfile, installing into a clean
image, gets the pinned version, so a local `make check` and `make docker` can
disagree about what the tool even is. The canonical form:
- compares the installed version against the pin over the **whole** version
token; a parser that stops at the first `-` reports `2.12.2` for a host
running `2.12.2-rc1` and skips the install;
- treats absent, non-zero, empty or unrecognised `--version` output as a
mismatch, so the failure direction is a redundant install and never a
skipped one;
- after installing, re-resolves the binary the way callers do — `hash -r`,
then through `PATH`, not through the directory the installer wrote to —
and fails naming the resolved path, since an install that a shadowing
binary hides succeeds while changing nothing any caller sees;
- is actually called, and prints the version on both success paths: a
function defined and never invoked has the same exit status and the same
empty output as one that worked.
Keep it POSIX sh: no arrays, no `[[`, no `grep -P`.
A Go tool a repo needs on the host is installed with `go install` pinned to
a commit hash (`go install <package>@<commit hash>`). It is never tracked as
a `go.mod` tool dependency or through a `tools.go` file, either of which
pulls the tool's own dependencies into the repo's `go.mod` and `go.sum`.
- When pinning images or packages by hash, add a comment above the reference - When pinning images or packages by hash, add a comment above the reference
with the version and date (YYYY-MM-DD). with the version and date (YYYY-MM-DD).
@@ -371,15 +644,21 @@ style conventions are in separate documents:
Never edit existing migrations after release. Never edit existing migrations after release.
- All repos should have an `.editorconfig` enforcing the project's indentation - All repos should have an `.editorconfig` enforcing the project's indentation
settings. settings: the standard file from
`https://git.eeqj.de/sneak/prompts/raw/branch/main/.editorconfig`, which sets
tabs for `Makefile` and Go files, followed by the repo's own sections, such as
one for another language it uses. A re-vendor replaces the standard part and
keeps those sections.
- Avoid putting files in the repo root unless necessary. Root should contain - Avoid putting files in the repo root unless necessary. Root should contain
only project-level config files (`README.md`, `Makefile`, `Dockerfile`, only project-level config files (`README.md`, `AGENTS.md`, `Makefile`,
`LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`, and `Dockerfile`, `LICENSE`, `.gitignore`, `.editorconfig`, `REPO_POLICIES.md`,
language-specific config). Everything else goes in a subdirectory. Canonical and language-specific config). Everything else goes in a subdirectory.
subdirectory names: Canonical subdirectory names:
- `bin/` — executable scripts and tools - `bin/` — executable scripts and tools
- `cmd/` — Go command entrypoints - `cmd/` — Go command entrypoints; thin only: one `main.go` per binary whose
body is a single call into `internal/` or `pkg/`, no project logic in
`cmd/`
- `configs/` — configuration templates and examples - `configs/` — configuration templates and examples
- `deploy/` — deployment manifests (k8s, compose, terraform) - `deploy/` — deployment manifests (k8s, compose, terraform)
- `docs/` — documentation and markdown (README.md stays in root) - `docs/` — documentation and markdown (README.md stays in root)
@@ -406,3 +685,7 @@ style conventions are in separate documents:
- Go: `go.mod`, `go.sum`, `.golangci.yml` - Go: `go.mod`, `go.sum`, `.golangci.yml`
- JS: `package.json`, `yarn.lock`, `.prettierrc`, `.prettierignore` - JS: `package.json`, `yarn.lock`, `.prettierrc`, `.prettierignore`
- Python: `pyproject.toml` - Python: `pyproject.toml`
- Guidance for coding agents lives in one `AGENTS.md` at the repository root. It
is never committed under a file or directory named after one agent tool, such
as `CLAUDE.md` or `.claude/`, and never split into separate memory files.
+40 -24
View File
@@ -1,12 +1,12 @@
# Workflow # Workflow
* branch (from `main`) - branch (from `main`)
* do the work in Next Step - do the work in Next Step
* move Next Step to the top of Completed Steps - move Next Step to the top of Completed Steps
* move the top item of Future Steps into Next Step - move the top item of Future Steps into Next Step
* commit (`TODO.md` changes in the same commit as the work) - commit (`TODO.md` changes in the same commit as the work)
* merge to `main` if the branch is not protected, otherwise open a PR - merge to `main` if the branch is not protected, otherwise open a PR
* push - push
# Status # Status
@@ -14,30 +14,46 @@ pre-1.0
# Next Step # Next Step
Bring the repo into policy compliance in one commit: add .gitignore, Expand the `internal/bsdaily` tests beyond the compilation smoke test: unit
.dockerignore, .editorconfig, and .golangci.yml. Verify `make check` stays tests for the extraction, verification, and atomic-publish paths.
green with the new lint config.
# Completed Steps # Completed Steps
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints, - 2026-10-06: Moved the command line (flags, the rules for combining them, date
Makefile shims, README Entrypoints section parsing) from `cmd/bsdaily` into `internal/cli`, with unit tests for the flag
- 2026-06-28: Fixed errcheck lint failures; added compilation smoke test; rules and the dates; `cmd/bsdaily/main.go` is now a single call into it
tidied go.mod. (https://git.eeqj.de/sneak/bsdaily/issues/8).
- 2026-06-28: Added repo scaffolding: README, LICENSE, Makefile, - 2026-10-06: A plain `docker build .` and a host `make` build stamp the git tag
Dockerfile, REPO_POLICIES.md, and Gitea CI. or short commit into the binary, which `bsdaily` logs on the first line of
- 2026-02-12: Fixed SQLite database locking by removing parallel every run and prints with `--version`
processing; fixed Linux build via golang.org/x/sys/unix Fadvise. (https://git.eeqj.de/sneak/bsdaily/issues/4).
- 2026-02-12: Optimized file copy for large databases; moved temp - 2026-10-06: Added the canonical `.golangci.yml`, moved the lint phase to
directory to NVMe scratch storage. golangci-lint v2.14.0, and fixed the code to pass it
(https://git.eeqj.de/sneak/bsdaily/issues/6).
- 2026-10-06: Formatted Markdown with prettier: `script/fmt` writes and
`script/fmt-check` checks every Markdown file; prettier pinned in
`package.json` and `yarn.lock`, installed by `script/bootstrap`, which
installs node and yarn at pinned versions when absent and otherwise uses the
ones already installed; existing Markdown reformatted.
- 2026-10-05: Brought the repo up to the standard layout: canonical
`.gitignore`, `.dockerignore` and `.editorconfig`; `lint` and `test` phases in
the `Dockerfile`, built by `script/lint` and `script/test`; canonical
`script/cibuild`, `script/docker` and CI workflow; no linter installed on the
host; re-vendored `REPO_POLICIES.md`.
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints, Makefile
shims, README Entrypoints section
- 2026-06-28: Fixed errcheck lint failures; added compilation smoke test; tidied
go.mod.
- 2026-06-28: Added repo scaffolding: README, LICENSE, Makefile, Dockerfile,
REPO_POLICIES.md, and Gitea CI.
- 2026-02-12: Fixed SQLite database locking by removing parallel processing;
fixed Linux build via golang.org/x/sys/unix Fadvise.
- 2026-02-12: Optimized file copy for large databases; moved temp directory to
NVMe scratch storage.
- 2026-02-11: Added date range support. - 2026-02-11: Added date range support.
- 2026-02-09: Initial implementation: single-day extraction, specific-date - 2026-02-09: Initial implementation: single-day extraction, specific-date
targeting, faster pruning of throwaway database copies. targeting, faster pruning of throwaway database copies.
# Future Steps # Future Steps
- Add .gitignore, .dockerignore, .editorconfig, .golangci.yml (the Next
Step).
- Expand tests beyond the compilation smoke test: unit tests for the
extraction, verification, and atomic-publish paths.
- Cut a first SemVer release once compliance and test coverage land. - Cut a first SemVer release once compliance and test coverage land.
+11 -75
View File
@@ -1,84 +1,20 @@
// Package main is the bsdaily command. It extracts one day, or a range
// of days, from the latest daily snapshot.
package main package main
import ( import (
"fmt"
"log/slog"
"os" "os"
"time"
"git.eeqj.de/sneak/bsdaily/internal/bsdaily" "git.eeqj.de/sneak/bsdaily/internal/cli"
"github.com/spf13/cobra"
) )
// Version is the git tag or short commit, set at link time with
// -X main.Version=... by the Dockerfile and the Makefile. A build that sets
// nothing, or sets it empty, reports dev.
//
//nolint:gochecknoglobals // -X can only set a package-level variable
var Version string
func main() { func main() {
logger := slog.New(slog.NewTextHandler(os.Stderr, &slog.HandlerOptions{ os.Exit(cli.Main(Version))
Level: slog.LevelInfo,
}))
slog.SetDefault(logger)
var dateFlag string
var fromFlag string
var toFlag string
rootCmd := &cobra.Command{
Use: "bsdaily",
Short: "Extract a single day's data from the latest daily snapshot",
SilenceUsage: true,
RunE: func(cmd *cobra.Command, args []string) error {
hasDate := dateFlag != ""
hasFrom := fromFlag != ""
hasTo := toFlag != ""
// Validate mutual exclusivity
if hasDate && (hasFrom || hasTo) {
return fmt.Errorf("--date and --from/--to are mutually exclusive")
}
if hasFrom != hasTo {
if hasFrom {
return fmt.Errorf("--from requires --to")
}
return fmt.Errorf("--to requires --from")
}
var targetDates []time.Time
if hasDate {
t, err := time.Parse("2006-01-02", dateFlag)
if err != nil {
return fmt.Errorf("invalid --date %q (expected YYYY-MM-DD): %w", dateFlag, err)
}
targetDates = []time.Time{t}
} else if hasFrom {
from, err := time.Parse("2006-01-02", fromFlag)
if err != nil {
return fmt.Errorf("invalid --from %q (expected YYYY-MM-DD): %w", fromFlag, err)
}
to, err := time.Parse("2006-01-02", toFlag)
if err != nil {
return fmt.Errorf("invalid --to %q (expected YYYY-MM-DD): %w", toFlag, err)
}
if from.After(to) {
return fmt.Errorf("--from %s is after --to %s", fromFlag, toFlag)
}
for d := from; !d.After(to); d = d.AddDate(0, 0, 1) {
targetDates = append(targetDates, d)
}
}
// else: targetDates remains nil → Run() defaults to snapshot date minus one
if err := bsdaily.Run(targetDates); err != nil {
return err
}
slog.Info("completed successfully")
return nil
},
}
rootCmd.Flags().StringVarP(&dateFlag, "date", "d", "", "target date to extract (YYYY-MM-DD); defaults to snapshot date minus one day")
rootCmd.Flags().StringVar(&fromFlag, "from", "", "start of date range to extract (YYYY-MM-DD, inclusive); use with --to")
rootCmd.Flags().StringVar(&toFlag, "to", "", "end of date range to extract (YYYY-MM-DD, inclusive); use with --from")
if err := rootCmd.Execute(); err != nil {
os.Exit(1)
}
} }
+17 -6
View File
@@ -1,21 +1,32 @@
package bsdaily package bsdaily_test
import "testing" import (
"testing"
"git.eeqj.de/sneak/bsdaily/internal/bsdaily"
)
// TestCompiles is a minimal smoke test that references the package's exported // TestCompiles is a minimal smoke test that references the package's exported
// surface so that `go test` fails if the package stops compiling. It does not // surface so that `go test` fails if the package stops compiling. It does not
// touch the filesystem or any of the hard-coded production paths. // touch the filesystem or any of the hard-coded production paths.
func TestCompiles(t *testing.T) { func TestCompiles(t *testing.T) {
if DBFilename == "" || WALFilename == "" || SHMFilename == "" { t.Parallel()
if bsdaily.DBFilename == "" || bsdaily.WALFilename == "" ||
bsdaily.SHMFilename == "" {
t.Fatal("expected database filename constants to be set") t.Fatal("expected database filename constants to be set")
} }
if SnapshotBase == "" || TmpBase == "" || DailiesBase == "" {
if bsdaily.SnapshotBase == "" || bsdaily.TmpBase == "" ||
bsdaily.DailiesBase == "" {
t.Fatal("expected base path constants to be set") t.Fatal("expected base path constants to be set")
} }
if MinTmpFreeBytes == 0 || MinDailiesFreeBytes == 0 {
if bsdaily.MinTmpFreeBytes == 0 || bsdaily.MinDailiesFreeBytes == 0 {
t.Fatal("expected free-space thresholds to be set") t.Fatal("expected free-space thresholds to be set")
} }
if ErrNoPosts == nil {
if bsdaily.ErrNoPosts == nil {
t.Fatal("expected ErrNoPosts sentinel to be set") t.Fatal("expected ErrNoPosts sentinel to be set")
} }
} }
+6 -1
View File
@@ -1,7 +1,11 @@
// Package bsdaily extracts single days of firehose data from the latest
// daily ZFS snapshot into zstd-compressed SQL dumps.
package bsdaily package bsdaily
import "regexp" import "regexp"
// Paths, file names and tuning for a run. They are set for one
// production host; see the README.
const ( const (
SnapshotBase = "/srv/berlin.sneak.fs.blueskyarchive/.zfs/snapshot" SnapshotBase = "/srv/berlin.sneak.fs.blueskyarchive/.zfs/snapshot"
TmpBase = "/srv/storage/tmp" TmpBase = "/srv/storage/tmp"
@@ -27,4 +31,5 @@ const (
verificationHeadLines = 20 verificationHeadLines = 20
) )
var snapshotPattern = regexp.MustCompile(`^zfs-auto-snap_daily-(\d{4}-\d{2}-\d{2})-\d{4}$`) var snapshotPattern = regexp.MustCompile(
`^zfs-auto-snap_daily-(\d{4}-\d{2}-\d{2})-\d{4}$`)
+28 -8
View File
@@ -1,6 +1,7 @@
package bsdaily package bsdaily
import ( import (
"errors"
"fmt" "fmt"
"io" "io"
"log/slog" "log/slog"
@@ -9,20 +10,30 @@ import (
) )
const ( const (
copyBufferSize = 256 * 1024 * 1024 // 256MB buffer for large file copies from fast storage // 256MB buffer for large file copies from fast storage
copyBufferSize = 256 * 1024 * 1024
oneMB = 1024 * 1024
oneGB = 1024 * 1024 * 1024 oneGB = 1024 * 1024 * 1024
) )
var errShortCopy = errors.New("short copy")
// CopyFile copies src to dst through a large buffer, pre-allocating dst
// and syncing it to disk before returning.
func CopyFile(src, dst string) (err error) { func CopyFile(src, dst string) (err error) {
startTime := time.Now() startTime := time.Now()
slog.Info("copying file", "src", src, "dst", dst) slog.Info("copying file", "src", src, "dst", dst)
//nolint:gosec // src is a path this package built
srcFile, err := os.Open(src) srcFile, err := os.Open(src)
if err != nil { if err != nil {
return fmt.Errorf("opening source %s: %w", src, err) return fmt.Errorf("opening source %s: %w", src, err)
} }
defer func() { defer func() {
if cerr := srcFile.Close(); cerr != nil { cerr := srcFile.Close()
if cerr != nil {
slog.Warn("failed to close source file", "src", src, "error", cerr) slog.Warn("failed to close source file", "src", src, "error", cerr)
} }
}() }()
@@ -37,40 +48,49 @@ func CopyFile(src, dst string) (err error) {
applyFileAdvice(srcFile, srcInfo.Size()) applyFileAdvice(srcFile, srcInfo.Size())
} }
//nolint:gosec // dst is a path this package built
dstFile, err := os.Create(dst) dstFile, err := os.Create(dst)
if err != nil { if err != nil {
return fmt.Errorf("creating destination %s: %w", dst, err) return fmt.Errorf("creating destination %s: %w", dst, err)
} }
defer func() { defer func() {
if cerr := dstFile.Close(); cerr != nil && err == nil { cerr := dstFile.Close()
if cerr != nil && err == nil {
err = fmt.Errorf("closing destination %s: %w", dst, cerr) err = fmt.Errorf("closing destination %s: %w", dst, cerr)
} }
}() }()
// Pre-allocate space for the destination file to avoid fragmentation // Pre-allocate space for the destination file to avoid fragmentation
if err := dstFile.Truncate(srcInfo.Size()); err != nil { truncErr := dstFile.Truncate(srcInfo.Size())
slog.Warn("failed to pre-allocate destination file", "error", err) if truncErr != nil {
slog.Warn("failed to pre-allocate destination file", "error", truncErr)
} }
// Use a much larger buffer for NVMe-speed copies // Use a much larger buffer for NVMe-speed copies
buf := make([]byte, copyBufferSize) buf := make([]byte, copyBufferSize)
written, err := io.CopyBuffer(dstFile, srcFile, buf) written, err := io.CopyBuffer(dstFile, srcFile, buf)
if err != nil { if err != nil {
return fmt.Errorf("copying data: %w", err) return fmt.Errorf("copying data: %w", err)
} }
if written != srcInfo.Size() { if written != srcInfo.Size() {
return fmt.Errorf("short copy: wrote %d bytes, expected %d", written, srcInfo.Size()) return fmt.Errorf("%w: wrote %d bytes, expected %d",
errShortCopy, written, srcInfo.Size())
} }
if err := dstFile.Sync(); err != nil { err = dstFile.Sync()
if err != nil {
return fmt.Errorf("syncing destination %s: %w", dst, err) return fmt.Errorf("syncing destination %s: %w", dst, err)
} }
elapsed := time.Since(startTime) elapsed := time.Since(startTime)
throughputMBps := float64(written) / elapsed.Seconds() / (1024 * 1024) throughputMBps := float64(written) / elapsed.Seconds() / oneMB
slog.Info("file copied", "dst", dst, "bytes", written, slog.Info("file copied", "dst", dst, "bytes", written,
"elapsed", elapsed.Round(time.Millisecond), "elapsed", elapsed.Round(time.Millisecond),
"throughput_mbps", fmt.Sprintf("%.1f", throughputMBps)) "throughput_mbps", fmt.Sprintf("%.1f", throughputMBps))
return nil return nil
} }
+3 -4
View File
@@ -9,8 +9,7 @@ import (
func applyFileAdvice(file *os.File, size int64) { func applyFileAdvice(file *os.File, size int64) {
fd := int(file.Fd()) fd := int(file.Fd())
// POSIX_FADV_SEQUENTIAL = 2 _ = unix.Fadvise(fd, 0, size, unix.FADV_SEQUENTIAL)
_ = unix.Fadvise(fd, 0, size, 2) // Prefetch the file into the page cache
// POSIX_FADV_WILLNEED = 3 - prefetch file into cache _ = unix.Fadvise(fd, 0, size, unix.FADV_WILLNEED)
_ = unix.Fadvise(fd, 0, size, 3)
} }
+16 -3
View File
@@ -1,26 +1,39 @@
package bsdaily package bsdaily
import ( import (
"errors"
"fmt" "fmt"
"log/slog" "log/slog"
"golang.org/x/sys/unix" "golang.org/x/sys/unix"
) )
var errInsufficientSpace = errors.New("insufficient disk space")
// CheckFreeSpace returns an error when the filesystem holding path has
// fewer than minBytes bytes available. label names the location in the
// log line and the error.
func CheckFreeSpace(path string, minBytes uint64, label string) error { func CheckFreeSpace(path string, minBytes uint64, label string) error {
var stat unix.Statfs_t var stat unix.Statfs_t
if err := unix.Statfs(path, &stat); err != nil {
err := unix.Statfs(path, &stat)
if err != nil {
return fmt.Errorf("statfs %s (%s): %w", path, label, err) return fmt.Errorf("statfs %s (%s): %w", path, label, err)
} }
//nolint:gosec,unconvert // Bsize is never negative; Bavail is signed on FreeBSD
free := uint64(stat.Bavail) * uint64(stat.Bsize) free := uint64(stat.Bavail) * uint64(stat.Bsize)
freeGB := float64(free) / float64(bytesPerGB) freeGB := float64(free) / float64(bytesPerGB)
minGB := float64(minBytes) / float64(bytesPerGB) minGB := float64(minBytes) / float64(bytesPerGB)
slog.Info("disk space check", "label", label, "path", path, slog.Info("disk space check", "label", label, "path", path,
"free_gb", fmt.Sprintf("%.1f", freeGB), "free_gb", fmt.Sprintf("%.1f", freeGB),
"required_gb", fmt.Sprintf("%.1f", minGB)) "required_gb", fmt.Sprintf("%.1f", minGB))
if free < minBytes { if free < minBytes {
return fmt.Errorf("insufficient disk space on %s (%s): %.1f GB free, need %.1f GB", return fmt.Errorf("%w on %s (%s): %.1f GB free, need %.1f GB",
path, label, freeGB, minGB) errInsufficientSpace, path, label, freeGB, minGB)
} }
return nil return nil
} }
+74 -35
View File
@@ -1,6 +1,8 @@
package bsdaily package bsdaily
import ( import (
"context"
"errors"
"fmt" "fmt"
"log/slog" "log/slog"
"os" "os"
@@ -9,61 +11,44 @@ import (
"strings" "strings"
) )
var errEmptyOutput = errors.New("compressed output is empty")
// DumpAndCompress writes a `sqlite3 .dump` of the database at dbPath,
// compressed by zstdmt, to outputPath.
func DumpAndCompress(dbPath, outputPath string) (err error) { func DumpAndCompress(dbPath, outputPath string) (err error) {
for _, tool := range []string{"sqlite3", "zstdmt"} { for _, tool := range []string{"sqlite3", "zstdmt"} {
if _, err := exec.LookPath(tool); err != nil { _, err = exec.LookPath(tool)
if err != nil {
return fmt.Errorf("required tool %q not found in PATH: %w", tool, err) return fmt.Errorf("required tool %q not found in PATH: %w", tool, err)
} }
} }
if err := CheckFreeSpace(filepath.Dir(outputPath), MinDailiesFreeBytes, "dailiesBase (pre-dump)"); err != nil { err = CheckFreeSpace(filepath.Dir(outputPath), MinDailiesFreeBytes,
"dailiesBase (pre-dump)")
if err != nil {
return err return err
} }
//nolint:gosec // outputPath is a path this package built
outFile, err := os.Create(outputPath) outFile, err := os.Create(outputPath)
if err != nil { if err != nil {
return fmt.Errorf("creating output file: %w", err) return fmt.Errorf("creating output file: %w", err)
} }
defer func() { defer func() {
if cerr := outFile.Close(); cerr != nil && err == nil { cerr := outFile.Close()
if cerr != nil && err == nil {
err = fmt.Errorf("closing output: %w", cerr) err = fmt.Errorf("closing output: %w", cerr)
} }
}() }()
// Dump all tables but use INSERT OR IGNORE for mergeable imports err = runDumpPipeline(context.Background(), dbPath, outFile)
// This preserves all data while allowing multiple dumps to be merged
// Users should import with: zstdcat *.sql.zst | sed 's/INSERT INTO/INSERT OR IGNORE INTO/g' | sqlite3 merged.db
dumpCmd := exec.Command("sqlite3", dbPath, ".dump")
zstdCmd := exec.Command("zstdmt", fmt.Sprintf("-%d", zstdCompressionLevel))
pipe, err := dumpCmd.StdoutPipe()
if err != nil { if err != nil {
return fmt.Errorf("creating dump stdout pipe: %w", err) return err
}
zstdCmd.Stdin = pipe
zstdCmd.Stdout = outFile
var dumpStderr, zstdStderr strings.Builder
dumpCmd.Stderr = &dumpStderr
zstdCmd.Stderr = &zstdStderr
slog.Info("starting sqlite3 dump and zstdmt compression")
if err := zstdCmd.Start(); err != nil {
return fmt.Errorf("starting zstdmt: %w", err)
}
if err := dumpCmd.Start(); err != nil {
return fmt.Errorf("starting sqlite3 dump: %w", err)
} }
if err := dumpCmd.Wait(); err != nil { err = outFile.Sync()
return fmt.Errorf("sqlite3 dump failed: %w; stderr: %s", err, dumpStderr.String()) if err != nil {
}
if err := zstdCmd.Wait(); err != nil {
return fmt.Errorf("zstdmt failed: %w; stderr: %s", err, zstdStderr.String())
}
if err := outFile.Sync(); err != nil {
return fmt.Errorf("syncing output: %w", err) return fmt.Errorf("syncing output: %w", err)
} }
@@ -71,12 +56,66 @@ func DumpAndCompress(dbPath, outputPath string) (err error) {
if err != nil { if err != nil {
return fmt.Errorf("stat output: %w", err) return fmt.Errorf("stat output: %w", err)
} }
const bytesPerMB = 1024 * 1024 const bytesPerMB = 1024 * 1024
slog.Info("compressed output written", "path", outputPath, slog.Info("compressed output written", "path", outputPath,
"size_bytes", info.Size(), "size_mb", info.Size()/bytesPerMB) "size_bytes", info.Size(), "size_mb", info.Size()/bytesPerMB)
if info.Size() == 0 { if info.Size() == 0 {
return fmt.Errorf("compressed output is empty") return errEmptyOutput
}
return nil
}
// runDumpPipeline runs `sqlite3 dbPath .dump | zstdmt` with the
// compressed stream going to outFile, and waits for both to finish.
func runDumpPipeline(ctx context.Context, dbPath string, outFile *os.File) error {
// The dump holds plain INSERT INTO statements. merge_daily_dumps.sh
// rewrites them to INSERT OR IGNORE INTO so that several dumps can be
// merged into one database.
//nolint:gosec // dbPath is a scratch file this package created
dumpCmd := exec.CommandContext(ctx, "sqlite3", dbPath, ".dump")
//nolint:gosec // the argument is built from a constant
zstdCmd := exec.CommandContext(ctx, "zstdmt",
fmt.Sprintf("-%d", zstdCompressionLevel))
pipe, err := dumpCmd.StdoutPipe()
if err != nil {
return fmt.Errorf("creating dump stdout pipe: %w", err)
}
zstdCmd.Stdin = pipe
zstdCmd.Stdout = outFile
var dumpStderr, zstdStderr strings.Builder
dumpCmd.Stderr = &dumpStderr
zstdCmd.Stderr = &zstdStderr
slog.Info("starting sqlite3 dump and zstdmt compression")
err = zstdCmd.Start()
if err != nil {
return fmt.Errorf("starting zstdmt: %w", err)
}
err = dumpCmd.Start()
if err != nil {
return fmt.Errorf("starting sqlite3 dump: %w", err)
}
err = dumpCmd.Wait()
if err != nil {
return fmt.Errorf("sqlite3 dump failed: %w; stderr: %s",
err, dumpStderr.String())
}
err = zstdCmd.Wait()
if err != nil {
return fmt.Errorf("zstdmt failed: %w; stderr: %s",
err, zstdStderr.String())
} }
return nil return nil
+251 -124
View File
@@ -1,22 +1,29 @@
package bsdaily package bsdaily
import ( import (
"context"
"database/sql" "database/sql"
"errors" "errors"
"fmt" "fmt"
"log/slog" "log/slog"
"time" "time"
// Registers the "sqlite" driver with database/sql.
_ "modernc.org/sqlite" _ "modernc.org/sqlite"
) )
// ErrNoPosts is returned by ExtractDay when the source holds no posts for
// the target day.
var ErrNoPosts = errors.New("no posts found for target day") var ErrNoPosts = errors.New("no posts found for target day")
var errPostCountMismatch = errors.New("post count mismatch")
// ExtractDay opens a new empty database at dstDBPath, attaches srcDBPath, // ExtractDay opens a new empty database at dstDBPath, attaches srcDBPath,
// and copies only the target day's data into it. This is much faster than // and copies only the target day's data into it. This is much faster than
// pruning a full copy because it only reads/writes the small slice of data // pruning a full copy because it only reads/writes the small slice of data
// being kept. // being kept.
func ExtractDay(srcDBPath, dstDBPath string, targetDay time.Time) error { func ExtractDay(srcDBPath, dstDBPath string, targetDay time.Time) error {
ctx := context.Background()
dayStart := targetDay.Format("2006-01-02") + "T00:00:00" dayStart := targetDay.Format("2006-01-02") + "T00:00:00"
dayEnd := targetDay.AddDate(0, 0, 1).Format("2006-01-02") + "T00:00:00" dayEnd := targetDay.AddDate(0, 0, 1).Format("2006-01-02") + "T00:00:00"
@@ -24,162 +31,282 @@ func ExtractDay(srcDBPath, dstDBPath string, targetDay time.Time) error {
// Maximum performance pragmas - we don't care about crash safety for temp files // Maximum performance pragmas - we don't care about crash safety for temp files
// Use WAL mode for the source attachment to avoid locking issues // Use WAL mode for the source attachment to avoid locking issues
pragmas := fmt.Sprintf("?_pragma=journal_mode(WAL)&_pragma=synchronous(OFF)&_pragma=cache_size(%d)&_pragma=foreign_keys(OFF)&_pragma=temp_store(MEMORY)&_pragma=busy_timeout(5000)", sqliteCacheSizeKB) pragmas := fmt.Sprintf("?_pragma=journal_mode(WAL)"+
"&_pragma=synchronous(OFF)&_pragma=cache_size(%d)"+
"&_pragma=foreign_keys(OFF)&_pragma=temp_store(MEMORY)"+
"&_pragma=busy_timeout(5000)", sqliteCacheSizeKB)
db, err := sql.Open("sqlite", dstDBPath+pragmas) db, err := sql.Open("sqlite", dstDBPath+pragmas)
if err != nil { if err != nil {
return fmt.Errorf("opening destination database: %w", err) return fmt.Errorf("opening destination database: %w", err)
} }
defer func() { defer func() {
if cerr := db.Close(); cerr != nil { cerr := db.Close()
slog.Warn("failed to close destination database", "path", dstDBPath, "error", cerr) if cerr != nil {
slog.Warn("failed to close destination database",
"path", dstDBPath, "error", cerr)
} }
}() }()
// Attach source database // Attach source database
if _, err := db.Exec("ATTACH DATABASE ? AS src", srcDBPath); err != nil { _, err = db.ExecContext(ctx, "ATTACH DATABASE ? AS src", srcDBPath)
if err != nil {
return fmt.Errorf("attaching source database: %w", err) return fmt.Errorf("attaching source database: %w", err)
} }
// Copy table DDL from source err = createTables(ctx, db)
slog.Info("copying table DDL from source")
rows, err := db.Query("SELECT sql FROM src.sqlite_master WHERE type='table' AND name NOT LIKE 'sqlite_%' ORDER BY name")
if err != nil { if err != nil {
return fmt.Errorf("reading source schema: %w", err) return err
}
defer func() {
if cerr := rows.Close(); cerr != nil {
slog.Warn("failed to close schema rows", "error", cerr)
}
}()
var ddlStatements []string
for rows.Next() {
var ddl string
if err := rows.Scan(&ddl); err != nil {
return fmt.Errorf("scanning DDL: %w", err)
}
ddlStatements = append(ddlStatements, ddl)
}
if err := rows.Err(); err != nil {
return fmt.Errorf("iterating DDL rows: %w", err)
} }
for _, ddl := range ddlStatements { postCount, err := insertDayRows(ctx, db, targetDay, dayStart, dayEnd)
if _, err := db.Exec(ddl); err != nil {
return fmt.Errorf("creating table: %w\nDDL: %s", err, ddl)
}
}
// Begin transaction for bulk inserts
tx, err := db.Begin()
if err != nil { if err != nil {
return fmt.Errorf("beginning transaction: %w", err) return err
} }
defer func() {
if err != nil {
if rerr := tx.Rollback(); rerr != nil && !errors.Is(rerr, sql.ErrTxDone) {
slog.Warn("failed to roll back transaction", "error", rerr)
}
}
}()
// Insert target day's data err = createIndexes(ctx, db)
slog.Info("inserting posts for target day")
result, err := tx.Exec("INSERT INTO posts SELECT * FROM src.posts WHERE timestamp >= ? AND timestamp < ?", dayStart, dayEnd)
if err != nil { if err != nil {
return fmt.Errorf("inserting posts: %w", err) return err
}
postCount, _ := result.RowsAffected()
slog.Info("inserted posts", "count", postCount)
if postCount == 0 {
return fmt.Errorf("%w %s - aborting to avoid producing empty output",
ErrNoPosts, targetDay.Format("2006-01-02"))
}
slog.Info("inserting junction and lookup tables")
if _, err := tx.Exec("INSERT INTO posts_hashtags SELECT * FROM src.posts_hashtags WHERE post_id IN (SELECT id FROM posts)"); err != nil {
return fmt.Errorf("inserting posts_hashtags: %w", err)
}
if _, err := tx.Exec("INSERT INTO posts_urls SELECT * FROM src.posts_urls WHERE post_id IN (SELECT id FROM posts)"); err != nil {
return fmt.Errorf("inserting posts_urls: %w", err)
}
if _, err := tx.Exec("INSERT INTO hashtags SELECT * FROM src.hashtags WHERE id IN (SELECT hashtag_id FROM posts_hashtags)"); err != nil {
return fmt.Errorf("inserting hashtags: %w", err)
}
if _, err := tx.Exec("INSERT INTO urls SELECT * FROM src.urls WHERE id IN (SELECT url_id FROM posts_urls)"); err != nil {
return fmt.Errorf("inserting urls: %w", err)
}
if _, err := tx.Exec("INSERT INTO users SELECT * FROM src.users WHERE did IN (SELECT user_did FROM posts)"); err != nil {
return fmt.Errorf("inserting users: %w", err)
}
// Check if media table exists in source and copy if present
var mediaTableExists int
if err := tx.QueryRow("SELECT COUNT(*) FROM src.sqlite_master WHERE type='table' AND name='media'").Scan(&mediaTableExists); err != nil {
slog.Warn("checking for media table", "error", err)
} else if mediaTableExists > 0 {
slog.Info("inserting media entries")
// Get post blob_cids for this day's posts
if _, err := tx.Exec("INSERT INTO media SELECT * FROM src.media WHERE content_hash IN (SELECT blob_cids FROM posts WHERE blob_cids IS NOT NULL)"); err != nil {
slog.Warn("inserting media (may not have matching entries)", "error", err)
}
}
// Commit the transaction before any further database operations
if err := tx.Commit(); err != nil {
return fmt.Errorf("committing transaction: %w", err)
}
tx = nil // Clear tx to ensure defer doesn't try to rollback
// Create indexes after bulk insert for speed
slog.Info("creating indexes")
idxRows, err := db.Query("SELECT sql FROM src.sqlite_master WHERE type='index' AND name NOT LIKE 'sqlite_%' AND sql IS NOT NULL ORDER BY name")
if err != nil {
return fmt.Errorf("reading source indexes: %w", err)
}
defer func() {
if cerr := idxRows.Close(); cerr != nil {
slog.Warn("failed to close index rows", "error", cerr)
}
}()
var idxStatements []string
for idxRows.Next() {
var idxSQL string
if err := idxRows.Scan(&idxSQL); err != nil {
return fmt.Errorf("scanning index DDL: %w", err)
}
idxStatements = append(idxStatements, idxSQL)
}
if err := idxRows.Err(); err != nil {
return fmt.Errorf("iterating index rows: %w", err)
}
for _, idxSQL := range idxStatements {
if _, err := db.Exec(idxSQL); err != nil {
return fmt.Errorf("creating index: %w\nDDL: %s", err, idxSQL)
}
} }
// Detach source // Detach source
if _, err := db.Exec("DETACH DATABASE src"); err != nil { _, err = db.ExecContext(ctx, "DETACH DATABASE src")
if err != nil {
return fmt.Errorf("detaching source database: %w", err) return fmt.Errorf("detaching source database: %w", err)
} }
// Verify post count // Verify post count
var verifyCount int64 var verifyCount int64
if err := db.QueryRow("SELECT COUNT(*) FROM posts").Scan(&verifyCount); err != nil {
err = db.QueryRowContext(ctx, "SELECT COUNT(*) FROM posts").Scan(&verifyCount)
if err != nil {
return fmt.Errorf("verifying post count: %w", err) return fmt.Errorf("verifying post count: %w", err)
} }
if verifyCount != postCount { if verifyCount != postCount {
return fmt.Errorf("post count mismatch: inserted %d but found %d", postCount, verifyCount) return fmt.Errorf("%w: inserted %d but found %d",
errPostCountMismatch, postCount, verifyCount)
} }
slog.Info("extraction complete", "posts", verifyCount) slog.Info("extraction complete", "posts", verifyCount)
return nil return nil
} }
// createTables creates every table of the attached source database, empty,
// in the destination database.
func createTables(ctx context.Context, db *sql.DB) error {
// Copy table DDL from source
slog.Info("copying table DDL from source")
rows, err := db.QueryContext(ctx, "SELECT sql FROM src.sqlite_master "+
"WHERE type='table' AND name NOT LIKE 'sqlite_%' ORDER BY name")
if err != nil {
return fmt.Errorf("reading source schema: %w", err)
}
defer func() {
cerr := rows.Close()
if cerr != nil {
slog.Warn("failed to close schema rows", "error", cerr)
}
}()
var ddlStatements []string
for rows.Next() {
var ddl string
err = rows.Scan(&ddl)
if err != nil {
return fmt.Errorf("scanning DDL: %w", err)
}
ddlStatements = append(ddlStatements, ddl)
}
err = rows.Err()
if err != nil {
return fmt.Errorf("iterating DDL rows: %w", err)
}
for _, ddl := range ddlStatements {
_, err = db.ExecContext(ctx, ddl)
if err != nil {
return fmt.Errorf("creating table: %w\nDDL: %s", err, ddl)
}
}
return nil
}
// insertDayRows copies the target day's posts, and the rows they refer
// to, from the source database in one transaction. It returns the number
// of posts copied, or ErrNoPosts when there are none.
func insertDayRows(ctx context.Context, db *sql.DB, targetDay time.Time,
dayStart, dayEnd string,
) (int64, error) {
// Begin transaction for bulk inserts
tx, err := db.BeginTx(ctx, nil)
if err != nil {
return 0, fmt.Errorf("beginning transaction: %w", err)
}
// After a successful Commit this does nothing and returns ErrTxDone.
defer func() {
rerr := tx.Rollback()
if rerr != nil && !errors.Is(rerr, sql.ErrTxDone) {
slog.Warn("failed to roll back transaction", "error", rerr)
}
}()
// Insert target day's data
slog.Info("inserting posts for target day")
//nolint:unqueryvet // copies whole rows; the source defines the columns
result, err := tx.ExecContext(ctx, "INSERT INTO posts SELECT * FROM src.posts "+
"WHERE timestamp >= ? AND timestamp < ?", dayStart, dayEnd)
if err != nil {
return 0, fmt.Errorf("inserting posts: %w", err)
}
postCount, _ := result.RowsAffected()
slog.Info("inserted posts", "count", postCount)
if postCount == 0 {
return 0, fmt.Errorf("%w %s - aborting to avoid producing empty output",
ErrNoPosts, targetDay.Format("2006-01-02"))
}
slog.Info("inserting junction and lookup tables")
err = insertRelatedRows(ctx, tx)
if err != nil {
return 0, err
}
copyMediaRows(ctx, tx)
// Commit the transaction before any further database operations
err = tx.Commit()
if err != nil {
return 0, fmt.Errorf("committing transaction: %w", err)
}
return postCount, nil
}
// insertRelatedRows copies the hashtag and URL links of the posts already
// inserted, the hashtags and URLs they link to, and the posting users.
//
//nolint:unqueryvet // copies whole rows; the source defines the columns
func insertRelatedRows(ctx context.Context, tx *sql.Tx) error {
_, err := tx.ExecContext(ctx, "INSERT INTO posts_hashtags "+
"SELECT * FROM src.posts_hashtags WHERE post_id IN (SELECT id FROM posts)")
if err != nil {
return fmt.Errorf("inserting posts_hashtags: %w", err)
}
_, err = tx.ExecContext(ctx, "INSERT INTO posts_urls "+
"SELECT * FROM src.posts_urls WHERE post_id IN (SELECT id FROM posts)")
if err != nil {
return fmt.Errorf("inserting posts_urls: %w", err)
}
_, err = tx.ExecContext(ctx, "INSERT INTO hashtags "+
"SELECT * FROM src.hashtags "+
"WHERE id IN (SELECT hashtag_id FROM posts_hashtags)")
if err != nil {
return fmt.Errorf("inserting hashtags: %w", err)
}
_, err = tx.ExecContext(ctx, "INSERT INTO urls "+
"SELECT * FROM src.urls WHERE id IN (SELECT url_id FROM posts_urls)")
if err != nil {
return fmt.Errorf("inserting urls: %w", err)
}
_, err = tx.ExecContext(ctx, "INSERT INTO users "+
"SELECT * FROM src.users WHERE did IN (SELECT user_did FROM posts)")
if err != nil {
return fmt.Errorf("inserting users: %w", err)
}
return nil
}
// copyMediaRows copies the media rows of the posts already inserted, when
// the source has a media table. Failures are logged, not returned.
func copyMediaRows(ctx context.Context, tx *sql.Tx) {
// Check if media table exists in source and copy if present
var mediaTableExists int
err := tx.QueryRowContext(ctx, "SELECT COUNT(*) FROM src.sqlite_master "+
"WHERE type='table' AND name='media'").Scan(&mediaTableExists)
if err != nil {
slog.Warn("checking for media table", "error", err)
} else if mediaTableExists > 0 {
slog.Info("inserting media entries")
// Get post blob_cids for this day's posts
//nolint:unqueryvet // copies whole rows; the source defines the columns
_, err = tx.ExecContext(ctx, "INSERT INTO media SELECT * FROM src.media "+
"WHERE content_hash IN "+
"(SELECT blob_cids FROM posts WHERE blob_cids IS NOT NULL)")
if err != nil {
slog.Warn("inserting media (may not have matching entries)",
"error", err)
}
}
}
// createIndexes creates every index of the attached source database in
// the destination database. Doing this after the bulk insert is faster
// than inserting into indexed tables.
func createIndexes(ctx context.Context, db *sql.DB) error {
// Create indexes after bulk insert for speed
slog.Info("creating indexes")
idxRows, err := db.QueryContext(ctx, "SELECT sql FROM src.sqlite_master "+
"WHERE type='index' AND name NOT LIKE 'sqlite_%' AND sql IS NOT NULL "+
"ORDER BY name")
if err != nil {
return fmt.Errorf("reading source indexes: %w", err)
}
defer func() {
cerr := idxRows.Close()
if cerr != nil {
slog.Warn("failed to close index rows", "error", cerr)
}
}()
var idxStatements []string
for idxRows.Next() {
var idxSQL string
err = idxRows.Scan(&idxSQL)
if err != nil {
return fmt.Errorf("scanning index DDL: %w", err)
}
idxStatements = append(idxStatements, idxSQL)
}
err = idxRows.Err()
if err != nil {
return fmt.Errorf("iterating index rows: %w", err)
}
for _, idxSQL := range idxStatements {
_, err = db.ExecContext(ctx, idxSQL)
if err != nil {
return fmt.Errorf("creating index: %w\nDDL: %s", err, idxSQL)
}
}
return nil
}
+174 -87
View File
@@ -9,21 +9,29 @@ import (
"time" "time"
) )
var errEmptySource = errors.New("source file is empty")
// cleanup removes a temporary file, logging a warning if removal fails so // cleanup removes a temporary file, logging a warning if removal fails so
// that leaked scratch files are surfaced rather than silently ignored. A // that leaked scratch files are surfaced rather than silently ignored. A
// missing file is not an error. // missing file is not an error.
func cleanup(path string) { func cleanup(path string) {
if err := os.Remove(path); err != nil && !os.IsNotExist(err) { err := os.Remove(path)
if err != nil && !os.IsNotExist(err) {
slog.Warn("failed to remove temporary file", "path", path, "error", err) slog.Warn("failed to remove temporary file", "path", path, "error", err)
} }
} }
// Run writes a compressed SQL dump of each day in targetDates, taken from
// the latest daily snapshot, skipping days that already have one or have
// no posts. With no dates it does the day before the snapshot date.
func Run(targetDates []time.Time) error { func Run(targetDates []time.Time) error {
snapshotDir, snapshotDate, err := FindLatestDailySnapshot() snapshotDir, snapshotDate, err := FindLatestDailySnapshot()
if err != nil { if err != nil {
return fmt.Errorf("finding latest snapshot: %w", err) return fmt.Errorf("finding latest snapshot: %w", err)
} }
slog.Info("found latest daily snapshot", "dir", snapshotDir, "snapshot_date", snapshotDate.Format("2006-01-02"))
slog.Info("found latest daily snapshot", "dir", snapshotDir,
"snapshot_date", snapshotDate.Format("2006-01-02"))
if len(targetDates) == 0 { if len(targetDates) == 0 {
targetDates = []time.Time{snapshotDate.AddDate(0, 0, -1)} targetDates = []time.Time{snapshotDate.AddDate(0, 0, -1)}
@@ -34,10 +42,13 @@ func Run(targetDates []time.Time) error {
"last", targetDates[len(targetDates)-1].Format("2006-01-02")) "last", targetDates[len(targetDates)-1].Format("2006-01-02"))
// Check disk space // Check disk space
if err := CheckFreeSpace(TmpBase, MinTmpFreeBytes, "tmpBase"); err != nil { err = CheckFreeSpace(TmpBase, MinTmpFreeBytes, "tmpBase")
if err != nil {
return err return err
} }
if err := CheckFreeSpace(DailiesBase, MinDailiesFreeBytes, "dailiesBase"); err != nil {
err = CheckFreeSpace(DailiesBase, MinDailiesFreeBytes, "dailiesBase")
if err != nil {
return err return err
} }
@@ -46,14 +57,53 @@ func Run(targetDates []time.Time) error {
if err != nil { if err != nil {
return fmt.Errorf("creating temp directory in %s: %w", TmpBase, err) return fmt.Errorf("creating temp directory in %s: %w", TmpBase, err)
} }
slog.Info("created temp directory", "path", tmpDir) slog.Info("created temp directory", "path", tmpDir)
defer func() { defer func() {
slog.Info("cleaning up temp directory", "path", tmpDir) slog.Info("cleaning up temp directory", "path", tmpDir)
if err := os.RemoveAll(tmpDir); err != nil {
slog.Error("failed to remove temp directory", "path", tmpDir, "error", err) rerr := os.RemoveAll(tmpDir)
if rerr != nil {
slog.Error("failed to remove temp directory",
"path", tmpDir, "error", rerr)
} }
}() }()
dstDB, err := copySnapshotFiles(snapshotDir, tmpDir)
if err != nil {
return err
}
// Process each day completely before moving to the next. This ensures
// we don't have multiple SQLite operations competing for the same
// source database.
processed := 0
skipped := 0
for _, targetDay := range targetDates {
written, err := processDay(tmpDir, dstDB, targetDay)
if err != nil {
return err
}
if written {
processed++
} else {
skipped++
}
}
slog.Info("run summary", "processed", processed, "skipped", skipped,
"total", len(targetDates))
return nil
}
// copySnapshotFiles copies the database, its WAL and, if present, its SHM
// file from snapshotDir into tmpDir, and returns the copied database's
// path.
func copySnapshotFiles(snapshotDir, tmpDir string) (string, error) {
// Copy database files from snapshot to temp // Copy database files from snapshot to temp
srcDB := filepath.Join(snapshotDir, DBFilename) srcDB := filepath.Join(snapshotDir, DBFilename)
srcWAL := filepath.Join(snapshotDir, WALFilename) srcWAL := filepath.Join(snapshotDir, WALFilename)
@@ -65,100 +115,137 @@ func Run(targetDates []time.Time) error {
for _, f := range []string{srcDB, srcWAL} { for _, f := range []string{srcDB, srcWAL} {
info, err := os.Stat(f) info, err := os.Stat(f)
if err != nil { if err != nil {
return fmt.Errorf("source file missing: %s: %w", f, err) return "", fmt.Errorf("source file missing: %s: %w", f, err)
} }
if info.Size() == 0 { if info.Size() == 0 {
return fmt.Errorf("source file is empty: %s", f) return "", fmt.Errorf("%w: %s", errEmptySource, f)
} }
slog.Info("source file", "path", f, "size_bytes", info.Size()) slog.Info("source file", "path", f, "size_bytes", info.Size())
} }
if err := CopyFile(srcDB, dstDB); err != nil { err := CopyFile(srcDB, dstDB)
return fmt.Errorf("copying database: %w", err) if err != nil {
} return "", fmt.Errorf("copying database: %w", err)
if err := CopyFile(srcWAL, dstWAL); err != nil {
return fmt.Errorf("copying WAL: %w", err)
}
if _, err := os.Stat(srcSHM); err == nil {
if err := CopyFile(srcSHM, dstSHM); err != nil {
return fmt.Errorf("copying SHM: %w", err)
}
} }
// Process each day completely before moving to the next err = CopyFile(srcWAL, dstWAL)
// This ensures we don't have multiple SQLite operations competing for the same source database if err != nil {
processed := 0 return "", fmt.Errorf("copying WAL: %w", err)
skipped := 0 }
for _, targetDay := range targetDates { _, err = os.Stat(srcSHM)
dayStr := targetDay.Format("2006-01-02") if err == nil {
slog.Info("processing day", "date", dayStr) err = CopyFile(srcSHM, dstSHM)
// Check if output already exists
outputDir := filepath.Join(DailiesBase, targetDay.Format("2006-01"))
outputFinal := filepath.Join(outputDir, dayStr+".sql.zst")
if _, err := os.Stat(outputFinal); err == nil {
slog.Info("output already exists, skipping", "path", outputFinal)
skipped++
continue
}
// Extract target day into a per-day database
extractedDB := filepath.Join(tmpDir, "extracted-"+dayStr+".db")
slog.Info("extracting target day", "src", dstDB, "dst", extractedDB)
if err := ExtractDay(dstDB, extractedDB, targetDay); err != nil {
if errors.Is(err, ErrNoPosts) {
slog.Warn("no posts found, skipping day", "date", dayStr)
cleanup(extractedDB)
skipped++
continue
}
return fmt.Errorf("extracting day %s: %w", dayStr, err)
}
// Dump to SQL and compress
if err := os.MkdirAll(outputDir, 0755); err != nil {
cleanup(extractedDB)
return fmt.Errorf("creating output directory %s: %w", outputDir, err)
}
outputTmp := filepath.Join(outputDir, "."+dayStr+".sql.zst.tmp")
slog.Info("dumping and compressing", "tmp_output", outputTmp)
if err := DumpAndCompress(extractedDB, outputTmp); err != nil {
cleanup(outputTmp)
cleanup(extractedDB)
return fmt.Errorf("dump and compress for %s: %w", dayStr, err)
}
slog.Info("verifying compressed output")
if err := VerifyOutput(outputTmp); err != nil {
cleanup(outputTmp)
cleanup(extractedDB)
return fmt.Errorf("verification failed for %s: %w", dayStr, err)
}
// Atomic rename to final path
slog.Info("renaming to final output", "from", outputTmp, "to", outputFinal)
if err := os.Rename(outputTmp, outputFinal); err != nil {
cleanup(outputTmp)
cleanup(extractedDB)
return fmt.Errorf("atomic rename for %s: %w", dayStr, err)
}
info, err := os.Stat(outputFinal)
if err != nil { if err != nil {
cleanup(extractedDB) return "", fmt.Errorf("copying SHM: %w", err)
return fmt.Errorf("stat final output: %w", err)
} }
slog.Info("day completed", "date", dayStr, "path", outputFinal, "size_bytes", info.Size())
// Remove extracted DB to reclaim space immediately
cleanup(extractedDB)
processed++
} }
slog.Info("run summary", "processed", processed, "skipped", skipped, "total", len(targetDates)) return dstDB, nil
}
// processDay extracts one day from the copied database at dstDB and
// publishes its compressed dump. It returns false, with no error, for a
// day it skips: one whose output already exists or that has no posts.
func processDay(tmpDir, dstDB string, targetDay time.Time) (bool, error) {
dayStr := targetDay.Format("2006-01-02")
slog.Info("processing day", "date", dayStr)
// Check if output already exists
outputDir := filepath.Join(DailiesBase, targetDay.Format("2006-01"))
outputFinal := filepath.Join(outputDir, dayStr+".sql.zst")
_, err := os.Stat(outputFinal)
if err == nil {
slog.Info("output already exists, skipping", "path", outputFinal)
return false, nil
}
// Extract target day into a per-day database
extractedDB := filepath.Join(tmpDir, "extracted-"+dayStr+".db")
slog.Info("extracting target day", "src", dstDB, "dst", extractedDB)
err = ExtractDay(dstDB, extractedDB, targetDay)
if err != nil {
if errors.Is(err, ErrNoPosts) {
slog.Warn("no posts found, skipping day", "date", dayStr)
cleanup(extractedDB)
return false, nil
}
return false, fmt.Errorf("extracting day %s: %w", dayStr, err)
}
// Dump to SQL and compress
//nolint:gosec,mnd // world-readable on purpose: the dailies tree is published
err = os.MkdirAll(outputDir, 0755)
if err != nil {
cleanup(extractedDB)
return false, fmt.Errorf("creating output directory %s: %w", outputDir, err)
}
outputTmp := filepath.Join(outputDir, "."+dayStr+".sql.zst.tmp")
slog.Info("dumping and compressing", "tmp_output", outputTmp)
err = DumpAndCompress(extractedDB, outputTmp)
if err != nil {
cleanup(outputTmp)
cleanup(extractedDB)
return false, fmt.Errorf("dump and compress for %s: %w", dayStr, err)
}
slog.Info("verifying compressed output")
err = VerifyOutput(outputTmp)
if err != nil {
cleanup(outputTmp)
cleanup(extractedDB)
return false, fmt.Errorf("verification failed for %s: %w", dayStr, err)
}
err = publishOutput(outputTmp, outputFinal, dayStr)
// Remove extracted DB to reclaim space immediately
cleanup(extractedDB)
if err != nil {
return false, err
}
return true, nil
}
// publishOutput renames the verified temporary output to its final path
// and logs the finished day. The rename is atomic, so the final path never
// holds a partial file.
func publishOutput(outputTmp, outputFinal, dayStr string) error {
// Atomic rename to final path
slog.Info("renaming to final output", "from", outputTmp, "to", outputFinal)
err := os.Rename(outputTmp, outputFinal)
if err != nil {
cleanup(outputTmp)
return fmt.Errorf("atomic rename for %s: %w", dayStr, err)
}
info, err := os.Stat(outputFinal)
if err != nil {
return fmt.Errorf("stat final output: %w", err)
}
slog.Info("day completed", "date", dayStr, "path", outputFinal,
"size_bytes", info.Size())
return nil return nil
} }
+23 -7
View File
@@ -1,6 +1,7 @@
package bsdaily package bsdaily
import ( import (
"errors"
"fmt" "fmt"
"log/slog" "log/slog"
"os" "os"
@@ -9,10 +10,16 @@ import (
"time" "time"
) )
func FindLatestDailySnapshot() (dir string, snapshotDate time.Time, err error) { var errNoSnapshots = errors.New("no daily snapshots found")
// FindLatestDailySnapshot returns the directory and date of the newest
// daily snapshot in SnapshotBase. It fails when there is none or when the
// newest one has no database file.
func FindLatestDailySnapshot() (string, time.Time, error) {
entries, err := os.ReadDir(SnapshotBase) entries, err := os.ReadDir(SnapshotBase)
if err != nil { if err != nil {
return "", time.Time{}, fmt.Errorf("reading snapshot directory %s: %w", SnapshotBase, err) return "", time.Time{}, fmt.Errorf("reading snapshot directory %s: %w",
SnapshotBase, err)
} }
type snapshot struct { type snapshot struct {
@@ -21,24 +28,30 @@ func FindLatestDailySnapshot() (dir string, snapshotDate time.Time, err error) {
} }
var snapshots []snapshot var snapshots []snapshot
for _, e := range entries { for _, e := range entries {
if !e.IsDir() { if !e.IsDir() {
continue continue
} }
m := snapshotPattern.FindStringSubmatch(e.Name()) m := snapshotPattern.FindStringSubmatch(e.Name())
if m == nil { if m == nil {
continue continue
} }
d, err := time.Parse("2006-01-02", m[1]) d, err := time.Parse("2006-01-02", m[1])
if err != nil { if err != nil {
slog.Warn("skipping snapshot with unparseable date", "name", e.Name(), "error", err) slog.Warn("skipping snapshot with unparseable date",
"name", e.Name(), "error", err)
continue continue
} }
snapshots = append(snapshots, snapshot{name: e.Name(), date: d}) snapshots = append(snapshots, snapshot{name: e.Name(), date: d})
} }
if len(snapshots) == 0 { if len(snapshots) == 0 {
return "", time.Time{}, fmt.Errorf("no daily snapshots found in %s", SnapshotBase) return "", time.Time{}, fmt.Errorf("%w in %s", errNoSnapshots, SnapshotBase)
} }
sort.Slice(snapshots, func(i, j int) bool { sort.Slice(snapshots, func(i, j int) bool {
@@ -46,11 +59,14 @@ func FindLatestDailySnapshot() (dir string, snapshotDate time.Time, err error) {
}) })
latest := snapshots[0] latest := snapshots[0]
dir = filepath.Join(SnapshotBase, latest.name) dir := filepath.Join(SnapshotBase, latest.name)
dbPath := filepath.Join(dir, DBFilename) dbPath := filepath.Join(dir, DBFilename)
if _, err := os.Stat(dbPath); err != nil {
return "", time.Time{}, fmt.Errorf("database not found in snapshot %s: %w", dir, err) _, err = os.Stat(dbPath)
if err != nil {
return "", time.Time{}, fmt.Errorf("database not found in snapshot %s: %w",
dir, err)
} }
return dir, latest.date, nil return dir, latest.date, nil
+81 -19
View File
@@ -1,6 +1,7 @@
package bsdaily package bsdaily
import ( import (
"context"
"errors" "errors"
"fmt" "fmt"
"log/slog" "log/slog"
@@ -9,6 +10,11 @@ import (
"strings" "strings"
) )
var (
errEmptyDecompressed = errors.New("decompressed content is empty")
errNotSQL = errors.New("decompressed content does not look like SQL")
)
// killCat terminates the zstdcat process, ignoring the benign case where it // killCat terminates the zstdcat process, ignoring the benign case where it
// has already exited (e.g. after receiving SIGPIPE when head closed the pipe) // has already exited (e.g. after receiving SIGPIPE when head closed the pipe)
// and logging any other failure. // and logging any other failure.
@@ -16,70 +22,126 @@ func killCat(cmd *exec.Cmd) {
if cmd.Process == nil { if cmd.Process == nil {
return return
} }
if err := cmd.Process.Kill(); err != nil && !errors.Is(err, os.ErrProcessDone) {
err := cmd.Process.Kill()
if err != nil && !errors.Is(err, os.ErrProcessDone) {
slog.Warn("failed to kill zstdcat process", "error", err) slog.Warn("failed to kill zstdcat process", "error", err)
} }
} }
// VerifyOutput checks that the compressed file at path passes zstdmt's
// integrity test and that its first lines look like SQL.
func VerifyOutput(path string) error { func VerifyOutput(path string) error {
ctx := context.Background()
slog.Info("running zstdmt integrity check") slog.Info("running zstdmt integrity check")
testCmd := exec.Command("zstdmt", "--test", path)
//nolint:gosec // path is an output file this package named
testCmd := exec.CommandContext(ctx, "zstdmt", "--test", path)
var testStderr strings.Builder var testStderr strings.Builder
testCmd.Stderr = &testStderr testCmd.Stderr = &testStderr
if err := testCmd.Run(); err != nil {
return fmt.Errorf("zstdmt --test failed: %w; stderr: %s", err, testStderr.String()) err := testCmd.Run()
if err != nil {
return fmt.Errorf("zstdmt --test failed: %w; stderr: %s",
err, testStderr.String())
} }
slog.Info("zstdmt integrity check passed") slog.Info("zstdmt integrity check passed")
slog.Info("verifying SQL content") slog.Info("verifying SQL content")
catCmd := exec.Command("zstdcat", path)
headCmd := exec.Command("head", fmt.Sprintf("-%d", verificationHeadLines)) content, err := readDecompressedHead(ctx, path)
if err != nil {
return err
}
err = checkLooksLikeSQL(content)
if err != nil {
return err
}
slog.Info("SQL content verification passed")
return nil
}
// readDecompressedHead returns the first verificationHeadLines lines of
// the decompressed file at path, read through `zstdcat path | head`.
func readDecompressedHead(ctx context.Context, path string) (string, error) {
//nolint:gosec // path is an output file this package named
catCmd := exec.CommandContext(ctx, "zstdcat", path)
//nolint:gosec // the argument is built from a constant
headCmd := exec.CommandContext(ctx, "head",
fmt.Sprintf("-%d", verificationHeadLines))
pipe, err := catCmd.StdoutPipe() pipe, err := catCmd.StdoutPipe()
if err != nil { if err != nil {
return fmt.Errorf("creating zstdcat pipe: %w", err) return "", fmt.Errorf("creating zstdcat pipe: %w", err)
} }
headCmd.Stdin = pipe headCmd.Stdin = pipe
var headOut strings.Builder var headOut strings.Builder
headCmd.Stdout = &headOut headCmd.Stdout = &headOut
if err := catCmd.Start(); err != nil { err = catCmd.Start()
return fmt.Errorf("starting zstdcat: %w", err) if err != nil {
return "", fmt.Errorf("starting zstdcat: %w", err)
} }
if err := headCmd.Start(); err != nil {
err = headCmd.Start()
if err != nil {
killCat(catCmd) // Clean up if head fails to start killCat(catCmd) // Clean up if head fails to start
return fmt.Errorf("starting head: %w", err)
return "", fmt.Errorf("starting head: %w", err)
} }
// Wait for head first (it will exit when it has enough lines) // Wait for head first (it will exit when it has enough lines)
if err := headCmd.Wait(); err != nil { err = headCmd.Wait()
if err != nil {
killCat(catCmd) killCat(catCmd)
return fmt.Errorf("head command failed: %w", err)
return "", fmt.Errorf("head command failed: %w", err)
} }
// Kill zstdcat since head closed the pipe (expected SIGPIPE) // Kill zstdcat since head closed the pipe (expected SIGPIPE)
killCat(catCmd) killCat(catCmd)
_ = catCmd.Wait() // Reap the process _ = catCmd.Wait() // Reap the process
content := headOut.String() return headOut.String(), nil
}
// checkLooksLikeSQL returns an error when content is empty or contains
// none of the keywords expected near the start of a `sqlite3 .dump`.
func checkLooksLikeSQL(content string) error {
if len(content) == 0 { if len(content) == 0 {
return fmt.Errorf("decompressed content is empty") return errEmptyDecompressed
} }
hasSQLMarker := false hasSQLMarker := false
for _, marker := range []string{"BEGIN TRANSACTION", "CREATE TABLE", "INSERT INTO", "PRAGMA"} {
for _, marker := range []string{
"BEGIN TRANSACTION", "CREATE TABLE", "INSERT INTO", "PRAGMA",
} {
if strings.Contains(content, marker) { if strings.Contains(content, marker) {
hasSQLMarker = true hasSQLMarker = true
break break
} }
} }
const verificationSampleBytes = 200 const verificationSampleBytes = 200
if !hasSQLMarker { if !hasSQLMarker {
return fmt.Errorf("decompressed content does not look like SQL; first %d bytes: %s", return fmt.Errorf("%w; first %d bytes: %s", errNotSQL,
verificationSampleBytes, content[:min(verificationSampleBytes, len(content))]) verificationSampleBytes,
content[:min(verificationSampleBytes, len(content))])
} }
slog.Info("SQL content verification passed")
return nil return nil
} }
+67
View File
@@ -0,0 +1,67 @@
// Package cli is the bsdaily command line: the command, its flags, the
// rules for combining them and the dates they name. The extraction itself
// is in package bsdaily.
package cli
import (
"log/slog"
"os"
"git.eeqj.de/sneak/bsdaily/internal/bsdaily"
"github.com/spf13/cobra"
)
// Main runs the bsdaily command on the program's command-line arguments
// and returns the status for the process to exit with. version is the
// build's version; an empty one is reported as dev.
func Main(version string) int {
if version == "" {
version = "dev"
}
logger := slog.New(slog.NewTextHandler(os.Stderr, &slog.HandlerOptions{
Level: slog.LevelInfo,
}))
slog.SetDefault(logger)
var dateFlag, fromFlag, toFlag string
rootCmd := &cobra.Command{
Use: "bsdaily",
Short: "Extract a single day's data from the latest daily snapshot",
Version: version,
SilenceUsage: true,
RunE: func(_ *cobra.Command, _ []string) error {
slog.Info("starting", "version", version)
targetDates, err := ParseTargetDates(dateFlag, fromFlag, toFlag)
if err != nil {
return err
}
err = bsdaily.Run(targetDates)
if err != nil {
return err
}
slog.Info("completed successfully")
return nil
},
}
rootCmd.Flags().StringVarP(&dateFlag, "date", "d", "",
"target date to extract (YYYY-MM-DD); "+
"defaults to snapshot date minus one day")
rootCmd.Flags().StringVar(&fromFlag, "from", "",
"start of date range to extract (YYYY-MM-DD, inclusive); use with --to")
rootCmd.Flags().StringVar(&toFlag, "to", "",
"end of date range to extract (YYYY-MM-DD, inclusive); use with --from")
err := rootCmd.Execute()
if err != nil {
return 1
}
return 0
}
+74
View File
@@ -0,0 +1,74 @@
package cli
import (
"errors"
"fmt"
"time"
)
var (
errDateExclusive = errors.New("--date and --from/--to are mutually exclusive")
errFromRequiresTo = errors.New("--from requires --to")
errToRequiresFrom = errors.New("--to requires --from")
errFromAfterTo = errors.New("is after --to")
)
// ParseTargetDates turns the --date, --from and --to flags into the days
// to extract. It returns nil when none of them is set, which bsdaily.Run
// takes to mean the snapshot date minus one day.
func ParseTargetDates(dateFlag, fromFlag, toFlag string) ([]time.Time, error) {
hasDate := dateFlag != ""
hasFrom := fromFlag != ""
hasTo := toFlag != ""
// Validate mutual exclusivity
if hasDate && (hasFrom || hasTo) {
return nil, errDateExclusive
}
if hasFrom != hasTo {
if hasFrom {
return nil, errFromRequiresTo
}
return nil, errToRequiresFrom
}
if hasDate {
t, err := time.Parse("2006-01-02", dateFlag)
if err != nil {
return nil, fmt.Errorf(
"invalid --date %q (expected YYYY-MM-DD): %w", dateFlag, err)
}
return []time.Time{t}, nil
}
if !hasFrom {
return nil, nil
}
from, err := time.Parse("2006-01-02", fromFlag)
if err != nil {
return nil, fmt.Errorf(
"invalid --from %q (expected YYYY-MM-DD): %w", fromFlag, err)
}
to, err := time.Parse("2006-01-02", toFlag)
if err != nil {
return nil, fmt.Errorf(
"invalid --to %q (expected YYYY-MM-DD): %w", toFlag, err)
}
if from.After(to) {
return nil, fmt.Errorf("--from %s %w %s", fromFlag, errFromAfterTo, toFlag)
}
var targetDates []time.Time
for d := from; !d.After(to); d = d.AddDate(0, 0, 1) {
targetDates = append(targetDates, d)
}
return targetDates, nil
}
+127
View File
@@ -0,0 +1,127 @@
package cli_test
import (
"slices"
"strings"
"testing"
"git.eeqj.de/sneak/bsdaily/internal/cli"
)
func TestParseTargetDates(t *testing.T) {
t.Parallel()
const oneDay = "2026-05-14"
// want lists the expected days as YYYY-MM-DD.
tests := []struct {
name string
date string
from string
to string
want []string
}{
{name: "no flags"},
{name: "single date", date: "2026-06-27", want: []string{"2026-06-27"}},
{
name: "range across a month end", from: "2026-06-29", to: "2026-07-02",
want: []string{"2026-06-29", "2026-06-30", "2026-07-01", "2026-07-02"},
},
{name: "range of one day", from: oneDay, to: oneDay, want: []string{oneDay}},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
got, err := cli.ParseTargetDates(tt.date, tt.from, tt.to)
if err != nil {
t.Fatalf("unexpected error: %v", err)
}
days := make([]string, 0, len(got))
for _, day := range got {
days = append(days, day.Format("2006-01-02"))
}
if !slices.Equal(days, tt.want) {
t.Errorf("days = %v, want %v", days, tt.want)
}
})
}
}
func TestParseTargetDatesErrors(t *testing.T) {
t.Parallel()
const dateExclusive = "--date and --from/--to are mutually exclusive"
// wantErr is the whole error message. With startsWith set it is only how
// the message starts: for a malformed date, the date parser's own
// explanation follows it.
tests := []struct {
name string
date string
from string
to string
wantErr string
startsWith bool
}{
{
name: "from after to", from: "2026-03-02", to: "2026-03-01",
wantErr: "--from 2026-03-02 is after --to 2026-03-01",
},
{name: "from without to", from: "2026-04-01", wantErr: "--from requires --to"},
{name: "to without from", to: "2026-04-02", wantErr: "--to requires --from"},
{
name: "date with from", date: "2026-01-05", from: "2026-01-06",
wantErr: dateExclusive,
},
{
name: "date with to", date: "2026-01-13", to: "2026-01-14",
wantErr: dateExclusive,
},
{
name: "date with from and to", date: "2026-01-07",
from: "2026-01-08", to: "2026-01-09",
wantErr: dateExclusive,
},
{
name: "malformed date", date: "27.06.2026",
wantErr: `invalid --date "27.06.2026" (expected YYYY-MM-DD): `,
startsWith: true,
},
{
name: "date that does not exist", date: "2026-02-30",
wantErr: `invalid --date "2026-02-30" (expected YYYY-MM-DD): `,
startsWith: true,
},
{
name: "malformed from", from: "2026-1-10", to: "2026-01-11",
wantErr: `invalid --from "2026-1-10" (expected YYYY-MM-DD): `,
startsWith: true,
},
{
name: "malformed to", from: "2026-01-12", to: "tomorrow",
wantErr: `invalid --to "tomorrow" (expected YYYY-MM-DD): `,
startsWith: true,
},
}
for _, tt := range tests {
t.Run(tt.name, func(t *testing.T) {
t.Parallel()
_, err := cli.ParseTargetDates(tt.date, tt.from, tt.to)
switch {
case err == nil:
t.Errorf("no error, want %q", tt.wantErr)
case tt.startsWith && !strings.HasPrefix(err.Error(), tt.wantErr):
t.Errorf("error = %q, want one starting %q", err, tt.wantErr)
case !tt.startsWith && err.Error() != tt.wantErr:
t.Errorf("error = %q, want %q", err, tt.wantErr)
}
})
}
}
+5
View File
@@ -0,0 +1,5 @@
{
"devDependencies": {
"prettier": "3.8.1"
}
}
+88 -9
View File
@@ -3,13 +3,23 @@
# this repo. Idempotent: every install is guarded by a check so already # this repo. Idempotent: every install is guarded by a check so already
# installed tools are skipped. Base tooling comes from nix, apt, brew, # installed tools are skipped. Base tooling comes from nix, apt, brew,
# or apk (detected in that order); assumes NOTHING is present (not git, # or apk (detected in that order); assumes NOTHING is present (not git,
# make, or go). # make, go, or node). Node is used directly if installed; otherwise it
# is installed at a pinned version via nvm (installing nvm itself first,
# from a hash-verified release archive, never curl | sh).
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Pinned versions, 2026-07-06
NODE_VERSION="22.17.0"
NVM_VERSION="0.40.3"
# sha256 of https://github.com/nvm-sh/nvm/archive/refs/tags/v0.40.3.tar.gz
NVM_SHA256="5f4d6aaa04a177dc93c985e31dbc411ab6b8c6e1e21d8015dbc1372625fcd1d0"
YARN_VERSION="1.22.22"
PKGMGR="" PKGMGR=""
SUDO="" SUDO=""
APT_UPDATED=""
detect_pkgmgr() { detect_pkgmgr() {
[ -n "$PKGMGR" ] && return 0 [ -n "$PKGMGR" ] && return 0
@@ -38,7 +48,14 @@ pkg_install() {
detect_pkgmgr detect_pkgmgr
case "$PKGMGR" in case "$PKGMGR" in
nix) nix-env -iA "nixpkgs.$1" ;; nix) nix-env -iA "nixpkgs.$1" ;;
apt) $SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2" ;; apt)
# Package lists may be empty (fresh images); refresh once per run.
if [ -z "$APT_UPDATED" ]; then
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get update
APT_UPDATED=1
fi
$SUDO env DEBIAN_FRONTEND=noninteractive apt-get install -y "$2"
;;
brew) brew install "$3" ;; brew) brew install "$3" ;;
apk) apk add --no-cache "$4" ;; apk) apk add --no-cache "$4" ;;
esac esac
@@ -48,6 +65,69 @@ missing() {
! command -v "$1" >/dev/null 2>&1 ! command -v "$1" >/dev/null 2>&1
} }
# verify_sha256 <file> <expected-hash>
verify_sha256() {
if command -v sha256sum >/dev/null 2>&1; then
actual="$(sha256sum "$1" | cut -d' ' -f1)"
else
actual="$(shasum -a 256 "$1" | cut -d' ' -f1)"
fi
if [ "$actual" != "$2" ]; then
echo "bootstrap: sha256 mismatch for $1" >&2
echo " expected: $2" >&2
echo " actual: $actual" >&2
exit 1
fi
}
# nvm is a bash script; run a command in a bash with nvm loaded
nvm_sh() {
bash -c ". \"\$HOME/.nvm/nvm.sh\" && $*"
}
ensure_nvm() {
[ -s "$HOME/.nvm/nvm.sh" ] && return 0
# nvm prerequisites; nvm itself requires bash
if missing bash; then pkg_install bash bash bash bash; fi
if missing curl; then pkg_install curl curl curl curl; fi
if missing git; then pkg_install git git git git; fi
tmp="$(mktemp -d)"
curl -fsSL -o "$tmp/nvm.tar.gz" \
"https://github.com/nvm-sh/nvm/archive/refs/tags/v${NVM_VERSION}.tar.gz"
verify_sha256 "$tmp/nvm.tar.gz" "$NVM_SHA256"
mkdir -p "$HOME/.nvm"
tar -xzf "$tmp/nvm.tar.gz" -C "$HOME/.nvm" --strip-components=1
rm -rf "$tmp"
}
ensure_node() {
if ! missing node; then return 0; fi
ensure_nvm
nvm_sh "nvm install $NODE_VERSION"
}
ensure_yarn() {
if ! missing yarn; then return 0; fi
if ! missing corepack; then
corepack enable
corepack prepare "yarn@$YARN_VERSION" --activate
elif [ -s "$HOME/.nvm/nvm.sh" ]; then
nvm_sh "nvm use $NODE_VERSION >/dev/null && corepack enable && \
corepack prepare yarn@$YARN_VERSION --activate"
else
npm install -g "yarn@$YARN_VERSION"
fi
}
install_js_deps() {
if missing yarn && [ -s "$HOME/.nvm/nvm.sh" ]; then
nvm_sh "nvm use $NODE_VERSION >/dev/null && cd \"$ROOT\" && \
yarn install --frozen-lockfile"
else
yarn install --frozen-lockfile
fi
}
main() { main() {
cd "$ROOT" cd "$ROOT"
@@ -58,15 +138,14 @@ main() {
# Go toolchain # Go toolchain
if missing go; then pkg_install go golang go go; fi if missing go; then pkg_install go golang go go; fi
# golangci-lint: packaged in nix, brew, and apk. There is no apt
# package; on apt systems install it manually from a hash-verified
# GitHub release archive (never curl | sh).
if missing golangci-lint; then
pkg_install golangci-lint golangci-lint golangci-lint golangci-lint
fi
go mod download go mod download
# Node and yarn, then the prettier pinned in package.json and
# yarn.lock, which script/fmt and script/fmt-check run
ensure_node
ensure_yarn
install_js_deps
echo "bootstrap complete" echo "bootstrap complete"
} }
+21 -5
View File
@@ -1,14 +1,30 @@
#!/bin/sh #!/bin/sh
# script/cibuild: run the CI build. The Dockerfile runs script/check # script/cibuild: run the CI build. It bootstraps first: a CI runner
# (via make check), so a successful build implies all checks pass. # checks out and runs this and nothing else, and script/fmt-check runs
# Generic: needs no adaptation. The Gitea workflow runs this on push. # the formatter on the host, which a pristine checkout cannot do.
# --no-cache for the same reason as script/docker: the gate phases the
# final stage depends on are RUN steps, and a cached one is a check that
# did not run.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
docker build . "$SCRIPT_DIR/bootstrap"
"$SCRIPT_DIR/check"
# The version and the tag each get their own line: a failing
# command substitution inside an argument does not trip `set -e`,
# so the inline form degrades silently to an empty constant. The
# VERSION build argument takes precedence over the version a build
# stage derives from the .git in the context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
tag="$("$SCRIPT_DIR/projectname")"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$tag" .
} }
main "$@" main "$@"
+13 -2
View File
@@ -1,7 +1,8 @@
#!/bin/sh #!/bin/sh
# script/docker: build the Docker image tagged with the project name. # script/docker: build the Docker image tagged with the project name.
# Identical in all repos; the tag comes from script/projectname. # Identical in all repos; the tag comes from script/projectname.
# Generic: needs no adaptation. # --no-cache because the gate phases the final stage depends on are RUN
# steps, and a cached one is a check that did not run.
set -eu set -eu
SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
@@ -9,7 +10,17 @@ ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
docker build -t "$("$SCRIPT_DIR/projectname")" . # The version and the tag each get their own line: a failing
# command substitution inside an argument does not trip `set -e`,
# so the inline form degrades silently to an empty constant. The
# VERSION build argument takes precedence over the version a build
# stage derives from the .git in the context.
version="$(git describe --tags --always --dirty 2>/dev/null || true)"
[ -n "$version" ] || version="unknown"
tag="$("$SCRIPT_DIR/projectname")"
docker build --no-cache \
--build-arg VERSION="$version" \
-t "$tag" .
} }
main "$@" main "$@"
+23 -1
View File
@@ -1,12 +1,34 @@
#!/bin/sh #!/bin/sh
# script/fmt: format all files (writes). # script/fmt: format all files (writes): the Go code with go fmt and
# every Markdown file with prettier.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Must match the pin in script/bootstrap.
NODE_VERSION="22.17.0"
# script/bootstrap installs node and yarn under nvm and leaves neither
# on the PATH of the shell that called it, so resolve the pinned
# toolchain here the way bootstrap's own install step does. nvm is a
# bash script, hence the subshell.
run_yarn() {
if command -v yarn >/dev/null 2>&1; then
exec yarn "$@"
fi
if [ ! -s "$HOME/.nvm/nvm.sh" ]; then
echo "fmt: no yarn; run script/bootstrap first" >&2
exit 1
fi
exec bash -c '. "$HOME/.nvm/nvm.sh" && nvm use "$1" >/dev/null &&
shift && exec yarn "$@"' bash "$NODE_VERSION" "$@"
}
main() { main() {
cd "$ROOT" cd "$ROOT"
go fmt ./... go fmt ./...
# run_yarn replaces this shell, so it stays the last step.
run_yarn run prettier --write '**/*.md' --tab-width 4 --prose-wrap always
} }
main "$@" main "$@"
+23 -1
View File
@@ -1,10 +1,30 @@
#!/bin/sh #!/bin/sh
# script/fmt-check: check formatting (read-only). Same scope as # script/fmt-check: check formatting (read-only). Same scope as
# script/fmt, but fails instead of writing. # script/fmt, the Go code and every Markdown file, but fails instead of
# writing.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Must match the pin in script/bootstrap.
NODE_VERSION="22.17.0"
# script/bootstrap installs node and yarn under nvm and leaves neither
# on the PATH of the shell that called it, so resolve the pinned
# toolchain here the way bootstrap's own install step does. nvm is a
# bash script, hence the subshell.
run_yarn() {
if command -v yarn >/dev/null 2>&1; then
exec yarn "$@"
fi
if [ ! -s "$HOME/.nvm/nvm.sh" ]; then
echo "fmt-check: no yarn; run script/bootstrap first" >&2
exit 1
fi
exec bash -c '. "$HOME/.nvm/nvm.sh" && nvm use "$1" >/dev/null &&
shift && exec yarn "$@"' bash "$NODE_VERSION" "$@"
}
main() { main() {
cd "$ROOT" cd "$ROOT"
unformatted="$(gofmt -l .)" unformatted="$(gofmt -l .)"
@@ -13,6 +33,8 @@ main() {
echo "$unformatted" >&2 echo "$unformatted" >&2
exit 1 exit 1
fi fi
# run_yarn replaces this shell, so it stays the last step.
run_yarn run prettier --check '**/*.md' --tab-width 4 --prose-wrap always
} }
main "$@" main "$@"
+18 -3
View File
@@ -1,12 +1,27 @@
#!/bin/sh #!/bin/sh
# script/lint: run the linter. # script/lint: run the linter. Linting is a phase of the Dockerfile and
# this builds that phase alone; the linter is never installed or run on
# a developer host, where a shared result cache and a host-global lock
# make its answer untrustworthy.
#
# The phase is not the last stage in the file, so it is built only when
# --target names it. --no-cache because a cached lint layer is a lint
# that did not run. The tag makes each build replace the previous image
# instead of leaving a dangling one behind.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
golangci-lint run ./... # The tag gets its own line: a failing command substitution inside
# an argument does not trip `set -e`, so the inline form degrades
# silently to an empty constant.
tag="$("$SCRIPT_DIR/projectname")"
docker build --no-cache \
--target lint \
-t "$tag-lint" .
} }
main "$@" main "$@"
+14 -9
View File
@@ -1,18 +1,23 @@
#!/bin/sh #!/bin/sh
# script/test: run the test suite. Quiet on success; on failure, rerun # script/test: run the test suite. Testing is a phase of the Dockerfile
# verbosely for full diagnostic output (the exit 1 ensures the rerun # and this builds that phase alone, on the same terms as script/lint:
# never turns a failure into a pass). # --target because a phase that is not the last stage is built only when
# named, --no-cache because a cached test layer is a test that did not
# run, and a tag so each build replaces the previous image.
set -eu set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)" SCRIPT_DIR="$(cd "$(dirname "$0")" && pwd -P)"
ROOT="$(cd "$SCRIPT_DIR/.." && pwd -P)"
main() { main() {
cd "$ROOT" cd "$ROOT"
go test -race -timeout 30s ./... || { # The tag gets its own line: a failing command substitution inside
echo "--- Rerunning with -v for details ---" # an argument does not trip `set -e`, so the inline form degrades
go test -race -timeout 30s -v ./... # silently to an empty constant.
exit 1 tag="$("$SCRIPT_DIR/projectname")"
} docker build --no-cache \
--target test \
-t "$tag-test" .
} }
main "$@" main "$@"
+8
View File
@@ -0,0 +1,8 @@
# THIS IS AN AUTOGENERATED FILE. DO NOT EDIT THIS FILE DIRECTLY.
# yarn lockfile v1
prettier@3.8.1:
version "3.8.1"
resolved "https://registry.yarnpkg.com/prettier/-/prettier-3.8.1.tgz#edf48977cf991558f4fcbd8a3ba6015ba2a3a173"
integrity sha512-UOnG6LftzbdaHZcKoPFtOcCKztrQ57WkHDeRD9t/PTQtmT0NHSeWWepj6pS0z/N7+08BHFDQVUrfmfMRcZwbMg==