Author SHA1 Message Date
sneak 9b0186c481 Say script/bootstrap pins node and yarn only when absent (refs #7)
check / check (push) Waiting to run
The README Entrypoints line and the TODO.md Completed Steps entry said
script/bootstrap installs a pinned node and yarn. It installs them at the
pinned versions only when they are missing and uses a node or yarn already
installed whatever its version, as its header comment says.

Model: opus-5-5
2026-10-06 08:25:55 +00:00
clawbot 8125d4aebf Reformat the existing Markdown with prettier (closes #7)
check / check (push) Waiting to run
The output of make fmt after the previous commit: README.md and TODO.md
are rewrapped at 80 columns, and the TODO.md Workflow list uses -
markers. No wording changes. REPO_POLICIES.md was already formatted.

Model: opus-5-5
2026-10-06 06:31:14 +00:00
clawbot 47e48a2529 Format Markdown with prettier in script/fmt and script/fmt-check (refs #7)
check / check (push) Waiting to run
script/fmt now also runs prettier over every Markdown file, and
script/fmt-check checks them without writing. prettier is pinned in
package.json and yarn.lock. script/bootstrap keeps its Go and apt
handling and gains the canonical pinned node (nvm from a hash-checked
archive) and yarn (corepack) install, then installs the locked
packages. Both fmt scripts find yarn the canonical way, sourcing nvm
when yarn is not on PATH. The new files and the yarn lookup come from
sneak/prompts at cc440118c876. The vendored REPO_POLICIES.md already
passes prettier, so it is not ignored. The Markdown is reformatted in
the next commit.

Deviation: package.json drops the canonical "license": "MIT" line; this
repo is WTFPL.

Model: opus-5-5
2026-10-06 06:30:59 +00:00
clawbot 17b6b84322 Bring the repo up to the standard layout (closes #1)
check / check (push) Waiting to run
Vendors .gitignore, .dockerignore, .editorconfig, REPO_POLICIES.md,
script/cibuild, script/docker, script/lint, script/test and the CI
workflow byte-identical from sneak/prompts commit cc440118c876; the
two ignore files then add this repo's own build output, root-anchored
so cmd/bsdaily/ stays in. The Dockerfile gains a lint phase on the
golangci-lint image already pinned and a test phase on the Debian Go
image with sqlite3 and zstd; the build stage copies a file from each,
so a plain docker build cannot skip them. Nothing installs
golangci-lint on the host. On apt, script/bootstrap refreshes the
package lists once before its first install.

.golangci.yml and the lint cleanup: #6

Model: opus-5-5
2026-10-06 08:25:03 +02:00
9 changed files with 295 additions and 152 deletions
+2
View File
@@ -0,0 +1,2 @@
node_modules/
yarn.lock
+4
View File
@@ -0,0 +1,4 @@
{
"tabWidth": 4,
"proseWrap": "always"
}
+121 -122
View File
@@ -1,28 +1,26 @@
# bsdaily
[bsdaily](https://git.eeqj.de/sneak/bsdaily) is a command-line utility
written in [Go](https://golang.org) that carves a single day (or a range of
days) of [Bluesky](https://bsky.app) firehose data out of a large,
continuously-growing SQLite database and writes it out as a self-contained,
[bsdaily](https://git.eeqj.de/sneak/bsdaily) is a command-line utility written
in [Go](https://golang.org) that carves a single day (or a range of days) of
[Bluesky](https://bsky.app) firehose data out of a large, continuously-growing
SQLite database and writes it out as a self-contained,
[zstd](https://facebook.github.io/zstd/)-compressed SQL dump. The dumps are
named by date (e.g. `2026-06-27.sql.zst`), organized into per-month
directories, and are designed to be published, archived, mirrored, and later
re-merged back into a single database.
named by date (e.g. `2026-06-27.sql.zst`), organized into per-month directories,
and are designed to be published, archived, mirrored, and later re-merged back
into a single database.
The source database is read from a read-only [ZFS](https://openzfs.org)
snapshot, so extraction never contends with the live firehose ingester that
is writing to the original database. The tool is operationally
conservative: it checks free disk space before starting, copies the snapshot
to fast scratch storage, processes one day at a time to avoid SQLite lock
contention, verifies every compressed output before publishing it, and
writes output atomically via a temp-file-and-rename so a partial run never
leaves a corrupt `.sql.zst` behind.
snapshot, so extraction never contends with the live firehose ingester that is
writing to the original database. The tool is operationally conservative: it
checks free disk space before starting, copies the snapshot to fast scratch
storage, processes one day at a time to avoid SQLite lock contention, verifies
every compressed output before publishing it, and writes output atomically via a
temp-file-and-rename so a partial run never leaves a corrupt `.sql.zst` behind.
This project was written by [@sneak](https://sneak.berlin) to produce a
daily, mergeable, publicly-mirrorable archive of the Bluesky firehose. It is
currently a one-person effort. The current version is pre-1.0 and there has
not yet been a versioned release; [SemVer](https://semver.org) will be used
for releases.
This project was written by [@sneak](https://sneak.berlin) to produce a daily,
mergeable, publicly-mirrorable archive of the Bluesky firehose. It is currently
a one-person effort. The current version is pre-1.0 and there has not yet been a
versioned release; [SemVer](https://semver.org) will be used for releases.
# Build Status
@@ -33,14 +31,14 @@ branch must always be green.
Primary development happens on a privately-run Gitea instance at
[https://git.eeqj.de/sneak/bsdaily](https://git.eeqj.de/sneak/bsdaily) and
issues are [tracked
there](https://git.eeqj.de/sneak/bsdaily/issues).
issues are [tracked there](https://git.eeqj.de/sneak/bsdaily/issues).
Changes must always be formatted with a standard `go fmt`, syntactically
valid, and must pass the linting defined in the repository (presently the
`golangci-lint` defaults), which can be run with a `make lint`. The `main`
branch is protected and all changes must be made via [pull
requests](https://git.eeqj.de/sneak/bsdaily/pulls) and pass CI to be merged.
Changes must always be formatted with `make fmt` (`go fmt` for Go, prettier for
Markdown), syntactically valid, and must pass the linting defined in the
repository (presently the `golangci-lint` defaults), which can be run with a
`make lint`. The `main` branch is protected and all changes must be made via
[pull requests](https://git.eeqj.de/sneak/bsdaily/pulls) and pass CI to be
merged.
See [`REPO_POLICIES.md`](REPO_POLICIES.md) for detailed coding standards,
tooling requirements, and workflow conventions.
@@ -50,55 +48,55 @@ tooling requirements, and workflow conventions.
This repository adheres to the
[Scripts to Rule Them All](https://github.com/github/scripts-to-rule-them-all)
standard: normalized scripts in `script/` are the entrypoints for the
development workflow, and the Makefile targets are thin shims that call
them. We provide:
development workflow, and the Makefile targets are thin shims that call them. We
provide:
- `script/bootstrap` — install all development dependencies (go, Go
module download); the linter is not installed on the host
- `script/bootstrap` — install all development dependencies (go, Go module
download, node and yarn at pinned versions when absent, and the prettier
pinned in `package.json` and `yarn.lock`); a node or yarn already installed is
used whatever its version; the linter is not installed on the host
- `script/setup` — make a fresh clone ready for development: runs
`script/bootstrap`, then `script/install-precommit`
- `script/projectname` — print the project name (used for the Docker
image tags)
- `script/test` — build the Dockerfile's `test` phase, which runs the
test suite with `-race` (verbose rerun on failure)
- `script/projectname` — print the project name (used for the Docker image tags)
- `script/test` — build the Dockerfile's `test` phase, which runs the test suite
with `-race` (verbose rerun on failure)
- `script/lint` — build the Dockerfile's `lint` phase, which runs
`golangci-lint run ./...`
- `script/fmt` — format the Go code with `go fmt` (writes)
- `script/fmt-check` — check the Go formatting with `gofmt` (read-only)
- `script/check` — run `script/test`, `script/lint`, and
`script/fmt-check`
- `script/docker` — build the Docker image tagged via
`script/projectname`; the build runs the `lint` and `test` phases
first
- `script/cibuild` — CI entrypoint: runs `script/bootstrap` and
`script/check`, then builds the Docker image tagged via
`script/projectname`
- `script/precommit` — pre-commit gate: `go mod tidy` (must not change
`go.mod` or `go.sum`) and `go fmt`, then `script/check`
- `script/install-precommit` — install the git pre-commit hook that
runs `script/precommit`
- `script/fmt` — format the Go code with `go fmt` and every Markdown file with
prettier (writes)
- `script/fmt-check` — check the Go formatting with `gofmt` and the Markdown
formatting with prettier (read-only)
- `script/check` — run `script/test`, `script/lint`, and `script/fmt-check`
- `script/docker` — build the Docker image tagged via `script/projectname`; the
build runs the `lint` and `test` phases first
- `script/cibuild` — CI entrypoint: runs `script/bootstrap` and `script/check`,
then builds the Docker image tagged via `script/projectname`
- `script/precommit` — pre-commit gate: `go mod tidy` (must not change `go.mod`
or `go.sum`) and `go fmt`, then `script/check`
- `script/install-precommit` — install the git pre-commit hook that runs
`script/precommit`
Every Docker build in `script/` is uncached, so the `lint` and `test`
phases always run rather than being served from the build cache.
Every Docker build in `script/` is uncached, so the `lint` and `test` phases
always run rather than being served from the build cache.
# Problem Statement
A Bluesky firehose ingester writes every observed post (and associated
users, hashtags, URLs, and media references) into a single ever-growing
SQLite database, `firehose.db`. This database has several properties that
make it awkward to publish or archive directly:
A Bluesky firehose ingester writes every observed post (and associated users,
hashtags, URLs, and media references) into a single ever-growing SQLite
database, `firehose.db`. This database has several properties that make it
awkward to publish or archive directly:
- It is **large and always growing**, so re-publishing the whole thing every
day is wasteful.
- It is **large and always growing**, so re-publishing the whole thing every day
is wasteful.
- It is **continuously written**, so reading from it directly risks lock
contention with the live ingester and inconsistent reads.
- It is **monolithic**, so there is no natural unit at which to mirror,
share, or distribute "just yesterday's posts".
- It is **monolithic**, so there is no natural unit at which to mirror, share,
or distribute "just yesterday's posts".
What is wanted instead is a stable, immutable, per-day artifact: a small
file containing exactly one calendar day of firehose data, cheap to publish,
cheap to mirror, and trivially re-mergeable into a full database by anyone
who collects a set of them.
What is wanted instead is a stable, immutable, per-day artifact: a small file
containing exactly one calendar day of firehose data, cheap to publish, cheap to
mirror, and trivially re-mergeable into a full database by anyone who collects a
set of them.
# Proposed Solution
@@ -108,81 +106,81 @@ A tool, `bsdaily`, that:
filesystem, so it reads from a consistent point-in-time copy that the live
ingester cannot be writing to;
- copies the snapshot's database files to fast scratch storage;
- **extracts** a single day's `posts` (and all rows reachable from them) into
a fresh, minimal per-day SQLite database;
- **dumps** that per-day database to SQL and pipes it through multithreaded
zstd compression;
- **verifies** the compressed output (zstd integrity check plus a sanity
check that the decompressed stream actually looks like SQL);
- **extracts** a single day's `posts` (and all rows reachable from them) into a
fresh, minimal per-day SQLite database;
- **dumps** that per-day database to SQL and pipes it through multithreaded zstd
compression;
- **verifies** the compressed output (zstd integrity check plus a sanity check
that the decompressed stream actually looks like SQL);
- **publishes** the result atomically as
`DailiesBase/YYYY-MM/YYYY-MM-DD.sql.zst`.
Each daily dump is emitted with `INSERT` statements over the full schema
(including the deduplicated `users`, `hashtags`, and `urls` lookup tables),
so any collection of daily dumps can be merged into a single database by
rewriting `INSERT INTO` to `INSERT OR IGNORE INTO` and replaying them in
sequence. Two helper scripts ([`merge_daily_dumps.sh`](merge_daily_dumps.sh)
and [`regenerate_auxiliary_tables.sql`](regenerate_auxiliary_tables.sql)) are
(including the deduplicated `users`, `hashtags`, and `urls` lookup tables), so
any collection of daily dumps can be merged into a single database by rewriting
`INSERT INTO` to `INSERT OR IGNORE INTO` and replaying them in sequence. Two
helper scripts ([`merge_daily_dumps.sh`](merge_daily_dumps.sh) and
[`regenerate_auxiliary_tables.sql`](regenerate_auxiliary_tables.sql)) are
included to do exactly this and to rebuild the aggregate statistics
(`use_count`, `first_seen`, user `resolved_at`/`updated_at`) afterward.
# Design Goals
- **Never disturb the live ingester.** All reads come from a ZFS snapshot,
never the live database.
- **Never disturb the live ingester.** All reads come from a ZFS snapshot, never
the live database.
- **Crash-safe, idempotent runs.** Output is written to a temp file and
atomically renamed; a day whose final output already exists is skipped, so
re-running a range is safe and resumable.
- **Mergeable output.** Daily dumps re-combine losslessly into a full
database via `INSERT OR IGNORE`.
- **Mergeable output.** Daily dumps re-combine losslessly into a full database
via `INSERT OR IGNORE`.
- **Operationally cautious.** Free-space preflight checks on both scratch and
output filesystems; explicit verification of every artifact before it is
published.
- **Fast where it's free.** Large snapshot copies use a 256MiB buffer,
pre-allocate the destination, and (on Linux) issue `posix_fadvise`
sequential/willneed hints; extraction uses aggressive,
crash-unsafe-by-design SQLite pragmas because the working data lives only
in disposable scratch space.
sequential/willneed hints; extraction uses aggressive, crash-unsafe-by-design
SQLite pragmas because the working data lives only in disposable scratch
space.
# Non-Goals
- **Real-time export.** `bsdaily` operates on daily snapshots; the freshest
day it can produce is the snapshot date minus one.
- **Schema ownership.** The schema is defined by the upstream firehose
ingester; [`schema.sql`](schema.sql) is included for reference only.
`bsdaily` copies whatever table and index DDL it finds in the source.
- **Cross-platform deployment.** It is built and run on Linux (the
free-space check and fadvise hints use `golang.org/x/sys/unix`; a non-Linux
build compiles but is a no-op for the fadvise hints). The hard-coded paths
assume the production host's ZFS layout.
- **Real-time export.** `bsdaily` operates on daily snapshots; the freshest day
it can produce is the snapshot date minus one.
- **Schema ownership.** The schema is defined by the upstream firehose ingester;
[`schema.sql`](schema.sql) is included for reference only. `bsdaily` copies
whatever table and index DDL it finds in the source.
- **Cross-platform deployment.** It is built and run on Linux (the free-space
check and fadvise hints use `golang.org/x/sys/unix`; a non-Linux build
compiles but is a no-op for the fadvise hints). The hard-coded paths assume
the production host's ZFS layout.
# How It Works
A single run proceeds as follows:
1. **Find the snapshot.** Scan `SnapshotBase` for directories matching
`zfs-auto-snap_daily-YYYY-MM-DD-NNNN`, pick the most recent, and confirm
it contains `firehose.db`.
2. **Determine target days.** Default to the snapshot date minus one day;
or use `--date`, or every day in the inclusive `--from`/`--to` range.
`zfs-auto-snap_daily-YYYY-MM-DD-NNNN`, pick the most recent, and confirm it
contains `firehose.db`.
2. **Determine target days.** Default to the snapshot date minus one day; or use
`--date`, or every day in the inclusive `--from`/`--to` range.
3. **Preflight disk space.** Require at least 500GiB free on the scratch
filesystem and 20GiB free on the output filesystem.
4. **Copy the database to scratch.** Copy `firehose.db`, its `-wal`, and (if
present) its `-shm` from the snapshot into a fresh temp directory under
`TmpBase`.
5. **Per day**, processed strictly one at a time to avoid SQLite contention:
- skip the day if its final output file already exists;
- `ATTACH` the copied source DB to a new empty per-day DB, recreate the
table DDL, and `INSERT ... SELECT` the target day's `posts` plus all
rows reachable from them (`posts_hashtags`, `posts_urls`, `hashtags`,
`urls`, `users`, and `media` if that table exists);
- abort the day cleanly if there are zero posts (`ErrNoPosts`), rather
than emitting an empty dump;
- recreate indexes, detach the source, and verify the inserted row count;
- `sqlite3 .dump | zstdmt` into a hidden temp file;
- run a zstd integrity check and confirm the decompressed head looks like
SQL;
- atomically rename into place and delete the per-day scratch DB.
- skip the day if its final output file already exists;
- `ATTACH` the copied source DB to a new empty per-day DB, recreate the
table DDL, and `INSERT ... SELECT` the target day's `posts` plus all rows
reachable from them (`posts_hashtags`, `posts_urls`, `hashtags`, `urls`,
`users`, and `media` if that table exists);
- abort the day cleanly if there are zero posts (`ErrNoPosts`), rather than
emitting an empty dump;
- recreate indexes, detach the source, and verify the inserted row count;
- `sqlite3 .dump | zstdmt` into a hidden temp file;
- run a zstd integrity check and confirm the decompressed head looks like
SQL;
- atomically rename into place and delete the per-day scratch DB.
6. **Clean up** the temp directory and log a processed/skipped/total summary.
# Usage
@@ -219,9 +217,9 @@ sqlite3 merged.db < regenerate_auxiliary_tables.sql
- **Go** (see [`go.mod`](go.mod) for the toolchain version) to build.
- **Linux** for production use (ZFS snapshots, `statfs` free-space checks,
`posix_fadvise` hints).
- The **`sqlite3`** and **`zstdmt`** (multithreaded zstd) binaries on
`PATH`; `zstdcat` is used for verification. SQLite reads/writes during
extraction use the pure-Go [`modernc.org/sqlite`](https://pkg.go.dev/modernc.org/sqlite)
- The **`sqlite3`** and **`zstdmt`** (multithreaded zstd) binaries on `PATH`;
`zstdcat` is used for verification. SQLite reads/writes during extraction use
the pure-Go [`modernc.org/sqlite`](https://pkg.go.dev/modernc.org/sqlite)
driver, so no cgo is required for that part.
# Configuration
@@ -241,21 +239,21 @@ elsewhere.
# Data Model
The firehose schema (reference copy in [`schema.sql`](schema.sql)) centers
on a `posts` table, with `users` keyed by DID and many-to-many junction
tables linking posts to deduplicated `hashtags` and `urls`. An optional
`media` table tracks downloaded blobs by content hash. `bsdaily` does not
own this schema; it reflects whatever DDL exists in the source snapshot and
selects forward from `posts` along the foreign-key relationships to produce a
referentially-complete per-day slice.
The firehose schema (reference copy in [`schema.sql`](schema.sql)) centers on a
`posts` table, with `users` keyed by DID and many-to-many junction tables
linking posts to deduplicated `hashtags` and `urls`. An optional `media` table
tracks downloaded blobs by content hash. `bsdaily` does not own this schema; it
reflects whatever DDL exists in the source snapshot and selects forward from
`posts` along the foreign-key relationships to produce a referentially-complete
per-day slice.
# Use Cases
## Daily public archive
Publish one small, immutable file per day to static HTTP (or IPFS, or a
mirror network) so that anyone can fetch exactly the day(s) they want and
re-merge them locally.
Publish one small, immutable file per day to static HTTP (or IPFS, or a mirror
network) so that anyone can fetch exactly the day(s) they want and re-merge them
locally.
## Backfilling a range
@@ -265,16 +263,17 @@ safe to re-run.
## Reconstituting a full database
Collect any set of daily dumps and merge them with `INSERT OR IGNORE` to
rebuild a complete, queryable SQLite database, then regenerate the aggregate
statistics tables.
Collect any set of daily dumps and merge them with `INSERT OR IGNORE` to rebuild
a complete, queryable SQLite database, then regenerate the aggregate statistics
tables.
# See Also
## Links
- Repo: [https://git.eeqj.de/sneak/bsdaily](https://git.eeqj.de/sneak/bsdaily)
- Issues: [https://git.eeqj.de/sneak/bsdaily/issues](https://git.eeqj.de/sneak/bsdaily/issues)
- Issues:
[https://git.eeqj.de/sneak/bsdaily/issues](https://git.eeqj.de/sneak/bsdaily/issues)
- Bluesky: [https://bsky.app](https://bsky.app)
- zstd: [https://facebook.github.io/zstd/](https://facebook.github.io/zstd/)
+30 -27
View File
@@ -1,12 +1,12 @@
# Workflow
* branch (from `main`)
* do the work in Next Step
* move Next Step to the top of Completed Steps
* move the top item of Future Steps into Next Step
* commit (`TODO.md` changes in the same commit as the work)
* merge to `main` if the branch is not protected, otherwise open a PR
* push
- branch (from `main`)
- do the work in Next Step
- move Next Step to the top of Completed Steps
- move the top item of Future Steps into Next Step
- commit (`TODO.md` changes in the same commit as the work)
- merge to `main` if the branch is not protected, otherwise open a PR
- push
# Status
@@ -14,35 +14,38 @@ pre-1.0
# Next Step
Add the canonical `.golangci.yml`, move the lint phase to golangci-lint
v2.14.0 in the same commit, and fix the findings it surfaces
Add the canonical `.golangci.yml`, move the lint phase to golangci-lint v2.14.0
in the same commit, and fix the findings it surfaces
(https://git.eeqj.de/sneak/bsdaily/issues/6).
# Completed Steps
- 2026-10-06: Formatted Markdown with prettier: `script/fmt` writes and
`script/fmt-check` checks every Markdown file; prettier pinned in
`package.json` and `yarn.lock`, installed by `script/bootstrap`, which
installs node and yarn at pinned versions when absent and otherwise uses the
ones already installed; existing Markdown reformatted.
- 2026-10-05: Brought the repo up to the standard layout: canonical
`.gitignore`, `.dockerignore` and `.editorconfig`; `lint` and `test`
phases in the `Dockerfile`, built by `script/lint` and `script/test`;
canonical `script/cibuild`, `script/docker` and CI workflow; no linter
installed on the host; re-vendored `REPO_POLICIES.md`.
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints,
Makefile shims, README Entrypoints section
- 2026-06-28: Fixed errcheck lint failures; added compilation smoke test;
tidied go.mod.
- 2026-06-28: Added repo scaffolding: README, LICENSE, Makefile,
Dockerfile, REPO_POLICIES.md, and Gitea CI.
- 2026-02-12: Fixed SQLite database locking by removing parallel
processing; fixed Linux build via golang.org/x/sys/unix Fadvise.
- 2026-02-12: Optimized file copy for large databases; moved temp
directory to NVMe scratch storage.
`.gitignore`, `.dockerignore` and `.editorconfig`; `lint` and `test` phases in
the `Dockerfile`, built by `script/lint` and `script/test`; canonical
`script/cibuild`, `script/docker` and CI workflow; no linter installed on the
host; re-vendored `REPO_POLICIES.md`.
- 2026-07-07 Adopted scripts-to-rule-them-all: `script/` entrypoints, Makefile
shims, README Entrypoints section
- 2026-06-28: Fixed errcheck lint failures; added compilation smoke test; tidied
go.mod.
- 2026-06-28: Added repo scaffolding: README, LICENSE, Makefile, Dockerfile,
REPO_POLICIES.md, and Gitea CI.
- 2026-02-12: Fixed SQLite database locking by removing parallel processing;
fixed Linux build via golang.org/x/sys/unix Fadvise.
- 2026-02-12: Optimized file copy for large databases; moved temp directory to
NVMe scratch storage.
- 2026-02-11: Added date range support.
- 2026-02-09: Initial implementation: single-day extraction, specific-date
targeting, faster pruning of throwaway database copies.
# Future Steps
- Format Markdown with prettier in `script/fmt` and `script/fmt-check`
(https://git.eeqj.de/sneak/bsdaily/issues/7).
- Expand tests beyond the compilation smoke test: unit tests for the
extraction, verification, and atomic-publish paths.
- Expand tests beyond the compilation smoke test: unit tests for the extraction,
verification, and atomic-publish paths.
- Cut a first SemVer release once compliance and test coverage land.
+5
View File
@@ -0,0 +1,5 @@
{
"devDependencies": {
"prettier": "3.8.1"
}
}
+79 -1
View File
@@ -3,11 +3,20 @@
# this repo. Idempotent: every install is guarded by a check so already
# installed tools are skipped. Base tooling comes from nix, apt, brew,
# or apk (detected in that order); assumes NOTHING is present (not git,
# make, or go).
# make, go, or node). Node is used directly if installed; otherwise it
# is installed at a pinned version via nvm (installing nvm itself first,
# from a hash-verified release archive, never curl | sh).
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Pinned versions, 2026-07-06
NODE_VERSION="22.17.0"
NVM_VERSION="0.40.3"
# sha256 of https://github.com/nvm-sh/nvm/archive/refs/tags/v0.40.3.tar.gz
NVM_SHA256="5f4d6aaa04a177dc93c985e31dbc411ab6b8c6e1e21d8015dbc1372625fcd1d0"
YARN_VERSION="1.22.22"
PKGMGR=""
SUDO=""
APT_UPDATED=""
@@ -56,6 +65,69 @@ missing() {
! command -v "$1" >/dev/null 2>&1
}
# verify_sha256 <file> <expected-hash>
verify_sha256() {
if command -v sha256sum >/dev/null 2>&1; then
actual="$(sha256sum "$1" | cut -d' ' -f1)"
else
actual="$(shasum -a 256 "$1" | cut -d' ' -f1)"
fi
if [ "$actual" != "$2" ]; then
echo "bootstrap: sha256 mismatch for $1" >&2
echo " expected: $2" >&2
echo " actual: $actual" >&2
exit 1
fi
}
# nvm is a bash script; run a command in a bash with nvm loaded
nvm_sh() {
bash -c ". \"\$HOME/.nvm/nvm.sh\" && $*"
}
ensure_nvm() {
[ -s "$HOME/.nvm/nvm.sh" ] && return 0
# nvm prerequisites; nvm itself requires bash
if missing bash; then pkg_install bash bash bash bash; fi
if missing curl; then pkg_install curl curl curl curl; fi
if missing git; then pkg_install git git git git; fi
tmp="$(mktemp -d)"
curl -fsSL -o "$tmp/nvm.tar.gz" \
"https://github.com/nvm-sh/nvm/archive/refs/tags/v${NVM_VERSION}.tar.gz"
verify_sha256 "$tmp/nvm.tar.gz" "$NVM_SHA256"
mkdir -p "$HOME/.nvm"
tar -xzf "$tmp/nvm.tar.gz" -C "$HOME/.nvm" --strip-components=1
rm -rf "$tmp"
}
ensure_node() {
if ! missing node; then return 0; fi
ensure_nvm
nvm_sh "nvm install $NODE_VERSION"
}
ensure_yarn() {
if ! missing yarn; then return 0; fi
if ! missing corepack; then
corepack enable
corepack prepare "yarn@$YARN_VERSION" --activate
elif [ -s "$HOME/.nvm/nvm.sh" ]; then
nvm_sh "nvm use $NODE_VERSION >/dev/null && corepack enable && \
corepack prepare yarn@$YARN_VERSION --activate"
else
npm install -g "yarn@$YARN_VERSION"
fi
}
install_js_deps() {
if missing yarn && [ -s "$HOME/.nvm/nvm.sh" ]; then
nvm_sh "nvm use $NODE_VERSION >/dev/null && cd \"$ROOT\" && \
yarn install --frozen-lockfile"
else
yarn install --frozen-lockfile
fi
}
main() {
cd "$ROOT"
@@ -68,6 +140,12 @@ main() {
go mod download
# Node and yarn, then the prettier pinned in package.json and
# yarn.lock, which script/fmt and script/fmt-check run
ensure_node
ensure_yarn
install_js_deps
echo "bootstrap complete"
}
+23 -1
View File
@@ -1,12 +1,34 @@
#!/bin/sh
# script/fmt: format all files (writes).
# script/fmt: format all files (writes): the Go code with go fmt and
# every Markdown file with prettier.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Must match the pin in script/bootstrap.
NODE_VERSION="22.17.0"
# script/bootstrap installs node and yarn under nvm and leaves neither
# on the PATH of the shell that called it, so resolve the pinned
# toolchain here the way bootstrap's own install step does. nvm is a
# bash script, hence the subshell.
run_yarn() {
if command -v yarn >/dev/null 2>&1; then
exec yarn "$@"
fi
if [ ! -s "$HOME/.nvm/nvm.sh" ]; then
echo "fmt: no yarn; run script/bootstrap first" >&2
exit 1
fi
exec bash -c '. "$HOME/.nvm/nvm.sh" && nvm use "$1" >/dev/null &&
shift && exec yarn "$@"' bash "$NODE_VERSION" "$@"
}
main() {
cd "$ROOT"
go fmt ./...
# run_yarn replaces this shell, so it stays the last step.
run_yarn run prettier --write '**/*.md' --tab-width 4 --prose-wrap always
}
main "$@"
+23 -1
View File
@@ -1,10 +1,30 @@
#!/bin/sh
# script/fmt-check: check formatting (read-only). Same scope as
# script/fmt, but fails instead of writing.
# script/fmt, the Go code and every Markdown file, but fails instead of
# writing.
set -eu
ROOT="$(cd "$(dirname "$0")/.." && pwd -P)"
# Must match the pin in script/bootstrap.
NODE_VERSION="22.17.0"
# script/bootstrap installs node and yarn under nvm and leaves neither
# on the PATH of the shell that called it, so resolve the pinned
# toolchain here the way bootstrap's own install step does. nvm is a
# bash script, hence the subshell.
run_yarn() {
if command -v yarn >/dev/null 2>&1; then
exec yarn "$@"
fi
if [ ! -s "$HOME/.nvm/nvm.sh" ]; then
echo "fmt-check: no yarn; run script/bootstrap first" >&2
exit 1
fi
exec bash -c '. "$HOME/.nvm/nvm.sh" && nvm use "$1" >/dev/null &&
shift && exec yarn "$@"' bash "$NODE_VERSION" "$@"
}
main() {
cd "$ROOT"
unformatted="$(gofmt -l .)"
@@ -13,6 +33,8 @@ main() {
echo "$unformatted" >&2
exit 1
fi
# run_yarn replaces this shell, so it stays the last step.
run_yarn run prettier --check '**/*.md' --tab-width 4 --prose-wrap always
}
main "$@"
+8
View File
@@ -0,0 +1,8 @@
# THIS IS AN AUTOGENERATED FILE. DO NOT EDIT THIS FILE DIRECTLY.
# yarn lockfile v1
prettier@3.8.1:
version "3.8.1"
resolved "https://registry.yarnpkg.com/prettier/-/prettier-3.8.1.tgz#edf48977cf991558f4fcbd8a3ba6015ba2a3a173"
integrity sha512-UOnG6LftzbdaHZcKoPFtOcCKztrQ57WkHDeRD9t/PTQtmT0NHSeWWepj6pS0z/N7+08BHFDQVUrfmfMRcZwbMg==