Files
hnblogs/README.md
sneak 7e1b010c4f embed blogs.json instead of fetching it at runtime (closes #1)
The dataset is vendored as blogs.json and compiled in with go:embed, so
the package does no network I/O on any code path.

FetchBlogs is removed rather than kept as a no-op wrapper, since nothing
is fetched any more: GetBlogs decodes the embedded bytes on first call
and memoizes the result, and RandomBlog, RandomBlogs and NthBlog go
through it. Callers of FetchBlogs have to switch to GetBlogs.

make update-data refreshes the vendored copy from the BlogsURL constant,
downloading to a temporary file and replacing blogs.json only once that
file parses as a non-empty JSON array of blog entries, so a
complete-but-wrong response cannot overwrite the good dataset. A test
walks the dependency graph of the non-test build and fails if any net/*
package is reachable.

Model: opus-5
2026-09-05 04:37:58 +00:00

1.5 KiB

hnblogs

A Go library for the blogs.hn dataset: a list of personal blogs collected from Hacker News.

import "sneak.berlin/go/hnblogs"

blog, err := hnblogs.RandomBlog()

Embedded data

The dataset is vendored into this repository as blogs.json and compiled into the package with go:embed. The library performs no network I/O: importing it does not reach out to anything, results do not change under a caller between runs of the same build, and go test works offline.

GetBlogs decodes the embedded bytes on first call and memoizes the result; every other accessor goes through it. Its error return is only reachable if the committed blogs.json is malformed.

The trade-off is that the dataset is a build-time artifact: it is roughly 8 MB of JSON, it lands in every binary that links the package, and it is only as fresh as the last commit that refreshed it.

Refreshing the dataset

make update-data

That target reads the upstream location from the BlogsURL constant in hnblogs.go — the single source of truth — and downloads it to a temporary file. It replaces blogs.json only once that file parses as a non-empty JSON array of blog entries, so a response that arrives complete but is not the dataset leaves the vendored copy untouched. It then runs the test suite against the new data. Requires curl and jq. Commit the resulting blogs.json to publish the update.

Development

make test
make lint
make docker

make docker runs lint and tests in containers.