Compare commits

..

1 Commits

Author SHA1 Message Date
d5ed473e31 embed blogs.json instead of fetching it at runtime (closes #1)
The library needed network access on first use and its results changed
under the caller between runs. blogs.json is now vendored and compiled in
with go:embed, so the dataset is fixed for a given build.

FetchBlogs keeps its name, signature and sync.Once memoization but now
decodes the embedded bytes; its error return is only reachable if the
committed blogs.json is malformed. net/http is gone from the package, and
a test asserts that no net/* package appears anywhere in the dependency
graph of the non-test build, not only in its direct imports.

make update-data refreshes the vendored file, reading the upstream
location from the BlogsURL constant so the URL has one definition. It
downloads to a temporary file and replaces blogs.json only once that file
parses as a non-empty JSON array of blog entries, so neither a truncated
transfer nor a complete-but-wrong response such as an error page can
overwrite the good dataset. It then runs the test suite against the new
data. That target now needs jq as well as curl.

blogs.json is an unmodified copy of a third party's file, redistributed
here and in every binary that links the package, and the upstream
repository publishes no licence. The README and the embed doc comment now
record where it came from and that this repository's LICENSE does not
extend to it. Whether that arrangement is acceptable is the owner's call.

The dataset is committed verbatim as upstream serves it, which is ~8 MB
of JSON in the repo and in every linking binary; most of that is per-blog
post history that the Blog struct does not expose.

Model: opus-5
2026-09-05 03:39:02 +00:00
4 changed files with 32 additions and 14 deletions

View File

@@ -31,7 +31,7 @@ update-data:
@command -v jq >/dev/null || { echo "update-data requires jq" >&2; exit 1; }
@url=$$(sed -n 's/^const BlogsURL = "\(.*\)"$$/\1/p' hnblogs.go); \
test -n "$$url" || { echo "could not parse BlogsURL from hnblogs.go" >&2; exit 1; }; \
echo "downloading $$url"; \
echo "fetching $$url"; \
curl -fsSL "$$url" -o blogs.json.tmp
@jq -e 'type == "array" and length > 0 and all(.[]; type == "object" and has("url"))' \
blogs.json.tmp >/dev/null 2>&1 \

View File

@@ -16,14 +16,25 @@ the package with `go:embed`. The library performs no network I/O: importing it
does not reach out to anything, results do not change under a caller between
runs of the same build, and `go test` works offline.
`GetBlogs` decodes the embedded bytes on first call and memoizes the result;
every other accessor goes through it. Its error return is only reachable if the
committed `blogs.json` is malformed.
`FetchBlogs` keeps its name and its `sync.Once` memoization, but on first call
it decodes the embedded bytes rather than issuing an HTTP request. Its error
return is now only reachable if the committed `blogs.json` is malformed.
The trade-off is that the dataset is a build-time artifact: it is roughly 8 MB
of JSON, it lands in every binary that links the package, and it is only as
fresh as the last commit that refreshed it.
### Where blogs.json came from
`blogs.json` is not this project's work. It is an unmodified copy of
<https://raw.githubusercontent.com/surprisetalk/blogs.hn/main/blogs.json>
from the [surprisetalk/blogs.hn](https://github.com/surprisetalk/blogs.hn)
repository, which publishes no licence. This repository's `LICENSE` covers the
code here and does not extend to `blogs.json`, and a binary that links this
package redistributes that file too.
## Refreshing the dataset
```sh

View File

@@ -14,11 +14,13 @@ import (
"sync"
)
// BlogsURL is the upstream source of blogs.json. Nothing reads it at runtime;
// it documents where "make update-data" downloads the vendored copy from.
// BlogsURL is the upstream source of blogs.json. It is not fetched at runtime;
// it documents where "make update-data" pulls the vendored copy from.
const BlogsURL = "https://raw.githubusercontent.com/surprisetalk/blogs.hn/main/blogs.json"
// blogsJSON is the vendored dataset, refreshed by "make update-data".
// blogsJSON is the vendored dataset, refreshed by "make update-data". It is an
// unmodified copy of a third party's file, not this project's work; see "Where
// blogs.json came from" in README.md.
//
//go:embed blogs.json
var blogsJSON []byte
@@ -39,13 +41,13 @@ type Blog struct {
Desc string `json:"desc"`
}
// GetBlogs returns the embedded list of blogs, decoding it on first call and
// FetchBlogs returns the embedded list of blogs, decoding it on first call and
// memoizing the result for subsequent calls.
//
// It performs no I/O: the data is compiled into the binary, so the only error
// it can return is a malformed embedded blogs.json, which would mean the
// committed dataset is broken.
func GetBlogs() ([]Blog, error) {
// Despite the name it performs no I/O: the data is compiled into the binary, so
// the only error it can return is a malformed embedded blogs.json, which would
// mean the committed dataset is broken.
func FetchBlogs() ([]Blog, error) {
once.Do(func() {
var decoded []Blog
if err := json.Unmarshal(blogsJSON, &decoded); err != nil {
@@ -59,6 +61,11 @@ func GetBlogs() ([]Blog, error) {
return blogs, loadError
}
// GetBlogs returns the memoized list of blogs.
func GetBlogs() ([]Blog, error) {
return FetchBlogs()
}
// RandomBlog returns a random blog from the list of blogs.
func RandomBlog() (Blog, error) {
blogs, err := GetBlogs()

View File

@@ -8,8 +8,8 @@ import (
"testing"
)
func TestGetBlogs(t *testing.T) {
blogs, err := GetBlogs()
func TestFetchBlogs(t *testing.T) {
blogs, err := FetchBlogs()
if err != nil {
t.Fatalf("Expected no error, got %v", err)
}