Compare commits
1
Commits
next
..
5f1971a293
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
5f1971a293 |
@@ -1,33 +1,30 @@
|
|||||||
NetWatch is an MIT-licensed JavaScript single-page application by
|
NetWatch is an MIT-licensed JavaScript single-page application by
|
||||||
[@sneak](https://sneak.berlin) that provides real-time network latency
|
[@sneak](https://sneak.berlin) that provides real-time network latency
|
||||||
monitoring to common internet hosts, displayed with color-coded figures and
|
monitoring to common internet hosts, displayed with color-coded figures and
|
||||||
sparkline graphs, served from a static bucket or from its Docker image, where a
|
sparkline graphs, served from a static bucket or Docker container.
|
||||||
small Go backend stores the measurements the page reports.
|
|
||||||
|
|
||||||
## Getting Started
|
## Getting Started
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
# Install the dependencies and the git pre-commit hook
|
# Install dependencies
|
||||||
make setup
|
yarn install
|
||||||
|
|
||||||
# Run the page on the Vite dev server
|
# Development server
|
||||||
make dev
|
yarn dev
|
||||||
|
|
||||||
# Run the tests, both linters and the format check
|
# Production build
|
||||||
make check
|
yarn build
|
||||||
|
|
||||||
# Build the page into dist/
|
# Preview production build
|
||||||
make build
|
yarn preview
|
||||||
|
|
||||||
# Build the image and run it
|
# Docker
|
||||||
make docker
|
docker build -t netwatch .
|
||||||
docker run -p 8080:8080 netwatch
|
docker run -p 8080:8080 netwatch
|
||||||
```
|
```
|
||||||
|
|
||||||
`make check` and `make docker` need Docker. `make dev` passes `/api` to
|
`yarn dev` proxies `/api` to `http://127.0.0.1:8080`, so a locally running
|
||||||
`http://127.0.0.1:8080`, where `make run` in `backend/` starts `netwatch-server`
|
`netwatch-server` (see `backend/`) receives the reports the page posts.
|
||||||
with its defaults, so the reports the page posts are stored in
|
|
||||||
`backend/data/reports`.
|
|
||||||
|
|
||||||
## Entrypoints
|
## Entrypoints
|
||||||
|
|
||||||
@@ -97,25 +94,22 @@ halves, so the root `make check` fails if either one is broken. We provide:
|
|||||||
The narrow-viewport layout lives in the `max-width: 768px` media block in
|
The narrow-viewport layout lives in the `max-width: 768px` media block in
|
||||||
`src/styles.css`. It is verified automatically by `make frontend-viewport-test`,
|
`src/styles.css`. It is verified automatically by `make frontend-viewport-test`,
|
||||||
which drives a digest-pinned headless Chrome against the built `dist/` and
|
which drives a digest-pinned headless Chrome against the built `dist/` and
|
||||||
asserts on computed layout at widths derived from that CSS — on every breakpoint
|
asserts on computed layout at widths derived from that CSS — one pixel either
|
||||||
it declares and one pixel either side of it, plus a 320px floor, a desktop
|
side of every breakpoint it declares, plus a 320px floor, a desktop baseline and
|
||||||
baseline and two landscape sizes. See
|
two landscape sizes. See [test/viewport/README.md](test/viewport/README.md) for
|
||||||
[test/viewport/README.md](test/viewport/README.md) for what it covers and what
|
what it covers and what it genuinely cannot.
|
||||||
it genuinely cannot.
|
|
||||||
|
|
||||||
## Rationale
|
## Rationale
|
||||||
|
|
||||||
When debugging network issues, it's useful to have a persistent at-a-glance view
|
When debugging network issues, it's useful to have a persistent at-a-glance view
|
||||||
of latency and reachability to multiple well-known internet endpoints. NetWatch
|
of latency and reachability to multiple well-known internet endpoints. NetWatch
|
||||||
provides this as a single page that does all its measuring in the browser, so it
|
provides this as a zero-dependency SPA that can be deployed anywhere static
|
||||||
can be served from anywhere static files are served. The backend in its Docker
|
files are served, with no backend required.
|
||||||
image only stores the measurements the page reports; without it, the page works
|
|
||||||
the same and nothing is stored.
|
|
||||||
|
|
||||||
## Design
|
## Design
|
||||||
|
|
||||||
The page is built with Vite and Tailwind CSS v4. Its code is all in
|
The application is a single-page app built with Vite and Tailwind CSS v4. All
|
||||||
`src/main.js`, with a class-based architecture:
|
code lives in `src/main.js` with a class-based architecture:
|
||||||
|
|
||||||
- **`CONFIG`**: Configuration object (update interval, timeouts, axis ticks,
|
- **`CONFIG`**: Configuration object (update interval, timeouts, axis ticks,
|
||||||
etc.). The interval menu sets `updateInterval`, the one value the page writes
|
etc.). The interval menu sets `updateInterval`, the one value the page writes
|
||||||
@@ -131,10 +125,10 @@ The page is built with Vite and Tailwind CSS v4. Its code is all in
|
|||||||
`updateSummary()` / `updateHealthBox()` handle incremental updates
|
`updateSummary()` / `updateHealthBox()` handle incremental updates
|
||||||
- **`tick()`**: Main loop — measures all hosts in parallel, pushing each host's
|
- **`tick()`**: Main loop — measures all hosts in parallel, pushing each host's
|
||||||
sample and redrawing its row as soon as its check ends, then redraws every
|
sample and redrawing its row as soon as its check ends, then redraws every
|
||||||
row, the summary and the health box once the last check ends. The first round,
|
row, the summary and the health box once the last check ends. The rows are
|
||||||
after loading or an interval change, is discarded. The rows are sorted when
|
sorted then too, after the first round that is not discarded and every tenth
|
||||||
the last check ends in round 2, the first one kept, and in rounds 11, 21, 31
|
round after that. When paused, pushes blank markers (no probes, no false
|
||||||
and so on. When paused, pushes blank markers (no probes, no false outage)
|
outage)
|
||||||
- **`Reporter`**: Posts collected samples to the backend
|
- **`Reporter`**: Posts collected samples to the backend
|
||||||
|
|
||||||
### Reporting
|
### Reporting
|
||||||
@@ -148,35 +142,12 @@ delivered report, and while paused nothing is sent. Delivery failure is quiet
|
|||||||
one debug-log line per outage, retried at the next interval, never blocking
|
one debug-log line per outage, retried at the next interval, never blocking
|
||||||
probing. The report-building step is a pure function of host state.
|
probing. The report-building step is a pure function of host state.
|
||||||
|
|
||||||
### Backend
|
|
||||||
|
|
||||||
`netwatch-server`, in `backend/`, is a small Go HTTP server that stores the
|
|
||||||
reports the page posts. It keeps them in memory and writes them to `DATA_DIR` as
|
|
||||||
zstd-compressed files of JSON lines: every minute, whenever 10 MiB are waiting,
|
|
||||||
and when it stops. Its routes:
|
|
||||||
|
|
||||||
- `POST /api/v1/reports` — takes a report, without credentials; each client
|
|
||||||
address may send a limited number a minute, and the report files are capped in
|
|
||||||
size, the oldest deleted first
|
|
||||||
- `GET /.well-known/healthcheck` — answers 200 with `"status":"ok"`, the
|
|
||||||
server's version and its uptime
|
|
||||||
- `GET /metrics` — Prometheus metrics behind basic auth, only when
|
|
||||||
`METRICS_USERNAME` and `METRICS_PASSWORD` are set; each client address may
|
|
||||||
make a limited number of requests to it a minute
|
|
||||||
|
|
||||||
In the image, the `builder` stage of `Dockerfile` tests it and builds it with
|
|
||||||
`backend/script/build`, and `bin/entrypoint.sh` runs it as user `netwatch` on
|
|
||||||
`127.0.0.1:8081`, behind nginx. Outside the image, `make run` in `backend/`
|
|
||||||
builds it and runs it on port 8080. Its settings, report storage and limits are
|
|
||||||
in [backend/README.md](backend/README.md).
|
|
||||||
|
|
||||||
### Monitoring targets
|
### Monitoring targets
|
||||||
|
|
||||||
- **26 WAN hosts**: datavi.be (pinned at start), Anthropic API, OpenAI API, AWS
|
- **22 WAN hosts**: datavi.be, Anthropic API, OpenAI API, AWS Console, GCP
|
||||||
Console, Google Cloud Console, Microsoft Azure, Cloudflare, Fastly CDN,
|
Console, Azure, Cloudflare, Fastly, Akamai, GitHub, B2, 7 S3 regional
|
||||||
Akamai, Google, GitHub, B2, 8 S3 regional endpoints (Cape Town, London,
|
endpoints (Cape Town, London, Bahrain, Tokyo, Sydney, Oregon, São Paulo), 4
|
||||||
Bahrain, Tokyo, Singapore, Sydney, Oregon, São Paulo) and 6 Hetzner speed test
|
GCS locational endpoints (Iowa, Belgium, Singapore, Sydney)
|
||||||
servers (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore)
|
|
||||||
- **Local CPE**: Cable modem at 192.168.100.1 (always monitored)
|
- **Local CPE**: Cable modem at 192.168.100.1 (always monitored)
|
||||||
- **Local Gateway**: Auto-detected on startup by probing common default gateway
|
- **Local Gateway**: Auto-detected on startup by probing common default gateway
|
||||||
addresses (192.168.1.1, 192.168.0.1, 192.168.8.1, 10.0.0.1); first responder
|
addresses (192.168.1.1, 192.168.0.1, 192.168.8.1, 10.0.0.1); first responder
|
||||||
@@ -188,17 +159,15 @@ Local hosts are tracked separately from WAN stats.
|
|||||||
|
|
||||||
### Latency measurement
|
### Latency measurement
|
||||||
|
|
||||||
GET requests with `mode: 'no-cors'`, `cache: 'no-store'` and a cache-busting
|
HEAD requests with `mode: 'no-cors'` and `cache: 'no-store'`, timed with
|
||||||
query parameter, timed with `performance.now()`. Each check times out after 80%
|
`performance.now()`. Each check times out after 80% of the refresh interval (24
|
||||||
of the refresh interval (24 seconds at 30 seconds) and is then recorded as a
|
seconds at 30 seconds) and is then recorded as a timeout, so a round's checks
|
||||||
timeout, so a round's checks have all finished before the next round is due.
|
have all finished before the next round is due. When no WAN host answers, a
|
||||||
When no WAN host answers, a recovery probe checks 4 WAN hosts, picked at random
|
recovery probe checks 4 random WAN hosts every half second, giving up the checks
|
||||||
when it starts, every half second, giving up the checks it started half a second
|
it started half a second before. As soon as one answers, a new round starts at
|
||||||
before. As soon as one answers, a new round starts at once, as it does after an
|
once, as it does after an interval change. A round started early gives up the
|
||||||
interval change. A round started early gives up the last round's checks if they
|
last round's checks if they are still waiting, and that round records nothing
|
||||||
are still waiting, and that round records nothing more, so rounds never overlap.
|
more, so rounds never overlap. IPv4 only.
|
||||||
The browser chooses between IPv4 and IPv6 for each target, as for any request;
|
|
||||||
the local targets are IPv4 addresses.
|
|
||||||
|
|
||||||
### Color coding
|
### Color coding
|
||||||
|
|
||||||
@@ -223,42 +192,25 @@ dist/
|
|||||||
|
|
||||||
## Features
|
## Features
|
||||||
|
|
||||||
- A round of checks every 3 seconds by default; the interval menu sets 1, 2, 3,
|
- Real-time monitoring with 2s update interval and 300s history sparklines
|
||||||
5, 10, 15, 30 or 60 seconds and clears the history
|
- Health indicator: green (HEALTHY) or red (DEGRADED) based on WAN reachability
|
||||||
- Sparklines of each target's last 100 rounds: 300 seconds at 3 seconds
|
- Summary stats: reachable count, min/max/avg latency across WAN hosts only
|
||||||
- The first round after loading or an interval change is discarded, as DNS and
|
- Fixed chart axes: Y-axis 0–1000ms, X-axis 0–300s
|
||||||
TLS setup inflate its latencies
|
|
||||||
- Health indicator from the WAN hosts' latest results: OFFLINE (red) when more
|
|
||||||
than 10 fail and at most 4 answer, otherwise DEGRADED (orange) when more than
|
|
||||||
4 fail, otherwise SLOW (yellow) when more than 3 take over 1000ms, otherwise
|
|
||||||
HEALTHY (green)
|
|
||||||
- Summary stats across WAN hosts only: how many answered, the min, median,
|
|
||||||
average and max of their latest latencies, the min and max over the whole
|
|
||||||
history, and the number of rounds run (`Checks`)
|
|
||||||
- Fixed chart axes: Y-axis 0–1000ms, higher latencies drawn at the top; X-axis
|
|
||||||
the time the history spans
|
|
||||||
- Color-coded latency figures and sparkline line segments
|
- Color-coded latency figures and sparkline line segments
|
||||||
- WAN host rows sorted by latest latency, unreachable last; pinned rows stay on
|
|
||||||
top, in name order
|
|
||||||
- Play/pause: pause stops probes but history keeps scrolling (blank gaps, no
|
- Play/pause: pause stops probes but history keeps scrolling (blank gaps, no
|
||||||
false outage)
|
false outage)
|
||||||
- Debug log panel, behind a checkbox in the footer, with five levels (error,
|
|
||||||
warning, notice, info, debug) and the last 1000 lines
|
|
||||||
- Local and UTC clocks
|
|
||||||
- Clickable service URLs
|
- Clickable service URLs
|
||||||
- A footer link to the commit the page was built from
|
|
||||||
- Canvas-based sparkline rendering with devicePixelRatio scaling
|
- Canvas-based sparkline rendering with devicePixelRatio scaling
|
||||||
- Zero runtime dependencies: all resources bundled into build artifacts
|
- Zero runtime dependencies: all resources bundled into build artifacts
|
||||||
|
|
||||||
## Deployment
|
## Deployment
|
||||||
|
|
||||||
`make build` writes the page to `dist/`, which any static file host (S3, GCS,
|
After running `yarn build`, deploy the contents of the `dist/` directory to any
|
||||||
Cloudflare Pages, Vercel, Netlify, GitHub Pages) can serve; with no backend
|
static file host (S3, GCS, Cloudflare Pages, Vercel, Netlify, GitHub Pages) or
|
||||||
there, its reports fail quietly and nothing is stored. Or run the Docker image
|
use the Docker image behind a reverse proxy.
|
||||||
behind a reverse proxy.
|
|
||||||
|
|
||||||
The Docker image, built from `Dockerfile` by `make docker`, is the whole service
|
The Docker image, built from `Dockerfile`, is the whole service in one
|
||||||
in one container: nginx serves the built frontend and passes `/api/`,
|
container: nginx serves the built frontend and passes `/api/`,
|
||||||
`/.well-known/healthcheck` and `/metrics` to the Go backend, `netwatch-server`,
|
`/.well-known/healthcheck` and `/metrics` to the Go backend, `netwatch-server`,
|
||||||
which listens only inside the container, on `127.0.0.1:8081`. The image:
|
which listens only inside the container, on `127.0.0.1:8081`. The image:
|
||||||
|
|
||||||
@@ -331,19 +283,19 @@ properties.
|
|||||||
|
|
||||||
## Limitations
|
## Limitations
|
||||||
|
|
||||||
- **CORS**: The checks are cross-origin requests in `no-cors` mode, so the page
|
- **CORS**: Some hosts may block cross-origin HEAD requests. The app uses
|
||||||
cannot read the answer, only time it: any answer counts as reachable, an error
|
`no-cors` mode which allows the request but provides opaque responses. Latency
|
||||||
page included.
|
is still measurable based on request timing.
|
||||||
- **Local targets**: The cable modem at 192.168.100.1 and the detected gateway
|
- **Local gateway**: The 192.168.100.1 endpoint requires the host to be
|
||||||
answer only on a network that has them, and only when NetWatch is served from
|
accessible from your network.
|
||||||
localhost or a private address (see Monitoring targets).
|
|
||||||
- **Network conditions**: Measurements reflect browser-to-endpoint latency,
|
- **Network conditions**: Measurements reflect browser-to-endpoint latency,
|
||||||
which includes your local network, ISP, and internet routing.
|
which includes your local network, ISP, and internet routing.
|
||||||
|
|
||||||
## TODO
|
## TODO
|
||||||
|
|
||||||
The to-do list is [TODO.md](TODO.md): where the work stands, the next step, the
|
- Add configurable host list (environment variable or config file)
|
||||||
open work, and what has been done.
|
- Add latency history export (CSV/JSON)
|
||||||
|
- Add notification/alert when status changes to DEGRADED
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|||||||
@@ -1,43 +1,28 @@
|
|||||||
# Workflow
|
# Workflow
|
||||||
|
|
||||||
- branch from `next`
|
- branch (from `main`)
|
||||||
- do the work in Next Step
|
- do the work in Next Step
|
||||||
- move Next Step to the top of Completed Steps
|
- move Next Step to the top of Completed Steps
|
||||||
- move the top item of Future Steps into Next Step
|
- move the top item of Future Steps into Next Step
|
||||||
- commit (`TODO.md` changes in the same commit as the work)
|
- commit (`TODO.md` changes in the same commit as the work)
|
||||||
- push the branch and open a PR against `next`
|
- merge to `main` if the branch is not protected, otherwise open a PR
|
||||||
|
- push
|
||||||
|
|
||||||
# Status
|
# Status
|
||||||
|
|
||||||
pre-1.0. No git tags. `main` is the stable branch and `next` the development
|
pre-1.0. No git tags. `feat/reportbuf-storage` is merged; the backend, the CI
|
||||||
branch, which every PR targets. The frontend and the Go backend ship as one
|
workflow, and the backend repo standard files are all on `main`. Frontend and
|
||||||
Docker image, and the Gitea workflow `.gitea/workflows/check.yml` runs
|
backend are both functional. Working toward the 1.0.0 milestone by closing the
|
||||||
`script/cibuild` on every push. Working toward 1.0.0.
|
remaining repo-compliance issues on the tracker.
|
||||||
|
|
||||||
# Next Step
|
# Next Step
|
||||||
|
|
||||||
Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
|
Confirm the `.gitea/workflows/check.yml` run is green (main always green
|
||||||
`backend/` no longer repeating files from the root
|
policy). The workflow file is already on `main`; what is unverified is that its
|
||||||
([#30](https://git.eeqj.de/sneak/netwatch/issues/30)).
|
latest run passes.
|
||||||
|
|
||||||
# Completed Steps
|
# Completed Steps
|
||||||
|
|
||||||
- 2026-10-04: the page's footer no longer says "IPv4 only"
|
|
||||||
([#111](https://git.eeqj.de/sneak/netwatch/issues/111)): each check is a
|
|
||||||
`fetch`, the browser picks IPv4 or IPv6 for each WAN host, and the local
|
|
||||||
targets are IPv4 addresses. The rest of the footer is unchanged
|
|
||||||
- 2026-10-04: `README.md`, `TODO.md` and `test/viewport/README.md` say what the
|
|
||||||
tree does (issue #24). The README's Getting Started leads with `make` targets;
|
|
||||||
a new Backend section says what `netwatch-server` stores, its routes and how
|
|
||||||
the image builds and runs it, and points to `backend/README.md` for its
|
|
||||||
settings; the checks are GET requests; the 26 WAN hosts, the four health
|
|
||||||
states, the summary's figures and the features the list lacked are described
|
|
||||||
as the page has them; and its TODO section points here, as does the one in
|
|
||||||
`backend/README.md`, whose open items moved to Future Steps. This file's
|
|
||||||
Workflow branches from `next` and opens the PR against `next`, Status says
|
|
||||||
where the repo stands, and Next Step and Future Steps hold only open work,
|
|
||||||
linked to its issue where one exists. The viewport harness README names Node's
|
|
||||||
test runner, not `vitest`
|
|
||||||
- 2026-10-04: in `src/main.js` (issue #102), a target's min, max, median and
|
- 2026-10-04: in `src/main.js` (issue #102), a target's min, max, median and
|
||||||
average latency come from one list of its answers, through the same function
|
average latency come from one list of its answers, through the same function
|
||||||
the summary's figures use, so the median is written once. The latency color
|
the summary's figures use, so the median is written once. The latency color
|
||||||
@@ -48,13 +33,6 @@ Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
|
|||||||
`CONFIG` entries too. A unit test checks the summary's figures. Nothing the
|
`CONFIG` entries too. A unit test checks the summary's figures. Nothing the
|
||||||
page does or shows changed; the footer's color legend still writes the limits
|
page does or shows changed; the footer's color legend still writes the limits
|
||||||
out as text
|
out as text
|
||||||
- 2026-10-04: password guesses at `/metrics` are rate limited (issue #104): each
|
|
||||||
client address, resolved through `TRUSTED_PROXIES` as for reports, may make 60
|
|
||||||
requests to `/metrics` a minute, counted by `go-chi/httprate` apart from its
|
|
||||||
reports; past that it gets 429 and its basic auth credentials are not checked.
|
|
||||||
The limit is a constant in `backend/internal/server/routes.go`. A test uses up
|
|
||||||
one client's allowance on wrong passwords, gets 429 with the right one, and
|
|
||||||
checks that another client behind the same nginx still gets in
|
|
||||||
- 2026-10-04: the backend reports errors to Sentry (issue #95). With
|
- 2026-10-04: the backend reports errors to Sentry (issue #95). With
|
||||||
`SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-`
|
`SENTRY_DSN` set, it sets up `sentry-go` with the release `netwatch-server-`
|
||||||
and its version, reports each panic in a handler through `sentryhttp`, the
|
and its version, reports each panic in a handler through `sentryhttp`, the
|
||||||
@@ -389,14 +367,12 @@ Decide whether the repo moves to the layout `REPO_POLICIES.md` gives, with
|
|||||||
|
|
||||||
# Future Steps
|
# Future Steps
|
||||||
|
|
||||||
- Run `make frontend-viewport-test` in CI as its own step; it is not part of
|
- Wire `script/frontend-viewport-test` into CI as its own step (deliberately not
|
||||||
`make check`, as it needs Docker and takes minutes
|
part of `make check` today; the decision has real CI-runtime cost and is
|
||||||
- A backend test that posts a report to `POST /api/v1/reports` and checks the
|
tracked separately)
|
||||||
compressed file it is written to
|
- Compliance top-up as one small commit: add .editorconfig and add the hooks
|
||||||
- A backend route that decompresses the stored reports and answers queries on
|
target to the Makefile
|
||||||
them
|
- After merge, confirm .gitea/workflows/check.yml is on main and CI is green
|
||||||
- Prometheus metrics for the backend's in-memory buffer: its size, the number of
|
(main always green policy)
|
||||||
flushes and the number of reports
|
- Decide what to do with untracked resume.sh: commit it, gitignore it, or delete
|
||||||
- A configurable host list (an environment variable or a config file)
|
it
|
||||||
- Export of the latency history (CSV or JSON)
|
|
||||||
- A notification when the health status changes to DEGRADED
|
|
||||||
|
|||||||
+3
-10
@@ -199,14 +199,6 @@ is recorded and `/metrics` answers 404. One without the other stops the server
|
|||||||
from starting, with an error naming both; so does a `METRICS_USERNAME`
|
from starting, with an error naming both; so does a `METRICS_USERNAME`
|
||||||
containing `:`, which basic auth cannot carry, with an error naming it.
|
containing `:`, which basic auth cannot carry, with an error naming it.
|
||||||
|
|
||||||
`/metrics` is rate limited, so that its password cannot be guessed quickly: each
|
|
||||||
client address, resolved through `TRUSTED_PROXIES`, may make 60 requests to it a
|
|
||||||
minute, whatever their credentials. Past that it gets 429 with
|
|
||||||
`Retry-After: 60`, and its credentials are not checked. The minute slides as it
|
|
||||||
does for reports (see [Report limits](#report-limits)), so a scraper polling
|
|
||||||
every 2 seconds or less often is never refused. This allowance is apart from the
|
|
||||||
one for reports.
|
|
||||||
|
|
||||||
### Sentry
|
### Sentry
|
||||||
|
|
||||||
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
|
With `SENTRY_DSN` set, the server sends its errors to that Sentry project: each
|
||||||
@@ -219,8 +211,9 @@ sent to it.
|
|||||||
|
|
||||||
## TODO
|
## TODO
|
||||||
|
|
||||||
The to-do list, this backend's open work included, is [TODO.md](../TODO.md) at
|
- Add integration test that POSTs a report and verifies the compressed output
|
||||||
the repo root.
|
- Add report decompression/query endpoint
|
||||||
|
- Add metrics (Prometheus) for buffer size, flush count, report count
|
||||||
|
|
||||||
## License
|
## License
|
||||||
|
|
||||||
|
|||||||
@@ -12,10 +12,6 @@ func (s *Server) Router() *chi.Mux {
|
|||||||
// external tests.
|
// external tests.
|
||||||
const MaxRequestBodyBytes = maxRequestBodyBytes
|
const MaxRequestBodyBytes = maxRequestBodyBytes
|
||||||
|
|
||||||
// MetricsRequestsPerMinute exposes the /metrics rate limit to the
|
|
||||||
// external tests.
|
|
||||||
const MetricsRequestsPerMinute = metricsRequestsPerMinute
|
|
||||||
|
|
||||||
// ListenAddr exposes the address the server listens on to the
|
// ListenAddr exposes the address the server listens on to the
|
||||||
// external tests.
|
// external tests.
|
||||||
func (s *Server) ListenAddr() string {
|
func (s *Server) ListenAddr() string {
|
||||||
|
|||||||
@@ -18,12 +18,6 @@ const (
|
|||||||
// can mount s.mw.MaxBodyBytes with a smaller value to lower
|
// can mount s.mw.MaxBodyBytes with a smaller value to lower
|
||||||
// its bound, but cannot raise it: this cap runs first.
|
// its bound, but cannot raise it: this cap runs first.
|
||||||
maxRequestBodyBytes int64 = 1 << 20 // 1 MiB
|
maxRequestBodyBytes int64 = 1 << 20 // 1 MiB
|
||||||
|
|
||||||
// metricsRequestsPerMinute is how many requests to /metrics each
|
|
||||||
// client address may make a minute, whatever their credentials. A
|
|
||||||
// scraper polling every 2 seconds sends half of it, which httprate
|
|
||||||
// never refuses.
|
|
||||||
metricsRequestsPerMinute = 60
|
|
||||||
)
|
)
|
||||||
|
|
||||||
// SetupRoutes configures the chi router with middleware and
|
// SetupRoutes configures the chi router with middleware and
|
||||||
@@ -72,13 +66,9 @@ func (s *Server) SetupRoutes() {
|
|||||||
Post("/api/v1/reports", s.h.HandleReport())
|
Post("/api/v1/reports", s.h.HandleReport())
|
||||||
})
|
})
|
||||||
|
|
||||||
// The rate limit comes before the basic auth, so a client past it
|
|
||||||
// gets 429 and its password is not checked.
|
|
||||||
if s.params.Config.MetricsUsername != "" {
|
if s.params.Config.MetricsUsername != "" {
|
||||||
s.router.With(
|
s.router.With(s.mw.MetricsAuth()).
|
||||||
s.mw.RateLimit(metricsRequestsPerMinute),
|
Get("/metrics", promhttp.HandlerFor(
|
||||||
s.mw.MetricsAuth(),
|
|
||||||
).Get("/metrics", promhttp.HandlerFor(
|
|
||||||
registry, promhttp.HandlerOpts{},
|
registry, promhttp.HandlerOpts{},
|
||||||
).ServeHTTP)
|
).ServeHTTP)
|
||||||
}
|
}
|
||||||
|
|||||||
@@ -188,50 +188,6 @@ func TestMetricsBehindBasicAuth(t *testing.T) {
|
|||||||
}
|
}
|
||||||
}
|
}
|
||||||
|
|
||||||
// TestMetricsAreRateLimited: a client that has used up its /metrics
|
|
||||||
// allowance on wrong passwords gets 429 even with the right one, which
|
|
||||||
// is then not checked, while another client behind the same nginx
|
|
||||||
// still gets in.
|
|
||||||
func TestMetricsAreRateLimited(t *testing.T) {
|
|
||||||
t.Setenv("METRICS_USERNAME", "prometheus")
|
|
||||||
t.Setenv("METRICS_PASSWORD", "right")
|
|
||||||
// As in the container: nginx connects from loopback and names the
|
|
||||||
// client in X-Forwarded-For.
|
|
||||||
t.Setenv("TRUSTED_PROXIES", "127.0.0.1/32")
|
|
||||||
|
|
||||||
srv := newServer(t)
|
|
||||||
srv.SetupRoutes()
|
|
||||||
|
|
||||||
get := func(client, password string) int {
|
|
||||||
rec := httptest.NewRecorder()
|
|
||||||
req := httptest.NewRequestWithContext(t.Context(),
|
|
||||||
http.MethodGet, "/metrics", http.NoBody)
|
|
||||||
req.RemoteAddr = "127.0.0.1:40000"
|
|
||||||
req.Header.Set("X-Forwarded-For", client)
|
|
||||||
req.SetBasicAuth("prometheus", password)
|
|
||||||
srv.ServeHTTP(rec, req)
|
|
||||||
|
|
||||||
return rec.Code
|
|
||||||
}
|
|
||||||
|
|
||||||
for i := range server.MetricsRequestsPerMinute {
|
|
||||||
if code := get("203.0.113.7", "wrong"); code != http.StatusUnauthorized {
|
|
||||||
t.Fatalf("guess %d: status = %d, want %d",
|
|
||||||
i+1, code, http.StatusUnauthorized)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
if code := get("203.0.113.7", "right"); code != http.StatusTooManyRequests {
|
|
||||||
t.Fatalf("right password past the limit: status = %d, want %d",
|
|
||||||
code, http.StatusTooManyRequests)
|
|
||||||
}
|
|
||||||
|
|
||||||
if code := get("203.0.113.8", "right"); code != http.StatusOK {
|
|
||||||
t.Fatalf("another client: status = %d, want %d",
|
|
||||||
code, http.StatusOK)
|
|
||||||
}
|
|
||||||
}
|
|
||||||
|
|
||||||
// TestMetricsInTwoServers: two servers in one process can both have
|
// TestMetricsInTwoServers: two servers in one process can both have
|
||||||
// metrics on.
|
// metrics on.
|
||||||
func TestMetricsInTwoServers(t *testing.T) {
|
func TestMetricsInTwoServers(t *testing.T) {
|
||||||
|
|||||||
+1
-1
@@ -857,7 +857,7 @@ function buildUI(state) {
|
|||||||
</div>
|
</div>
|
||||||
|
|
||||||
<footer class="mt-8 text-center text-gray-600 text-xs">
|
<footer class="mt-8 text-center text-gray-600 text-xs">
|
||||||
<p>Latency measured via GET requests | CORS restrictions may affect some measurements</p>
|
<p>Latency measured via GET requests | IPv4 only | CORS restrictions may affect some measurements</p>
|
||||||
<p class="mt-2">
|
<p class="mt-2">
|
||||||
<span class="inline-block w-3 h-3 rounded-full bg-green-500 mr-1 align-middle"></span><50ms
|
<span class="inline-block w-3 h-3 rounded-full bg-green-500 mr-1 align-middle"></span><50ms
|
||||||
<span class="inline-block w-3 h-3 rounded-full bg-lime-500 mr-1 ml-3 align-middle"></span><100ms
|
<span class="inline-block w-3 h-3 rounded-full bg-lime-500 mr-1 ml-3 align-middle"></span><100ms
|
||||||
|
|||||||
@@ -77,7 +77,7 @@ the internet, so the app's latency probes cannot reach anything real. The
|
|||||||
harness answers them itself from a fixed delay table, with a deterministic
|
harness answers them itself from a fixed delay table, with a deterministic
|
||||||
fraction failed outright, so the rows render a realistic spread of one-, two-
|
fraction failed outright, so the rows render a realistic spread of one-, two-
|
||||||
and three-digit latencies plus some unreachable rows. That spread is what the
|
and three-digit latencies plus some unreachable rows. That spread is what the
|
||||||
layout has to survive; a `---` placeholder in every row would not exercise it.
|
layout has to survive; 24 identical `---` placeholders would not exercise it.
|
||||||
|
|
||||||
## What this cannot verify
|
## What this cannot verify
|
||||||
|
|
||||||
@@ -104,11 +104,10 @@ Everything else this issue was actually about — does the layout reflow, does
|
|||||||
anything overflow, is content clipped, are the controls big enough — is a
|
anything overflow, is content clipped, are the controls big enough — is a
|
||||||
function of viewport width and CSS, and is covered above.
|
function of viewport width and CSS, and is covered above.
|
||||||
|
|
||||||
## Relation to the unit tests
|
## Relation to the unit test framework (#21)
|
||||||
|
|
||||||
Complementary layers, not two stacks. The unit tests in `test/unit/`, which
|
Complementary layers, not two stacks. `vitest` (#21) will exercise module-level
|
||||||
`make test` runs with Node's built-in test runner, exercise the functions
|
logic in-process with no browser. This harness exercises rendered layout in a
|
||||||
`src/main.js` exports in-process with no browser. This harness exercises
|
real engine and is the only thing here that can see a media query. Neither
|
||||||
rendered layout in a real engine and is the only thing here that can see a media
|
replaces the other; assertions about computed styles and element geometry belong
|
||||||
query. Neither replaces the other; assertions about computed styles and element
|
here, assertions about functions belong in `vitest`.
|
||||||
geometry belong here, assertions about functions belong in `test/unit/`.
|
|
||||||
|
|||||||
Reference in New Issue
Block a user