please have the manager run netwatch in a browser and figure out why the hetzner hosts are failing. i'd like detailed logging to console added for these types of failures to make clear in the console exactly what is failing and when and why.
The six Hetzner speed-test targets (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore) show as failing.
Definition of done:
netwatch is run in a real headless browser (Chrome in a container, the way the page is served for sneak) and the Hetzner targets are watched failing. The cause is found and stated on this issue in plain language: which URL is requested, what the browser does with it (timeout, network error, blocked request, DNS, TLS, IPv6, or another cause), and why. The evidence is a short excerpt of console output, not a log dump.
If the cause is in netwatch (wrong URL, wrong request mode, a timeout too short), it is fixed. If it is outside netwatch (the Hetzner servers themselves, or the browser's rules), this issue says so, and what netwatch should do about it is put to sneak as one question with a recommendation.
Every failed check logs one line to the browser console that says exactly what failed and when and why: the target's name and URL, the time, the kind of failure (timeout after how long, network error, or the error's own message), and how long the request took. A target that recovers logs that too. Nothing is logged for checks that succeed.
Tests cover the logging. Lands on next after an independent review.
Model: opus-5-5
sneak (chat, 2026-10-07 ~08:33 UTC):
> please have the manager run netwatch in a browser and figure out why the hetzner hosts are failing. i'd like detailed logging to console added for these types of failures to make clear in the console exactly what is failing and when and why.
The six Hetzner speed-test targets (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore) show as failing.
Definition of done:
- netwatch is run in a real headless browser (Chrome in a container, the way the page is served for sneak) and the Hetzner targets are watched failing. The cause is found and stated on this issue in plain language: which URL is requested, what the browser does with it (timeout, network error, blocked request, DNS, TLS, IPv6, or another cause), and why. The evidence is a short excerpt of console output, not a log dump.
- If the cause is in netwatch (wrong URL, wrong request mode, a timeout too short), it is fixed. If it is outside netwatch (the Hetzner servers themselves, or the browser's rules), this issue says so, and what netwatch should do about it is put to sneak as one question with a recommendation.
- Every failed check logs one line to the browser console that says exactly what failed and when and why: the target's name and URL, the time, the kind of failure (timeout after how long, network error, or the error's own message), and how long the request took. A target that recovers logs that too. Nothing is logged for checks that succeed.
- Tests cover the logging. Lands on `next` after an independent review.
Model: opus-5-5
A worker builds the image from next, serves it as sneak gets it, opens it in headless Chrome in a container and watches several rounds of the six Hetzner checks: which URL is requested, what the browser does with it, and why. It posts the cause here with a short console excerpt. No code changes in this step.
Lead from a shell on the build host, to be confirmed or refuted in the browser: each Hetzner speed-test server answers https://fsn1-speed.hetzner.com/ with 200 but drops the connection when the URL carries any query string (/?a=1 is reset, over http as well as https). Every check netwatch makes adds ?_cb=<time> to the target's URL (measureLatency in src/main.js).
A second worker, one PR to next: the fix, if the cause is in netwatch, and one console line for each failed check and each recovery as the definition of done above says, with tests. Today the failure messages go only to the page's debug panel. If the query string is confirmed as the cause, the reading taken is to stop adding it for every target: the checks already use cache: "no-store", which keeps the browser's cache out of the measurement and sends Cache-Control: no-cache to caches on the way, and an exception for one group of hosts would be a second code path. The README sentence that describes the query parameter changes with it. An independent reviewer gates the PR before it is merged.
Model: opus-5-5
Plan, two steps.
1. A worker builds the image from `next`, serves it as sneak gets it, opens it in headless Chrome in a container and watches several rounds of the six Hetzner checks: which URL is requested, what the browser does with it, and why. It posts the cause here with a short console excerpt. No code changes in this step.
Lead from a shell on the build host, to be confirmed or refuted in the browser: each Hetzner speed-test server answers `https://fsn1-speed.hetzner.com/` with 200 but drops the connection when the URL carries any query string (`/?a=1` is reset, over http as well as https). Every check netwatch makes adds `?_cb=<time>` to the target's URL (`measureLatency` in `src/main.js`).
2. A second worker, one PR to `next`: the fix, if the cause is in netwatch, and one console line for each failed check and each recovery as the definition of done above says, with tests. Today the failure messages go only to the page's debug panel. If the query string is confirmed as the cause, the reading taken is to stop adding it for every target: the checks already use `cache: "no-store"`, which keeps the browser's cache out of the measurement and sends `Cache-Control: no-cache` to caches on the way, and an exception for one group of hosts would be a second code path. The README sentence that describes the query parameter changes with it. An independent reviewer gates the PR before it is merged.
Model: opus-5-5
clawbot
self-assigned this 2026-10-07 10:36:17 +02:00
Cause found: the Hetzner speed-test servers drop every request whose URL has a query string, and netwatch adds one to every check.
What netwatch requests. Each check fetches the target's URL with ?_cb= and the current time added (measureLatency in src/main.js), for example https://fsn1-speed.hetzner.com/?_cb=1791362479236.
What the browser does with it. Headless Chrome 154, loading the image built from next with its security headers, sent that request every round to all six Hetzner hosts. Each server accepted the connection, completed the encryption handshake, read the request, then closed the connection without sending any answer. Chrome reports this as net::ERR_EMPTY_RESPONSE. The fetch call fails with TypeError: Failed to fetch, and netwatch marks the row "unreachable". Nothing timed out, and the page's security rules blocked nothing: each failure arrives about one network round trip after the request is sent.
Without the query string. Inside the same page, the same request to each of the six hosts without ?_cb=… got status 200 every time, and with ?_cb=1 it failed every time. From a shell on the build host the servers do the same for any non-empty query string (?a=1, ?x), on any path, over http as well as https, so the browser plays no part in it.
Other targets. Every other target also answered without the query string in the same browser. S3 me-south-1 (Bahrain) timed out and Local CPE (http://192.168.100.1, which a container cannot reach) failed both with and without it, so neither has to do with this cause.
Where the cause lies. Dropping these requests is the servers' own behaviour, which netwatch cannot change. But netwatch is what adds the query string, so it can be fixed in netwatch by not adding it, as step 2 of the plan says. This needs no decision from sneak.
REQFAILED net::ERR_EMPTY_RESPONSE https://fsn1-speed.hetzner.com/?_cb=1791362479236
REQFAILED net::ERR_EMPTY_RESPONSE https://sin-speed.hetzner.com/?_cb=1791362479236
CONSOLE error Failed to load resource: net::ERR_EMPTY_RESPONSE
debug panel: ERROR https://fsn1-speed.hetzner.com unreachable
fetch https://fsn1-speed.hetzner.com/ -> resolved, status 200 from 78.46.170.2:443
fetch https://fsn1-speed.hetzner.com/?_cb=1 -> rejected "TypeError: Failed to fetch"
REQFAILED net::ERR_EMPTY_RESPONSE https://fsn1-speed.hetzner.com/?_cb=1
IPv6 not tested: this build host has no IPv6 route, so the browser used IPv4. The servers drop the request only after reading it, so IPv6 should behave the same, but this was not checked.
Deviation: the browser container ran without the --memory 2g and --cpus 2 limits, because this host's Docker refuses both flags. One browser with one page ran.
Model: opus-5-5
Cause found: the Hetzner speed-test servers drop every request whose URL has a query string, and netwatch adds one to every check.
**What netwatch requests.** Each check fetches the target's URL with `?_cb=` and the current time added (`measureLatency` in `src/main.js`), for example `https://fsn1-speed.hetzner.com/?_cb=1791362479236`.
**What the browser does with it.** Headless Chrome 154, loading the image built from `next` with its security headers, sent that request every round to all six Hetzner hosts. Each server accepted the connection, completed the encryption handshake, read the request, then closed the connection without sending any answer. Chrome reports this as `net::ERR_EMPTY_RESPONSE`. The `fetch` call fails with `TypeError: Failed to fetch`, and netwatch marks the row "unreachable". Nothing timed out, and the page's security rules blocked nothing: each failure arrives about one network round trip after the request is sent.
**Without the query string.** Inside the same page, the same request to each of the six hosts without `?_cb=…` got status 200 every time, and with `?_cb=1` it failed every time. From a shell on the build host the servers do the same for any non-empty query string (`?a=1`, `?x`), on any path, over http as well as https, so the browser plays no part in it.
**Other targets.** Every other target also answered without the query string in the same browser. S3 me-south-1 (Bahrain) timed out and Local CPE (`http://192.168.100.1`, which a container cannot reach) failed both with and without it, so neither has to do with this cause.
**Where the cause lies.** Dropping these requests is the servers' own behaviour, which netwatch cannot change. But netwatch is what adds the query string, so it can be fixed in netwatch by not adding it, as step 2 of the plan says. This needs no decision from sneak.
```
REQFAILED net::ERR_EMPTY_RESPONSE https://fsn1-speed.hetzner.com/?_cb=1791362479236
REQFAILED net::ERR_EMPTY_RESPONSE https://sin-speed.hetzner.com/?_cb=1791362479236
CONSOLE error Failed to load resource: net::ERR_EMPTY_RESPONSE
debug panel: ERROR https://fsn1-speed.hetzner.com unreachable
fetch https://fsn1-speed.hetzner.com/ -> resolved, status 200 from 78.46.170.2:443
fetch https://fsn1-speed.hetzner.com/?_cb=1 -> rejected "TypeError: Failed to fetch"
REQFAILED net::ERR_EMPTY_RESPONSE https://fsn1-speed.hetzner.com/?_cb=1
```
- IPv6 not tested: this build host has no IPv6 route, so the browser used IPv4. The servers drop the request only after reading it, so IPv6 should behave the same, but this was not checked.
- Deviation: the browser container ran without the `--memory 2g` and `--cpus 2` limits, because this host's Docker refuses both flags. One browser with one page ran.
Model: opus-5-5
Built in #116, waiting for review. Checks now fetch each target's URL without adding a query string, so the six Hetzner targets answer again. Each failed check the page records writes one line to the browser console, and the same line to the debug log: the target's name and URL, the time, what failed and how long the request took. A target that answers after failed checks writes one line saying it recovered and how many checks in a row had failed.
Model: opus-5-5
Built in https://git.eeqj.de/sneak/netwatch/pulls/116, waiting for review. Checks now fetch each target's URL without adding a query string, so the six Hetzner targets answer again. Each failed check the page records writes one line to the browser console, and the same line to the debug log: the target's name and URL, the time, what failed and how long the request took. A target that answers after failed checks writes one line saying it recovered and how many checks in a row had failed.
Model: opus-5-5
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.
sneak (chat, 2026-10-07 ~08:33 UTC):
The six Hetzner speed-test targets (Nuremberg, Falkenstein, Helsinki, Ashburn, Hillsboro, Singapore) show as failing.
Definition of done:
nextafter an independent review.Model: opus-5-5
Plan, two steps.
A worker builds the image from
next, serves it as sneak gets it, opens it in headless Chrome in a container and watches several rounds of the six Hetzner checks: which URL is requested, what the browser does with it, and why. It posts the cause here with a short console excerpt. No code changes in this step.Lead from a shell on the build host, to be confirmed or refuted in the browser: each Hetzner speed-test server answers
https://fsn1-speed.hetzner.com/with 200 but drops the connection when the URL carries any query string (/?a=1is reset, over http as well as https). Every check netwatch makes adds?_cb=<time>to the target's URL (measureLatencyinsrc/main.js).A second worker, one PR to
next: the fix, if the cause is in netwatch, and one console line for each failed check and each recovery as the definition of done above says, with tests. Today the failure messages go only to the page's debug panel. If the query string is confirmed as the cause, the reading taken is to stop adding it for every target: the checks already usecache: "no-store", which keeps the browser's cache out of the measurement and sendsCache-Control: no-cacheto caches on the way, and an exception for one group of hosts would be a second code path. The README sentence that describes the query parameter changes with it. An independent reviewer gates the PR before it is merged.Model: opus-5-5
Cause found: the Hetzner speed-test servers drop every request whose URL has a query string, and netwatch adds one to every check.
What netwatch requests. Each check fetches the target's URL with
?_cb=and the current time added (measureLatencyinsrc/main.js), for examplehttps://fsn1-speed.hetzner.com/?_cb=1791362479236.What the browser does with it. Headless Chrome 154, loading the image built from
nextwith its security headers, sent that request every round to all six Hetzner hosts. Each server accepted the connection, completed the encryption handshake, read the request, then closed the connection without sending any answer. Chrome reports this asnet::ERR_EMPTY_RESPONSE. Thefetchcall fails withTypeError: Failed to fetch, and netwatch marks the row "unreachable". Nothing timed out, and the page's security rules blocked nothing: each failure arrives about one network round trip after the request is sent.Without the query string. Inside the same page, the same request to each of the six hosts without
?_cb=…got status 200 every time, and with?_cb=1it failed every time. From a shell on the build host the servers do the same for any non-empty query string (?a=1,?x), on any path, over http as well as https, so the browser plays no part in it.Other targets. Every other target also answered without the query string in the same browser. S3 me-south-1 (Bahrain) timed out and Local CPE (
http://192.168.100.1, which a container cannot reach) failed both with and without it, so neither has to do with this cause.Where the cause lies. Dropping these requests is the servers' own behaviour, which netwatch cannot change. But netwatch is what adds the query string, so it can be fixed in netwatch by not adding it, as step 2 of the plan says. This needs no decision from sneak.
--memory 2gand--cpus 2limits, because this host's Docker refuses both flags. One browser with one page ran.Model: opus-5-5
Built in #116, waiting for review. Checks now fetch each target's URL without adding a query string, so the six Hetzner targets answer again. Each failed check the page records writes one line to the browser console, and the same line to the debug log: the target's name and URL, the time, what failed and how long the request took. A target that answers after failed checks writes one line saying it recovered and how many checks in a row had failed.
Model: opus-5-5