harness-probe-false-negative-wrong-json-field-path
Done. I diagnosed the four probe defects, reproduced each against a live build, and wrote the solution.
Deliverable: ~/<project>-probe-false-negative-solution.md
Cloned and built the Go server (go build ./cmd/server), started it on scratch ports 18801/18802 with CR_REQUIRE_AGENT_SIG=false / =true, and ran the probe scenarios with uv --with websockets and curl:
root_code=None, error.code=CONTROLLER_OFFLINE, request_id echoes the REQUEST message_id, in 0.3 ms. The probe's body['code'] read caused the phantom HUNG >5s."target":"ghost-1" → confirmed silent drop (no ERROR, no log), matching the documented behavior; the fix is "target":{"agent_id":"ghost-1"}.register without key → 201 on the sig-off server, 400 "public_key is required" on the default sig-on server. Also verified GET /version does not expose the flag, so the harness must own the configuration.alerts.> receives both events, and the relay stream carries only the event body (no topic), so a single recv() and topic-matching are both wrong — drain until a payload marker or deadline.The doc contains root-cause analysis with source citations (internal/mesh/message.go, internal/mesh/peer.go, internal/registry/handler.go, config/config.go), copy-pasteable corrected probe code, a port-ownership guard, and a verification section with the observed raw output. No product change — product was correct throughout. Scratch ports and processes cleaned up.
# Evidence - Problem class: harness-probe-false-negative-wrong-json-field-path - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-17T16:52:08.338Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a live-probe harness reported a product defect that did not exist. Three probes against a running <project> server failed in one run: (1) 'mesh REQUEST to an unconnected peer hangs >5s', (2) 'alerts.> did not receive the deep topic', (3) 'register returned 400'. Running the same server by hand showed all three were fine.\n\nROOT CAUSE (all three were PROBE bugs, each a different class):\n1. Field-path error. The mesh ERROR frame nests its code: {\"type\":\"ERROR\",\"request_id\":...,\"error\":{\"code\":\"CONTROLLER_OFFLINE\",\"message\":...}}. The probe asserted body['code']=='CONTROLLER_OFFLINE' at the ROOT, so a correct frame arriving in 0ms was scored as 'HUNG >5s'. Rule: when a probe reports a timeout/hang but the server log shows it answered, check the frame's nested shape (error.code, result.*) before believing the hang. The protocol reference showed the nesting; the probe was written from intuition.\n2. Malformed request silently dropped. The same probe sent target as a bare string (\"target\":\"ghost-1\") instead of the PeerRef object {\"agent_id\":\"ghost-1\"}. The server drops malformed mesh frames with no ERROR and no log (documented behaviour), so an invalid probe is indistinguishable from an unresponsive peer. Envelope-shaped protocols REQUIRE object refs; validate your own frame against the spec example before concluding the peer is down.\n3. Wrong server configuration for the probe's premise. 'register without public_key -> 201' only holds on a server started CR_REQUIRE_AGENT_SIG=false. The probe ran against a signature-enforcing server (the default) and got 400 - a correct rejection scored as a failure. Assert the server's configuration (probe /version or an unauthenticated endpoint, or the documented effect of the flag) inside the harness before asserting flag-dependent behaviour.\nAlso: a wildcard drain bug - the .> subscriber legitimately received the earlier one-segment topic first, and the probe read exactly one frame, so it scored the deep-topic assertion against the previous frame. Drain until the expected frame or a deadline.\n\nFIX: corrected the probes (nested error.code; PeerRef objects; sig-off server for that cell; drain-until-match). Rerun: 8/9 then 9/9 - the only remaining 'failure' was the same field-path bug in probe A, fixed by reading error.code. No product change was made: the product had been correct the whole time.\n\nVERIFICATION: fresh build of the Go server at HEAD 60cf425 on scratch ports 18779/18791/18792/18793, each start preceded by a port pre-check and followed by a holder-pid assertion (ss -tlnp pid == started pid). Corrected probe: ERROR error.code=CONTROLLER_OFFLINE in 0ms with request_id echoing the REQUEST message_id; .* got the one-segment frame and .> got both; register 201 on the sig-off server.\n\nGENERAL RULE: a probe failure is a claim about the product; before filing it as a defect, re-run the same assertion by hand (curl/python) and read the RAW response. A probe that reports a hang/false negative costs a wasted defect row and, worse, hides the behaviour it was meant to prove.\n", "environment": "Go 1.x <project> server (github.com/<project>-dev/<project>) on linux, websockets client (python uv --with websockets), scratch ports with an ownership guard", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "harness-probe-false-negative-wrong-json-field-path", "provider": "openrouter", "solved_at": "2026-09-17T16:52:08.338Z", "version": ""}Done. I diagnosed the four probe defects, reproduced each against a live build, and wrote the solution.
Deliverable: ~/<project>-probe-false-negative-solution.md
Cloned and built the Go server (go build ./cmd/server), started it on scratch ports 18801/18802 with CR_REQUIRE_AGENT_SIG=false / =true, and ran the probe scenarios with uv --with websockets and curl:
root_code=None, error.code=CONTROLLER_OFFLINE, request_id echoes the REQUEST message_id, in 0.3 ms. The probe's body['code'] read caused the phantom HUNG >5s."target":"ghost-1" → confirmed silent drop (no ERROR, no log), matching the documented behavior; the fix is "target":{"agent_id":"ghost-1"}.register without key → 201 on the sig-off server, 400 "public_key is required" on the default sig-on server. Also verified GET /version does not expose the flag, so the harness must own the configuration.alerts.> receives both events, and the relay stream carries only the event body (no topic), so a single recv() and topic-matching are both wrong — drain until a payload marker or deadline.The doc contains root-cause analysis with source citations (internal/mesh/message.go, internal/mesh/peer.go, internal/registry/handler.go, config/config.go), copy-pasteable corrected probe code, a port-ownership guard, and a verification section with the observed raw output. No product change — product was correct throughout. Scratch ports and processes cleaned up.
# Evidence - Problem class: harness-probe-false-negative-wrong-json-field-path - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-17T16:52:08.338Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a live-probe harness reported a product defect that did not exist. Three probes against a running <project> server failed in one run: (1) 'mesh REQUEST to an unconnected peer hangs >5s', (2) 'alerts.> did not receive the deep topic', (3) 'register returned 400'. Running the same server by hand showed all three were fine.\n\nROOT CAUSE (all three were PROBE bugs, each a different class):\n1. Field-path error. The mesh ERROR frame nests its code: {\"type\":\"ERROR\",\"request_id\":...,\"error\":{\"code\":\"CONTROLLER_OFFLINE\",\"message\":...}}. The probe asserted body['code']=='CONTROLLER_OFFLINE' at the ROOT, so a correct frame arriving in 0ms was scored as 'HUNG >5s'. Rule: when a probe reports a timeout/hang but the server log shows it answered, check the frame's nested shape (error.code, result.*) before believing the hang. The protocol reference showed the nesting; the probe was written from intuition.\n2. Malformed request silently dropped. The same probe sent target as a bare string (\"target\":\"ghost-1\") instead of the PeerRef object {\"agent_id\":\"ghost-1\"}. The server drops malformed mesh frames with no ERROR and no log (documented behaviour), so an invalid probe is indistinguishable from an unresponsive peer. Envelope-shaped protocols REQUIRE object refs; validate your own frame against the spec example before concluding the peer is down.\n3. Wrong server configuration for the probe's premise. 'register without public_key -> 201' only holds on a server started CR_REQUIRE_AGENT_SIG=false. The probe ran against a signature-enforcing server (the default) and got 400 - a correct rejection scored as a failure. Assert the server's configuration (probe /version or an unauthenticated endpoint, or the documented effect of the flag) inside the harness before asserting flag-dependent behaviour.\nAlso: a wildcard drain bug - the .> subscriber legitimately received the earlier one-segment topic first, and the probe read exactly one frame, so it scored the deep-topic assertion against the previous frame. Drain until the expected frame or a deadline.\n\nFIX: corrected the probes (nested error.code; PeerRef objects; sig-off server for that cell; drain-until-match). Rerun: 8/9 then 9/9 - the only remaining 'failure' was the same field-path bug in probe A, fixed by reading error.code. No product change was made: the product had been correct the whole time.\n\nVERIFICATION: fresh build of the Go server at HEAD 60cf425 on scratch ports 18779/18791/18792/18793, each start preceded by a port pre-check and followed by a holder-pid assertion (ss -tlnp pid == started pid). Corrected probe: ERROR error.code=CONTROLLER_OFFLINE in 0ms with request_id echoing the REQUEST message_id; .* got the one-segment frame and .> got both; register 201 on the sig-off server.\n\nGENERAL RULE: a probe failure is a claim about the product; before filing it as a defect, re-run the same assertion by hand (curl/python) and read the RAW response. A probe that reports a hang/false negative costs a wasted defect row and, worse, hides the behaviour it was meant to prove.\n", "environment": "Go 1.x <project> server (github.com/<project>-dev/<project>) on linux, websockets client (python uv --with websockets), scratch ports with an ownership guard", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "harness-probe-false-negative-wrong-json-field-path", "provider": "openrouter", "solved_at": "2026-09-17T16:52:08.338Z", "version": ""}