Problem class: sse-heartbeat-interval-equals-server-write-timeout
I diagnosed the issue, built a faithful real-server reproduction, and verified both fixes end-to-end. Solution written to /workspace/solution.md:
http.Server.WriteTimeoutProblem class: sse-heartbeat-interval-equals-server-write-timeout
Project: hermes-canopy / canopyd (Go, chi, net/http)
Symptom: A long-lived SSE stream closes cleanly at exactly WriteTimeout; the client receives snapshot/replay frames but zero heartbeat frames. The consumer's liveness flag never latches, and because the stream closes without an error, a retry-on-error transport never fires — the UI silently stops updating.
Affected anchors: internal/server/server.go:148 (WriteTimeout: 30s), internal/handler/card_events_handler.go:40 (heartbeat 30s), internal/handler/gateway_handler.go:359-361 (same repo already uses 20s and documents why).
net/http installs one absolute write deadline per connection, at request-read time, and never refreshes it:
// net/http/server.go, (*conn).readRequest
if d := c.server.WriteTimeout; d > 0 {
defer func() { c.rwc.SetWriteDeadline(time.Now().Add(d)) }()
}
request_start + 30s, fixed for the whole stream.w.Write then returns i/o timeout (os.ErrDeadlineExceeded); the handler returns, silently.200 in 30.0015s; client sees a clean EOF. Any heartbeat >= WriteTimeout has this property.Why the unit test can't see it: it injects a short interval and drives the handler directly/httptest.NewRecorder, bypassing http.Server entirely, so no connection deadline exists. Only a real streaming client against a real http.Server with WriteTimeout set, asserting a heartbeat frame, can detect this.
A (recommended): refresh the deadline per write with http.ResponseController.SetWriteDeadline (Go 1.20+, HTTP/1+HTTP/2), keeping spec cadence even when heartbeat == WriteTimeout:
import (
"errors"
"log/slog"
"net/http"
"time"
)
rc := http.NewResponseController(w)
writeFrame := func(frame []byte) error {
if err := rc.SetWriteDeadline(time.Now().Add(h.writeTimeout)); err != nil &&
!errors.Is(err, http.ErrNotSupported) {
return err
}
if _, err := w.Write(frame); err != nil {
return err
}
return rc.Flush()
}
// initial snapshot, then:
ticker := time.NewTicker(h.heartbeatInterval)
defer ticker.Stop()
for {
select {
case <-r.Context().Done():
return
case <-ticker.C:
if err := writeFrame([]byte(": heartbeat\n\n")); err != nil {
slog.WarnContext(r.Context(), "sse heartbeat write failed", "err", err)
return // never silently
}
}
}
Pass the server's WriteTimeout into the handler so the values can't drift.
B (immediate mitigation): set the card heartbeat strictly below WriteTimeout (20s vs 30s, matching the gateway feed) and guard the invariant in an init() that panics if sseHeartbeatInterval >= serverWriteTimeout. Prefer A.
Also: always log write/flush errors — the silent return is what hid this.
Built and ran a real httptest server with Config.WriteTimeout + a real streaming client (Go 1.26):
| Case | Handler | heartbeat | WriteTimeout | Result |
|---|---|---|---|---|
| Broken (HEAD) | no deadline refresh | 300ms | 300ms | closes at deadline, no heartbeat ❌ |
| Minimal fix (B) | no refresh | 100ms | 300ms | heartbeat delivered ✅ |
| Robust fix (A) | SetWriteDeadline per write |
300ms | 300ms | heartbeat delivered ✅ |
--- PASS: TestBrokenHandlerMissesHeartbeat
reproduced: stream closed at WriteTimeout with only ["event: snapshot" "data: {}" ""]
--- PASS: TestIntervalBelowTimeoutDeliversHeartbeat
--- PASS: TestFixedHandlerDeliversHeartbeat
ok sseproof 1.510s
The full runnable regression test (TestSSEHeartbeatArrivesOnRealServer), manual curl -sN/python timestamped reproduction commands, rollout checklist, and the complete proof-harness files are in the document.
Key takeaway: the regression test must be a real server + real streaming client asserting a heartbeat frame — the existing injected-interval unit test bypasses the deadline mechanism and cannot see this class.
# Evidence - Problem class: sse-heartbeat-interval-equals-server-write-timeout - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T08:20:26.039Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A long-lived SSE stream closes at exactly the server WriteTimeout and the client never receives a single heartbeat frame, even though the handler's heartbeat branch is reached. Root cause: the heartbeat ticker interval equals the http.Server WriteTimeout, so the heartbeat write lands AT the deadline, fails, and the handler returns (silently, no log). Evidence: canopyd sets WriteTimeout 30s (internal/server/server.go:148) while the card SSE handler fixes the heartbeat at 30s (internal/handler/card_events_handler.go:40); a curl -sN stream and a python urllib capture on a HEAD-built instance both ended at 30.0015s with the request logged as '200 in 30.0015s' and only the snapshot/replay frames delivered. Fix: set the heartbeat interval strictly BELOW the write timeout (the same repo's gateway feed uses 20s and documents exactly this reason at internal/handler/gateway_handler.go:359-361), or extend the write deadline per-response with http.ResponseController.SetWriteDeadline inside the stream loop and keep the spec cadence. Detection: the unit test that injects a short heartbeat interval cannot see this class at all \u2014 the regression test must be a real streaming client that asserts a heartbeat FRAME arrives on a real server. Client-side symptom to look for: the consumer's liveness flag (set only by a heartbeat) never latches, and the stream closes CLEANLY, so a retry-on-error transport never fires and the UI silently stops updating.", "environment": "canopyd (Go, chi router, http.Server WriteTimeout) on Linux; HEAD-built instance on a scratch port/DB", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "sse-heartbeat-interval-equals-server-write-timeout", "provider": "openrouter", "solved_at": "2026-09-19T08:20:26.039Z", "version": ""}I diagnosed the issue, built a faithful real-server reproduction, and verified both fixes end-to-end. Solution written to /workspace/solution.md:
http.Server.WriteTimeoutProblem class: sse-heartbeat-interval-equals-server-write-timeout
Project: hermes-canopy / canopyd (Go, chi, net/http)
Symptom: A long-lived SSE stream closes cleanly at exactly WriteTimeout; the client receives snapshot/replay frames but zero heartbeat frames. The consumer's liveness flag never latches, and because the stream closes without an error, a retry-on-error transport never fires — the UI silently stops updating.
Affected anchors: internal/server/server.go:148 (WriteTimeout: 30s), internal/handler/card_events_handler.go:40 (heartbeat 30s), internal/handler/gateway_handler.go:359-361 (same repo already uses 20s and documents why).
net/http installs one absolute write deadline per connection, at request-read time, and never refreshes it:
// net/http/server.go, (*conn).readRequest
if d := c.server.WriteTimeout; d > 0 {
defer func() { c.rwc.SetWriteDeadline(time.Now().Add(d)) }()
}
request_start + 30s, fixed for the whole stream.w.Write then returns i/o timeout (os.ErrDeadlineExceeded); the handler returns, silently.200 in 30.0015s; client sees a clean EOF. Any heartbeat >= WriteTimeout has this property.Why the unit test can't see it: it injects a short interval and drives the handler directly/httptest.NewRecorder, bypassing http.Server entirely, so no connection deadline exists. Only a real streaming client against a real http.Server with WriteTimeout set, asserting a heartbeat frame, can detect this.
A (recommended): refresh the deadline per write with http.ResponseController.SetWriteDeadline (Go 1.20+, HTTP/1+HTTP/2), keeping spec cadence even when heartbeat == WriteTimeout:
import (
"errors"
"log/slog"
"net/http"
"time"
)
rc := http.NewResponseController(w)
writeFrame := func(frame []byte) error {
if err := rc.SetWriteDeadline(time.Now().Add(h.writeTimeout)); err != nil &&
!errors.Is(err, http.ErrNotSupported) {
return err
}
if _, err := w.Write(frame); err != nil {
return err
}
return rc.Flush()
}
// initial snapshot, then:
ticker := time.NewTicker(h.heartbeatInterval)
defer ticker.Stop()
for {
select {
case <-r.Context().Done():
return
case <-ticker.C:
if err := writeFrame([]byte(": heartbeat\n\n")); err != nil {
slog.WarnContext(r.Context(), "sse heartbeat write failed", "err", err)
return // never silently
}
}
}
Pass the server's WriteTimeout into the handler so the values can't drift.
B (immediate mitigation): set the card heartbeat strictly below WriteTimeout (20s vs 30s, matching the gateway feed) and guard the invariant in an init() that panics if sseHeartbeatInterval >= serverWriteTimeout. Prefer A.
Also: always log write/flush errors — the silent return is what hid this.
Built and ran a real httptest server with Config.WriteTimeout + a real streaming client (Go 1.26):
| Case | Handler | heartbeat | WriteTimeout | Result |
|---|---|---|---|---|
| Broken (HEAD) | no deadline refresh | 300ms | 300ms | closes at deadline, no heartbeat ❌ |
| Minimal fix (B) | no refresh | 100ms | 300ms | heartbeat delivered ✅ |
| Robust fix (A) | SetWriteDeadline per write |
300ms | 300ms | heartbeat delivered ✅ |
--- PASS: TestBrokenHandlerMissesHeartbeat
reproduced: stream closed at WriteTimeout with only ["event: snapshot" "data: {}" ""]
--- PASS: TestIntervalBelowTimeoutDeliversHeartbeat
--- PASS: TestFixedHandlerDeliversHeartbeat
ok sseproof 1.510s
The full runnable regression test (TestSSEHeartbeatArrivesOnRealServer), manual curl -sN/python timestamped reproduction commands, rollout checklist, and the complete proof-harness files are in the document.
Key takeaway: the regression test must be a real server + real streaming client asserting a heartbeat frame — the existing injected-interval unit test bypasses the deadline mechanism and cannot see this class.
# Evidence - Problem class: sse-heartbeat-interval-equals-server-write-timeout - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T08:20:26.039Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A long-lived SSE stream closes at exactly the server WriteTimeout and the client never receives a single heartbeat frame, even though the handler's heartbeat branch is reached. Root cause: the heartbeat ticker interval equals the http.Server WriteTimeout, so the heartbeat write lands AT the deadline, fails, and the handler returns (silently, no log). Evidence: canopyd sets WriteTimeout 30s (internal/server/server.go:148) while the card SSE handler fixes the heartbeat at 30s (internal/handler/card_events_handler.go:40); a curl -sN stream and a python urllib capture on a HEAD-built instance both ended at 30.0015s with the request logged as '200 in 30.0015s' and only the snapshot/replay frames delivered. Fix: set the heartbeat interval strictly BELOW the write timeout (the same repo's gateway feed uses 20s and documents exactly this reason at internal/handler/gateway_handler.go:359-361), or extend the write deadline per-response with http.ResponseController.SetWriteDeadline inside the stream loop and keep the spec cadence. Detection: the unit test that injects a short heartbeat interval cannot see this class at all \u2014 the regression test must be a real streaming client that asserts a heartbeat FRAME arrives on a real server. Client-side symptom to look for: the consumer's liveness flag (set only by a heartbeat) never latches, and the stream closes CLEANLY, so a retry-on-error transport never fires and the UI silently stops updating.", "environment": "canopyd (Go, chi router, http.Server WriteTimeout) on Linux; HEAD-built instance on a scratch port/DB", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "sse-heartbeat-interval-equals-server-write-timeout", "provider": "openrouter", "solved_at": "2026-09-19T08:20:26.039Z", "version": ""}