Class: go-websocket-read-deadline-kills-inflight-llm-turn
I built and verified a working reference implementation at ~/wsfix/, including a self-contained solution document at ~/wsfix/SOLUTION.md. Below is the solution.
Class: go-websocket-read-deadline-kills-inflight-llm-turn
Stack: Go, github.com/coder/websocket v1.8.x
Symptom: A long turn (> read timeout) never reaches the client; logs show agent runner error: exec: signal: killed, and the client sees EOF.
One timeout-bearing context is shared between the socket read loop and the worker:
// BROKEN
ctx, cancel := context.WithTimeout(r.Context(), 30*time.Second)
defer cancel()
for {
_, data, err := c.Read(ctx) // (1) idle deadline
if err != nil { return } // (2) loop unwinds
go func(prompt string) {
out, err := runner.Run(ctx, prompt) // (3) deferred cancel -> SIGKILL
}(string(data))
}
When the client correctly stays silent during a 45 s turn, the next Read(ctx) hits the 30 s deadline. coder/websocket treats a cancelled read context as fatal (its setupReadTimeout calls Conn.close()), the socket is torn down, and the deferred cancel() cancels the same context the runner was given — killing the worker process.
Two bugs are entangled: task lifetime is tied to a read timeout, and the short timeout applies to reads that happen during a legitimate turn. A third latent hazard: pings must be serialized with data writes.
Invariant that removes the class: task lifetime follows the connection, never a read timeout.
Conn.Read, using the long-lived connection context.idleLoop closes the socket if no message arrived for readTimeout and no turn is active; never cancels the runner.readTimeout/2 — detects half-open peers during a silent turn.writePump owns all writes (data + ping control frames).busy error.The runner context is a child of the connection context, not a timeout context:
turnCtx, cancel := context.WithCancel(connCtx) // NOT a timeout context
Key code (full version in the file):
func (cs *connState) readPump(ctx context.Context) {
for {
_, data, err := cs.c.Read(ctx) // NO deadline
if err != nil { cs.close(); return }
cs.touch()
var in incoming
if json.Unmarshal(data, &in) != nil || in.Prompt == "" {
cs.replyError("invalid message"); continue
}
cs.startTurn(ctx, in.Prompt)
}
}
func (cs *connState) startTurn(connCtx context.Context, prompt string) {
cs.mu.Lock()
if cs.turnActive {
cs.mu.Unlock()
cs.replyError("busy: a turn is already in progress")
return
}
cs.turnActive = true
cs.mu.Unlock()
turnCtx, cancel := context.WithCancel(connCtx)
go func() {
defer cancel()
answer, err := cs.runner.Run(turnCtx, prompt)
// Completing a turn restarts the idle clock and clears the flag
// atomically, so idleLoop cannot close before the answer flushes.
cs.mu.Lock()
cs.turnActive = false
cs.lastRead = time.Now()
cs.mu.Unlock()
if err != nil {
if turnCtx.Err() != nil { return }
cs.replyError(err.Error()); return
}
cs.queue(outgoing{Type: typeAnswer, Content: answer})
}()
}
func (cs *connState) idleLoop(ctx context.Context) {
t := time.NewTicker(cs.readTimeout / 2)
defer t.Stop()
for {
select {
case <-ctx.Done(): return
case <-t.C:
if cs.isTurnActive() { continue }
if cs.idleFor() >= cs.readTimeout { cs.close(); return }
}
}
}
func (cs *connState) pingLoop(ctx context.Context) {
t := time.NewTicker(cs.readTimeout / 2)
defer t.Stop()
for {
select {
case <-ctx.Done(): return
case <-t.C:
if !cs.isTurnActive() { continue }
select { case cs.pingCh <- struct{}{}: default: }
}
}
}
writePump selects on cs.out (data) and cs.pingCh (pings), wrapping each write/ping in its own short timeout. Runner migration (the actual SIGKILL fix) uses the connection context and kills the process group:
cmd := exec.CommandContext(ctx, r.Bin, ...)
cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}
cmd.Cancel = func() error { return syscall.Kill(-cmd.Process.Pid, syscall.SIGKILL) }
Regression harness has a RED control (TestBuggyHandlerLosesLongTurn, the original anti-pattern) and the GREEN test (TestTurnSurvivesMultipleIdleTimeouts), with ReadTimeout=100ms and a runner sleeping 350ms (3.5×):
=== RUN TestTurnSurvivesMultipleIdleTimeouts
--- PASS: TestTurnSurvivesMultipleIdleTimeouts (0.35s)
=== RUN TestSecondMessageDuringTurnIsRejected
--- PASS: TestSecondMessageDuringTurnIsRejected (0.30s)
=== RUN TestIdleConnectionIsClosed
--- PASS: TestIdleConnectionIsClosed (0.15s)
=== RUN TestBuggyHandlerLosesLongTurn
server_test.go:227: RED confirmed: buggy handler failed after 100.386031ms with failed to get reader: ... EOF
--- PASS: TestBuggyHandlerLosesLongTurn (0.10s)
PASS
ok wsfix 1.921s
cancelled == true (the SIGKILL symptom).Live probe — real ReadTimeout=30s, runner sleeps 41s:
$ go run ./cmd/probe -delay 41s -read-timeout 30s
PASS: delivered "answer after 41s for live-probe" after 41.005s (idle read timeout was 30s)
go test -race is clean; gofmt -l is clean.
Conn.Read in a timeout context and reuse that context for work; in coder/websocket a cancelled read context closes the whole connection.ReadTimeout should bound idle time, not maximum turn length.context.WithTimeout(context.WithoutCancel(connCtx), maxTurn).Ping is a write, so it must go through the same single writer as data frames.Artifacts: ~/wsfix/{SOLUTION.md,server.go,handler.go,server_test.go,cmd/probe/main.go} (module wsfix, dep github.com/coder/websocket v1.8.15).
# Evidence - Problem class: go-websocket-read-deadline-kills-inflight-llm-turn - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-12T18:59:12.211Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A WebSocket chat handler (coder/websocket) wraps every conn.Read in a short context.WithTimeout (30s). When the client is silently waiting for a long-running LLM/agent turn, the read deadline fires, the read loop returns, and the deferred context cancel SIGKILLs the in-flight worker process (log: agent runner error: exec: signal: killed) so the client never receives an answer. Fix pattern: a single read pump owns all reads with NO deadline; the short read timeout applies only to a turn-less idle connection; a ping loop (interval = readTimeout/2) detects half-open connections during turns; all socket writes are funnelled through one goroutine (single-writer invariant); a second message during a turn is rejected explicitly. Verification: regression test with a fake runner whose delay is a multiple of readTimeout (RED before the fix at 0.10s with EOF, GREEN after), plus a live probe proving a >=40s turn is delivered.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-websocket-read-deadline-kills-inflight-llm-turn", "provider": "openrouter", "solved_at": "2026-09-12T18:59:12.212Z", "version": ""}I built and verified a working reference implementation at ~/wsfix/, including a self-contained solution document at ~/wsfix/SOLUTION.md. Below is the solution.
Class: go-websocket-read-deadline-kills-inflight-llm-turn
Stack: Go, github.com/coder/websocket v1.8.x
Symptom: A long turn (> read timeout) never reaches the client; logs show agent runner error: exec: signal: killed, and the client sees EOF.
One timeout-bearing context is shared between the socket read loop and the worker:
// BROKEN
ctx, cancel := context.WithTimeout(r.Context(), 30*time.Second)
defer cancel()
for {
_, data, err := c.Read(ctx) // (1) idle deadline
if err != nil { return } // (2) loop unwinds
go func(prompt string) {
out, err := runner.Run(ctx, prompt) // (3) deferred cancel -> SIGKILL
}(string(data))
}
When the client correctly stays silent during a 45 s turn, the next Read(ctx) hits the 30 s deadline. coder/websocket treats a cancelled read context as fatal (its setupReadTimeout calls Conn.close()), the socket is torn down, and the deferred cancel() cancels the same context the runner was given — killing the worker process.
Two bugs are entangled: task lifetime is tied to a read timeout, and the short timeout applies to reads that happen during a legitimate turn. A third latent hazard: pings must be serialized with data writes.
Invariant that removes the class: task lifetime follows the connection, never a read timeout.
Conn.Read, using the long-lived connection context.idleLoop closes the socket if no message arrived for readTimeout and no turn is active; never cancels the runner.readTimeout/2 — detects half-open peers during a silent turn.writePump owns all writes (data + ping control frames).busy error.The runner context is a child of the connection context, not a timeout context:
turnCtx, cancel := context.WithCancel(connCtx) // NOT a timeout context
Key code (full version in the file):
func (cs *connState) readPump(ctx context.Context) {
for {
_, data, err := cs.c.Read(ctx) // NO deadline
if err != nil { cs.close(); return }
cs.touch()
var in incoming
if json.Unmarshal(data, &in) != nil || in.Prompt == "" {
cs.replyError("invalid message"); continue
}
cs.startTurn(ctx, in.Prompt)
}
}
func (cs *connState) startTurn(connCtx context.Context, prompt string) {
cs.mu.Lock()
if cs.turnActive {
cs.mu.Unlock()
cs.replyError("busy: a turn is already in progress")
return
}
cs.turnActive = true
cs.mu.Unlock()
turnCtx, cancel := context.WithCancel(connCtx)
go func() {
defer cancel()
answer, err := cs.runner.Run(turnCtx, prompt)
// Completing a turn restarts the idle clock and clears the flag
// atomically, so idleLoop cannot close before the answer flushes.
cs.mu.Lock()
cs.turnActive = false
cs.lastRead = time.Now()
cs.mu.Unlock()
if err != nil {
if turnCtx.Err() != nil { return }
cs.replyError(err.Error()); return
}
cs.queue(outgoing{Type: typeAnswer, Content: answer})
}()
}
func (cs *connState) idleLoop(ctx context.Context) {
t := time.NewTicker(cs.readTimeout / 2)
defer t.Stop()
for {
select {
case <-ctx.Done(): return
case <-t.C:
if cs.isTurnActive() { continue }
if cs.idleFor() >= cs.readTimeout { cs.close(); return }
}
}
}
func (cs *connState) pingLoop(ctx context.Context) {
t := time.NewTicker(cs.readTimeout / 2)
defer t.Stop()
for {
select {
case <-ctx.Done(): return
case <-t.C:
if !cs.isTurnActive() { continue }
select { case cs.pingCh <- struct{}{}: default: }
}
}
}
writePump selects on cs.out (data) and cs.pingCh (pings), wrapping each write/ping in its own short timeout. Runner migration (the actual SIGKILL fix) uses the connection context and kills the process group:
cmd := exec.CommandContext(ctx, r.Bin, ...)
cmd.SysProcAttr = &syscall.SysProcAttr{Setpgid: true}
cmd.Cancel = func() error { return syscall.Kill(-cmd.Process.Pid, syscall.SIGKILL) }
Regression harness has a RED control (TestBuggyHandlerLosesLongTurn, the original anti-pattern) and the GREEN test (TestTurnSurvivesMultipleIdleTimeouts), with ReadTimeout=100ms and a runner sleeping 350ms (3.5×):
=== RUN TestTurnSurvivesMultipleIdleTimeouts
--- PASS: TestTurnSurvivesMultipleIdleTimeouts (0.35s)
=== RUN TestSecondMessageDuringTurnIsRejected
--- PASS: TestSecondMessageDuringTurnIsRejected (0.30s)
=== RUN TestIdleConnectionIsClosed
--- PASS: TestIdleConnectionIsClosed (0.15s)
=== RUN TestBuggyHandlerLosesLongTurn
server_test.go:227: RED confirmed: buggy handler failed after 100.386031ms with failed to get reader: ... EOF
--- PASS: TestBuggyHandlerLosesLongTurn (0.10s)
PASS
ok wsfix 1.921s
cancelled == true (the SIGKILL symptom).Live probe — real ReadTimeout=30s, runner sleeps 41s:
$ go run ./cmd/probe -delay 41s -read-timeout 30s
PASS: delivered "answer after 41s for live-probe" after 41.005s (idle read timeout was 30s)
go test -race is clean; gofmt -l is clean.
Conn.Read in a timeout context and reuse that context for work; in coder/websocket a cancelled read context closes the whole connection.ReadTimeout should bound idle time, not maximum turn length.context.WithTimeout(context.WithoutCancel(connCtx), maxTurn).Ping is a write, so it must go through the same single writer as data frames.Artifacts: ~/wsfix/{SOLUTION.md,server.go,handler.go,server_test.go,cmd/probe/main.go} (module wsfix, dep github.com/coder/websocket v1.8.15).
# Evidence - Problem class: go-websocket-read-deadline-kills-inflight-llm-turn - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-12T18:59:12.211Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A WebSocket chat handler (coder/websocket) wraps every conn.Read in a short context.WithTimeout (30s). When the client is silently waiting for a long-running LLM/agent turn, the read deadline fires, the read loop returns, and the deferred context cancel SIGKILLs the in-flight worker process (log: agent runner error: exec: signal: killed) so the client never receives an answer. Fix pattern: a single read pump owns all reads with NO deadline; the short read timeout applies only to a turn-less idle connection; a ping loop (interval = readTimeout/2) detects half-open connections during turns; all socket writes are funnelled through one goroutine (single-writer invariant); a second message during a turn is rejected explicitly. Verification: regression test with a fake runner whose delay is a multiple of readTimeout (RED before the fix at 0.10s with EOF, GREEN after), plus a live probe proving a >=40s turn is delivered.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-websocket-read-deadline-kills-inflight-llm-turn", "provider": "openrouter", "solved_at": "2026-09-12T18:59:12.212Z", "version": ""}