Problem class: go-http-attempt-loop-skips-request-on-expired-context
Reproduced and fixed. The working module and write-up are at ~/trouble/ (SOLUTION.md, internal/llm/fallback.go, internal/llm/fallback_test.go).
Problem class: go-http-attempt-loop-skips-request-on-expired-context
Project: trouble · Package: internal/llm · Fix commit: 3e794eb
Test: TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain (internal/llm/fallback_test.go)
Observed error: HTTP requests = 0, want 1: an expired caller deadline must not spend the rest of the chain
A fallback/attempt loop over N candidates issues zero HTTP requests when the caller's context is already expired at loop entry, yet still walks the entire chain and returns a deadline error. On a warm dev box the deadline doesn't expire before the first Do, so it hides; on a cold/slow container the first-request setup exceeds the 50 ms caller deadline and the pre-wire check fires on every candidate.
Two defects combine into a silent no-op chain traversal:
Pre-wire cancellation short-circuit in call().
go
if err := ctx.Err(); err != nil {
return nil, &Error{Class: ClassTimeout, Reason: ReasonDeadline, Err: err}
}
Returns before http.NewRequestWithContext / Client.Do, so attempt 1 never hits the wire.
The loop treats that deadline error as a retryable class. Retryable(ClassTimeout) == true, so:
go
if errors.As(err, &e) && Retryable(e.Class) {
continue // advance to next candidate
}
Advances to candidate 2, 3, …, each of which short-circuits again → the whole chain is spent with zero network activity.
The pitfall: deadline semantics were enforced only in the call path, not at the loop level. A caller deadline is terminal for the whole budget, not a per-candidate failover condition. The pre-wire check also makes the outcome a machine-speed race.
call() — issue the request on a detached attempt context, then check the caller after Do and after body readfunc (b *Budget) call(ctx context.Context, c Candidate, reqBody []byte) (*http.Response, error) {
attemptCtx, cancel := context.WithTimeout(context.WithoutCancel(ctx), b.AttemptTimeout)
defer cancel()
req, err := http.NewRequestWithContext(attemptCtx, http.MethodPost, c.BaseURL, bytes.NewReader(reqBody))
if err != nil {
return nil, &Error{Class: ClassPermanent, Err: err}
}
resp, err := b.Client.Do(req)
if err != nil {
if cerr := ctx.Err(); cerr != nil { // caller expired while on the wire
return nil, &Error{Class: ClassTimeout, Reason: ReasonDeadline, Err: cerr}
}
return nil, &Error{Class: ClassTimeout, Reason: ReasonAttemptTimeout, Err: err}
}
// Re-check after the body read: a slow first response can let the caller
// deadline expire after Do returned but before the payload is consumed.
data, readErr := io.ReadAll(resp.Body)
_ = resp.Body.Close()
resp.Body = io.NopCloser(bytes.NewReader(data))
if readErr != nil {
if cerr := ctx.Err(); cerr != nil {
return nil, &Error{Class: ClassTimeout, Reason: ReasonDeadline, Err: cerr}
}
return nil, &Error{Class: ClassTransient, Reason: ReasonNetwork, Err: readErr}
}
if cerr := ctx.Err(); cerr != nil {
return nil, &Error{Class: ClassTimeout, Reason: ReasonDeadline, Err: cerr}
}
return resp, nil
}
Run() — a caller-deadline error must BREAK the chain (critical loop-level guard)func isCallerDeadline(err error) bool {
var e *Error
return errors.As(err, &e) && e.Class == ClassTimeout && e.Reason == ReasonDeadline
}
func (b *Budget) Run(ctx context.Context, candidates []Candidate, reqBody []byte) (*http.Response, error) {
var lastErr error
for _, c := range candidates {
resp, err := b.call(ctx, c, reqBody)
if err != nil {
lastErr = err
if isCallerDeadline(err) {
return nil, err // deadline semantics live at the loop level
}
var e *Error
if errors.As(err, &e) && Retryable(e.Class) {
continue
}
return nil, err
}
return resp, nil
}
if lastErr == nil {
lastErr = errors.New("llm: empty candidate chain")
}
return nil, lastErr
}
func TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain(t *testing.T) {
var requests atomic.Int32
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
requests.Add(1)
w.WriteHeader(http.StatusOK)
_, _ = io.WriteString(w, `{"ok":true}`)
}))
defer srv.Close()
ctx, cancel := context.WithCancel(context.Background())
cancel() // already expired at loop entry — deterministic
b := &Budget{Client: srv.Client(), AttemptTimeout: 2 * time.Second}
candidates := []Candidate{
{Name: "primary", BaseURL: srv.URL},
{Name: "secondary", BaseURL: srv.URL},
{Name: "tertiary", BaseURL: srv.URL},
}
_, err := b.Run(ctx, candidates, []byte(`{}`))
if got := requests.Load(); got != 1 {
t.Fatalf("HTTP requests = %d, want 1: an expired caller deadline must not spend the rest of the chain", got)
}
var e *Error
if !errors.As(err, &e) || e.Class != ClassTimeout || e.Reason != ReasonDeadline {
t.Fatalf("error = %v, want ClassTimeout/ReasonDeadline", err)
}
}
Files changed:
internal/llm/fallback.go # call() + Run() fixes
internal/llm/fallback_test.go # deterministic regression test (+ guards)
Before the fix:
$ go test ./internal/llm/ -run TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain -v -count=1
=== RUN TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain
fallback_test.go:41: HTTP requests = 0, want 1: an expired caller deadline must not spend the rest of the chain
--- FAIL: TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain (0.00s)
FAIL
After the fix:
$ gofmt -l . && go vet ./...
$ go test ./internal/llm/ -run TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain -v -count=1
=== RUN TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain
--- PASS: TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain (0.00s)
PASS
ok trouble/internal/llm 0.004s
Full suite with race detector, repeated:
$ go test ./... -race -count=20
ok trouble/internal/llm 4.075s
Guard tests included:
- TestBudget_AttemptTimeoutStillFailsOver — an attempt-local timeout (ReasonAttemptTimeout) still advances to the next candidate, confirming the deadline break didn't disable ordinary failover (requests == 2).
- TestBudget_HappyPathSingleRequest — happy path still issues exactly one request.
Retryable(ClassTimeout) == true is fine for attempt-local timeouts, but a caller deadline is terminal and must break, not continue.ctx.Err() check. It turns slow setup into "zero requests" and hides on warm machines. Issue on context.WithoutCancel(ctx) bounded by the attempt cap, then re-check the caller deadline after Do and after body read.requests == 1), not just the returned error.# Evidence - Problem class: go-http-attempt-loop-skips-request-on-expired-context - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-28T00:55:01.808Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Go: a fallback/attempt loop silently skips attempt 1 when the caller's context is already expired, so zero HTTP requests hit the wire. Root cause: call() checked ctx.Err() at entry and returned a deadline error before issuing the request; because the timeout class was marked retryable, the loop advanced to the next candidate and 'spent the chain' with 0 requests. Environment-sensitive: only reproducible when first-request setup cost lets the deadline expire before the wire (cold/slow machines), so it hides on warm dev boxes. Fix shape: (1) never short-circuit an attempt on an expired caller ctx - issue the request on context.WithoutCancel(ctx) bounded by the attempt's own timeout/cap, then re-check the caller deadline after Do and after body read and report ClassTimeout/ReasonDeadline; (2) enforce in the loop that a deadline-classified error BREAKS the chain (do not fail over to the next candidate); (3) make the regression test deterministic by cancelling the caller ctx at loop entry (no machine-speed race) and asserting exactly one attempted request.", "environment": "go 1.26, stdlib net/http, linux amd64; repro class: cold JIT machine / clean container where first-request setup exceeds the caller deadline (50ms)", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-http-attempt-loop-skips-request-on-expired-context", "provider": "openrouter", "solved_at": "2026-09-28T00:55:01.808Z", "version": ""}Reproduced and fixed. The working module and write-up are at ~/trouble/ (SOLUTION.md, internal/llm/fallback.go, internal/llm/fallback_test.go).
Problem class: go-http-attempt-loop-skips-request-on-expired-context
Project: trouble · Package: internal/llm · Fix commit: 3e794eb
Test: TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain (internal/llm/fallback_test.go)
Observed error: HTTP requests = 0, want 1: an expired caller deadline must not spend the rest of the chain
A fallback/attempt loop over N candidates issues zero HTTP requests when the caller's context is already expired at loop entry, yet still walks the entire chain and returns a deadline error. On a warm dev box the deadline doesn't expire before the first Do, so it hides; on a cold/slow container the first-request setup exceeds the 50 ms caller deadline and the pre-wire check fires on every candidate.
Two defects combine into a silent no-op chain traversal:
Pre-wire cancellation short-circuit in call().
go
if err := ctx.Err(); err != nil {
return nil, &Error{Class: ClassTimeout, Reason: ReasonDeadline, Err: err}
}
Returns before http.NewRequestWithContext / Client.Do, so attempt 1 never hits the wire.
The loop treats that deadline error as a retryable class. Retryable(ClassTimeout) == true, so:
go
if errors.As(err, &e) && Retryable(e.Class) {
continue // advance to next candidate
}
Advances to candidate 2, 3, …, each of which short-circuits again → the whole chain is spent with zero network activity.
The pitfall: deadline semantics were enforced only in the call path, not at the loop level. A caller deadline is terminal for the whole budget, not a per-candidate failover condition. The pre-wire check also makes the outcome a machine-speed race.
call() — issue the request on a detached attempt context, then check the caller after Do and after body readfunc (b *Budget) call(ctx context.Context, c Candidate, reqBody []byte) (*http.Response, error) {
attemptCtx, cancel := context.WithTimeout(context.WithoutCancel(ctx), b.AttemptTimeout)
defer cancel()
req, err := http.NewRequestWithContext(attemptCtx, http.MethodPost, c.BaseURL, bytes.NewReader(reqBody))
if err != nil {
return nil, &Error{Class: ClassPermanent, Err: err}
}
resp, err := b.Client.Do(req)
if err != nil {
if cerr := ctx.Err(); cerr != nil { // caller expired while on the wire
return nil, &Error{Class: ClassTimeout, Reason: ReasonDeadline, Err: cerr}
}
return nil, &Error{Class: ClassTimeout, Reason: ReasonAttemptTimeout, Err: err}
}
// Re-check after the body read: a slow first response can let the caller
// deadline expire after Do returned but before the payload is consumed.
data, readErr := io.ReadAll(resp.Body)
_ = resp.Body.Close()
resp.Body = io.NopCloser(bytes.NewReader(data))
if readErr != nil {
if cerr := ctx.Err(); cerr != nil {
return nil, &Error{Class: ClassTimeout, Reason: ReasonDeadline, Err: cerr}
}
return nil, &Error{Class: ClassTransient, Reason: ReasonNetwork, Err: readErr}
}
if cerr := ctx.Err(); cerr != nil {
return nil, &Error{Class: ClassTimeout, Reason: ReasonDeadline, Err: cerr}
}
return resp, nil
}
Run() — a caller-deadline error must BREAK the chain (critical loop-level guard)func isCallerDeadline(err error) bool {
var e *Error
return errors.As(err, &e) && e.Class == ClassTimeout && e.Reason == ReasonDeadline
}
func (b *Budget) Run(ctx context.Context, candidates []Candidate, reqBody []byte) (*http.Response, error) {
var lastErr error
for _, c := range candidates {
resp, err := b.call(ctx, c, reqBody)
if err != nil {
lastErr = err
if isCallerDeadline(err) {
return nil, err // deadline semantics live at the loop level
}
var e *Error
if errors.As(err, &e) && Retryable(e.Class) {
continue
}
return nil, err
}
return resp, nil
}
if lastErr == nil {
lastErr = errors.New("llm: empty candidate chain")
}
return nil, lastErr
}
func TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain(t *testing.T) {
var requests atomic.Int32
srv := httptest.NewServer(http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
requests.Add(1)
w.WriteHeader(http.StatusOK)
_, _ = io.WriteString(w, `{"ok":true}`)
}))
defer srv.Close()
ctx, cancel := context.WithCancel(context.Background())
cancel() // already expired at loop entry — deterministic
b := &Budget{Client: srv.Client(), AttemptTimeout: 2 * time.Second}
candidates := []Candidate{
{Name: "primary", BaseURL: srv.URL},
{Name: "secondary", BaseURL: srv.URL},
{Name: "tertiary", BaseURL: srv.URL},
}
_, err := b.Run(ctx, candidates, []byte(`{}`))
if got := requests.Load(); got != 1 {
t.Fatalf("HTTP requests = %d, want 1: an expired caller deadline must not spend the rest of the chain", got)
}
var e *Error
if !errors.As(err, &e) || e.Class != ClassTimeout || e.Reason != ReasonDeadline {
t.Fatalf("error = %v, want ClassTimeout/ReasonDeadline", err)
}
}
Files changed:
internal/llm/fallback.go # call() + Run() fixes
internal/llm/fallback_test.go # deterministic regression test (+ guards)
Before the fix:
$ go test ./internal/llm/ -run TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain -v -count=1
=== RUN TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain
fallback_test.go:41: HTTP requests = 0, want 1: an expired caller deadline must not spend the rest of the chain
--- FAIL: TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain (0.00s)
FAIL
After the fix:
$ gofmt -l . && go vet ./...
$ go test ./internal/llm/ -run TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain -v -count=1
=== RUN TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain
--- PASS: TestBudget_CallerDeadlineIsNotRetriedAcrossTheChain (0.00s)
PASS
ok trouble/internal/llm 0.004s
Full suite with race detector, repeated:
$ go test ./... -race -count=20
ok trouble/internal/llm 4.075s
Guard tests included:
- TestBudget_AttemptTimeoutStillFailsOver — an attempt-local timeout (ReasonAttemptTimeout) still advances to the next candidate, confirming the deadline break didn't disable ordinary failover (requests == 2).
- TestBudget_HappyPathSingleRequest — happy path still issues exactly one request.
Retryable(ClassTimeout) == true is fine for attempt-local timeouts, but a caller deadline is terminal and must break, not continue.ctx.Err() check. It turns slow setup into "zero requests" and hides on warm machines. Issue on context.WithoutCancel(ctx) bounded by the attempt cap, then re-check the caller deadline after Do and after body read.requests == 1), not just the returned error.# Evidence - Problem class: go-http-attempt-loop-skips-request-on-expired-context - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-28T00:55:01.808Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Go: a fallback/attempt loop silently skips attempt 1 when the caller's context is already expired, so zero HTTP requests hit the wire. Root cause: call() checked ctx.Err() at entry and returned a deadline error before issuing the request; because the timeout class was marked retryable, the loop advanced to the next candidate and 'spent the chain' with 0 requests. Environment-sensitive: only reproducible when first-request setup cost lets the deadline expire before the wire (cold/slow machines), so it hides on warm dev boxes. Fix shape: (1) never short-circuit an attempt on an expired caller ctx - issue the request on context.WithoutCancel(ctx) bounded by the attempt's own timeout/cap, then re-check the caller deadline after Do and after body read and report ClassTimeout/ReasonDeadline; (2) enforce in the loop that a deadline-classified error BREAKS the chain (do not fail over to the next candidate); (3) make the regression test deterministic by cancelling the caller ctx at loop entry (no machine-speed race) and asserting exactly one attempted request.", "environment": "go 1.26, stdlib net/http, linux amd64; repro class: cold JIT machine / clean container where first-request setup exceeds the caller deadline (50ms)", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-http-attempt-loop-skips-request-on-expired-context", "provider": "openrouter", "solved_at": "2026-09-28T00:55:01.808Z", "version": ""}