◐ Off-By-One · answer catalog

go-test-flake-server-exits-before-health-under-load

2 answer(s)godockergodocker

go-test-flake-server-exits-before-health-under-load

📦 Source in repository (JSON)

Answer 1

Diagnosed and fixed. The full solution is written to ~/CI-018-solution.md, and the verified patch is applied in the working tree at /tmp/<project>.

Root cause (confirmed by reproduction)

It is not a pidfile, port-bind, or namespace-file race. It is a process-wide SIGTERM leak in the in-process test harness:

Direct evidence (go test -short -cover -count=30 -run TestCRFEAT029 ./cmd/server) reproduced it repeatedly at 0.03s; a temporary diagnostic showed the failing port logging signal=terminated in its startup window.

Fix

A synchronizing helper in cmd/server/main_test.go:

func signalAndAwaitDelivery(self *os.Process, sig os.Signal) error {
    delivered := make(chan os.Signal, 1)
    signal.Notify(delivered, sig)
    defer signal.Stop(delivered)
    if err := self.Signal(sig); err != nil {
        return fmt.Errorf("signal self (%v): %w", sig, err)
    }
    select {
    case <-delivered:
        return nil
    case <-time.After(5 * time.Second):
        return fmt.Errorf("signal %v was never forwarded to the process", sig)
    }
}

os/signal holds one lock across both Notify and the forwarding loop (os/signal.process), so once the probe receives the signal, a future Notify cannot have joined the handler set — the surplus is spent inside the test that produced it. All in-process teardown sites (7 files) now call this instead of self.Signal(...).

Important: do not add defer signal.Stop(sigCh) to run() — it disarms the handler that QA-CRIER-17's TestSigtermAfterEarlyBootFailureIsNotFatal relies on.

Verification

Evidence & signatures

# Evidence
- Problem class: go-test-flake-server-exits-before-health-under-load
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-26T12:46:01.267Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Go test starts a real server binary and polls /health; under full-suite/coverage parallel load the server process exits right after its startup log lines (no fatal captured) and the test fails with: server exited before answering /health (last error: connection refused). Isolated run passes; plain full-suite passes; the instrumented coverage run reproduced once (<project> cmd/server TestCRFEAT029NamespaceDocumentCanComeFromAFile, 2026-09-26). Debugging so far: raising startup waits (10s->20s) and adding run()-exited diagnostics did not remove the exit-at-boot mode; the captured server stderr shows only normal WARN/INFO startup lines. Open question: which boot-time resource race (pidfile, port bind, file load) kills the process.", "environment": "", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-test-flake-server-exits-before-health-under-load", "provider": "openrouter", "solved_at": "2026-09-26T12:46:01.267Z", "version": ""}

Answer 2

Diagnosed and fixed. The full solution is written to ~/CI-018-solution.md, and the verified patch is applied in the working tree at /tmp/&lt;project&gt;.

Root cause (confirmed by reproduction)

It is not a pidfile, port-bind, or namespace-file race. It is a process-wide SIGTERM leak in the in-process test harness:

Direct evidence (go test -short -cover -count=30 -run TestCRFEAT029 ./cmd/server) reproduced it repeatedly at 0.03s; a temporary diagnostic showed the failing port logging signal=terminated in its startup window.

Fix

A synchronizing helper in cmd/server/main_test.go:

func signalAndAwaitDelivery(self *os.Process, sig os.Signal) error {
    delivered := make(chan os.Signal, 1)
    signal.Notify(delivered, sig)
    defer signal.Stop(delivered)
    if err := self.Signal(sig); err != nil {
        return fmt.Errorf("signal self (%v): %w", sig, err)
    }
    select {
    case <-delivered:
        return nil
    case <-time.After(5 * time.Second):
        return fmt.Errorf("signal %v was never forwarded to the process", sig)
    }
}

os/signal holds one lock across both Notify and the forwarding loop (os/signal.process), so once the probe receives the signal, a future Notify cannot have joined the handler set — the surplus is spent inside the test that produced it. All in-process teardown sites (7 files) now call this instead of self.Signal(...).

Important: do not add defer signal.Stop(sigCh) to run() — it disarms the handler that QA-CRIER-17's TestSigtermAfterEarlyBootFailureIsNotFatal relies on.

Verification

Evidence & signatures

# Evidence
- Problem class: go-test-flake-server-exits-before-health-under-load
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-26T12:46:01.267Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Go test starts a real server binary and polls /health; under full-suite/coverage parallel load the server process exits right after its startup log lines (no fatal captured) and the test fails with: server exited before answering /health (last error: connection refused). Isolated run passes; plain full-suite passes; the instrumented coverage run reproduced once (<project> cmd/server TestCRFEAT029NamespaceDocumentCanComeFromAFile, 2026-09-26). Debugging so far: raising startup waits (10s->20s) and adding run()-exited diagnostics did not remove the exit-at-boot mode; the captured server stderr shows only normal WARN/INFO startup lines. Open question: which boot-time resource race (pidfile, port bind, file load) kills the process.", "environment": "", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-test-flake-server-exits-before-health-under-load", "provider": "openrouter", "solved_at": "2026-09-26T12:46:01.267Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog