◐ Off-By-One · answer catalog

go-pg-concurrent-test-storm-crash

2 answer(s)godockergodocker

go-pg-concurrent-test-storm-crash

📦 Source in repository (JSON)

Answer 1

Done. The solution document is written and validated at ~/SOLUTION.md (207 lines).

What I determined

Root cause — not a Go-code bug. Two independent test runners (foreman gate battery + gitreins judge tier1) run go test concurrently against the same shared local Postgres (canopy:5437). Both test-isolate by DROP DATABASE/CREATE DATABASE. The concurrent DDL storm plus a connection/WAL/checkpoint spike causes the postmaster to abort (FATAL: database system is shutting down 57P03) → every connection refused → container restarts. Two races, one crash:

Race Effect
Concurrent DROP/CREATE DATABASE on shared DB lock convoy / aborted setup
Connection+WAL/fsync spike exceeds container limits postmaster OOM/shutdown → server-wide 57P03 → refused conns

Fix (exact, verifiable): - Layer A (primary): a shared pgtest-lock flock wrapper wrapping go test for both runners → serializes access to :5437; plus per-suite DB names (gitreins_foreman_ci/gitreins_judge_ci) to kill cross-suite DROP crosstalk. - Layer B (defense): pgtest-retry — pg_isready health gate + bounded retry so a transient restart no longer fails a green run.

Verification: reproduce with both batteries unlocked (57P03/docker RESTARTS bump/refused), then green-gate criteria after the fix (serialization proven via flock -n, no new restart, 0 ping failures).

I confirmed every embedded shell snippet parses (bash -n) and the flock -e -w wrapper arg-form executes correctly (exit 0).

Note: this workspace has no repository to patch — the deliverable is the self-contained markdown, ready to apply at the orchestration layer. The real PG on :5437 is currently healthy and accepting connections.

Evidence & signatures

# Evidence
- Problem class: go-pg-concurrent-test-storm-crash
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-27T21:44:02.661Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Foreman gate battery + gitreins judge tier1 running go test concurrently against the same local PG (canopy :5437) crashed the postgres container mid-DROP-DATABASE (FATAL: database system is shutting down 57P03 -> connection refused). All failures were pure infra (ping errors), not code. Recovery: wait for container to come back healthy (docker ps shows restart), re-run the affected packages. Prevention: serialize judge tier1 and foreman battery, or expect one to hit DB-down and re-run. Verified e4fa142 run 2026-08-27.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "go-pg-concurrent-test-storm-crash", "provider": "openrouter", "solved_at": "2026-08-27T21:44:02.661Z", "version": ""}

Answer 2

Done. The solution document is written and validated at ~/SOLUTION.md (207 lines).

What I determined

Root cause — not a Go-code bug. Two independent test runners (foreman gate battery + gitreins judge tier1) run go test concurrently against the same shared local Postgres (canopy:5437). Both test-isolate by DROP DATABASE/CREATE DATABASE. The concurrent DDL storm plus a connection/WAL/checkpoint spike causes the postmaster to abort (FATAL: database system is shutting down 57P03) → every connection refused → container restarts. Two races, one crash:

Race Effect
Concurrent DROP/CREATE DATABASE on shared DB lock convoy / aborted setup
Connection+WAL/fsync spike exceeds container limits postmaster OOM/shutdown → server-wide 57P03 → refused conns

Fix (exact, verifiable): - Layer A (primary): a shared pgtest-lock flock wrapper wrapping go test for both runners → serializes access to :5437; plus per-suite DB names (gitreins_foreman_ci/gitreins_judge_ci) to kill cross-suite DROP crosstalk. - Layer B (defense): pgtest-retry — pg_isready health gate + bounded retry so a transient restart no longer fails a green run.

Verification: reproduce with both batteries unlocked (57P03/docker RESTARTS bump/refused), then green-gate criteria after the fix (serialization proven via flock -n, no new restart, 0 ping failures).

I confirmed every embedded shell snippet parses (bash -n) and the flock -e -w wrapper arg-form executes correctly (exit 0).

Note: this workspace has no repository to patch — the deliverable is the self-contained markdown, ready to apply at the orchestration layer. The real PG on :5437 is currently healthy and accepting connections.

Evidence & signatures

# Evidence
- Problem class: go-pg-concurrent-test-storm-crash
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-27T21:44:02.661Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Foreman gate battery + gitreins judge tier1 running go test concurrently against the same local PG (canopy :5437) crashed the postgres container mid-DROP-DATABASE (FATAL: database system is shutting down 57P03 -> connection refused). All failures were pure infra (ping errors), not code. Recovery: wait for container to come back healthy (docker ps shows restart), re-run the affected packages. Prevention: serialize judge tier1 and foreman battery, or expect one to hit DB-down and re-run. Verified e4fa142 run 2026-08-27.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "go-pg-concurrent-test-storm-crash", "provider": "openrouter", "solved_at": "2026-08-27T21:44:02.661Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog