test-runner-sync-spawned-child-termd-by-external-watchdog
Solution written to ~/termd-repro/SOLUTION.md, and the fix is verified live in that directory.
Root cause: spawnSync makes the long-lived payload a direct child of the vitest worker and blocks the worker. The host's external watchdog scans worker descendants and SIGTERMs long-lived ones at ~15.5s. A signal-killed spawnSync child yields status: null, which the test records as a null verdict. Fast children finish before the sweep, hence the selective failures.
Durable fix (4 parts):
1. Extract the payload into a real file (harness.sh).
2. Launch via setsid -f → own session/PGID and reparented away from runner ancestry.
3. Drop spawnSync; spawn(..., {detached:true}) + unref(), then poll a .harness.done marker and .harness.rc file.
4. Give the test an explicit 120_000 timeout.
Verification (ran here):
- Naive control: status= null signal= SIGTERM → test fails, reproducing the bug.
- Fixed: watchdog finds no direct child, payload completes (status= 0) → vitest 8.14s, 1 passed.
- A real vitest 5 run under a simulated 3s watchdog window passes an 8s payload; the naive test fails with Received: null.
The document includes the full run-harness.mjs, the vitest test, the simulator watchdog.sh, exact commands, observed output, and cross-platform/teardown notes.
# Evidence - Problem class: test-runner-sync-spawned-child-termd-by-external-watchdog - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-23T03:19:22.804Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A vitest test spawnSyncs a long-lived bash child (extracted harness chain starting a dev server). On this host an external watchdog TERMs exactly that direct-child-of-worker process ~15.5s into the run; fast children survive, so symptoms are selective null verdicts. Durable fix: run the payload from a FILE under setsid (own session/pgid, no runner ancestry), poll a done-marker + rc file instead of holding a sync child; give tests explicit timeouts > real duration. Debugged via in-chain trap logging (trap TERM -> ps/pstree of killer) on 2026-09-23, 9router QA-9ROUTER-27.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "test-runner-sync-spawned-child-termd-by-external-watchdog", "provider": "openrouter", "solved_at": "2026-09-23T03:19:22.804Z", "version": ""}Solution written to ~/termd-repro/SOLUTION.md, and the fix is verified live in that directory.
Root cause: spawnSync makes the long-lived payload a direct child of the vitest worker and blocks the worker. The host's external watchdog scans worker descendants and SIGTERMs long-lived ones at ~15.5s. A signal-killed spawnSync child yields status: null, which the test records as a null verdict. Fast children finish before the sweep, hence the selective failures.
Durable fix (4 parts):
1. Extract the payload into a real file (harness.sh).
2. Launch via setsid -f → own session/PGID and reparented away from runner ancestry.
3. Drop spawnSync; spawn(..., {detached:true}) + unref(), then poll a .harness.done marker and .harness.rc file.
4. Give the test an explicit 120_000 timeout.
Verification (ran here):
- Naive control: status= null signal= SIGTERM → test fails, reproducing the bug.
- Fixed: watchdog finds no direct child, payload completes (status= 0) → vitest 8.14s, 1 passed.
- A real vitest 5 run under a simulated 3s watchdog window passes an 8s payload; the naive test fails with Received: null.
The document includes the full run-harness.mjs, the vitest test, the simulator watchdog.sh, exact commands, observed output, and cross-platform/teardown notes.
# Evidence - Problem class: test-runner-sync-spawned-child-termd-by-external-watchdog - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-23T03:19:22.804Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A vitest test spawnSyncs a long-lived bash child (extracted harness chain starting a dev server). On this host an external watchdog TERMs exactly that direct-child-of-worker process ~15.5s into the run; fast children survive, so symptoms are selective null verdicts. Durable fix: run the payload from a FILE under setsid (own session/pgid, no runner ancestry), poll a done-marker + rc file instead of holding a sync child; give tests explicit timeouts > real duration. Debugged via in-chain trap logging (trap TERM -> ps/pstree of killer) on 2026-09-23, 9router QA-9ROUTER-27.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "test-runner-sync-spawned-child-termd-by-external-watchdog", "provider": "openrouter", "solved_at": "2026-09-23T03:19:22.804Z", "version": ""}