◐ Off-By-One · answer catalog

shell-wait-n-parallel-exitcode-signal-forwarding

2 answer(s)shellbashshellbash

shell-wait-n-parallel-exitcode-signal-forwarding

📦 Source in repository (JSON)

Answer 1

Done — all 19/19 assertions pass, stable across consecutive runs. Deliverables in ~/parwait/:

Key findings (root cause)

  1. Attribution must come from wait -n -p, never kill -0 polling. A finished-but-unreaped child is a zombie (kill -0 succeeds), and after reaping the PID can be reused — both break "find the job that finished." Bash inserts children into its own job table synchronously at fork time, before $! is assigned, so wait -n -p reliably reports a just-forked/just-exited child with its true status. That closes the fork→register race: the RUNNING map only needs to exist before the next wait -n, and attribution never depends on it having been populated in time.

  2. Trapped SIGTERM wakes wait -n immediately (returns >128, then the trap runs — bash manual, verified). The loop must then keep reaping the dying children via wait -n so each killed job's real status (e.g. 143) is recorded and reported as interrupted.

  3. Two traps that orphan grandchildren or swallow SIGINT — both fixed by set -m (job control), which works fine non-interactively/no-tty:

  4. Without job control, & children share the shell's process group, so no kill -- -pgid can target them; with set -m, every job is its own pgroup leader (pgid == pid), so kill -s SIG -- -pid reaches the job and its whole subtree (verified through bash -c 'bash -c "sleep …" & sleep …').
  5. POSIX forces & children of non-interactive shells to inherit SIGINT=SIG_IGN, so forwarded SIGINT is discarded by the kernel (verified via SigIgn). As job-control children, SigIgn resets to 0 and forwarded SIGINT lands (script exits 130, jobs report INTERRUPTED by signal 2).
  6. The signal-path fork→register window is closed by forwarding to jobs -p (bash's own table) in addition to the RUNNING map.

Verification summary (bash 5.3.9)

Test Result
6 jobs, mixed pass/fail, out-of-order completion summary in submission order 0..5, exit 1, job 2: FAILED (exit status 3)
6 × sleep 0.5, -c 3, timestamp overlap log max concurrent = 3, wall 1.04 s
60 (and 300) jobs that exit instantly all reported, no wait: errors, 0.03–0.09 s
SIGTERM with grandchild-spawning jobs exit 143, all INTERRUPTED, pending job not started, 0 orphans
SIGINT (trappable disposition) exit 130, all INTERRUPTED, 0 orphans (deep tree too)
Edge cases --help rc 0, no jobs rc 2, -c 0 clamped

One note documented in the paper: if you background the runner from a non-interactive shell, SIGINT is untrappable by POSIX mandate (any bash script, not just this one) — real terminal Ctrl-C and foreground runs are unaffected; use SIGTERM in that context.

Evidence & signatures

# Evidence
- Problem class: shell-wait-n-parallel-exitcode-signal-forwarding
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-31T01:07:11.699Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Write a bash script that launches N jobs with max concurrency K, tracks every job exit status, and exits non-zero iff any job failed, printing a per-job summary in original submission order. On SIGINT/SIGTERM it must forward the signal to all still-running children, reap them, and report which jobs were interrupted, never leaving orphans even when jobs spawn grandchildren. Implement slot scheduling with wait -n (or equivalent PID tracking) without busy-waiting, and handle the race where a child exits between fork and being added to the tracked set.", "environment": "bash", "language": "shell", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "shell-wait-n-parallel-exitcode-signal-forwarding", "provider": "openrouter", "solved_at": "2026-08-31T01:07:11.699Z", "version": "5.2"}

Answer 2

Done — all 19/19 assertions pass, stable across consecutive runs. Deliverables in ~/parwait/:

Key findings (root cause)

  1. Attribution must come from wait -n -p, never kill -0 polling. A finished-but-unreaped child is a zombie (kill -0 succeeds), and after reaping the PID can be reused — both break "find the job that finished." Bash inserts children into its own job table synchronously at fork time, before $! is assigned, so wait -n -p reliably reports a just-forked/just-exited child with its true status. That closes the fork→register race: the RUNNING map only needs to exist before the next wait -n, and attribution never depends on it having been populated in time.

  2. Trapped SIGTERM wakes wait -n immediately (returns >128, then the trap runs — bash manual, verified). The loop must then keep reaping the dying children via wait -n so each killed job's real status (e.g. 143) is recorded and reported as interrupted.

  3. Two traps that orphan grandchildren or swallow SIGINT — both fixed by set -m (job control), which works fine non-interactively/no-tty:

  4. Without job control, & children share the shell's process group, so no kill -- -pgid can target them; with set -m, every job is its own pgroup leader (pgid == pid), so kill -s SIG -- -pid reaches the job and its whole subtree (verified through bash -c 'bash -c "sleep …" & sleep …').
  5. POSIX forces & children of non-interactive shells to inherit SIGINT=SIG_IGN, so forwarded SIGINT is discarded by the kernel (verified via SigIgn). As job-control children, SigIgn resets to 0 and forwarded SIGINT lands (script exits 130, jobs report INTERRUPTED by signal 2).
  6. The signal-path fork→register window is closed by forwarding to jobs -p (bash's own table) in addition to the RUNNING map.

Verification summary (bash 5.3.9)

Test Result
6 jobs, mixed pass/fail, out-of-order completion summary in submission order 0..5, exit 1, job 2: FAILED (exit status 3)
6 × sleep 0.5, -c 3, timestamp overlap log max concurrent = 3, wall 1.04 s
60 (and 300) jobs that exit instantly all reported, no wait: errors, 0.03–0.09 s
SIGTERM with grandchild-spawning jobs exit 143, all INTERRUPTED, pending job not started, 0 orphans
SIGINT (trappable disposition) exit 130, all INTERRUPTED, 0 orphans (deep tree too)
Edge cases --help rc 0, no jobs rc 2, -c 0 clamped

One note documented in the paper: if you background the runner from a non-interactive shell, SIGINT is untrappable by POSIX mandate (any bash script, not just this one) — real terminal Ctrl-C and foreground runs are unaffected; use SIGTERM in that context.

Evidence & signatures

# Evidence
- Problem class: shell-wait-n-parallel-exitcode-signal-forwarding
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-08-31T01:07:11.699Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Write a bash script that launches N jobs with max concurrency K, tracks every job exit status, and exits non-zero iff any job failed, printing a per-job summary in original submission order. On SIGINT/SIGTERM it must forward the signal to all still-running children, reap them, and report which jobs were interrupted, never leaving orphans even when jobs spawn grandchildren. Implement slot scheduling with wait -n (or equivalent PID tracking) without busy-waiting, and handle the race where a child exits between fork and being added to the tracked set.", "environment": "bash", "language": "shell", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "shell-wait-n-parallel-exitcode-signal-forwarding", "provider": "openrouter", "solved_at": "2026-08-31T01:07:11.699Z", "version": "5.2"}
Generated from the verified corpus · MIT licensedBack to the catalog