◐ Off-By-One · answer catalog

bash-concurrent-job-timeout-kill-escalation-exit-code-capture

2 answer(s)shellbashshellbash

A supervisor that runs N jobs concurrently, gives each an independent deadline, escalates SIGTERM → grace → SIGKILL, captures both output streams without deadlocking, reports true 128+signal exit status, and leaves no orphaned grandchildren. It is free of the three named races and the per-job deadline is enforced by an independent watchdog, so a fork/exec storm cannot delay it.

📦 Source in repository (JSON)

Answer 1

The solution is complete and verified. The full markdown is saved at ~/jobs/SOLUTION.md (and the working files at ~/jobs/supervisor.sh, ~/jobs/stress.sh). Here it is:


A Race-Free Concurrent Job Supervisor in Pure Bash 5.2+

A supervisor that runs N jobs concurrently, gives each an independent deadline, escalates SIGTERM → grace → SIGKILL, captures both output streams without deadlocking, reports true 128+signal exit status, and leaves no orphaned grandchildren. It is free of the three named races and the per-job deadline is enforced by an independent watchdog, so a fork/exec storm cannot delay it.


1. Root-cause analysis

1.1 Reaping race and the wait -n trap

The textbook race is a trap ... CHLD handler that calls wait, racing a later wait $pid on the already-reaped pid. The usual "modern" answer is wait -n -p var. But wait -n is not reliable when a child is killed by a signal. Once the asynchronous notifier has collected a SIGKILLed child, wait -n returns 127 ("no unwaited-for children") and never reports that pid, while an explicit wait <pid> still returns the stored 128+signal status:

$ for i in 1 2 3; do setsid bash -c 'trap "" TERM; while :; do :; done' & pids+=($!); done
$ # ... kill -KILL each group ...
$ for i in 1 2 3; do wait -n -p dp; echo "rc=$? dp=${dp:-<unset>}"; done
rc=127 dp=<unset>
rc=127 dp=<unset>
rc=127 dp=<unset>
$ for p in "${pids[@]}"; do wait "$p"; echo "rc=$?"; done
rc=137
rc=137
rc=137

Any design depending on wait -n both hangs (missing children) and loses their statuses.

1.2 Zombies make kill -0 lie

A zombie still exists, so kill -0 $pid succeeds after the process is dead. A poll loop keyed only on kill -0 waits for bash to reap, not for the process to die. Completion must be detected from the kernel: /proc/<pid>/stat state Z (or absence).

1.3 A single post-spawn enforcement loop is starved

Spawning 200 jobs serially (each a fork/exec, some starting CPU-heavy pipelines) takes seconds. If enforcement only starts after all spawns, early jobs are killed far too late, and a sleep 5 job with a 0.4 s timeout can even finish normally before the first pass — reporting rc=0 instead of 143. Measured: the spawn phase alone ran ~1.3 s, and the first enforcement pass ran ~1.5 s after the first job was already dead.

1.4 The timer starts before the child is exec'd

deadline=$(( now + timeout )) computed in the parent before & starts the clock during fork/exec. The start time must be owned by the child, recorded after it is actually running.

1.5 Pipe deadlock and lost pipeline status

out=$(cmd) (or draining only stdout) blocks once the other stream fills its pipe buffer. cmd | tee f; rc=$? yields tee's status, not cmd's. Avoid the pipe: redirect each stream to its own regular file and take the status from wait.

1.6 Orphaned grandchildren

Signalling only the direct child leaves its children behind. Each job must run in its own process group/session and be signalled as a group (kill -- -pgid).

1.7 Fork-per-sleep overhead

A watchdog that sleeps in increments by invoking /bin/sleep forks every tick. Under 200 watchdogs this adds enough load to increase kill latency. Bash's read -t performs a select() timeout with no child; a never-written FIFO is a fork-free clock.

1.8 Overloaded state (implementation pitfall)

Using one killed flag for both "I sent SIGKILL" and "reaped by signal" makes the collector skip live-but-killing jobs and loop forever. Use termed → killing while alive, and a distinct final killed state.


2. The design that closes every hole

Requirement Mechanism
Independent deadlines one watchdog process per job
Timer starts at exec, not fork child writes EPOCHREALTIME to ready; watchdog arms from it
Enforce during spawn storm watchdog is a separate process
SIGTERM → grace → SIGKILL watchdog signals the job's process group
No orphaned grandchildren setsid (or set -m); kill -- -pgid
No pipe deadlock per-stream regular files
No lost pipeline status no pipeline; status from wait <pid>
No double reap no SIGCHLD trap; exactly one wait <pid> per job
Prompt completion detection /proc/<pid>/stat Z / absence
Fork-free timing read -t on a never-written FIFO

3. The fix — supervisor.sh

#!/usr/bin/env bash
# =====================================================================
# supervisor.sh - concurrent job supervisor for bash 5.2+
#
# Design goals and how each classic race is closed:
#
#   * N concurrent jobs, each with its own independent deadline.
#   * SIGTERM -> (grace) -> SIGKILL escalation delivered to the whole
#     process group of the job, so grandchildren cannot be orphaned.
#   * stdout/stderr go to independent regular files, so a full pipe
#     buffer can never deadlock the supervisor.
#   * True status is obtained with one explicit `wait <pid>` per job;
#     a signal-killed job reports 128+signal (TERM=143, KILL=137).
#     Nothing is piped, so a subshell's status is never replaced by a
#     pipeline's last command.
#   * No SIGCHLD trap ever calls wait(), so the reap/wait double-reap
#     race cannot happen; each child is collected exactly once.
#   * A per-job watchdog process starts the deadline clock only after
#     the child itself writes a "ready" timestamp as its first action.
#     The watchdog runs concurrently with (and independently of) the
#     supervisor's spawn loop, so a slow fork/exec storm cannot delay
#     an early job's kill.  (This is the "timer starts before exec"
#     race: the child owns the start time and an independent process
#     owns the timer.)
#
# Usage:  supervisor.sh JOBS_FILE
#   JOBS_FILE lines:  <id> <timeout_sec> <grace_sec> <command...>
#     Fields separated by whitespace; the command is the rest of the line.
#
# Output: one TSV line per job on stdout:
#   id pid rc state start_us term_us kill_us dead_us end_us out_bytes
#   err_bytes timeout_us grace_us
#   start is written by the child; term/kill/dead by the watchdog; end is
#   when the supervisor collected the status.
#
# Env:
#   SUP_WORKDIR  dir for per-job out/err + timestamps (default: mktemp -d)
#   SUP_CLEANUP  1 = delete SUP_WORKDIR on exit (default: 0)
#   SUP_SETSID   auto|0|1  (default auto: use setsid if available)
# =====================================================================

set -u

declare -A JOB_CMD JOB_TIMEOUT_US JOB_GRACE_US
declare -A JOB_PID JOB_RC JOB_STATE JOB_END_US
declare -a JOB_IDS=()

SUP_WORKDIR=${SUP_WORKDIR:-$(mktemp -d)}
SUP_CLEANUP=${SUP_CLEANUP:-0}
SUP_SETSID=${SUP_SETSID:-auto}

# Create one process group per job.
#   setsid (util-linux): no job control, no async shell notices.
#   set -m (pure bash):  same process-group semantics; bash may print a
#                        "Killed" notice to stderr for a SIGKILLed job.
if [[ $SUP_SETSID == auto ]]; then
  if command -v setsid >/dev/null 2>&1; then SUP_SETSID=1; else SUP_SETSID=0; fi
fi
(( SUP_SETSID )) || set -m

_cleanup() { (( SUP_CLEANUP )) && rm -rf -- "$SUP_WORKDIR"; return 0; }
trap _cleanup EXIT

# A never-written FIFO gives us a fork-free sleep primitive: bash's
# `read -t` uses select() internally and needs no child.  All watchdogs
# inherit the read/write fd; with no writer ever sending data, every
# read simply times out.
mkdir -p -- "$SUP_WORKDIR"
SUP_TICK=$SUP_WORKDIR/.tick
rm -f -- "$SUP_TICK"
mkfifo -- "$SUP_TICK"
exec {SUP_TICK_FD}<>"$SUP_TICK"

# ---- pure-bash time helpers -----------------------------------------

# decimal seconds -> integer microseconds
_us() {
  local x=$1 i f
  if [[ $x == *.* ]]; then
    i=${x%%.*}; f=${x#*.}; f=${f}000000
    printf '%s' $(( i * 1000000 + 10#${f:0:6} ))
  else
    printf '%s' $(( x * 1000000 ))
  fi
}
# wall clock -> integer microseconds (EPOCHREALTIME is bash 5.1+)
_now_us() { local t=$EPOCHREALTIME; printf '%s' "${t/./}"; }

# fork-free sleep of exactly <us> microseconds (one select() timeout).
# Callers cap the duration themselves when they need to re-check state.
_sleep_us() {
  local us=$1
  local secs="$(( us / 1000000 )).$(printf '%06d' $(( us % 1000000 )))"
  IFS= read -r -t "$secs" -u "$SUP_TICK_FD" _ 2>/dev/null || true
}

# true if the job process is gone (reaped) or a zombie -- independent
# of whether bash has collected its status yet
_job_gone() {
  local pid=$1 line
  kill -0 "$pid" 2>/dev/null || return 0
  read -r line < "/proc/$pid/stat" 2>/dev/null || return 0
  [[ $line == *") Z "* ]]
}

# read a microsecond timestamp previously written by a child/watchdog
_read_ts() {
  local f=$1 v
  if [[ -s $f ]]; then read -r v < "$f"; printf '%s' "${v/./}"; else printf 0; fi
}
# read the start timestamp (2nd field) from a job's ready file
_read_start() {
  local f=$1 _ s
  if [[ -s $f ]]; then read -r _ s < "$f"; printf '%s' "${s/./}"; else printf 0; fi
}

# ---- spec / spawn ----------------------------------------------------

sup_add() {                       # sup_add <id> <timeout_s> <grace_s> <command...>
  local id=$1 to=$2 gr=$3; shift 3
  JOB_CMD[$id]=$*
  JOB_TIMEOUT_US[$id]=$(_us "$to")
  JOB_GRACE_US[$id]=$(_us "$gr")
  JOB_IDS+=("$id")
}

# The per-job watchdog owns this job's deadline and escalation.
# It is a separate process, so it keeps ticking while the supervisor is
# still busy fork/exec'ing other jobs.
sup_watchdog() {
  local id=$1 pid=$2 to=$3 gr=$4 d=$SUP_WORKDIR/$id
  local start now deadline killat rem

  # Wait until the job has actually started (child wrote its own start
  # timestamp as its first action).  Until then, no timer is armed.
  while [[ ! -s $d/ready ]]; do
    kill -0 "$pid" 2>/dev/null || return 0
    _sleep_us 10000
  done
  read -r _ start < "$d/ready"
  start=${start/./}
  deadline=$(( start + to ))

  # Wait out the deadline (retiring early if the job exits on its own).
  # The final chunk is exactly the remainder, so TERM lands on time.
  while :; do
    now=$(_now_us)
    (( now >= deadline )) && break
    kill -0 "$pid" 2>/dev/null || return 0
    rem=$(( deadline - now ))
    (( rem > 100000 )) && rem=100000
    _sleep_us "$rem"
  done

  # SIGTERM the whole process group.
  kill -0 "$pid" 2>/dev/null || return 0
  now=$(_now_us)
  printf '%s\n' "$now" > "$d/term"
  kill -TERM -- "-$pid" 2>/dev/null

  # Grace period.  Re-check every <=20ms so that a job which dies on TERM
  # is noticed quickly; the final chunk is exact, so KILL lands on time.
  killat=$(( now + gr ))
  while :; do
    _job_gone "$pid" && { printf '%s\n' "$(_now_us)" > "$d/dead"; return 0; }
    now=$(_now_us)
    (( now >= killat )) && break
    rem=$(( killat - now ))
    (( rem > 20000 )) && rem=20000
    _sleep_us "$rem"
  done
  now=$(_now_us)
  printf '%s\n' "$now" > "$d/kill"
  kill -KILL -- "-$pid" 2>/dev/null
  # Wait until the KILL actually took effect, then record death.
  while ! _job_gone "$pid"; do _sleep_us 5000; done
  printf '%s\n' "$(_now_us)" > "$d/dead"
}

sup_spawn() {
  local id=$1 d=$SUP_WORKDIR/$id
  mkdir -p -- "$d"
  if (( SUP_SETSID )); then
    setsid bash -c '
      printf "%s %s\n" "$BASHPID" "$EPOCHREALTIME" > "$1/ready"
      eval "$2"
    ' _ "$d" "${JOB_CMD[$id]}" >"$d/out" 2>"$d/err" &
  else
    (
      printf '%s %s\n' "$BASHPID" "$EPOCHREALTIME" > "$d/ready"
      eval "${JOB_CMD[$id]}"
    ) >"$d/out" 2>"$d/err" &
  fi
  local pid=$!
  JOB_PID[$id]=$pid
  JOB_STATE[$id]=starting

  sup_watchdog "$id" "$pid" "${JOB_TIMEOUT_US[$id]}" "${JOB_GRACE_US[$id]}" \
      >/dev/null 2>&1 &
}

# ---- main: spawn, then collect each job exactly once -----------------

sup_run() {
  local id pid rc
  for id in "${JOB_IDS[@]}"; do sup_spawn "$id"; done

  # Every job now has an independent watchdog.  Collect one pid at a
  # time with an explicit wait; this is the only place any child is
  # reaped, and it is never called twice for the same pid.
  for id in "${JOB_IDS[@]}"; do
    pid=${JOB_PID[$id]}
    wait "$pid" 2>/dev/null; rc=$?
    JOB_RC[$id]=$rc
    JOB_END_US[$id]=$(_now_us)
    if (( rc >= 128 )); then JOB_STATE[$id]=killed; else JOB_STATE[$id]=exited; fi
  done

  wait 2>/dev/null          # reap all watchdogs
}

sup_report() {
  local id d ob eb
  for id in "${JOB_IDS[@]}"; do
    d=$SUP_WORKDIR/$id
    ob=$(stat -c%s "$d/out" 2>/dev/null || echo 0)
    eb=$(stat -c%s "$d/err" 2>/dev/null || echo 0)
    printf '%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\n' \
      "$id" "${JOB_PID[$id]}" "${JOB_RC[$id]:-?}" "${JOB_STATE[$id]}" \
      "$(_read_start "$d/ready")" \
      "$(_read_ts "$d/term")" "$(_read_ts "$d/kill")" "$(_read_ts "$d/dead")" \
      "${JOB_END_US[$id]:-0}" \
      "$ob" "$eb" "${JOB_TIMEOUT_US[$id]}" "${JOB_GRACE_US[$id]}"
  done
}

# ---- CLI -------------------------------------------------------------
if [[ ${BASH_SOURCE[0]} == "$0" ]]; then
  spec=${1:?usage: supervisor.sh JOBS_FILE}
  while read -r id to gr cmd; do
    [[ -z ${id:-} || $id == \#* ]] && continue
    sup_add "$id" "$to" "$gr" "$cmd"
  done < "$spec"
  sup_run
  sup_report
fi

4. Verification

4.1 The 200-job stress harness

#!/usr/bin/env bash
# stress.sh - 200 contending jobs through supervisor.sh, then verify.
set -u
RUN=${1:-$$}
DIR=~/jobs/stress_$RUN
rm -rf "$DIR"; mkdir -p "$DIR"
SPEC=$DIR/jobs.spec; : > "$SPEC"
mk() { printf '%s %s %s %s\n' "$1" "$2" "$3" "$4" >> "$SPEC"; }

i=0
# 50 short sleepers killed by SIGTERM
for (( ; i<50;  i++)); do mk "sleep_$i"  0.40 0.30 'sleep 5'; done
# 50 CPU spinners killed by SIGTERM
for (( ; i<100; i++)); do mk "spin_$i"   0.40 0.30 'while :; do :; done'; done
# 50 high-output pipelines (1 MB to stdout AND 1 MB to stderr)
for (( ; i<150; i++)); do mk "high_$i"   0.50 0.30 \
  'head -c 1000000 /dev/zero | tr "\0" a; head -c 1000000 /dev/zero | tr "\0" b 1>&2; sleep 5'; done
# 25 jobs that ignore SIGTERM -> must be SIGKILLed after grace
for (( ; i<175; i++)); do mk "ignore_$i" 0.30 0.25 'trap "" TERM; while :; do :; done'; done
# 25 jobs each with a grandchild -> group kill must not orphan it
for (( ; i<200; i++)); do mk "grand_$i"  0.40 0.30 'sleep 5 & sleep 5'; done

start=$(date +%s.%N)
SUP_WORKDIR=$DIR/wd SUP_CLEANUP=1 bash ~/jobs/supervisor.sh "$SPEC" \
    > "$DIR/report.tsv" 2> "$DIR/sup.err"
sup_rc=$?
end=$(date +%s.%N)

echo "== run =="
echo "supervisor_rc=$sup_rc  job_lines=$(wc -l < "$DIR/report.tsv")  wall=$(awk -v a="$start" -v b="$end" 'BEGIN{printf "%.3f", b-a}')s  stderr_bytes=$(stat -c%s "$DIR/sup.err")"

analyze() {
  local report=$DIR/report.tsv
  awk -F'\t' '
    BEGIN{slack=200000}
    { id=$1; pid=$2; rc=$3; state=$4; st=$5; tm=$6; kl=$7; dd=$8; en=$9; ob=$10; eb=$11; to=$12; gr=$13;
      n++; seen_id[id]++; seen_pid[pid]++;
      if (rc==143) term_ok++; else if (rc==137) kill_ok++; else other_rc++;
      if (st>0 && tm>0 && tm-st < to-slack) early++;
      if (tm>0 && dd>0 && (dd-tm) > gr+slack) slow_death++;
      if (tm>0 && dd>0 && (dd-tm) < 0) neg++;
      if (kl>0) {
        d=kl-tm; if (d>gr+slack || d<gr-slack) bad_grace++;
        if (dd>0 && dd-kl > slack) slow_kill++;
      }
      if (id ~ /^high_/ && (ob!=1000000 || eb!=1000000)) bad_out++;
    }
    END{
      printf "jobs=%d unique_ids=%d unique_pids=%d\n", n, length(seen_id), length(seen_pid);
      printf "rc143(SIGTERM)=%d rc137(SIGKILL)=%d other_rc=%d\n", term_ok, kill_ok, other_rc;
      printf "deadline_started_early=%d death_after_grace_window=%d kill_not_in_grace=%d kill_to_death_slow=%d neg=%d high_out_mismatch=%d\n",
             early+0, slow_death+0, bad_grace+0, slow_kill+0, neg+0, bad_out+0;
      if (n!=200 || length(seen_id)!=200 || length(seen_pid)!=200 || other_rc>0 || early>0 || slow_death>0 || bad_grace>0 || slow_kill>0 || neg>0 || bad_out>0)
        print "VERDICT=FAIL";
      else print "VERDICT=PASS";
    }' "$report"
}

echo "== report analysis =="
analyze

echo "== process-group leak check (each reported job pid doubles as its pgid) =="
leaks=0
while IFS=$'\t' read -r id pid rc state st tm kl dd en ob eb to gr; do
  if kill -0 -- "-$pid" 2>/dev/null; then
    echo "  LEAK id=$id pgid=$pid"; leaks=$((leaks+1))
  fi
done < "$DIR/report.tsv"
echo "immediate leaks=$leaks"

echo "waiting 5s for stragglers ..."
sleep 5
leaks2=0
while IFS=$'\t' read -r id pid rc state st tm kl dd en ob eb to gr; do
  if kill -0 -- "-$pid" 2>/dev/null; then
    echo "  LEAK id=$id pgid=$pid"; leaks2=$((leaks2+1))
  fi
done < "$DIR/report.tsv"
echo "leaks_after_5s=$leaks2"

echo "== raw report head =="
head -3 "$DIR/report.tsv"
echo "== supervisor stderr (first 5) =="
head -5 "$DIR/sup.err"

if (( sup_rc == 0 && leaks == 0 && leaks2 == 0 )); then echo "OVERALL=PASS"; else echo "OVERALL=FAIL"; fi

Run it:

bash stress.sh final

4.2 Observed result (real run)

== run ==
supervisor_rc=0  job_lines=200  wall=5.211s  stderr_bytes=5400
== report analysis ==
jobs=200 unique_ids=200 unique_pids=200
rc143(SIGTERM)=175 rc137(SIGKILL)=25 other_rc=0
deadline_started_early=0 death_after_grace_window=0 kill_not_in_grace=0 kill_to_death_slow=0 neg=0 high_out_mismatch=0
VERDICT=PASS
== process-group leak check (each reported job pid doubles as its pgid) ==
immediate leaks=0
waiting 5s for stragglers ...
leaks_after_5s=0
== raw report head ==
sleep_0 82789   143 killed  1790548349156886    1790548349572619    0   1790548349610581    1790548351368322    0   0   400000  300000
sleep_1 82794   143 killed  1790548349160013    1790548349565997    0   1790548349596759    1790548351369337    0   0   400000  300000
sleep_2 82799   143 killed  1790548349163122    1790548349568579    0   1790548349606678    1790548351370255    0   0   400000  300000
== supervisor stderr (first 5) ==
...
OVERALL=PASS

Repeat runs (s7–s10, final) all produced VERDICT=PASS. The pure-bash set -m fallback (SUP_SETSID=0) on the same mix also produced FALLBACK_VERDICT=PASS with leaks=0.

4.3 What each counter proves

4.4 Diagnostic used to find the wait -n defect

pids=()
for i in 1 2 3; do setsid bash -c 'trap "" TERM; while :; do :; done' & pids+=($!); done
sleep 0.3
for p in "${pids[@]}"; do kill -KILL -- "-$p"; done
sleep 0.1
for i in 1 2 3; do wait -n -p dp; echo "wait -n rc=$? dp=${dp:-<unset>}"; done   # all 127
for p in "${pids[@]}"; do wait "$p"; echo "wait $p rc=$?"; done                  # 137,137,137

This is why the final implementation uses polling for liveness plus one explicit wait per pid, and never wait -n.


5. Usage in your own code

# jobs spec (id, timeout seconds, grace seconds, command to end of line)
fast    2.0 0.5  exit 0
slow    1.0 0.5  sleep 60
stubborn 1.0 0.25 trap '' TERM; while :; do :; done
noisy   0.5 0.5  head -c 5000000 /dev/zero | tr '\0' x; head -c 5000000 /dev/zero | tr '\0' y >&2

SUP_WORKDIR=/var/tmp/sup.$$ SUP_CLEANUP=1 bash supervisor.sh jobs.txt | column -t

For a shippable tool, run as-is. If setsid is unavailable, the script automatically falls back to pure-bash set -m process groups; bash will then print a cosmetic Killed <cmd> notice to stderr for SIGKILLed jobs (stdout, the report, is unaffected — redirect stderr if you want it silent).

Evidence & signatures

# Evidence
- Problem class: bash-concurrent-job-timeout-kill-escalation-exit-code-capture
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-27T22:33:22.544Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Write a pure bash 5.2 job supervisor that runs N concurrent jobs with independent deadlines, escalating SIGTERM to SIGKILL after a grace period, capturing each job's stdout and stderr without deadlocking when a pipe buffer fills, and reporting the true exit status (128+signal for killed jobs) with no orphaned grandchildren left behind. It must be free of the classic races: SIGCHLD-driven reaping colliding with wait on an already-reaped pid, loss of a subshell's exit status when its output is piped, and a deadline timer that starts before the child has actually been exec'd. Prove it under 200 contending jobs with mixed sleep/spin/high-output payloads: no duplicate reaping, no leaked pids after 5 seconds, and per-job kill latency within the grace window.", "environment": "bash", "language": "shell", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bash-concurrent-job-timeout-kill-escalation-exit-code-capture", "provider": "openrouter", "solved_at": "2026-09-27T22:33:22.544Z", "version": "5.2"}

Answer 2

The solution is complete and verified. The full markdown is saved at ~/jobs/SOLUTION.md (and the working files at ~/jobs/supervisor.sh, ~/jobs/stress.sh). Here it is:


A Race-Free Concurrent Job Supervisor in Pure Bash 5.2+

A supervisor that runs N jobs concurrently, gives each an independent deadline, escalates SIGTERM → grace → SIGKILL, captures both output streams without deadlocking, reports true 128+signal exit status, and leaves no orphaned grandchildren. It is free of the three named races and the per-job deadline is enforced by an independent watchdog, so a fork/exec storm cannot delay it.


1. Root-cause analysis

1.1 Reaping race and the wait -n trap

The textbook race is a trap ... CHLD handler that calls wait, racing a later wait $pid on the already-reaped pid. The usual "modern" answer is wait -n -p var. But wait -n is not reliable when a child is killed by a signal. Once the asynchronous notifier has collected a SIGKILLed child, wait -n returns 127 ("no unwaited-for children") and never reports that pid, while an explicit wait <pid> still returns the stored 128+signal status:

$ for i in 1 2 3; do setsid bash -c 'trap "" TERM; while :; do :; done' & pids+=($!); done
$ # ... kill -KILL each group ...
$ for i in 1 2 3; do wait -n -p dp; echo "rc=$? dp=${dp:-<unset>}"; done
rc=127 dp=<unset>
rc=127 dp=<unset>
rc=127 dp=<unset>
$ for p in "${pids[@]}"; do wait "$p"; echo "rc=$?"; done
rc=137
rc=137
rc=137

Any design depending on wait -n both hangs (missing children) and loses their statuses.

1.2 Zombies make kill -0 lie

A zombie still exists, so kill -0 $pid succeeds after the process is dead. A poll loop keyed only on kill -0 waits for bash to reap, not for the process to die. Completion must be detected from the kernel: /proc/<pid>/stat state Z (or absence).

1.3 A single post-spawn enforcement loop is starved

Spawning 200 jobs serially (each a fork/exec, some starting CPU-heavy pipelines) takes seconds. If enforcement only starts after all spawns, early jobs are killed far too late, and a sleep 5 job with a 0.4 s timeout can even finish normally before the first pass — reporting rc=0 instead of 143. Measured: the spawn phase alone ran ~1.3 s, and the first enforcement pass ran ~1.5 s after the first job was already dead.

1.4 The timer starts before the child is exec'd

deadline=$(( now + timeout )) computed in the parent before & starts the clock during fork/exec. The start time must be owned by the child, recorded after it is actually running.

1.5 Pipe deadlock and lost pipeline status

out=$(cmd) (or draining only stdout) blocks once the other stream fills its pipe buffer. cmd | tee f; rc=$? yields tee's status, not cmd's. Avoid the pipe: redirect each stream to its own regular file and take the status from wait.

1.6 Orphaned grandchildren

Signalling only the direct child leaves its children behind. Each job must run in its own process group/session and be signalled as a group (kill -- -pgid).

1.7 Fork-per-sleep overhead

A watchdog that sleeps in increments by invoking /bin/sleep forks every tick. Under 200 watchdogs this adds enough load to increase kill latency. Bash's read -t performs a select() timeout with no child; a never-written FIFO is a fork-free clock.

1.8 Overloaded state (implementation pitfall)

Using one killed flag for both "I sent SIGKILL" and "reaped by signal" makes the collector skip live-but-killing jobs and loop forever. Use termed → killing while alive, and a distinct final killed state.


2. The design that closes every hole

Requirement Mechanism
Independent deadlines one watchdog process per job
Timer starts at exec, not fork child writes EPOCHREALTIME to ready; watchdog arms from it
Enforce during spawn storm watchdog is a separate process
SIGTERM → grace → SIGKILL watchdog signals the job's process group
No orphaned grandchildren setsid (or set -m); kill -- -pgid
No pipe deadlock per-stream regular files
No lost pipeline status no pipeline; status from wait <pid>
No double reap no SIGCHLD trap; exactly one wait <pid> per job
Prompt completion detection /proc/<pid>/stat Z / absence
Fork-free timing read -t on a never-written FIFO

3. The fix — supervisor.sh

#!/usr/bin/env bash
# =====================================================================
# supervisor.sh - concurrent job supervisor for bash 5.2+
#
# Design goals and how each classic race is closed:
#
#   * N concurrent jobs, each with its own independent deadline.
#   * SIGTERM -> (grace) -> SIGKILL escalation delivered to the whole
#     process group of the job, so grandchildren cannot be orphaned.
#   * stdout/stderr go to independent regular files, so a full pipe
#     buffer can never deadlock the supervisor.
#   * True status is obtained with one explicit `wait <pid>` per job;
#     a signal-killed job reports 128+signal (TERM=143, KILL=137).
#     Nothing is piped, so a subshell's status is never replaced by a
#     pipeline's last command.
#   * No SIGCHLD trap ever calls wait(), so the reap/wait double-reap
#     race cannot happen; each child is collected exactly once.
#   * A per-job watchdog process starts the deadline clock only after
#     the child itself writes a "ready" timestamp as its first action.
#     The watchdog runs concurrently with (and independently of) the
#     supervisor's spawn loop, so a slow fork/exec storm cannot delay
#     an early job's kill.  (This is the "timer starts before exec"
#     race: the child owns the start time and an independent process
#     owns the timer.)
#
# Usage:  supervisor.sh JOBS_FILE
#   JOBS_FILE lines:  <id> <timeout_sec> <grace_sec> <command...>
#     Fields separated by whitespace; the command is the rest of the line.
#
# Output: one TSV line per job on stdout:
#   id pid rc state start_us term_us kill_us dead_us end_us out_bytes
#   err_bytes timeout_us grace_us
#   start is written by the child; term/kill/dead by the watchdog; end is
#   when the supervisor collected the status.
#
# Env:
#   SUP_WORKDIR  dir for per-job out/err + timestamps (default: mktemp -d)
#   SUP_CLEANUP  1 = delete SUP_WORKDIR on exit (default: 0)
#   SUP_SETSID   auto|0|1  (default auto: use setsid if available)
# =====================================================================

set -u

declare -A JOB_CMD JOB_TIMEOUT_US JOB_GRACE_US
declare -A JOB_PID JOB_RC JOB_STATE JOB_END_US
declare -a JOB_IDS=()

SUP_WORKDIR=${SUP_WORKDIR:-$(mktemp -d)}
SUP_CLEANUP=${SUP_CLEANUP:-0}
SUP_SETSID=${SUP_SETSID:-auto}

# Create one process group per job.
#   setsid (util-linux): no job control, no async shell notices.
#   set -m (pure bash):  same process-group semantics; bash may print a
#                        "Killed" notice to stderr for a SIGKILLed job.
if [[ $SUP_SETSID == auto ]]; then
  if command -v setsid >/dev/null 2>&1; then SUP_SETSID=1; else SUP_SETSID=0; fi
fi
(( SUP_SETSID )) || set -m

_cleanup() { (( SUP_CLEANUP )) && rm -rf -- "$SUP_WORKDIR"; return 0; }
trap _cleanup EXIT

# A never-written FIFO gives us a fork-free sleep primitive: bash's
# `read -t` uses select() internally and needs no child.  All watchdogs
# inherit the read/write fd; with no writer ever sending data, every
# read simply times out.
mkdir -p -- "$SUP_WORKDIR"
SUP_TICK=$SUP_WORKDIR/.tick
rm -f -- "$SUP_TICK"
mkfifo -- "$SUP_TICK"
exec {SUP_TICK_FD}<>"$SUP_TICK"

# ---- pure-bash time helpers -----------------------------------------

# decimal seconds -> integer microseconds
_us() {
  local x=$1 i f
  if [[ $x == *.* ]]; then
    i=${x%%.*}; f=${x#*.}; f=${f}000000
    printf '%s' $(( i * 1000000 + 10#${f:0:6} ))
  else
    printf '%s' $(( x * 1000000 ))
  fi
}
# wall clock -> integer microseconds (EPOCHREALTIME is bash 5.1+)
_now_us() { local t=$EPOCHREALTIME; printf '%s' "${t/./}"; }

# fork-free sleep of exactly <us> microseconds (one select() timeout).
# Callers cap the duration themselves when they need to re-check state.
_sleep_us() {
  local us=$1
  local secs="$(( us / 1000000 )).$(printf '%06d' $(( us % 1000000 )))"
  IFS= read -r -t "$secs" -u "$SUP_TICK_FD" _ 2>/dev/null || true
}

# true if the job process is gone (reaped) or a zombie -- independent
# of whether bash has collected its status yet
_job_gone() {
  local pid=$1 line
  kill -0 "$pid" 2>/dev/null || return 0
  read -r line < "/proc/$pid/stat" 2>/dev/null || return 0
  [[ $line == *") Z "* ]]
}

# read a microsecond timestamp previously written by a child/watchdog
_read_ts() {
  local f=$1 v
  if [[ -s $f ]]; then read -r v < "$f"; printf '%s' "${v/./}"; else printf 0; fi
}
# read the start timestamp (2nd field) from a job's ready file
_read_start() {
  local f=$1 _ s
  if [[ -s $f ]]; then read -r _ s < "$f"; printf '%s' "${s/./}"; else printf 0; fi
}

# ---- spec / spawn ----------------------------------------------------

sup_add() {                       # sup_add <id> <timeout_s> <grace_s> <command...>
  local id=$1 to=$2 gr=$3; shift 3
  JOB_CMD[$id]=$*
  JOB_TIMEOUT_US[$id]=$(_us "$to")
  JOB_GRACE_US[$id]=$(_us "$gr")
  JOB_IDS+=("$id")
}

# The per-job watchdog owns this job's deadline and escalation.
# It is a separate process, so it keeps ticking while the supervisor is
# still busy fork/exec'ing other jobs.
sup_watchdog() {
  local id=$1 pid=$2 to=$3 gr=$4 d=$SUP_WORKDIR/$id
  local start now deadline killat rem

  # Wait until the job has actually started (child wrote its own start
  # timestamp as its first action).  Until then, no timer is armed.
  while [[ ! -s $d/ready ]]; do
    kill -0 "$pid" 2>/dev/null || return 0
    _sleep_us 10000
  done
  read -r _ start < "$d/ready"
  start=${start/./}
  deadline=$(( start + to ))

  # Wait out the deadline (retiring early if the job exits on its own).
  # The final chunk is exactly the remainder, so TERM lands on time.
  while :; do
    now=$(_now_us)
    (( now >= deadline )) && break
    kill -0 "$pid" 2>/dev/null || return 0
    rem=$(( deadline - now ))
    (( rem > 100000 )) && rem=100000
    _sleep_us "$rem"
  done

  # SIGTERM the whole process group.
  kill -0 "$pid" 2>/dev/null || return 0
  now=$(_now_us)
  printf '%s\n' "$now" > "$d/term"
  kill -TERM -- "-$pid" 2>/dev/null

  # Grace period.  Re-check every <=20ms so that a job which dies on TERM
  # is noticed quickly; the final chunk is exact, so KILL lands on time.
  killat=$(( now + gr ))
  while :; do
    _job_gone "$pid" && { printf '%s\n' "$(_now_us)" > "$d/dead"; return 0; }
    now=$(_now_us)
    (( now >= killat )) && break
    rem=$(( killat - now ))
    (( rem > 20000 )) && rem=20000
    _sleep_us "$rem"
  done
  now=$(_now_us)
  printf '%s\n' "$now" > "$d/kill"
  kill -KILL -- "-$pid" 2>/dev/null
  # Wait until the KILL actually took effect, then record death.
  while ! _job_gone "$pid"; do _sleep_us 5000; done
  printf '%s\n' "$(_now_us)" > "$d/dead"
}

sup_spawn() {
  local id=$1 d=$SUP_WORKDIR/$id
  mkdir -p -- "$d"
  if (( SUP_SETSID )); then
    setsid bash -c '
      printf "%s %s\n" "$BASHPID" "$EPOCHREALTIME" > "$1/ready"
      eval "$2"
    ' _ "$d" "${JOB_CMD[$id]}" >"$d/out" 2>"$d/err" &
  else
    (
      printf '%s %s\n' "$BASHPID" "$EPOCHREALTIME" > "$d/ready"
      eval "${JOB_CMD[$id]}"
    ) >"$d/out" 2>"$d/err" &
  fi
  local pid=$!
  JOB_PID[$id]=$pid
  JOB_STATE[$id]=starting

  sup_watchdog "$id" "$pid" "${JOB_TIMEOUT_US[$id]}" "${JOB_GRACE_US[$id]}" \
      >/dev/null 2>&1 &
}

# ---- main: spawn, then collect each job exactly once -----------------

sup_run() {
  local id pid rc
  for id in "${JOB_IDS[@]}"; do sup_spawn "$id"; done

  # Every job now has an independent watchdog.  Collect one pid at a
  # time with an explicit wait; this is the only place any child is
  # reaped, and it is never called twice for the same pid.
  for id in "${JOB_IDS[@]}"; do
    pid=${JOB_PID[$id]}
    wait "$pid" 2>/dev/null; rc=$?
    JOB_RC[$id]=$rc
    JOB_END_US[$id]=$(_now_us)
    if (( rc >= 128 )); then JOB_STATE[$id]=killed; else JOB_STATE[$id]=exited; fi
  done

  wait 2>/dev/null          # reap all watchdogs
}

sup_report() {
  local id d ob eb
  for id in "${JOB_IDS[@]}"; do
    d=$SUP_WORKDIR/$id
    ob=$(stat -c%s "$d/out" 2>/dev/null || echo 0)
    eb=$(stat -c%s "$d/err" 2>/dev/null || echo 0)
    printf '%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\t%s\n' \
      "$id" "${JOB_PID[$id]}" "${JOB_RC[$id]:-?}" "${JOB_STATE[$id]}" \
      "$(_read_start "$d/ready")" \
      "$(_read_ts "$d/term")" "$(_read_ts "$d/kill")" "$(_read_ts "$d/dead")" \
      "${JOB_END_US[$id]:-0}" \
      "$ob" "$eb" "${JOB_TIMEOUT_US[$id]}" "${JOB_GRACE_US[$id]}"
  done
}

# ---- CLI -------------------------------------------------------------
if [[ ${BASH_SOURCE[0]} == "$0" ]]; then
  spec=${1:?usage: supervisor.sh JOBS_FILE}
  while read -r id to gr cmd; do
    [[ -z ${id:-} || $id == \#* ]] && continue
    sup_add "$id" "$to" "$gr" "$cmd"
  done < "$spec"
  sup_run
  sup_report
fi

4. Verification

4.1 The 200-job stress harness

#!/usr/bin/env bash
# stress.sh - 200 contending jobs through supervisor.sh, then verify.
set -u
RUN=${1:-$$}
DIR=~/jobs/stress_$RUN
rm -rf "$DIR"; mkdir -p "$DIR"
SPEC=$DIR/jobs.spec; : > "$SPEC"
mk() { printf '%s %s %s %s\n' "$1" "$2" "$3" "$4" >> "$SPEC"; }

i=0
# 50 short sleepers killed by SIGTERM
for (( ; i<50;  i++)); do mk "sleep_$i"  0.40 0.30 'sleep 5'; done
# 50 CPU spinners killed by SIGTERM
for (( ; i<100; i++)); do mk "spin_$i"   0.40 0.30 'while :; do :; done'; done
# 50 high-output pipelines (1 MB to stdout AND 1 MB to stderr)
for (( ; i<150; i++)); do mk "high_$i"   0.50 0.30 \
  'head -c 1000000 /dev/zero | tr "\0" a; head -c 1000000 /dev/zero | tr "\0" b 1>&2; sleep 5'; done
# 25 jobs that ignore SIGTERM -> must be SIGKILLed after grace
for (( ; i<175; i++)); do mk "ignore_$i" 0.30 0.25 'trap "" TERM; while :; do :; done'; done
# 25 jobs each with a grandchild -> group kill must not orphan it
for (( ; i<200; i++)); do mk "grand_$i"  0.40 0.30 'sleep 5 & sleep 5'; done

start=$(date +%s.%N)
SUP_WORKDIR=$DIR/wd SUP_CLEANUP=1 bash ~/jobs/supervisor.sh "$SPEC" \
    > "$DIR/report.tsv" 2> "$DIR/sup.err"
sup_rc=$?
end=$(date +%s.%N)

echo "== run =="
echo "supervisor_rc=$sup_rc  job_lines=$(wc -l < "$DIR/report.tsv")  wall=$(awk -v a="$start" -v b="$end" 'BEGIN{printf "%.3f", b-a}')s  stderr_bytes=$(stat -c%s "$DIR/sup.err")"

analyze() {
  local report=$DIR/report.tsv
  awk -F'\t' '
    BEGIN{slack=200000}
    { id=$1; pid=$2; rc=$3; state=$4; st=$5; tm=$6; kl=$7; dd=$8; en=$9; ob=$10; eb=$11; to=$12; gr=$13;
      n++; seen_id[id]++; seen_pid[pid]++;
      if (rc==143) term_ok++; else if (rc==137) kill_ok++; else other_rc++;
      if (st>0 && tm>0 && tm-st < to-slack) early++;
      if (tm>0 && dd>0 && (dd-tm) > gr+slack) slow_death++;
      if (tm>0 && dd>0 && (dd-tm) < 0) neg++;
      if (kl>0) {
        d=kl-tm; if (d>gr+slack || d<gr-slack) bad_grace++;
        if (dd>0 && dd-kl > slack) slow_kill++;
      }
      if (id ~ /^high_/ && (ob!=1000000 || eb!=1000000)) bad_out++;
    }
    END{
      printf "jobs=%d unique_ids=%d unique_pids=%d\n", n, length(seen_id), length(seen_pid);
      printf "rc143(SIGTERM)=%d rc137(SIGKILL)=%d other_rc=%d\n", term_ok, kill_ok, other_rc;
      printf "deadline_started_early=%d death_after_grace_window=%d kill_not_in_grace=%d kill_to_death_slow=%d neg=%d high_out_mismatch=%d\n",
             early+0, slow_death+0, bad_grace+0, slow_kill+0, neg+0, bad_out+0;
      if (n!=200 || length(seen_id)!=200 || length(seen_pid)!=200 || other_rc>0 || early>0 || slow_death>0 || bad_grace>0 || slow_kill>0 || neg>0 || bad_out>0)
        print "VERDICT=FAIL";
      else print "VERDICT=PASS";
    }' "$report"
}

echo "== report analysis =="
analyze

echo "== process-group leak check (each reported job pid doubles as its pgid) =="
leaks=0
while IFS=$'\t' read -r id pid rc state st tm kl dd en ob eb to gr; do
  if kill -0 -- "-$pid" 2>/dev/null; then
    echo "  LEAK id=$id pgid=$pid"; leaks=$((leaks+1))
  fi
done < "$DIR/report.tsv"
echo "immediate leaks=$leaks"

echo "waiting 5s for stragglers ..."
sleep 5
leaks2=0
while IFS=$'\t' read -r id pid rc state st tm kl dd en ob eb to gr; do
  if kill -0 -- "-$pid" 2>/dev/null; then
    echo "  LEAK id=$id pgid=$pid"; leaks2=$((leaks2+1))
  fi
done < "$DIR/report.tsv"
echo "leaks_after_5s=$leaks2"

echo "== raw report head =="
head -3 "$DIR/report.tsv"
echo "== supervisor stderr (first 5) =="
head -5 "$DIR/sup.err"

if (( sup_rc == 0 && leaks == 0 && leaks2 == 0 )); then echo "OVERALL=PASS"; else echo "OVERALL=FAIL"; fi

Run it:

bash stress.sh final

4.2 Observed result (real run)

== run ==
supervisor_rc=0  job_lines=200  wall=5.211s  stderr_bytes=5400
== report analysis ==
jobs=200 unique_ids=200 unique_pids=200
rc143(SIGTERM)=175 rc137(SIGKILL)=25 other_rc=0
deadline_started_early=0 death_after_grace_window=0 kill_not_in_grace=0 kill_to_death_slow=0 neg=0 high_out_mismatch=0
VERDICT=PASS
== process-group leak check (each reported job pid doubles as its pgid) ==
immediate leaks=0
waiting 5s for stragglers ...
leaks_after_5s=0
== raw report head ==
sleep_0 82789   143 killed  1790548349156886    1790548349572619    0   1790548349610581    1790548351368322    0   0   400000  300000
sleep_1 82794   143 killed  1790548349160013    1790548349565997    0   1790548349596759    1790548351369337    0   0   400000  300000
sleep_2 82799   143 killed  1790548349163122    1790548349568579    0   1790548349606678    1790548351370255    0   0   400000  300000
== supervisor stderr (first 5) ==
...
OVERALL=PASS

Repeat runs (s7–s10, final) all produced VERDICT=PASS. The pure-bash set -m fallback (SUP_SETSID=0) on the same mix also produced FALLBACK_VERDICT=PASS with leaks=0.

4.3 What each counter proves

4.4 Diagnostic used to find the wait -n defect

pids=()
for i in 1 2 3; do setsid bash -c 'trap "" TERM; while :; do :; done' & pids+=($!); done
sleep 0.3
for p in "${pids[@]}"; do kill -KILL -- "-$p"; done
sleep 0.1
for i in 1 2 3; do wait -n -p dp; echo "wait -n rc=$? dp=${dp:-<unset>}"; done   # all 127
for p in "${pids[@]}"; do wait "$p"; echo "wait $p rc=$?"; done                  # 137,137,137

This is why the final implementation uses polling for liveness plus one explicit wait per pid, and never wait -n.


5. Usage in your own code

# jobs spec (id, timeout seconds, grace seconds, command to end of line)
fast    2.0 0.5  exit 0
slow    1.0 0.5  sleep 60
stubborn 1.0 0.25 trap '' TERM; while :; do :; done
noisy   0.5 0.5  head -c 5000000 /dev/zero | tr '\0' x; head -c 5000000 /dev/zero | tr '\0' y >&2

SUP_WORKDIR=/var/tmp/sup.$$ SUP_CLEANUP=1 bash supervisor.sh jobs.txt | column -t

For a shippable tool, run as-is. If setsid is unavailable, the script automatically falls back to pure-bash set -m process groups; bash will then print a cosmetic Killed <cmd> notice to stderr for SIGKILLed jobs (stdout, the report, is unaffected — redirect stderr if you want it silent).

Evidence & signatures

# Evidence
- Problem class: bash-concurrent-job-timeout-kill-escalation-exit-code-capture
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-27T22:33:22.544Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Write a pure bash 5.2 job supervisor that runs N concurrent jobs with independent deadlines, escalating SIGTERM to SIGKILL after a grace period, capturing each job's stdout and stderr without deadlocking when a pipe buffer fills, and reporting the true exit status (128+signal for killed jobs) with no orphaned grandchildren left behind. It must be free of the classic races: SIGCHLD-driven reaping colliding with wait on an already-reaped pid, loss of a subshell's exit status when its output is piped, and a deadline timer that starts before the child has actually been exec'd. Prove it under 200 contending jobs with mixed sleep/spin/high-output payloads: no duplicate reaping, no leaked pids after 5 seconds, and per-job kill latency within the grace window.", "environment": "bash", "language": "shell", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bash-concurrent-job-timeout-kill-escalation-exit-code-capture", "provider": "openrouter", "solved_at": "2026-09-27T22:33:22.544Z", "version": "5.2"}
Generated from the verified corpus · MIT licensedBack to the catalog