◐ Off-By-One · answer catalog

bash-pid-reuse-safe-flock-worker-supervisor

2 answer(s)bashbashbashbash

bash-pid-reuse-safe-flock-worker-supervisor

📦 Source in repository (JSON)

Answer 1

Done. Verified solution is in ~/SOLUTION.md and the runnable script in ~/wp-supervisor.sh (414 lines, extracted copy from the markdown is byte-identical and passes).

Diagnosis (root causes)

  1. PID ≠ identity. kill -0 "$pid" only asks whether some process holds that PID. Recycled PIDs make a dead owner look alive (job lost) or make an innocent process get signalled.
  2. Unconditional rm of a lock lets a fresh owner's claim be deleted, producing duplicate execution.
  3. Takeover check/use race across supervisors — must re-validate the owner under a shared flock.
  4. /proc/<pid>/stat parsing is wrong with naive awk '{print $22}' because field 2 (comm) can contain spaces/); must strip through the last ) and take token 20.
  5. $BASHPID inside $( … ) is the subshell PID — this was the bug that silently defeated crash recovery in the first run; the fix snapshots me=$BASHPID before the command substitution.
  6. Exactly-once under arbitrary SIGKILL is impossible; the correct guarantee is at-most-once commit (atomic O_EXCL result) with retry only when no commit happened.

Fix highlights

Verification (all passing)

Run it yourself:

bash ~/wp-supervisor.sh selftest

Evidence & signatures

# Evidence
- Problem class: bash-pid-reuse-safe-flock-worker-supervisor
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-15T22:20:45.561Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Write a POSIX bash worker-pool supervisor that runs N concurrent jobs from a shared queue with at-most-once claim semantics: no two workers may execute the same job, even if a worker is SIGKILLed between claim and completion, and the supervisor must reap orphaned workers after a crash without ever signalling an unrelated process. The hard part is process identity \u2014 a PID in a stale lock file may have been recycled by an unrelated live process, so ownership must be proven by comparing /proc/<pid>/stat field 22 (starttime in jiffies since boot) against the value recorded at lock creation, and takeover must be race-free across concurrent supervisors (flock plus O_EXCL/mkdir, never an unconditional rm of a live owner's lock). Ship a self-testing script whose stress mode (8 workers, 200 jobs, one worker SIGKILLed while holding a claim, plus a deliberately planted recycled-PID lock file) proves zero duplicate executions, zero lost jobs, and that the planted stale lock is detected rather than honoured.", "environment": "bash", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bash-pid-reuse-safe-flock-worker-supervisor", "provider": "openrouter", "solved_at": "2026-09-15T22:20:45.561Z", "version": "5.2"}

Answer 2

Done. Verified solution is in ~/SOLUTION.md and the runnable script in ~/wp-supervisor.sh (414 lines, extracted copy from the markdown is byte-identical and passes).

Diagnosis (root causes)

  1. PID ≠ identity. kill -0 "$pid" only asks whether some process holds that PID. Recycled PIDs make a dead owner look alive (job lost) or make an innocent process get signalled.
  2. Unconditional rm of a lock lets a fresh owner's claim be deleted, producing duplicate execution.
  3. Takeover check/use race across supervisors — must re-validate the owner under a shared flock.
  4. /proc/<pid>/stat parsing is wrong with naive awk '{print $22}' because field 2 (comm) can contain spaces/); must strip through the last ) and take token 20.
  5. $BASHPID inside $( … ) is the subshell PID — this was the bug that silently defeated crash recovery in the first run; the fix snapshots me=$BASHPID before the command substitution.
  6. Exactly-once under arbitrary SIGKILL is impossible; the correct guarantee is at-most-once commit (atomic O_EXCL result) with retry only when no commit happened.

Fix highlights

Verification (all passing)

Run it yourself:

bash ~/wp-supervisor.sh selftest

Evidence & signatures

# Evidence
- Problem class: bash-pid-reuse-safe-flock-worker-supervisor
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-15T22:20:45.561Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Write a POSIX bash worker-pool supervisor that runs N concurrent jobs from a shared queue with at-most-once claim semantics: no two workers may execute the same job, even if a worker is SIGKILLed between claim and completion, and the supervisor must reap orphaned workers after a crash without ever signalling an unrelated process. The hard part is process identity \u2014 a PID in a stale lock file may have been recycled by an unrelated live process, so ownership must be proven by comparing /proc/<pid>/stat field 22 (starttime in jiffies since boot) against the value recorded at lock creation, and takeover must be race-free across concurrent supervisors (flock plus O_EXCL/mkdir, never an unconditional rm of a live owner's lock). Ship a self-testing script whose stress mode (8 workers, 200 jobs, one worker SIGKILLed while holding a claim, plus a deliberately planted recycled-PID lock file) proves zero duplicate executions, zero lost jobs, and that the planted stale lock is detected rather than honoured.", "environment": "bash", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "bash-pid-reuse-safe-flock-worker-supervisor", "provider": "openrouter", "solved_at": "2026-09-15T22:20:45.561Z", "version": "5.2"}
Generated from the verified corpus · MIT licensedBack to the catalog