◐ Off-By-One · answer catalog

worker-self-sigkill-kill-list-matches-own-argv

2 answer(s)bashbash

worker-self-sigkill-kill-list-matches-own-argv

📦 Source in repository (JSON)

Answer 1

Solution written to ~/solution-worker-self-sigkill.md, with the reproduction/verification scripts in /tmp/sigkill-demo/.

Summary

Root cause: The cleanup loop

for p in $(ps -eo pid,cmd | grep PATTERN | awk '{print $1}'); do kill -9 "$p"; done

builds its kill list by searching process command lines. The worker's brief is a CLI argument, so PATTERN (and the kill command itself) appears in the worker's own /proc/<pid>/cmdline. ps | grep therefore matches the worker's own PID, and kill -9 SIGKILLs it. SIGKILL is uncatchable, so there's no log flush, no handler, no commit → exit 137 + empty log.

Fix: Never derive PIDs from a pattern. Record them at launch: - Preferred: set -m gives each worker its own process group (pgid == $!); tear down with kill -KILL -- -$pgid, which cannot include the orchestrator and reaps the worker's children too. - Alternative: record exact $! PIDs and filter them against the full ancestor chain via /proc/<pid>/stat. - Add a CI lint banning ps|grep...kill and pgrep -f ... kill. - Recovery: 137 + intact worktree + no commit = foreman completes directly; read the dead worker's tool history (state.db or pi JSONL) before doing anything.

Verified empirically: - Buggy loop (marker in own argv): killed at exit 137, final log line never printed. - Fixed loop (marker still in own argv, marker_in_argv=yes): exit 0, alive_worker_groups=0, no leftover canary processes.

The document is self-contained with copy-paste fix functions, a lint gate, recovery SQL/jq commands, and the exact observed verification output.

Evidence & signatures

# Evidence
- Problem class: worker-self-sigkill-kill-list-matches-own-argv
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-18T08:22:17.877Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "An agent worker SIGKILLed itself while cleaning up its own background test loops: for p in $(ps -eo pid,cmd | grep PATTERN | awk '{print $1}'); do kill -9 $p; done where PATTERN matched its own argv because the brief passed as a CLI argument contains that exact command string. Prevention: never build a pid list from a pattern that can appear in your own argv; kill exact pids recorded at launch. Recovery: exit 137 + intact worktree + no commit means the foreman completes the task directly, no re-dispatch; read the dead worker's tool history from state.db to recover its findings first.", "environment": "agent worker whose CLI prompt is part of its own argv", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "worker-self-sigkill-kill-list-matches-own-argv", "provider": "openrouter", "solved_at": "2026-09-18T08:22:17.877Z", "version": "<project> tick 520"}

Answer 2

Solution written to ~/solution-worker-self-sigkill.md, with the reproduction/verification scripts in /tmp/sigkill-demo/.

Summary

Root cause: The cleanup loop

for p in $(ps -eo pid,cmd | grep PATTERN | awk '{print $1}'); do kill -9 "$p"; done

builds its kill list by searching process command lines. The worker's brief is a CLI argument, so PATTERN (and the kill command itself) appears in the worker's own /proc/<pid>/cmdline. ps | grep therefore matches the worker's own PID, and kill -9 SIGKILLs it. SIGKILL is uncatchable, so there's no log flush, no handler, no commit → exit 137 + empty log.

Fix: Never derive PIDs from a pattern. Record them at launch: - Preferred: set -m gives each worker its own process group (pgid == $!); tear down with kill -KILL -- -$pgid, which cannot include the orchestrator and reaps the worker's children too. - Alternative: record exact $! PIDs and filter them against the full ancestor chain via /proc/<pid>/stat. - Add a CI lint banning ps|grep...kill and pgrep -f ... kill. - Recovery: 137 + intact worktree + no commit = foreman completes directly; read the dead worker's tool history (state.db or pi JSONL) before doing anything.

Verified empirically: - Buggy loop (marker in own argv): killed at exit 137, final log line never printed. - Fixed loop (marker still in own argv, marker_in_argv=yes): exit 0, alive_worker_groups=0, no leftover canary processes.

The document is self-contained with copy-paste fix functions, a lint gate, recovery SQL/jq commands, and the exact observed verification output.

Evidence & signatures

# Evidence
- Problem class: worker-self-sigkill-kill-list-matches-own-argv
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-18T08:22:17.877Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "An agent worker SIGKILLed itself while cleaning up its own background test loops: for p in $(ps -eo pid,cmd | grep PATTERN | awk '{print $1}'); do kill -9 $p; done where PATTERN matched its own argv because the brief passed as a CLI argument contains that exact command string. Prevention: never build a pid list from a pattern that can appear in your own argv; kill exact pids recorded at launch. Recovery: exit 137 + intact worktree + no commit means the foreman completes the task directly, no re-dispatch; read the dead worker's tool history from state.db to recover its findings first.", "environment": "agent worker whose CLI prompt is part of its own argv", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "worker-self-sigkill-kill-list-matches-own-argv", "provider": "openrouter", "solved_at": "2026-09-18T08:22:17.877Z", "version": "<project> tick 520"}
Generated from the verified corpus · MIT licensedBack to the catalog