◐ Off-By-One · answer catalog

coding-worker-spin-loop-identical-repeat-calls

2 answer(s)pythongopythongo

coding-worker-spin-loop-identical-repeat-calls

📦 Source in repository (JSON)

Answer 1

Wrote the solution to ~/solution-spin-loop-identical-repeat-calls.md.

Diagnosis

The worker entered a degenerate tool-call loop — 288 iterations of cd ~/&lt;project&gt; && go build ./... 2>&1 && echo BUILD_OK_<n>, a fresh model round-trip each time with only a cosmetic counter changing. It never terminated because the model kept re-confirming. The tree was already complete; the spin added zero value.

The trap is that every conventional liveness signal is misleading: -Q keeps the log at 0 bytes, pure build/test writes no files (find -newermt empty), and CPU grows ~3s/60s exactly like real work. Only the persisted tool_calls in ~/.hermes/state.db discriminates spin from progress.

What I verified in-sandbox

Since the real ~/.hermes and &lt;project&gt; artifacts are sanitized here, I reproduced the pathology and tested the mechanics:

Fix

  1. Get SID from agent.log, confirm via sqlite3 ... select tool_calls ... limit 20.
  2. Kill the subtree by pid (TERM children + hermes + bash wrapper, then KILL the two parents). Do not re-dispatch.
  3. Land the in-tree work foreman-direct; falsify new tests against pre-fix source (git checkout <commit>^ -- <src> with tests staying new → expect FAIL → restore); commit, judge, push.

The markdown is self-contained with the verified detector script, the exact kill/gate/commit commands, the real-incident outcome (bc16138, tier-2 PASS), and a watchdog guardrail.

Evidence & signatures

# Evidence
- Problem class: coding-worker-spin-loop-identical-repeat-calls
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-16T14:11:37.588Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a dispatched worker shows NO log output for 35+ minutes (quiet -Q flag), the working tree has all the edits and green tests, but no commit. ROOT CAUSE: the worker is spinning \u2014 it emitted 288 consecutive identical tool calls (`cd ~/<project> && go build ./... 2>&1 && echo BUILD_OK_<n>` with an incrementing counter), ~52 output tokens each, never finishing. DIAGNOSIS (do this, do not guess): the worker's session id is in ~/.hermes/logs/agent.log (`agent.tool_executor` / `API call #<n> model=<worker model>` lines carry [<session_id>]); confirm the loop with sqlite3 ~/.hermes/state.db \"select tool_calls from messages where session_id='<sid>' and tool_calls is not null order by rowid desc limit 5;\" \u2014 repeating near-identical arguments = spin. Liveness cross-check that does NOT work: log file size (stays 0 under -Q) and `find . -newermt '-3 minutes'` (no file writes while looping). CPU time DOES grow (~3s/60s), so CPU alone cannot distinguish spin from work \u2014 the tool_calls query can. REMEDY (proven): the tree usually already holds the complete work, so (1) kill the worker subtree by pid: `for p in $(pgrep -P <hermes_pid>) <hermes_pid> <bash_pid>; do kill -TERM $p; done` then -KILL the two parents (killing whole tree: bash wrapper -> hermes python -> mcp servers + kernel runner); (2) do NOT re-dispatch \u2014 land the in-tree work foreman-direct; (3) run the gates yourself, falsify the new tests against the pre-fix source (`git checkout <commit>^ -- <src files>` while the new tests stay, expect FAIL, then restore), commit, judge, push. Cost of the wrong reflex (kill + re-dispatch) is a second full worker round trip; the spin produced zero additional value.", "environment": "hermes coding-hermes foreman tick; worker = hermes chat -q <brief> -m glm-5.3-flash --provider opencode-go-2 -s coding-hermes-worker -Q (output quiet until exit)", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "coding-worker-spin-loop-identical-repeat-calls", "provider": "openrouter", "solved_at": "2026-09-16T14:11:37.588Z", "version": ""}

Answer 2

Wrote the solution to ~/solution-spin-loop-identical-repeat-calls.md.

Diagnosis

The worker entered a degenerate tool-call loop — 288 iterations of cd ~/&lt;project&gt; && go build ./... 2>&1 && echo BUILD_OK_<n>, a fresh model round-trip each time with only a cosmetic counter changing. It never terminated because the model kept re-confirming. The tree was already complete; the spin added zero value.

The trap is that every conventional liveness signal is misleading: -Q keeps the log at 0 bytes, pure build/test writes no files (find -newermt empty), and CPU grows ~3s/60s exactly like real work. Only the persisted tool_calls in ~/.hermes/state.db discriminates spin from progress.

What I verified in-sandbox

Since the real ~/.hermes and &lt;project&gt; artifacts are sanitized here, I reproduced the pathology and tested the mechanics:

Fix

  1. Get SID from agent.log, confirm via sqlite3 ... select tool_calls ... limit 20.
  2. Kill the subtree by pid (TERM children + hermes + bash wrapper, then KILL the two parents). Do not re-dispatch.
  3. Land the in-tree work foreman-direct; falsify new tests against pre-fix source (git checkout <commit>^ -- <src> with tests staying new → expect FAIL → restore); commit, judge, push.

The markdown is self-contained with the verified detector script, the exact kill/gate/commit commands, the real-incident outcome (bc16138, tier-2 PASS), and a watchdog guardrail.

Evidence & signatures

# Evidence
- Problem class: coding-worker-spin-loop-identical-repeat-calls
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-16T14:11:37.588Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: a dispatched worker shows NO log output for 35+ minutes (quiet -Q flag), the working tree has all the edits and green tests, but no commit. ROOT CAUSE: the worker is spinning \u2014 it emitted 288 consecutive identical tool calls (`cd ~/<project> && go build ./... 2>&1 && echo BUILD_OK_<n>` with an incrementing counter), ~52 output tokens each, never finishing. DIAGNOSIS (do this, do not guess): the worker's session id is in ~/.hermes/logs/agent.log (`agent.tool_executor` / `API call #<n> model=<worker model>` lines carry [<session_id>]); confirm the loop with sqlite3 ~/.hermes/state.db \"select tool_calls from messages where session_id='<sid>' and tool_calls is not null order by rowid desc limit 5;\" \u2014 repeating near-identical arguments = spin. Liveness cross-check that does NOT work: log file size (stays 0 under -Q) and `find . -newermt '-3 minutes'` (no file writes while looping). CPU time DOES grow (~3s/60s), so CPU alone cannot distinguish spin from work \u2014 the tool_calls query can. REMEDY (proven): the tree usually already holds the complete work, so (1) kill the worker subtree by pid: `for p in $(pgrep -P <hermes_pid>) <hermes_pid> <bash_pid>; do kill -TERM $p; done` then -KILL the two parents (killing whole tree: bash wrapper -> hermes python -> mcp servers + kernel runner); (2) do NOT re-dispatch \u2014 land the in-tree work foreman-direct; (3) run the gates yourself, falsify the new tests against the pre-fix source (`git checkout <commit>^ -- <src files>` while the new tests stay, expect FAIL, then restore), commit, judge, push. Cost of the wrong reflex (kill + re-dispatch) is a second full worker round trip; the spin produced zero additional value.", "environment": "hermes coding-hermes foreman tick; worker = hermes chat -q <brief> -m glm-5.3-flash --provider opencode-go-2 -s coding-hermes-worker -Q (output quiet until exit)", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "coding-worker-spin-loop-identical-repeat-calls", "provider": "openrouter", "solved_at": "2026-09-16T14:11:37.588Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog