hermes-chat-concurrent-spawn-state-db-contention
Solution written to ~/hermes-concurrent-spawn-state-db-contention.md.
1. Root cause — Each hermes chat worker performs a session-registration write against the shared SQLite state.db at boot. With the default busy_timeout=0, the two workers that lose the write-lock race abort immediately with SQLITE_BUSY rather than waiting. Because this happens before any task work, nothing is left behind.
2. Boot-death vs mid-task-death triage (log-free):
- launcher process gone from ps
- log contains only the storage-busy warning
- worktree git status clean → no edits → never started work
- explicit warning that --cli -Q logs stay 0 bytes until completion, so they're not a liveness signal.
3. Fix
- Immediate: a stagger_redispatch.sh script dispatching workers at strictly increasing 5/25/45 s offsets, plus a no-script one-liner variant.
- Durable: wrap dispatch in flock, set state.db to WAL with busy_timeout=30000, and preserve the scheduler's --min-interval.
4. Verification
- Authoritative ps --ppid <worker-pid> child-mapping loop over dispatch.pids (verified the field parsing — the initial greedy sed mis-read ppid= as pid=, now fixed with prefix stripping).
- Independent checks for busy warnings, worktree state, and PRAGMA journal_mode.
- Post-completion poll on -Q logs growing from empty as the final success signal.
The embedded dispatch script passed bash -n, and the verification parser was tested against a representative dispatch.pids file.
# Evidence - Problem class: hermes-chat-concurrent-spawn-state-db-contention - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-25T18:16:21.821Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "3 hermes chat workers launched in the same second: 2 died at boot with session storage busy state.db write contention. Diagnosis without logs: launcher process gone from ps + log contains only the storage-busy warning + worktree git status clean = death at boot, not mid-task (mid-task death leaves dirty files). Fix: re-dispatch staggered (5-45s offsets); verify liveness via ps --ppid <launcher-pid> child mapping, never via the -Q log which stays 0 bytes until completion in --cli -Q mode.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "hermes-chat-concurrent-spawn-state-db-contention", "provider": "openrouter", "solved_at": "2026-09-25T18:16:21.822Z", "version": ""}Solution written to ~/hermes-concurrent-spawn-state-db-contention.md.
1. Root cause — Each hermes chat worker performs a session-registration write against the shared SQLite state.db at boot. With the default busy_timeout=0, the two workers that lose the write-lock race abort immediately with SQLITE_BUSY rather than waiting. Because this happens before any task work, nothing is left behind.
2. Boot-death vs mid-task-death triage (log-free):
- launcher process gone from ps
- log contains only the storage-busy warning
- worktree git status clean → no edits → never started work
- explicit warning that --cli -Q logs stay 0 bytes until completion, so they're not a liveness signal.
3. Fix
- Immediate: a stagger_redispatch.sh script dispatching workers at strictly increasing 5/25/45 s offsets, plus a no-script one-liner variant.
- Durable: wrap dispatch in flock, set state.db to WAL with busy_timeout=30000, and preserve the scheduler's --min-interval.
4. Verification
- Authoritative ps --ppid <worker-pid> child-mapping loop over dispatch.pids (verified the field parsing — the initial greedy sed mis-read ppid= as pid=, now fixed with prefix stripping).
- Independent checks for busy warnings, worktree state, and PRAGMA journal_mode.
- Post-completion poll on -Q logs growing from empty as the final success signal.
The embedded dispatch script passed bash -n, and the verification parser was tested against a representative dispatch.pids file.
# Evidence - Problem class: hermes-chat-concurrent-spawn-state-db-contention - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-25T18:16:21.821Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "3 hermes chat workers launched in the same second: 2 died at boot with session storage busy state.db write contention. Diagnosis without logs: launcher process gone from ps + log contains only the storage-busy warning + worktree git status clean = death at boot, not mid-task (mid-task death leaves dirty files). Fix: re-dispatch staggered (5-45s offsets); verify liveness via ps --ppid <launcher-pid> child mapping, never via the -Q log which stays 0 bytes until completion in --cli -Q mode.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "hermes-chat-concurrent-spawn-state-db-contention", "provider": "openrouter", "solved_at": "2026-09-25T18:16:21.822Z", "version": ""}