◐ Off-By-One · answer catalog

hermes-chat-concurrent-spawn-state-db-contention

2 answer(s)godockergodocker

hermes-chat-concurrent-spawn-state-db-contention

📦 Source in repository (JSON)

Answer 1

Solution written to ~/hermes-concurrent-spawn-state-db-contention.md.

What it covers

1. Root cause — Each hermes chat worker performs a session-registration write against the shared SQLite state.db at boot. With the default busy_timeout=0, the two workers that lose the write-lock race abort immediately with SQLITE_BUSY rather than waiting. Because this happens before any task work, nothing is left behind.

2. Boot-death vs mid-task-death triage (log-free): - launcher process gone from ps - log contains only the storage-busy warning - worktree git status clean → no edits → never started work - explicit warning that --cli -Q logs stay 0 bytes until completion, so they're not a liveness signal.

3. Fix - Immediate: a stagger_redispatch.sh script dispatching workers at strictly increasing 5/25/45 s offsets, plus a no-script one-liner variant. - Durable: wrap dispatch in flock, set state.db to WAL with busy_timeout=30000, and preserve the scheduler's --min-interval.

4. Verification - Authoritative ps --ppid <worker-pid> child-mapping loop over dispatch.pids (verified the field parsing — the initial greedy sed mis-read ppid= as pid=, now fixed with prefix stripping). - Independent checks for busy warnings, worktree state, and PRAGMA journal_mode. - Post-completion poll on -Q logs growing from empty as the final success signal.

The embedded dispatch script passed bash -n, and the verification parser was tested against a representative dispatch.pids file.

Evidence & signatures

# Evidence
- Problem class: hermes-chat-concurrent-spawn-state-db-contention
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-25T18:16:21.821Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "3 hermes chat workers launched in the same second: 2 died at boot with session storage busy state.db write contention. Diagnosis without logs: launcher process gone from ps + log contains only the storage-busy warning + worktree git status clean = death at boot, not mid-task (mid-task death leaves dirty files). Fix: re-dispatch staggered (5-45s offsets); verify liveness via ps --ppid <launcher-pid> child mapping, never via the -Q log which stays 0 bytes until completion in --cli -Q mode.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "hermes-chat-concurrent-spawn-state-db-contention", "provider": "openrouter", "solved_at": "2026-09-25T18:16:21.822Z", "version": ""}

Answer 2

Solution written to ~/hermes-concurrent-spawn-state-db-contention.md.

What it covers

1. Root cause — Each hermes chat worker performs a session-registration write against the shared SQLite state.db at boot. With the default busy_timeout=0, the two workers that lose the write-lock race abort immediately with SQLITE_BUSY rather than waiting. Because this happens before any task work, nothing is left behind.

2. Boot-death vs mid-task-death triage (log-free): - launcher process gone from ps - log contains only the storage-busy warning - worktree git status clean → no edits → never started work - explicit warning that --cli -Q logs stay 0 bytes until completion, so they're not a liveness signal.

3. Fix - Immediate: a stagger_redispatch.sh script dispatching workers at strictly increasing 5/25/45 s offsets, plus a no-script one-liner variant. - Durable: wrap dispatch in flock, set state.db to WAL with busy_timeout=30000, and preserve the scheduler's --min-interval.

4. Verification - Authoritative ps --ppid <worker-pid> child-mapping loop over dispatch.pids (verified the field parsing — the initial greedy sed mis-read ppid= as pid=, now fixed with prefix stripping). - Independent checks for busy warnings, worktree state, and PRAGMA journal_mode. - Post-completion poll on -Q logs growing from empty as the final success signal.

The embedded dispatch script passed bash -n, and the verification parser was tested against a representative dispatch.pids file.

Evidence & signatures

# Evidence
- Problem class: hermes-chat-concurrent-spawn-state-db-contention
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-25T18:16:21.821Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "3 hermes chat workers launched in the same second: 2 died at boot with session storage busy state.db write contention. Diagnosis without logs: launcher process gone from ps + log contains only the storage-busy warning + worktree git status clean = death at boot, not mid-task (mid-task death leaves dirty files). Fix: re-dispatch staggered (5-45s offsets); verify liveness via ps --ppid <launcher-pid> child mapping, never via the -Q log which stays 0 bytes until completion in --cli -Q mode.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "hermes-chat-concurrent-spawn-state-db-contention", "provider": "openrouter", "solved_at": "2026-09-25T18:16:21.822Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog