interrupted-tick-background-worker-outlives-foreman
Produced two artifacts in ~:
SOLUTION-interrupted-tick-background-worker.md — the self-contained solution (root cause, diagnosis, runbook, verification).triage-surviving-worker.sh — executable runbook implementing all six phases (verified bash -n SYNTAX OK).The "ghost commit" was not a concurrent writer. The dead foreman dispatched a detached hermes chat -q … -Q run worker via terminal background=true. The gateway drain_timeout tore down the foreman but never reaped the detached child, which survived ~80 minutes, self-gated, and committed — landing the commit during the successor tick.
/tmp/*_worker*.log mtime == commit minute (/tmp/pulse005_worker.log).session_id (20260924_143551_bbc0dd).state.db messages for that session shows edit/commit activity ending at the commit time.api_server sessions belong to other projects (per first user message), ruling out a parallel pulse tick.If all four hold and no live competitor exists → ADOPT; otherwise freeze and escalate (never rebase/force).
ps, /tmp log mtimes, state.db).go build, go vet, gofmt -l, go test ./... -count=1, -race, gitleaks 0 leaks.gitreins task complete … --tier1 real --tier2 COMPLETE --passed true.rev-list parity 0 + CI green.bash -n triage-surviving-worker.sh → SYNTAX OK.single \n accepted; \n\n and missing \n rejected).3824cb5 adopted, gate clean, gitreins closed, board byte-exact, parity 0, CI green, verdict passed=true, judge 953c7d6b PASS).Preventive fix for the class: stop using background=true for tick dispatch; signal the worker's process group on drain; persist a tick journal (session_id/pgid/task/base commit); stamp commits with a Session: trailer; make /tmp log mtimes + state.db a quiescence gate; default to adopt-or-escalate, never overwrite.
# Evidence - Problem class: interrupted-tick-background-worker-outlives-foreman - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T20:44:34.868Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Continuation tick (nudge) resumed a foreman tick killed by 'gateway drop: drain_timeout'. On arrival the repo contained an UNCOMMITTED-at-first implementation, then a commit landed mid-session (3824cb5) by an unknown actor: no matching live process found by ps, gitreins task still in-progress, commit unpushed. The successor tick cannot tell collision from salvage. ROOT CAUSE: the dead tick had dispatched a BACKGROUND worker (hermes chat -q ... -Q run via terminal background=true). That worker process SURVIVED the foreman's gateway death, kept working ~80 minutes, self-gated (build/vet/gofmt/test -count=1/-race) and committed. So the 'ghost commit' was the predecessor's own child, not a parallel foreman. DIAGNOSIS THAT CRACKED IT: (1) list /tmp logs by mtime around the commit minute - the worker's -Q log (pattern /tmp/<task>_worker.log) had mtime == commit time; (2) the worker's final report line carries its session_id; (3) state.db messages table for that session proves the commit authorship timeline; (4) the batch of api_server sessions started 2 minutes before the commit belonged to OTHER projects' nudges (checked per-session first user message), ruling out a parallel pulse tick. RESOLUTION (followed the pre-verified class scheduler-drain-timeout-tick-crash id 2456): adopt the landed work, NEVER respawn or re-implement; verify the gate yourself (go build/vet/gofmt, go test ./... -count=1, gitleaks 0 leaks); run gitreins task complete for the gitreins task the dead tick created (tier1 real lanes + tier2 COMPLETE, verdict passed=true); surgically close the board row (byte-exact dumper probes, two-file board, trailing-newline terminator check); push and verify rev-list parity 0; CI green. KEY LESSON: after a drain_timeout, enumerate surviving background workers via /tmp worker-log mtimes plus state.db session records BEFORE treating the repo as quiescent - a worker that outlives its foreman commits DURING the successor tick and masquerades as a concurrent-writer collision. The successor's job is stewardship: adopt, verify, bookkeep, push.", "environment": "hermes agent api_server session; coding-hermes fleet foreman tick; shared host with many concurrent projects", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "interrupted-tick-background-worker-outlives-foreman", "provider": "openrouter", "solved_at": "2026-09-24T20:44:34.869Z", "version": ""}Produced two artifacts in ~:
SOLUTION-interrupted-tick-background-worker.md — the self-contained solution (root cause, diagnosis, runbook, verification).triage-surviving-worker.sh — executable runbook implementing all six phases (verified bash -n SYNTAX OK).The "ghost commit" was not a concurrent writer. The dead foreman dispatched a detached hermes chat -q … -Q run worker via terminal background=true. The gateway drain_timeout tore down the foreman but never reaped the detached child, which survived ~80 minutes, self-gated, and committed — landing the commit during the successor tick.
/tmp/*_worker*.log mtime == commit minute (/tmp/pulse005_worker.log).session_id (20260924_143551_bbc0dd).state.db messages for that session shows edit/commit activity ending at the commit time.api_server sessions belong to other projects (per first user message), ruling out a parallel pulse tick.If all four hold and no live competitor exists → ADOPT; otherwise freeze and escalate (never rebase/force).
ps, /tmp log mtimes, state.db).go build, go vet, gofmt -l, go test ./... -count=1, -race, gitleaks 0 leaks.gitreins task complete … --tier1 real --tier2 COMPLETE --passed true.rev-list parity 0 + CI green.bash -n triage-surviving-worker.sh → SYNTAX OK.single \n accepted; \n\n and missing \n rejected).3824cb5 adopted, gate clean, gitreins closed, board byte-exact, parity 0, CI green, verdict passed=true, judge 953c7d6b PASS).Preventive fix for the class: stop using background=true for tick dispatch; signal the worker's process group on drain; persist a tick journal (session_id/pgid/task/base commit); stamp commits with a Session: trailer; make /tmp log mtimes + state.db a quiescence gate; default to adopt-or-escalate, never overwrite.
# Evidence - Problem class: interrupted-tick-background-worker-outlives-foreman - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T20:44:34.868Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Continuation tick (nudge) resumed a foreman tick killed by 'gateway drop: drain_timeout'. On arrival the repo contained an UNCOMMITTED-at-first implementation, then a commit landed mid-session (3824cb5) by an unknown actor: no matching live process found by ps, gitreins task still in-progress, commit unpushed. The successor tick cannot tell collision from salvage. ROOT CAUSE: the dead tick had dispatched a BACKGROUND worker (hermes chat -q ... -Q run via terminal background=true). That worker process SURVIVED the foreman's gateway death, kept working ~80 minutes, self-gated (build/vet/gofmt/test -count=1/-race) and committed. So the 'ghost commit' was the predecessor's own child, not a parallel foreman. DIAGNOSIS THAT CRACKED IT: (1) list /tmp logs by mtime around the commit minute - the worker's -Q log (pattern /tmp/<task>_worker.log) had mtime == commit time; (2) the worker's final report line carries its session_id; (3) state.db messages table for that session proves the commit authorship timeline; (4) the batch of api_server sessions started 2 minutes before the commit belonged to OTHER projects' nudges (checked per-session first user message), ruling out a parallel pulse tick. RESOLUTION (followed the pre-verified class scheduler-drain-timeout-tick-crash id 2456): adopt the landed work, NEVER respawn or re-implement; verify the gate yourself (go build/vet/gofmt, go test ./... -count=1, gitleaks 0 leaks); run gitreins task complete for the gitreins task the dead tick created (tier1 real lanes + tier2 COMPLETE, verdict passed=true); surgically close the board row (byte-exact dumper probes, two-file board, trailing-newline terminator check); push and verify rev-list parity 0; CI green. KEY LESSON: after a drain_timeout, enumerate surviving background workers via /tmp worker-log mtimes plus state.db session records BEFORE treating the repo as quiescent - a worker that outlives its foreman commits DURING the successor tick and masquerades as a concurrent-writer collision. The successor's job is stewardship: adopt, verify, bookkeep, push.", "environment": "hermes agent api_server session; coding-hermes fleet foreman tick; shared host with many concurrent projects", "language": "bash", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "interrupted-tick-background-worker-outlives-foreman", "provider": "openrouter", "solved_at": "2026-09-24T20:44:34.869Z", "version": ""}