foreman-interrupted-wave-closeout-recovery
Done. The solution is written to /workspace/SOLUTION.md and is backed by a runnable, verified harness in /workspace/recovery-lab/.
There was no live repo/fixture in the environment (only the problem-class descriptor at /workspace/problem.json), so I built a faithful reproduction of the interrupted state and verified the fix against it.
Root-cause analysis (in the doc):
1. Non-atomic tick lifecycle — merges commit, then closeout bookkeeping crashes; nothing reconciles forward.
2. Persistence pitfall — closeout mutates board rows in memory but never calls write_text, so flips silently vanish and retries loop/re-dispatch.
3. Escape-only diff noise — naive full-board re-serialization with ensure_ascii=True rewrites untouched CJK rows as \uXXXX.
4. passed:false is a real fail — a tier-2 failure must become failed/rework, not be laundered to done because gates looked green.
The fix — recover_closeout.py, a drop-in 7-step reconciler: verify merges as ancestors of HEAD → read verdicts → re-run gates → surgical line-preserving board update with atomic persist + --self-check disk re-read → deduped events + header refresh → pathspec commit/push → detached rework worktree at HEAD. It also classifies the board diff per id as semantic vs escape-only, and is idempotent on re-run.
The synthetic repo (3 merged worker branches, one tier-2 passed:false, unrelated T-900, CJK rows) demonstrates:
- buggy in-memory closeout run twice → flips lost, git status clean (the pitfall is real);
- classifier distinguishes escape-only from semantic by id;
- real recovery persists T-010/T-011=done, T-012=failed, leaves T-900 untouched, refreshes header, appends 3 events, commits with a 3-file pathspec, stages the rework worktree;
- RESULT: 22 passed, 0 failed.
The embedded script in SOLUTION.md is byte-identical to the verified recover_closeout.py (confirmed by diff). Reproduce anytime with:
cd /workspace/recovery-lab && bash verify.sh
# Evidence - Problem class: foreman-interrupted-wave-closeout-recovery - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-24T12:35:32.572Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A scheduler tick crashed after merging wave workers but before board closeout. Detection: wave manifest .coding-hermes/waves/<tick-id>.json exists with recovery_status field + board rows still in_progress while git log already shows the merge commits. Recovery order: (1) verify merges landed via git log + manifest merged_branches, (2) read each worker verdict.json (passed field) \u2014 a passed:false tier2 verdict is a REAL fail needing rework, not a stale premise, (3) re-run gates on the merged tree yourself (build/vet/gofmt/golangci-lint/-race), (4) surgical per-line board row updates (one flip per touched row; expect ~8-9 extra escape-only diff lines from ensure_ascii re-serialization, classify by id before commit), (5) append missing events, refresh header, commit with pathspec, push, (6) leave a staged worktree at HEAD for the rework dispatch instead of re-dispatching into a 2h-exceeded tick. Key pitfall: the closeout script must WRITE tasks.jsonl back to disk \u2014 modifying rows in memory without persisting silently loses the flips on re-run.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "foreman-interrupted-wave-closeout-recovery", "provider": "openrouter", "solved_at": "2026-09-24T12:35:32.572Z", "version": ""}