cxx-foreman-worker-resume-after-iteration-cap
Root cause. Tracked-file edits exist only in the worktree. When a worker hits the 200-iteration cap before committing, the only durable copies of that work are commits, stashes, or reflog-reachable objects — none of which exist. The scheduler's tree reset (git reset --hard / clean checkout) then destroys tracked edits outright; untracked new files survive because reset doesn't touch them. Recovery therefore cannot come from the tree or the transcript — it must come from a structured handoff artifact written at cap time, plus git forensics to confirm what was actually lost.
Fix. Implemented as ~/foreman_resume.py (stdlib only, git CLI as the sole dependency). The recovery pipeline:
1. Worker emits a structured HandoffSummary at cap time — design decisions, known failing tests (with expected/observed), and exact fixes (file/anchor/before/after). This is the only artifact that survives the reset:
@dataclass
class HandoffSummary:
task: str
iteration_cap: int
design_decisions: List[DesignDecision]
failing_tests: List[FailingTest]
exact_fixes: List[ExactFix] # file, anchor, before, after
new_files: List[str] # untracked survivors
remaining_steps: List[str]
source: str = "handoff" # "transcript" is banned
2. Verify edit loss before re-dispatching — inspect git stash list, git fsck --lost-found (dangling commits), git reflog, and git status --porcelain, then classify per handoff file: applied (still present), lost (wiped, irrecoverable), stashed/dangling (recoverable from git), conflicting (modified but no anchor matches), or untracked survivor:
def classify_unmodified(p):
text = (repo / p).read_text(...) if (repo / p).exists() else ""
if f.after in text: applied.append(p) # fix already committed
elif in_stash(p): stashed.append(p) # recoverable: git stash show --name-only
elif in_dangling(p): dangled.append(p) # recoverable: git show <hash>:<path>
else: lost.append(p)
3. Compile the resume prompt from the handoff only — carries design decisions, failing tests, and exact fixes verbatim so the fresh worker re-applies instead of re-derives; markers come from the LossReport buckets ([REAPPLY] vs [PRESENT] vs [RECOVERABLE (git stash)]), with "Definition of done: COMMIT".
4. Dispatch a fresh worker with that prompt. --restore attempts stash-apply / files-only cherry-pick first, then re-verifies before compiling; without it, a recoverable-work WARNING gates the fresh-worker spend.
Transcript ban is enforced, not just documented — assert_not_transcript() rejects source == "transcript", transcript-shaped dumps (no structured sections), and missing required fields (task, or any of file/anchor/before/after on a fix).
Verified two ways — a 23-test suite against **real temp git repos** (`git 2.53.0`, `Python 3.14.4`), plus a live CLI reproduction of the exact incident:
- **Incident reproduced live:** worker edited `solver.cpp`, created `newfile.h`, hit cap uncommitted; `git reset --hard` → `git status` shows only `?? newfile.h`. `foreman_resume.py verify` reported `lost_files: ["solver.cpp"]`, `untracked_survivors: ["newfile.h"]`, `recoverable: false`, stash/fsck empty, with the cap-era reflog captured as context. The resume prompt then marked `[REAPPLY] solver.cpp` and `[PRESENT — verify only, do not rewrite] newfile.h`.
- **Stash path live:** edited + `git stash` + reset → resume printed `WARNING: recoverable work exists...`; with `--restore` the stash was applied and `solver.cpp` came back as `return a * b;`.
- **Test coverage (23/23 pass):** wiped-tracked/survived-untracked classification; uncommitted edit leaves no reflog salvage; stash detected, listed, restored, and multi-stash parsed (`stash@{0}`/`stash@{1}`); dangling commit found via `fsck` after `reset HEAD~1` and restored; already-applied fixes marked PRESENT not lost; conflicting edits (matches neither anchor) detected; deleted tracked file → lost; prompt carries every handoff section verbatim (cap count, decisions, failing test + expected, before/after, remaining steps, commit DoD); missing-new-file marked `REAPPLY: create`; transcript source, transcript-shaped dump, missing `task`, and fix missing `after` all rejected with `ValueError`; CLI `verify`→`resume` pipeline and recoverable-work warning gate.
Two real bugs were caught by the tests and fixed during verification: (a) `compile_resume_prompt` initially read mutated `fix.status` off a *different* handoff object than the one `verify` annotated — now markers derive from `LossReport` buckets; (b) unmodified files that already contain the fix (committed at HEAD) were misreported as lost — now classified `applied`.{"model": "deepseek-v4-flash", "problem_class": "cxx-foreman-worker-resume-after-iteration-cap", "result": "passed", "tests": 23}Root cause. Tracked-file edits exist only in the worktree. When a worker hits the 200-iteration cap before committing, the only durable copies of that work are commits, stashes, or reflog-reachable objects — none of which exist. The scheduler's tree reset (git reset --hard / clean checkout) then destroys tracked edits outright; untracked new files survive because reset doesn't touch them. Recovery therefore cannot come from the tree or the transcript — it must come from a structured handoff artifact written at cap time, plus git forensics to confirm what was actually lost.
Fix. Implemented as ~/foreman_resume.py (stdlib only, git CLI as the sole dependency). The recovery pipeline:
1. Worker emits a structured HandoffSummary at cap time — design decisions, known failing tests (with expected/observed), and exact fixes (file/anchor/before/after). This is the only artifact that survives the reset:
@dataclass
class HandoffSummary:
task: str
iteration_cap: int
design_decisions: List[DesignDecision]
failing_tests: List[FailingTest]
exact_fixes: List[ExactFix] # file, anchor, before, after
new_files: List[str] # untracked survivors
remaining_steps: List[str]
source: str = "handoff" # "transcript" is banned
2. Verify edit loss before re-dispatching — inspect git stash list, git fsck --lost-found (dangling commits), git reflog, and git status --porcelain, then classify per handoff file: applied (still present), lost (wiped, irrecoverable), stashed/dangling (recoverable from git), conflicting (modified but no anchor matches), or untracked survivor:
def classify_unmodified(p):
text = (repo / p).read_text(...) if (repo / p).exists() else ""
if f.after in text: applied.append(p) # fix already committed
elif in_stash(p): stashed.append(p) # recoverable: git stash show --name-only
elif in_dangling(p): dangled.append(p) # recoverable: git show <hash>:<path>
else: lost.append(p)
3. Compile the resume prompt from the handoff only — carries design decisions, failing tests, and exact fixes verbatim so the fresh worker re-applies instead of re-derives; markers come from the LossReport buckets ([REAPPLY] vs [PRESENT] vs [RECOVERABLE (git stash)]), with "Definition of done: COMMIT".
4. Dispatch a fresh worker with that prompt. --restore attempts stash-apply / files-only cherry-pick first, then re-verifies before compiling; without it, a recoverable-work WARNING gates the fresh-worker spend.
Transcript ban is enforced, not just documented — assert_not_transcript() rejects source == "transcript", transcript-shaped dumps (no structured sections), and missing required fields (task, or any of file/anchor/before/after on a fix).
Verified two ways — a 23-test suite against **real temp git repos** (`git 2.53.0`, `Python 3.14.4`), plus a live CLI reproduction of the exact incident:
- **Incident reproduced live:** worker edited `solver.cpp`, created `newfile.h`, hit cap uncommitted; `git reset --hard` → `git status` shows only `?? newfile.h`. `foreman_resume.py verify` reported `lost_files: ["solver.cpp"]`, `untracked_survivors: ["newfile.h"]`, `recoverable: false`, stash/fsck empty, with the cap-era reflog captured as context. The resume prompt then marked `[REAPPLY] solver.cpp` and `[PRESENT — verify only, do not rewrite] newfile.h`.
- **Stash path live:** edited + `git stash` + reset → resume printed `WARNING: recoverable work exists...`; with `--restore` the stash was applied and `solver.cpp` came back as `return a * b;`.
- **Test coverage (23/23 pass):** wiped-tracked/survived-untracked classification; uncommitted edit leaves no reflog salvage; stash detected, listed, restored, and multi-stash parsed (`stash@{0}`/`stash@{1}`); dangling commit found via `fsck` after `reset HEAD~1` and restored; already-applied fixes marked PRESENT not lost; conflicting edits (matches neither anchor) detected; deleted tracked file → lost; prompt carries every handoff section verbatim (cap count, decisions, failing test + expected, before/after, remaining steps, commit DoD); missing-new-file marked `REAPPLY: create`; transcript source, transcript-shaped dump, missing `task`, and fix missing `after` all rejected with `ValueError`; CLI `verify`→`resume` pipeline and recoverable-work warning gate.
Two real bugs were caught by the tests and fixed during verification: (a) `compile_resume_prompt` initially read mutated `fix.status` off a *different* handoff object than the one `verify` annotated — now markers derive from `LossReport` buckets; (b) unmodified files that already contain the fix (committed at HEAD) were misreported as lost — now classified `applied`.{"model": "deepseek-v4-flash", "problem_class": "cxx-foreman-worker-resume-after-iteration-cap", "result": "passed", "tests": 23}