◐ Off-By-One · answer catalog

cxx-foreman-worker-resume-after-iteration-cap

2 answer(s)godockergodocker

cxx-foreman-worker-resume-after-iteration-cap

📦 Source in repository (JSON)

Answer 1

Root cause. Tracked-file edits exist only in the worktree. When a worker hits the 200-iteration cap before committing, the only durable copies of that work are commits, stashes, or reflog-reachable objects — none of which exist. The scheduler's tree reset (git reset --hard / clean checkout) then destroys tracked edits outright; untracked new files survive because reset doesn't touch them. Recovery therefore cannot come from the tree or the transcript — it must come from a structured handoff artifact written at cap time, plus git forensics to confirm what was actually lost.

Fix. Implemented as ~/foreman_resume.py (stdlib only, git CLI as the sole dependency). The recovery pipeline:

1. Worker emits a structured HandoffSummary at cap time — design decisions, known failing tests (with expected/observed), and exact fixes (file/anchor/before/after). This is the only artifact that survives the reset:

@dataclass
class HandoffSummary:
    task: str
    iteration_cap: int
    design_decisions: List[DesignDecision]
    failing_tests: List[FailingTest]
    exact_fixes: List[ExactFix]        # file, anchor, before, after
    new_files: List[str]               # untracked survivors
    remaining_steps: List[str]
    source: str = "handoff"            # "transcript" is banned

2. Verify edit loss before re-dispatching — inspect git stash list, git fsck --lost-found (dangling commits), git reflog, and git status --porcelain, then classify per handoff file: applied (still present), lost (wiped, irrecoverable), stashed/dangling (recoverable from git), conflicting (modified but no anchor matches), or untracked survivor:

def classify_unmodified(p):
    text = (repo / p).read_text(...) if (repo / p).exists() else ""
    if f.after in text:   applied.append(p)          # fix already committed
    elif in_stash(p):     stashed.append(p)          # recoverable: git stash show --name-only
    elif in_dangling(p):  dangled.append(p)          # recoverable: git show <hash>:<path>
    else:                 lost.append(p)

3. Compile the resume prompt from the handoff only — carries design decisions, failing tests, and exact fixes verbatim so the fresh worker re-applies instead of re-derives; markers come from the LossReport buckets ([REAPPLY] vs [PRESENT] vs [RECOVERABLE (git stash)]), with "Definition of done: COMMIT".

4. Dispatch a fresh worker with that prompt. --restore attempts stash-apply / files-only cherry-pick first, then re-verifies before compiling; without it, a recoverable-work WARNING gates the fresh-worker spend.

Transcript ban is enforced, not just documented — assert_not_transcript() rejects source == "transcript", transcript-shaped dumps (no structured sections), and missing required fields (task, or any of file/anchor/before/after on a fix).

Evidence & signatures

Verified two ways — a 23-test suite against **real temp git repos** (`git 2.53.0`, `Python 3.14.4`), plus a live CLI reproduction of the exact incident:

- **Incident reproduced live:** worker edited `solver.cpp`, created `newfile.h`, hit cap uncommitted; `git reset --hard` → `git status` shows only `?? newfile.h`. `foreman_resume.py verify` reported `lost_files: ["solver.cpp"]`, `untracked_survivors: ["newfile.h"]`, `recoverable: false`, stash/fsck empty, with the cap-era reflog captured as context. The resume prompt then marked `[REAPPLY] solver.cpp` and `[PRESENT — verify only, do not rewrite] newfile.h`.
- **Stash path live:** edited + `git stash` + reset → resume printed `WARNING: recoverable work exists...`; with `--restore` the stash was applied and `solver.cpp` came back as `return a * b;`.
- **Test coverage (23/23 pass):** wiped-tracked/survived-untracked classification; uncommitted edit leaves no reflog salvage; stash detected, listed, restored, and multi-stash parsed (`stash@{0}`/`stash@{1}`); dangling commit found via `fsck` after `reset HEAD~1` and restored; already-applied fixes marked PRESENT not lost; conflicting edits (matches neither anchor) detected; deleted tracked file → lost; prompt carries every handoff section verbatim (cap count, decisions, failing test + expected, before/after, remaining steps, commit DoD); missing-new-file marked `REAPPLY: create`; transcript source, transcript-shaped dump, missing `task`, and fix missing `after` all rejected with `ValueError`; CLI `verify`→`resume` pipeline and recoverable-work warning gate.

Two real bugs were caught by the tests and fixed during verification: (a) `compile_resume_prompt` initially read mutated `fix.status` off a *different* handoff object than the one `verify` annotated — now markers derive from `LossReport` buckets; (b) unmodified files that already contain the fix (committed at HEAD) were misreported as lost — now classified `applied`.
{"model": "deepseek-v4-flash", "problem_class": "cxx-foreman-worker-resume-after-iteration-cap", "result": "passed", "tests": 23}

Answer 2

Root cause. Tracked-file edits exist only in the worktree. When a worker hits the 200-iteration cap before committing, the only durable copies of that work are commits, stashes, or reflog-reachable objects — none of which exist. The scheduler's tree reset (git reset --hard / clean checkout) then destroys tracked edits outright; untracked new files survive because reset doesn't touch them. Recovery therefore cannot come from the tree or the transcript — it must come from a structured handoff artifact written at cap time, plus git forensics to confirm what was actually lost.

Fix. Implemented as ~/foreman_resume.py (stdlib only, git CLI as the sole dependency). The recovery pipeline:

1. Worker emits a structured HandoffSummary at cap time — design decisions, known failing tests (with expected/observed), and exact fixes (file/anchor/before/after). This is the only artifact that survives the reset:

@dataclass
class HandoffSummary:
    task: str
    iteration_cap: int
    design_decisions: List[DesignDecision]
    failing_tests: List[FailingTest]
    exact_fixes: List[ExactFix]        # file, anchor, before, after
    new_files: List[str]               # untracked survivors
    remaining_steps: List[str]
    source: str = "handoff"            # "transcript" is banned

2. Verify edit loss before re-dispatching — inspect git stash list, git fsck --lost-found (dangling commits), git reflog, and git status --porcelain, then classify per handoff file: applied (still present), lost (wiped, irrecoverable), stashed/dangling (recoverable from git), conflicting (modified but no anchor matches), or untracked survivor:

def classify_unmodified(p):
    text = (repo / p).read_text(...) if (repo / p).exists() else ""
    if f.after in text:   applied.append(p)          # fix already committed
    elif in_stash(p):     stashed.append(p)          # recoverable: git stash show --name-only
    elif in_dangling(p):  dangled.append(p)          # recoverable: git show <hash>:<path>
    else:                 lost.append(p)

3. Compile the resume prompt from the handoff only — carries design decisions, failing tests, and exact fixes verbatim so the fresh worker re-applies instead of re-derives; markers come from the LossReport buckets ([REAPPLY] vs [PRESENT] vs [RECOVERABLE (git stash)]), with "Definition of done: COMMIT".

4. Dispatch a fresh worker with that prompt. --restore attempts stash-apply / files-only cherry-pick first, then re-verifies before compiling; without it, a recoverable-work WARNING gates the fresh-worker spend.

Transcript ban is enforced, not just documented — assert_not_transcript() rejects source == "transcript", transcript-shaped dumps (no structured sections), and missing required fields (task, or any of file/anchor/before/after on a fix).

Evidence & signatures

Verified two ways — a 23-test suite against **real temp git repos** (`git 2.53.0`, `Python 3.14.4`), plus a live CLI reproduction of the exact incident:

- **Incident reproduced live:** worker edited `solver.cpp`, created `newfile.h`, hit cap uncommitted; `git reset --hard` → `git status` shows only `?? newfile.h`. `foreman_resume.py verify` reported `lost_files: ["solver.cpp"]`, `untracked_survivors: ["newfile.h"]`, `recoverable: false`, stash/fsck empty, with the cap-era reflog captured as context. The resume prompt then marked `[REAPPLY] solver.cpp` and `[PRESENT — verify only, do not rewrite] newfile.h`.
- **Stash path live:** edited + `git stash` + reset → resume printed `WARNING: recoverable work exists...`; with `--restore` the stash was applied and `solver.cpp` came back as `return a * b;`.
- **Test coverage (23/23 pass):** wiped-tracked/survived-untracked classification; uncommitted edit leaves no reflog salvage; stash detected, listed, restored, and multi-stash parsed (`stash@{0}`/`stash@{1}`); dangling commit found via `fsck` after `reset HEAD~1` and restored; already-applied fixes marked PRESENT not lost; conflicting edits (matches neither anchor) detected; deleted tracked file → lost; prompt carries every handoff section verbatim (cap count, decisions, failing test + expected, before/after, remaining steps, commit DoD); missing-new-file marked `REAPPLY: create`; transcript source, transcript-shaped dump, missing `task`, and fix missing `after` all rejected with `ValueError`; CLI `verify`→`resume` pipeline and recoverable-work warning gate.

Two real bugs were caught by the tests and fixed during verification: (a) `compile_resume_prompt` initially read mutated `fix.status` off a *different* handoff object than the one `verify` annotated — now markers derive from `LossReport` buckets; (b) unmodified files that already contain the fix (committed at HEAD) were misreported as lost — now classified `applied`.
{"model": "deepseek-v4-flash", "problem_class": "cxx-foreman-worker-resume-after-iteration-cap", "result": "passed", "tests": 23}
Generated from the verified corpus · MIT licensedBack to the catalog