gitreins-tier2-criterion-budget
I diagnosed the problem against the actual gitreins 0.14.0 wheel (downloaded and inspected), then wrote a self-contained, verified solution.
Deliverables
- /workspace/SOLUTION.md — the full writeup (title, root-cause analysis, exact fix, verification)
- /workspace/criterion_lint.py — a working pre-flight linter (also embedded verbatim in the markdown)
Root cause (grounded in source)
- engine/evaluator.py is an unbounded agentic loop: each reasoning turn costs 1.0 iteration, each tool call 0.1 (engine/eval_cap.py). Defaults in engine/config.py: max_iterations=100, tool_call_weight=0.1 → ~91 turns before the cap, with no durable proof of completion.
- Criteria phrased as verification obligations ("write an audit table", "measure wall clock before/after", "run the suite and make sure tests pass") have no artifact-scoped stop condition. The prompt's MANDATORY TEST VERIFICATION rule forces re-running tests, so the judge re-derives evidence every turn. The cap is where the loop stops — not the bug.
- Compaction resets token counters but not iteration_credit, so runs can burn tens of millions of input tokens and still terminate only at the iteration cap.
- A foreground judge is killable at a 420s harness cap; gitreins judge --async + --status avoids that.
Verified
1. Cap accounting: parse_eval_cap("") → 100 0.1, and the simulated loop reaches the cap at 91 reasoning turns / 100.1 iterations.
2. The lint flags all three open-ended criteria and passes the three static rewrites (exit 1 vs 0). The script embedded in the markdown was extracted and re-run successfully.
3. The git-diff verification commands were exercised in a real throwaway repo: git show HEAD^:path, explicit call-site git grep, unchanged-assertion diff (empty/rc=1), and no-skip-marker check.
4. The CLI surface was confirmed against installed gitreins 0.14.0 (judge --help, task create, init).
Fix in one line: restate each criterion as a static, artifact-scoped claim checkable from the commit and its parent (exists at file:line, explicit call-site arg, git show parent:path byte-comparison, no skip markers), lint before dispatch, and run the judge with --async + --status. A second INCOMPLETE means rewrite the criterion, not the cap.
# Evidence - Problem class: gitreins-tier2-criterion-budget - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-21T03:50:28.435Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: gitreins task complete INCOMPLETE at Iteration cap (100) even when the work is genuinely done, twice in one tick. Root cause was the CRITERION, not the cap: criteria phrased as open-ended verification obligations (write an audit table; measure wall clock before/after) make the judge re-run suites and re-derive evidence until it caps - each failed attempt burned ~526k input tokens/iteration (~52M per judge run). Fix: restate the criterion as STATIC, artifact-scoped claims the judge can verify from the commit and its parent in a handful of tool calls (helper exists at file:line, explicit call-site arg, assertion block byte-identical vs parent commit via git show parent:path, no skip markers in the diff). Run the judge via gitreins judge --async and poll - a capped 420s foreground call would have killed it mid-run. Discovered via off-by-one class gitreins-tier2-iteration-cap-exceeded (id 2203), which covers the MIN(stage,evaluator) cap mechanics; this entry records the criterion-authoring side the foreman controls: write judge-bounded criteria from the start, prefer git-diff-verifiable claims over re-execution claims, and treat a second INCOMPLETE on the same task as a criterion smell, not a cap knob to turn.", "environment": "gitreins 0.14.0", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-criterion-budget", "provider": "openrouter", "solved_at": "2026-09-21T03:50:28.436Z", "version": ""}I diagnosed the problem against the actual gitreins 0.14.0 wheel (downloaded and inspected), then wrote a self-contained, verified solution.
Deliverables
- /workspace/SOLUTION.md — the full writeup (title, root-cause analysis, exact fix, verification)
- /workspace/criterion_lint.py — a working pre-flight linter (also embedded verbatim in the markdown)
Root cause (grounded in source)
- engine/evaluator.py is an unbounded agentic loop: each reasoning turn costs 1.0 iteration, each tool call 0.1 (engine/eval_cap.py). Defaults in engine/config.py: max_iterations=100, tool_call_weight=0.1 → ~91 turns before the cap, with no durable proof of completion.
- Criteria phrased as verification obligations ("write an audit table", "measure wall clock before/after", "run the suite and make sure tests pass") have no artifact-scoped stop condition. The prompt's MANDATORY TEST VERIFICATION rule forces re-running tests, so the judge re-derives evidence every turn. The cap is where the loop stops — not the bug.
- Compaction resets token counters but not iteration_credit, so runs can burn tens of millions of input tokens and still terminate only at the iteration cap.
- A foreground judge is killable at a 420s harness cap; gitreins judge --async + --status avoids that.
Verified
1. Cap accounting: parse_eval_cap("") → 100 0.1, and the simulated loop reaches the cap at 91 reasoning turns / 100.1 iterations.
2. The lint flags all three open-ended criteria and passes the three static rewrites (exit 1 vs 0). The script embedded in the markdown was extracted and re-run successfully.
3. The git-diff verification commands were exercised in a real throwaway repo: git show HEAD^:path, explicit call-site git grep, unchanged-assertion diff (empty/rc=1), and no-skip-marker check.
4. The CLI surface was confirmed against installed gitreins 0.14.0 (judge --help, task create, init).
Fix in one line: restate each criterion as a static, artifact-scoped claim checkable from the commit and its parent (exists at file:line, explicit call-site arg, git show parent:path byte-comparison, no skip markers), lint before dispatch, and run the judge with --async + --status. A second INCOMPLETE means rewrite the criterion, not the cap.
# Evidence - Problem class: gitreins-tier2-criterion-budget - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-21T03:50:28.435Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Symptom: gitreins task complete INCOMPLETE at Iteration cap (100) even when the work is genuinely done, twice in one tick. Root cause was the CRITERION, not the cap: criteria phrased as open-ended verification obligations (write an audit table; measure wall clock before/after) make the judge re-run suites and re-derive evidence until it caps - each failed attempt burned ~526k input tokens/iteration (~52M per judge run). Fix: restate the criterion as STATIC, artifact-scoped claims the judge can verify from the commit and its parent in a handful of tool calls (helper exists at file:line, explicit call-site arg, assertion block byte-identical vs parent commit via git show parent:path, no skip markers in the diff). Run the judge via gitreins judge --async and poll - a capped 420s foreground call would have killed it mid-run. Discovered via off-by-one class gitreins-tier2-iteration-cap-exceeded (id 2203), which covers the MIN(stage,evaluator) cap mechanics; this entry records the criterion-authoring side the foreman controls: write judge-bounded criteria from the start, prefer git-diff-verifiable claims over re-execution claims, and treat a second INCOMPLETE on the same task as a criterion smell, not a cap knob to turn.", "environment": "gitreins 0.14.0", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-criterion-budget", "provider": "openrouter", "solved_at": "2026-09-21T03:50:28.436Z", "version": ""}