Class: A Tier-2 evaluator that starves on a different knob after the token rung was raised. The rung is no longer the driver — the criterion is. The fix is to bound the criterion, not to climb a third rung.
I inspected the actual gitreins==0.14.0 source (downloaded the wheel and read engine/eval_cap.py, engine/pipeline.py, engine/judge.py, engine/task_manager.py, gitreins/cli.py) to ground the fix in the real cap-resolution hierarchy and CLI verbs. The solution is written to /workspace/solution.md:
gitreins-tier2-iteration-cap-after-token-raiseClass: A Tier-2 evaluator that starves on a different knob after the token rung was raised. The rung is no longer the driver — the criterion is. The fix is to bound the criterion, not to climb a third rung.
Stack: h3-sdk-python on gitreins 0.14.0; judge deepseek-v4-flash @ deepseek-foreman.
Evidence: verdicts 038d93be (original starvation), spy run /tmp/spy_gap055_pre.jsonl, bc9aaac8 (iteration cap at 3,148,006 tokens on the raised rung), 827a214f (COMPLETE/PASS at 1,380,265 tokens, no cap line).
| Run | Rung (iter / time / input) | Burn | Terminal line |
|---|---|---|---|
| original ×2 | 50 / ? / 1M | 1,039,907 & 1,029,521 | Input token budget (1.0M) exceeded (tier1 PASS, tier2 INCOMPLETE) |
after raise (bc9aaac8) |
100 / 20m / 8M | 3,148,006 (39 % of new token rung) | Iteration cap (100) reached (100.2 used) |
after criterion fix (827a214f) |
100 / 20m / 8M | 1,380,265 | no cap line — tier1 PASS + tier2 COMPLETE/PASS |
Two starvations at the same rung licensed exactly one raise. On the very next run the token knob stopped binding and the iteration knob bound. That is the expected hand-off, not a failed raise.
eval_cap_from_config() (engine/eval_cap.py) builds the effective cap from the defaults: section overlaid by the evaluator: section only — it never reads pipeline.stages[]. But the AI-eval pipeline step (engine/pipeline.py, _run_ai_eval) does:
max_iterations = step_def.get("max_iterations", -1)
...
if max_iterations not in (None, -1):
explicit_caps["max_iterations"] = max_iterations
...
if explicit_caps:
base = eval_cap_from_config(self.config)
eval_cap = EvalCap(
max_iterations=float(explicit_caps.get("max_iterations", base.max_iterations)),
max_input_tokens=int(explicit_caps.get("max_input_tokens", base.max_input_tokens)),
...
)
evaluator = AgenticEvaluator(self._llm, self.workdir, eval_cap=eval_cap)
else:
evaluator = AgenticEvaluator(self._llm, self.workdir) # defers to config.yaml
So the effective rung is pipeline.stages[tier2].max_iterations if it is set (≠ -1), else evaluator.max_iterations. Raising only the top-level evaluator: block while the tier2 stage still pins a number is inert — the stage value wins. (This is the "MIN(top-level, stage)" gate in the incident write-up: the practical ceiling is the stage pin.) Always confirm the effective rung before believing a raise took:
python - <<'PY'
import yaml
from engine.eval_cap import eval_cap_from_config
cfg = yaml.safe_load(open(".gitreins/config.yaml"))
c = eval_cap_from_config(cfg)
print("evaluator section ->", c.max_iterations, c.max_seconds,
c.max_input_tokens, c.max_output_tokens)
for s in cfg.get("pipeline", {}).get("stages", []):
if s.get("type") == "ai_eval" or s.get("id") == "tier2":
print("stage", s.get("id"), "pins ->",
s.get("max_iterations"), s.get("max_time"),
s.get("max_input_tokens"), s.get("max_output_tokens"))
PY
EvalCap.record_llm_call charges 1.0 iteration per LLM reasoning turn and record_tool_call charges tool_call_weight (default 0.1). The iteration cap is checked before the call (lenient — hence 100.2 used), while the input token cap is a hard limit checked after. After the raise, the judge reached 100 reasoning turns having spent only 3.1M input tokens: the judge was doing a lot of turns, each re-reading context, not a lot of distinct tokens. The iteration knob therefore fired first — exactly what the burn-to-cap ratio (3.1M / 8M = 39 %) predicted.
The tier2 criterion read:
the prose names every convention the battery asserts
That asks the judge to enumerate every assertion inside a cross-repo test battery. The judge can re-read the battery and re-derive the enumeration on every iteration, and the enumeration never closes because there is no canonical "done" set. Every turn costs 1.0 iteration, so the criterion — not the rung — consumed the 100 iterations. The growth law from the spy run (P0 = 3,831 + 1,064 tokens/call, 0 compactions) confirms the judge was in a reading loop, not converging.
Rule: at most two rung raises per incident. Once the same eval starves after two raises, the judged context (criterion / prompt) is the driver. A third rung is forbidden; amend the criterion instead.
gitreins 0.14.0 has no task update verb (only task create | start | complete | list | delete). To change the criteria you must delete → create → start. The amendment is stated inside the criterion so the judge sees why the wording changed.
TASK_ID="<task-id>" # the id judged by bc9aaac8
TITLE="<original task title>"
REPO="<path to h3-sdk-python checkout>" # cd there; must be the unchanged commit
cd "$REPO"
python - <<'PY'
import yaml
d = yaml.safe_load(open(".gitreins/tasks.yaml"))
for t in d.get("tasks", []):
if t["id"] == "<task-id>":
for i, c in enumerate(t["criteria"], 1):
print(f"{i}. {c}")
PY
Keep every other criterion byte-for-byte identical. Replace only the unbounded enumeration criterion:
AMENDED='[AMENDED 2026-09-19: the previous form "names every convention the battery asserts" required the judge to enumerate every assertion in a cross-repo test battery; that enumeration re-reads the battery every turn, never closes, and exhausted the iteration cap. This bounded form is decidable and is what must be judged from now on.] The prose names every convention GROUP declared by the cross-repo battery, one bullet per group. The judge MUST NOT enumerate individual test assertions. Verify once by listing the battery group headings (test-module docstrings / test class names) and confirming a one-to-one match with the bullet headings in the prose. FAIL only if a group heading has no matching bullet.'
OTHER_CRITERION_1='<unchanged criterion 2 verbatim>'
OTHER_CRITERION_2='<unchanged criterion 3 verbatim>'
gitreins task delete "$TASK_ID"
gitreins task create "$TASK_ID" "$TITLE" \
"$AMENDED" "$OTHER_CRITERION_1" "$OTHER_CRITERION_2"
gitreins task start "$TASK_ID"
Notes:
- Each positional argument is one criterion (criteria is nargs="*"), so quote each separately.
- If any criterion text begins with -, insert -- before the first criterion so argparse stops option parsing.
- task delete does not care that the task is complete; verdict history is stored separately and is preserved.
gitreins judge "$TASK_ID" 2>&1 | tee /tmp/judge_after_amend.log
gitreins judge re-runs Tier 1 guards + the Tier 2 evaluator against the current tree. Do not commit a new tree; the point is that the criterion changed, not the code.
So a top-level raise is never silently neutralised by a stale stage pin, either set the tier2 stage to defer, or keep the two numbers equal:
# .gitreins/config.yaml
evaluator:
max_iterations: 100 # effective when the stage defers
max_time: "20m"
max_input_tokens: "8M"
max_output_tokens: "1M"
pipeline:
stages:
- id: tier2
type: ai_eval
"on": ["pre-eval"] # quote it: YAML 1.1 turns bare `on:` into boolean True
condition: "true"
max_iterations: -1 # -1 = defer to the evaluator: block (raises then take effect)
# OR pin it to 100 here and raise BOTH places together — never just one.
python - <<'PY'
import yaml
from engine.eval_cap import eval_cap_from_config
c = eval_cap_from_config(yaml.safe_load(open(".gitreins/config.yaml")))
assert c.max_iterations == 100.0, c.max_iterations
assert c.max_seconds == 1200.0, c.max_seconds
assert c.max_input_tokens == 8_000_000, c.max_input_tokens
print("effective rung OK:", c.max_iterations, c.max_seconds, c.max_input_tokens)
PY
Expected: effective rung OK: 100.0 1200.0 8000000 (this is the exact check quoted in the incident evidence).
python -c "import yaml;print([t['criteria'] for t in yaml.safe_load(open('.gitreins/tasks.yaml'))['tasks'] if t['id']=='$TASK_ID'])"
Confirm the string contains AMENDED and no unbounded every convention the battery asserts phrase.
grep -E "Iteration cap|Input token budget|Time cap" /tmp/judge_after_amend.log \
&& { echo "REGRESSION: still starving"; exit 1; } \
|| echo "OK: no cap line"
grep -Ei "tier *1|tier *2|COMPLETE|PASS|FAIL" /tmp/judge_after_amend.log
Expected: Tier 1 PASS, Tier 2 COMPLETE with all criteria PASS, and no cap line.
tail -1 .gitreins/usage.jsonl | python -m json.tool
Expected (matches 827a214f): tokens_in ≈ 1,380,265 — 1.38× the original 1M cap, with no cap line. This is the key proof that the first token raise was necessary and that the remaining driver was the criterion, not the rung.
gitreins report -n 5
Confirm the newest verdict for $TASK_ID is COMPLETE/PASS and the previous one (bc9aaac8) is the iteration-cap INCOMPLETE.
| New terminal line | Meaning | Next action |
|---|---|---|
Iteration cap (100) reached |
criterion still unbounded | split the criterion into two single-decision tasks; do not raise |
Input token budget (8.0M) exceeded |
genuine context volume | only now is a raise licensed; size it from the measured burn |
Time cap (20m) exceeded |
wall-clock workload | raise max_time (or move work to Tier 1), not iterations/tokens |
| another criterion unsatisfiable | judged context | amend / split that criterion |
A raise should be sized from an instrumented spy run, never from the verdict text: P0 + tokens/call × expected calls, rounded up, with the next-smallest knob checked on the next run.
The token rung was raised and stopped binding; the iteration cap then bound because the criterion (names every convention the battery asserts) made the judge enumerate an unclosable cross-repo artifact once per turn. Since gitreins has no task update, delete + create + start the task with the criterion rewritten to a bounded one-to-one group mapping (with the amendment stated in the text), then gitreins judge <id> on the unchanged commit. Result: tier1 PASS + tier2 COMPLETE/PASS, no cap line, 1,380,265 tokens — 1.38× the original cap.
Files: solution written to /workspace/solution.md. The only repo-specific substitutions are TASK_ID, TITLE, the unchanged criteria, and the criterion's battery/group path wording.
# Evidence - Problem class: gitreins-tier2-iteration-cap-after-token-raise - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T09:56:30.097Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "", "environment": "Linux; h3-sdk-python (Python SDK, 51 source files, 192 tests) on gitreins 0.14.0; .gitreins/config.yaml carries a top-level evaluator block and a pipeline.stages[tier2] block that carries its own max_iterations (MIN gate). Judge model deepseek-v4-flash @ deepseek-foreman.", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-iteration-cap-after-token-raise", "provider": "openrouter", "solved_at": "2026-09-19T09:56:30.098Z", "version": "gitreins 0.14.0"}I inspected the actual gitreins==0.14.0 source (downloaded the wheel and read engine/eval_cap.py, engine/pipeline.py, engine/judge.py, engine/task_manager.py, gitreins/cli.py) to ground the fix in the real cap-resolution hierarchy and CLI verbs. The solution is written to /workspace/solution.md:
gitreins-tier2-iteration-cap-after-token-raiseClass: A Tier-2 evaluator that starves on a different knob after the token rung was raised. The rung is no longer the driver — the criterion is. The fix is to bound the criterion, not to climb a third rung.
Stack: h3-sdk-python on gitreins 0.14.0; judge deepseek-v4-flash @ deepseek-foreman.
Evidence: verdicts 038d93be (original starvation), spy run /tmp/spy_gap055_pre.jsonl, bc9aaac8 (iteration cap at 3,148,006 tokens on the raised rung), 827a214f (COMPLETE/PASS at 1,380,265 tokens, no cap line).
| Run | Rung (iter / time / input) | Burn | Terminal line |
|---|---|---|---|
| original ×2 | 50 / ? / 1M | 1,039,907 & 1,029,521 | Input token budget (1.0M) exceeded (tier1 PASS, tier2 INCOMPLETE) |
after raise (bc9aaac8) |
100 / 20m / 8M | 3,148,006 (39 % of new token rung) | Iteration cap (100) reached (100.2 used) |
after criterion fix (827a214f) |
100 / 20m / 8M | 1,380,265 | no cap line — tier1 PASS + tier2 COMPLETE/PASS |
Two starvations at the same rung licensed exactly one raise. On the very next run the token knob stopped binding and the iteration knob bound. That is the expected hand-off, not a failed raise.
eval_cap_from_config() (engine/eval_cap.py) builds the effective cap from the defaults: section overlaid by the evaluator: section only — it never reads pipeline.stages[]. But the AI-eval pipeline step (engine/pipeline.py, _run_ai_eval) does:
max_iterations = step_def.get("max_iterations", -1)
...
if max_iterations not in (None, -1):
explicit_caps["max_iterations"] = max_iterations
...
if explicit_caps:
base = eval_cap_from_config(self.config)
eval_cap = EvalCap(
max_iterations=float(explicit_caps.get("max_iterations", base.max_iterations)),
max_input_tokens=int(explicit_caps.get("max_input_tokens", base.max_input_tokens)),
...
)
evaluator = AgenticEvaluator(self._llm, self.workdir, eval_cap=eval_cap)
else:
evaluator = AgenticEvaluator(self._llm, self.workdir) # defers to config.yaml
So the effective rung is pipeline.stages[tier2].max_iterations if it is set (≠ -1), else evaluator.max_iterations. Raising only the top-level evaluator: block while the tier2 stage still pins a number is inert — the stage value wins. (This is the "MIN(top-level, stage)" gate in the incident write-up: the practical ceiling is the stage pin.) Always confirm the effective rung before believing a raise took:
python - <<'PY'
import yaml
from engine.eval_cap import eval_cap_from_config
cfg = yaml.safe_load(open(".gitreins/config.yaml"))
c = eval_cap_from_config(cfg)
print("evaluator section ->", c.max_iterations, c.max_seconds,
c.max_input_tokens, c.max_output_tokens)
for s in cfg.get("pipeline", {}).get("stages", []):
if s.get("type") == "ai_eval" or s.get("id") == "tier2":
print("stage", s.get("id"), "pins ->",
s.get("max_iterations"), s.get("max_time"),
s.get("max_input_tokens"), s.get("max_output_tokens"))
PY
EvalCap.record_llm_call charges 1.0 iteration per LLM reasoning turn and record_tool_call charges tool_call_weight (default 0.1). The iteration cap is checked before the call (lenient — hence 100.2 used), while the input token cap is a hard limit checked after. After the raise, the judge reached 100 reasoning turns having spent only 3.1M input tokens: the judge was doing a lot of turns, each re-reading context, not a lot of distinct tokens. The iteration knob therefore fired first — exactly what the burn-to-cap ratio (3.1M / 8M = 39 %) predicted.
The tier2 criterion read:
the prose names every convention the battery asserts
That asks the judge to enumerate every assertion inside a cross-repo test battery. The judge can re-read the battery and re-derive the enumeration on every iteration, and the enumeration never closes because there is no canonical "done" set. Every turn costs 1.0 iteration, so the criterion — not the rung — consumed the 100 iterations. The growth law from the spy run (P0 = 3,831 + 1,064 tokens/call, 0 compactions) confirms the judge was in a reading loop, not converging.
Rule: at most two rung raises per incident. Once the same eval starves after two raises, the judged context (criterion / prompt) is the driver. A third rung is forbidden; amend the criterion instead.
gitreins 0.14.0 has no task update verb (only task create | start | complete | list | delete). To change the criteria you must delete → create → start. The amendment is stated inside the criterion so the judge sees why the wording changed.
TASK_ID="<task-id>" # the id judged by bc9aaac8
TITLE="<original task title>"
REPO="<path to h3-sdk-python checkout>" # cd there; must be the unchanged commit
cd "$REPO"
python - <<'PY'
import yaml
d = yaml.safe_load(open(".gitreins/tasks.yaml"))
for t in d.get("tasks", []):
if t["id"] == "<task-id>":
for i, c in enumerate(t["criteria"], 1):
print(f"{i}. {c}")
PY
Keep every other criterion byte-for-byte identical. Replace only the unbounded enumeration criterion:
AMENDED='[AMENDED 2026-09-19: the previous form "names every convention the battery asserts" required the judge to enumerate every assertion in a cross-repo test battery; that enumeration re-reads the battery every turn, never closes, and exhausted the iteration cap. This bounded form is decidable and is what must be judged from now on.] The prose names every convention GROUP declared by the cross-repo battery, one bullet per group. The judge MUST NOT enumerate individual test assertions. Verify once by listing the battery group headings (test-module docstrings / test class names) and confirming a one-to-one match with the bullet headings in the prose. FAIL only if a group heading has no matching bullet.'
OTHER_CRITERION_1='<unchanged criterion 2 verbatim>'
OTHER_CRITERION_2='<unchanged criterion 3 verbatim>'
gitreins task delete "$TASK_ID"
gitreins task create "$TASK_ID" "$TITLE" \
"$AMENDED" "$OTHER_CRITERION_1" "$OTHER_CRITERION_2"
gitreins task start "$TASK_ID"
Notes:
- Each positional argument is one criterion (criteria is nargs="*"), so quote each separately.
- If any criterion text begins with -, insert -- before the first criterion so argparse stops option parsing.
- task delete does not care that the task is complete; verdict history is stored separately and is preserved.
gitreins judge "$TASK_ID" 2>&1 | tee /tmp/judge_after_amend.log
gitreins judge re-runs Tier 1 guards + the Tier 2 evaluator against the current tree. Do not commit a new tree; the point is that the criterion changed, not the code.
So a top-level raise is never silently neutralised by a stale stage pin, either set the tier2 stage to defer, or keep the two numbers equal:
# .gitreins/config.yaml
evaluator:
max_iterations: 100 # effective when the stage defers
max_time: "20m"
max_input_tokens: "8M"
max_output_tokens: "1M"
pipeline:
stages:
- id: tier2
type: ai_eval
"on": ["pre-eval"] # quote it: YAML 1.1 turns bare `on:` into boolean True
condition: "true"
max_iterations: -1 # -1 = defer to the evaluator: block (raises then take effect)
# OR pin it to 100 here and raise BOTH places together — never just one.
python - <<'PY'
import yaml
from engine.eval_cap import eval_cap_from_config
c = eval_cap_from_config(yaml.safe_load(open(".gitreins/config.yaml")))
assert c.max_iterations == 100.0, c.max_iterations
assert c.max_seconds == 1200.0, c.max_seconds
assert c.max_input_tokens == 8_000_000, c.max_input_tokens
print("effective rung OK:", c.max_iterations, c.max_seconds, c.max_input_tokens)
PY
Expected: effective rung OK: 100.0 1200.0 8000000 (this is the exact check quoted in the incident evidence).
python -c "import yaml;print([t['criteria'] for t in yaml.safe_load(open('.gitreins/tasks.yaml'))['tasks'] if t['id']=='$TASK_ID'])"
Confirm the string contains AMENDED and no unbounded every convention the battery asserts phrase.
grep -E "Iteration cap|Input token budget|Time cap" /tmp/judge_after_amend.log \
&& { echo "REGRESSION: still starving"; exit 1; } \
|| echo "OK: no cap line"
grep -Ei "tier *1|tier *2|COMPLETE|PASS|FAIL" /tmp/judge_after_amend.log
Expected: Tier 1 PASS, Tier 2 COMPLETE with all criteria PASS, and no cap line.
tail -1 .gitreins/usage.jsonl | python -m json.tool
Expected (matches 827a214f): tokens_in ≈ 1,380,265 — 1.38× the original 1M cap, with no cap line. This is the key proof that the first token raise was necessary and that the remaining driver was the criterion, not the rung.
gitreins report -n 5
Confirm the newest verdict for $TASK_ID is COMPLETE/PASS and the previous one (bc9aaac8) is the iteration-cap INCOMPLETE.
| New terminal line | Meaning | Next action |
|---|---|---|
Iteration cap (100) reached |
criterion still unbounded | split the criterion into two single-decision tasks; do not raise |
Input token budget (8.0M) exceeded |
genuine context volume | only now is a raise licensed; size it from the measured burn |
Time cap (20m) exceeded |
wall-clock workload | raise max_time (or move work to Tier 1), not iterations/tokens |
| another criterion unsatisfiable | judged context | amend / split that criterion |
A raise should be sized from an instrumented spy run, never from the verdict text: P0 + tokens/call × expected calls, rounded up, with the next-smallest knob checked on the next run.
The token rung was raised and stopped binding; the iteration cap then bound because the criterion (names every convention the battery asserts) made the judge enumerate an unclosable cross-repo artifact once per turn. Since gitreins has no task update, delete + create + start the task with the criterion rewritten to a bounded one-to-one group mapping (with the amendment stated in the text), then gitreins judge <id> on the unchanged commit. Result: tier1 PASS + tier2 COMPLETE/PASS, no cap line, 1,380,265 tokens — 1.38× the original cap.
Files: solution written to /workspace/solution.md. The only repo-specific substitutions are TASK_ID, TITLE, the unchanged criteria, and the criterion's battery/group path wording.
# Evidence - Problem class: gitreins-tier2-iteration-cap-after-token-raise - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T09:56:30.097Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "", "environment": "Linux; h3-sdk-python (Python SDK, 51 source files, 192 tests) on gitreins 0.14.0; .gitreins/config.yaml carries a top-level evaluator block and a pipeline.stages[tier2] block that carries its own max_iterations (MIN gate). Judge model deepseek-v4-flash @ deepseek-foreman.", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-iteration-cap-after-token-raise", "provider": "openrouter", "solved_at": "2026-09-19T09:56:30.098Z", "version": "gitreins 0.14.0"}