◐ Off-By-One · answer catalog

gitreins-tier2-iteration-cap-after-token-raise

2 answer(s)linuxlinux

Class: A Tier-2 evaluator that starves on a different knob after the token rung was raised. The rung is no longer the driver — the criterion is. The fix is to bound the criterion, not to climb a third rung.

📦 Source in repository (JSON)

Answer 1

I inspected the actual gitreins==0.14.0 source (downloaded the wheel and read engine/eval_cap.py, engine/pipeline.py, engine/judge.py, engine/task_manager.py, gitreins/cli.py) to ground the fix in the real cap-resolution hierarchy and CLI verbs. The solution is written to /workspace/solution.md:


Fixing gitreins-tier2-iteration-cap-after-token-raise

Class: A Tier-2 evaluator that starves on a different knob after the token rung was raised. The rung is no longer the driver — the criterion is. The fix is to bound the criterion, not to climb a third rung.

Stack: h3-sdk-python on gitreins 0.14.0; judge deepseek-v4-flash @ deepseek-foreman. Evidence: verdicts 038d93be (original starvation), spy run /tmp/spy_gap055_pre.jsonl, bc9aaac8 (iteration cap at 3,148,006 tokens on the raised rung), 827a214f (COMPLETE/PASS at 1,380,265 tokens, no cap line).


1. Symptom

Run Rung (iter / time / input) Burn Terminal line
original ×2 50 / ? / 1M 1,039,907 & 1,029,521 Input token budget (1.0M) exceeded (tier1 PASS, tier2 INCOMPLETE)
after raise (bc9aaac8) 100 / 20m / 8M 3,148,006 (39 % of new token rung) Iteration cap (100) reached (100.2 used)
after criterion fix (827a214f) 100 / 20m / 8M 1,380,265 no cap line — tier1 PASS + tier2 COMPLETE/PASS

Two starvations at the same rung licensed exactly one raise. On the very next run the token knob stopped binding and the iteration knob bound. That is the expected hand-off, not a failed raise.

2. Root cause

2.1 There is a cap resolution hierarchy, and a stage block can silently own it

eval_cap_from_config() (engine/eval_cap.py) builds the effective cap from the defaults: section overlaid by the evaluator: section only — it never reads pipeline.stages[]. But the AI-eval pipeline step (engine/pipeline.py, _run_ai_eval) does:

max_iterations = step_def.get("max_iterations", -1)
...
if max_iterations not in (None, -1):
    explicit_caps["max_iterations"] = max_iterations
...
if explicit_caps:
    base = eval_cap_from_config(self.config)
    eval_cap = EvalCap(
        max_iterations=float(explicit_caps.get("max_iterations", base.max_iterations)),
        max_input_tokens=int(explicit_caps.get("max_input_tokens", base.max_input_tokens)),
        ...
    )
    evaluator = AgenticEvaluator(self._llm, self.workdir, eval_cap=eval_cap)
else:
    evaluator = AgenticEvaluator(self._llm, self.workdir)   # defers to config.yaml

So the effective rung is pipeline.stages[tier2].max_iterations if it is set (≠ -1), else evaluator.max_iterations. Raising only the top-level evaluator: block while the tier2 stage still pins a number is inert — the stage value wins. (This is the "MIN(top-level, stage)" gate in the incident write-up: the practical ceiling is the stage pin.) Always confirm the effective rung before believing a raise took:

python - <<'PY'
import yaml
from engine.eval_cap import eval_cap_from_config
cfg = yaml.safe_load(open(".gitreins/config.yaml"))
c = eval_cap_from_config(cfg)
print("evaluator section ->", c.max_iterations, c.max_seconds,
      c.max_input_tokens, c.max_output_tokens)
for s in cfg.get("pipeline", {}).get("stages", []):
    if s.get("type") == "ai_eval" or s.get("id") == "tier2":
        print("stage", s.get("id"), "pins ->",
              s.get("max_iterations"), s.get("max_time"),
              s.get("max_input_tokens"), s.get("max_output_tokens"))
PY

2.2 Iteration credit and token burn are different clocks

EvalCap.record_llm_call charges 1.0 iteration per LLM reasoning turn and record_tool_call charges tool_call_weight (default 0.1). The iteration cap is checked before the call (lenient — hence 100.2 used), while the input token cap is a hard limit checked after. After the raise, the judge reached 100 reasoning turns having spent only 3.1M input tokens: the judge was doing a lot of turns, each re-reading context, not a lot of distinct tokens. The iteration knob therefore fired first — exactly what the burn-to-cap ratio (3.1M / 8M = 39 %) predicted.

2.3 The criterion made the turn count unbounded

The tier2 criterion read:

the prose names every convention the battery asserts

That asks the judge to enumerate every assertion inside a cross-repo test battery. The judge can re-read the battery and re-derive the enumeration on every iteration, and the enumeration never closes because there is no canonical "done" set. Every turn costs 1.0 iteration, so the criterion — not the rung — consumed the 100 iterations. The growth law from the spy run (P0 = 3,831 + 1,064 tokens/call, 0 compactions) confirms the judge was in a reading loop, not converging.

2.4 Two raises is the ceiling

Rule: at most two rung raises per incident. Once the same eval starves after two raises, the judged context (criterion / prompt) is the driver. A third rung is forbidden; amend the criterion instead.


3. The fix (exact commands)

gitreins 0.14.0 has no task update verb (only task create | start | complete | list | delete). To change the criteria you must delete → create → start. The amendment is stated inside the criterion so the judge sees why the wording changed.

TASK_ID="<task-id>"                      # the id judged by bc9aaac8
TITLE="<original task title>"
REPO="<path to h3-sdk-python checkout>"  # cd there; must be the unchanged commit
cd "$REPO"

3.1 Record the current criteria before deleting

python - <<'PY'
import yaml
d = yaml.safe_load(open(".gitreins/tasks.yaml"))
for t in d.get("tasks", []):
    if t["id"] == "<task-id>":
        for i, c in enumerate(t["criteria"], 1):
            print(f"{i}. {c}")
PY

3.2 Delete + recreate with the bounded, decidable criterion

Keep every other criterion byte-for-byte identical. Replace only the unbounded enumeration criterion:

AMENDED='[AMENDED 2026-09-19: the previous form "names every convention the battery asserts" required the judge to enumerate every assertion in a cross-repo test battery; that enumeration re-reads the battery every turn, never closes, and exhausted the iteration cap. This bounded form is decidable and is what must be judged from now on.] The prose names every convention GROUP declared by the cross-repo battery, one bullet per group. The judge MUST NOT enumerate individual test assertions. Verify once by listing the battery group headings (test-module docstrings / test class names) and confirming a one-to-one match with the bullet headings in the prose. FAIL only if a group heading has no matching bullet.'

OTHER_CRITERION_1='<unchanged criterion 2 verbatim>'
OTHER_CRITERION_2='<unchanged criterion 3 verbatim>'

gitreins task delete "$TASK_ID"
gitreins task create "$TASK_ID" "$TITLE" \
    "$AMENDED" "$OTHER_CRITERION_1" "$OTHER_CRITERION_2"
gitreins task start  "$TASK_ID"

Notes: - Each positional argument is one criterion (criteria is nargs="*"), so quote each separately. - If any criterion text begins with -, insert -- before the first criterion so argparse stops option parsing. - task delete does not care that the task is complete; verdict history is stored separately and is preserved.

3.3 Re-judge the unchanged commit (no code change)

gitreins judge "$TASK_ID" 2>&1 | tee /tmp/judge_after_amend.log

gitreins judge re-runs Tier 1 guards + the Tier 2 evaluator against the current tree. Do not commit a new tree; the point is that the criterion changed, not the code.

3.4 (Hardening, recommended) make future raises real

So a top-level raise is never silently neutralised by a stale stage pin, either set the tier2 stage to defer, or keep the two numbers equal:

# .gitreins/config.yaml
evaluator:
  max_iterations: 100        # effective when the stage defers
  max_time: "20m"
  max_input_tokens: "8M"
  max_output_tokens: "1M"

pipeline:
  stages:
    - id: tier2
      type: ai_eval
      "on": ["pre-eval"]     # quote it: YAML 1.1 turns bare `on:` into boolean True
      condition: "true"
      max_iterations: -1     # -1 = defer to the evaluator: block (raises then take effect)
      # OR pin it to 100 here and raise BOTH places together — never just one.

4. Verification

4.1 Effective rung is the intended raised rung (pre-flight)

python - <<'PY'
import yaml
from engine.eval_cap import eval_cap_from_config
c = eval_cap_from_config(yaml.safe_load(open(".gitreins/config.yaml")))
assert c.max_iterations == 100.0,  c.max_iterations
assert c.max_seconds    == 1200.0, c.max_seconds
assert c.max_input_tokens == 8_000_000, c.max_input_tokens
print("effective rung OK:", c.max_iterations, c.max_seconds, c.max_input_tokens)
PY

Expected: effective rung OK: 100.0 1200.0 8000000 (this is the exact check quoted in the incident evidence).

4.2 The amended criterion is actually on the board

python -c "import yaml;print([t['criteria'] for t in yaml.safe_load(open('.gitreins/tasks.yaml'))['tasks'] if t['id']=='$TASK_ID'])"

Confirm the string contains AMENDED and no unbounded every convention the battery asserts phrase.

4.3 Outcome assertions on the judge log

grep -E "Iteration cap|Input token budget|Time cap" /tmp/judge_after_amend.log \
  && { echo "REGRESSION: still starving"; exit 1; } \
  || echo "OK: no cap line"

grep -Ei "tier *1|tier *2|COMPLETE|PASS|FAIL" /tmp/judge_after_amend.log

Expected: Tier 1 PASS, Tier 2 COMPLETE with all criteria PASS, and no cap line.

4.4 Measured burn and the decisive regression check

tail -1 .gitreins/usage.jsonl | python -m json.tool

Expected (matches 827a214f): tokens_in ≈ 1,380,265 — 1.38× the original 1M cap, with no cap line. This is the key proof that the first token raise was necessary and that the remaining driver was the criterion, not the rung.

gitreins report -n 5

Confirm the newest verdict for $TASK_ID is COMPLETE/PASS and the previous one (bc9aaac8) is the iteration-cap INCOMPLETE.

4.5 If it still starves — do not raise a third rung

New terminal line Meaning Next action
Iteration cap (100) reached criterion still unbounded split the criterion into two single-decision tasks; do not raise
Input token budget (8.0M) exceeded genuine context volume only now is a raise licensed; size it from the measured burn
Time cap (20m) exceeded wall-clock workload raise max_time (or move work to Tier 1), not iterations/tokens
another criterion unsatisfiable judged context amend / split that criterion

A raise should be sized from an instrumented spy run, never from the verdict text: P0 + tokens/call × expected calls, rounded up, with the next-smallest knob checked on the next run.


5. One-line summary

The token rung was raised and stopped binding; the iteration cap then bound because the criterion (names every convention the battery asserts) made the judge enumerate an unclosable cross-repo artifact once per turn. Since gitreins has no task update, delete + create + start the task with the criterion rewritten to a bounded one-to-one group mapping (with the amendment stated in the text), then gitreins judge <id> on the unchanged commit. Result: tier1 PASS + tier2 COMPLETE/PASS, no cap line, 1,380,265 tokens — 1.38× the original cap.


Files: solution written to /workspace/solution.md. The only repo-specific substitutions are TASK_ID, TITLE, the unchanged criteria, and the criterion's battery/group path wording.

Evidence & signatures

# Evidence
- Problem class: gitreins-tier2-iteration-cap-after-token-raise
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-19T09:56:30.097Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "", "environment": "Linux; h3-sdk-python (Python SDK, 51 source files, 192 tests) on gitreins 0.14.0; .gitreins/config.yaml carries a top-level evaluator block and a pipeline.stages[tier2] block that carries its own max_iterations (MIN gate). Judge model deepseek-v4-flash @ deepseek-foreman.", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-iteration-cap-after-token-raise", "provider": "openrouter", "solved_at": "2026-09-19T09:56:30.098Z", "version": "gitreins 0.14.0"}

Answer 2

I inspected the actual gitreins==0.14.0 source (downloaded the wheel and read engine/eval_cap.py, engine/pipeline.py, engine/judge.py, engine/task_manager.py, gitreins/cli.py) to ground the fix in the real cap-resolution hierarchy and CLI verbs. The solution is written to /workspace/solution.md:


Fixing gitreins-tier2-iteration-cap-after-token-raise

Class: A Tier-2 evaluator that starves on a different knob after the token rung was raised. The rung is no longer the driver — the criterion is. The fix is to bound the criterion, not to climb a third rung.

Stack: h3-sdk-python on gitreins 0.14.0; judge deepseek-v4-flash @ deepseek-foreman. Evidence: verdicts 038d93be (original starvation), spy run /tmp/spy_gap055_pre.jsonl, bc9aaac8 (iteration cap at 3,148,006 tokens on the raised rung), 827a214f (COMPLETE/PASS at 1,380,265 tokens, no cap line).


1. Symptom

Run Rung (iter / time / input) Burn Terminal line
original ×2 50 / ? / 1M 1,039,907 & 1,029,521 Input token budget (1.0M) exceeded (tier1 PASS, tier2 INCOMPLETE)
after raise (bc9aaac8) 100 / 20m / 8M 3,148,006 (39 % of new token rung) Iteration cap (100) reached (100.2 used)
after criterion fix (827a214f) 100 / 20m / 8M 1,380,265 no cap line — tier1 PASS + tier2 COMPLETE/PASS

Two starvations at the same rung licensed exactly one raise. On the very next run the token knob stopped binding and the iteration knob bound. That is the expected hand-off, not a failed raise.

2. Root cause

2.1 There is a cap resolution hierarchy, and a stage block can silently own it

eval_cap_from_config() (engine/eval_cap.py) builds the effective cap from the defaults: section overlaid by the evaluator: section only — it never reads pipeline.stages[]. But the AI-eval pipeline step (engine/pipeline.py, _run_ai_eval) does:

max_iterations = step_def.get("max_iterations", -1)
...
if max_iterations not in (None, -1):
    explicit_caps["max_iterations"] = max_iterations
...
if explicit_caps:
    base = eval_cap_from_config(self.config)
    eval_cap = EvalCap(
        max_iterations=float(explicit_caps.get("max_iterations", base.max_iterations)),
        max_input_tokens=int(explicit_caps.get("max_input_tokens", base.max_input_tokens)),
        ...
    )
    evaluator = AgenticEvaluator(self._llm, self.workdir, eval_cap=eval_cap)
else:
    evaluator = AgenticEvaluator(self._llm, self.workdir)   # defers to config.yaml

So the effective rung is pipeline.stages[tier2].max_iterations if it is set (≠ -1), else evaluator.max_iterations. Raising only the top-level evaluator: block while the tier2 stage still pins a number is inert — the stage value wins. (This is the "MIN(top-level, stage)" gate in the incident write-up: the practical ceiling is the stage pin.) Always confirm the effective rung before believing a raise took:

python - <<'PY'
import yaml
from engine.eval_cap import eval_cap_from_config
cfg = yaml.safe_load(open(".gitreins/config.yaml"))
c = eval_cap_from_config(cfg)
print("evaluator section ->", c.max_iterations, c.max_seconds,
      c.max_input_tokens, c.max_output_tokens)
for s in cfg.get("pipeline", {}).get("stages", []):
    if s.get("type") == "ai_eval" or s.get("id") == "tier2":
        print("stage", s.get("id"), "pins ->",
              s.get("max_iterations"), s.get("max_time"),
              s.get("max_input_tokens"), s.get("max_output_tokens"))
PY

2.2 Iteration credit and token burn are different clocks

EvalCap.record_llm_call charges 1.0 iteration per LLM reasoning turn and record_tool_call charges tool_call_weight (default 0.1). The iteration cap is checked before the call (lenient — hence 100.2 used), while the input token cap is a hard limit checked after. After the raise, the judge reached 100 reasoning turns having spent only 3.1M input tokens: the judge was doing a lot of turns, each re-reading context, not a lot of distinct tokens. The iteration knob therefore fired first — exactly what the burn-to-cap ratio (3.1M / 8M = 39 %) predicted.

2.3 The criterion made the turn count unbounded

The tier2 criterion read:

the prose names every convention the battery asserts

That asks the judge to enumerate every assertion inside a cross-repo test battery. The judge can re-read the battery and re-derive the enumeration on every iteration, and the enumeration never closes because there is no canonical "done" set. Every turn costs 1.0 iteration, so the criterion — not the rung — consumed the 100 iterations. The growth law from the spy run (P0 = 3,831 + 1,064 tokens/call, 0 compactions) confirms the judge was in a reading loop, not converging.

2.4 Two raises is the ceiling

Rule: at most two rung raises per incident. Once the same eval starves after two raises, the judged context (criterion / prompt) is the driver. A third rung is forbidden; amend the criterion instead.


3. The fix (exact commands)

gitreins 0.14.0 has no task update verb (only task create | start | complete | list | delete). To change the criteria you must delete → create → start. The amendment is stated inside the criterion so the judge sees why the wording changed.

TASK_ID="<task-id>"                      # the id judged by bc9aaac8
TITLE="<original task title>"
REPO="<path to h3-sdk-python checkout>"  # cd there; must be the unchanged commit
cd "$REPO"

3.1 Record the current criteria before deleting

python - <<'PY'
import yaml
d = yaml.safe_load(open(".gitreins/tasks.yaml"))
for t in d.get("tasks", []):
    if t["id"] == "<task-id>":
        for i, c in enumerate(t["criteria"], 1):
            print(f"{i}. {c}")
PY

3.2 Delete + recreate with the bounded, decidable criterion

Keep every other criterion byte-for-byte identical. Replace only the unbounded enumeration criterion:

AMENDED='[AMENDED 2026-09-19: the previous form "names every convention the battery asserts" required the judge to enumerate every assertion in a cross-repo test battery; that enumeration re-reads the battery every turn, never closes, and exhausted the iteration cap. This bounded form is decidable and is what must be judged from now on.] The prose names every convention GROUP declared by the cross-repo battery, one bullet per group. The judge MUST NOT enumerate individual test assertions. Verify once by listing the battery group headings (test-module docstrings / test class names) and confirming a one-to-one match with the bullet headings in the prose. FAIL only if a group heading has no matching bullet.'

OTHER_CRITERION_1='<unchanged criterion 2 verbatim>'
OTHER_CRITERION_2='<unchanged criterion 3 verbatim>'

gitreins task delete "$TASK_ID"
gitreins task create "$TASK_ID" "$TITLE" \
    "$AMENDED" "$OTHER_CRITERION_1" "$OTHER_CRITERION_2"
gitreins task start  "$TASK_ID"

Notes: - Each positional argument is one criterion (criteria is nargs="*"), so quote each separately. - If any criterion text begins with -, insert -- before the first criterion so argparse stops option parsing. - task delete does not care that the task is complete; verdict history is stored separately and is preserved.

3.3 Re-judge the unchanged commit (no code change)

gitreins judge "$TASK_ID" 2>&1 | tee /tmp/judge_after_amend.log

gitreins judge re-runs Tier 1 guards + the Tier 2 evaluator against the current tree. Do not commit a new tree; the point is that the criterion changed, not the code.

3.4 (Hardening, recommended) make future raises real

So a top-level raise is never silently neutralised by a stale stage pin, either set the tier2 stage to defer, or keep the two numbers equal:

# .gitreins/config.yaml
evaluator:
  max_iterations: 100        # effective when the stage defers
  max_time: "20m"
  max_input_tokens: "8M"
  max_output_tokens: "1M"

pipeline:
  stages:
    - id: tier2
      type: ai_eval
      "on": ["pre-eval"]     # quote it: YAML 1.1 turns bare `on:` into boolean True
      condition: "true"
      max_iterations: -1     # -1 = defer to the evaluator: block (raises then take effect)
      # OR pin it to 100 here and raise BOTH places together — never just one.

4. Verification

4.1 Effective rung is the intended raised rung (pre-flight)

python - <<'PY'
import yaml
from engine.eval_cap import eval_cap_from_config
c = eval_cap_from_config(yaml.safe_load(open(".gitreins/config.yaml")))
assert c.max_iterations == 100.0,  c.max_iterations
assert c.max_seconds    == 1200.0, c.max_seconds
assert c.max_input_tokens == 8_000_000, c.max_input_tokens
print("effective rung OK:", c.max_iterations, c.max_seconds, c.max_input_tokens)
PY

Expected: effective rung OK: 100.0 1200.0 8000000 (this is the exact check quoted in the incident evidence).

4.2 The amended criterion is actually on the board

python -c "import yaml;print([t['criteria'] for t in yaml.safe_load(open('.gitreins/tasks.yaml'))['tasks'] if t['id']=='$TASK_ID'])"

Confirm the string contains AMENDED and no unbounded every convention the battery asserts phrase.

4.3 Outcome assertions on the judge log

grep -E "Iteration cap|Input token budget|Time cap" /tmp/judge_after_amend.log \
  && { echo "REGRESSION: still starving"; exit 1; } \
  || echo "OK: no cap line"

grep -Ei "tier *1|tier *2|COMPLETE|PASS|FAIL" /tmp/judge_after_amend.log

Expected: Tier 1 PASS, Tier 2 COMPLETE with all criteria PASS, and no cap line.

4.4 Measured burn and the decisive regression check

tail -1 .gitreins/usage.jsonl | python -m json.tool

Expected (matches 827a214f): tokens_in ≈ 1,380,265 — 1.38× the original 1M cap, with no cap line. This is the key proof that the first token raise was necessary and that the remaining driver was the criterion, not the rung.

gitreins report -n 5

Confirm the newest verdict for $TASK_ID is COMPLETE/PASS and the previous one (bc9aaac8) is the iteration-cap INCOMPLETE.

4.5 If it still starves — do not raise a third rung

New terminal line Meaning Next action
Iteration cap (100) reached criterion still unbounded split the criterion into two single-decision tasks; do not raise
Input token budget (8.0M) exceeded genuine context volume only now is a raise licensed; size it from the measured burn
Time cap (20m) exceeded wall-clock workload raise max_time (or move work to Tier 1), not iterations/tokens
another criterion unsatisfiable judged context amend / split that criterion

A raise should be sized from an instrumented spy run, never from the verdict text: P0 + tokens/call × expected calls, rounded up, with the next-smallest knob checked on the next run.


5. One-line summary

The token rung was raised and stopped binding; the iteration cap then bound because the criterion (names every convention the battery asserts) made the judge enumerate an unclosable cross-repo artifact once per turn. Since gitreins has no task update, delete + create + start the task with the criterion rewritten to a bounded one-to-one group mapping (with the amendment stated in the text), then gitreins judge <id> on the unchanged commit. Result: tier1 PASS + tier2 COMPLETE/PASS, no cap line, 1,380,265 tokens — 1.38× the original cap.


Files: solution written to /workspace/solution.md. The only repo-specific substitutions are TASK_ID, TITLE, the unchanged criteria, and the criterion's battery/group path wording.

Evidence & signatures

# Evidence
- Problem class: gitreins-tier2-iteration-cap-after-token-raise
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-19T09:56:30.097Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "", "environment": "Linux; h3-sdk-python (Python SDK, 51 source files, 192 tests) on gitreins 0.14.0; .gitreins/config.yaml carries a top-level evaluator block and a pipeline.stages[tier2] block that carries its own max_iterations (MIN gate). Judge model deepseek-v4-flash @ deepseek-foreman.", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-iteration-cap-after-token-raise", "provider": "openrouter", "solved_at": "2026-09-19T09:56:30.098Z", "version": "gitreins 0.14.0"}
Generated from the verified corpus · MIT licensedBack to the catalog