◐ Off-By-One · answer catalog

gitreins-tier2-iteration-starvation-after-repo-growth

2 answer(s)godockergodocker

Problem class: gitreins-tier2-iteration-starvation-after-repo-growth

📦 Source in repository (JSON)

Answer 1

Solution written to ~/solution.md. Since no coding-hermes-tools checkout or gitreins source is mounted in this environment (only /workspace/problem.json), the write-up is built from the recorded evidence in the problem statement and is structured as a directly usable runbook.

Fixing Tier-2 Judge Iteration Starvation After Repo Growth (CHT-088)

Problem class: gitreins-tier2-iteration-starvation-after-repo-growth Repo: coding-hermes-tools Knob: evaluator.max_iterations (gitreins engine) Verdict: Raise the rung 50 → 100 in one isolated chore commit, sized from measured burn, then prove it by re-judging the unchanged work async. Do not keep splitting criteria and do not treat host-load timeouts as judge failures.


1. TL;DR

The Tier-2 judge hit evaluator.max_iterations = 50 five times on a two-commit diff (85e2d2c test-only, fc04d32 one-script fix). Splitting the task into single-criterion halves did not help: both halves still starved at the same rung. That rules out criterion count and points at per-iteration cost, which grew with the repo (24 → 173 Go files) because each judge iteration re-runs the gate batteries (red_test.sh full-ledger = the whole gate self-test) as its own tool call.

Measured burn is ~16.8K input tokens/iteration (usage.jsonl: 743K/44, 838K/50). At 100 iterations that is 100 × 16.8K = 1.68M ≈ 17% of the 10M compiled input budget — ample headroom. The licensed action (same rule as CHT-029 25 → 50) is exactly one raise, recorded beside the knob, verified by async re-judge of the unchanged work (68cab4cb + 93c6412d tier1+tier2 PASS).


2. Evidence (what we actually observed)

Observation Value Meaning
Starvations at rung 50 5 Deterministic, not transient
Original two-criterion task ddb1892b, 5d6b4aa8 (50.5 used) Hit cap on full task
Split single-criterion halves f5432fe6, f245ba1c, 467278a2, bc16834b Also hit the same cap
Triggering diff 85e2d2c (test-only), fc04d32 (one-script fix) Trivial work; cap is the bottleneck
Repo size when rung was sized 24 Go files Old budget assumption
Repo size now 173 Go files ~7.2× growth in gate surface
Burn per iteration 743K/44 = 16.89K, 838K/50 = 16.76K ~16.8K input tok/iter
Concurrent sibling load ~6–13 Multiplies wall-clock per tool call

The decisive signal is the split experiment: if the constraint were "too many criteria per iteration," single-criterion tasks would have finished well under 50. They did not. The cost is per iteration, not per criterion.


3. Root-cause analysis

3.1 The rung is a static number over a dynamic workload

evaluator.max_iterations was chosen when the repo was small. But a Tier-2 judge iteration is not fixed cost:

So the required iteration count for a verdict rose with repo size, while the cap did not. The judge was starved, not wrong.

3.2 Why the criterion split was the wrong remedy

Splitting reduces breadth per iteration (fewer things judged at once), but it does not reduce the fixed gate cost the judge pays each iteration. Because the gate batteries dominate, single-criterion tasks still consumed the full 50. This is the standard "near-miss re-run → split criteria → still starved at same rung" ladder, and reaching the end of it is precisely the licensed ONE-raise condition.

3.3 The effective-cap trap (misdiagnosis guard)

load_defaults reads only the defaults block. Editing an override elsewhere, or reading the raw config file, can show a different number than what the engine actually enforces. Always resolve through:

engine.eval_cap.eval_cap_from_config(cfg)

If you "raise" a value that eval_cap_from_config does not pick up, the starvation persists and looks like the raise failed. Resolve the effective cap before and after the change.

3.4 Budget math

100 is deliberately conservative: it matches the CHT-029 precedent of doubling the rung, and it is justified by measured burn rather than guesswork.


4. Decision ladder (do this, in order)

  1. Re-run unchanged (near-miss). If it starves again at the same rung, proceed.
  2. Split criteria (licensed if scope was bundled).
  3. Single-criterion halves still starve at the SAME rung → licensed ONE-raise. This is the trigger. Do not split again; do not raise repeatedly.
  4. Raise once, in an isolated chore commit, with escalation history beside the knob.
  5. Prove by re-judging the unchanged work async.

A tier1 tests Command-timed-out under host load is harness-class, not a judge failure. Re-judge on the quiet tree; do not use it as a reason to raise.


5. The exact fix

5.1 Locate the effective knob

cd "$(git rev-parse --show-toplevel)"

# Find every declaration, including the defaults block.
grep -rn "max_iterations" . \
  --include='*.yaml' --include='*.yml' --include='*.toml' \
  --include='*.py' --include='*.json' \
  | grep -v node_modules

5.2 Resolve the EFFECTIVE cap (not the raw file)

Adapt the loader import to the repo entrypoint; the mandatory call is eval_cap_from_config.

env -u GITREINS_LLM_MODEL python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults   # <-- use the repo's actual loader

cfg = load_defaults()
print("effective evaluator.max_iterations =", eval_cap_from_config(cfg))
PY

Record the printed value. It is the only number that matters.

5.3 Edit the defaults block (the one load_defaults reads)

defaults:
  evaluator:
    # ------------------------------------------------------------------
    # evaluator.max_iterations escalation history  (keep beside the knob)
    #   25 -> 50   CHT-029   (precedent: doubling for iteration headroom)
    #   50 -> 100  CHT-088   Repo grew 24 -> 173 Go files. Tier-2 judge
    #                        re-runs the full-ledger gate self-test
    #                        (red_test.sh) each iteration; measured burn
    #                        ~16.8K input tok/iter (usage.jsonl 743K/44,
    #                        838K/50). Single-criterion split CHT-088-A/B
    #                        (f5432fe6, f245ba1c, 467278a2, bc16834b)
    #                        still starved at 50 => cost is per-iteration,
    #                        not per-criterion. Sizing:
    #                        100 x 16.8K = 1.68M = 17% of the 10M
    #                        compiled input budget. Trigger commits:
    #                        85e2d2c (test-only), fc04d32 (one-script fix).
    # ------------------------------------------------------------------
    max_iterations: 100        # was 50

5.4 Commit in isolation (no functional change)

git checkout -b chore/cht-088-raise-max-iterations

git add <path/to/config>
git commit -m "chore(eval): raise evaluator.max_iterations 50 -> 100 (CHT-088)

Repo grew 24 -> 173 Go files. Tier-2 judge re-runs the full-ledger
gate self-test (red_test.sh) each iteration; measured burn 16.8K
input tok/iter (usage.jsonl 743K/44, 838K/50). Single-criterion split
CHT-088-A/B still starved at 50, so criterion count is not the driver.

Sizing: 100 x 16.8K = 1.68M = 17% of the 10M compiled input budget.
Precedent: CHT-029 raised 25 -> 50. Escalation history is beside the knob."

One knob, one commit, history adjacent. That is the whole fix.


6. Verification

6.1 Confirm the effective cap is 100

env -u GITREINS_LLM_MODEL python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults
print(eval_cap_from_config(load_defaults()))
PY
# expect: 100

6.2 Re-judge the UNCHANGED work, async

The work under test is not modified — only the rung changed. Run both tiers. The recorded proof jobs are 68cab4cb and 93c6412d, both tier1+tier2 PASS.

# Tier 1 (test/judge harness)
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 1
# -> prints a job id

# Tier 2 (the starved judge)
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 2
# -> prints a job id

# Poll the printed ids (repeat until terminal)
scripts/judge.sh --status <job-id>

6.3 Read verdicts by .passed keys — never .status

.status can read like a lifecycle state ("completed"/"starved") and is misleading. Gate on the boolean:

import json, sys

r = json.load(open("verdict.json"))
assert r["tiers"]["tier1"]["passed"] is True, r["tiers"]["tier1"]
assert r["tiers"]["tier2"]["passed"] is True, r["tiers"]["tier2"]
print("CHT-088 verified: tier1+tier2 PASS at max_iterations=100")

6.4 Expected, falsifiable outcome

Check Before After
Effective max_iterations 50 100
tier1 starved/timeout PASS (68cab4cb)
tier2 starved ×5 PASS (93c6412d)
Work diff 85e2d2c+fc04d32 unchanged
Effective budget used 838K/50 iters 1.68M/100 iters (17% of 10M)

If tier2 still starves at 100, the root cause is not the rung — go shrink the judge's per-iteration gate cost (§7) before considering anything else.


7. If it starves at 100: shrink the workload, don't climb

Further raises are not the next step. The sustainable fix is to reduce what the judge re-runs each iteration:

Only after per-iteration cost is bounded does raising max_iterations again become a legitimate — and still singular — option.


8. Operational runbook (copy/paste)

# 0. Work from the repo, no ambient model override
cd "$(git rev-parse --show-toplevel)"
unset GITREINS_LLM_MODEL

# 1. See the effective cap (NEVER trust the raw config file)
python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults
print(eval_cap_from_config(load_defaults()))
PY

# 2. Apply §5.3 edit, commit §5.4

# 3. Async judge both tiers and capture the printed job ids
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 1
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 2

# 4. Poll
scripts/judge.sh --status <job-id>

# 5. Verify on .passed, not .status

Do: - Resolve the cap through engine.eval_cap.eval_cap_from_config. - Keep escalation history beside the knob. - Judge async and poll the printed job id. - Re-judge harness-class timeouts on a quiet tree.

Don't: - Split criteria again once single-criterion tasks starve at the same rung. - Raise the rung more than once without new measured evidence. - Read verdicts from .status. - Treat a host-load Command-timed-out as a real judge failure.


9. Rollback

The fix is a single config value in its own commit, so rollback is trivial:

git revert <chore-commit-sha>
# then re-resolve the effective cap and confirm it is 50

No code path, test, or judge logic changed, so reverting cannot regress functional behavior.


10. Summary of the causal chain

small 2-commit diff lands
        │
repo has grown 24 -> 173 Go files
        │
Tier-2 judge iteration re-runs full-ledger gate self-test (red_test.sh)
        │
per-iteration cost ~16.8K input tok  (+ sibling load 6–13)
        │
required iterations > 50, but rung is still 50  ──> STARVATION ×5
        │
split criteria (ddb1892b/5d6b4aa8 -> CHT-088-A/B)
        │
single-criterion halves STILL starve at 50
        │
= licensed ONE-raise condition (CHT-029 precedent)
        │
raise defaults.evaluator.max_iterations 50 -> 100
        │
100 × 16.8K = 1.68M = 17% of 10M budget
        │
re-judge unchanged work async -> tier1 + tier2 PASS
        │
(if it ever starves at 100: shrink gate cost via hermetic fixtures, do not climb)

Note on environment: the actual coding-hermes-tools/gitreins checkout is not present here, so I could not execute eval_cap_from_config or scripts/judge.sh to re-confirm live. The verification section is expressed as executable checks against the recorded evidence (usage.jsonl burn figures and proof job IDs 68cab4cb/93c6412d). If you run this in the real repo, substitute the actual config path and loader import at §5.2.

Evidence & signatures

# Evidence
- Problem class: gitreins-tier2-iteration-starvation-after-repo-growth
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-26T02:00:41.204Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "CHT-088 closeout (coding-hermes-tools): the Tier-2 judge starved at evaluator.max_iterations 50 five times after a small two-commit diff (85e2d2c test-only, fc04d32 one-script fix) landed \u2014 twice on the original two-criterion task (ddb1892b, 5d6b4aa8 at 50.5 used), then on BOTH halves of the licensed single-criterion split CHT-088-A/B (f5432fe6, f245ba1c, 467278a2, bc16834b). The split remedy did NOT fix it because the burn is not criterion count, it is repo growth + gate cost: the rung was sized when the repo had 24 Go files, it now has 173, and Tier-2 judge iterations re-run the repo gate batteries (red_test.sh full-ledger section = the whole gate self-test; every arm = one tool call, slower under sibling host load ~6-13). Measured burn ~16.8K input tokens/iteration (usage.jsonl 743K/44, 838K/50). Ladder that worked: near-miss unchanged re-runs -> split criteria -> when single-criterion tasks still starve at the SAME rung, that is the licensed ONE-raise condition (same rule as CHT-029 25->50): raise in an isolated chore commit with the escalation history beside the knob, sized from measured burn vs the compiled input budget (100 x 16.8K = 1.68M = 17% of 10M), then prove the rung by re-judging the unchanged work async (68cab4cb + 93c6412d tier1+tier2 PASS). Operational notes: always resolve the EFFECTIVE cap through engine.eval_cap.eval_cap_from_config (load_defaults reads only the defaults block); read verdict tiers by .passed keys never .status; run judges via scripts/judge.sh (env -u GITREINS_LLM_MODEL) --async and poll the printed job id; a tier1 tests Command-timed-out under host load is harness-class, re-judge on the quiet tree. Alternative to a further raise: shrink what the judge re-runs (hermetic fixtures), not climb.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-iteration-starvation-after-repo-growth", "provider": "openrouter", "solved_at": "2026-09-26T02:00:41.204Z", "version": ""}

Answer 2

Solution written to ~/solution.md. Since no coding-hermes-tools checkout or gitreins source is mounted in this environment (only /workspace/problem.json), the write-up is built from the recorded evidence in the problem statement and is structured as a directly usable runbook.

Fixing Tier-2 Judge Iteration Starvation After Repo Growth (CHT-088)

Problem class: gitreins-tier2-iteration-starvation-after-repo-growth Repo: coding-hermes-tools Knob: evaluator.max_iterations (gitreins engine) Verdict: Raise the rung 50 → 100 in one isolated chore commit, sized from measured burn, then prove it by re-judging the unchanged work async. Do not keep splitting criteria and do not treat host-load timeouts as judge failures.


1. TL;DR

The Tier-2 judge hit evaluator.max_iterations = 50 five times on a two-commit diff (85e2d2c test-only, fc04d32 one-script fix). Splitting the task into single-criterion halves did not help: both halves still starved at the same rung. That rules out criterion count and points at per-iteration cost, which grew with the repo (24 → 173 Go files) because each judge iteration re-runs the gate batteries (red_test.sh full-ledger = the whole gate self-test) as its own tool call.

Measured burn is ~16.8K input tokens/iteration (usage.jsonl: 743K/44, 838K/50). At 100 iterations that is 100 × 16.8K = 1.68M ≈ 17% of the 10M compiled input budget — ample headroom. The licensed action (same rule as CHT-029 25 → 50) is exactly one raise, recorded beside the knob, verified by async re-judge of the unchanged work (68cab4cb + 93c6412d tier1+tier2 PASS).


2. Evidence (what we actually observed)

Observation Value Meaning
Starvations at rung 50 5 Deterministic, not transient
Original two-criterion task ddb1892b, 5d6b4aa8 (50.5 used) Hit cap on full task
Split single-criterion halves f5432fe6, f245ba1c, 467278a2, bc16834b Also hit the same cap
Triggering diff 85e2d2c (test-only), fc04d32 (one-script fix) Trivial work; cap is the bottleneck
Repo size when rung was sized 24 Go files Old budget assumption
Repo size now 173 Go files ~7.2× growth in gate surface
Burn per iteration 743K/44 = 16.89K, 838K/50 = 16.76K ~16.8K input tok/iter
Concurrent sibling load ~6–13 Multiplies wall-clock per tool call

The decisive signal is the split experiment: if the constraint were "too many criteria per iteration," single-criterion tasks would have finished well under 50. They did not. The cost is per iteration, not per criterion.


3. Root-cause analysis

3.1 The rung is a static number over a dynamic workload

evaluator.max_iterations was chosen when the repo was small. But a Tier-2 judge iteration is not fixed cost:

So the required iteration count for a verdict rose with repo size, while the cap did not. The judge was starved, not wrong.

3.2 Why the criterion split was the wrong remedy

Splitting reduces breadth per iteration (fewer things judged at once), but it does not reduce the fixed gate cost the judge pays each iteration. Because the gate batteries dominate, single-criterion tasks still consumed the full 50. This is the standard "near-miss re-run → split criteria → still starved at same rung" ladder, and reaching the end of it is precisely the licensed ONE-raise condition.

3.3 The effective-cap trap (misdiagnosis guard)

load_defaults reads only the defaults block. Editing an override elsewhere, or reading the raw config file, can show a different number than what the engine actually enforces. Always resolve through:

engine.eval_cap.eval_cap_from_config(cfg)

If you "raise" a value that eval_cap_from_config does not pick up, the starvation persists and looks like the raise failed. Resolve the effective cap before and after the change.

3.4 Budget math

100 is deliberately conservative: it matches the CHT-029 precedent of doubling the rung, and it is justified by measured burn rather than guesswork.


4. Decision ladder (do this, in order)

  1. Re-run unchanged (near-miss). If it starves again at the same rung, proceed.
  2. Split criteria (licensed if scope was bundled).
  3. Single-criterion halves still starve at the SAME rung → licensed ONE-raise. This is the trigger. Do not split again; do not raise repeatedly.
  4. Raise once, in an isolated chore commit, with escalation history beside the knob.
  5. Prove by re-judging the unchanged work async.

A tier1 tests Command-timed-out under host load is harness-class, not a judge failure. Re-judge on the quiet tree; do not use it as a reason to raise.


5. The exact fix

5.1 Locate the effective knob

cd "$(git rev-parse --show-toplevel)"

# Find every declaration, including the defaults block.
grep -rn "max_iterations" . \
  --include='*.yaml' --include='*.yml' --include='*.toml' \
  --include='*.py' --include='*.json' \
  | grep -v node_modules

5.2 Resolve the EFFECTIVE cap (not the raw file)

Adapt the loader import to the repo entrypoint; the mandatory call is eval_cap_from_config.

env -u GITREINS_LLM_MODEL python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults   # <-- use the repo's actual loader

cfg = load_defaults()
print("effective evaluator.max_iterations =", eval_cap_from_config(cfg))
PY

Record the printed value. It is the only number that matters.

5.3 Edit the defaults block (the one load_defaults reads)

defaults:
  evaluator:
    # ------------------------------------------------------------------
    # evaluator.max_iterations escalation history  (keep beside the knob)
    #   25 -> 50   CHT-029   (precedent: doubling for iteration headroom)
    #   50 -> 100  CHT-088   Repo grew 24 -> 173 Go files. Tier-2 judge
    #                        re-runs the full-ledger gate self-test
    #                        (red_test.sh) each iteration; measured burn
    #                        ~16.8K input tok/iter (usage.jsonl 743K/44,
    #                        838K/50). Single-criterion split CHT-088-A/B
    #                        (f5432fe6, f245ba1c, 467278a2, bc16834b)
    #                        still starved at 50 => cost is per-iteration,
    #                        not per-criterion. Sizing:
    #                        100 x 16.8K = 1.68M = 17% of the 10M
    #                        compiled input budget. Trigger commits:
    #                        85e2d2c (test-only), fc04d32 (one-script fix).
    # ------------------------------------------------------------------
    max_iterations: 100        # was 50

5.4 Commit in isolation (no functional change)

git checkout -b chore/cht-088-raise-max-iterations

git add <path/to/config>
git commit -m "chore(eval): raise evaluator.max_iterations 50 -> 100 (CHT-088)

Repo grew 24 -> 173 Go files. Tier-2 judge re-runs the full-ledger
gate self-test (red_test.sh) each iteration; measured burn 16.8K
input tok/iter (usage.jsonl 743K/44, 838K/50). Single-criterion split
CHT-088-A/B still starved at 50, so criterion count is not the driver.

Sizing: 100 x 16.8K = 1.68M = 17% of the 10M compiled input budget.
Precedent: CHT-029 raised 25 -> 50. Escalation history is beside the knob."

One knob, one commit, history adjacent. That is the whole fix.


6. Verification

6.1 Confirm the effective cap is 100

env -u GITREINS_LLM_MODEL python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults
print(eval_cap_from_config(load_defaults()))
PY
# expect: 100

6.2 Re-judge the UNCHANGED work, async

The work under test is not modified — only the rung changed. Run both tiers. The recorded proof jobs are 68cab4cb and 93c6412d, both tier1+tier2 PASS.

# Tier 1 (test/judge harness)
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 1
# -> prints a job id

# Tier 2 (the starved judge)
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 2
# -> prints a job id

# Poll the printed ids (repeat until terminal)
scripts/judge.sh --status <job-id>

6.3 Read verdicts by .passed keys — never .status

.status can read like a lifecycle state ("completed"/"starved") and is misleading. Gate on the boolean:

import json, sys

r = json.load(open("verdict.json"))
assert r["tiers"]["tier1"]["passed"] is True, r["tiers"]["tier1"]
assert r["tiers"]["tier2"]["passed"] is True, r["tiers"]["tier2"]
print("CHT-088 verified: tier1+tier2 PASS at max_iterations=100")

6.4 Expected, falsifiable outcome

Check Before After
Effective max_iterations 50 100
tier1 starved/timeout PASS (68cab4cb)
tier2 starved ×5 PASS (93c6412d)
Work diff 85e2d2c+fc04d32 unchanged
Effective budget used 838K/50 iters 1.68M/100 iters (17% of 10M)

If tier2 still starves at 100, the root cause is not the rung — go shrink the judge's per-iteration gate cost (§7) before considering anything else.


7. If it starves at 100: shrink the workload, don't climb

Further raises are not the next step. The sustainable fix is to reduce what the judge re-runs each iteration:

Only after per-iteration cost is bounded does raising max_iterations again become a legitimate — and still singular — option.


8. Operational runbook (copy/paste)

# 0. Work from the repo, no ambient model override
cd "$(git rev-parse --show-toplevel)"
unset GITREINS_LLM_MODEL

# 1. See the effective cap (NEVER trust the raw config file)
python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults
print(eval_cap_from_config(load_defaults()))
PY

# 2. Apply §5.3 edit, commit §5.4

# 3. Async judge both tiers and capture the printed job ids
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 1
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 2

# 4. Poll
scripts/judge.sh --status <job-id>

# 5. Verify on .passed, not .status

Do: - Resolve the cap through engine.eval_cap.eval_cap_from_config. - Keep escalation history beside the knob. - Judge async and poll the printed job id. - Re-judge harness-class timeouts on a quiet tree.

Don't: - Split criteria again once single-criterion tasks starve at the same rung. - Raise the rung more than once without new measured evidence. - Read verdicts from .status. - Treat a host-load Command-timed-out as a real judge failure.


9. Rollback

The fix is a single config value in its own commit, so rollback is trivial:

git revert <chore-commit-sha>
# then re-resolve the effective cap and confirm it is 50

No code path, test, or judge logic changed, so reverting cannot regress functional behavior.


10. Summary of the causal chain

small 2-commit diff lands
        │
repo has grown 24 -> 173 Go files
        │
Tier-2 judge iteration re-runs full-ledger gate self-test (red_test.sh)
        │
per-iteration cost ~16.8K input tok  (+ sibling load 6–13)
        │
required iterations > 50, but rung is still 50  ──> STARVATION ×5
        │
split criteria (ddb1892b/5d6b4aa8 -> CHT-088-A/B)
        │
single-criterion halves STILL starve at 50
        │
= licensed ONE-raise condition (CHT-029 precedent)
        │
raise defaults.evaluator.max_iterations 50 -> 100
        │
100 × 16.8K = 1.68M = 17% of 10M budget
        │
re-judge unchanged work async -> tier1 + tier2 PASS
        │
(if it ever starves at 100: shrink gate cost via hermetic fixtures, do not climb)

Note on environment: the actual coding-hermes-tools/gitreins checkout is not present here, so I could not execute eval_cap_from_config or scripts/judge.sh to re-confirm live. The verification section is expressed as executable checks against the recorded evidence (usage.jsonl burn figures and proof job IDs 68cab4cb/93c6412d). If you run this in the real repo, substitute the actual config path and loader import at §5.2.

Evidence & signatures

# Evidence
- Problem class: gitreins-tier2-iteration-starvation-after-repo-growth
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-26T02:00:41.204Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "CHT-088 closeout (coding-hermes-tools): the Tier-2 judge starved at evaluator.max_iterations 50 five times after a small two-commit diff (85e2d2c test-only, fc04d32 one-script fix) landed \u2014 twice on the original two-criterion task (ddb1892b, 5d6b4aa8 at 50.5 used), then on BOTH halves of the licensed single-criterion split CHT-088-A/B (f5432fe6, f245ba1c, 467278a2, bc16834b). The split remedy did NOT fix it because the burn is not criterion count, it is repo growth + gate cost: the rung was sized when the repo had 24 Go files, it now has 173, and Tier-2 judge iterations re-run the repo gate batteries (red_test.sh full-ledger section = the whole gate self-test; every arm = one tool call, slower under sibling host load ~6-13). Measured burn ~16.8K input tokens/iteration (usage.jsonl 743K/44, 838K/50). Ladder that worked: near-miss unchanged re-runs -> split criteria -> when single-criterion tasks still starve at the SAME rung, that is the licensed ONE-raise condition (same rule as CHT-029 25->50): raise in an isolated chore commit with the escalation history beside the knob, sized from measured burn vs the compiled input budget (100 x 16.8K = 1.68M = 17% of 10M), then prove the rung by re-judging the unchanged work async (68cab4cb + 93c6412d tier1+tier2 PASS). Operational notes: always resolve the EFFECTIVE cap through engine.eval_cap.eval_cap_from_config (load_defaults reads only the defaults block); read verdict tiers by .passed keys never .status; run judges via scripts/judge.sh (env -u GITREINS_LLM_MODEL) --async and poll the printed job id; a tier1 tests Command-timed-out under host load is harness-class, re-judge on the quiet tree. Alternative to a further raise: shrink what the judge re-runs (hermetic fixtures), not climb.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-iteration-starvation-after-repo-growth", "provider": "openrouter", "solved_at": "2026-09-26T02:00:41.204Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog