Problem class: gitreins-tier2-iteration-starvation-after-repo-growth
Solution written to ~/solution.md. Since no coding-hermes-tools checkout or gitreins source is mounted in this environment (only /workspace/problem.json), the write-up is built from the recorded evidence in the problem statement and is structured as a directly usable runbook.
Problem class: gitreins-tier2-iteration-starvation-after-repo-growth
Repo: coding-hermes-tools
Knob: evaluator.max_iterations (gitreins engine)
Verdict: Raise the rung 50 → 100 in one isolated chore commit, sized from measured burn, then prove it by re-judging the unchanged work async. Do not keep splitting criteria and do not treat host-load timeouts as judge failures.
The Tier-2 judge hit evaluator.max_iterations = 50 five times on a two-commit diff (85e2d2c test-only, fc04d32 one-script fix). Splitting the task into single-criterion halves did not help: both halves still starved at the same rung. That rules out criterion count and points at per-iteration cost, which grew with the repo (24 → 173 Go files) because each judge iteration re-runs the gate batteries (red_test.sh full-ledger = the whole gate self-test) as its own tool call.
Measured burn is ~16.8K input tokens/iteration (usage.jsonl: 743K/44, 838K/50). At 100 iterations that is 100 × 16.8K = 1.68M ≈ 17% of the 10M compiled input budget — ample headroom. The licensed action (same rule as CHT-029 25 → 50) is exactly one raise, recorded beside the knob, verified by async re-judge of the unchanged work (68cab4cb + 93c6412d tier1+tier2 PASS).
| Observation | Value | Meaning |
|---|---|---|
| Starvations at rung 50 | 5 | Deterministic, not transient |
| Original two-criterion task | ddb1892b, 5d6b4aa8 (50.5 used) |
Hit cap on full task |
| Split single-criterion halves | f5432fe6, f245ba1c, 467278a2, bc16834b |
Also hit the same cap |
| Triggering diff | 85e2d2c (test-only), fc04d32 (one-script fix) |
Trivial work; cap is the bottleneck |
| Repo size when rung was sized | 24 Go files | Old budget assumption |
| Repo size now | 173 Go files | ~7.2× growth in gate surface |
| Burn per iteration | 743K/44 = 16.89K, 838K/50 = 16.76K |
~16.8K input tok/iter |
| Concurrent sibling load | ~6–13 | Multiplies wall-clock per tool call |
The decisive signal is the split experiment: if the constraint were "too many criteria per iteration," single-criterion tasks would have finished well under 50. They did not. The cost is per iteration, not per criterion.
evaluator.max_iterations was chosen when the repo was small. But a Tier-2 judge iteration is not fixed cost:
red_test.sh full-ledger section runs the entire gate self-test, so one tool call scales with the number of Go files and tests in the tree.6–13 concurrent jobs) each call is slower, so the judge spends longer per iteration and needs more of them before converging.So the required iteration count for a verdict rose with repo size, while the cap did not. The judge was starved, not wrong.
Splitting reduces breadth per iteration (fewer things judged at once), but it does not reduce the fixed gate cost the judge pays each iteration. Because the gate batteries dominate, single-criterion tasks still consumed the full 50. This is the standard "near-miss re-run → split criteria → still starved at same rung" ladder, and reaching the end of it is precisely the licensed ONE-raise condition.
load_defaults reads only the defaults block. Editing an override elsewhere, or reading the raw config file, can show a different number than what the engine actually enforces. Always resolve through:
engine.eval_cap.eval_cap_from_config(cfg)
If you "raise" a value that eval_cap_from_config does not pick up, the starvation persists and looks like the raise failed. Resolve the effective cap before and after the change.
10M / 16.8K ≈ 595.100 × 16.8K = 1.68M = 16.8% ≈ 17% of budget.100 is deliberately conservative: it matches the CHT-029 precedent of doubling the rung, and it is justified by measured burn rather than guesswork.
A
tier1testsCommand-timed-out under host load is harness-class, not a judge failure. Re-judge on the quiet tree; do not use it as a reason to raise.
cd "$(git rev-parse --show-toplevel)"
# Find every declaration, including the defaults block.
grep -rn "max_iterations" . \
--include='*.yaml' --include='*.yml' --include='*.toml' \
--include='*.py' --include='*.json' \
| grep -v node_modules
Adapt the loader import to the repo entrypoint; the mandatory call is eval_cap_from_config.
env -u GITREINS_LLM_MODEL python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults # <-- use the repo's actual loader
cfg = load_defaults()
print("effective evaluator.max_iterations =", eval_cap_from_config(cfg))
PY
Record the printed value. It is the only number that matters.
defaults block (the one load_defaults reads)defaults:
evaluator:
# ------------------------------------------------------------------
# evaluator.max_iterations escalation history (keep beside the knob)
# 25 -> 50 CHT-029 (precedent: doubling for iteration headroom)
# 50 -> 100 CHT-088 Repo grew 24 -> 173 Go files. Tier-2 judge
# re-runs the full-ledger gate self-test
# (red_test.sh) each iteration; measured burn
# ~16.8K input tok/iter (usage.jsonl 743K/44,
# 838K/50). Single-criterion split CHT-088-A/B
# (f5432fe6, f245ba1c, 467278a2, bc16834b)
# still starved at 50 => cost is per-iteration,
# not per-criterion. Sizing:
# 100 x 16.8K = 1.68M = 17% of the 10M
# compiled input budget. Trigger commits:
# 85e2d2c (test-only), fc04d32 (one-script fix).
# ------------------------------------------------------------------
max_iterations: 100 # was 50
git checkout -b chore/cht-088-raise-max-iterations
git add <path/to/config>
git commit -m "chore(eval): raise evaluator.max_iterations 50 -> 100 (CHT-088)
Repo grew 24 -> 173 Go files. Tier-2 judge re-runs the full-ledger
gate self-test (red_test.sh) each iteration; measured burn 16.8K
input tok/iter (usage.jsonl 743K/44, 838K/50). Single-criterion split
CHT-088-A/B still starved at 50, so criterion count is not the driver.
Sizing: 100 x 16.8K = 1.68M = 17% of the 10M compiled input budget.
Precedent: CHT-029 raised 25 -> 50. Escalation history is beside the knob."
One knob, one commit, history adjacent. That is the whole fix.
env -u GITREINS_LLM_MODEL python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults
print(eval_cap_from_config(load_defaults()))
PY
# expect: 100
The work under test is not modified — only the rung changed. Run both tiers. The recorded proof jobs are 68cab4cb and 93c6412d, both tier1+tier2 PASS.
# Tier 1 (test/judge harness)
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 1
# -> prints a job id
# Tier 2 (the starved judge)
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 2
# -> prints a job id
# Poll the printed ids (repeat until terminal)
scripts/judge.sh --status <job-id>
.passed keys — never .status.status can read like a lifecycle state ("completed"/"starved") and is misleading. Gate on the boolean:
import json, sys
r = json.load(open("verdict.json"))
assert r["tiers"]["tier1"]["passed"] is True, r["tiers"]["tier1"]
assert r["tiers"]["tier2"]["passed"] is True, r["tiers"]["tier2"]
print("CHT-088 verified: tier1+tier2 PASS at max_iterations=100")
| Check | Before | After |
|---|---|---|
Effective max_iterations |
50 | 100 |
| tier1 | starved/timeout | PASS (68cab4cb) |
| tier2 | starved ×5 | PASS (93c6412d) |
| Work diff | 85e2d2c+fc04d32 |
unchanged |
| Effective budget used | 838K/50 iters | 1.68M/100 iters (17% of 10M) |
If tier2 still starves at 100, the root cause is not the rung — go shrink the judge's per-iteration gate cost (§7) before considering anything else.
Further raises are not the next step. The sustainable fix is to reduce what the judge re-runs each iteration:
red_test.sh full-ledger — the gate self-test should be a fixture, not an every-iteration cost.6–13 sibling load inflates per-call time and triggers harness timeouts).Only after per-iteration cost is bounded does raising max_iterations again become a legitimate — and still singular — option.
# 0. Work from the repo, no ambient model override
cd "$(git rev-parse --show-toplevel)"
unset GITREINS_LLM_MODEL
# 1. See the effective cap (NEVER trust the raw config file)
python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults
print(eval_cap_from_config(load_defaults()))
PY
# 2. Apply §5.3 edit, commit §5.4
# 3. Async judge both tiers and capture the printed job ids
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 1
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 2
# 4. Poll
scripts/judge.sh --status <job-id>
# 5. Verify on .passed, not .status
Do:
- Resolve the cap through engine.eval_cap.eval_cap_from_config.
- Keep escalation history beside the knob.
- Judge async and poll the printed job id.
- Re-judge harness-class timeouts on a quiet tree.
Don't:
- Split criteria again once single-criterion tasks starve at the same rung.
- Raise the rung more than once without new measured evidence.
- Read verdicts from .status.
- Treat a host-load Command-timed-out as a real judge failure.
The fix is a single config value in its own commit, so rollback is trivial:
git revert <chore-commit-sha>
# then re-resolve the effective cap and confirm it is 50
No code path, test, or judge logic changed, so reverting cannot regress functional behavior.
small 2-commit diff lands
│
repo has grown 24 -> 173 Go files
│
Tier-2 judge iteration re-runs full-ledger gate self-test (red_test.sh)
│
per-iteration cost ~16.8K input tok (+ sibling load 6–13)
│
required iterations > 50, but rung is still 50 ──> STARVATION ×5
│
split criteria (ddb1892b/5d6b4aa8 -> CHT-088-A/B)
│
single-criterion halves STILL starve at 50
│
= licensed ONE-raise condition (CHT-029 precedent)
│
raise defaults.evaluator.max_iterations 50 -> 100
│
100 × 16.8K = 1.68M = 17% of 10M budget
│
re-judge unchanged work async -> tier1 + tier2 PASS
│
(if it ever starves at 100: shrink gate cost via hermetic fixtures, do not climb)
Note on environment: the actual coding-hermes-tools/gitreins checkout is not present here, so I could not execute eval_cap_from_config or scripts/judge.sh to re-confirm live. The verification section is expressed as executable checks against the recorded evidence (usage.jsonl burn figures and proof job IDs 68cab4cb/93c6412d). If you run this in the real repo, substitute the actual config path and loader import at §5.2.
# Evidence - Problem class: gitreins-tier2-iteration-starvation-after-repo-growth - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-26T02:00:41.204Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "CHT-088 closeout (coding-hermes-tools): the Tier-2 judge starved at evaluator.max_iterations 50 five times after a small two-commit diff (85e2d2c test-only, fc04d32 one-script fix) landed \u2014 twice on the original two-criterion task (ddb1892b, 5d6b4aa8 at 50.5 used), then on BOTH halves of the licensed single-criterion split CHT-088-A/B (f5432fe6, f245ba1c, 467278a2, bc16834b). The split remedy did NOT fix it because the burn is not criterion count, it is repo growth + gate cost: the rung was sized when the repo had 24 Go files, it now has 173, and Tier-2 judge iterations re-run the repo gate batteries (red_test.sh full-ledger section = the whole gate self-test; every arm = one tool call, slower under sibling host load ~6-13). Measured burn ~16.8K input tokens/iteration (usage.jsonl 743K/44, 838K/50). Ladder that worked: near-miss unchanged re-runs -> split criteria -> when single-criterion tasks still starve at the SAME rung, that is the licensed ONE-raise condition (same rule as CHT-029 25->50): raise in an isolated chore commit with the escalation history beside the knob, sized from measured burn vs the compiled input budget (100 x 16.8K = 1.68M = 17% of 10M), then prove the rung by re-judging the unchanged work async (68cab4cb + 93c6412d tier1+tier2 PASS). Operational notes: always resolve the EFFECTIVE cap through engine.eval_cap.eval_cap_from_config (load_defaults reads only the defaults block); read verdict tiers by .passed keys never .status; run judges via scripts/judge.sh (env -u GITREINS_LLM_MODEL) --async and poll the printed job id; a tier1 tests Command-timed-out under host load is harness-class, re-judge on the quiet tree. Alternative to a further raise: shrink what the judge re-runs (hermetic fixtures), not climb.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-iteration-starvation-after-repo-growth", "provider": "openrouter", "solved_at": "2026-09-26T02:00:41.204Z", "version": ""}Solution written to ~/solution.md. Since no coding-hermes-tools checkout or gitreins source is mounted in this environment (only /workspace/problem.json), the write-up is built from the recorded evidence in the problem statement and is structured as a directly usable runbook.
Problem class: gitreins-tier2-iteration-starvation-after-repo-growth
Repo: coding-hermes-tools
Knob: evaluator.max_iterations (gitreins engine)
Verdict: Raise the rung 50 → 100 in one isolated chore commit, sized from measured burn, then prove it by re-judging the unchanged work async. Do not keep splitting criteria and do not treat host-load timeouts as judge failures.
The Tier-2 judge hit evaluator.max_iterations = 50 five times on a two-commit diff (85e2d2c test-only, fc04d32 one-script fix). Splitting the task into single-criterion halves did not help: both halves still starved at the same rung. That rules out criterion count and points at per-iteration cost, which grew with the repo (24 → 173 Go files) because each judge iteration re-runs the gate batteries (red_test.sh full-ledger = the whole gate self-test) as its own tool call.
Measured burn is ~16.8K input tokens/iteration (usage.jsonl: 743K/44, 838K/50). At 100 iterations that is 100 × 16.8K = 1.68M ≈ 17% of the 10M compiled input budget — ample headroom. The licensed action (same rule as CHT-029 25 → 50) is exactly one raise, recorded beside the knob, verified by async re-judge of the unchanged work (68cab4cb + 93c6412d tier1+tier2 PASS).
| Observation | Value | Meaning |
|---|---|---|
| Starvations at rung 50 | 5 | Deterministic, not transient |
| Original two-criterion task | ddb1892b, 5d6b4aa8 (50.5 used) |
Hit cap on full task |
| Split single-criterion halves | f5432fe6, f245ba1c, 467278a2, bc16834b |
Also hit the same cap |
| Triggering diff | 85e2d2c (test-only), fc04d32 (one-script fix) |
Trivial work; cap is the bottleneck |
| Repo size when rung was sized | 24 Go files | Old budget assumption |
| Repo size now | 173 Go files | ~7.2× growth in gate surface |
| Burn per iteration | 743K/44 = 16.89K, 838K/50 = 16.76K |
~16.8K input tok/iter |
| Concurrent sibling load | ~6–13 | Multiplies wall-clock per tool call |
The decisive signal is the split experiment: if the constraint were "too many criteria per iteration," single-criterion tasks would have finished well under 50. They did not. The cost is per iteration, not per criterion.
evaluator.max_iterations was chosen when the repo was small. But a Tier-2 judge iteration is not fixed cost:
red_test.sh full-ledger section runs the entire gate self-test, so one tool call scales with the number of Go files and tests in the tree.6–13 concurrent jobs) each call is slower, so the judge spends longer per iteration and needs more of them before converging.So the required iteration count for a verdict rose with repo size, while the cap did not. The judge was starved, not wrong.
Splitting reduces breadth per iteration (fewer things judged at once), but it does not reduce the fixed gate cost the judge pays each iteration. Because the gate batteries dominate, single-criterion tasks still consumed the full 50. This is the standard "near-miss re-run → split criteria → still starved at same rung" ladder, and reaching the end of it is precisely the licensed ONE-raise condition.
load_defaults reads only the defaults block. Editing an override elsewhere, or reading the raw config file, can show a different number than what the engine actually enforces. Always resolve through:
engine.eval_cap.eval_cap_from_config(cfg)
If you "raise" a value that eval_cap_from_config does not pick up, the starvation persists and looks like the raise failed. Resolve the effective cap before and after the change.
10M / 16.8K ≈ 595.100 × 16.8K = 1.68M = 16.8% ≈ 17% of budget.100 is deliberately conservative: it matches the CHT-029 precedent of doubling the rung, and it is justified by measured burn rather than guesswork.
A
tier1testsCommand-timed-out under host load is harness-class, not a judge failure. Re-judge on the quiet tree; do not use it as a reason to raise.
cd "$(git rev-parse --show-toplevel)"
# Find every declaration, including the defaults block.
grep -rn "max_iterations" . \
--include='*.yaml' --include='*.yml' --include='*.toml' \
--include='*.py' --include='*.json' \
| grep -v node_modules
Adapt the loader import to the repo entrypoint; the mandatory call is eval_cap_from_config.
env -u GITREINS_LLM_MODEL python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults # <-- use the repo's actual loader
cfg = load_defaults()
print("effective evaluator.max_iterations =", eval_cap_from_config(cfg))
PY
Record the printed value. It is the only number that matters.
defaults block (the one load_defaults reads)defaults:
evaluator:
# ------------------------------------------------------------------
# evaluator.max_iterations escalation history (keep beside the knob)
# 25 -> 50 CHT-029 (precedent: doubling for iteration headroom)
# 50 -> 100 CHT-088 Repo grew 24 -> 173 Go files. Tier-2 judge
# re-runs the full-ledger gate self-test
# (red_test.sh) each iteration; measured burn
# ~16.8K input tok/iter (usage.jsonl 743K/44,
# 838K/50). Single-criterion split CHT-088-A/B
# (f5432fe6, f245ba1c, 467278a2, bc16834b)
# still starved at 50 => cost is per-iteration,
# not per-criterion. Sizing:
# 100 x 16.8K = 1.68M = 17% of the 10M
# compiled input budget. Trigger commits:
# 85e2d2c (test-only), fc04d32 (one-script fix).
# ------------------------------------------------------------------
max_iterations: 100 # was 50
git checkout -b chore/cht-088-raise-max-iterations
git add <path/to/config>
git commit -m "chore(eval): raise evaluator.max_iterations 50 -> 100 (CHT-088)
Repo grew 24 -> 173 Go files. Tier-2 judge re-runs the full-ledger
gate self-test (red_test.sh) each iteration; measured burn 16.8K
input tok/iter (usage.jsonl 743K/44, 838K/50). Single-criterion split
CHT-088-A/B still starved at 50, so criterion count is not the driver.
Sizing: 100 x 16.8K = 1.68M = 17% of the 10M compiled input budget.
Precedent: CHT-029 raised 25 -> 50. Escalation history is beside the knob."
One knob, one commit, history adjacent. That is the whole fix.
env -u GITREINS_LLM_MODEL python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults
print(eval_cap_from_config(load_defaults()))
PY
# expect: 100
The work under test is not modified — only the rung changed. Run both tiers. The recorded proof jobs are 68cab4cb and 93c6412d, both tier1+tier2 PASS.
# Tier 1 (test/judge harness)
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 1
# -> prints a job id
# Tier 2 (the starved judge)
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 2
# -> prints a job id
# Poll the printed ids (repeat until terminal)
scripts/judge.sh --status <job-id>
.passed keys — never .status.status can read like a lifecycle state ("completed"/"starved") and is misleading. Gate on the boolean:
import json, sys
r = json.load(open("verdict.json"))
assert r["tiers"]["tier1"]["passed"] is True, r["tiers"]["tier1"]
assert r["tiers"]["tier2"]["passed"] is True, r["tiers"]["tier2"]
print("CHT-088 verified: tier1+tier2 PASS at max_iterations=100")
| Check | Before | After |
|---|---|---|
Effective max_iterations |
50 | 100 |
| tier1 | starved/timeout | PASS (68cab4cb) |
| tier2 | starved ×5 | PASS (93c6412d) |
| Work diff | 85e2d2c+fc04d32 |
unchanged |
| Effective budget used | 838K/50 iters | 1.68M/100 iters (17% of 10M) |
If tier2 still starves at 100, the root cause is not the rung — go shrink the judge's per-iteration gate cost (§7) before considering anything else.
Further raises are not the next step. The sustainable fix is to reduce what the judge re-runs each iteration:
red_test.sh full-ledger — the gate self-test should be a fixture, not an every-iteration cost.6–13 sibling load inflates per-call time and triggers harness timeouts).Only after per-iteration cost is bounded does raising max_iterations again become a legitimate — and still singular — option.
# 0. Work from the repo, no ambient model override
cd "$(git rev-parse --show-toplevel)"
unset GITREINS_LLM_MODEL
# 1. See the effective cap (NEVER trust the raw config file)
python - <<'PY'
from engine.eval_cap import eval_cap_from_config
from engine.config import load_defaults
print(eval_cap_from_config(load_defaults()))
PY
# 2. Apply §5.3 edit, commit §5.4
# 3. Async judge both tiers and capture the printed job ids
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 1
env -u GITREINS_LLM_MODEL scripts/judge.sh --async --task CHT-088 --tier 2
# 4. Poll
scripts/judge.sh --status <job-id>
# 5. Verify on .passed, not .status
Do:
- Resolve the cap through engine.eval_cap.eval_cap_from_config.
- Keep escalation history beside the knob.
- Judge async and poll the printed job id.
- Re-judge harness-class timeouts on a quiet tree.
Don't:
- Split criteria again once single-criterion tasks starve at the same rung.
- Raise the rung more than once without new measured evidence.
- Read verdicts from .status.
- Treat a host-load Command-timed-out as a real judge failure.
The fix is a single config value in its own commit, so rollback is trivial:
git revert <chore-commit-sha>
# then re-resolve the effective cap and confirm it is 50
No code path, test, or judge logic changed, so reverting cannot regress functional behavior.
small 2-commit diff lands
│
repo has grown 24 -> 173 Go files
│
Tier-2 judge iteration re-runs full-ledger gate self-test (red_test.sh)
│
per-iteration cost ~16.8K input tok (+ sibling load 6–13)
│
required iterations > 50, but rung is still 50 ──> STARVATION ×5
│
split criteria (ddb1892b/5d6b4aa8 -> CHT-088-A/B)
│
single-criterion halves STILL starve at 50
│
= licensed ONE-raise condition (CHT-029 precedent)
│
raise defaults.evaluator.max_iterations 50 -> 100
│
100 × 16.8K = 1.68M = 17% of 10M budget
│
re-judge unchanged work async -> tier1 + tier2 PASS
│
(if it ever starves at 100: shrink gate cost via hermetic fixtures, do not climb)
Note on environment: the actual coding-hermes-tools/gitreins checkout is not present here, so I could not execute eval_cap_from_config or scripts/judge.sh to re-confirm live. The verification section is expressed as executable checks against the recorded evidence (usage.jsonl burn figures and proof job IDs 68cab4cb/93c6412d). If you run this in the real repo, substitute the actual config path and loader import at §5.2.
# Evidence - Problem class: gitreins-tier2-iteration-starvation-after-repo-growth - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-26T02:00:41.204Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "CHT-088 closeout (coding-hermes-tools): the Tier-2 judge starved at evaluator.max_iterations 50 five times after a small two-commit diff (85e2d2c test-only, fc04d32 one-script fix) landed \u2014 twice on the original two-criterion task (ddb1892b, 5d6b4aa8 at 50.5 used), then on BOTH halves of the licensed single-criterion split CHT-088-A/B (f5432fe6, f245ba1c, 467278a2, bc16834b). The split remedy did NOT fix it because the burn is not criterion count, it is repo growth + gate cost: the rung was sized when the repo had 24 Go files, it now has 173, and Tier-2 judge iterations re-run the repo gate batteries (red_test.sh full-ledger section = the whole gate self-test; every arm = one tool call, slower under sibling host load ~6-13). Measured burn ~16.8K input tokens/iteration (usage.jsonl 743K/44, 838K/50). Ladder that worked: near-miss unchanged re-runs -> split criteria -> when single-criterion tasks still starve at the SAME rung, that is the licensed ONE-raise condition (same rule as CHT-029 25->50): raise in an isolated chore commit with the escalation history beside the knob, sized from measured burn vs the compiled input budget (100 x 16.8K = 1.68M = 17% of 10M), then prove the rung by re-judging the unchanged work async (68cab4cb + 93c6412d tier1+tier2 PASS). Operational notes: always resolve the EFFECTIVE cap through engine.eval_cap.eval_cap_from_config (load_defaults reads only the defaults block); read verdict tiers by .passed keys never .status; run judges via scripts/judge.sh (env -u GITREINS_LLM_MODEL) --async and poll the printed job id; a tier1 tests Command-timed-out under host load is harness-class, re-judge on the quiet tree. Alternative to a further raise: shrink what the judge re-runs (hermetic fixtures), not climb.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-iteration-starvation-after-repo-growth", "provider": "openrouter", "solved_at": "2026-09-26T02:00:41.204Z", "version": ""}