Problem class: gitreins-judge-input-cap-starvation
Written to ~/gitreins-judge-input-cap-starvation.md. Full solution:
INCOMPLETE: Judge Input-Token Budget StarvationProblem class: gitreins-judge-input-cap-starvation
Affects: gitreins 0.11.0 CLI, large JS repos (≈800 test files / 2507 tests), evaluator deepseek-v4-flash
Signature: used == configured cap, tier1 PASS, tier2 FAIL INCOMPLETE, no merit failure
gitreins.evaluator: WARNING: Eval cap exceeded: Input token budget (1.0M) exceeded (1.0M used).
Increase max_input_tokens or reduce message context.
-> Stage tier1: PASS / Stage tier2: FAIL INCOMPLETE / Overall: FAIL
Same 5-criterion task, same code, only budget changed:
Configured max_input_tokens |
Result |
|---|---|
| 1M | starved (INCOMPLETE) |
| 2M | starved (INCOMPLETE) |
| 4M | PASS (verdict ce3762f6) |
The evaluator enforces a cumulative input-token budget across the whole judge run (max_input_tokens). On a large repo, repo context is re-sent across many round-trips (~100 iterations for a 5-criterion task), so the sum exceeds the configured budget and the run aborts mid-eval as INCOMPLETE, never reaching merit judgments.
The diagnostic is the equality used == configured cap (e.g. 1.0M used against 1.0M):
- Budget starvation, not a model-context ceiling.
- A true ceiling reads as a cap above the model window (config 10M, error shows model's 1.0M).
- Cumulative budget ≠ per-call window: 4M over ~100 iterations is ~40K/iteration against a 1M per-call context.
.gitreins/tasks.yaml status to complete — status is not a pass.--force can re-trigger a run never evaluated on merits. Read verdict.json.Raise max_input_tokens. Some repos have only the top-level evaluator block; others also carry a stage-level tier2 value. Both must move.
1. Locate every occurrence:
cd <repo-root>
grep -rn "max_input_tokens" .gitreins/config.yaml
2. Raise both to 4M:
evaluator:
model: deepseek-v4-flash
max_input_tokens: 4000000 # was 1000000
stages:
tier2:
evaluator:
max_input_tokens: 4000000 # must also move if present
Non-interactive:
sed -i 's/max_input_tokens: 1000000/max_input_tokens: 4000000/g' .gitreins/config.yaml
grep -n "max_input_tokens" .gitreins/config.yaml # confirm both
git add .gitreins/config.yaml
git commit -m "chore(gitreins): raise evaluator max_input_tokens 1M->4M"
3. Force fresh judge run (prior starved run already marked task complete):
.venv/bin/gitreins task complete <id> --force
Sizing: for >500 test files, start at 4M instead of laddering 1M→2M→4M (each wasted run ~3–5 min).
Do not trust .gitreins/tasks.yaml. Inspect the newest verdict.
TASK=<id>
V=$(ls -1dt .gitreins/history/$(date -u +%F)/$TASK/*/verdict.json | head -1)
echo "$V"
# midnight fallback:
# V=$(ls -1dt .gitreins/history/*/$TASK/*/verdict.json | head -1)
jq -e '
.passed == true
and .stages.tier1.passed == true
and .stages.tier2.passed == true
' "$V" >/dev/null && echo "VERIFIED PASS" || { echo "NOT A PASS"; jq '{passed, stages: (.stages | {tier1: .tier1.passed, tier2: .tier2.passed})}' "$V"; }
Expected: Overall PASS; verdict.json has passed: true, tier1.passed: true, tier2.passed: true; no Eval cap exceeded warning. Reference: totalwindupflightsystems/9router, verdict ce3762f6.
| Observation | Meaning | Action |
|---|---|---|
used == configured cap |
Cumulative input budget exhausted | Raise max_input_tokens (both locations), --force re-run |
| Error cap > model window (10M vs 1.0M) | Model context ceiling | Reduce context, don't raise config |
passed=false, tier2 fails on merits |
Real quality failure | Fix code/criteria |
.gitreins/tasks.yaml: status complete |
Not authoritative | Read verdict.json |
Note: no gitreins install or repo was present in this sandbox (~/gitreins-poc does not exist), so commands above are grounded in the reported 0.11.0 behavior/artifacts rather than executed here.
# Evidence - Problem class: gitreins-judge-input-cap-starvation - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-14T01:53:03.721Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "GitReins Tier 2 judge returns INCOMPLETE with 'Eval cap exceeded: Input token budget (N) exceeded (N used)' where N is the CONFIGURED evaluator max_input_tokens in .gitreins/config.yaml. On a large repo a 5-criterion task starves at 1M AND at 2M, then PASSES at 4M with identical criteria and identical code. The used==configured-cap equality is the diagnostic: it means budget starvation mid-eval, NOT the model context ceiling (a ceiling reads as a cap above the model window, e.g. config 10M but error shows 1.0M). The fix is to RAISE max_input_tokens (there is only one location in this config: the top-level evaluator block, though some repos also carry a stage-level tier2 max_input_tokens and BOTH must move), then re-run 'gitreins task complete <id> --force'. Do NOT split criteria or switch models on the used==configured signature. Note the cumulative input budget is not the per-call model window: 4M burned over ~100 iterations is ~40K/iteration against a 1M per-call context. Sizing guidance for repos with >500 test files: start at 4M rather than laddering 1M -> 2M -> 4M, which costs two wasted judge runs (~3-5 min each) per task. Also: a cap-starved run still flips .gitreins/tasks.yaml status to 'complete', so the status is NOT a pass \u2014 read the newest .gitreins/history/<UTC-date>/<id>/verdict.json and require passed=true with stages.tier1.passed AND stages.tier2.passed.", "environment": "gitreins 0.11.0 CLI (~/gitreins-poc/.venv/bin/gitreins), evaluator model deepseek-v4-flash, large JS repo (800 test files, 2507 tests)", "language": "yaml", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-judge-input-cap-starvation", "provider": "openrouter", "solved_at": "2026-09-14T01:53:03.721Z", "version": "0.11.0"}Written to ~/gitreins-judge-input-cap-starvation.md. Full solution:
INCOMPLETE: Judge Input-Token Budget StarvationProblem class: gitreins-judge-input-cap-starvation
Affects: gitreins 0.11.0 CLI, large JS repos (≈800 test files / 2507 tests), evaluator deepseek-v4-flash
Signature: used == configured cap, tier1 PASS, tier2 FAIL INCOMPLETE, no merit failure
gitreins.evaluator: WARNING: Eval cap exceeded: Input token budget (1.0M) exceeded (1.0M used).
Increase max_input_tokens or reduce message context.
-> Stage tier1: PASS / Stage tier2: FAIL INCOMPLETE / Overall: FAIL
Same 5-criterion task, same code, only budget changed:
Configured max_input_tokens |
Result |
|---|---|
| 1M | starved (INCOMPLETE) |
| 2M | starved (INCOMPLETE) |
| 4M | PASS (verdict ce3762f6) |
The evaluator enforces a cumulative input-token budget across the whole judge run (max_input_tokens). On a large repo, repo context is re-sent across many round-trips (~100 iterations for a 5-criterion task), so the sum exceeds the configured budget and the run aborts mid-eval as INCOMPLETE, never reaching merit judgments.
The diagnostic is the equality used == configured cap (e.g. 1.0M used against 1.0M):
- Budget starvation, not a model-context ceiling.
- A true ceiling reads as a cap above the model window (config 10M, error shows model's 1.0M).
- Cumulative budget ≠ per-call window: 4M over ~100 iterations is ~40K/iteration against a 1M per-call context.
.gitreins/tasks.yaml status to complete — status is not a pass.--force can re-trigger a run never evaluated on merits. Read verdict.json.Raise max_input_tokens. Some repos have only the top-level evaluator block; others also carry a stage-level tier2 value. Both must move.
1. Locate every occurrence:
cd <repo-root>
grep -rn "max_input_tokens" .gitreins/config.yaml
2. Raise both to 4M:
evaluator:
model: deepseek-v4-flash
max_input_tokens: 4000000 # was 1000000
stages:
tier2:
evaluator:
max_input_tokens: 4000000 # must also move if present
Non-interactive:
sed -i 's/max_input_tokens: 1000000/max_input_tokens: 4000000/g' .gitreins/config.yaml
grep -n "max_input_tokens" .gitreins/config.yaml # confirm both
git add .gitreins/config.yaml
git commit -m "chore(gitreins): raise evaluator max_input_tokens 1M->4M"
3. Force fresh judge run (prior starved run already marked task complete):
.venv/bin/gitreins task complete <id> --force
Sizing: for >500 test files, start at 4M instead of laddering 1M→2M→4M (each wasted run ~3–5 min).
Do not trust .gitreins/tasks.yaml. Inspect the newest verdict.
TASK=<id>
V=$(ls -1dt .gitreins/history/$(date -u +%F)/$TASK/*/verdict.json | head -1)
echo "$V"
# midnight fallback:
# V=$(ls -1dt .gitreins/history/*/$TASK/*/verdict.json | head -1)
jq -e '
.passed == true
and .stages.tier1.passed == true
and .stages.tier2.passed == true
' "$V" >/dev/null && echo "VERIFIED PASS" || { echo "NOT A PASS"; jq '{passed, stages: (.stages | {tier1: .tier1.passed, tier2: .tier2.passed})}' "$V"; }
Expected: Overall PASS; verdict.json has passed: true, tier1.passed: true, tier2.passed: true; no Eval cap exceeded warning. Reference: totalwindupflightsystems/9router, verdict ce3762f6.
| Observation | Meaning | Action |
|---|---|---|
used == configured cap |
Cumulative input budget exhausted | Raise max_input_tokens (both locations), --force re-run |
| Error cap > model window (10M vs 1.0M) | Model context ceiling | Reduce context, don't raise config |
passed=false, tier2 fails on merits |
Real quality failure | Fix code/criteria |
.gitreins/tasks.yaml: status complete |
Not authoritative | Read verdict.json |
Note: no gitreins install or repo was present in this sandbox (~/gitreins-poc does not exist), so commands above are grounded in the reported 0.11.0 behavior/artifacts rather than executed here.
# Evidence - Problem class: gitreins-judge-input-cap-starvation - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-14T01:53:03.721Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "GitReins Tier 2 judge returns INCOMPLETE with 'Eval cap exceeded: Input token budget (N) exceeded (N used)' where N is the CONFIGURED evaluator max_input_tokens in .gitreins/config.yaml. On a large repo a 5-criterion task starves at 1M AND at 2M, then PASSES at 4M with identical criteria and identical code. The used==configured-cap equality is the diagnostic: it means budget starvation mid-eval, NOT the model context ceiling (a ceiling reads as a cap above the model window, e.g. config 10M but error shows 1.0M). The fix is to RAISE max_input_tokens (there is only one location in this config: the top-level evaluator block, though some repos also carry a stage-level tier2 max_input_tokens and BOTH must move), then re-run 'gitreins task complete <id> --force'. Do NOT split criteria or switch models on the used==configured signature. Note the cumulative input budget is not the per-call model window: 4M burned over ~100 iterations is ~40K/iteration against a 1M per-call context. Sizing guidance for repos with >500 test files: start at 4M rather than laddering 1M -> 2M -> 4M, which costs two wasted judge runs (~3-5 min each) per task. Also: a cap-starved run still flips .gitreins/tasks.yaml status to 'complete', so the status is NOT a pass \u2014 read the newest .gitreins/history/<UTC-date>/<id>/verdict.json and require passed=true with stages.tier1.passed AND stages.tier2.passed.", "environment": "gitreins 0.11.0 CLI (~/gitreins-poc/.venv/bin/gitreins), evaluator model deepseek-v4-flash, large JS repo (800 test files, 2507 tests)", "language": "yaml", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-judge-input-cap-starvation", "provider": "openrouter", "solved_at": "2026-09-14T01:53:03.721Z", "version": "0.11.0"}