Class id: gitreins-tier2-incomplete-empty-bell-input-cap-starvation
Wrote the solution to ~/gitreins-tier2-empty-verdict-input-cap.md. Full contents below.
items: []” — Agentic Evaluator Input-Token Cap StarvationClass id: gitreins-tier2-incomplete-empty-bell-input-cap-starvation
Area: gitreins task complete → Tier 2 (Agentic Evaluator)
Severity: High (false “harness death” reports; task silently never judged)
Fix cost: One config value, ~30 seconds
Do NOT file a harness bug. This is evaluator context starvation, not a crash.
gitreins task complete returns Tier 2 (Agentic Evaluator): INCOMPLETE, and the produced verdict.json has zero criteria results:
{
"task_id": "...",
"task_title": "...",
"evaluated_at": "...",
"passed": false,
"summary": "...",
"task_criteria": [...],
"items": []
}
Key tell: top-level key set [evaluated_at, items, passed, summary, task_criteria, task_id, task_title] with items: []. This is the partial-or-INCOMPLETE return path, not a crash.
Other characteristics:
- Deterministic on retry — same empty verdict, same ~2 min duration.
- Evaluator model is healthy and reachable.
- Evaluator starts, reads broadly across its iteration loop, exits before emitting any per-criterion items.
It looks like the harness killed the evaluator. It did not.
The Tier 2 evaluator runs an agentic loop that reads many files to independently verify each criterion. Each iteration appends tool output to context; the engine enforces a hard input-token budget per evaluation.
evaluator.max_input_tokens).task-router/registry.json is 1.2 MB of JSON.engine/evaluator.py (~lines 1080–1270) returns partial-or-INCOMPLETE with summary beginning Cap exceeded: ... (or the LLM-call-failed variant).The smoking gun is a logging.WARNING on stderr:
Eval cap exceeded: Input token budget (4.0M) exceeded (4.0M used).
Increase max_input_tokens or reduce message context.
The CLI sets logging to WARNING, so the line only appears on stderr. Piping (... 2>&1 | tail) drops/truncates it, leaving a blank INCOMPLETE that masquerades as harness death.
| Signal | Cap starvation (this class) | Harness death |
|---|---|---|
| Upstream error | none | identical error (e.g. model-name resolution) in promise/real-use |
| Product contact | evaluator reads files | dies before contact |
| Stderr | WARNING ... Eval cap exceeded ... Input token budget |
harness/model failure, no cap warning |
| Runtime | normal (~2 min) | early death |
verdict.json |
items: [], summary Cap exceeded: ... |
no usable verdict / crash |
| Model health | healthy | unhealthy / misconfigured |
Rule: cap-exceeded warning present ⇒ this class. Add headroom; do not file a harness bug.
cd /path/to/repo
gitreins task complete 2>&1 | tee /tmp/gr-complete.log
# or:
gitreins task complete >/tmp/gr-complete.out 2>/tmp/gr-complete.err; echo "exit=$?"
grep -n 'WARNING.*Eval cap exceeded.*Input token budget' /tmp/gr-complete.err /tmp/gr-complete.log
If absent, stop — it's a different class.
Edit git-tracked .gitreins/config.yaml; bump evaluator.max_input_tokens 4000000 (4 M) → 12000000 (12 M). Engine default is 10 M; 12 M is safe for multi-MB tables.
yq (recommended):
cd /path/to/repo
yq -i '.evaluator.max_input_tokens = 12000000' .gitreins/config.yaml
Portable sed fallback:
cd /path/to/repo
grep -q '^\s*max_input_tokens:' .gitreins/config.yaml \
&& sed -i -E 's/^(\s*max_input_tokens:\s*).*/\112000000/' .gitreins/config.yaml \
|| printf '\nevaluator:\n max_input_tokens: 12000000\n' >> .gitreins/config.yaml
Sanity check:
python -c "import yaml; c=yaml.safe_load(open('.gitreins/config.yaml')); print(c['evaluator']['max_input_tokens'])"
# -> 12000000
If headroom can't be raised, reduce evaluator context instead — exclude the giant generated file (e.g. add
registry.jsonto evaluator ignore/glob exclusions). Raising headroom is the validated fix.
cd /path/to/repo
git add .gitreins/config.yaml
git commit -m "chore(gitreins): raise evaluator max_input_tokens 4M -> 12M
Tier 2 agentic evaluator exhausted the 4M input budget on the 1.2MB
generated registry.json before emitting a verdict, producing
INCOMPLETE with items:[]. Bump headroom to 12M."
git push
cd /path/to/repo
gitreins task complete 2>&1 | tee /tmp/gr-complete-after.log
grep -n 'Eval cap exceeded' /tmp/gr-complete-after.log || echo "OK: no cap warning"
The new verdict supersedes the old; both remain under .gitreins/history/<date>/.
LATEST=$(ls -dt .gitreins/history/*/ | head -1)
V="$LATEST/verdict.json"
echo "$V"
python - "$V" <<'PY'
import json, sys
v = json.load(open(sys.argv[1]))
print("keys:", sorted(v))
print("passed:", v.get("passed"))
print("summary:", v.get("summary"))
print("items:", len(v.get("items", [])))
assert v.get("items"), "STILL EMPTY - cap not resolved"
assert not str(v.get("summary","")).startswith("Cap exceeded"), "cap warning persisted"
print("VERDICT COMPLETE")
PY
Success: items non-empty (one per criterion), summary no longer starts with Cap exceeded:, passed: true for the task-router case (all 6 criteria).
grep -c 'Eval cap exceeded' /tmp/gr-complete-after.log # expect 0
2486ca54 + 7678325a were INCOMPLETE / items: [] under the 4 M cap.dfdd01c set evaluator.max_input_tokens: 12000000.0b1eaddf returned PASS with all 6 criteria independently verified, including a live router_spawn re-run and the full 287-test suite.gitreins router_spawn ... # live product path
pytest -q # expect 287 passed
grep -E 'router_spawn|pytest|tool_call' /tmp/gr-complete.log | head # touched product?
grep 'Eval cap exceeded' /tmp/gr-complete.log # cap warning?
jq '.items | length' .gitreins/history/*/verdict.json # empty items?
cd /path/to/repo
gitreins task complete >/tmp/gr.out 2>/tmp/gr.err; echo "exit=$?"
grep -n 'WARNING.*Eval cap exceeded.*Input token budget' /tmp/gr.err /tmp/gr.out
yq -i '.evaluator.max_input_tokens = 12000000' .gitreins/config.yaml
git add .gitreins/config.yaml
git commit -m "chore(gitreins): raise evaluator max_input_tokens 4M -> 12M"
gitreins task complete 2>&1 | tee /tmp/gr.after.log
grep 'Eval cap exceeded' /tmp/gr.after.log || echo "OK: no cap warning"
V=$(ls -dt .gitreins/history/*/verdict.json | head -1)
jq '{passed, summary, n_items: (.items|length)}' "$V"
| tail a gitreins task complete — the #1 cause of misdiagnosing this as harness death..gitreins/history/<date>/; pick the newest by evaluated_at. Non-destructive.items + healthy model ⇒ this class. Missing WARNING + duplicated upstream error + pre-product death ⇒ harness death.Note: the gitreins symlink in this sandbox is broken (its pipx venv directory is empty), so I could not introspect engine/evaluator.py directly. The write-up is built from the authoritative class diagnosis and is executable as-is in a real repo.
# Evidence - Problem class: gitreins-tier2-incomplete-empty-bell-input-cap-starvation - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-11T20:16:24.580Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: gitreins task complete returns Tier 2 (Agentic Evaluator): INCOMPLETE with verdict.json items:[] and zero criteria results; deterministic on retry (~2min to die). Looks like harness death but is NOT. ROOT CAUSE: evaluator input-token cap \u2014 stderr (logging.WARNING, silently dropped when output is tail-piped) says: Eval cap exceeded: Input token budget (4.0M) exceeded (4.0M used). Increase max_input_tokens or reduce message context. Trigger: repos with big generated tables \u2014 task-router registry.json is 1.2MB JSON; the agentic evaluator reads broadly across its iteration loop and exhausts the default 4M input budget before emitting a verdict. FIX: (1) re-run task complete with FULL stderr captured (redirect 2>&1 to a file; never tail-pipe \u2014 the WARNING line is what confirms the class). (2) grep 'WARNING.*Eval cap exceeded.*Input token budget'. (3) bump evaluator.max_input_tokens in .gitreins/config.yaml (4M -> 12M here; engine default is 10M); config.yaml is git-tracked, commit it. (4) re-run gitreins task complete \u2014 new verdict supersedes the old (both stay in .gitreins/history/<date>/). DIAGNOSIS CHAIN: verdict.json keys are [evaluated_at, items, passed, summary, task_criteria, task_id, task_title] with items empty; cap-exceeded path in engine/evaluator.py (~lines 1080-1270) returns partial-or-INCOMPLETE with summary 'Cap exceeded: ...' or the LLM-call-failed variant; CLI sets logging to WARNING so the cap warning only appears in unpiped stderr. DISTINGUISH from evaluator-harness-death (<project> DF class): harness death carries an identical upstream error (e.g. model-name resolution) in promise/real-use fields and dies before product contact; cap starvation shows the cap-exceeded warning, normal-duration death, empty items, and a healthy evaluator model. Do NOT file a harness bug \u2014 just add headroom. EVIDENCE: task-router 2026-09-11 \u2014 verdicts 2486ca54 + 7678325a INCOMPLETE items:[] under 4M cap; fix commit dfdd01c (12M); re-judge verdict 0b1eaddf PASS with all 6 criteria independently verified (incl. live router_spawn re-run + 287-test suite).", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-incomplete-empty-bell-input-cap-starvation", "provider": "openrouter", "solved_at": "2026-09-11T20:16:24.580Z", "version": ""}Wrote the solution to ~/gitreins-tier2-empty-verdict-input-cap.md. Full contents below.
items: []” — Agentic Evaluator Input-Token Cap StarvationClass id: gitreins-tier2-incomplete-empty-bell-input-cap-starvation
Area: gitreins task complete → Tier 2 (Agentic Evaluator)
Severity: High (false “harness death” reports; task silently never judged)
Fix cost: One config value, ~30 seconds
Do NOT file a harness bug. This is evaluator context starvation, not a crash.
gitreins task complete returns Tier 2 (Agentic Evaluator): INCOMPLETE, and the produced verdict.json has zero criteria results:
{
"task_id": "...",
"task_title": "...",
"evaluated_at": "...",
"passed": false,
"summary": "...",
"task_criteria": [...],
"items": []
}
Key tell: top-level key set [evaluated_at, items, passed, summary, task_criteria, task_id, task_title] with items: []. This is the partial-or-INCOMPLETE return path, not a crash.
Other characteristics:
- Deterministic on retry — same empty verdict, same ~2 min duration.
- Evaluator model is healthy and reachable.
- Evaluator starts, reads broadly across its iteration loop, exits before emitting any per-criterion items.
It looks like the harness killed the evaluator. It did not.
The Tier 2 evaluator runs an agentic loop that reads many files to independently verify each criterion. Each iteration appends tool output to context; the engine enforces a hard input-token budget per evaluation.
evaluator.max_input_tokens).task-router/registry.json is 1.2 MB of JSON.engine/evaluator.py (~lines 1080–1270) returns partial-or-INCOMPLETE with summary beginning Cap exceeded: ... (or the LLM-call-failed variant).The smoking gun is a logging.WARNING on stderr:
Eval cap exceeded: Input token budget (4.0M) exceeded (4.0M used).
Increase max_input_tokens or reduce message context.
The CLI sets logging to WARNING, so the line only appears on stderr. Piping (... 2>&1 | tail) drops/truncates it, leaving a blank INCOMPLETE that masquerades as harness death.
| Signal | Cap starvation (this class) | Harness death |
|---|---|---|
| Upstream error | none | identical error (e.g. model-name resolution) in promise/real-use |
| Product contact | evaluator reads files | dies before contact |
| Stderr | WARNING ... Eval cap exceeded ... Input token budget |
harness/model failure, no cap warning |
| Runtime | normal (~2 min) | early death |
verdict.json |
items: [], summary Cap exceeded: ... |
no usable verdict / crash |
| Model health | healthy | unhealthy / misconfigured |
Rule: cap-exceeded warning present ⇒ this class. Add headroom; do not file a harness bug.
cd /path/to/repo
gitreins task complete 2>&1 | tee /tmp/gr-complete.log
# or:
gitreins task complete >/tmp/gr-complete.out 2>/tmp/gr-complete.err; echo "exit=$?"
grep -n 'WARNING.*Eval cap exceeded.*Input token budget' /tmp/gr-complete.err /tmp/gr-complete.log
If absent, stop — it's a different class.
Edit git-tracked .gitreins/config.yaml; bump evaluator.max_input_tokens 4000000 (4 M) → 12000000 (12 M). Engine default is 10 M; 12 M is safe for multi-MB tables.
yq (recommended):
cd /path/to/repo
yq -i '.evaluator.max_input_tokens = 12000000' .gitreins/config.yaml
Portable sed fallback:
cd /path/to/repo
grep -q '^\s*max_input_tokens:' .gitreins/config.yaml \
&& sed -i -E 's/^(\s*max_input_tokens:\s*).*/\112000000/' .gitreins/config.yaml \
|| printf '\nevaluator:\n max_input_tokens: 12000000\n' >> .gitreins/config.yaml
Sanity check:
python -c "import yaml; c=yaml.safe_load(open('.gitreins/config.yaml')); print(c['evaluator']['max_input_tokens'])"
# -> 12000000
If headroom can't be raised, reduce evaluator context instead — exclude the giant generated file (e.g. add
registry.jsonto evaluator ignore/glob exclusions). Raising headroom is the validated fix.
cd /path/to/repo
git add .gitreins/config.yaml
git commit -m "chore(gitreins): raise evaluator max_input_tokens 4M -> 12M
Tier 2 agentic evaluator exhausted the 4M input budget on the 1.2MB
generated registry.json before emitting a verdict, producing
INCOMPLETE with items:[]. Bump headroom to 12M."
git push
cd /path/to/repo
gitreins task complete 2>&1 | tee /tmp/gr-complete-after.log
grep -n 'Eval cap exceeded' /tmp/gr-complete-after.log || echo "OK: no cap warning"
The new verdict supersedes the old; both remain under .gitreins/history/<date>/.
LATEST=$(ls -dt .gitreins/history/*/ | head -1)
V="$LATEST/verdict.json"
echo "$V"
python - "$V" <<'PY'
import json, sys
v = json.load(open(sys.argv[1]))
print("keys:", sorted(v))
print("passed:", v.get("passed"))
print("summary:", v.get("summary"))
print("items:", len(v.get("items", [])))
assert v.get("items"), "STILL EMPTY - cap not resolved"
assert not str(v.get("summary","")).startswith("Cap exceeded"), "cap warning persisted"
print("VERDICT COMPLETE")
PY
Success: items non-empty (one per criterion), summary no longer starts with Cap exceeded:, passed: true for the task-router case (all 6 criteria).
grep -c 'Eval cap exceeded' /tmp/gr-complete-after.log # expect 0
2486ca54 + 7678325a were INCOMPLETE / items: [] under the 4 M cap.dfdd01c set evaluator.max_input_tokens: 12000000.0b1eaddf returned PASS with all 6 criteria independently verified, including a live router_spawn re-run and the full 287-test suite.gitreins router_spawn ... # live product path
pytest -q # expect 287 passed
grep -E 'router_spawn|pytest|tool_call' /tmp/gr-complete.log | head # touched product?
grep 'Eval cap exceeded' /tmp/gr-complete.log # cap warning?
jq '.items | length' .gitreins/history/*/verdict.json # empty items?
cd /path/to/repo
gitreins task complete >/tmp/gr.out 2>/tmp/gr.err; echo "exit=$?"
grep -n 'WARNING.*Eval cap exceeded.*Input token budget' /tmp/gr.err /tmp/gr.out
yq -i '.evaluator.max_input_tokens = 12000000' .gitreins/config.yaml
git add .gitreins/config.yaml
git commit -m "chore(gitreins): raise evaluator max_input_tokens 4M -> 12M"
gitreins task complete 2>&1 | tee /tmp/gr.after.log
grep 'Eval cap exceeded' /tmp/gr.after.log || echo "OK: no cap warning"
V=$(ls -dt .gitreins/history/*/verdict.json | head -1)
jq '{passed, summary, n_items: (.items|length)}' "$V"
| tail a gitreins task complete — the #1 cause of misdiagnosing this as harness death..gitreins/history/<date>/; pick the newest by evaluated_at. Non-destructive.items + healthy model ⇒ this class. Missing WARNING + duplicated upstream error + pre-product death ⇒ harness death.Note: the gitreins symlink in this sandbox is broken (its pipx venv directory is empty), so I could not introspect engine/evaluator.py directly. The write-up is built from the authoritative class diagnosis and is executable as-is in a real repo.
# Evidence - Problem class: gitreins-tier2-incomplete-empty-bell-input-cap-starvation - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-11T20:16:24.580Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "SYMPTOM: gitreins task complete returns Tier 2 (Agentic Evaluator): INCOMPLETE with verdict.json items:[] and zero criteria results; deterministic on retry (~2min to die). Looks like harness death but is NOT. ROOT CAUSE: evaluator input-token cap \u2014 stderr (logging.WARNING, silently dropped when output is tail-piped) says: Eval cap exceeded: Input token budget (4.0M) exceeded (4.0M used). Increase max_input_tokens or reduce message context. Trigger: repos with big generated tables \u2014 task-router registry.json is 1.2MB JSON; the agentic evaluator reads broadly across its iteration loop and exhausts the default 4M input budget before emitting a verdict. FIX: (1) re-run task complete with FULL stderr captured (redirect 2>&1 to a file; never tail-pipe \u2014 the WARNING line is what confirms the class). (2) grep 'WARNING.*Eval cap exceeded.*Input token budget'. (3) bump evaluator.max_input_tokens in .gitreins/config.yaml (4M -> 12M here; engine default is 10M); config.yaml is git-tracked, commit it. (4) re-run gitreins task complete \u2014 new verdict supersedes the old (both stay in .gitreins/history/<date>/). DIAGNOSIS CHAIN: verdict.json keys are [evaluated_at, items, passed, summary, task_criteria, task_id, task_title] with items empty; cap-exceeded path in engine/evaluator.py (~lines 1080-1270) returns partial-or-INCOMPLETE with summary 'Cap exceeded: ...' or the LLM-call-failed variant; CLI sets logging to WARNING so the cap warning only appears in unpiped stderr. DISTINGUISH from evaluator-harness-death (<project> DF class): harness death carries an identical upstream error (e.g. model-name resolution) in promise/real-use fields and dies before product contact; cap starvation shows the cap-exceeded warning, normal-duration death, empty items, and a healthy evaluator model. Do NOT file a harness bug \u2014 just add headroom. EVIDENCE: task-router 2026-09-11 \u2014 verdicts 2486ca54 + 7678325a INCOMPLETE items:[] under 4M cap; fix commit dfdd01c (12M); re-judge verdict 0b1eaddf PASS with all 6 criteria independently verified (incl. live router_spawn re-run + 287-test suite).", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-incomplete-empty-bell-input-cap-starvation", "provider": "openrouter", "solved_at": "2026-09-11T20:16:24.580Z", "version": ""}