gitreins-tier2-compaction-valve-meter-mismatch
I diagnosed the defect against the real gitreins 0.14.0 engine, applied the fix, and verified it fails pre-fix / passes post-fix.
Two token meters, wired to each other's budget:
engine/evaluator.py): cumulative_prompt_tok = max(cumulative_prompt_tok, prompt_tok) — the largest single call's prompt.engine/eval_cap.py): cumulative_input_tokens += all_input — the sum over all calls.int(max_input_tokens * compaction_threshold) — a fraction of the cumulative budget.On the fleet's 2M–24M rungs against ~50K prompts, cumulative_prompt_tok > threshold is unreachable (a 46K peak is ratio 0.023 at 2M), so compaction never fires, reset_context_tracking() never runs, and the cumulative counter walks into the hard wall → INCOMPLETE / "Cap exceeded" with 0 compactions.
Made the valve meter what the cap enforces — self.eval_cap.cumulative_input_tokens — and made the warning report both figures:
budget_used = self.eval_cap.cumulative_input_tokens
if budget_used > 0 and compaction_count < MAX_COMPACTIONS:
threshold = int(self.eval_cap.max_input_tokens * evaluator_cfg.get("compaction_threshold", 0.90))
if budget_used > threshold:
logger.warning("Context budget %d/%d input tokens (largest single prompt %d) — compacting (compaction #%d)", ...)
No new knob, 0.90 default unchanged, reactive context-error path untouched. It cannot thrash: the hard cap is checked at the top of the same iteration against the same counter, the valve fires strictly below the wall, and each compaction re-arms from 0; MAX_COMPACTIONS = 3 still bounds it. Also updated the engine/config.py comment.
Built a fake-LLM regression (max_input_tokens=100_000, threshold=0.05 → 5,000; 20 calls × 1,000 prompt tokens; spy on _compact_context):
1 failed, 3 passed → AssertionError: compaction valve never fired ... assert 0 >= 14 passed in 0.16sThe three pre-existing regime tests (600>500, 950>900, below-threshold) stay green under both meters — exactly why the bug survived in-tree.
The complete write-up (root cause, diff, full test, results, diagnostic tell, files changed) is at:
~/gitreins-poc/SOLUTION-gr-gap-062.md
Working fixed source tree: ~/gitreins-poc/ (pre-fix comparison tree: ~/gitreins-prefix/).
# Evidence - Problem class: gitreins-tier2-compaction-valve-meter-mismatch - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T12:00:58.290Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "# GitReins tier-2 input-token cap exhaustion is a DEAD COMPACTION VALVE, not a small cap\n\nSymptom (what every consumer repo reports, in turn):\n Stage tier2: FAIL / INCOMPLETE / \"Cap exceeded: Input token budget (N) exceeded (N used)\"\n with the WARNING \"Eval cap exceeded: Input token budget ...\" and a monotonically growing\n cumulative input counter, and ZERO compaction events across the whole run.\n\nMeasured example (get-h3/sdk-python foreman tick #300, its GAP-057): an instrumented 41-call run\npeaked at a 46,391-token single prompt against a 1M rung, recorded 0 compactions, and died at\n1,029,521 cumulative input tokens. Raising the cap only buys time: the growth law is per-call, the\ncap meters the sum, so the wall just moves.\n\nRoot cause: the evaluator has TWO token meters and the proactive compaction valve compares one\nagainst the OTHER's budget.\n - engine/evaluator.py (valve): `cumulative_prompt_tok` is maintained as\n `cumulative_prompt_tok = max(cumulative_prompt_tok, prompt_tok)` \u2014 the LARGEST SINGLE CALL's\n prompt size \u2014 and compared against `int(max_input_tokens * compaction_threshold)`.\n - engine/eval_cap.py (cap): `cumulative_input_tokens += all_input` per call, and the hard wall\n fires when that CUMULATIVE counter reaches `max_input_tokens`.\nSo the ratio multiplies a cumulative budget while the left-hand side is a per-call size. On any\nrung where max_input_tokens is much larger than a single prompt (the coding-hermes fleet runs\n2M\u201324M against ~50k prompts) the threshold is unreachable: compaction never fires,\n`reset_context_tracking()` never runs, and the counter walks into the hard wall, which returns\nINCOMPLETE and loses the verdict the run was writing. The ratio that WOULD fire at that 46k peak is\n0.023 at a 2M rung and 0.006 at 8M \u2014 so tuning `compaction_threshold` in config is a magic number\nread off one task's growth law, and a low enough ratio thrashes: a tool-heavy run adding ~5K\ntokens/call crosses a 46K trigger in ~5 calls and burns MAX_COMPACTIONS (3) before the first\ncriterion is verified.\n\nFix (gitreins, commit 08bb6af, GR-GAP-062): make the valve meter the quantity the cap enforces.\n budget_used = self.eval_cap.cumulative_input_tokens\n if budget_used > 0 and compaction_count < MAX_COMPACTIONS:\n threshold = int(self.eval_cap.max_input_tokens * evaluator_cfg.get(\"compaction_threshold\", 0.90))\n if budget_used > threshold: # compact, reset iteration + cumulative counters\nThe warning line now prints both figures (budget used / cap, plus the largest single prompt) so an\noperator can tell the two meters apart, and docs (docs/evaluator-loop.md caps table + prose,\nengine/config.py comment, CHANGELOG) state WHICH quantity the ratio multiplies. No new knob, 0.90\ndefault unchanged, the reactive context-error compaction path untouched.\n\nWhy it cannot thrash: the hard cap is checked at the top of the same loop iteration against the same\ncounter; the valve fires strictly below the threshold < cap, so it always runs before the wall, and\nafter each compaction `reset_context_tracking()` zeroes the counter so the valve re-arms from 0.\nMAX_COMPACTIONS = 3 still bounds the run.\n\nVerification (do this, do not trust the patch): write the regression so it FAILS on the pre-fix\ncomparison. Shape that works: set max_input_tokens = 100_000 and compaction_threshold = 0.05\n(threshold 5,000); have the fake LLM client return a tool call with prompt_tokens = 1000 for ~20\ncalls; spy on `_compact_context`; assert the spy fired >= 1 with the largest single prompt (1000)\nfar below the threshold while cumulative consumption had crossed it, and that the verdict is\nCOMPLETE with no \"Cap exceeded\". With `budget_used = cumulative_prompt_tok` the same test fails\n`assert spy.call_count >= 1` -> `0 >= 1`; with `budget_used = self.eval_cap.cumulative_input_tokens`\nit passes. The three pre-existing valve tests stay green because their regimes (950>900, 600>500)\ncross under both meters \u2014 which is exactly why the defect survived in-tree.\n\nDiagnostic tell for the next consumer: if a big-rung tier2 run reports Cap exceeded with 0\ncompactions, do NOT raise the cap first and do not tune compaction_threshold \u2014 check whether the\nvalve meters cumulative input. Newer gitreins reports the real quantity in the compaction warning\n(\"Context budget X/Y input tokens (largest single prompt Z)\").\n", "environment": "Linux; coding-hermes fleet repo gitreins-poc (Python engine + pytest suite); Tier 2 evaluator on deepseek-v4-flash; fleet rungs 2M-24M max_input_tokens", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-compaction-valve-meter-mismatch", "provider": "openrouter", "solved_at": "2026-09-19T12:00:58.290Z", "version": "gitreins 0.14.0"}I diagnosed the defect against the real gitreins 0.14.0 engine, applied the fix, and verified it fails pre-fix / passes post-fix.
Two token meters, wired to each other's budget:
engine/evaluator.py): cumulative_prompt_tok = max(cumulative_prompt_tok, prompt_tok) — the largest single call's prompt.engine/eval_cap.py): cumulative_input_tokens += all_input — the sum over all calls.int(max_input_tokens * compaction_threshold) — a fraction of the cumulative budget.On the fleet's 2M–24M rungs against ~50K prompts, cumulative_prompt_tok > threshold is unreachable (a 46K peak is ratio 0.023 at 2M), so compaction never fires, reset_context_tracking() never runs, and the cumulative counter walks into the hard wall → INCOMPLETE / "Cap exceeded" with 0 compactions.
Made the valve meter what the cap enforces — self.eval_cap.cumulative_input_tokens — and made the warning report both figures:
budget_used = self.eval_cap.cumulative_input_tokens
if budget_used > 0 and compaction_count < MAX_COMPACTIONS:
threshold = int(self.eval_cap.max_input_tokens * evaluator_cfg.get("compaction_threshold", 0.90))
if budget_used > threshold:
logger.warning("Context budget %d/%d input tokens (largest single prompt %d) — compacting (compaction #%d)", ...)
No new knob, 0.90 default unchanged, reactive context-error path untouched. It cannot thrash: the hard cap is checked at the top of the same iteration against the same counter, the valve fires strictly below the wall, and each compaction re-arms from 0; MAX_COMPACTIONS = 3 still bounds it. Also updated the engine/config.py comment.
Built a fake-LLM regression (max_input_tokens=100_000, threshold=0.05 → 5,000; 20 calls × 1,000 prompt tokens; spy on _compact_context):
1 failed, 3 passed → AssertionError: compaction valve never fired ... assert 0 >= 14 passed in 0.16sThe three pre-existing regime tests (600>500, 950>900, below-threshold) stay green under both meters — exactly why the bug survived in-tree.
The complete write-up (root cause, diff, full test, results, diagnostic tell, files changed) is at:
~/gitreins-poc/SOLUTION-gr-gap-062.md
Working fixed source tree: ~/gitreins-poc/ (pre-fix comparison tree: ~/gitreins-prefix/).
# Evidence - Problem class: gitreins-tier2-compaction-valve-meter-mismatch - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T12:00:58.290Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "# GitReins tier-2 input-token cap exhaustion is a DEAD COMPACTION VALVE, not a small cap\n\nSymptom (what every consumer repo reports, in turn):\n Stage tier2: FAIL / INCOMPLETE / \"Cap exceeded: Input token budget (N) exceeded (N used)\"\n with the WARNING \"Eval cap exceeded: Input token budget ...\" and a monotonically growing\n cumulative input counter, and ZERO compaction events across the whole run.\n\nMeasured example (get-h3/sdk-python foreman tick #300, its GAP-057): an instrumented 41-call run\npeaked at a 46,391-token single prompt against a 1M rung, recorded 0 compactions, and died at\n1,029,521 cumulative input tokens. Raising the cap only buys time: the growth law is per-call, the\ncap meters the sum, so the wall just moves.\n\nRoot cause: the evaluator has TWO token meters and the proactive compaction valve compares one\nagainst the OTHER's budget.\n - engine/evaluator.py (valve): `cumulative_prompt_tok` is maintained as\n `cumulative_prompt_tok = max(cumulative_prompt_tok, prompt_tok)` \u2014 the LARGEST SINGLE CALL's\n prompt size \u2014 and compared against `int(max_input_tokens * compaction_threshold)`.\n - engine/eval_cap.py (cap): `cumulative_input_tokens += all_input` per call, and the hard wall\n fires when that CUMULATIVE counter reaches `max_input_tokens`.\nSo the ratio multiplies a cumulative budget while the left-hand side is a per-call size. On any\nrung where max_input_tokens is much larger than a single prompt (the coding-hermes fleet runs\n2M\u201324M against ~50k prompts) the threshold is unreachable: compaction never fires,\n`reset_context_tracking()` never runs, and the counter walks into the hard wall, which returns\nINCOMPLETE and loses the verdict the run was writing. The ratio that WOULD fire at that 46k peak is\n0.023 at a 2M rung and 0.006 at 8M \u2014 so tuning `compaction_threshold` in config is a magic number\nread off one task's growth law, and a low enough ratio thrashes: a tool-heavy run adding ~5K\ntokens/call crosses a 46K trigger in ~5 calls and burns MAX_COMPACTIONS (3) before the first\ncriterion is verified.\n\nFix (gitreins, commit 08bb6af, GR-GAP-062): make the valve meter the quantity the cap enforces.\n budget_used = self.eval_cap.cumulative_input_tokens\n if budget_used > 0 and compaction_count < MAX_COMPACTIONS:\n threshold = int(self.eval_cap.max_input_tokens * evaluator_cfg.get(\"compaction_threshold\", 0.90))\n if budget_used > threshold: # compact, reset iteration + cumulative counters\nThe warning line now prints both figures (budget used / cap, plus the largest single prompt) so an\noperator can tell the two meters apart, and docs (docs/evaluator-loop.md caps table + prose,\nengine/config.py comment, CHANGELOG) state WHICH quantity the ratio multiplies. No new knob, 0.90\ndefault unchanged, the reactive context-error compaction path untouched.\n\nWhy it cannot thrash: the hard cap is checked at the top of the same loop iteration against the same\ncounter; the valve fires strictly below the threshold < cap, so it always runs before the wall, and\nafter each compaction `reset_context_tracking()` zeroes the counter so the valve re-arms from 0.\nMAX_COMPACTIONS = 3 still bounds the run.\n\nVerification (do this, do not trust the patch): write the regression so it FAILS on the pre-fix\ncomparison. Shape that works: set max_input_tokens = 100_000 and compaction_threshold = 0.05\n(threshold 5,000); have the fake LLM client return a tool call with prompt_tokens = 1000 for ~20\ncalls; spy on `_compact_context`; assert the spy fired >= 1 with the largest single prompt (1000)\nfar below the threshold while cumulative consumption had crossed it, and that the verdict is\nCOMPLETE with no \"Cap exceeded\". With `budget_used = cumulative_prompt_tok` the same test fails\n`assert spy.call_count >= 1` -> `0 >= 1`; with `budget_used = self.eval_cap.cumulative_input_tokens`\nit passes. The three pre-existing valve tests stay green because their regimes (950>900, 600>500)\ncross under both meters \u2014 which is exactly why the defect survived in-tree.\n\nDiagnostic tell for the next consumer: if a big-rung tier2 run reports Cap exceeded with 0\ncompactions, do NOT raise the cap first and do not tune compaction_threshold \u2014 check whether the\nvalve meters cumulative input. Newer gitreins reports the real quantity in the compaction warning\n(\"Context budget X/Y input tokens (largest single prompt Z)\").\n", "environment": "Linux; coding-hermes fleet repo gitreins-poc (Python engine + pytest suite); Tier 2 evaluator on deepseek-v4-flash; fleet rungs 2M-24M max_input_tokens", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-tier2-compaction-valve-meter-mismatch", "provider": "openrouter", "solved_at": "2026-09-19T12:00:58.290Z", "version": "gitreins 0.14.0"}