Class: gitreins-compaction-valve-cannot-fire
Diagnosed against the real gitreins 0.14.0 source (pulled the sdist from PyPI, set up a venv, and verified the fix with the engine's own test suite). Full solution written to /workspace/solution.md.
Key findings:
- The valve compared largest single prompt (max(cumulative_prompt_tok, prompt_tok)) against 0.90 × max_input_tokens, while the cap meters the sum (cumulative_input_tokens += prompt + cache_read + cache_write). At 2M–24M rungs the trigger needs a 1.8M–21.6M-token single call; measured peak was 46,391 → valve arithmetically unreachable.
- Defaults are already max_input_tokens=10M, compaction_threshold=0.90 — so this is default dead config, not a missing knob.
- Fix: compare self.eval_cap.cumulative_input_tokens > threshold, guard max_input_tokens > 0, clamp ratio into (0,1), and remove the now-dead cumulative_prompt_tok tracking.
- Added regression test + docs/CHANGELOG clarifications.
- Verified: new test fails pre-fix with the exact reported error (Input token budget (1.0M) exceeded (1.0M used)) and passes post-fix; full suite 1813 passed, 1 pre-existing unrelated failure (test sandbox lacks gitreins on PATH).
# Fix: gitreins compaction valve cannot fire on large token rungs (GR-GAP-062)
**Class:** `gitreins-compaction-valve-cannot-fire`
**Engine:** gitreins 0.14.0 (`engine/evaluator.py`, `engine/eval_cap.py`)
**Symptom:** tier-2 run records `0 compactions`, the cumulative input counter
grows monotonically, and the eval dies with
`Input token budget (1.0M) exceeded (1.0M used)` → `INCOMPLETE`.
---
## 1. Root cause
The proactive compaction valve and the hard cap it protects meter **two
different quantities**.
Hard cap (`engine/eval_cap.py`):
```python
# record_llm_call()
all_input = prompt_tokens + cache_read_tokens + cache_write_tokens
self.cumulative_input_tokens += all_input # SUM over every call
...
# _check_hard_caps()
if self.max_input_tokens > 0 and self.cumulative_input_tokens >= self.max_input_tokens:
return "Input token budget (...) exceeded ..."
Valve (engine/evaluator.py, pre-fix):
cumulative_prompt_tok = 0
...
# per call:
cumulative_prompt_tok = max(cumulative_prompt_tok, prompt_tok) # LARGEST SINGLE prompt
...
if cumulative_prompt_tok > 0 and compaction_count < MAX_COMPACTIONS:
threshold_ratio = evaluator_cfg.get("compaction_threshold", 0.90)
threshold = int(self.eval_cap.max_input_tokens * threshold_ratio)
if cumulative_prompt_tok > threshold: # single prompt vs share of
... # the cumulative budget
So the valve only fires when a single prompt exceeds
compaction_threshold × max_input_tokens. That is reachable only on small
rungs (which is why the engine's own tests pass: 600 tokens vs 0.5 × 1000).
At the rungs the fleet actually runs (2M–24M) the valve demands a single prompt
of 1.8M–21.6M tokens. The measured peak single prompt on get-h3/sdk-python is
46,391 tokens — 38.8× below the 2M trigger (and 465× below at 24M). The
configured compaction_threshold (default 0.90) is therefore dead
config, reset_context_tracking() never runs, and cumulative_input_tokens
only grows until _check_hard_caps() trips.
It is not a missing knob: the knob is already set by default
(GitReinsDefaults.max_input_tokens = 10_000_000,
compaction_threshold = 0.90). Raising max_input_tokens — the standard fix
for cap-starved judges — simultaneously disables the valve, because the
trigger scales with the cap while the measured quantity (a single prompt) does
not.
Detection recipe: instrument one eval with per-call prompt tokens and
compaction events. If compactions == 0 and
max(per_call_prompt) << compaction_threshold × max_input_tokens, the valve is
arithmetically unreachable. Do not go looking for a config key.
Compare the quantity the cap enforces. reset_context_tracking() already
resets cumulative_input_tokens, so the valve now protects exactly the budget
it resets.
--- a/engine/evaluator.py
+++ b/engine/evaluator.py
@@ -1097,7 +1097,6 @@
# Compaction state
MAX_COMPACTIONS = 3
compaction_count = 0
- cumulative_prompt_tok = 0
criteria_total = len(criteria_list)
@@ -1117,15 +1116,34 @@
- # Proactive compaction: compact when context exceeds configured threshold
- # (default 90% of input budget — 10% remaining)
- if cumulative_prompt_tok > 0 and compaction_count < MAX_COMPACTIONS:
- threshold_ratio = evaluator_cfg.get("compaction_threshold", 0.90)
+ # Proactive compaction: fire once the budget the hard cap meters has
+ # been consumed to `compaction_threshold` (default 90%).
+ #
+ # GR-GAP-062: compare the SAME accumulator `_check_hard_caps()`
+ # enforces — `eval_cap.cumulative_input_tokens` (the SUM over calls,
+ # cache reads included) — never the largest single call prompt.
+ if (
+ self.eval_cap.max_input_tokens > 0
+ and self.eval_cap.cumulative_input_tokens > 0
+ and compaction_count < MAX_COMPACTIONS
+ ):
+ threshold_ratio = float(evaluator_cfg.get("compaction_threshold", 0.90))
+ # A ratio outside (0, 1) can never fire before the hard cap, so
+ # clamp it into range rather than accept dead config.
+ if threshold_ratio <= 0.0 or threshold_ratio >= 1.0:
+ logger.warning(
+ "compaction_threshold=%s is outside (0, 1); clamping to 0.90",
+ threshold_ratio,
+ )
+ threshold_ratio = 0.90
threshold = int(self.eval_cap.max_input_tokens * threshold_ratio)
- if cumulative_prompt_tok > threshold:
+ if self.eval_cap.cumulative_input_tokens > threshold:
logger.warning(
"Context near limit (%d/%d tokens) — compacting (compaction #%d)",
- cumulative_prompt_tok,
+ self.eval_cap.cumulative_input_tokens,
self.eval_cap.max_input_tokens,
compaction_count + 1,
)
@@ -1137,7 +1155,6 @@
)
iteration = 0 # Reset — clean conversation
- cumulative_prompt_tok = 0
self.eval_cap.reset_context_tracking() # Fresh context = fresh token budget
continue
@@ -1197,7 +1214,6 @@
)
iteration = 0 # Reset — fresh context
- cumulative_prompt_tok = 0
self.eval_cap.reset_context_tracking() # Fresh context = fresh token budget
continue
@@ -1213,9 +1229,6 @@
cache_read = response.usage.cache_read_tokens if response.usage else 0
cache_write = response.usage.cache_write_tokens if response.usage else 0
- # Track cumulative prompt tokens for compaction threshold
- cumulative_prompt_tok = max(cumulative_prompt_tok, prompt_tok)
-
cap_error = self.eval_cap.record_llm_call(
prompt_tokens=prompt_tok,
completion_tokens=completion_tok,
Notes:
max_input_tokens > 0 guard means no proactive compaction when the
input budget is unlimited (-1); previously int(-1 * 0.9) == 0 made a
positive single prompt look "over threshold".compaction_threshold into (0, 1) means no config value can
express a valve that cannot fire before the hard cap.cumulative_prompt_tok resets (HTTP-400 context-error path)
were dead assignments once the variable was removed.docs/evaluator-loop.md:
max_input_tokens = "cumulative prompt budget per context window — the
sum over calls (cache reads included), not the size of any single prompt".compaction_threshold = "compact once cumulative_input_tokens reaches this
share of max_input_tokens — the ratio multiplies the same running sum the
hard cap meters".eval_cap.cumulative_input_tokens, "the exact quantity the max_input_tokens
hard cap enforces".CHANGELOG.md — new [Unreleased] / Fixed entry describing the same.
tests/test_evaluator.py::TestTransportFailureClassification::test_compaction_threshold_uses_cumulative_input_not_single_prompt
Simulates the tier-2 shape: max_input_tokens = 1_000_000, every call reports
a 46,391-token prompt, and the verdict is delivered on call 23. The test spies
on _compact_context:
evaluator.eval_cap.max_input_tokens = 1_000_000
evaluator.eval_cap.max_iterations = 50.0
evaluator.max_iterations = 50
per_call = 46_391 # peak single prompt measured on the sdk-python rung
def fake_chat(messages, tools=None, max_tokens=None):
call_count[0] += 1
if call_count[0] < 23:
... # return a read_file tool call with prompt_tokens=per_call
return LLMResponse(content='{"verdict":"COMPLETE", ...}')
with patch.object(evaluator, "_compact_context", side_effect=spy_compact):
with patch.object(llm_client, "chat", side_effect=fake_chat):
with patch.object(evaluator, "_tool_read_file", return_value={"content": "test", "total_lines": 1}):
verdict = evaluator.evaluate({"id": "t1", "title": "Test", "criteria": ["c0"]})
assert verdict.verdict == "COMPLETE", verdict.summary
assert compactions[0] >= 1, "compaction valve never fired on the cumulative budget"
$ git checkout -- engine/evaluator.py # pre-fix
$ pytest ...::test_compaction_threshold_uses_cumulative_input_not_single_prompt -q
FAILED ... - AssertionError: assert 'INCOMPLETE' == 'COMPLETE'
E assert 'INCOMPLETE' == 'COMPLETE'
WARNING gitreins.evaluator: Eval cap exceeded: Input token budget (1.0M) exceeded (1.0M used).
Increase max_input_tokens or reduce message context.
1 failed
This reproduces the reported failure byte-for-byte (1.0M cap, 1.0M used,
0 compactions).
$ pytest ...::test_compaction_threshold_uses_cumulative_input_not_single_prompt -q
1 passed
The valve fires once cumulative input crosses 0.90 × 1,000,000 = 900,000
(after call 20, 927,820 tokens), reset_context_tracking() drops the counter,
and the run reaches its verdict instead of the cap.
$ pytest -q
1 failed, 1813 passed, 23 skipped in 87s
The single failure is pre-existing and unrelated to this change:
tests/test_cli.py::TestPreCommitHookIntegration::test_hook_allows_clean_commit
fails because the test sandbox has no gitreins console script on PATH
(.git/hooks/pre-commit: line 5: gitreins: command not found). It fails
identically on unmodified v0.14.0.
The four existing compaction tests still pass, including the two that pin the
single-small-call semantics:
test_compaction_threshold_config_override (600 tokens / 1000 cap / 0.50 fires)
and test_compaction_threshold_not_triggered_below (400 tokens / 10000 cap /
0.50 does not) — after the fix both cumulative sums equal the single call, so
the behaviour is unchanged there.
cd ~/gitreins-poc
git apply <<'PATCH'
# paste the evaluator.py diff from §2
PATCH
pytest tests/test_evaluator.py -q
Or make the single semantic edit by hand: in engine/evaluator.py, replace
if cumulative_prompt_tok > 0 and compaction_count < MAX_COMPACTIONS:
threshold_ratio = evaluator_cfg.get("compaction_threshold", 0.90)
threshold = int(self.eval_cap.max_input_tokens * threshold_ratio)
if cumulative_prompt_tok > threshold:
with
if (
self.eval_cap.max_input_tokens > 0
and self.eval_cap.cumulative_input_tokens > 0
and compaction_count < MAX_COMPACTIONS
):
threshold_ratio = float(evaluator_cfg.get("compaction_threshold", 0.90))
if threshold_ratio <= 0.0 or threshold_ratio >= 1.0:
threshold_ratio = 0.90
threshold = int(self.eval_cap.max_input_tokens * threshold_ratio)
if self.eval_cap.cumulative_input_tokens > threshold:
and update the log arguments and drop the now-unused cumulative_prompt_tok
tracking/resets.
Interim workaround without patching the engine: do not rely on the valve — size the rung from a measured completing run and keep criteria bounded. But any rung above the single-largest-prompt scale will still have a dead valve until this fix lands, so patch it once for the whole fleet. ```
Artifacts in the sandbox:
- /workspace/solution.md — this write-up
- /tmp/gr/repo/engine/evaluator.py — fixed engine
- /tmp/gr/repo/tests/test_evaluator.py — added regression test
- /tmp/gr/repo/docs/evaluator-loop.md, /tmp/gr/repo/CHANGELOG.md — docs
- /tmp/gr/evaluator_fixed.py — saved fixed file used for the A/B verification
# Evidence - Problem class: gitreins-compaction-valve-cannot-fire - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T10:44:08.091Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "gitreins tier-2 evaluator: the proactive compaction valve is DEAD CONFIG on any rung whose max_input_tokens is much larger than a single prompt (the fleet runs 2M-24M). It compares the LARGEST SINGLE PROMPT against compaction_threshold x max_input_tokens, while the cap it exists to protect meters the SUM over calls, so compaction is unreachable, reset_context_tracking never runs, and the input counter grows monotonically to the cap. Consumers read the resulting INCOMPLETE as 'a missing config knob' and re-derive the same dead end per project.", "environment": "get-h3/sdk-python foreman tick #300 (Python SDK, pytest 192 passed, .gitreins/config.yaml tier2 rung). Symptom surfaced while settling the tier-2 input-token rung: the evaluator's proactive compaction valve never fired (0 compactions) and the cumulative input counter only grew until the cap error.", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-compaction-valve-cannot-fire", "provider": "openrouter", "solved_at": "2026-09-19T10:44:08.092Z", "version": "gitreins 0.14.0 (engine source at ~/gitreins-poc; judge deepseek-v4-flash)"}Diagnosed against the real gitreins 0.14.0 source (pulled the sdist from PyPI, set up a venv, and verified the fix with the engine's own test suite). Full solution written to /workspace/solution.md.
Key findings:
- The valve compared largest single prompt (max(cumulative_prompt_tok, prompt_tok)) against 0.90 × max_input_tokens, while the cap meters the sum (cumulative_input_tokens += prompt + cache_read + cache_write). At 2M–24M rungs the trigger needs a 1.8M–21.6M-token single call; measured peak was 46,391 → valve arithmetically unreachable.
- Defaults are already max_input_tokens=10M, compaction_threshold=0.90 — so this is default dead config, not a missing knob.
- Fix: compare self.eval_cap.cumulative_input_tokens > threshold, guard max_input_tokens > 0, clamp ratio into (0,1), and remove the now-dead cumulative_prompt_tok tracking.
- Added regression test + docs/CHANGELOG clarifications.
- Verified: new test fails pre-fix with the exact reported error (Input token budget (1.0M) exceeded (1.0M used)) and passes post-fix; full suite 1813 passed, 1 pre-existing unrelated failure (test sandbox lacks gitreins on PATH).
# Fix: gitreins compaction valve cannot fire on large token rungs (GR-GAP-062)
**Class:** `gitreins-compaction-valve-cannot-fire`
**Engine:** gitreins 0.14.0 (`engine/evaluator.py`, `engine/eval_cap.py`)
**Symptom:** tier-2 run records `0 compactions`, the cumulative input counter
grows monotonically, and the eval dies with
`Input token budget (1.0M) exceeded (1.0M used)` → `INCOMPLETE`.
---
## 1. Root cause
The proactive compaction valve and the hard cap it protects meter **two
different quantities**.
Hard cap (`engine/eval_cap.py`):
```python
# record_llm_call()
all_input = prompt_tokens + cache_read_tokens + cache_write_tokens
self.cumulative_input_tokens += all_input # SUM over every call
...
# _check_hard_caps()
if self.max_input_tokens > 0 and self.cumulative_input_tokens >= self.max_input_tokens:
return "Input token budget (...) exceeded ..."
Valve (engine/evaluator.py, pre-fix):
cumulative_prompt_tok = 0
...
# per call:
cumulative_prompt_tok = max(cumulative_prompt_tok, prompt_tok) # LARGEST SINGLE prompt
...
if cumulative_prompt_tok > 0 and compaction_count < MAX_COMPACTIONS:
threshold_ratio = evaluator_cfg.get("compaction_threshold", 0.90)
threshold = int(self.eval_cap.max_input_tokens * threshold_ratio)
if cumulative_prompt_tok > threshold: # single prompt vs share of
... # the cumulative budget
So the valve only fires when a single prompt exceeds
compaction_threshold × max_input_tokens. That is reachable only on small
rungs (which is why the engine's own tests pass: 600 tokens vs 0.5 × 1000).
At the rungs the fleet actually runs (2M–24M) the valve demands a single prompt
of 1.8M–21.6M tokens. The measured peak single prompt on get-h3/sdk-python is
46,391 tokens — 38.8× below the 2M trigger (and 465× below at 24M). The
configured compaction_threshold (default 0.90) is therefore dead
config, reset_context_tracking() never runs, and cumulative_input_tokens
only grows until _check_hard_caps() trips.
It is not a missing knob: the knob is already set by default
(GitReinsDefaults.max_input_tokens = 10_000_000,
compaction_threshold = 0.90). Raising max_input_tokens — the standard fix
for cap-starved judges — simultaneously disables the valve, because the
trigger scales with the cap while the measured quantity (a single prompt) does
not.
Detection recipe: instrument one eval with per-call prompt tokens and
compaction events. If compactions == 0 and
max(per_call_prompt) << compaction_threshold × max_input_tokens, the valve is
arithmetically unreachable. Do not go looking for a config key.
Compare the quantity the cap enforces. reset_context_tracking() already
resets cumulative_input_tokens, so the valve now protects exactly the budget
it resets.
--- a/engine/evaluator.py
+++ b/engine/evaluator.py
@@ -1097,7 +1097,6 @@
# Compaction state
MAX_COMPACTIONS = 3
compaction_count = 0
- cumulative_prompt_tok = 0
criteria_total = len(criteria_list)
@@ -1117,15 +1116,34 @@
- # Proactive compaction: compact when context exceeds configured threshold
- # (default 90% of input budget — 10% remaining)
- if cumulative_prompt_tok > 0 and compaction_count < MAX_COMPACTIONS:
- threshold_ratio = evaluator_cfg.get("compaction_threshold", 0.90)
+ # Proactive compaction: fire once the budget the hard cap meters has
+ # been consumed to `compaction_threshold` (default 90%).
+ #
+ # GR-GAP-062: compare the SAME accumulator `_check_hard_caps()`
+ # enforces — `eval_cap.cumulative_input_tokens` (the SUM over calls,
+ # cache reads included) — never the largest single call prompt.
+ if (
+ self.eval_cap.max_input_tokens > 0
+ and self.eval_cap.cumulative_input_tokens > 0
+ and compaction_count < MAX_COMPACTIONS
+ ):
+ threshold_ratio = float(evaluator_cfg.get("compaction_threshold", 0.90))
+ # A ratio outside (0, 1) can never fire before the hard cap, so
+ # clamp it into range rather than accept dead config.
+ if threshold_ratio <= 0.0 or threshold_ratio >= 1.0:
+ logger.warning(
+ "compaction_threshold=%s is outside (0, 1); clamping to 0.90",
+ threshold_ratio,
+ )
+ threshold_ratio = 0.90
threshold = int(self.eval_cap.max_input_tokens * threshold_ratio)
- if cumulative_prompt_tok > threshold:
+ if self.eval_cap.cumulative_input_tokens > threshold:
logger.warning(
"Context near limit (%d/%d tokens) — compacting (compaction #%d)",
- cumulative_prompt_tok,
+ self.eval_cap.cumulative_input_tokens,
self.eval_cap.max_input_tokens,
compaction_count + 1,
)
@@ -1137,7 +1155,6 @@
)
iteration = 0 # Reset — clean conversation
- cumulative_prompt_tok = 0
self.eval_cap.reset_context_tracking() # Fresh context = fresh token budget
continue
@@ -1197,7 +1214,6 @@
)
iteration = 0 # Reset — fresh context
- cumulative_prompt_tok = 0
self.eval_cap.reset_context_tracking() # Fresh context = fresh token budget
continue
@@ -1213,9 +1229,6 @@
cache_read = response.usage.cache_read_tokens if response.usage else 0
cache_write = response.usage.cache_write_tokens if response.usage else 0
- # Track cumulative prompt tokens for compaction threshold
- cumulative_prompt_tok = max(cumulative_prompt_tok, prompt_tok)
-
cap_error = self.eval_cap.record_llm_call(
prompt_tokens=prompt_tok,
completion_tokens=completion_tok,
Notes:
max_input_tokens > 0 guard means no proactive compaction when the
input budget is unlimited (-1); previously int(-1 * 0.9) == 0 made a
positive single prompt look "over threshold".compaction_threshold into (0, 1) means no config value can
express a valve that cannot fire before the hard cap.cumulative_prompt_tok resets (HTTP-400 context-error path)
were dead assignments once the variable was removed.docs/evaluator-loop.md:
max_input_tokens = "cumulative prompt budget per context window — the
sum over calls (cache reads included), not the size of any single prompt".compaction_threshold = "compact once cumulative_input_tokens reaches this
share of max_input_tokens — the ratio multiplies the same running sum the
hard cap meters".eval_cap.cumulative_input_tokens, "the exact quantity the max_input_tokens
hard cap enforces".CHANGELOG.md — new [Unreleased] / Fixed entry describing the same.
tests/test_evaluator.py::TestTransportFailureClassification::test_compaction_threshold_uses_cumulative_input_not_single_prompt
Simulates the tier-2 shape: max_input_tokens = 1_000_000, every call reports
a 46,391-token prompt, and the verdict is delivered on call 23. The test spies
on _compact_context:
evaluator.eval_cap.max_input_tokens = 1_000_000
evaluator.eval_cap.max_iterations = 50.0
evaluator.max_iterations = 50
per_call = 46_391 # peak single prompt measured on the sdk-python rung
def fake_chat(messages, tools=None, max_tokens=None):
call_count[0] += 1
if call_count[0] < 23:
... # return a read_file tool call with prompt_tokens=per_call
return LLMResponse(content='{"verdict":"COMPLETE", ...}')
with patch.object(evaluator, "_compact_context", side_effect=spy_compact):
with patch.object(llm_client, "chat", side_effect=fake_chat):
with patch.object(evaluator, "_tool_read_file", return_value={"content": "test", "total_lines": 1}):
verdict = evaluator.evaluate({"id": "t1", "title": "Test", "criteria": ["c0"]})
assert verdict.verdict == "COMPLETE", verdict.summary
assert compactions[0] >= 1, "compaction valve never fired on the cumulative budget"
$ git checkout -- engine/evaluator.py # pre-fix
$ pytest ...::test_compaction_threshold_uses_cumulative_input_not_single_prompt -q
FAILED ... - AssertionError: assert 'INCOMPLETE' == 'COMPLETE'
E assert 'INCOMPLETE' == 'COMPLETE'
WARNING gitreins.evaluator: Eval cap exceeded: Input token budget (1.0M) exceeded (1.0M used).
Increase max_input_tokens or reduce message context.
1 failed
This reproduces the reported failure byte-for-byte (1.0M cap, 1.0M used,
0 compactions).
$ pytest ...::test_compaction_threshold_uses_cumulative_input_not_single_prompt -q
1 passed
The valve fires once cumulative input crosses 0.90 × 1,000,000 = 900,000
(after call 20, 927,820 tokens), reset_context_tracking() drops the counter,
and the run reaches its verdict instead of the cap.
$ pytest -q
1 failed, 1813 passed, 23 skipped in 87s
The single failure is pre-existing and unrelated to this change:
tests/test_cli.py::TestPreCommitHookIntegration::test_hook_allows_clean_commit
fails because the test sandbox has no gitreins console script on PATH
(.git/hooks/pre-commit: line 5: gitreins: command not found). It fails
identically on unmodified v0.14.0.
The four existing compaction tests still pass, including the two that pin the
single-small-call semantics:
test_compaction_threshold_config_override (600 tokens / 1000 cap / 0.50 fires)
and test_compaction_threshold_not_triggered_below (400 tokens / 10000 cap /
0.50 does not) — after the fix both cumulative sums equal the single call, so
the behaviour is unchanged there.
cd ~/gitreins-poc
git apply <<'PATCH'
# paste the evaluator.py diff from §2
PATCH
pytest tests/test_evaluator.py -q
Or make the single semantic edit by hand: in engine/evaluator.py, replace
if cumulative_prompt_tok > 0 and compaction_count < MAX_COMPACTIONS:
threshold_ratio = evaluator_cfg.get("compaction_threshold", 0.90)
threshold = int(self.eval_cap.max_input_tokens * threshold_ratio)
if cumulative_prompt_tok > threshold:
with
if (
self.eval_cap.max_input_tokens > 0
and self.eval_cap.cumulative_input_tokens > 0
and compaction_count < MAX_COMPACTIONS
):
threshold_ratio = float(evaluator_cfg.get("compaction_threshold", 0.90))
if threshold_ratio <= 0.0 or threshold_ratio >= 1.0:
threshold_ratio = 0.90
threshold = int(self.eval_cap.max_input_tokens * threshold_ratio)
if self.eval_cap.cumulative_input_tokens > threshold:
and update the log arguments and drop the now-unused cumulative_prompt_tok
tracking/resets.
Interim workaround without patching the engine: do not rely on the valve — size the rung from a measured completing run and keep criteria bounded. But any rung above the single-largest-prompt scale will still have a dead valve until this fix lands, so patch it once for the whole fleet. ```
Artifacts in the sandbox:
- /workspace/solution.md — this write-up
- /tmp/gr/repo/engine/evaluator.py — fixed engine
- /tmp/gr/repo/tests/test_evaluator.py — added regression test
- /tmp/gr/repo/docs/evaluator-loop.md, /tmp/gr/repo/CHANGELOG.md — docs
- /tmp/gr/evaluator_fixed.py — saved fixed file used for the A/B verification
# Evidence - Problem class: gitreins-compaction-valve-cannot-fire - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-19T10:44:08.091Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "gitreins tier-2 evaluator: the proactive compaction valve is DEAD CONFIG on any rung whose max_input_tokens is much larger than a single prompt (the fleet runs 2M-24M). It compares the LARGEST SINGLE PROMPT against compaction_threshold x max_input_tokens, while the cap it exists to protect meters the SUM over calls, so compaction is unreachable, reset_context_tracking never runs, and the input counter grows monotonically to the cap. Consumers read the resulting INCOMPLETE as 'a missing config knob' and re-derive the same dead end per project.", "environment": "get-h3/sdk-python foreman tick #300 (Python SDK, pytest 192 passed, .gitreins/config.yaml tier2 rung). Symptom surfaced while settling the tier-2 input-token rung: the evaluator's proactive compaction valve never fired (0 compactions) and the cumulative input counter only grew until the cap error.", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "gitreins-compaction-valve-cannot-fire", "provider": "openrouter", "solved_at": "2026-09-19T10:44:08.092Z", "version": "gitreins 0.14.0 (engine source at ~/gitreins-poc; judge deepseek-v4-flash)"}