if step.get("maxiterations") == UNLIMITED:
The bug: _run_ai_eval forwarded only max_iterations into AgenticEvaluator; every token cap silently defaulted to -1, so the compaction threshold became int(-1 * 0.9) == 0 — the evaluator compacted on every turn and never emitted a verdict (fleet-wide tier2 INCOMPLETE). The fix is a single build function, _build_eval_cap, implementing the correct pattern — config base + explicit step overrides merged, strings parsed, deferring steps pass eval_cap=None:
UNLIMITED = -1
COMPACTION_FACTOR = 0.9
def _parse_tokens(value):
"""'200k' -> 200_000, '1.5m' -> 1_500_000, '500' -> 500, -1/None -> -1."""
if value is None or isinstance(value, int) and not isinstance(value, bool):
return value if value is not None else UNLIMITED
if isinstance(value, str):
m = re.fullmatch(r"\s*(\d+(?:\.\d+)?)\s*([km]?)\s*", value, re.IGNORECASE)
if not m:
raise ValueError(f"invalid token cap string: {value!r}")
amount = float(m.group(1)) * {"": 1, "k": 1_000, "m": 1_000_000}[m.group(2).lower()]
if amount != int(amount):
raise ValueError(f"token cap must be a whole number: {value!r}")
return int(amount)
raise TypeError(f"token cap must be int or str, got {type(value).__name__}: {value!r}")
@dataclass(frozen=True)
class EvalCap:
max_iterations: int = UNLIMITED # all fields int-typed
input_token_cap: int = UNLIMITED
output_token_cap: int = UNLIMITED
total_token_cap: int = UNLIMITED
def __post_init__(self): # Trap 2: strings re-parsed on construction
for f in ("input_token_cap", "output_token_cap", "total_token_cap"):
object.__setattr__(self, f, _parse_tokens(getattr(self, f)))
object.__setattr__(self, "max_iterations", int(self.max_iterations))
@property
def is_unlimited(self):
return all(getattr(self, f) == UNLIMITED
for f in ("max_iterations", *("input_token_cap", "output_token_cap", "total_token_cap")))
@property
def compaction_threshold(self):
cap = self.total_token_cap if self.total_token_cap != UNLIMITED else self.input_token_cap
return UNLIMITED if cap == UNLIMITED else int(cap * 0.9) # never 0 from -1
def _build_eval_cap(config, step):
# Trap 1: a deferring step must never materialize an explicit all--1
# EvalCap (is_unlimited -> same compaction loop). None => read config.yaml.
if step.get("max_iterations") == UNLIMITED:
return None
# Correct pattern: config base + explicit step overrides merged.
step_caps = dict(step.get("eval_cap") or {})
if "max_iterations" in step and "max_iterations" not in step_caps:
step_caps["max_iterations"] = step["max_iterations"]
merged = dict(config.get("eval_cap") or {})
merged.update(step_caps) # step wins over base
return EvalCap(
max_iterations=int(merged.get("max_iterations", UNLIMITED)),
input_token_cap=_parse_tokens(merged.get("input_token_cap", UNLIMITED)),
output_token_cap=_parse_tokens(merged.get("output_token_cap", UNLIMITED)),
total_token_cap=_parse_tokens(merged.get("total_token_cap", UNLIMITED)),
)
def _run_ai_eval(config, step, *, config_path="config.yaml", **kw):
return AgenticEvaluator(eval_cap=_build_eval_cap(config, step),
config_path=config_path).run(**kw)
Two traps handled: (1) a step with max_iterations: -1 yields None — never an explicit EvalCap whose all--1 caps are is_unlimited and trip the same loop; the evaluator falls back to reading its own config.yaml (total_token_cap: "200k" → threshold 180 000). (2) string caps ("200k") pass through _parse_tokens (also enforced in EvalCap.__post_init__) so int-typed fields can never receive a string.
Full files: ~/eval_cap_fix/eval_caps.py, config.yaml, test_eval_caps.py.
`python3 -m pytest -q` → **32 passed, 0 failed** (Python 3.14.4, pytest 9.0.2).
- **Loop reproduced before fix:** `BuggyPipeline` (old forwarding: only `max_iterations`, naive `int(-1*0.9)==0` threshold) → `INCOMPLETE` for *both* an explicit step (`max_iterations: 5`) and a deferring step — fleet-wide as reported.
- **Fix 1 verified:** `_build_eval_cap({}, {"max_iterations": -1}) is None`; `_run_ai_eval` then reads `config.yaml` (total `"200k"`) → threshold 180 000 → verdict `passed` on turn 1. Also: a deferring step returns `None` even when the step carries `eval_cap` or a config base exists.
- **Fix 2 verified:** `_parse_tokens`: `"200k"→200_000`, `"1.5m"→1_500_000`, `"500"→500`, `" 200k "→200_000`, int passthrough, `None/-1→-1`; raises on `""`, `"200kb"`, `"1,000"`, bools, floats; `EvalCap(total_token_cap="200k").total_token_cap == 200_000` and `isinstance(..., int)`; threshold `int(200_000*0.9) == 180_000 > 0`; all-`-1` cap → threshold `-1` (no forced compaction), `is_unlimited` True.
- **Merge verified:** base `{max_iterations: 5, input: "100k", total: "200k"}` + step `{max_iterations: 10, eval_cap: {output: "50k"}}` → `max_iterations=10` (step wins), `input=100_000`/`total=200_000` (inherited & parsed), `output=50_000` → end-to-end `passed`.
- **Edge cases:** empty config+step → `None`; `EvalCap(max_iterations=-2)` → `ValueError`; `eval_cap=None` without `config_path` → `ValueError`. Degenerate real caps (`cap=1` → threshold 0) are a config error, not the `-1` default bug.{"model": "deepseek-v4-flash", "problem_class": "python-evaluator-compaction-loop-cap-forwarding", "result": "passed", "tests": 32}The bug: _run_ai_eval forwarded only max_iterations into AgenticEvaluator; every token cap silently defaulted to -1, so the compaction threshold became int(-1 * 0.9) == 0 — the evaluator compacted on every turn and never emitted a verdict (fleet-wide tier2 INCOMPLETE). The fix is a single build function, _build_eval_cap, implementing the correct pattern — config base + explicit step overrides merged, strings parsed, deferring steps pass eval_cap=None:
UNLIMITED = -1
COMPACTION_FACTOR = 0.9
def _parse_tokens(value):
"""'200k' -> 200_000, '1.5m' -> 1_500_000, '500' -> 500, -1/None -> -1."""
if value is None or isinstance(value, int) and not isinstance(value, bool):
return value if value is not None else UNLIMITED
if isinstance(value, str):
m = re.fullmatch(r"\s*(\d+(?:\.\d+)?)\s*([km]?)\s*", value, re.IGNORECASE)
if not m:
raise ValueError(f"invalid token cap string: {value!r}")
amount = float(m.group(1)) * {"": 1, "k": 1_000, "m": 1_000_000}[m.group(2).lower()]
if amount != int(amount):
raise ValueError(f"token cap must be a whole number: {value!r}")
return int(amount)
raise TypeError(f"token cap must be int or str, got {type(value).__name__}: {value!r}")
@dataclass(frozen=True)
class EvalCap:
max_iterations: int = UNLIMITED # all fields int-typed
input_token_cap: int = UNLIMITED
output_token_cap: int = UNLIMITED
total_token_cap: int = UNLIMITED
def __post_init__(self): # Trap 2: strings re-parsed on construction
for f in ("input_token_cap", "output_token_cap", "total_token_cap"):
object.__setattr__(self, f, _parse_tokens(getattr(self, f)))
object.__setattr__(self, "max_iterations", int(self.max_iterations))
@property
def is_unlimited(self):
return all(getattr(self, f) == UNLIMITED
for f in ("max_iterations", *("input_token_cap", "output_token_cap", "total_token_cap")))
@property
def compaction_threshold(self):
cap = self.total_token_cap if self.total_token_cap != UNLIMITED else self.input_token_cap
return UNLIMITED if cap == UNLIMITED else int(cap * 0.9) # never 0 from -1
def _build_eval_cap(config, step):
# Trap 1: a deferring step must never materialize an explicit all--1
# EvalCap (is_unlimited -> same compaction loop). None => read config.yaml.
if step.get("max_iterations") == UNLIMITED:
return None
# Correct pattern: config base + explicit step overrides merged.
step_caps = dict(step.get("eval_cap") or {})
if "max_iterations" in step and "max_iterations" not in step_caps:
step_caps["max_iterations"] = step["max_iterations"]
merged = dict(config.get("eval_cap") or {})
merged.update(step_caps) # step wins over base
return EvalCap(
max_iterations=int(merged.get("max_iterations", UNLIMITED)),
input_token_cap=_parse_tokens(merged.get("input_token_cap", UNLIMITED)),
output_token_cap=_parse_tokens(merged.get("output_token_cap", UNLIMITED)),
total_token_cap=_parse_tokens(merged.get("total_token_cap", UNLIMITED)),
)
def _run_ai_eval(config, step, *, config_path="config.yaml", **kw):
return AgenticEvaluator(eval_cap=_build_eval_cap(config, step),
config_path=config_path).run(**kw)
Two traps handled: (1) a step with max_iterations: -1 yields None — never an explicit EvalCap whose all--1 caps are is_unlimited and trip the same loop; the evaluator falls back to reading its own config.yaml (total_token_cap: "200k" → threshold 180 000). (2) string caps ("200k") pass through _parse_tokens (also enforced in EvalCap.__post_init__) so int-typed fields can never receive a string.
Full files: ~/eval_cap_fix/eval_caps.py, config.yaml, test_eval_caps.py.
`python3 -m pytest -q` → **32 passed, 0 failed** (Python 3.14.4, pytest 9.0.2).
- **Loop reproduced before fix:** `BuggyPipeline` (old forwarding: only `max_iterations`, naive `int(-1*0.9)==0` threshold) → `INCOMPLETE` for *both* an explicit step (`max_iterations: 5`) and a deferring step — fleet-wide as reported.
- **Fix 1 verified:** `_build_eval_cap({}, {"max_iterations": -1}) is None`; `_run_ai_eval` then reads `config.yaml` (total `"200k"`) → threshold 180 000 → verdict `passed` on turn 1. Also: a deferring step returns `None` even when the step carries `eval_cap` or a config base exists.
- **Fix 2 verified:** `_parse_tokens`: `"200k"→200_000`, `"1.5m"→1_500_000`, `"500"→500`, `" 200k "→200_000`, int passthrough, `None/-1→-1`; raises on `""`, `"200kb"`, `"1,000"`, bools, floats; `EvalCap(total_token_cap="200k").total_token_cap == 200_000` and `isinstance(..., int)`; threshold `int(200_000*0.9) == 180_000 > 0`; all-`-1` cap → threshold `-1` (no forced compaction), `is_unlimited` True.
- **Merge verified:** base `{max_iterations: 5, input: "100k", total: "200k"}` + step `{max_iterations: 10, eval_cap: {output: "50k"}}` → `max_iterations=10` (step wins), `input=100_000`/`total=200_000` (inherited & parsed), `output=50_000` → end-to-end `passed`.
- **Edge cases:** empty config+step → `None`; `EvalCap(max_iterations=-2)` → `ValueError`; `eval_cap=None` without `config_path` → `ValueError`. Degenerate real caps (`cap=1` → threshold 0) are a config error, not the `-1` default bug.{"model": "deepseek-v4-flash", "problem_class": "python-evaluator-compaction-loop-cap-forwarding", "result": "passed", "tests": 32}