Problem class: model-sync-score-concatenated-json-parse
All 16 tests pass and the end-to-end demo confirms both acceptance branches. Here is the verified solution.
model-sync --score aborts on concatenated JSON and exits 0 with no score fileProblem class: model-sync-score-concatenated-json-parse
Files: scripts/model_sync.py (_extract_json_object, _llm_score_candidates, scoring exception handler), tests/test_model_sync_scoring.py
Shipped in: merge into main tick 279 (work 2bbb130)
Two defects chained into a silent-success hole. A 225-candidate "lane-namespace flood" made the scorer reply with two concatenated JSON objects ({...}{...}) and finish_reason=stop.
Defect 1 — brace-slice fallback re-fed the whole concatenation to json.loads.
The old extractor did a text.find("{") … text.rfind("}") slice and called json.loads on it. For {"a":1}{"b":2} the slice is the entire string, so json.loads raised JSONDecodeError: Extra data: line 2 column 1 (char 24). Because JSONDecodeError subclasses ValueError, it propagated out as an ordinary parse error, indistinguishable from a non-truncation failure.
Defect 2 — the token-doubling ladder was gated on finish_reason == length.
_score_llm_reply() raised immediately unless finish_reason == length, so an 8192→16384→32768 retry never fired. The concatenated reply was thrown away instead of being parsed from its first object, and truncated replies never got the larger budget.
Why the wrapper still exited 0 with no score file: the scoring exception was caught somewhere up the stack and swallowed (or the process returned normally after printing), so the cron wrapper observed success while model_scores_*.yaml was never written.
The error observed:
LLM scoring failed: Extra data: line 2 column 1 (char 24)
(model=deepseek-v4-flash, finish_reason=stop, max_tokens=8192);
no model_scores_*.yaml written
{ scan + raw_decode from each positionscripts/model_sync.py — replace _extract_json_object:
def _iter_top_level_object_starts(text: str):
"""Yield indices of top-level '{' characters, ignoring string contents.
Only depth-0 braces are yielded. This keeps a truncated outer object
(``{"a": {"b": 1}``) from being "rescued" by parsing its nested inner
object; such a reply must fall through to the length-doubling ladder.
"""
depth = 0
in_string = False
escaped = False
for i, ch in enumerate(text):
if in_string:
if escaped:
escaped = False
elif ch == "\\":
escaped = True
elif ch == '"':
in_string = False
continue
if ch == '"':
in_string = True
elif ch == "{":
if depth == 0:
yield i
depth += 1
elif ch == "}":
if depth > 0:
depth -= 1
def _extract_json_object(text: str):
"""Return the first top-level JSON object in ``text`` or ``None``.
Handles clean JSON, ```json fenced blocks, prose-wrapped objects, and
concatenated replies (``{...}{...}`` -> first object). Returns ``None`` for
truncated outer objects so callers can trigger the token ladder.
"""
if not text:
return None
decoder = json.JSONDecoder()
for start in _iter_top_level_object_starts(text):
try:
obj, _end = decoder.raw_decode(text, start)
except ValueError:
continue # not valid from here; try the next top-level '{'
if isinstance(obj, dict):
return obj
return None
Key properties:
- raw_decode(text, start) parses exactly one value and ignores trailing data → concatenated objects resolve from the first object.
- Fenced/prose-wrapped replies work because we search for the first parseable top-level {.
- Braces inside strings are skipped.
- Only depth-0 starts are tried, so a truncated outer object returns None rather than silently scoring a nested partial object — it reaches the ladder instead.
Replace the body of _llm_score_candidates (and the raise in _score_llm_reply):
INITIAL_SCORING_TOKENS = 8192
MAX_SCORING_TOKENS = 32768
class ScoringError(RuntimeError):
def __init__(self, reason, **details):
self.reason = reason
self.details = details
detail_str = ", ".join(f"{k}={v!r}" for k, v in details.items())
super().__init__(f"{reason} ({detail_str})" if detail_str else reason)
def _llm_score_candidates(candidates, scorer, limit=5):
selected = _top_candidates_by_recency(candidates, limit)
max_tokens = INITIAL_SCORING_TOKENS
while True:
result = scorer(_build_prompt(selected), max_tokens=max_tokens)
obj = _extract_json_object(getattr(result, "text", "") or "")
if obj is not None:
return obj
if getattr(result, "finish_reason", None) == "length" and max_tokens < MAX_SCORING_TOKENS:
max_tokens = min(max_tokens * 2, MAX_SCORING_TOKENS)
continue
raise ScoringError(
"unparseable_llm_reply",
finish_reason=getattr(result, "finish_reason", None),
max_tokens=max_tokens,
preview=(getattr(result, "text", "") or "")[:120],
)
The scoring exception handler must print a named reason to stderr and sys.exit(1):
try:
scores = _llm_score_candidates(candidates, scorer)
except ScoringError as exc:
print(f"LLM scoring failed: {exc.reason} ({exc.details})", file=sys.stderr)
sys.exit(1)
# only write after a successful parse
write_scores(scores)
MAX_SCORING_CANDIDATES = 5
def _top_candidates_by_recency(candidates, limit=MAX_SCORING_CANDIDATES):
return sorted(candidates, key=lambda c: c.get("mtime", 0), reverse=True)[:limit]
Reproduced the scoring path in scripts/model_sync.py and added tests/test_model_sync_scoring.py. Run:
python3 -m pytest -q tests/test_model_sync_scoring.py
Result:
................ [100%]
16 passed in 0.04s
Covered cases:
| Test | Assertion |
|---|---|
test_concatenated_objects_parse_from_first |
{"a":1}{"b":2} → first object |
test_old_impl_raises_extra_data_on_concatenation |
old code raises Extra data (root cause #1) |
test_fenced_reply_parses / test_prose_wrapped_reply_parses |
fenced + prose parse |
test_braces_inside_strings_do_not_break_scan |
braces in string literals ignored |
test_truncated_outer_object_returns_none |
top-level-only scan; no nested rescue |
test_ladder_doubles_on_length_then_succeeds |
token calls [8192, 16384] |
test_ladder_exhausts_and_raises_named_reason |
[8192, 16384, 32768] → ScoringError |
test_non_length_unparseable_raises_named_reason |
stop + garbage → named error, not silent |
test_top_five_by_recency_cap |
225 candidates → newest 5 |
test_run_scoring_exits_nonzero_and_writes_nothing_on_failure |
exit 1, no output written, named reason on stderr |
test_run_scoring_writes_scores_on_success |
success writes scores |
End-to-end acceptance run:
python3 /tmp/acceptance_demo.py
LLM scoring failed: unparseable_llm_reply
({'finish_reason': 'stop', 'max_tokens': 8192, 'preview': 'I cannot score these.'})
A: exit ok, score file contents = [{'scores': {'x': 88}}]
B: exit code = 1 | score file written = []
1, writes nothing, prints unparseable_llm_reply to stderr.Acceptance held: a concatenated reply either scores from the first object, or the process exits non-zero with a named reason — it can never exit ok with no score file.
# Evidence - Problem class: model-sync-score-concatenated-json-parse - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-25T21:02:07.400Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "model-sync --score aborts on a non-truncated unparsable LLM reply (two concatenated JSON objects: 'Extra data: line 2 column 1') and the wrapper still exits ok with NO score file. Two chained defects: _extract_json_object()'s brace-slice fallback fed the whole concatenation back to json.loads (JSONDecodeError subclasses ValueError, so it propagated as a non-truncation error), and _score_llm_reply() raised immediately unless finish_reason==length, so the 8192->32768 doubling ladder never fired. Trigger: 225 candidates embedded in the scoring prompt (lane-namespace flood). FIX (shipped this tick): string-aware top-level '{' scan + json.JSONDecoder().raw_decode() from each position \u2014 concatenated/fenced/prose-wrapped replies parse from the FIRST object (top-level-only so truncated outer objects still reach the length ladder); the scoring exception handler now prints a named reason to stderr and sys.exit(1) \u2014 silent-success hole closed; top-5-by-recency candidate cap pinned by test. ACCEPTANCE held: concatenated reply scores from first object OR exits non-zero with a named reason \u2014 never exit ok with no score file.", "environment": "<project> scripts/model_sync.py --score-from (cron), deepseek-v4-flash scorer", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "model-sync-score-concatenated-json-parse", "provider": "openrouter", "solved_at": "2026-09-25T21:02:07.400Z", "version": ""}All 16 tests pass and the end-to-end demo confirms both acceptance branches. Here is the verified solution.
model-sync --score aborts on concatenated JSON and exits 0 with no score fileProblem class: model-sync-score-concatenated-json-parse
Files: scripts/model_sync.py (_extract_json_object, _llm_score_candidates, scoring exception handler), tests/test_model_sync_scoring.py
Shipped in: merge into main tick 279 (work 2bbb130)
Two defects chained into a silent-success hole. A 225-candidate "lane-namespace flood" made the scorer reply with two concatenated JSON objects ({...}{...}) and finish_reason=stop.
Defect 1 — brace-slice fallback re-fed the whole concatenation to json.loads.
The old extractor did a text.find("{") … text.rfind("}") slice and called json.loads on it. For {"a":1}{"b":2} the slice is the entire string, so json.loads raised JSONDecodeError: Extra data: line 2 column 1 (char 24). Because JSONDecodeError subclasses ValueError, it propagated out as an ordinary parse error, indistinguishable from a non-truncation failure.
Defect 2 — the token-doubling ladder was gated on finish_reason == length.
_score_llm_reply() raised immediately unless finish_reason == length, so an 8192→16384→32768 retry never fired. The concatenated reply was thrown away instead of being parsed from its first object, and truncated replies never got the larger budget.
Why the wrapper still exited 0 with no score file: the scoring exception was caught somewhere up the stack and swallowed (or the process returned normally after printing), so the cron wrapper observed success while model_scores_*.yaml was never written.
The error observed:
LLM scoring failed: Extra data: line 2 column 1 (char 24)
(model=deepseek-v4-flash, finish_reason=stop, max_tokens=8192);
no model_scores_*.yaml written
{ scan + raw_decode from each positionscripts/model_sync.py — replace _extract_json_object:
def _iter_top_level_object_starts(text: str):
"""Yield indices of top-level '{' characters, ignoring string contents.
Only depth-0 braces are yielded. This keeps a truncated outer object
(``{"a": {"b": 1}``) from being "rescued" by parsing its nested inner
object; such a reply must fall through to the length-doubling ladder.
"""
depth = 0
in_string = False
escaped = False
for i, ch in enumerate(text):
if in_string:
if escaped:
escaped = False
elif ch == "\\":
escaped = True
elif ch == '"':
in_string = False
continue
if ch == '"':
in_string = True
elif ch == "{":
if depth == 0:
yield i
depth += 1
elif ch == "}":
if depth > 0:
depth -= 1
def _extract_json_object(text: str):
"""Return the first top-level JSON object in ``text`` or ``None``.
Handles clean JSON, ```json fenced blocks, prose-wrapped objects, and
concatenated replies (``{...}{...}`` -> first object). Returns ``None`` for
truncated outer objects so callers can trigger the token ladder.
"""
if not text:
return None
decoder = json.JSONDecoder()
for start in _iter_top_level_object_starts(text):
try:
obj, _end = decoder.raw_decode(text, start)
except ValueError:
continue # not valid from here; try the next top-level '{'
if isinstance(obj, dict):
return obj
return None
Key properties:
- raw_decode(text, start) parses exactly one value and ignores trailing data → concatenated objects resolve from the first object.
- Fenced/prose-wrapped replies work because we search for the first parseable top-level {.
- Braces inside strings are skipped.
- Only depth-0 starts are tried, so a truncated outer object returns None rather than silently scoring a nested partial object — it reaches the ladder instead.
Replace the body of _llm_score_candidates (and the raise in _score_llm_reply):
INITIAL_SCORING_TOKENS = 8192
MAX_SCORING_TOKENS = 32768
class ScoringError(RuntimeError):
def __init__(self, reason, **details):
self.reason = reason
self.details = details
detail_str = ", ".join(f"{k}={v!r}" for k, v in details.items())
super().__init__(f"{reason} ({detail_str})" if detail_str else reason)
def _llm_score_candidates(candidates, scorer, limit=5):
selected = _top_candidates_by_recency(candidates, limit)
max_tokens = INITIAL_SCORING_TOKENS
while True:
result = scorer(_build_prompt(selected), max_tokens=max_tokens)
obj = _extract_json_object(getattr(result, "text", "") or "")
if obj is not None:
return obj
if getattr(result, "finish_reason", None) == "length" and max_tokens < MAX_SCORING_TOKENS:
max_tokens = min(max_tokens * 2, MAX_SCORING_TOKENS)
continue
raise ScoringError(
"unparseable_llm_reply",
finish_reason=getattr(result, "finish_reason", None),
max_tokens=max_tokens,
preview=(getattr(result, "text", "") or "")[:120],
)
The scoring exception handler must print a named reason to stderr and sys.exit(1):
try:
scores = _llm_score_candidates(candidates, scorer)
except ScoringError as exc:
print(f"LLM scoring failed: {exc.reason} ({exc.details})", file=sys.stderr)
sys.exit(1)
# only write after a successful parse
write_scores(scores)
MAX_SCORING_CANDIDATES = 5
def _top_candidates_by_recency(candidates, limit=MAX_SCORING_CANDIDATES):
return sorted(candidates, key=lambda c: c.get("mtime", 0), reverse=True)[:limit]
Reproduced the scoring path in scripts/model_sync.py and added tests/test_model_sync_scoring.py. Run:
python3 -m pytest -q tests/test_model_sync_scoring.py
Result:
................ [100%]
16 passed in 0.04s
Covered cases:
| Test | Assertion |
|---|---|
test_concatenated_objects_parse_from_first |
{"a":1}{"b":2} → first object |
test_old_impl_raises_extra_data_on_concatenation |
old code raises Extra data (root cause #1) |
test_fenced_reply_parses / test_prose_wrapped_reply_parses |
fenced + prose parse |
test_braces_inside_strings_do_not_break_scan |
braces in string literals ignored |
test_truncated_outer_object_returns_none |
top-level-only scan; no nested rescue |
test_ladder_doubles_on_length_then_succeeds |
token calls [8192, 16384] |
test_ladder_exhausts_and_raises_named_reason |
[8192, 16384, 32768] → ScoringError |
test_non_length_unparseable_raises_named_reason |
stop + garbage → named error, not silent |
test_top_five_by_recency_cap |
225 candidates → newest 5 |
test_run_scoring_exits_nonzero_and_writes_nothing_on_failure |
exit 1, no output written, named reason on stderr |
test_run_scoring_writes_scores_on_success |
success writes scores |
End-to-end acceptance run:
python3 /tmp/acceptance_demo.py
LLM scoring failed: unparseable_llm_reply
({'finish_reason': 'stop', 'max_tokens': 8192, 'preview': 'I cannot score these.'})
A: exit ok, score file contents = [{'scores': {'x': 88}}]
B: exit code = 1 | score file written = []
1, writes nothing, prints unparseable_llm_reply to stderr.Acceptance held: a concatenated reply either scores from the first object, or the process exits non-zero with a named reason — it can never exit ok with no score file.
# Evidence - Problem class: model-sync-score-concatenated-json-parse - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-25T21:02:07.400Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "model-sync --score aborts on a non-truncated unparsable LLM reply (two concatenated JSON objects: 'Extra data: line 2 column 1') and the wrapper still exits ok with NO score file. Two chained defects: _extract_json_object()'s brace-slice fallback fed the whole concatenation back to json.loads (JSONDecodeError subclasses ValueError, so it propagated as a non-truncation error), and _score_llm_reply() raised immediately unless finish_reason==length, so the 8192->32768 doubling ladder never fired. Trigger: 225 candidates embedded in the scoring prompt (lane-namespace flood). FIX (shipped this tick): string-aware top-level '{' scan + json.JSONDecoder().raw_decode() from each position \u2014 concatenated/fenced/prose-wrapped replies parse from the FIRST object (top-level-only so truncated outer objects still reach the length ladder); the scoring exception handler now prints a named reason to stderr and sys.exit(1) \u2014 silent-success hole closed; top-5-by-recency candidate cap pinned by test. ACCEPTANCE held: concatenated reply scores from first object OR exits non-zero with a named reason \u2014 never exit ok with no score file.", "environment": "<project> scripts/model_sync.py --score-from (cron), deepseek-v4-flash scorer", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "model-sync-score-concatenated-json-parse", "provider": "openrouter", "solved_at": "2026-09-25T21:02:07.400Z", "version": ""}