jsonl-glued-two-objects-one-line
Solution written to ~/solution.md. The <project> repo isn't present in this environment (no .coding-hermes/board/tasks.jsonl), so I verified the approach against a byte-level reproduction of the glued line (DF-CRIER-278 + QA-CRIER-30) in /tmp/repro.
Diagnosis: The board line is
{"id":"DF-CRIER-278", ... }{"id":"QA-CRIER-30","status":"pending", ...}
json.loads is all-or-nothing, so Extra data fails the entire line and every line-oriented consumer drops both rows.
Root cause: an appender concatenated a new object onto a line that did not end in \n, violating the one-JSON-value-per-line invariant of JSONL.
Fix: use json.JSONDecoder().raw_decode from pos 0 — it returns each object plus its end offset. Assert ids/status and end2 == len(line), then write:
- line 1 = line[:end1] (raw slice → object 1 stays byte-identical, never re-serialized)
- line 2 = line[end1:].lstrip(" ") (verbatim tail, only leading spaces stripped)
Verification (demonstrated on the reproduction):
- every line parses individually
- jq -e . file clean (confirmed with real jq)
- git diff --numstat = 2 1 (confirmed in a throwaway git repo)
- first added line is a byte-prefix of the deleted line up to }
- both DF-CRIER-278 and QA-CRIER-30 appear in a fresh scan
Note: I corrected the problem text's "exactly +1/-1" — splitting one physical line into two newline-terminated lines is necessarily +2/-1 (one deletion, two insertions), which matches the diff +2/-1 in the provided context. The script refuses to write if it finds zero or multiple glued lines, and aborts before writing on any assertion failure, so no partial repair is possible.
# Evidence - Problem class: jsonl-glued-two-objects-one-line - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-21T00:03:26.894Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Git-tracked JSONL board line contains TWO concatenated JSON objects: a scripted row injection lost the separating newline, so a complete row (DF-CRIER-278) and an appended pending row (QA-CRIER-30) share one physical line. json.loads rejects the line with 'Extra data', so EVERY line-oriented consumer (boardctl list, python per-line parse, jq) silently sees NEITHER row: the pending task is invisible to scans and duplicate-id detection cannot fire. Repair without churning the file: run json.JSONDecoder().raw_decode across the malformed line from pos 0; it returns each object plus its end offset. Assert object 1's id is the expected complete row and object 2's id/status is the expected missing row, and that end2 equals len(rest) (no trailing junk). Then repair BYTE-IDENTICALLY: line1 = original_line[:end1] as a raw SLICE (never re-serialize - key order, escaping and separators of the untouched complete row are preserved exactly), and insert the verbatim tail (only leading spaces stripped) as its own line. Verify: all lines parse individually, jq -e accepts the whole file, git diff is exactly +1/-1 with the first line unchanged up to its closing brace, and both ids appear in a fresh scan. Root cause to prevent: injectors/appenders that build the file with string concatenation instead of appending each row as its own newline-terminated line; any append path must guarantee the previous line ended with a newline before writing the next object.", "environment": "linux, python3 stdlib, git-tracked JSONL foreman board (.coding-hermes/board/tasks.jsonl)", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "jsonl-glued-two-objects-one-line", "provider": "openrouter", "solved_at": "2026-09-21T00:03:26.894Z", "version": ""}Solution written to ~/solution.md. The <project> repo isn't present in this environment (no .coding-hermes/board/tasks.jsonl), so I verified the approach against a byte-level reproduction of the glued line (DF-CRIER-278 + QA-CRIER-30) in /tmp/repro.
Diagnosis: The board line is
{"id":"DF-CRIER-278", ... }{"id":"QA-CRIER-30","status":"pending", ...}
json.loads is all-or-nothing, so Extra data fails the entire line and every line-oriented consumer drops both rows.
Root cause: an appender concatenated a new object onto a line that did not end in \n, violating the one-JSON-value-per-line invariant of JSONL.
Fix: use json.JSONDecoder().raw_decode from pos 0 — it returns each object plus its end offset. Assert ids/status and end2 == len(line), then write:
- line 1 = line[:end1] (raw slice → object 1 stays byte-identical, never re-serialized)
- line 2 = line[end1:].lstrip(" ") (verbatim tail, only leading spaces stripped)
Verification (demonstrated on the reproduction):
- every line parses individually
- jq -e . file clean (confirmed with real jq)
- git diff --numstat = 2 1 (confirmed in a throwaway git repo)
- first added line is a byte-prefix of the deleted line up to }
- both DF-CRIER-278 and QA-CRIER-30 appear in a fresh scan
Note: I corrected the problem text's "exactly +1/-1" — splitting one physical line into two newline-terminated lines is necessarily +2/-1 (one deletion, two insertions), which matches the diff +2/-1 in the provided context. The script refuses to write if it finds zero or multiple glued lines, and aborts before writing on any assertion failure, so no partial repair is possible.
# Evidence - Problem class: jsonl-glued-two-objects-one-line - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-21T00:03:26.894Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Git-tracked JSONL board line contains TWO concatenated JSON objects: a scripted row injection lost the separating newline, so a complete row (DF-CRIER-278) and an appended pending row (QA-CRIER-30) share one physical line. json.loads rejects the line with 'Extra data', so EVERY line-oriented consumer (boardctl list, python per-line parse, jq) silently sees NEITHER row: the pending task is invisible to scans and duplicate-id detection cannot fire. Repair without churning the file: run json.JSONDecoder().raw_decode across the malformed line from pos 0; it returns each object plus its end offset. Assert object 1's id is the expected complete row and object 2's id/status is the expected missing row, and that end2 equals len(rest) (no trailing junk). Then repair BYTE-IDENTICALLY: line1 = original_line[:end1] as a raw SLICE (never re-serialize - key order, escaping and separators of the untouched complete row are preserved exactly), and insert the verbatim tail (only leading spaces stripped) as its own line. Verify: all lines parse individually, jq -e accepts the whole file, git diff is exactly +1/-1 with the first line unchanged up to its closing brace, and both ids appear in a fresh scan. Root cause to prevent: injectors/appenders that build the file with string concatenation instead of appending each row as its own newline-terminated line; any append path must guarantee the previous line ended with a newline before writing the next object.", "environment": "linux, python3 stdlib, git-tracked JSONL foreman board (.coding-hermes/board/tasks.jsonl)", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "jsonl-glued-two-objects-one-line", "provider": "openrouter", "solved_at": "2026-09-21T00:03:26.894Z", "version": ""}