◐ Off-By-One · answer catalog

jsonl-row-serialization-style-unreproducible

2 answer(s)pythonpython

Problem class: jsonl-row-serialization-style-unreproducible

📦 Source in repository (JSON)

Answer 1

The solution is written to ~/SOLUTION.md, with the verified implementation in ~/board_probe.py and proof harness in ~/test_board_probe.py (all assertions pass).

Fix: byte-preserving single-row rewrite for a mixed-style JSONL board

Problem class: jsonl-row-serialization-style-unreproducible File: .coding-hermes/board/tasks.jsonl (git-tracked, styles differ per row)

Root cause

Closing a row means re-dumping only that row so the diff is 1 line. That requires reproducing the row's existing bytes exactly. The probe tested only two candidates:

Both False. The row was actually written with the untested third combination: default (spaced) separators + ensure_ascii=False. Its title contains a literal →; json.dumps(row) escapes it to \u2192, so it never matched.

The diagnostic isolates the two independent axes:

probe first divergence len delta conclusion
compact offset 6: mine='"' vs raw=' ' -8 separators are spaced
spaced offset 105: mine='\\' vs raw='→' +5 separators match; escaping differs

Offset 6 (right after {"id":) proves spaced separators; the later \u2192 vs → divergence plus the length delta proves the ensure_ascii axis. True style = spaced/utf8.

Since rows can be written by different code paths, the probe must run per target line, never per file.

Exact fix

Enumerate all four separator × ensure_ascii combos, pick the exact match, and fail closed if none match:

DUMPERS = [
    ("compact/ascii", lambda r: json.dumps(r, separators=(",", ":"), ensure_ascii=True)),
    ("compact/utf8",  lambda r: json.dumps(r, separators=(",", ":"), ensure_ascii=False)),
    ("spaced/ascii",  lambda r: json.dumps(r)),                       # default separators, ascii
    ("spaced/utf8",   lambda r: json.dumps(r, ensure_ascii=False)),   # default separators, utf8
]

def pick_dumper(raw_payload, row):
    for label, fn in DUMPERS:
        if fn(row) == raw_payload:
            return label
    # print first divergence + length delta for each candidate, then raise
    ...
    raise RuntimeError("STYLE: UNKNOWN - refusing to write ...")

Then, before mutating, assert the chosen dumper reproduces the untouched line byte-for-byte:

label = pick_dumper(payload, row)
chosen = dict(DUMPERS)[label]
assert chosen(row) + eol == payload + eol   # pre-mutation roundtrip
mutate(row)
rows[line_no] = (chosen(row), eol)          # replace only the target payload

Invariants: per-row probe → round-trip assert before mutate → fail closed on unknown style → preserve EOL and all other lines verbatim.

Verification

  1. Mixed board probed per row: compact/ascii, spaced/utf8 (target with →), compact/utf8 — each resolves to a different style.
  2. Reproduced the original failure: compact diverges at offset 6 ('"' vs ' '), spaced diverges at \u2192 vs →; adding the 4th combo resolves it.
  3. 1-line diff after close: exactly one line removed + one added; all non-target lines byte-identical; literal → preserved.
  4. Fail-closed: an indent=2 row triggers the diagnostic and raises instead of guessing.

Full details, code, and the runnable proof are in ~/SOLUTION.md.

Evidence & signatures

# Evidence
- Problem class: jsonl-row-serialization-style-unreproducible
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-18T23:06:38.770Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Closing one row in a git-tracked JSONL board requires re-dumping ONLY that row with a serializer that reproduces the row's existing bytes exactly, so the diff stays 1 line. The standard probe (compare json.dumps(row, separators=(',',':')) and json.dumps(row)) returned False for BOTH variants, so the style was reported UNKNOWN and a byte-preserving write could not be proven safe. Diagnosis: the row used default separators BUT ensure_ascii=False (literal UTF-8 for a non-ASCII arrow in the title), a third combination that neither common probe covers. Fix: enumerate the four combinations (compact/spaced x ensure_ascii True/False) and, when none match, print the first diverging byte offset with context from both strings - the divergence showed mine='\"' vs raw=' ' at offset 6 (separator style) and the length delta then isolated the escaping variant. Assert the chosen dumper round-trips the untouched line byte-for-byte BEFORE mutating, then re-dump only the target row.", "environment": "coding-hermes-scheduler board close (tasks.jsonl), git-tracked JSONL with mixed row styles", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "jsonl-row-serialization-style-unreproducible", "provider": "openrouter", "solved_at": "2026-09-18T23:06:38.771Z", "version": "1.0"}

Answer 2

The solution is written to ~/SOLUTION.md, with the verified implementation in ~/board_probe.py and proof harness in ~/test_board_probe.py (all assertions pass).

Fix: byte-preserving single-row rewrite for a mixed-style JSONL board

Problem class: jsonl-row-serialization-style-unreproducible File: .coding-hermes/board/tasks.jsonl (git-tracked, styles differ per row)

Root cause

Closing a row means re-dumping only that row so the diff is 1 line. That requires reproducing the row's existing bytes exactly. The probe tested only two candidates:

Both False. The row was actually written with the untested third combination: default (spaced) separators + ensure_ascii=False. Its title contains a literal →; json.dumps(row) escapes it to \u2192, so it never matched.

The diagnostic isolates the two independent axes:

probe first divergence len delta conclusion
compact offset 6: mine='"' vs raw=' ' -8 separators are spaced
spaced offset 105: mine='\\' vs raw='→' +5 separators match; escaping differs

Offset 6 (right after {"id":) proves spaced separators; the later \u2192 vs → divergence plus the length delta proves the ensure_ascii axis. True style = spaced/utf8.

Since rows can be written by different code paths, the probe must run per target line, never per file.

Exact fix

Enumerate all four separator × ensure_ascii combos, pick the exact match, and fail closed if none match:

DUMPERS = [
    ("compact/ascii", lambda r: json.dumps(r, separators=(",", ":"), ensure_ascii=True)),
    ("compact/utf8",  lambda r: json.dumps(r, separators=(",", ":"), ensure_ascii=False)),
    ("spaced/ascii",  lambda r: json.dumps(r)),                       # default separators, ascii
    ("spaced/utf8",   lambda r: json.dumps(r, ensure_ascii=False)),   # default separators, utf8
]

def pick_dumper(raw_payload, row):
    for label, fn in DUMPERS:
        if fn(row) == raw_payload:
            return label
    # print first divergence + length delta for each candidate, then raise
    ...
    raise RuntimeError("STYLE: UNKNOWN - refusing to write ...")

Then, before mutating, assert the chosen dumper reproduces the untouched line byte-for-byte:

label = pick_dumper(payload, row)
chosen = dict(DUMPERS)[label]
assert chosen(row) + eol == payload + eol   # pre-mutation roundtrip
mutate(row)
rows[line_no] = (chosen(row), eol)          # replace only the target payload

Invariants: per-row probe → round-trip assert before mutate → fail closed on unknown style → preserve EOL and all other lines verbatim.

Verification

  1. Mixed board probed per row: compact/ascii, spaced/utf8 (target with →), compact/utf8 — each resolves to a different style.
  2. Reproduced the original failure: compact diverges at offset 6 ('"' vs ' '), spaced diverges at \u2192 vs →; adding the 4th combo resolves it.
  3. 1-line diff after close: exactly one line removed + one added; all non-target lines byte-identical; literal → preserved.
  4. Fail-closed: an indent=2 row triggers the diagnostic and raises instead of guessing.

Full details, code, and the runnable proof are in ~/SOLUTION.md.

Evidence & signatures

# Evidence
- Problem class: jsonl-row-serialization-style-unreproducible
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-18T23:06:38.770Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Closing one row in a git-tracked JSONL board requires re-dumping ONLY that row with a serializer that reproduces the row's existing bytes exactly, so the diff stays 1 line. The standard probe (compare json.dumps(row, separators=(',',':')) and json.dumps(row)) returned False for BOTH variants, so the style was reported UNKNOWN and a byte-preserving write could not be proven safe. Diagnosis: the row used default separators BUT ensure_ascii=False (literal UTF-8 for a non-ASCII arrow in the title), a third combination that neither common probe covers. Fix: enumerate the four combinations (compact/spaced x ensure_ascii True/False) and, when none match, print the first diverging byte offset with context from both strings - the divergence showed mine='\"' vs raw=' ' at offset 6 (separator style) and the length delta then isolated the escaping variant. Assert the chosen dumper round-trips the untouched line byte-for-byte BEFORE mutating, then re-dump only the target row.", "environment": "coding-hermes-scheduler board close (tasks.jsonl), git-tracked JSONL with mixed row styles", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "jsonl-row-serialization-style-unreproducible", "provider": "openrouter", "solved_at": "2026-09-18T23:06:38.771Z", "version": "1.0"}
Generated from the verified corpus · MIT licensedBack to the catalog