Problem class: jsonl-parity-probe-legacy-idless-rows
Diagnosis complete and verified. I created the fix and a test harness, then confirmed the acceptance criteria.
Files written:
- <project>-foreman-ops/scripts/parity_check.py — robust probe
- <project>-foreman-ops/tests/verify_parity_fix.py — 767-event/11-id-less verification harness
- <project>-foreman-ops/SOLUTION.md — the markdown deliverable
Verification result:
parity run: jsonl rows=767 ids=756 max_id=756 id_less_rows_skipped=11 mirror_ids=756 mirror_max_id=756 parity=ok
missing-mirror run: ... mirror=skipped reason=mirror_not_found:...
ALL CHECKS PASSED
And a real mismatch still exits 1 with parity=MISMATCH.
Problem class: jsonl-parity-probe-legacy-idless-rows
File: <project>-foreman-ops/scripts/parity_check.py
Verified on: Python 3.11+ (reproduced on 3.14 with duckdb 1.5.5)
The probe made two unsafe assumptions.
id. j_ids = [r["id"] for r in rows] (line 36) raises KeyError: 'id' on legacy/external rows that predate the canonical schema, killing the whole scan before any aggregate is computed.duckdb.connect(DB, read_only=True) (line 43) raises duckdb.IOError: ... database does not exist when board.db is absent. A missing mirror is an expected, reportable condition, not fatal.The fix: keep every row, aggregate numeric ids only, count skipped rows as id_less_rows_skipped=<n>, and report (never raise) a missing mirror with an opt-in --strict-mirror.
#!/usr/bin/env python3
"""JSONL <-> DuckDB parity probe (robust to legacy / external id-less rows)."""
from __future__ import annotations
import argparse
import json
import os
import sys
from typing import Any
try:
import duckdb
except ImportError:
duckdb = None
def is_numeric_id(value: Any) -> bool:
"""True for values usable as a numeric event id (bool is not an id)."""
if isinstance(value, bool):
return False
if isinstance(value, (int, float)):
return True
if isinstance(value, str):
try:
float(value)
except ValueError:
return False
return True
return False
def load_rows(jsonl_path: str) -> list[dict[str, Any]]:
"""Load every JSONL row, preserving id-less legacy/external rows."""
rows: list[dict[str, Any]] = []
with open(jsonl_path, "r", encoding="utf-8") as fh:
for lineno, line in enumerate(fh, 1):
line = line.strip()
if not line:
continue
try:
obj = json.loads(line)
except json.JSONDecodeError:
print(f"[warn] skipping malformed JSON at {jsonl_path}:{lineno}",
file=sys.stderr)
continue
rows.append(obj if isinstance(obj, dict) else {"_non_object": obj})
return rows
def numeric_ids(rows: list[dict[str, Any]]) -> tuple[list[int], int]:
"""(numeric ids, count of rows skipped for lacking a numeric id)."""
ids: list[int] = []
skipped = 0
for row in rows:
value = row.get("id")
if value is not None and is_numeric_id(value):
ids.append(int(float(value)))
else:
skipped += 1
return ids, skipped
def read_mirror(db_path: str) -> tuple[list[int] | None, str | None]:
"""Read ids from the mirror. Returns (ids, None) or (None, reason). Never raises."""
if not os.path.exists(db_path):
return None, f"mirror_not_found:{db_path}"
if duckdb is None:
return None, "duckdb_not_installed"
try:
con = duckdb.connect(db_path, read_only=True)
except Exception as exc:
return None, f"mirror_open_error:{exc}"
try:
try:
raw = con.execute("SELECT id FROM events").fetchall()
except Exception:
raw = con.execute("SELECT id FROM board_events").fetchall()
except Exception as exc:
return None, f"mirror_query_error:{exc}"
finally:
con.close()
ids: list[int] = []
for (value,) in raw:
if value is not None and is_numeric_id(value):
ids.append(int(float(value)))
return ids, None
def ids_summary(ids: list[int]) -> tuple[int, int | None]:
return len(ids), (max(ids) if ids else None)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--jsonl", default=os.environ.get("PARITY_JSONL", "events.jsonl"))
parser.add_argument("--db", default=os.environ.get("PARITY_DB", "board.db"))
parser.add_argument("--strict-mirror", action="store_true",
help="Exit non-zero when the DB mirror is missing/unreadable.")
args = parser.parse_args(argv)
if not os.path.exists(args.jsonl):
print(f"[error] JSONL board not found: {args.jsonl}", file=sys.stderr)
return 2
rows = load_rows(args.jsonl)
j_ids, id_less_skipped = numeric_ids(rows)
j_count, j_max = ids_summary(j_ids)
prefix = (f"jsonl rows={len(rows)} ids={j_count} max_id={j_max} "
f"id_less_rows_skipped={id_less_skipped}")
mirror_ids, reason = read_mirror(args.db)
if mirror_ids is None:
print(f"{prefix} mirror=skipped reason={reason}")
return 3 if args.strict_mirror else 0
d_count, d_max = ids_summary(mirror_ids)
parity = (j_count == d_count) and (j_max == d_max)
print(f"{prefix} mirror_ids={d_count} mirror_max_id={d_max} "
f"parity={'ok' if parity else 'MISMATCH'}")
return 0 if parity else 1
if __name__ == "__main__":
raise SystemExit(main())
| Symptom | Old | New |
|---|---|---|
KeyError: 'id' |
[r["id"] for r in rows] |
row.get("id") + is_numeric_id(); skipped rows counted |
| Rows lost | no-id rows broke aggregation | all rows retained; only ids filtered |
DuckDB IOError |
direct connect() |
os.path.exists() pre-check; errors captured as reason |
| No visibility | crash | id_less_rows_skipped=<n> always printed |
| CI signal | always crash | exit 0 + mirror=skipped, or 3 with --strict-mirror |
cd <project>-foreman-ops
python3 tests/verify_parity_fix.py
parity run: jsonl rows=767 ids=756 max_id=756 id_less_rows_skipped=11 mirror_ids=756 mirror_max_id=756 parity=ok
missing-mirror run: jsonl rows=767 ids=756 max_id=756 id_less_rows_skipped=11 mirror=skipped reason=mirror_not_found:/tmp/.../does-not-exist.db
ALL CHECKS PASSED
Checklist:
- [x] 767-event corpus with 11 id-less legacy rows scanned with no error
- [x] rows=767 confirms every row retained
- [x] max_id/ids computed over numeric ids only
- [x] parity print carries id_less_rows_skipped=11
- [x] missing mirror reported (mirror=skipped reason=mirror_not_found:...), exit 0 / strict 3, not aborted
- [x] real mismatch still yields parity=MISMATCH, exit 1
# Evidence - Problem class: jsonl-parity-probe-legacy-idless-rows - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-11T06:50:19.167Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A JSONL-vs-DuckDB parity probe died with KeyError mid-scan because legacy and external event rows predate the canonical schema and carry no id field. Separately, the probe hard-fails when the DB mirror file does not exist. The robust probe must compute max and count over numeric ids only, keep every row in the file, and either skip or clearly report the missing-mirror case instead of aborting.", "environment": "Hermes coding-hermes JSONL-canonical board", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "jsonl-parity-probe-legacy-idless-rows", "provider": "openrouter", "solved_at": "2026-09-11T06:50:19.167Z", "version": "3.11"}Diagnosis complete and verified. I created the fix and a test harness, then confirmed the acceptance criteria.
Files written:
- <project>-foreman-ops/scripts/parity_check.py — robust probe
- <project>-foreman-ops/tests/verify_parity_fix.py — 767-event/11-id-less verification harness
- <project>-foreman-ops/SOLUTION.md — the markdown deliverable
Verification result:
parity run: jsonl rows=767 ids=756 max_id=756 id_less_rows_skipped=11 mirror_ids=756 mirror_max_id=756 parity=ok
missing-mirror run: ... mirror=skipped reason=mirror_not_found:...
ALL CHECKS PASSED
And a real mismatch still exits 1 with parity=MISMATCH.
Problem class: jsonl-parity-probe-legacy-idless-rows
File: <project>-foreman-ops/scripts/parity_check.py
Verified on: Python 3.11+ (reproduced on 3.14 with duckdb 1.5.5)
The probe made two unsafe assumptions.
id. j_ids = [r["id"] for r in rows] (line 36) raises KeyError: 'id' on legacy/external rows that predate the canonical schema, killing the whole scan before any aggregate is computed.duckdb.connect(DB, read_only=True) (line 43) raises duckdb.IOError: ... database does not exist when board.db is absent. A missing mirror is an expected, reportable condition, not fatal.The fix: keep every row, aggregate numeric ids only, count skipped rows as id_less_rows_skipped=<n>, and report (never raise) a missing mirror with an opt-in --strict-mirror.
#!/usr/bin/env python3
"""JSONL <-> DuckDB parity probe (robust to legacy / external id-less rows)."""
from __future__ import annotations
import argparse
import json
import os
import sys
from typing import Any
try:
import duckdb
except ImportError:
duckdb = None
def is_numeric_id(value: Any) -> bool:
"""True for values usable as a numeric event id (bool is not an id)."""
if isinstance(value, bool):
return False
if isinstance(value, (int, float)):
return True
if isinstance(value, str):
try:
float(value)
except ValueError:
return False
return True
return False
def load_rows(jsonl_path: str) -> list[dict[str, Any]]:
"""Load every JSONL row, preserving id-less legacy/external rows."""
rows: list[dict[str, Any]] = []
with open(jsonl_path, "r", encoding="utf-8") as fh:
for lineno, line in enumerate(fh, 1):
line = line.strip()
if not line:
continue
try:
obj = json.loads(line)
except json.JSONDecodeError:
print(f"[warn] skipping malformed JSON at {jsonl_path}:{lineno}",
file=sys.stderr)
continue
rows.append(obj if isinstance(obj, dict) else {"_non_object": obj})
return rows
def numeric_ids(rows: list[dict[str, Any]]) -> tuple[list[int], int]:
"""(numeric ids, count of rows skipped for lacking a numeric id)."""
ids: list[int] = []
skipped = 0
for row in rows:
value = row.get("id")
if value is not None and is_numeric_id(value):
ids.append(int(float(value)))
else:
skipped += 1
return ids, skipped
def read_mirror(db_path: str) -> tuple[list[int] | None, str | None]:
"""Read ids from the mirror. Returns (ids, None) or (None, reason). Never raises."""
if not os.path.exists(db_path):
return None, f"mirror_not_found:{db_path}"
if duckdb is None:
return None, "duckdb_not_installed"
try:
con = duckdb.connect(db_path, read_only=True)
except Exception as exc:
return None, f"mirror_open_error:{exc}"
try:
try:
raw = con.execute("SELECT id FROM events").fetchall()
except Exception:
raw = con.execute("SELECT id FROM board_events").fetchall()
except Exception as exc:
return None, f"mirror_query_error:{exc}"
finally:
con.close()
ids: list[int] = []
for (value,) in raw:
if value is not None and is_numeric_id(value):
ids.append(int(float(value)))
return ids, None
def ids_summary(ids: list[int]) -> tuple[int, int | None]:
return len(ids), (max(ids) if ids else None)
def main(argv: list[str] | None = None) -> int:
parser = argparse.ArgumentParser(description=__doc__)
parser.add_argument("--jsonl", default=os.environ.get("PARITY_JSONL", "events.jsonl"))
parser.add_argument("--db", default=os.environ.get("PARITY_DB", "board.db"))
parser.add_argument("--strict-mirror", action="store_true",
help="Exit non-zero when the DB mirror is missing/unreadable.")
args = parser.parse_args(argv)
if not os.path.exists(args.jsonl):
print(f"[error] JSONL board not found: {args.jsonl}", file=sys.stderr)
return 2
rows = load_rows(args.jsonl)
j_ids, id_less_skipped = numeric_ids(rows)
j_count, j_max = ids_summary(j_ids)
prefix = (f"jsonl rows={len(rows)} ids={j_count} max_id={j_max} "
f"id_less_rows_skipped={id_less_skipped}")
mirror_ids, reason = read_mirror(args.db)
if mirror_ids is None:
print(f"{prefix} mirror=skipped reason={reason}")
return 3 if args.strict_mirror else 0
d_count, d_max = ids_summary(mirror_ids)
parity = (j_count == d_count) and (j_max == d_max)
print(f"{prefix} mirror_ids={d_count} mirror_max_id={d_max} "
f"parity={'ok' if parity else 'MISMATCH'}")
return 0 if parity else 1
if __name__ == "__main__":
raise SystemExit(main())
| Symptom | Old | New |
|---|---|---|
KeyError: 'id' |
[r["id"] for r in rows] |
row.get("id") + is_numeric_id(); skipped rows counted |
| Rows lost | no-id rows broke aggregation | all rows retained; only ids filtered |
DuckDB IOError |
direct connect() |
os.path.exists() pre-check; errors captured as reason |
| No visibility | crash | id_less_rows_skipped=<n> always printed |
| CI signal | always crash | exit 0 + mirror=skipped, or 3 with --strict-mirror |
cd <project>-foreman-ops
python3 tests/verify_parity_fix.py
parity run: jsonl rows=767 ids=756 max_id=756 id_less_rows_skipped=11 mirror_ids=756 mirror_max_id=756 parity=ok
missing-mirror run: jsonl rows=767 ids=756 max_id=756 id_less_rows_skipped=11 mirror=skipped reason=mirror_not_found:/tmp/.../does-not-exist.db
ALL CHECKS PASSED
Checklist:
- [x] 767-event corpus with 11 id-less legacy rows scanned with no error
- [x] rows=767 confirms every row retained
- [x] max_id/ids computed over numeric ids only
- [x] parity print carries id_less_rows_skipped=11
- [x] missing mirror reported (mirror=skipped reason=mirror_not_found:...), exit 0 / strict 3, not aborted
- [x] real mismatch still yields parity=MISMATCH, exit 1
# Evidence - Problem class: jsonl-parity-probe-legacy-idless-rows - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-11T06:50:19.167Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A JSONL-vs-DuckDB parity probe died with KeyError mid-scan because legacy and external event rows predate the canonical schema and carry no id field. Separately, the probe hard-fails when the DB mirror file does not exist. The robust probe must compute max and count over numeric ids only, keep every row in the file, and either skip or clearly report the missing-mirror case instead of aborting.", "environment": "Hermes coding-hermes JSONL-canonical board", "language": "python", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "jsonl-parity-probe-legacy-idless-rows", "provider": "openrouter", "solved_at": "2026-09-11T06:50:19.167Z", "version": "3.11"}