board-hilo-metric-staleness
Root cause: the foreman's per-tick hook copied the prior BoardEntry forward — header and metrics — without ever invoking hilo graph stats. So the header froze at 8,939/1,612 for 105+ ticks while the live graph grew to 9,646/1,625.
Fix contract (enforced in one shared module, ~/fix/board_header.py):
1. Fresh stats every tick, unconditionally — run_hilo_stats() is called at the top of run_tick(); the previous entry is never a source of metric values.
2. Divergence → provenance note — prev is read only to compare against fresh stats; when they differ, the header appends [provenance: ... corrected stale X -> live Y].
3. Failure → explicit UNAVAILABLE — a failed/unparseable hilo graph stats writes an explicit marker, never a silent reuse of yesterday's numbers.
def run_hilo_stats(runner=None) -> Optional[HiloStats]:
"""Run `hilo graph stats` FRESH. No cache, no fallback to prior tick."""
if runner is not None:
return parse_stats(runner())
try:
proc = subprocess.run(["hilo", "graph", "stats"],
capture_output=True, text=True,
timeout=30.0, check=True)
except (OSError, subprocess.SubprocessError):
return None # surfaces as UNAVAILABLE, never stale
return parse_stats(proc.stdout)
def build_entry(tick, stats, prev=None) -> BoardEntry:
if stats is None: # failure: explicit, no carry-forward
return BoardEntry(tick, f"[board] Hilo: UNAVAILABLE (tick {tick}) -- "
"`hilo graph stats` failed; metrics NOT carried forward", None)
note = None
if prev is not None and prev.stats is not None and prev.stats != stats:
note = (f"provenance: re-ran `hilo graph stats` live (tick {tick}); "
f"corrected stale {prev.stats.render()} -> live {stats.render()}")
header = f"[board] Hilo {stats.render()} (tick {tick})"
return BoardEntry(tick, header + (f" [{note}]" if note else ""), stats, note)
def run_tick(tick, board_path, prev=None, runner=None):
fresh = run_hilo_stats(runner) # 1) FRESH stats EVERY tick
entry = build_entry(tick, fresh, prev) # 2) never reuse prev.stats
_persist(board_path, entry) # 3) write header
return entry
parse_stats() pairs numbers with edge/file keywords via one-pass token claiming, so "edges: 9646 files: 1625", "9646 edges / 1625 files", and "edges = 9,646 files = 1,625" all parse correctly. The only persisted state is the previously posted header (tick number + old stats for divergence detection).
**Exact incident reproduced & fixed** (`demo.py`): under the old loop, after 105 ticks the header still read `[board] Hilo 8,939 edges/1,612 files (tick 1)` while live was 9,646/1,625. With the fix deployed, the first tick produced: ``` tick 106: [board] Hilo 9,646 edges/1,625 files (tick 106) [provenance: re-ran `hilo graph stats` live (tick 106); corrected stale 8,939 edges/1,612 files -> live 9,646 edges/1,625 files] tick 107: [board] Hilo 9,646 edges/1,625 files (tick 107) <- no note while live is stable ``` **12/12 tests pass** (`pytest test_board_header.py`, ~/fix/), covering edge cases: - **The 105-tick regression**: stale 8,939/1,612 vs live 9,646/1,625 → corrected header with provenance note naming *both* old and new values; no note when live stays stable; note reappears at the next real divergence. - **Freshness guarantee**: exactly one `hilo graph stats` invocation per tick, including when fresh values *coincidentally equal* `prev` (the old bug reused; the fix re-runs). - **Failure paths**: nonzero exit, empty/unparseable output, missing numbers, negative numbers → `UNAVAILABLE` header, stale numbers never carried forward. - **Parser**: comma-formatted (`9,646`), `N edges / M files`, `edges: N files: M`, `EDGES N FILES M`, mixed-case; garbage rejected as `None`. - **Subprocess integration**: fake `hilo` binary on `PATH` (exit 0 and exit 1) → correct parse and fresh re-run per call. - **Tick continuity**: tick numbers monotonic even as metrics refresh. One real bug was caught during verification: a naive number-adjacent-to-keyword regex misattributed 9646 to the `files` field in `"edges: 9646 files: 1625"`; fixed with one-pass keyword→number claiming (each number consumed by exactly one keyword).
{"model": "deepseek-v4-flash", "problem_class": "board-hilo-metric-staleness", "result": "passed", "tests": 12}Root cause: the foreman's per-tick hook copied the prior BoardEntry forward — header and metrics — without ever invoking hilo graph stats. So the header froze at 8,939/1,612 for 105+ ticks while the live graph grew to 9,646/1,625.
Fix contract (enforced in one shared module, ~/fix/board_header.py):
1. Fresh stats every tick, unconditionally — run_hilo_stats() is called at the top of run_tick(); the previous entry is never a source of metric values.
2. Divergence → provenance note — prev is read only to compare against fresh stats; when they differ, the header appends [provenance: ... corrected stale X -> live Y].
3. Failure → explicit UNAVAILABLE — a failed/unparseable hilo graph stats writes an explicit marker, never a silent reuse of yesterday's numbers.
def run_hilo_stats(runner=None) -> Optional[HiloStats]:
"""Run `hilo graph stats` FRESH. No cache, no fallback to prior tick."""
if runner is not None:
return parse_stats(runner())
try:
proc = subprocess.run(["hilo", "graph", "stats"],
capture_output=True, text=True,
timeout=30.0, check=True)
except (OSError, subprocess.SubprocessError):
return None # surfaces as UNAVAILABLE, never stale
return parse_stats(proc.stdout)
def build_entry(tick, stats, prev=None) -> BoardEntry:
if stats is None: # failure: explicit, no carry-forward
return BoardEntry(tick, f"[board] Hilo: UNAVAILABLE (tick {tick}) -- "
"`hilo graph stats` failed; metrics NOT carried forward", None)
note = None
if prev is not None and prev.stats is not None and prev.stats != stats:
note = (f"provenance: re-ran `hilo graph stats` live (tick {tick}); "
f"corrected stale {prev.stats.render()} -> live {stats.render()}")
header = f"[board] Hilo {stats.render()} (tick {tick})"
return BoardEntry(tick, header + (f" [{note}]" if note else ""), stats, note)
def run_tick(tick, board_path, prev=None, runner=None):
fresh = run_hilo_stats(runner) # 1) FRESH stats EVERY tick
entry = build_entry(tick, fresh, prev) # 2) never reuse prev.stats
_persist(board_path, entry) # 3) write header
return entry
parse_stats() pairs numbers with edge/file keywords via one-pass token claiming, so "edges: 9646 files: 1625", "9646 edges / 1625 files", and "edges = 9,646 files = 1,625" all parse correctly. The only persisted state is the previously posted header (tick number + old stats for divergence detection).
**Exact incident reproduced & fixed** (`demo.py`): under the old loop, after 105 ticks the header still read `[board] Hilo 8,939 edges/1,612 files (tick 1)` while live was 9,646/1,625. With the fix deployed, the first tick produced: ``` tick 106: [board] Hilo 9,646 edges/1,625 files (tick 106) [provenance: re-ran `hilo graph stats` live (tick 106); corrected stale 8,939 edges/1,612 files -> live 9,646 edges/1,625 files] tick 107: [board] Hilo 9,646 edges/1,625 files (tick 107) <- no note while live is stable ``` **12/12 tests pass** (`pytest test_board_header.py`, ~/fix/), covering edge cases: - **The 105-tick regression**: stale 8,939/1,612 vs live 9,646/1,625 → corrected header with provenance note naming *both* old and new values; no note when live stays stable; note reappears at the next real divergence. - **Freshness guarantee**: exactly one `hilo graph stats` invocation per tick, including when fresh values *coincidentally equal* `prev` (the old bug reused; the fix re-runs). - **Failure paths**: nonzero exit, empty/unparseable output, missing numbers, negative numbers → `UNAVAILABLE` header, stale numbers never carried forward. - **Parser**: comma-formatted (`9,646`), `N edges / M files`, `edges: N files: M`, `EDGES N FILES M`, mixed-case; garbage rejected as `None`. - **Subprocess integration**: fake `hilo` binary on `PATH` (exit 0 and exit 1) → correct parse and fresh re-run per call. - **Tick continuity**: tick numbers monotonic even as metrics refresh. One real bug was caught during verification: a naive number-adjacent-to-keyword regex misattributed 9646 to the `files` field in `"edges: 9646 files: 1625"`; fixed with one-pass keyword→number claiming (each number consumed by exactly one keyword).
{"model": "deepseek-v4-flash", "problem_class": "board-hilo-metric-staleness", "result": "passed", "tests": 12}