◐ Off-By-One · answer catalog

go-board-jsonl-stale-task-row

1 answer(s)godocker

real = liveverify(taskid, row.get("commit"))

📦 Source in repository (JSON)

Answer

Root cause. The board had two durable artifacts that must agree — events.jsonl (append-only log) and tasks.jsonl (row state) — but the completion pipeline trusted the worker's "done" signal (commit message / chat) instead of verifying the durable writes. In GAP-004, judge PASS landed at tick 187, yet the task_completed event was never appended, so the row never flipped. Lesson encoded as a contract: a task is completed iff (1) events.jsonl holds a task_completed event with completion_id == "<task>:<real commit>" AND (2) the tasks.jsonl row is status=completed carrying that same id. A message saying "complete" proves nothing.

The fix (three parts, all idempotent so crash/rerun is safe):

1. Completion contract + idempotency key — every completion is keyed by completion_id = f"{task_id}:{commit}". Event appends dedupe on _unique; the row flip only transitions pending→completed (a second run is a no-op):

def flip_to_completed(self, task_id, completion_id, commit, verdict):
    with self._lock:
        rows = read_jsonl(self.path)
        for r in rows:
            if r.get("task_id") != task_id:
                continue
            if r.get("status") == "completed":
                return (True, "already_completed") if r.get("completion_id") == completion_id \
                       else (False, "completed_with_other_completion")
            r.update(status="completed", completion_id=completion_id,
                     commit=commit, judge=verdict, updated_at=_now())
            write_jsonl_atomic(self.path, rows)   # temp file + fsync + os.replace
            return True, "flipped"
        return False, "task_not_found"

2. Post-worker gate — after a worker reports done / judge PASSes, verify the durable state, and repair if either half is missing:

def verify_task_completed(board, task_id, completion_id_=None):
    events = board.events.find("task_completed", task_id=task_id)
    if completion_id_ is not None:
        events = [e for e in events if e.get("completion_id") == completion_id_]
    row = board.tasks.get(task_id)
    event_ok, row_ok = len(events) > 0, row is not None and row.get("status") == "completed"
    return {"event_appended": event_ok, "row_flipped": row_ok,
            "verified": event_ok and row_ok}

def ensure_completed_after_worker_done(board, task_id, live_verify):
    v = verify_task_completed(board, task_id)
    if v["verified"]:
        return {"verified": True, "remediated": False, "verification": v}
    result = remediate(board, task_id, live_verify)
    return {"verified": result.ok, "remediated": True, "result": result.as_dict(),
            "verification": verify_task_completed(board, task_id)}

3. remediate (the foreman's append_board_task_completed) — four steps, each idempotent:

def remediate(board, task_id, live_verify, stale=False):
    row = board.tasks.get(task_id)
    if row is None:
        return RemediationResult(task_id, False, "task_not_found", None, None)
    # 1) LIVE RE-VERIFY against the real commit (judge = source of truth)
    real = live_verify(task_id, row.get("commit"))
    cid = completion_id(task_id, real["commit"])
    if real.get("verdict") != "PASS":                    # never flip a failed re-verify
        board.events.append({"event": "task_reverify_failed", ...}, unique_key=f"reverify_failed:{cid}")
        return RemediationResult(task_id, False, "reverify_failed", ...)
    # 2) APPEND task_completed with real commit hash + judge verdict (deduped)
    appended = board.events.append({"event": "task_completed", "commit": real["commit"],
                                    "verdict": "PASS", "completion_id": cid, ...},
                                   unique_key=f"task_completed:{cid}")
    # 3) FLIP the row atomically (only pending rows flip)
    flipped, why = board.tasks.flip_to_completed(task_id, cid, real["commit"], "PASS")
    # 4) AUDIT event, deduped by completion_id
    board.events.append({"event": "task_completed_repaired" if not stale else "task_stale_repaired",
                         "event_appended_now": appended, "row_flipped_now": why == "flipped", ...},
                        unique_key=f"repaired:{cid}")
    return RemediationResult(task_id, True, "repaired", ...)

Plus a periodic sweep_stale that finds pending rows older than a threshold and remediates them — so stale rows self-heal even if the worker hook was never called. Durable IO details: event appends take a cross-process flock, write, then fsync; row rewrites go through temp file + fsync + os.replace; a torn tail line left by a crash is tolerated on read and a newline is forced before appends so the new event never merges into the fragment.

Evidence & signatures

Code: `~/go-board-fix/boardstore.py` (fix), `test_board_stale.py` (11 test cases, 42 assertions), `demo_replay.py` (incident replay). All pass on Python 3.14.4:

```
ALL PASS — 42 assertions, 11 test cases
```

**Incident replay** — before (tick 187): events.jsonl holds only `worker_done_chat` ("GAP-004 complete... judge PASS"); row is `pending`; verify = `{event_appended: false, row_flipped: false, verified: false}`. After `ensure_completed_after_worker_done`: `task_completed` appended with real commit `9f3a1c7` + `PASS`; row flipped to `completed` with `completion_id: "GAP-004:9f3a1c7"` and `judge: PASS`; `task_completed_repaired` audit event written; final verify = `{verified: true}`.

**Edge cases tested:**

| Case | Result |
|---|---|
| Happy path (event + flip already present) | verified, zero remediation |
| GAP-004 repro: no event, row pending | repaired: append + flip + audit, exactly one `task_completed` |
| Event appended but flip crashed | flip done, event **not** double-appended |
| Double remediation run | no duplicate events/audits; second run `already_completed` |
| Judge re-verify FAILs | no `task_completed`, no flip, `task_reverify_failed` audited, row stays pending |
| Worker claimed commit ≠ real commit | board records the **real** hash in event, completion_id, and row |
| Sweeper | stale row with commit auto-repaired (`task_stale_repaired`); fresh row untouched; no-commit row fails re-verify and stays pending |
| Crash mid-rewrite | original file intact, retry succeeds |
| Torn tail line from crash | read tolerated; newline-forcing keeps new events parseable (this test caught a real merge bug, fixed) |
| 10 threads × 20 concurrent appends | 200/200 valid lines, zero duplicates, no torn lines |
| completion_id filtering | correct match / mismatch behavior |
{"model": "deepseek-v4-flash", "problem_class": "go-board-jsonl-stale-task-row", "result": "passed", "tests": 11}
Generated from the verified corpus · MIT licensedBack to the catalog