fleet-parallel-tick-collision-verification
Root cause of the collision: Tick 116's scope silently overlapped Tick 115's mid-tick commit (115 landed at 12:36:49, 3 min after 116 fired). The worker had no gate between trusting sibling claims and writing its own entry, and the umbrella task (UI-01-PARITY, 14 ACs / 9 tickets) let the judge score work the worker never implemented. Ephemeral /tmp mockups could vanish between ticks.
The fix is a 5-step verification protocol implemented as a runnable tool: ~/fleet_verify.py (stdlib-only, hermetic runner injection).
(1) Verify all sibling claims before writing your own entry — three independent signals must agree (judge verdict on disk + build/vet/tsc re-run + board counts re-query). Any contradiction ⇒ REJECT (no entry); board drift past the tick-fire snapshot ⇒ WAIT (resnapshot, re-verify):
def verify_sibling_claim(claim, judge_verdict, build_results, board_state, board_snapshot):
reject, wait = [], []
if judge_verdict is None:
reject.append("judge verdict file missing on disk -> unverified")
elif judge_verdict.get("result") != "passed":
reject.append(f"judge verdict = {judge_verdict.get('result')!r}")
for check, ok in build_results.items():
if not ok:
reject.append(f"re-run {check} failed (rc != 0)")
if board_state.status.get(claim.ticket) != "done":
reject.append(f"board re-query: {claim.ticket} not done")
if board_snapshot is not None and board_state.counts != board_snapshot:
wait.append("board counts drifted since tick fired; resnapshot and re-verify")
if reject: return VerdictDecision("REJECT", reject)
if wait: return VerdictDecision("WAIT", wait)
return VerdictDecision("ACCEPT", ["all independent signals agree"])
One real bug this surfaced: the board store keeps tickets as objects ({"status": "done"}), so the naive status read compared a dict to a string and always failed. read_board() now normalizes both shapes:
status = {t: (e["status"] if isinstance(e, dict) else e) for t, e in tickets.items()}
(2) Never run test suites in parallel with sibling verification — an exclusive flock shared by the verification phase and run_test_suite(); a contender raises TestLockBusy until the holder drains:
@contextlib.contextmanager
def acquire_test_lock(lock_path, timeout=0.0):
with lock_path.open("w") as fh:
deadline = time.monotonic() + timeout
while True:
try:
fcntl.flock(fh, fcntl.LOCK_EX | fcntl.LOCK_NB); break
except OSError:
if time.monotonic() >= deadline: raise TestLockBusy(...)
time.sleep(0.05)
try: yield
finally: fcntl.flock(fh, fcntl.LOCK_UN)
(3) Execute the sibling handoff explicitly instead of re-doing work — verified claims are dispatched, never re-implemented (rework: false), and the board is updated:
record = {"handoff": "UI-01", "from_tick": claim.tick_id, "writer_tick": writer_tick,
"action": "dispatch", "rework": False, "dispatched_at": claim.committed_at}
(4) Rescope the umbrella task to the dispatchable unit — UI-01-PARITY (14 ACs / 9 tickets) → UI-01.task.json with exactly 8 ACs, scope: ["UI-01"], the other 6 ACs in out_of_scope, so the judge scores only implemented work. Idempotent (re-run never duplicates ACs).
(5) Copy ephemeral /tmp assets into the repo with a checksummed manifest.json (stable repo paths, sha256, staged_from) so future ticks never depend on /tmp; missing source is a non-fatal skip.
Runbook (run_verification) orders the steps and writes the worker's own entry only after ACCEPT: lock → verify (judge/build/board) → handoff → rescope → stage assets → write entries/tick-N.entry.json. exit 0 = dispatched, exit 3 = held.
Verified by running the suite and the live incident workspace: - **`~/test_fleet_verify.py`: 11/11 tests pass** (`python3 test_fleet_verify.py` → exit 0) - **End-to-end runbook** against `demo-incident-12-36-49` (sibling 115 committed 12:36:49): `--runner green` → `ACCEPT` → handoff record `dispatch/UI-01.handoff.115.json`, board `UI-01` → `done`, `tasks/UI-01.task.json` with 8 ACs, mockups staged with manifest, entry written only after verification. Exit 0. - **Hold path**: same workspace with the real `shell` runner (no toolchain present) → `REJECT` with `re-run build/vet/tsc failed (rc != 0)`, exit 3, **no** entry/dispatch artifacts — the gate works. Edge cases tested: accepted claim (full protocol), broken tsc re-run → reject, missing judge verdict → reject, failed judge verdict → reject, board contradicts claim → reject, board counts drifted past snapshot → wait (no entry), lock held by sibling → `TestLockBusy` (no parallel suite, then drains), rescope idempotency (8 ACs, 6 out-of-scope, no duplication), missing `/tmp` assets → non-fatal skip, sha256 manifest integrity, handoff board update.
{"model": "deepseek-v4-flash", "problem_class": "fleet-parallel-tick-collision-verification", "result": "passed", "tests": 11}Root cause of the collision: Tick 116's scope silently overlapped Tick 115's mid-tick commit (115 landed at 12:36:49, 3 min after 116 fired). The worker had no gate between trusting sibling claims and writing its own entry, and the umbrella task (UI-01-PARITY, 14 ACs / 9 tickets) let the judge score work the worker never implemented. Ephemeral /tmp mockups could vanish between ticks.
The fix is a 5-step verification protocol implemented as a runnable tool: ~/fleet_verify.py (stdlib-only, hermetic runner injection).
(1) Verify all sibling claims before writing your own entry — three independent signals must agree (judge verdict on disk + build/vet/tsc re-run + board counts re-query). Any contradiction ⇒ REJECT (no entry); board drift past the tick-fire snapshot ⇒ WAIT (resnapshot, re-verify):
def verify_sibling_claim(claim, judge_verdict, build_results, board_state, board_snapshot):
reject, wait = [], []
if judge_verdict is None:
reject.append("judge verdict file missing on disk -> unverified")
elif judge_verdict.get("result") != "passed":
reject.append(f"judge verdict = {judge_verdict.get('result')!r}")
for check, ok in build_results.items():
if not ok:
reject.append(f"re-run {check} failed (rc != 0)")
if board_state.status.get(claim.ticket) != "done":
reject.append(f"board re-query: {claim.ticket} not done")
if board_snapshot is not None and board_state.counts != board_snapshot:
wait.append("board counts drifted since tick fired; resnapshot and re-verify")
if reject: return VerdictDecision("REJECT", reject)
if wait: return VerdictDecision("WAIT", wait)
return VerdictDecision("ACCEPT", ["all independent signals agree"])
One real bug this surfaced: the board store keeps tickets as objects ({"status": "done"}), so the naive status read compared a dict to a string and always failed. read_board() now normalizes both shapes:
status = {t: (e["status"] if isinstance(e, dict) else e) for t, e in tickets.items()}
(2) Never run test suites in parallel with sibling verification — an exclusive flock shared by the verification phase and run_test_suite(); a contender raises TestLockBusy until the holder drains:
@contextlib.contextmanager
def acquire_test_lock(lock_path, timeout=0.0):
with lock_path.open("w") as fh:
deadline = time.monotonic() + timeout
while True:
try:
fcntl.flock(fh, fcntl.LOCK_EX | fcntl.LOCK_NB); break
except OSError:
if time.monotonic() >= deadline: raise TestLockBusy(...)
time.sleep(0.05)
try: yield
finally: fcntl.flock(fh, fcntl.LOCK_UN)
(3) Execute the sibling handoff explicitly instead of re-doing work — verified claims are dispatched, never re-implemented (rework: false), and the board is updated:
record = {"handoff": "UI-01", "from_tick": claim.tick_id, "writer_tick": writer_tick,
"action": "dispatch", "rework": False, "dispatched_at": claim.committed_at}
(4) Rescope the umbrella task to the dispatchable unit — UI-01-PARITY (14 ACs / 9 tickets) → UI-01.task.json with exactly 8 ACs, scope: ["UI-01"], the other 6 ACs in out_of_scope, so the judge scores only implemented work. Idempotent (re-run never duplicates ACs).
(5) Copy ephemeral /tmp assets into the repo with a checksummed manifest.json (stable repo paths, sha256, staged_from) so future ticks never depend on /tmp; missing source is a non-fatal skip.
Runbook (run_verification) orders the steps and writes the worker's own entry only after ACCEPT: lock → verify (judge/build/board) → handoff → rescope → stage assets → write entries/tick-N.entry.json. exit 0 = dispatched, exit 3 = held.
Verified by running the suite and the live incident workspace: - **`~/test_fleet_verify.py`: 11/11 tests pass** (`python3 test_fleet_verify.py` → exit 0) - **End-to-end runbook** against `demo-incident-12-36-49` (sibling 115 committed 12:36:49): `--runner green` → `ACCEPT` → handoff record `dispatch/UI-01.handoff.115.json`, board `UI-01` → `done`, `tasks/UI-01.task.json` with 8 ACs, mockups staged with manifest, entry written only after verification. Exit 0. - **Hold path**: same workspace with the real `shell` runner (no toolchain present) → `REJECT` with `re-run build/vet/tsc failed (rc != 0)`, exit 3, **no** entry/dispatch artifacts — the gate works. Edge cases tested: accepted claim (full protocol), broken tsc re-run → reject, missing judge verdict → reject, failed judge verdict → reject, board contradicts claim → reject, board counts drifted past snapshot → wait (no entry), lock held by sibling → `TestLockBusy` (no parallel suite, then drains), rescope idempotency (8 ACs, 6 out-of-scope, no duplication), missing `/tmp` assets → non-fatal skip, sha256 manifest integrity, handoff board update.
{"model": "deepseek-v4-flash", "problem_class": "fleet-parallel-tick-collision-verification", "result": "passed", "tests": 11}