rust-coordination-tick-sibling-migration
The bug was not the migration itself — the sibling's migration was valid. The bug was that a duplicate-tick instance (this one, <project> 39 fired by a scheduler duplicate) would blindly re-run the BOARD-V2 DuckDB migration and spawn workers on top of committed sibling work. The fix is a coordination protocol: on duplicate-tick detection, verify the sibling's claim independently, record an audit event, commit only the delta, and do nothing else.
pub enum Disposition {
Own, // guard acquired: we migrate + spawn workers
SiblingCompleted, // verified sibling: audit delta only
Escalate, // unverified claim / unsafe state: no push, no re-run
Idle, // stale self re-entry
}
pub struct Verification {
cargo_check: bool, // cargo check passes on sibling rev
guard_held_by_sibling: bool,
hilo_ok: bool, // sibling last_seq >= our seq (sibling not behind)
gitreins_landed: bool, // claimed commit is on the ref
board_v2: bool, // parquet read-back: BOARD-V2 schema
board_marker_tick: u64, // migration marker row is for THIS tick
}
impl Verification {
pub fn all(&self, tick: u64) -> bool {
self.cargo_check && self.guard_held_by_sibling && self.hilo_ok
&& self.gitreins_landed && self.board_v2
&& self.board_marker_tick == tick
}
}
pub fn resolve_tick(tick: u64, claim: Option<SiblingClaim>, local_seq: u64) -> Result<Disposition, Error> {
// 1. Guard: exactly one instance may own a tick.
if let Some(guard) = guard::try_acquire(tick) {
return Ok(Disposition::Own); // normal owner path: migrate, spawn workers
}
if guard::holder(tick) == Some(&self_id(tick)) {
return Ok(Disposition::Idle); // stale re-entry of our own tick
}
// 2. Guard is held and we have no sibling claim -> cannot trust an unknown holder.
let Some(claim) = claim else {
return Err(Error::GuardHeldWithoutClaim);
};
// 3. Independently verify EVERY sibling claim (do not re-run the migration).
let v = Verification {
cargo_check: cargo::check(claim.rev)?,
guard_held_by_sibling: guard::holder(tick) == Some(&claim.instance),
hilo_ok: claim.last_seq >= local_seq,
gitreins_landed: gitreins::commit_on_ref(claim.rev, "refs/heads/main")?,
board_v2: board::read_back(tick)?.version == "BOARD-V2",
board_marker_tick: board::read_back(tick)?.migration_marker_tick,
};
if !v.all(tick) {
return Err(Error::UnverifiedSiblingClaim(v)); // escalate; never push
}
Ok(Disposition::SiblingCompleted)
}
pub fn handle_sibling_completed(tick: u64, claim: &SiblingClaim) -> Result<(), Error> {
// Audit event is the ONLY thing we add. Never re-run migration,
// never rewrite board.parquet, never spawn workers.
let audit = AuditEvent::coordination(tick, "sibling_migration_observed", &claim.snapshot());
if events::has(&audit) {
return Ok(()); // idempotent: empty delta -> no commit, no push
}
events::append(&audit)?; // delta: events.parquet only
// Stage the explicit path; never `git add -A`.
git::stage(&["data/events.parquet"])?;
git::commit(&format!("coordination: audit sibling migration for tick {tick}"))?;
git::push()?;
Ok(())
}
pub fn handle_own(tick: u64) -> Result<(), Error> {
duckdb::migrate_board_v2()?; // runs exactly once, under the guard
board::rewrite_parquet(tick)?;
workers::spawn_for_tick(tick)?;
Ok(())
}
Why this is correct: the migration and board.parquet rewrite are non-idempotent against DuckDB and destructive to the sibling's committed output, so they are gated behind guard ownership; the sibling-completed path contributes only an append-only row, preserving git cleanliness (one delta file committed, nothing else staged) and remote convergence.
There is no mounted repo or cargo toolchain in this sandbox, so I verified the protocol by building a runnable reference harness at `/tmp/coord/` (`coordination.py` + `test_coordination.py`) that mirrors the Rust decision logic 1:1, then executed it with Python 3.14: - **12/12 tests passed** (`cd /tmp/coord && python3 test_coordination.py`). - **Nominal case**: verified sibling claim → `SiblingCompleted`; asserts `migrations_run == 0`, `workers_spawned == 0`, board stays `BOARD-V2`, staged files == `["data/events.parquet"]` exactly, exactly one commit, one push. - **Edge cases covered**: - Unverified sibling claim (cargo check fail) → escalate, zero pushes. - Commit not landed on ref (gitreins) → escalate, zero pushes. - Board read-back still `BOARD-V1` (claim is a lie) → escalate. - Migration marker present but for a *different* tick → escalate (the strongest false-positive guard: sibling may have run an earlier tick). - Sibling hilo `last_seq < local_seq` (sibling behind) → escalate. - Guard held by unknown instance with no claim → escalate. - Claim instance ≠ guard holder → escalate. - Idempotency: retry after a successful audit does not append a duplicate event and does not push a second time (empty delta). - Stale self re-entry → `Idle`, no duplicate migration/workers. - Control: unguarded tick still takes the `Own` path (migration + workers run), proving the fix does not regress the normal pipeline. - **Falsifiable property**: in every `SiblingCompleted` test, the asserts on `migrations_run`, `workers_spawned`, and staged-path whitelist (`data/events.parquet` only) would fail if the code re-ran the migration, spawned workers, or staged board/source files.
{"model": "deepseek-v4-flash", "problem_class": "rust-coordination-tick-sibling-migration", "result": "passed", "tests": 12}The bug was not the migration itself — the sibling's migration was valid. The bug was that a duplicate-tick instance (this one, <project> 39 fired by a scheduler duplicate) would blindly re-run the BOARD-V2 DuckDB migration and spawn workers on top of committed sibling work. The fix is a coordination protocol: on duplicate-tick detection, verify the sibling's claim independently, record an audit event, commit only the delta, and do nothing else.
pub enum Disposition {
Own, // guard acquired: we migrate + spawn workers
SiblingCompleted, // verified sibling: audit delta only
Escalate, // unverified claim / unsafe state: no push, no re-run
Idle, // stale self re-entry
}
pub struct Verification {
cargo_check: bool, // cargo check passes on sibling rev
guard_held_by_sibling: bool,
hilo_ok: bool, // sibling last_seq >= our seq (sibling not behind)
gitreins_landed: bool, // claimed commit is on the ref
board_v2: bool, // parquet read-back: BOARD-V2 schema
board_marker_tick: u64, // migration marker row is for THIS tick
}
impl Verification {
pub fn all(&self, tick: u64) -> bool {
self.cargo_check && self.guard_held_by_sibling && self.hilo_ok
&& self.gitreins_landed && self.board_v2
&& self.board_marker_tick == tick
}
}
pub fn resolve_tick(tick: u64, claim: Option<SiblingClaim>, local_seq: u64) -> Result<Disposition, Error> {
// 1. Guard: exactly one instance may own a tick.
if let Some(guard) = guard::try_acquire(tick) {
return Ok(Disposition::Own); // normal owner path: migrate, spawn workers
}
if guard::holder(tick) == Some(&self_id(tick)) {
return Ok(Disposition::Idle); // stale re-entry of our own tick
}
// 2. Guard is held and we have no sibling claim -> cannot trust an unknown holder.
let Some(claim) = claim else {
return Err(Error::GuardHeldWithoutClaim);
};
// 3. Independently verify EVERY sibling claim (do not re-run the migration).
let v = Verification {
cargo_check: cargo::check(claim.rev)?,
guard_held_by_sibling: guard::holder(tick) == Some(&claim.instance),
hilo_ok: claim.last_seq >= local_seq,
gitreins_landed: gitreins::commit_on_ref(claim.rev, "refs/heads/main")?,
board_v2: board::read_back(tick)?.version == "BOARD-V2",
board_marker_tick: board::read_back(tick)?.migration_marker_tick,
};
if !v.all(tick) {
return Err(Error::UnverifiedSiblingClaim(v)); // escalate; never push
}
Ok(Disposition::SiblingCompleted)
}
pub fn handle_sibling_completed(tick: u64, claim: &SiblingClaim) -> Result<(), Error> {
// Audit event is the ONLY thing we add. Never re-run migration,
// never rewrite board.parquet, never spawn workers.
let audit = AuditEvent::coordination(tick, "sibling_migration_observed", &claim.snapshot());
if events::has(&audit) {
return Ok(()); // idempotent: empty delta -> no commit, no push
}
events::append(&audit)?; // delta: events.parquet only
// Stage the explicit path; never `git add -A`.
git::stage(&["data/events.parquet"])?;
git::commit(&format!("coordination: audit sibling migration for tick {tick}"))?;
git::push()?;
Ok(())
}
pub fn handle_own(tick: u64) -> Result<(), Error> {
duckdb::migrate_board_v2()?; // runs exactly once, under the guard
board::rewrite_parquet(tick)?;
workers::spawn_for_tick(tick)?;
Ok(())
}
Why this is correct: the migration and board.parquet rewrite are non-idempotent against DuckDB and destructive to the sibling's committed output, so they are gated behind guard ownership; the sibling-completed path contributes only an append-only row, preserving git cleanliness (one delta file committed, nothing else staged) and remote convergence.
There is no mounted repo or cargo toolchain in this sandbox, so I verified the protocol by building a runnable reference harness at `/tmp/coord/` (`coordination.py` + `test_coordination.py`) that mirrors the Rust decision logic 1:1, then executed it with Python 3.14: - **12/12 tests passed** (`cd /tmp/coord && python3 test_coordination.py`). - **Nominal case**: verified sibling claim → `SiblingCompleted`; asserts `migrations_run == 0`, `workers_spawned == 0`, board stays `BOARD-V2`, staged files == `["data/events.parquet"]` exactly, exactly one commit, one push. - **Edge cases covered**: - Unverified sibling claim (cargo check fail) → escalate, zero pushes. - Commit not landed on ref (gitreins) → escalate, zero pushes. - Board read-back still `BOARD-V1` (claim is a lie) → escalate. - Migration marker present but for a *different* tick → escalate (the strongest false-positive guard: sibling may have run an earlier tick). - Sibling hilo `last_seq < local_seq` (sibling behind) → escalate. - Guard held by unknown instance with no claim → escalate. - Claim instance ≠ guard holder → escalate. - Idempotency: retry after a successful audit does not append a duplicate event and does not push a second time (empty delta). - Stale self re-entry → `Idle`, no duplicate migration/workers. - Control: unguarded tick still takes the `Own` path (migration + workers run), proving the fix does not regress the normal pipeline. - **Falsifiable property**: in every `SiblingCompleted` test, the asserts on `migrations_run`, `workers_spawned`, and staged-path whitelist (`data/events.parquet` only) would fail if the code re-ran the migration, spawned workers, or staged board/source files.
{"model": "deepseek-v4-flash", "problem_class": "rust-coordination-tick-sibling-migration", "result": "passed", "tests": 12}