cron-double-fire-detection
The fix is a detector + foreman policy that turns "likely cron double-fire" into an explicit, auditable classification while deliberately never letting the redundancy assumption skip work. Files: ~/solution/double_fire.py, ~/solution/test_double_fire.py.
1. Canonical fingerprint of the tick's work input (not run metadata):
def fingerprint(payload: dict[str, Any]) -> str:
blob = json.dumps(_canon(payload), sort_keys=True, default=str)
return hashlib.sha256(blob.encode("utf-8")).hexdigest()
_canon sorts dict keys recursively and rounds floats to 6 places, so 05:26 vs 05:26.000001, or different key orders, don't break "same state" detection. tick_id/collected_at live on TickInput, outside the payload — a re-fire is the same job with the same input, not the same run name.
2. Detection — double-fire requires BOTH rapid interval AND identical state:
class DoubleFireDetector:
def classify(self, current, prior):
fp = fingerprint(current.payload)
if prior is None:
return "normal", fp, [] # first tick can't re-fire
delta = current.collected_at - prior.collected_at
if delta < 0: # clock skew / replay
return "normal", fp, ["out-of-order tick ..."]
same_state = fp == prior.fingerprint
if delta <= self.config.double_fire_window_s and same_state:
return "probable_double_fire", fp, [f"rapid re-fire {delta:.0f}s after {prior.tick_id} ..."]
return "normal", fp, [] # changed input = real event
Window is 360s (T68 landed 180s after T67). Either condition failing means it's a genuine event, not redundancy.
3. Foreman handling — audit always runs, gates never skipped:
class Foreman:
def process(self, tick):
classification, fp, notes = self.detector.classify(tick, prior)
audit_result = self.audit.run(tick.payload) # FULL audit, always executed
gate_results = run_gates(tick.payload, self.gates, self.live) # identical gate set
...
if classification == "probable_double_fire":
rec.board_notes.append(
"board entry: probable cron double-fire flagged; "
"full audit executed (cache-backed); gates NOT skipped")
The Audit is cache-backed by fingerprint (T68 hits T67's cache — that's the "cheap" part) but is invoked every tick; a miss is computed fresh. Gates receive (work_input, live) — live probes like sources_fresh can fail on the re-fire tick (a source went stale between 05:26 and 05:29) without changing the fingerprint, so a real failure is never masked by the double-fire label.
Verified by running the full suite (`pytest`, Python 3.14.4): **11/11 passed**. Direct replay of the incident:
```
T67: normal | gates: {'sources_fresh': True, 'invariants_hold': True, 'board_consistent': True}
T68: probable_double_fire
fingerprint identical: True
audit_ran: True | cache: 1 compute / 1 hit
note: rapid re-fire 180s after T67 with identical fingerprint 7de1dd8fea21 — probable cron double-fire
note: board entry: probable cron double-fire flagged; full audit executed (cache-backed); gates NOT skipped
```
Edge cases tested:
- **Exact incident** — T68 180s after T67, identical input → flagged; audit ran (cache hit, 1 compute/1 hit); all 3 gates evaluated, identical set.
- **Rapid fire + state change** (rows 42→43) → `normal` + "real event" note; redundancy assumption never applied.
- **Slow no-op rerun** (10 min later, same state) → `normal`, no double-fire note.
- **First-ever tick** → `normal` (nothing to re-fire).
- **Out-of-order clock / replay** (negative interval) → `normal` with explicit note.
- **Window boundary** — exactly 360s → flagged; 361s → not.
- **Fingerprint robustness** — dict key order and float noise don't change the hash.
- **Failing live gate on a double-fire tick** — classified double-fire *and* gate `sources_fresh: False` surfaces (`ok` is False); nothing masked.
- **Audit failure** — board note present on both the original and the double-fire tick.
- **3-tick cycle** — normal → double-fire → normal.{"model": "deepseek-v4-flash", "problem_class": "cron-double-fire-detection", "result": "passed", "tests": 11}The bug: the tick loop had no idempotency boundary, so the 04:35 re-fire of tick 176 re-ran gates (against a contaminated parallel workspace), spawned a duplicate glm-5.2 worker, mutated state twice, and short-circuited the loop. The fix layers atomic claims in front of every destructive action, makes gate verdicts isolation-based, and turns duplicate fires into a verify/account/finish path instead of a re-execution path.
1. Claim layers (tick-level + work-key-level), so a duplicate can never spawn a worker:
def run(self, raw_tick, commit, scope, signal="lint"):
work_key = f"{commit}::{scope}::{signal}"
expected = self.ledger.max_tick() + 1
# Layer 0: raw tick already accounted -> double-fire or stale replay.
if raw_tick <= self.ledger.max_tick():
return self._duplicate(raw_tick, work_key, expected, cause="tick", prior=None)
# Layer 1: O_CREAT|O_EXCL sentinel — atomic across sibling foremen.
# The 04:35 fire of tick 176 loses here against the 04:32 claim.
if not self.ledger.claim_tick(raw_tick):
return self._duplicate(raw_tick, work_key, expected, cause="tick", prior=None)
# Layer 2: work-key claim under fcntl lock — an *advanced* counter
# (177) for the same logical work is still deduplicated.
won, claim = self.ledger.claim_work(work_key)
if not won:
return self._duplicate(raw_tick, work_key, expected, cause="work_key", prior=claim)
return self._process(raw_tick, work_key, commit, scope, signal)
claim_tick uses os.open(sentinel, O_CREAT|O_EXCL) — exactly one foreman wins per tick; claim_work is read-append under fcntl.flock, so sibling foremen racing the same work can't double-claim.
2. Gates verified independently — a parallel-run FAIL is a contention artifact, not a verdict:
def resolve(self, scope):
"""Parallel FAILs are re-run in isolation; a FAIL that passes in
isolation is recorded as a contention artifact and excluded."""
parallel = self.run_parallel(scope)
final, artifacts = [], []
for r in parallel:
if r.passed:
final.append(r); continue
iso = self.run_isolated(r.name, scope) # private scratch dir
if iso.passed:
artifacts.append(f"{r.name}: parallel FAIL -> isolated PASS (contention artifact)")
final.append(iso)
return final, artifacts
Result: {build: True, vet: True, tests: True, lint: False} — controller tests PASS in isolation (42.3s); the parallel FAIL is recorded as an artifact. Lint is genuinely RED on e45b287 (10 new issues in SPEC-GAP-002 files) → LINT-REG-001 created exactly once.
3. Duplicate fire → verify, account, finish — never re-execute:
def _duplicate(self, raw_tick, work_key, expected, cause, prior):
before = self.store.snapshot() # hashes of registrations+workers
# NO gates, NO registration, NO worker spawn.
self.obo.submit({ # submit to off-by-one
"event": "off_by_one", "cause": cause, "work_key": work_key,
"expected_tick": expected, "observed_tick": raw_tick,
"delta": raw_tick - expected, "corrected": True, ...})
steps = self._loop_steps(work_key, ...) # DuckBrain write + signal scan
after = self.store.snapshot()
unchanged = before == after # verify no state change
self.ledger.append({"event": "duplicate", "tick": self.ledger.max_tick(),
"cause": cause, "state_unchanged": unchanged})
return TickResult(handled="duplicate", worker=None, state_unchanged=unchanged, ...)
The corrected logical tick comes from the ledger (max_tick), never from the corruptible cron counter — that's the off-by-one correction (expected_tick=177, observed_tick=176, delta=-1). Worker spawn is additionally idempotent (WorkerPool.spawn returns the existing worker for a claimed key), a second backstop.
Verified by running `python3 -m unittest test_double_fire -v` → **5/5 pass** (0.008s), plus the timeline demo:
```
[04:32] fire #1: verdict={'build':True,'vet':True,'tests':True,'lint':False}
contention_artifacts=['tests: parallel FAIL -> isolated PASS (contention artifact)']
registration=LINT-REG-001 worker=glm-5.2@tick-176
[04:35] fire #2: handled=duplicate cause=tick worker=None
state_unchanged=True off_by_one={expected_tick:177, observed_tick:176, delta:-1, corrected:True}
loop_steps=['DuckBrain write','signal scan']
workers=1 registrations=1 processed_events=1
```
Edge cases tested:
- **Same-tick double-fire (main scenario):** second fire → `duplicate/cause=tick`; no worker, no `LINT-REG-002`, state fingerprints unchanged, off-by-one record `{expected:177, observed:176, delta:-1, corrected:true}` submitted, DuckBrain write + signal scan completed. Exactly 1 worker / 1 registration / 1 `processed` ledger event; a third fire stays idempotent.
- **Contention artifact:** `tests` fails in parallel pass, passes isolated (42.3s) → artifact, not a verdict; `vet`/`build` unaffected; lint RED only for SPEC-GAP-002 scope.
- **Counter advanced (177) for the same work:** deduped via `cause=work_key`, logical tick corrected to 176, no second worker.
- **Sibling foreman race:** two `Foreman` instances on one state root firing tick 176 → exactly one `processed`, one `duplicate`, one worker (O_EXCL sentinel).
- **Distinct work not over-deduped:** `e45b287::SPEC-GAP-001::vet` proceeds normally → 2 workers, still 1 registration.{"model": "deepseek-v4-flash", "problem_class": "cron-double-fire-detection", "result": "passed", "tests": 5}The fix is a detector + foreman policy that turns "likely cron double-fire" into an explicit, auditable classification while deliberately never letting the redundancy assumption skip work. Files: ~/solution/double_fire.py, ~/solution/test_double_fire.py.
1. Canonical fingerprint of the tick's work input (not run metadata):
def fingerprint(payload: dict[str, Any]) -> str:
blob = json.dumps(_canon(payload), sort_keys=True, default=str)
return hashlib.sha256(blob.encode("utf-8")).hexdigest()
_canon sorts dict keys recursively and rounds floats to 6 places, so 05:26 vs 05:26.000001, or different key orders, don't break "same state" detection. tick_id/collected_at live on TickInput, outside the payload — a re-fire is the same job with the same input, not the same run name.
2. Detection — double-fire requires BOTH rapid interval AND identical state:
class DoubleFireDetector:
def classify(self, current, prior):
fp = fingerprint(current.payload)
if prior is None:
return "normal", fp, [] # first tick can't re-fire
delta = current.collected_at - prior.collected_at
if delta < 0: # clock skew / replay
return "normal", fp, ["out-of-order tick ..."]
same_state = fp == prior.fingerprint
if delta <= self.config.double_fire_window_s and same_state:
return "probable_double_fire", fp, [f"rapid re-fire {delta:.0f}s after {prior.tick_id} ..."]
return "normal", fp, [] # changed input = real event
Window is 360s (T68 landed 180s after T67). Either condition failing means it's a genuine event, not redundancy.
3. Foreman handling — audit always runs, gates never skipped:
class Foreman:
def process(self, tick):
classification, fp, notes = self.detector.classify(tick, prior)
audit_result = self.audit.run(tick.payload) # FULL audit, always executed
gate_results = run_gates(tick.payload, self.gates, self.live) # identical gate set
...
if classification == "probable_double_fire":
rec.board_notes.append(
"board entry: probable cron double-fire flagged; "
"full audit executed (cache-backed); gates NOT skipped")
The Audit is cache-backed by fingerprint (T68 hits T67's cache — that's the "cheap" part) but is invoked every tick; a miss is computed fresh. Gates receive (work_input, live) — live probes like sources_fresh can fail on the re-fire tick (a source went stale between 05:26 and 05:29) without changing the fingerprint, so a real failure is never masked by the double-fire label.
Verified by running the full suite (`pytest`, Python 3.14.4): **11/11 passed**. Direct replay of the incident:
```
T67: normal | gates: {'sources_fresh': True, 'invariants_hold': True, 'board_consistent': True}
T68: probable_double_fire
fingerprint identical: True
audit_ran: True | cache: 1 compute / 1 hit
note: rapid re-fire 180s after T67 with identical fingerprint 7de1dd8fea21 — probable cron double-fire
note: board entry: probable cron double-fire flagged; full audit executed (cache-backed); gates NOT skipped
```
Edge cases tested:
- **Exact incident** — T68 180s after T67, identical input → flagged; audit ran (cache hit, 1 compute/1 hit); all 3 gates evaluated, identical set.
- **Rapid fire + state change** (rows 42→43) → `normal` + "real event" note; redundancy assumption never applied.
- **Slow no-op rerun** (10 min later, same state) → `normal`, no double-fire note.
- **First-ever tick** → `normal` (nothing to re-fire).
- **Out-of-order clock / replay** (negative interval) → `normal` with explicit note.
- **Window boundary** — exactly 360s → flagged; 361s → not.
- **Fingerprint robustness** — dict key order and float noise don't change the hash.
- **Failing live gate on a double-fire tick** — classified double-fire *and* gate `sources_fresh: False` surfaces (`ok` is False); nothing masked.
- **Audit failure** — board note present on both the original and the double-fire tick.
- **3-tick cycle** — normal → double-fire → normal.{"model": "deepseek-v4-flash", "problem_class": "cron-double-fire-detection", "result": "passed", "tests": 11}