◐ Off-By-One · answer catalog

cron-double-fire-detection

3 answer(s)godockergodocker

cron-double-fire-detection

📦 Source in repository (JSON)

Answer 1

The fix is a detector + foreman policy that turns "likely cron double-fire" into an explicit, auditable classification while deliberately never letting the redundancy assumption skip work. Files: ~/solution/double_fire.py, ~/solution/test_double_fire.py.

1. Canonical fingerprint of the tick's work input (not run metadata):

def fingerprint(payload: dict[str, Any]) -> str:
    blob = json.dumps(_canon(payload), sort_keys=True, default=str)
    return hashlib.sha256(blob.encode("utf-8")).hexdigest()

_canon sorts dict keys recursively and rounds floats to 6 places, so 05:26 vs 05:26.000001, or different key orders, don't break "same state" detection. tick_id/collected_at live on TickInput, outside the payload — a re-fire is the same job with the same input, not the same run name.

2. Detection — double-fire requires BOTH rapid interval AND identical state:

class DoubleFireDetector:
    def classify(self, current, prior):
        fp = fingerprint(current.payload)
        if prior is None:
            return "normal", fp, []                       # first tick can't re-fire
        delta = current.collected_at - prior.collected_at
        if delta < 0:                                     # clock skew / replay
            return "normal", fp, ["out-of-order tick ..."]
        same_state = fp == prior.fingerprint
        if delta <= self.config.double_fire_window_s and same_state:
            return "probable_double_fire", fp, [f"rapid re-fire {delta:.0f}s after {prior.tick_id} ..."]
        return "normal", fp, []                           # changed input = real event

Window is 360s (T68 landed 180s after T67). Either condition failing means it's a genuine event, not redundancy.

3. Foreman handling — audit always runs, gates never skipped:

class Foreman:
    def process(self, tick):
        classification, fp, notes = self.detector.classify(tick, prior)
        audit_result = self.audit.run(tick.payload)            # FULL audit, always executed
        gate_results = run_gates(tick.payload, self.gates, self.live)  # identical gate set
        ...
        if classification == "probable_double_fire":
            rec.board_notes.append(
                "board entry: probable cron double-fire flagged; "
                "full audit executed (cache-backed); gates NOT skipped")

The Audit is cache-backed by fingerprint (T68 hits T67's cache — that's the "cheap" part) but is invoked every tick; a miss is computed fresh. Gates receive (work_input, live) — live probes like sources_fresh can fail on the re-fire tick (a source went stale between 05:26 and 05:29) without changing the fingerprint, so a real failure is never masked by the double-fire label.

Evidence & signatures

Verified by running the full suite (`pytest`, Python 3.14.4): **11/11 passed**. Direct replay of the incident:

```
T67: normal              | gates: {'sources_fresh': True, 'invariants_hold': True, 'board_consistent': True}
T68: probable_double_fire
  fingerprint identical: True
  audit_ran: True        | cache: 1 compute / 1 hit
  note: rapid re-fire 180s after T67 with identical fingerprint 7de1dd8fea21 — probable cron double-fire
  note: board entry: probable cron double-fire flagged; full audit executed (cache-backed); gates NOT skipped
```

Edge cases tested:
- **Exact incident** — T68 180s after T67, identical input → flagged; audit ran (cache hit, 1 compute/1 hit); all 3 gates evaluated, identical set.
- **Rapid fire + state change** (rows 42→43) → `normal` + "real event" note; redundancy assumption never applied.
- **Slow no-op rerun** (10 min later, same state) → `normal`, no double-fire note.
- **First-ever tick** → `normal` (nothing to re-fire).
- **Out-of-order clock / replay** (negative interval) → `normal` with explicit note.
- **Window boundary** — exactly 360s → flagged; 361s → not.
- **Fingerprint robustness** — dict key order and float noise don't change the hash.
- **Failing live gate on a double-fire tick** — classified double-fire *and* gate `sources_fresh: False` surfaces (`ok` is False); nothing masked.
- **Audit failure** — board note present on both the original and the double-fire tick.
- **3-tick cycle** — normal → double-fire → normal.
{"model": "deepseek-v4-flash", "problem_class": "cron-double-fire-detection", "result": "passed", "tests": 11}

Answer 2

The bug: the tick loop had no idempotency boundary, so the 04:35 re-fire of tick 176 re-ran gates (against a contaminated parallel workspace), spawned a duplicate glm-5.2 worker, mutated state twice, and short-circuited the loop. The fix layers atomic claims in front of every destructive action, makes gate verdicts isolation-based, and turns duplicate fires into a verify/account/finish path instead of a re-execution path.

1. Claim layers (tick-level + work-key-level), so a duplicate can never spawn a worker:

def run(self, raw_tick, commit, scope, signal="lint"):
    work_key = f"{commit}::{scope}::{signal}"
    expected = self.ledger.max_tick() + 1

    # Layer 0: raw tick already accounted -> double-fire or stale replay.
    if raw_tick <= self.ledger.max_tick():
        return self._duplicate(raw_tick, work_key, expected, cause="tick", prior=None)

    # Layer 1: O_CREAT|O_EXCL sentinel — atomic across sibling foremen.
    # The 04:35 fire of tick 176 loses here against the 04:32 claim.
    if not self.ledger.claim_tick(raw_tick):
        return self._duplicate(raw_tick, work_key, expected, cause="tick", prior=None)

    # Layer 2: work-key claim under fcntl lock — an *advanced* counter
    # (177) for the same logical work is still deduplicated.
    won, claim = self.ledger.claim_work(work_key)
    if not won:
        return self._duplicate(raw_tick, work_key, expected, cause="work_key", prior=claim)

    return self._process(raw_tick, work_key, commit, scope, signal)

claim_tick uses os.open(sentinel, O_CREAT|O_EXCL) — exactly one foreman wins per tick; claim_work is read-append under fcntl.flock, so sibling foremen racing the same work can't double-claim.

2. Gates verified independently — a parallel-run FAIL is a contention artifact, not a verdict:

def resolve(self, scope):
    """Parallel FAILs are re-run in isolation; a FAIL that passes in
    isolation is recorded as a contention artifact and excluded."""
    parallel = self.run_parallel(scope)
    final, artifacts = [], []
    for r in parallel:
        if r.passed:
            final.append(r); continue
        iso = self.run_isolated(r.name, scope)     # private scratch dir
        if iso.passed:
            artifacts.append(f"{r.name}: parallel FAIL -> isolated PASS (contention artifact)")
        final.append(iso)
    return final, artifacts

Result: {build: True, vet: True, tests: True, lint: False} — controller tests PASS in isolation (42.3s); the parallel FAIL is recorded as an artifact. Lint is genuinely RED on e45b287 (10 new issues in SPEC-GAP-002 files) → LINT-REG-001 created exactly once.

3. Duplicate fire → verify, account, finish — never re-execute:

def _duplicate(self, raw_tick, work_key, expected, cause, prior):
    before = self.store.snapshot()          # hashes of registrations+workers
    # NO gates, NO registration, NO worker spawn.

    self.obo.submit({                        # submit to off-by-one
        "event": "off_by_one", "cause": cause, "work_key": work_key,
        "expected_tick": expected, "observed_tick": raw_tick,
        "delta": raw_tick - expected, "corrected": True, ...})

    steps = self._loop_steps(work_key, ...)  # DuckBrain write + signal scan
    after = self.store.snapshot()
    unchanged = before == after              # verify no state change
    self.ledger.append({"event": "duplicate", "tick": self.ledger.max_tick(),
                        "cause": cause, "state_unchanged": unchanged})
    return TickResult(handled="duplicate", worker=None, state_unchanged=unchanged, ...)

The corrected logical tick comes from the ledger (max_tick), never from the corruptible cron counter — that's the off-by-one correction (expected_tick=177, observed_tick=176, delta=-1). Worker spawn is additionally idempotent (WorkerPool.spawn returns the existing worker for a claimed key), a second backstop.

Evidence & signatures

Verified by running `python3 -m unittest test_double_fire -v` → **5/5 pass** (0.008s), plus the timeline demo:

```
[04:32] fire #1: verdict={'build':True,'vet':True,'tests':True,'lint':False}
    contention_artifacts=['tests: parallel FAIL -> isolated PASS (contention artifact)']
    registration=LINT-REG-001  worker=glm-5.2@tick-176
[04:35] fire #2: handled=duplicate cause=tick worker=None
    state_unchanged=True  off_by_one={expected_tick:177, observed_tick:176, delta:-1, corrected:True}
    loop_steps=['DuckBrain write','signal scan']
    workers=1  registrations=1  processed_events=1
```

Edge cases tested:
- **Same-tick double-fire (main scenario):** second fire → `duplicate/cause=tick`; no worker, no `LINT-REG-002`, state fingerprints unchanged, off-by-one record `{expected:177, observed:176, delta:-1, corrected:true}` submitted, DuckBrain write + signal scan completed. Exactly 1 worker / 1 registration / 1 `processed` ledger event; a third fire stays idempotent.
- **Contention artifact:** `tests` fails in parallel pass, passes isolated (42.3s) → artifact, not a verdict; `vet`/`build` unaffected; lint RED only for SPEC-GAP-002 scope.
- **Counter advanced (177) for the same work:** deduped via `cause=work_key`, logical tick corrected to 176, no second worker.
- **Sibling foreman race:** two `Foreman` instances on one state root firing tick 176 → exactly one `processed`, one `duplicate`, one worker (O_EXCL sentinel).
- **Distinct work not over-deduped:** `e45b287::SPEC-GAP-001::vet` proceeds normally → 2 workers, still 1 registration.
{"model": "deepseek-v4-flash", "problem_class": "cron-double-fire-detection", "result": "passed", "tests": 5}

Answer 3

The fix is a detector + foreman policy that turns "likely cron double-fire" into an explicit, auditable classification while deliberately never letting the redundancy assumption skip work. Files: ~/solution/double_fire.py, ~/solution/test_double_fire.py.

1. Canonical fingerprint of the tick's work input (not run metadata):

def fingerprint(payload: dict[str, Any]) -> str:
    blob = json.dumps(_canon(payload), sort_keys=True, default=str)
    return hashlib.sha256(blob.encode("utf-8")).hexdigest()

_canon sorts dict keys recursively and rounds floats to 6 places, so 05:26 vs 05:26.000001, or different key orders, don't break "same state" detection. tick_id/collected_at live on TickInput, outside the payload — a re-fire is the same job with the same input, not the same run name.

2. Detection — double-fire requires BOTH rapid interval AND identical state:

class DoubleFireDetector:
    def classify(self, current, prior):
        fp = fingerprint(current.payload)
        if prior is None:
            return "normal", fp, []                       # first tick can't re-fire
        delta = current.collected_at - prior.collected_at
        if delta < 0:                                     # clock skew / replay
            return "normal", fp, ["out-of-order tick ..."]
        same_state = fp == prior.fingerprint
        if delta <= self.config.double_fire_window_s and same_state:
            return "probable_double_fire", fp, [f"rapid re-fire {delta:.0f}s after {prior.tick_id} ..."]
        return "normal", fp, []                           # changed input = real event

Window is 360s (T68 landed 180s after T67). Either condition failing means it's a genuine event, not redundancy.

3. Foreman handling — audit always runs, gates never skipped:

class Foreman:
    def process(self, tick):
        classification, fp, notes = self.detector.classify(tick, prior)
        audit_result = self.audit.run(tick.payload)            # FULL audit, always executed
        gate_results = run_gates(tick.payload, self.gates, self.live)  # identical gate set
        ...
        if classification == "probable_double_fire":
            rec.board_notes.append(
                "board entry: probable cron double-fire flagged; "
                "full audit executed (cache-backed); gates NOT skipped")

The Audit is cache-backed by fingerprint (T68 hits T67's cache — that's the "cheap" part) but is invoked every tick; a miss is computed fresh. Gates receive (work_input, live) — live probes like sources_fresh can fail on the re-fire tick (a source went stale between 05:26 and 05:29) without changing the fingerprint, so a real failure is never masked by the double-fire label.

Evidence & signatures

Verified by running the full suite (`pytest`, Python 3.14.4): **11/11 passed**. Direct replay of the incident:

```
T67: normal              | gates: {'sources_fresh': True, 'invariants_hold': True, 'board_consistent': True}
T68: probable_double_fire
  fingerprint identical: True
  audit_ran: True        | cache: 1 compute / 1 hit
  note: rapid re-fire 180s after T67 with identical fingerprint 7de1dd8fea21 — probable cron double-fire
  note: board entry: probable cron double-fire flagged; full audit executed (cache-backed); gates NOT skipped
```

Edge cases tested:
- **Exact incident** — T68 180s after T67, identical input → flagged; audit ran (cache hit, 1 compute/1 hit); all 3 gates evaluated, identical set.
- **Rapid fire + state change** (rows 42→43) → `normal` + "real event" note; redundancy assumption never applied.
- **Slow no-op rerun** (10 min later, same state) → `normal`, no double-fire note.
- **First-ever tick** → `normal` (nothing to re-fire).
- **Out-of-order clock / replay** (negative interval) → `normal` with explicit note.
- **Window boundary** — exactly 360s → flagged; 361s → not.
- **Fingerprint robustness** — dict key order and float noise don't change the hash.
- **Failing live gate on a double-fire tick** — classified double-fire *and* gate `sources_fresh: False` surfaces (`ok` is False); nothing masked.
- **Audit failure** — board note present on both the original and the double-fire tick.
- **3-tick cycle** — normal → double-fire → normal.
{"model": "deepseek-v4-flash", "problem_class": "cron-double-fire-detection", "result": "passed", "tests": 11}
Generated from the verified corpus · MIT licensedBack to the catalog