◐ Off-By-One · answer catalog

go-verification-double-fire

3 answer(s)godockergodocker

go-verification-double-fire

📦 Source in repository (JSON)

Answer 1

Root cause: the <project> ran its verification pipeline and wrote a board entry without first checking a durable, single-writer claim registry — so when a sibling foreman had already committed tick 39 (083750d), a second full verification and second board entry fired.

Fix (~/&lt;project&gt;/): a &lt;project&gt;.Store enforcing one committed board entry per tick number, plus a VerifyOnce entry point that implements the protocol:

  1. Claim before run — Acquire checks the durable ledger first; only if the tick is absent does it claim an in-progress token via O_CREATE|O_EXCL (atomic across processes, no lock server). Live leases are never stolen; stale leases (crashed owner) are stolen only via compare-and-delete.
// &lt;project&gt;.go
func (s *Store) Acquire(tick int, owner string, now time.Time) (bool, error) {
    // 1. Already on the board? → ErrDuplicate, never run again.
    if _, ok, err := s.GetEntry(tick); err != nil { return false, err } else if ok {
        return false, ErrDuplicate
    }
    // 2. Atomic per-tick claim: only one process wins.
    claimed, err := s.tryClaim(s.tokenPath(tick), Token{Tick: tick, Owner: owner, LeaseUntil: now.Add(s.lease)})
    ...
}
  1. One entry, never two — Commit serializes with flock and refuses a second entry: identical evidence → ErrDuplicate; differing evidence → ErrConflict (original preserved, escalated to humans, never overwritten). The ledger is written temp-file + fsync + atomic rename.
want := HashEvidence(e.Commit, e.Evidence)      // canonical digest, computed here
if cur, ok := entries[e.Tick]; ok {
    if cur.EvidenceSHA == want && cur.Commit == e.Commit {
        return fmt.Errorf("%w: ... identical — second entry refused", ErrDuplicate)
    }
    return fmt.Errorf("%w: ... refusing to overwrite", ErrConflict)
}
  1. Reconcile, don't re-run-and-publish — the losing/sibling path independently re-runs the gates (build, vet, 8+8 tests, guard, govulncheck, hilo-stats, gitreins-count, CI-runs), hashes them canonically (sorted key=commit:result lines, so map order can't alter the digest), compares against the sibling's committed hash, and writes only an attestation sidecar — never a second board entry.

  2. No duplicate workers — ClaimWorker registers worker IDs with O_EXCL (duplicate IDs refused) and additionally refuses to spawn any worker for an already-committed tick. Only the single acquiring process ever reaches the spawn step.

The <project> loop collapses to one call:

outcome, err := &lt;project&gt;.VerifyOnce(s, tick, owner, commit, gates, now)
// outcome: ran-and-committed | duplicate-confirmed | conflict-escalated

Evidence & signatures

`go vet`, `gofmt`, and `go test -race` all pass; 10/10 tests green:

| Test | Verifies |
|---|---|
| `TestIncident_AlreadyCommitted_SiblingWins` | Regression for the exact incident: `Acquire` → `ErrDuplicate`, second `Commit` refused, worker spawn refused, board has **exactly 1** entry owned by `foreman-a` |
| `TestIncident_VerifyOnceReconciles` | My run re-runs gates, matches sibling claims exactly, writes 1 attestation, board untouched |
| `TestRace_ManyCriersOneBoardEntry` | 8 racing <project>s on tick 39 → exactly 1 `ran`, 7 `duplicate-confirmed`, 1 board entry |
| `TestConflict_DifferentEvidenceNeverOverwrites` | Differing evidence → `ErrConflict`, original entry intact |
| `TestVerifyOnce_ConflictEscalated` | Conflict path through public API; board not clobbered |
| `TestLeaseExpiry_StaleTokenStolen` | Live token not stealable; expired (crashed) owner's token stolen; still 1 entry |
| `TestHashEvidence_Deterministic` | Digest order-independent; sensitive to gate result and commit sha |
| `TestWorkers_NoDuplicateSpawns` | Duplicate worker ID refused; committed tick closed to new workers |
| `TestDistinctTicks_Independent` | Different tick numbers proceed independently |
| `TestCorruptLedger_ErrorsNotSilenced` | Torn board errors loudly, never silently double-fires |

Example run (replays incident): `foreman-a → ran-and-committed`; `<project>-b → duplicate-confirmed`; `ci-recheck → duplicate-confirmed`; worker dup refused; **1** board entry for tick 39, sha `9e3f6693e0…`.

Edge cases covered: same-tick races, genuine conflicts (escalate, never overwrite), crashed-owner leases, duplicate worker processes, corrupt/torn ledger, and distinct-tick independence.
{"model": "deepseek-v4-flash", "problem_class": "go-verification-double-fire", "result": "passed", "tests": 10}

Answer 2

Root cause. Helix tick #60 found the tick #59 verification set double-fired for commit 6012714: two sibling sessions raced, and the loser's stale output was written as truth. Three concrete failures:

  1. Forgejo "DOWN" was a stale-port hallucination — the gate probed a hardcoded :8080 (a proxy answering 404) instead of the pinned :3030.
  2. Cooldown 7200 was a stale claim — fleet.toml pins cooldown = 900; 7200 came from an older fleet generation.
  3. CI-FIX-E2E-SKIP closed green with no verdict.json — judge transport failed, the artifact was never produced, yet CI reported green.

The fix makes each gate read pinned config (never sibling claims), refuse green without the judge artifact, and enforce single-fire idempotency per (tick, commit, gate) with the cooldown pin. Implementation (stdlib-only Go, built and tested at /tmp/gvf):

Fix 1 — healthcheck: pinned port + content sniff (kills the stale :8080 claim)

func HealthCheck(ctx context.Context, base string) error {
    u, err := url.Parse(base)
    if err != nil { return fmt.Errorf("forgejo_url parse: %w", err) }
    if u.Scheme == "" || u.Host == "" { return fmt.Errorf("forgejo_url %q is not absolute", base) }
    if u.Port() == "" {
        return fmt.Errorf("forgejo_url %q has no port: pin it (3030), never fall back to 8080", base)
    }
    req, _ := http.NewRequestWithContext(ctx, http.MethodGet, strings.TrimRight(base, "/")+"/", nil)
    resp, err := (&http.Client{Timeout: 5 * time.Second}).Do(req)
    if err != nil { return fmt.Errorf("forgejo unreachable at %s: %w", base, err) }
    defer resp.Body.Close()
    if resp.StatusCode != http.StatusOK { return fmt.Errorf("forgejo at %s returned %s", base, resp.Status) }
    body, _ := io.ReadAll(io.LimitReader(resp.Body, 64<<10))
    if !bytes.Contains(bytes.ToLower(body), []byte("forgejo")) {
        return fmt.Errorf("%w: %s answers %s but body does not identify Forgejo", ErrWrongService, base, resp.Status)
    }
    return nil
}

Fix 2 — fleet.toml cooldown pin (900, not 7200)

const DefaultCooldown = 900 * time.Second // fleet.toml pin

func LoadConfig(path string) (*Config, error) {
    cfg := &Config{Cooldown: DefaultCooldown, Forgejo: "http://<ip-address>:3030",
        Verdict: "verdict.json", MaxRefire: 3}
    b, err := os.ReadFile(path)
    if err != nil {
        if errors.Is(err, os.ErrNotExist) { return cfg, nil } // pins win when file absent
        return nil, err
    }
    for i, line := range strings.Split(string(b), "\n") {
        line = strings.TrimSpace(strings.SplitN(line, "#", 2)[0])
        if line == "" { continue }
        key, val, ok := strings.Cut(line, "=")
        if !ok { continue }
        key, val = strings.TrimSpace(key), strings.Trim(strings.TrimSpace(val), `"'`)
        switch key {
        case "cooldown":
            n, err := strconv.Atoi(val)
            if err != nil || n <= 0 { return nil, fmt.Errorf("fleet.toml:%d: bad cooldown", i+1) }
            cfg.Cooldown = time.Duration(n) * time.Second
        case "forgejo_url":
            if !strings.HasPrefix(val, "http") { return nil, fmt.Errorf("fleet.toml:%d: forgejo_url must be absolute", i+1) }
            cfg.Forgejo = val
        case "verdict_file":
            cfg.Verdict = val
        }
    }
    return cfg, nil
}

Fix 3 — CI-FIX-E2E-SKIP: green is unrepresentable without verdict.json

func LoadVerdict(path string) (*Verdict, error) {
    b, err := os.ReadFile(path)
    if err != nil {
        return nil, fmt.Errorf("verdict.json missing (%v): judge transport failure must be retried, never reported green", err)
    }
    var v Verdict
    if err := json.Unmarshal(b, &v); err != nil { return nil, fmt.Errorf("verdict.json corrupt: %w", err) }
    if v.Commit == "" { return nil, fmt.Errorf("verdict.json incomplete: commit required") }
    return &v, nil
}

func Green(v *Verdict) (bool, string) {
    switch v.Status {
    case "pass":
        return true, "verdict.json present: status=pass"
    case "fail":
        return false, "verdict.json present: status=fail"
    case "skip":
        if v.SkipWhy == "" { return false, "verdict.json status=skip without skip_why is NOT green" }
        return true, "verdict.json explicit skip: " + v.SkipWhy
    default:
        return false, "verdict.json unknown status " + v.Status
    }
}

Fix 4 — single-fire claims + cooldown + one DuckBrain record (anti-double-fire)

func ClaimKey(tick int, commit, gate string) string {
    return fmt.Sprintf("%d/%s/%s", tick, commit, gate) // deterministic across sessions
}

func Acquire(store ClaimStore, c Claim, cooldown time.Duration) error {
    if err := store.Insert(c); err != nil {
        if errors.Is(err, ErrAlreadyClaimed) {
            if prev, ok := store.Lookup(c.Key); ok && time.Since(prev.At) < cooldown {
                return ErrCooldown // inside the 900s pin: refire blocked
            }
            return ErrAlreadyClaimed // already ran this tick: permanent
        }
        return err
    }
    return nil
}

func WriteOne(path string, r Record) error { // one DuckBrain record per tick
    if b, err := os.ReadFile(path); err == nil {
        var prev Record
        if json.Unmarshal(b, &prev) == nil && prev.Tick == r.Tick {
            return errors.New("duckbrain record already exists for this tick: refusing to overwrite")
        }
    }
    r.At = time.Now().UTC().Format(time.RFC3339)
    b, err := json.MarshalIndent(r, "", "  ")
    if err != nil { return err }
    return os.WriteFile(path, append(b, '\n'), 0o644)
}

Per-tick order: Acquire → run gate → CIGate(verdict.json) → HealthCheck(cfg.Forgejo) → write one DuckBrain record. The second fire of the same slot is rejected before publishing; the record writer refuses overwrites.

Evidence & signatures

**Live verification of the two wrong sibling claims (probed in this session):**

| Probe | Result | Meaning |
|---|---|---|
| `GET http://<ip-address>:3030/` | HTTP 200, `<title>Forgejo: Beyond coding. We forge.</title>` | Forgejo **UP on :3030** — sibling's "DOWN" claim is wrong |
| `GET http://<ip-address>:8080/` | HTTP 404 text/plain | :8080 is a different service — the stale check probed the wrong port |
| `GET http://<ip-address>:9090/api/v1/health` | `{"active_ticks":6,"db":"connected","status":"ok",...}` | Fleet scheduler live |
| `GET http://<ip-address>:9224/` | `{"status":"ok","active":1,...}` | Process runner live |

**Gate results (real commands, real exit codes):**

```
gofmt -l .                 -> (empty)          lint gate: clean
go build ./...             -> exit 0           build gate: ok
go vet ./...               -> exit 0           vet gate: ok
go test -race -count=1 ./... -> ok gvf/gate 1.033s   tests+race: 17/17 PASS
```

**Edge cases tested (17/17 pass):** Forgejo up on real port; 200 from wrong service caught by content sniff (`ErrWrongService`); 404 is down; URL without pinned port rejected; cooldown pin 900 from fleet.toml; missing fleet.toml falls back to pins; stale file value honoured but stale claim ignored by construction; bad cooldown errors; missing verdict.json → hard error (never green); pass → green; fail → red; skip without `skip_why` → red, with reason → green; incomplete verdict (no commit) rejected; first claim OK + sibling duplicate rejected + new tick never blocked; refire inside 900s window blocked; claim key deterministic across sessions; second DuckBrain record for tick 60 refused.

**Protocol compliance:** commit NOTHING (no repository was touched); exactly one DuckBrain record written for tick 60 (`forgejo: "up"`, `cooldown: 900`, `verdict: pass`, `event_id: 4`, `commit: 6012714`). Deliverables written to `/workspace/solution.md`, `/workspace/evidence.md`, `/workspace/signatures.json`.
{"model": "deepseek-v4-flash", "problem_class": "go-verification-double-fire", "result": "passed", "tests": 17}

Answer 3

Root cause: the <project> ran its verification pipeline and wrote a board entry without first checking a durable, single-writer claim registry — so when a sibling foreman had already committed tick 39 (083750d), a second full verification and second board entry fired.

Fix (~/&lt;project&gt;/): a &lt;project&gt;.Store enforcing one committed board entry per tick number, plus a VerifyOnce entry point that implements the protocol:

  1. Claim before run — Acquire checks the durable ledger first; only if the tick is absent does it claim an in-progress token via O_CREATE|O_EXCL (atomic across processes, no lock server). Live leases are never stolen; stale leases (crashed owner) are stolen only via compare-and-delete.
// &lt;project&gt;.go
func (s *Store) Acquire(tick int, owner string, now time.Time) (bool, error) {
    // 1. Already on the board? → ErrDuplicate, never run again.
    if _, ok, err := s.GetEntry(tick); err != nil { return false, err } else if ok {
        return false, ErrDuplicate
    }
    // 2. Atomic per-tick claim: only one process wins.
    claimed, err := s.tryClaim(s.tokenPath(tick), Token{Tick: tick, Owner: owner, LeaseUntil: now.Add(s.lease)})
    ...
}
  1. One entry, never two — Commit serializes with flock and refuses a second entry: identical evidence → ErrDuplicate; differing evidence → ErrConflict (original preserved, escalated to humans, never overwritten). The ledger is written temp-file + fsync + atomic rename.
want := HashEvidence(e.Commit, e.Evidence)      // canonical digest, computed here
if cur, ok := entries[e.Tick]; ok {
    if cur.EvidenceSHA == want && cur.Commit == e.Commit {
        return fmt.Errorf("%w: ... identical — second entry refused", ErrDuplicate)
    }
    return fmt.Errorf("%w: ... refusing to overwrite", ErrConflict)
}
  1. Reconcile, don't re-run-and-publish — the losing/sibling path independently re-runs the gates (build, vet, 8+8 tests, guard, govulncheck, hilo-stats, gitreins-count, CI-runs), hashes them canonically (sorted key=commit:result lines, so map order can't alter the digest), compares against the sibling's committed hash, and writes only an attestation sidecar — never a second board entry.

  2. No duplicate workers — ClaimWorker registers worker IDs with O_EXCL (duplicate IDs refused) and additionally refuses to spawn any worker for an already-committed tick. Only the single acquiring process ever reaches the spawn step.

The <project> loop collapses to one call:

outcome, err := &lt;project&gt;.VerifyOnce(s, tick, owner, commit, gates, now)
// outcome: ran-and-committed | duplicate-confirmed | conflict-escalated

Evidence & signatures

`go vet`, `gofmt`, and `go test -race` all pass; 10/10 tests green:

| Test | Verifies |
|---|---|
| `TestIncident_AlreadyCommitted_SiblingWins` | Regression for the exact incident: `Acquire` → `ErrDuplicate`, second `Commit` refused, worker spawn refused, board has **exactly 1** entry owned by `foreman-a` |
| `TestIncident_VerifyOnceReconciles` | My run re-runs gates, matches sibling claims exactly, writes 1 attestation, board untouched |
| `TestRace_ManyCriersOneBoardEntry` | 8 racing <project>s on tick 39 → exactly 1 `ran`, 7 `duplicate-confirmed`, 1 board entry |
| `TestConflict_DifferentEvidenceNeverOverwrites` | Differing evidence → `ErrConflict`, original entry intact |
| `TestVerifyOnce_ConflictEscalated` | Conflict path through public API; board not clobbered |
| `TestLeaseExpiry_StaleTokenStolen` | Live token not stealable; expired (crashed) owner's token stolen; still 1 entry |
| `TestHashEvidence_Deterministic` | Digest order-independent; sensitive to gate result and commit sha |
| `TestWorkers_NoDuplicateSpawns` | Duplicate worker ID refused; committed tick closed to new workers |
| `TestDistinctTicks_Independent` | Different tick numbers proceed independently |
| `TestCorruptLedger_ErrorsNotSilenced` | Torn board errors loudly, never silently double-fires |

Example run (replays incident): `foreman-a → ran-and-committed`; `<project>-b → duplicate-confirmed`; `ci-recheck → duplicate-confirmed`; worker dup refused; **1** board entry for tick 39, sha `9e3f6693e0…`.

Edge cases covered: same-tick races, genuine conflicts (escalate, never overwrite), crashed-owner leases, duplicate worker processes, corrupt/torn ledger, and distinct-tick independence.
{"model": "deepseek-v4-flash", "problem_class": "go-verification-double-fire", "result": "passed", "tests": 10}
Generated from the verified corpus · MIT licensedBack to the catalog