go-verification-double-fire
Root cause: the <project> ran its verification pipeline and wrote a board entry without first checking a durable, single-writer claim registry — so when a sibling foreman had already committed tick 39 (083750d), a second full verification and second board entry fired.
Fix (~/<project>/): a <project>.Store enforcing one committed board entry per tick number, plus a VerifyOnce entry point that implements the protocol:
Acquire checks the durable ledger first; only if the tick is absent does it claim an in-progress token via O_CREATE|O_EXCL (atomic across processes, no lock server). Live leases are never stolen; stale leases (crashed owner) are stolen only via compare-and-delete.// <project>.go
func (s *Store) Acquire(tick int, owner string, now time.Time) (bool, error) {
// 1. Already on the board? → ErrDuplicate, never run again.
if _, ok, err := s.GetEntry(tick); err != nil { return false, err } else if ok {
return false, ErrDuplicate
}
// 2. Atomic per-tick claim: only one process wins.
claimed, err := s.tryClaim(s.tokenPath(tick), Token{Tick: tick, Owner: owner, LeaseUntil: now.Add(s.lease)})
...
}
Commit serializes with flock and refuses a second entry: identical evidence → ErrDuplicate; differing evidence → ErrConflict (original preserved, escalated to humans, never overwritten). The ledger is written temp-file + fsync + atomic rename.want := HashEvidence(e.Commit, e.Evidence) // canonical digest, computed here
if cur, ok := entries[e.Tick]; ok {
if cur.EvidenceSHA == want && cur.Commit == e.Commit {
return fmt.Errorf("%w: ... identical — second entry refused", ErrDuplicate)
}
return fmt.Errorf("%w: ... refusing to overwrite", ErrConflict)
}
Reconcile, don't re-run-and-publish — the losing/sibling path independently re-runs the gates (build, vet, 8+8 tests, guard, govulncheck, hilo-stats, gitreins-count, CI-runs), hashes them canonically (sorted key=commit:result lines, so map order can't alter the digest), compares against the sibling's committed hash, and writes only an attestation sidecar — never a second board entry.
No duplicate workers — ClaimWorker registers worker IDs with O_EXCL (duplicate IDs refused) and additionally refuses to spawn any worker for an already-committed tick. Only the single acquiring process ever reaches the spawn step.
The <project> loop collapses to one call:
outcome, err := <project>.VerifyOnce(s, tick, owner, commit, gates, now)
// outcome: ran-and-committed | duplicate-confirmed | conflict-escalated
`go vet`, `gofmt`, and `go test -race` all pass; 10/10 tests green: | Test | Verifies | |---|---| | `TestIncident_AlreadyCommitted_SiblingWins` | Regression for the exact incident: `Acquire` → `ErrDuplicate`, second `Commit` refused, worker spawn refused, board has **exactly 1** entry owned by `foreman-a` | | `TestIncident_VerifyOnceReconciles` | My run re-runs gates, matches sibling claims exactly, writes 1 attestation, board untouched | | `TestRace_ManyCriersOneBoardEntry` | 8 racing <project>s on tick 39 → exactly 1 `ran`, 7 `duplicate-confirmed`, 1 board entry | | `TestConflict_DifferentEvidenceNeverOverwrites` | Differing evidence → `ErrConflict`, original entry intact | | `TestVerifyOnce_ConflictEscalated` | Conflict path through public API; board not clobbered | | `TestLeaseExpiry_StaleTokenStolen` | Live token not stealable; expired (crashed) owner's token stolen; still 1 entry | | `TestHashEvidence_Deterministic` | Digest order-independent; sensitive to gate result and commit sha | | `TestWorkers_NoDuplicateSpawns` | Duplicate worker ID refused; committed tick closed to new workers | | `TestDistinctTicks_Independent` | Different tick numbers proceed independently | | `TestCorruptLedger_ErrorsNotSilenced` | Torn board errors loudly, never silently double-fires | Example run (replays incident): `foreman-a → ran-and-committed`; `<project>-b → duplicate-confirmed`; `ci-recheck → duplicate-confirmed`; worker dup refused; **1** board entry for tick 39, sha `9e3f6693e0…`. Edge cases covered: same-tick races, genuine conflicts (escalate, never overwrite), crashed-owner leases, duplicate worker processes, corrupt/torn ledger, and distinct-tick independence.
{"model": "deepseek-v4-flash", "problem_class": "go-verification-double-fire", "result": "passed", "tests": 10}Root cause. Helix tick #60 found the tick #59 verification set double-fired for commit 6012714: two sibling sessions raced, and the loser's stale output was written as truth. Three concrete failures:
:8080 (a proxy answering 404) instead of the pinned :3030.fleet.toml pins cooldown = 900; 7200 came from an older fleet generation.The fix makes each gate read pinned config (never sibling claims), refuse green without the judge artifact, and enforce single-fire idempotency per (tick, commit, gate) with the cooldown pin. Implementation (stdlib-only Go, built and tested at /tmp/gvf):
Fix 1 — healthcheck: pinned port + content sniff (kills the stale :8080 claim)
func HealthCheck(ctx context.Context, base string) error {
u, err := url.Parse(base)
if err != nil { return fmt.Errorf("forgejo_url parse: %w", err) }
if u.Scheme == "" || u.Host == "" { return fmt.Errorf("forgejo_url %q is not absolute", base) }
if u.Port() == "" {
return fmt.Errorf("forgejo_url %q has no port: pin it (3030), never fall back to 8080", base)
}
req, _ := http.NewRequestWithContext(ctx, http.MethodGet, strings.TrimRight(base, "/")+"/", nil)
resp, err := (&http.Client{Timeout: 5 * time.Second}).Do(req)
if err != nil { return fmt.Errorf("forgejo unreachable at %s: %w", base, err) }
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK { return fmt.Errorf("forgejo at %s returned %s", base, resp.Status) }
body, _ := io.ReadAll(io.LimitReader(resp.Body, 64<<10))
if !bytes.Contains(bytes.ToLower(body), []byte("forgejo")) {
return fmt.Errorf("%w: %s answers %s but body does not identify Forgejo", ErrWrongService, base, resp.Status)
}
return nil
}
Fix 2 — fleet.toml cooldown pin (900, not 7200)
const DefaultCooldown = 900 * time.Second // fleet.toml pin
func LoadConfig(path string) (*Config, error) {
cfg := &Config{Cooldown: DefaultCooldown, Forgejo: "http://<ip-address>:3030",
Verdict: "verdict.json", MaxRefire: 3}
b, err := os.ReadFile(path)
if err != nil {
if errors.Is(err, os.ErrNotExist) { return cfg, nil } // pins win when file absent
return nil, err
}
for i, line := range strings.Split(string(b), "\n") {
line = strings.TrimSpace(strings.SplitN(line, "#", 2)[0])
if line == "" { continue }
key, val, ok := strings.Cut(line, "=")
if !ok { continue }
key, val = strings.TrimSpace(key), strings.Trim(strings.TrimSpace(val), `"'`)
switch key {
case "cooldown":
n, err := strconv.Atoi(val)
if err != nil || n <= 0 { return nil, fmt.Errorf("fleet.toml:%d: bad cooldown", i+1) }
cfg.Cooldown = time.Duration(n) * time.Second
case "forgejo_url":
if !strings.HasPrefix(val, "http") { return nil, fmt.Errorf("fleet.toml:%d: forgejo_url must be absolute", i+1) }
cfg.Forgejo = val
case "verdict_file":
cfg.Verdict = val
}
}
return cfg, nil
}
Fix 3 — CI-FIX-E2E-SKIP: green is unrepresentable without verdict.json
func LoadVerdict(path string) (*Verdict, error) {
b, err := os.ReadFile(path)
if err != nil {
return nil, fmt.Errorf("verdict.json missing (%v): judge transport failure must be retried, never reported green", err)
}
var v Verdict
if err := json.Unmarshal(b, &v); err != nil { return nil, fmt.Errorf("verdict.json corrupt: %w", err) }
if v.Commit == "" { return nil, fmt.Errorf("verdict.json incomplete: commit required") }
return &v, nil
}
func Green(v *Verdict) (bool, string) {
switch v.Status {
case "pass":
return true, "verdict.json present: status=pass"
case "fail":
return false, "verdict.json present: status=fail"
case "skip":
if v.SkipWhy == "" { return false, "verdict.json status=skip without skip_why is NOT green" }
return true, "verdict.json explicit skip: " + v.SkipWhy
default:
return false, "verdict.json unknown status " + v.Status
}
}
Fix 4 — single-fire claims + cooldown + one DuckBrain record (anti-double-fire)
func ClaimKey(tick int, commit, gate string) string {
return fmt.Sprintf("%d/%s/%s", tick, commit, gate) // deterministic across sessions
}
func Acquire(store ClaimStore, c Claim, cooldown time.Duration) error {
if err := store.Insert(c); err != nil {
if errors.Is(err, ErrAlreadyClaimed) {
if prev, ok := store.Lookup(c.Key); ok && time.Since(prev.At) < cooldown {
return ErrCooldown // inside the 900s pin: refire blocked
}
return ErrAlreadyClaimed // already ran this tick: permanent
}
return err
}
return nil
}
func WriteOne(path string, r Record) error { // one DuckBrain record per tick
if b, err := os.ReadFile(path); err == nil {
var prev Record
if json.Unmarshal(b, &prev) == nil && prev.Tick == r.Tick {
return errors.New("duckbrain record already exists for this tick: refusing to overwrite")
}
}
r.At = time.Now().UTC().Format(time.RFC3339)
b, err := json.MarshalIndent(r, "", " ")
if err != nil { return err }
return os.WriteFile(path, append(b, '\n'), 0o644)
}
Per-tick order: Acquire → run gate → CIGate(verdict.json) → HealthCheck(cfg.Forgejo) → write one DuckBrain record. The second fire of the same slot is rejected before publishing; the record writer refuses overwrites.
**Live verification of the two wrong sibling claims (probed in this session):**
| Probe | Result | Meaning |
|---|---|---|
| `GET http://<ip-address>:3030/` | HTTP 200, `<title>Forgejo: Beyond coding. We forge.</title>` | Forgejo **UP on :3030** — sibling's "DOWN" claim is wrong |
| `GET http://<ip-address>:8080/` | HTTP 404 text/plain | :8080 is a different service — the stale check probed the wrong port |
| `GET http://<ip-address>:9090/api/v1/health` | `{"active_ticks":6,"db":"connected","status":"ok",...}` | Fleet scheduler live |
| `GET http://<ip-address>:9224/` | `{"status":"ok","active":1,...}` | Process runner live |
**Gate results (real commands, real exit codes):**
```
gofmt -l . -> (empty) lint gate: clean
go build ./... -> exit 0 build gate: ok
go vet ./... -> exit 0 vet gate: ok
go test -race -count=1 ./... -> ok gvf/gate 1.033s tests+race: 17/17 PASS
```
**Edge cases tested (17/17 pass):** Forgejo up on real port; 200 from wrong service caught by content sniff (`ErrWrongService`); 404 is down; URL without pinned port rejected; cooldown pin 900 from fleet.toml; missing fleet.toml falls back to pins; stale file value honoured but stale claim ignored by construction; bad cooldown errors; missing verdict.json → hard error (never green); pass → green; fail → red; skip without `skip_why` → red, with reason → green; incomplete verdict (no commit) rejected; first claim OK + sibling duplicate rejected + new tick never blocked; refire inside 900s window blocked; claim key deterministic across sessions; second DuckBrain record for tick 60 refused.
**Protocol compliance:** commit NOTHING (no repository was touched); exactly one DuckBrain record written for tick 60 (`forgejo: "up"`, `cooldown: 900`, `verdict: pass`, `event_id: 4`, `commit: 6012714`). Deliverables written to `/workspace/solution.md`, `/workspace/evidence.md`, `/workspace/signatures.json`.{"model": "deepseek-v4-flash", "problem_class": "go-verification-double-fire", "result": "passed", "tests": 17}Root cause: the <project> ran its verification pipeline and wrote a board entry without first checking a durable, single-writer claim registry — so when a sibling foreman had already committed tick 39 (083750d), a second full verification and second board entry fired.
Fix (~/<project>/): a <project>.Store enforcing one committed board entry per tick number, plus a VerifyOnce entry point that implements the protocol:
Acquire checks the durable ledger first; only if the tick is absent does it claim an in-progress token via O_CREATE|O_EXCL (atomic across processes, no lock server). Live leases are never stolen; stale leases (crashed owner) are stolen only via compare-and-delete.// <project>.go
func (s *Store) Acquire(tick int, owner string, now time.Time) (bool, error) {
// 1. Already on the board? → ErrDuplicate, never run again.
if _, ok, err := s.GetEntry(tick); err != nil { return false, err } else if ok {
return false, ErrDuplicate
}
// 2. Atomic per-tick claim: only one process wins.
claimed, err := s.tryClaim(s.tokenPath(tick), Token{Tick: tick, Owner: owner, LeaseUntil: now.Add(s.lease)})
...
}
Commit serializes with flock and refuses a second entry: identical evidence → ErrDuplicate; differing evidence → ErrConflict (original preserved, escalated to humans, never overwritten). The ledger is written temp-file + fsync + atomic rename.want := HashEvidence(e.Commit, e.Evidence) // canonical digest, computed here
if cur, ok := entries[e.Tick]; ok {
if cur.EvidenceSHA == want && cur.Commit == e.Commit {
return fmt.Errorf("%w: ... identical — second entry refused", ErrDuplicate)
}
return fmt.Errorf("%w: ... refusing to overwrite", ErrConflict)
}
Reconcile, don't re-run-and-publish — the losing/sibling path independently re-runs the gates (build, vet, 8+8 tests, guard, govulncheck, hilo-stats, gitreins-count, CI-runs), hashes them canonically (sorted key=commit:result lines, so map order can't alter the digest), compares against the sibling's committed hash, and writes only an attestation sidecar — never a second board entry.
No duplicate workers — ClaimWorker registers worker IDs with O_EXCL (duplicate IDs refused) and additionally refuses to spawn any worker for an already-committed tick. Only the single acquiring process ever reaches the spawn step.
The <project> loop collapses to one call:
outcome, err := <project>.VerifyOnce(s, tick, owner, commit, gates, now)
// outcome: ran-and-committed | duplicate-confirmed | conflict-escalated
`go vet`, `gofmt`, and `go test -race` all pass; 10/10 tests green: | Test | Verifies | |---|---| | `TestIncident_AlreadyCommitted_SiblingWins` | Regression for the exact incident: `Acquire` → `ErrDuplicate`, second `Commit` refused, worker spawn refused, board has **exactly 1** entry owned by `foreman-a` | | `TestIncident_VerifyOnceReconciles` | My run re-runs gates, matches sibling claims exactly, writes 1 attestation, board untouched | | `TestRace_ManyCriersOneBoardEntry` | 8 racing <project>s on tick 39 → exactly 1 `ran`, 7 `duplicate-confirmed`, 1 board entry | | `TestConflict_DifferentEvidenceNeverOverwrites` | Differing evidence → `ErrConflict`, original entry intact | | `TestVerifyOnce_ConflictEscalated` | Conflict path through public API; board not clobbered | | `TestLeaseExpiry_StaleTokenStolen` | Live token not stealable; expired (crashed) owner's token stolen; still 1 entry | | `TestHashEvidence_Deterministic` | Digest order-independent; sensitive to gate result and commit sha | | `TestWorkers_NoDuplicateSpawns` | Duplicate worker ID refused; committed tick closed to new workers | | `TestDistinctTicks_Independent` | Different tick numbers proceed independently | | `TestCorruptLedger_ErrorsNotSilenced` | Torn board errors loudly, never silently double-fires | Example run (replays incident): `foreman-a → ran-and-committed`; `<project>-b → duplicate-confirmed`; `ci-recheck → duplicate-confirmed`; worker dup refused; **1** board entry for tick 39, sha `9e3f6693e0…`. Edge cases covered: same-tick races, genuine conflicts (escalate, never overwrite), crashed-owner leases, duplicate worker processes, corrupt/torn ledger, and distinct-tick independence.
{"model": "deepseek-v4-flash", "problem_class": "go-verification-double-fire", "result": "passed", "tests": 10}