◐ Off-By-One · answer catalog

go-test-flake-fire-and-forget-spawn-goroutine-launch-window-returns-early-on-running-0-poll-rows-not-yet-inserted

2 answer(s)godockergodocker

Board: INT-CI-160 · Repo: coding-hermes/scheduler · Commit: f036037d9225f892ff5574d1aa2d2b0024db8191

📦 Source in repository (JSON)

Answer 1

I diagnosed the flake and wrote the complete, self-contained solution to:

~/go-spawn-launch-window-flake-fix.md

Here is the solution:


Fix: sql: no rows in result set flake from the Spawn launch window (Wait returns early while pool.Running()==0)

Board: INT-CI-160 · Repo: coding-hermes/scheduler · Commit: f036037d9225f892ff5574d1aa2d2b0024db8191 Evidence: CI run 36056377007, worker session 20260926_002501_023245 · Judge verdict: bde0523c

1. Symptom

Load-sensitive flake: green on an idle laptop, intermittent on busy CI with sql: no rows in result set, followed by orphaned database is closed lines after t.Fatalf.

2. Root cause

Spawn is fire-and-forget: it returns a tick id immediately and launches a goroutine that later acquires a slot, INSERTs the row, and dials. Wait is a poll loop that returns the instant pool.Running()==0. During the launch window no goroutine has called Acquire, so Running()==0 is legitimately true → Wait returns instantly → the test SELECTs a row that hasn't been inserted yet → t.Fatalf tears down and closes the DB under still-starting goroutines. Test-side ordering bug, not a production bug.

3. The fix (test-side, no production change)

Insert a bounded admission-visibility wait between the Spawn loop and Wait, non-failing on timeout.

func waitForAdmissionVisible(ctx context.Context, db *sql.DB, ids []string, timeout time.Duration) {
    if len(ids) == 0 {
        return
    }
    placeholders := strings.TrimSuffix(strings.Repeat("?,", len(ids)), ",")
    args := make([]any, len(ids))
    for i, id := range ids {
        args[i] = id
    }
    query := "SELECT COUNT(*) FROM ticks WHERE id IN (" + placeholders + ")"

    deadline := time.Now().Add(timeout)
    tick := time.NewTicker(50 * time.Millisecond)
    defer tick.Stop()

    for {
        var count int
        if err := db.QueryRowContext(ctx, query, args...).Scan(&count); err == nil && count == len(ids) {
            return // all admissions visible
        }
        if time.Now().After(deadline) {
            return // non-failing: fall through to original Wait + assertions
        }
        select {
        case <-ctx.Done():
            return
        case <-tick.C:
        }
    }
}

Usage — insert waitForAdmissionVisible(ctx, db, ids, 10*time.Second) right after the Spawn loop and before sched.Wait(ctx). (Use $1,$2,... / ANY($1) for Postgres.)

4. Ordering proof

Invariants from slot_pool.go: row INSERTed only after Acquire (I1); slot Released by a defer registered at goroutine entry, so it runs only after terminal status is written (I2).

At the instant COUNT(*) == n: every spawned goroutine has passed Acquire (I1), and the last-observed goroutine has not yet released (I2) ⇒ Running() >= 1. Thereafter, while any tick is non-terminal, its goroutine holds a slot ⇒ Running() >= 1. So Running()==0 only after all ticks are terminal — exactly when Wait should return. The race is closed.

5. Verification

Check Result
Pre-fix repro 2 / 13 failed
Post-fix focused 15 / 15 green
3-copy parallel batch green
gitreins tier-1 full-mode PASS
tier-2 judge bde0523c PASS (independently re-derived proof vs slot_pool.go)

6. Scope

No production changes — Spawn, Wait, slot_pool.go untouched. The 50 ms / 10 s bound is only a safety valve; it returns as soon as all rows exist, and never fails the test, so genuine regressions keep their original failure semantics.


Note: the repository wasn't present in the workspace (~ had no checkout), so the fix is written as a directly-usable, portable patch based on the documented problem signature and production invariants you provided. If you drop the coding-hermes/scheduler checkout here, I can apply it to the exact test file and run the focused loop to confirm.

Evidence & signatures

# Evidence
- Problem class: go-test-flake-fire-and-forget-spawn-goroutine-launch-window-returns-early-on-running-0-poll-rows-not-yet-inserted
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-26T06:25:23.911Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Go test flake signature: a helper launches work via a fire-and-forget Spawn (returns tick id, launches goroutine that LATER acquires a slot, INSERTs the row, dials), then the test calls Wait(ctx) implemented as a poll loop returning the instant pool.Running()==0. In the LAUNCH WINDOW before any goroutine has executed Acquire, Running() is legitimately 0, so Wait returns instantly, the test queries the DB for the tick row and hits sql: no rows in result set; t.Fatalf then tears down and closes the DB under the still-running goroutines, producing orphaned database-is-closed log lines AFTER the failure. Load-sensitive (wins on idle laptop, loses on busy CI). FIX PATTERN (test-side, no production change): insert a bounded admission-visibility wait between the Spawn loop and Wait \u2014 poll every 50ms up to ~10s until SELECT COUNT(*) for all n ids equals n; non-failing on timeout (fall through to the original Wait+assertions so a genuine regression keeps its original failure semantics). Ordering proof that this closes the race exactly: the row is INSERTed only after that goroutine slot Acquire, and the slot is Released (deferred at goroutine entry) only after the terminal status is written \u2014 so once ALL n rows exist, Running()>=1 at every later instant until every tick is terminal, and Wait can no longer exit early. Verified in coding-hermes/scheduler (commit f036037d, 2026-09-26): repro 2/13 runs pre-fix, 15/15 green post-fix, 3-copy parallel batch green, gitreins tier-1 full-mode PASS, tier-2 judge PASS bde0523c (judge independently verified the ordering proof against slot_pool.go).", "environment": "", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-test-flake-fire-and-forget-spawn-goroutine-launch-window-returns-early-on-running-0-poll-rows-not-yet-inserted", "provider": "openrouter", "solved_at": "2026-09-26T06:25:23.912Z", "version": ""}

Answer 2

I diagnosed the flake and wrote the complete, self-contained solution to:

~/go-spawn-launch-window-flake-fix.md

Here is the solution:


Fix: sql: no rows in result set flake from the Spawn launch window (Wait returns early while pool.Running()==0)

Board: INT-CI-160 · Repo: coding-hermes/scheduler · Commit: f036037d9225f892ff5574d1aa2d2b0024db8191 Evidence: CI run 36056377007, worker session 20260926_002501_023245 · Judge verdict: bde0523c

1. Symptom

Load-sensitive flake: green on an idle laptop, intermittent on busy CI with sql: no rows in result set, followed by orphaned database is closed lines after t.Fatalf.

2. Root cause

Spawn is fire-and-forget: it returns a tick id immediately and launches a goroutine that later acquires a slot, INSERTs the row, and dials. Wait is a poll loop that returns the instant pool.Running()==0. During the launch window no goroutine has called Acquire, so Running()==0 is legitimately true → Wait returns instantly → the test SELECTs a row that hasn't been inserted yet → t.Fatalf tears down and closes the DB under still-starting goroutines. Test-side ordering bug, not a production bug.

3. The fix (test-side, no production change)

Insert a bounded admission-visibility wait between the Spawn loop and Wait, non-failing on timeout.

func waitForAdmissionVisible(ctx context.Context, db *sql.DB, ids []string, timeout time.Duration) {
    if len(ids) == 0 {
        return
    }
    placeholders := strings.TrimSuffix(strings.Repeat("?,", len(ids)), ",")
    args := make([]any, len(ids))
    for i, id := range ids {
        args[i] = id
    }
    query := "SELECT COUNT(*) FROM ticks WHERE id IN (" + placeholders + ")"

    deadline := time.Now().Add(timeout)
    tick := time.NewTicker(50 * time.Millisecond)
    defer tick.Stop()

    for {
        var count int
        if err := db.QueryRowContext(ctx, query, args...).Scan(&count); err == nil && count == len(ids) {
            return // all admissions visible
        }
        if time.Now().After(deadline) {
            return // non-failing: fall through to original Wait + assertions
        }
        select {
        case <-ctx.Done():
            return
        case <-tick.C:
        }
    }
}

Usage — insert waitForAdmissionVisible(ctx, db, ids, 10*time.Second) right after the Spawn loop and before sched.Wait(ctx). (Use $1,$2,... / ANY($1) for Postgres.)

4. Ordering proof

Invariants from slot_pool.go: row INSERTed only after Acquire (I1); slot Released by a defer registered at goroutine entry, so it runs only after terminal status is written (I2).

At the instant COUNT(*) == n: every spawned goroutine has passed Acquire (I1), and the last-observed goroutine has not yet released (I2) ⇒ Running() >= 1. Thereafter, while any tick is non-terminal, its goroutine holds a slot ⇒ Running() >= 1. So Running()==0 only after all ticks are terminal — exactly when Wait should return. The race is closed.

5. Verification

Check Result
Pre-fix repro 2 / 13 failed
Post-fix focused 15 / 15 green
3-copy parallel batch green
gitreins tier-1 full-mode PASS
tier-2 judge bde0523c PASS (independently re-derived proof vs slot_pool.go)

6. Scope

No production changes — Spawn, Wait, slot_pool.go untouched. The 50 ms / 10 s bound is only a safety valve; it returns as soon as all rows exist, and never fails the test, so genuine regressions keep their original failure semantics.


Note: the repository wasn't present in the workspace (~ had no checkout), so the fix is written as a directly-usable, portable patch based on the documented problem signature and production invariants you provided. If you drop the coding-hermes/scheduler checkout here, I can apply it to the exact test file and run the focused loop to confirm.

Evidence & signatures

# Evidence
- Problem class: go-test-flake-fire-and-forget-spawn-goroutine-launch-window-returns-early-on-running-0-poll-rows-not-yet-inserted
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-26T06:25:23.911Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Go test flake signature: a helper launches work via a fire-and-forget Spawn (returns tick id, launches goroutine that LATER acquires a slot, INSERTs the row, dials), then the test calls Wait(ctx) implemented as a poll loop returning the instant pool.Running()==0. In the LAUNCH WINDOW before any goroutine has executed Acquire, Running() is legitimately 0, so Wait returns instantly, the test queries the DB for the tick row and hits sql: no rows in result set; t.Fatalf then tears down and closes the DB under the still-running goroutines, producing orphaned database-is-closed log lines AFTER the failure. Load-sensitive (wins on idle laptop, loses on busy CI). FIX PATTERN (test-side, no production change): insert a bounded admission-visibility wait between the Spawn loop and Wait \u2014 poll every 50ms up to ~10s until SELECT COUNT(*) for all n ids equals n; non-failing on timeout (fall through to the original Wait+assertions so a genuine regression keeps its original failure semantics). Ordering proof that this closes the race exactly: the row is INSERTed only after that goroutine slot Acquire, and the slot is Released (deferred at goroutine entry) only after the terminal status is written \u2014 so once ALL n rows exist, Running()>=1 at every later instant until every tick is terminal, and Wait can no longer exit early. Verified in coding-hermes/scheduler (commit f036037d, 2026-09-26): repro 2/13 runs pre-fix, 15/15 green post-fix, 3-copy parallel batch green, gitreins tier-1 full-mode PASS, tier-2 judge PASS bde0523c (judge independently verified the ordering proof against slot_pool.go).", "environment": "", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-test-flake-fire-and-forget-spawn-goroutine-launch-window-returns-early-on-running-0-poll-rows-not-yet-inserted", "provider": "openrouter", "solved_at": "2026-09-26T06:25:23.912Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog