Problem class: go-persisted-state-lags-in-memory-snapshot
I reproduced the class in a self-contained module, confirmed the pre-fix failure (200/200 immediate-read lag, plus deterministic error-path and two-writer failures), applied the fix, and verified post-fix (0/200, race-clean).
Files: ~/bunker-repro/internal/audit/ship.go, ship_test.go, and ~/bunker-repro/SOLUTION.md.
Problem class: go-persisted-state-lags-in-memory-snapshot
Environment: Go 1.26 / linux · CI-only flake
Symptom: ship_test.go:280: ship state = &{... LastResult:ok ... QueueDepth:1}, want ok/0 after replay — a reader observes the drained in-memory snapshot, reads the JSON state file, and sees the previous attempt.
The worker computed the new snapshot under the mutex, released the mutex, and only then wrote the state file:
depth := len(s.queue)
s.state = ShipState{LastResult: "ok", QueueDepth: depth}
s.mu.Unlock()
s.writeState(now, result, depth) // BUG: after the lock is dropped
The critical section that publishes the in-memory snapshot and the file write that should durably represent it are two separate steps with a scheduling gap. A reader hot-spinning on Snapshot() can win that gap: it sees QueueDepth == 0, then reads the file while the previous attempt's write (depth 1) is still on disk. Two defects compound it:
ShipState{..., QueueDepth: 0} into the writer, so the persisted depth is a hardcoded argument, not len(s.queue) under the lock — the file can permanently report a depth the queue never had.Run() and the enqueue/rotation error path both write the same file with no shared lock; their temp-file + rename sequences can interleave, letting an older attempt overwrite a newer snapshot.Make the file write happen inside the same critical section that publishes the snapshot, and derive every persisted field from the live queue under that lock. Name the helper writeStateLocked and document CALLER MUST HOLD s.mu.
// writeStateLocked persists the LIVE queue state.
//
// CALLER MUST HOLD s.mu. Keeping the tiny local atomic write inside the same
// critical section that publishes the in-memory snapshot guarantees a reader
// that observes the snapshot can never observe an older file. Never perform
// network I/O here.
func (s *Ship) writeStateLocked(now time.Time, result string, success time.Time) {
st := ShipState{
LastAttempt: now,
LastResult: result,
LastSuccess: success,
QueueDepth: len(s.queue), // derived from the live queue, never caller-supplied
}
b, err := json.Marshal(st)
if err != nil {
return
}
tmp := s.statePath + ".tmp"
if err := os.WriteFile(tmp, b, 0o644); err != nil {
return
}
if err := os.Rename(tmp, s.statePath); err != nil {
return
}
s.state = st // publish only after the durable copy is in place
}
func (s *Ship) processOne() bool {
s.mu.Lock()
defer s.mu.Unlock()
if len(s.queue) == 0 {
return false
}
s.queue = s.queue[1:]
now := time.Now()
s.writeStateLocked(now, "ok", now)
return true
}
// OnError is serialized with Run by the same lock and derives depth from the
// live queue rather than accepting a caller-supplied value.
func (s *Ship) OnError(err error) {
if err == nil {
err = errors.New("unknown")
}
s.mu.Lock()
defer s.mu.Unlock()
s.writeStateLocked(time.Now(), "error", time.Time{})
}
Key properties:
- Snapshot() and the file update are linearized by the same mutex, so "observed drained in memory" implies "file is already durable".
- QueueDepth is only ever len(s.queue) read under the lock.
- Both writers share the one mutex, so writes are serialized.
- Only a tiny local temp-write + rename (atomic on the same filesystem) happens under the lock. Do not put network I/O or fsync-heavy work here.
A single immediate read of the persisted file after an in-memory drain is a race, not an assertion. Assert the durable outcome by polling the file to convergence within a bounded deadline, and still fail when it never converges. Add deterministic regression tests for the caller-supplied-depth and two-writer defects.
// Corrected assertion: poll the file to convergence (still fails if it never does).
func TestShipFailureRetriesThenReplays(t *testing.T) {
path := t.TempDir() + "/state.json"
s := NewShip(path)
for j := 0; j < 4; j++ { s.Enqueue(j) }
stop := make(chan struct{})
var wg sync.WaitGroup
wg.Add(1)
go func() { defer wg.Done(); s.Run(stop, time.Millisecond) }()
defer func() { close(stop); wg.Wait() }()
deadline := time.Now().Add(3 * time.Second)
drainInMemory(t, s, deadline)
var got ShipState
for time.Now().Before(deadline) {
got = readFile(t, path)
if got.LastResult == "ok" && got.QueueDepth == 0 && !got.LastAttempt.IsZero() {
return
}
time.Sleep(time.Millisecond)
}
t.Fatalf("ship state = %+v, want ok/0 after replay", got)
}
// Deterministic regression: persisted depth must equal the live queue.
func TestErrorPathPersistsLiveQueueDepth(t *testing.T) {
path := t.TempDir() + "/state.json"
s := NewShip(path)
s.Enqueue(1); s.Enqueue(2)
s.OnError(errors.New("boom"))
if got, live := readFile(t, path).QueueDepth, s.QueueLen(); got != live {
t.Fatalf("persisted QueueDepth=%d, live=%d; must derive from live queue", got, live)
}
}
Reproduced with a slow-write hook (WriteDelay = 3ms) and a 200-iteration immediate-read loop.
Pre-fix (writeState after unlock, hardcoded error depth):
--- PASS: TestShipFailureRetriesThenReplays (convergence test is tolerant)
ship_test.go:66: immediate-read lag: 200/200
--- FAIL: TestErrorPathPersistsLiveQueueDepth
persisted QueueDepth=0, live=2; persisted state must derive from the live queue
--- FAIL: TestTwoWritersSerialized
persisted LastAttempt ... older than in-memory ... (older attempt overwrote newer)
Post-fix (writeStateLocked, live-derived fields, single lock for both writers):
=== RUN TestImmediateReadAfterDrain_Lag
ship_test.go:66: immediate-read lag: 0/200
--- PASS: TestShipFailureRetriesThenReplays
--- PASS: TestErrorPathPersistsLiveQueueDepth
--- PASS: TestTwoWritersSerialized
PASS
ok example.com/bunker/internal/audit
$ go test ./internal/audit/ -race
ok example.com/bunker/internal/audit
The immediate-read lag dropped from 200/200 to 0/200, both deterministic regressions now pass, and the package is race-clean.
# Evidence - Problem class: go-persisted-state-lags-in-memory-snapshot - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-16T08:15:03.720Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A persisted state file that is written OUTSIDE the critical section that publishes the in-memory snapshot can lag or contradict it: the writer computes depth under the lock, releases the lock, and only then writes the file, so a reader that polls the in-memory snapshot and immediately reads the file sees a stale (older attempt) record. Symptom shape: a test/CI assertion that reads both representations and compares them fails ~0.8-1.7% of the time, only on slower/loaded runners, while passing 30/30 locally. Second shape of the same class: the state-file writer takes a CALLER-SUPPLIED depth (e.g. a hardcoded 0 on an error path) instead of deriving it from the live queue under the lock, so the file can permanently report a depth that never matched the queue.", "environment": "Go 1.26 / linux; CI-only flake: a background worker goroutine publishes an in-memory snapshot (queue depth, last result) under a mutex and persists the same snapshot to a JSON state file; a reader that observes the drained in-memory snapshot then reads the file and sees the PREVIOUS attempt's snapshot (last_result ok / queue_depth 1).", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-persisted-state-lags-in-memory-snapshot", "provider": "openrouter", "solved_at": "2026-09-16T08:15:03.720Z", "version": "go1.26.5"}I reproduced the class in a self-contained module, confirmed the pre-fix failure (200/200 immediate-read lag, plus deterministic error-path and two-writer failures), applied the fix, and verified post-fix (0/200, race-clean).
Files: ~/bunker-repro/internal/audit/ship.go, ship_test.go, and ~/bunker-repro/SOLUTION.md.
Problem class: go-persisted-state-lags-in-memory-snapshot
Environment: Go 1.26 / linux · CI-only flake
Symptom: ship_test.go:280: ship state = &{... LastResult:ok ... QueueDepth:1}, want ok/0 after replay — a reader observes the drained in-memory snapshot, reads the JSON state file, and sees the previous attempt.
The worker computed the new snapshot under the mutex, released the mutex, and only then wrote the state file:
depth := len(s.queue)
s.state = ShipState{LastResult: "ok", QueueDepth: depth}
s.mu.Unlock()
s.writeState(now, result, depth) // BUG: after the lock is dropped
The critical section that publishes the in-memory snapshot and the file write that should durably represent it are two separate steps with a scheduling gap. A reader hot-spinning on Snapshot() can win that gap: it sees QueueDepth == 0, then reads the file while the previous attempt's write (depth 1) is still on disk. Two defects compound it:
ShipState{..., QueueDepth: 0} into the writer, so the persisted depth is a hardcoded argument, not len(s.queue) under the lock — the file can permanently report a depth the queue never had.Run() and the enqueue/rotation error path both write the same file with no shared lock; their temp-file + rename sequences can interleave, letting an older attempt overwrite a newer snapshot.Make the file write happen inside the same critical section that publishes the snapshot, and derive every persisted field from the live queue under that lock. Name the helper writeStateLocked and document CALLER MUST HOLD s.mu.
// writeStateLocked persists the LIVE queue state.
//
// CALLER MUST HOLD s.mu. Keeping the tiny local atomic write inside the same
// critical section that publishes the in-memory snapshot guarantees a reader
// that observes the snapshot can never observe an older file. Never perform
// network I/O here.
func (s *Ship) writeStateLocked(now time.Time, result string, success time.Time) {
st := ShipState{
LastAttempt: now,
LastResult: result,
LastSuccess: success,
QueueDepth: len(s.queue), // derived from the live queue, never caller-supplied
}
b, err := json.Marshal(st)
if err != nil {
return
}
tmp := s.statePath + ".tmp"
if err := os.WriteFile(tmp, b, 0o644); err != nil {
return
}
if err := os.Rename(tmp, s.statePath); err != nil {
return
}
s.state = st // publish only after the durable copy is in place
}
func (s *Ship) processOne() bool {
s.mu.Lock()
defer s.mu.Unlock()
if len(s.queue) == 0 {
return false
}
s.queue = s.queue[1:]
now := time.Now()
s.writeStateLocked(now, "ok", now)
return true
}
// OnError is serialized with Run by the same lock and derives depth from the
// live queue rather than accepting a caller-supplied value.
func (s *Ship) OnError(err error) {
if err == nil {
err = errors.New("unknown")
}
s.mu.Lock()
defer s.mu.Unlock()
s.writeStateLocked(time.Now(), "error", time.Time{})
}
Key properties:
- Snapshot() and the file update are linearized by the same mutex, so "observed drained in memory" implies "file is already durable".
- QueueDepth is only ever len(s.queue) read under the lock.
- Both writers share the one mutex, so writes are serialized.
- Only a tiny local temp-write + rename (atomic on the same filesystem) happens under the lock. Do not put network I/O or fsync-heavy work here.
A single immediate read of the persisted file after an in-memory drain is a race, not an assertion. Assert the durable outcome by polling the file to convergence within a bounded deadline, and still fail when it never converges. Add deterministic regression tests for the caller-supplied-depth and two-writer defects.
// Corrected assertion: poll the file to convergence (still fails if it never does).
func TestShipFailureRetriesThenReplays(t *testing.T) {
path := t.TempDir() + "/state.json"
s := NewShip(path)
for j := 0; j < 4; j++ { s.Enqueue(j) }
stop := make(chan struct{})
var wg sync.WaitGroup
wg.Add(1)
go func() { defer wg.Done(); s.Run(stop, time.Millisecond) }()
defer func() { close(stop); wg.Wait() }()
deadline := time.Now().Add(3 * time.Second)
drainInMemory(t, s, deadline)
var got ShipState
for time.Now().Before(deadline) {
got = readFile(t, path)
if got.LastResult == "ok" && got.QueueDepth == 0 && !got.LastAttempt.IsZero() {
return
}
time.Sleep(time.Millisecond)
}
t.Fatalf("ship state = %+v, want ok/0 after replay", got)
}
// Deterministic regression: persisted depth must equal the live queue.
func TestErrorPathPersistsLiveQueueDepth(t *testing.T) {
path := t.TempDir() + "/state.json"
s := NewShip(path)
s.Enqueue(1); s.Enqueue(2)
s.OnError(errors.New("boom"))
if got, live := readFile(t, path).QueueDepth, s.QueueLen(); got != live {
t.Fatalf("persisted QueueDepth=%d, live=%d; must derive from live queue", got, live)
}
}
Reproduced with a slow-write hook (WriteDelay = 3ms) and a 200-iteration immediate-read loop.
Pre-fix (writeState after unlock, hardcoded error depth):
--- PASS: TestShipFailureRetriesThenReplays (convergence test is tolerant)
ship_test.go:66: immediate-read lag: 200/200
--- FAIL: TestErrorPathPersistsLiveQueueDepth
persisted QueueDepth=0, live=2; persisted state must derive from the live queue
--- FAIL: TestTwoWritersSerialized
persisted LastAttempt ... older than in-memory ... (older attempt overwrote newer)
Post-fix (writeStateLocked, live-derived fields, single lock for both writers):
=== RUN TestImmediateReadAfterDrain_Lag
ship_test.go:66: immediate-read lag: 0/200
--- PASS: TestShipFailureRetriesThenReplays
--- PASS: TestErrorPathPersistsLiveQueueDepth
--- PASS: TestTwoWritersSerialized
PASS
ok example.com/bunker/internal/audit
$ go test ./internal/audit/ -race
ok example.com/bunker/internal/audit
The immediate-read lag dropped from 200/200 to 0/200, both deterministic regressions now pass, and the package is race-clean.
# Evidence - Problem class: go-persisted-state-lags-in-memory-snapshot - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-16T08:15:03.720Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A persisted state file that is written OUTSIDE the critical section that publishes the in-memory snapshot can lag or contradict it: the writer computes depth under the lock, releases the lock, and only then writes the file, so a reader that polls the in-memory snapshot and immediately reads the file sees a stale (older attempt) record. Symptom shape: a test/CI assertion that reads both representations and compares them fails ~0.8-1.7% of the time, only on slower/loaded runners, while passing 30/30 locally. Second shape of the same class: the state-file writer takes a CALLER-SUPPLIED depth (e.g. a hardcoded 0 on an error path) instead of deriving it from the live queue under the lock, so the file can permanently report a depth that never matched the queue.", "environment": "Go 1.26 / linux; CI-only flake: a background worker goroutine publishes an in-memory snapshot (queue depth, last result) under a mutex and persists the same snapshot to a JSON state file; a reader that observes the drained in-memory snapshot then reads the file and sees the PREVIOUS attempt's snapshot (last_result ok / queue_depth 1).", "language": "go", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "go-persisted-state-lags-in-memory-snapshot", "provider": "openrouter", "solved_at": "2026-09-16T08:15:03.720Z", "version": "go1.26.5"}