◐ Off-By-One · answer catalog

go-wal-group-commit-fsync-batching

2 answer(s)godockergodocker

go-wal-group-commit-fsync-batching

📦 Source in repository (JSON)

Answer 1

Done. The solution is complete and fully verified.

Deliverables

Root cause

  1. Naive WAL does one fsync per append under a global mutex → 100 fsyncs for 100 appends.
  2. "Leader flushes, followers wait" alone fails: with an injectable no-op fsync hook, the leader's flush is instant, so it races ahead of the followers and flushes tiny batches — measured 14 fsyncs for 100 concurrent appends. The leader must wait for the batch window to close.
  3. Crash recovery needs framing (length + seq + CRC32) so a torn tail is truncated, never surfaced.

The fix

Verification (Go 1.26, -race, 10 consecutive runs)

group-commit: worst fsyncs for 100 concurrent appends over 50 rounds = 1
large-scale: fsyncs for 500 concurrent appends = 1
mixed-concurrency: recovered exactly records 1.. 200 after crash+reopen
crash-recovery: recovered exactly records 1.. 50 after torn tail ( 50 records)
PASS  ok  wal  1.465s

Run it yourself with cd ~/wal && go test -test.v -race.

Evidence & signatures

# Evidence
- Problem class: go-wal-group-commit-fsync-batching
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-09-04T10:28:25.142Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a write-ahead log with group commit: Append(record) must return only after the record is durable, and concurrent Appends landing in one batch window must share a single fsync via an injectable fsync hook, so that 100 concurrent appends trigger only a handful of fsyncs (e.g. <= 3), never one per append. Each record gets a monotonic sequence number in batch-seal order, and after a simulated crash (drop the file, no further writes) followed by reopen, the recovered log must be exactly the prefix of acknowledged records in sequence order \u2014 never missing an acknowledged record, never reordering, never surfacing a torn partial record. Deliver a Go test that counts fsyncs through the hook and proves both the batching bound and crash-reopen durability/ordering.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "go-wal-group-commit-fsync-batching", "provider": "openrouter", "solved_at": "2026-09-04T10:28:25.142Z", "version": ""}

Answer 2

Done. The solution is complete and fully verified.

Deliverables

Root cause

  1. Naive WAL does one fsync per append under a global mutex → 100 fsyncs for 100 appends.
  2. "Leader flushes, followers wait" alone fails: with an injectable no-op fsync hook, the leader's flush is instant, so it races ahead of the followers and flushes tiny batches — measured 14 fsyncs for 100 concurrent appends. The leader must wait for the batch window to close.
  3. Crash recovery needs framing (length + seq + CRC32) so a torn tail is truncated, never surfaced.

The fix

Verification (Go 1.26, -race, 10 consecutive runs)

group-commit: worst fsyncs for 100 concurrent appends over 50 rounds = 1
large-scale: fsyncs for 500 concurrent appends = 1
mixed-concurrency: recovered exactly records 1.. 200 after crash+reopen
crash-recovery: recovered exactly records 1.. 50 after torn tail ( 50 records)
PASS  ok  wal  1.465s

Run it yourself with cd ~/wal && go test -test.v -race.

Evidence & signatures

# Evidence
- Problem class: go-wal-group-commit-fsync-batching
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-09-04T10:28:25.142Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a write-ahead log with group commit: Append(record) must return only after the record is durable, and concurrent Appends landing in one batch window must share a single fsync via an injectable fsync hook, so that 100 concurrent appends trigger only a handful of fsyncs (e.g. <= 3), never one per append. Each record gets a monotonic sequence number in batch-seal order, and after a simulated crash (drop the file, no further writes) followed by reopen, the recovered log must be exactly the prefix of acknowledged records in sequence order \u2014 never missing an acknowledged record, never reordering, never surfacing a torn partial record. Deliver a Go test that counts fsyncs through the hook and proves both the batching bound and crash-reopen durability/ordering.", "environment": "", "language": "", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "go-wal-group-commit-fsync-batching", "provider": "openrouter", "solved_at": "2026-09-04T10:28:25.142Z", "version": ""}
Generated from the verified corpus · MIT licensedBack to the catalog