go-raft-snapshot-stale-term-install-divergence
Done. The solution is verified and written up in ~/SOLUTION.md, with the working code in ~/raft-snapshot-fix/.
Root cause — the naive InstallSnapshot handler violates the Raft §7 receiver rules in four ways, in the exact order protocol correctness depends on:
1. Applies the snapshot before checking term staleness — a partitioned old leader's snapshot wipes committed state (the term-regression divergence).
2. Never checks LastIncludedIndex against commitIndex — blindly applying un-commits entries the follower already acknowledged.
3. Drops the leader's term bump — a snapshot with term > currentTerm must advance the term even if the body is ignorable.
4. Mishandles log reconciliation — must keep a matching suffix or discard the whole log, then reset the state machine and advance commit/apply to LastIncludedIndex.
The fix — a corrected HandleInstallSnapshot with four ordered rules: reject snap.Term < currentTerm before touching anything; adopt the term bump; ignore (ack) LastIncludedIndex ≤ commitIndex; otherwise reconcile/log-subsumes/reset. commitIndex can only move forward, nothing committed is ever dropped.
Verification — a self-contained Go 1.26 module with 7 tests, all passing under go test -race -v (plus 10× stress of the concurrent race):
- 4 unit tests, one per rule (stale-term reject, below-commit ignore, subsumes-with-term-bump, matching-suffix retention)
- TestTermRegressionRace: deterministic end-to-end race — old leader L0 (term 1) commits 1..3, gets partitioned while self-appending uncommittable 4..5; F1 elected at term 2 commits entry 4; L0 reconnects with a stale snapshot. Correct handler rejects it; committed entry 4 and k=v4 survive; replication continues.
- TestTermRegressionRaceConcurrent: 900 stale/stale-term mutations racing 300 real leader appends under the race detector — commit index never regresses, log never diverges.
- TestBuggyHandlerRegressesState: the "before" evidence, demonstrating the naive handler really does regress the replica to k=vStale and wipe committed entry 4.
# Evidence - Problem class: go-raft-snapshot-stale-term-install-divergence - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-08-28T16:20:06.614Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A Raft follower receives an InstallSnapshot RPC from a leader whose term is lower than the follower current term, but a naive handler applies the snapshot before checking staleness, wiping committed state and diverging the log; it also blindly applies snapshots whose last-included index is below the followers own commit index. Implement the correct handler: reject stale-term snapshots, refuse snapshots whose last-included-index is below commitIndex, and reconcile a snapshot that subsumes the follower log while still honoring the leaders term bump. Verify against a term-regression race where a partitioned old leader sends a stale snapshot after a new leader already committed entries.", "environment": "go1.26", "language": "go", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "go-raft-snapshot-stale-term-install-divergence", "provider": "openrouter", "solved_at": "2026-08-28T16:20:06.614Z", "version": "1.26"}Done. The solution is verified and written up in ~/SOLUTION.md, with the working code in ~/raft-snapshot-fix/.
Root cause — the naive InstallSnapshot handler violates the Raft §7 receiver rules in four ways, in the exact order protocol correctness depends on:
1. Applies the snapshot before checking term staleness — a partitioned old leader's snapshot wipes committed state (the term-regression divergence).
2. Never checks LastIncludedIndex against commitIndex — blindly applying un-commits entries the follower already acknowledged.
3. Drops the leader's term bump — a snapshot with term > currentTerm must advance the term even if the body is ignorable.
4. Mishandles log reconciliation — must keep a matching suffix or discard the whole log, then reset the state machine and advance commit/apply to LastIncludedIndex.
The fix — a corrected HandleInstallSnapshot with four ordered rules: reject snap.Term < currentTerm before touching anything; adopt the term bump; ignore (ack) LastIncludedIndex ≤ commitIndex; otherwise reconcile/log-subsumes/reset. commitIndex can only move forward, nothing committed is ever dropped.
Verification — a self-contained Go 1.26 module with 7 tests, all passing under go test -race -v (plus 10× stress of the concurrent race):
- 4 unit tests, one per rule (stale-term reject, below-commit ignore, subsumes-with-term-bump, matching-suffix retention)
- TestTermRegressionRace: deterministic end-to-end race — old leader L0 (term 1) commits 1..3, gets partitioned while self-appending uncommittable 4..5; F1 elected at term 2 commits entry 4; L0 reconnects with a stale snapshot. Correct handler rejects it; committed entry 4 and k=v4 survive; replication continues.
- TestTermRegressionRaceConcurrent: 900 stale/stale-term mutations racing 300 real leader appends under the race detector — commit index never regresses, log never diverges.
- TestBuggyHandlerRegressesState: the "before" evidence, demonstrating the naive handler really does regress the replica to k=vStale and wipe committed entry 4.
# Evidence - Problem class: go-raft-snapshot-stale-term-install-divergence - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-08-28T16:20:06.614Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A Raft follower receives an InstallSnapshot RPC from a leader whose term is lower than the follower current term, but a naive handler applies the snapshot before checking staleness, wiping committed state and diverging the log; it also blindly applies snapshots whose last-included index is below the followers own commit index. Implement the correct handler: reject stale-term snapshots, refuse snapshots whose last-included-index is below commitIndex, and reconcile a snapshot that subsumes the follower log while still honoring the leaders term bump. Verify against a term-regression race where a partitioned old leader sends a stale snapshot after a new leader already committed entries.", "environment": "go1.26", "language": "go", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "go-raft-snapshot-stale-term-install-divergence", "provider": "openrouter", "solved_at": "2026-08-28T16:20:06.614Z", "version": "1.26"}