node-redis-cluster-moved-ask-reshard-race
The complete, self-contained solution is written to ~/redis-cluster-clients/SOLUTION.md (embedded code is byte-identical to the runnable .js files in that directory, and I verified it by extracting the fenced blocks into a clean dir and running them).
Root causes (each independent):
1. RESP has no request IDs — correct matching requires one strict FIFO per connection, one entry per frame written. Retargeting must create a new entry and settled-latch the request so a late frame can't double-deliver.
2. ASK is transient (slot still owned by source → don't update the map); MOVED is permanent (→ do update).
3. ASKING authorizes only the next command in the byte stream. Under pipelining it must be written atomically with its command (Buffer.concat([ASKING, cmd]) in one socket.write).
4. Background CLUSTER SLOTS reconciliation races MOVED — version the slot map so an older snapshot can't clobber a newer owner.
5. One per-request attempts budget + deadline guarantees termination (no "re-routed forever").
6. RESP3 > push frames must not consume a FIFO slot.
7. MOVED/ASK retries are exactly-once safe; socket-drop retries are ambiguous and need idempotency keys for INCR-style commands.
The fix (resp.js, client.js, mock-cluster.js, test.js):
- Connection.pending — the only match table; parser pops exactly one entry per non-push reply.
- sendAsking pushes +OK and command slots, then writes both frames in one buffer.
- onReply handles MOVED (update map + retry), ASK (retry with ASKING, map untouched), CLUSTERDOWN/TRYAGAIN (bounded backoff).
- applySlots(ranges, seq) rejects any slot write with seq < slot.ver; MOVED writes with ver = ++topoSeq.
- Request.settled + attempts/deadline make succeed/fail idempotent and bounded.
Verification (no Redis install needed — 3 mock RESP3 nodes on <ip-address>:17000-17002 really migrate a hot slot):
[1] Reply-to-caller matching under pipelining ............ PASS (300 GETs, 0 mismatches)
[2] No drop/duplicate across live migration .............. PASS (400 SETs, 0 dup, ASK+=172)
[3] INCR applied exactly once during a storm ............. PASS (final=300)
[4] Slot-map versioning rejects stale reconciliations .... PASS
[5] ASKING frame written atomically ...................... PASS (single write)
RESULT: PASS
Run it with:
cd ~/redis-cluster-clients && node test.js
The document is directly usable as-is: copy the four fenced files, node test.js.
# Evidence - Problem class: node-redis-cluster-moved-ask-reshard-race - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-14T22:25:18.719Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a RESP3 pipelining client for a Redis Cluster that survives a live slot migration: it must follow MOVED, retarget individual commands with ASKING during ASK redirects, and reconcile its slot map while hundreds of in-flight requests span the migration. Guarantee per-connection reply to caller matching so no response is delivered to the wrong request, and show that no command is dropped, duplicated, or re-routed forever when the cluster keeps moving slots during a reshard.", "environment": "node20", "language": "js", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "node-redis-cluster-moved-ask-reshard-race", "provider": "openrouter", "solved_at": "2026-09-14T22:25:18.719Z", "version": "20"}The complete, self-contained solution is written to ~/redis-cluster-clients/SOLUTION.md (embedded code is byte-identical to the runnable .js files in that directory, and I verified it by extracting the fenced blocks into a clean dir and running them).
Root causes (each independent):
1. RESP has no request IDs — correct matching requires one strict FIFO per connection, one entry per frame written. Retargeting must create a new entry and settled-latch the request so a late frame can't double-deliver.
2. ASK is transient (slot still owned by source → don't update the map); MOVED is permanent (→ do update).
3. ASKING authorizes only the next command in the byte stream. Under pipelining it must be written atomically with its command (Buffer.concat([ASKING, cmd]) in one socket.write).
4. Background CLUSTER SLOTS reconciliation races MOVED — version the slot map so an older snapshot can't clobber a newer owner.
5. One per-request attempts budget + deadline guarantees termination (no "re-routed forever").
6. RESP3 > push frames must not consume a FIFO slot.
7. MOVED/ASK retries are exactly-once safe; socket-drop retries are ambiguous and need idempotency keys for INCR-style commands.
The fix (resp.js, client.js, mock-cluster.js, test.js):
- Connection.pending — the only match table; parser pops exactly one entry per non-push reply.
- sendAsking pushes +OK and command slots, then writes both frames in one buffer.
- onReply handles MOVED (update map + retry), ASK (retry with ASKING, map untouched), CLUSTERDOWN/TRYAGAIN (bounded backoff).
- applySlots(ranges, seq) rejects any slot write with seq < slot.ver; MOVED writes with ver = ++topoSeq.
- Request.settled + attempts/deadline make succeed/fail idempotent and bounded.
Verification (no Redis install needed — 3 mock RESP3 nodes on <ip-address>:17000-17002 really migrate a hot slot):
[1] Reply-to-caller matching under pipelining ............ PASS (300 GETs, 0 mismatches)
[2] No drop/duplicate across live migration .............. PASS (400 SETs, 0 dup, ASK+=172)
[3] INCR applied exactly once during a storm ............. PASS (final=300)
[4] Slot-map versioning rejects stale reconciliations .... PASS
[5] ASKING frame written atomically ...................... PASS (single write)
RESULT: PASS
Run it with:
cd ~/redis-cluster-clients && node test.js
The document is directly usable as-is: copy the four fenced files, node test.js.
# Evidence - Problem class: node-redis-cluster-moved-ask-reshard-race - Model: openrouter/deepseek/deepseek-v4.1-flash - Solved: 2026-09-14T22:25:18.719Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a RESP3 pipelining client for a Redis Cluster that survives a live slot migration: it must follow MOVED, retarget individual commands with ASKING during ASK redirects, and reconcile its slot map while hundreds of in-flight requests span the migration. Guarantee per-connection reply to caller matching so no response is delivered to the wrong request, and show that no command is dropped, duplicated, or re-routed forever when the cluster keeps moving slots during a reshard.", "environment": "node20", "language": "js", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "node-redis-cluster-moved-ask-reshard-race", "provider": "openrouter", "solved_at": "2026-09-14T22:25:18.719Z", "version": "20"}