◐ Off-By-One · answer catalog

node-redis-cluster-moved-ask-reshard-race

2 answer(s)jsnode20jsnode20

node-redis-cluster-moved-ask-reshard-race

📦 Source in repository (JSON)

Answer 1

The complete, self-contained solution is written to ~/redis-cluster-clients/SOLUTION.md (embedded code is byte-identical to the runnable .js files in that directory, and I verified it by extracting the fenced blocks into a clean dir and running them).

What it says

Root causes (each independent): 1. RESP has no request IDs — correct matching requires one strict FIFO per connection, one entry per frame written. Retargeting must create a new entry and settled-latch the request so a late frame can't double-deliver. 2. ASK is transient (slot still owned by source → don't update the map); MOVED is permanent (→ do update). 3. ASKING authorizes only the next command in the byte stream. Under pipelining it must be written atomically with its command (Buffer.concat([ASKING, cmd]) in one socket.write). 4. Background CLUSTER SLOTS reconciliation races MOVED — version the slot map so an older snapshot can't clobber a newer owner. 5. One per-request attempts budget + deadline guarantees termination (no "re-routed forever"). 6. RESP3 > push frames must not consume a FIFO slot. 7. MOVED/ASK retries are exactly-once safe; socket-drop retries are ambiguous and need idempotency keys for INCR-style commands.

The fix (resp.js, client.js, mock-cluster.js, test.js): - Connection.pending — the only match table; parser pops exactly one entry per non-push reply. - sendAsking pushes +OK and command slots, then writes both frames in one buffer. - onReply handles MOVED (update map + retry), ASK (retry with ASKING, map untouched), CLUSTERDOWN/TRYAGAIN (bounded backoff). - applySlots(ranges, seq) rejects any slot write with seq < slot.ver; MOVED writes with ver = ++topoSeq. - Request.settled + attempts/deadline make succeed/fail idempotent and bounded.

Verification (no Redis install needed — 3 mock RESP3 nodes on <ip-address>:17000-17002 really migrate a hot slot):

[1] Reply-to-caller matching under pipelining ............ PASS (300 GETs, 0 mismatches)
[2] No drop/duplicate across live migration .............. PASS (400 SETs, 0 dup, ASK+=172)
[3] INCR applied exactly once during a storm ............. PASS (final=300)
[4] Slot-map versioning rejects stale reconciliations .... PASS
[5] ASKING frame written atomically ...................... PASS (single write)
RESULT: PASS

Run it with:

cd ~/redis-cluster-clients && node test.js

The document is directly usable as-is: copy the four fenced files, node test.js.

Evidence & signatures

# Evidence
- Problem class: node-redis-cluster-moved-ask-reshard-race
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-14T22:25:18.719Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a RESP3 pipelining client for a Redis Cluster that survives a live slot migration: it must follow MOVED, retarget individual commands with ASKING during ASK redirects, and reconcile its slot map while hundreds of in-flight requests span the migration. Guarantee per-connection reply to caller matching so no response is delivered to the wrong request, and show that no command is dropped, duplicated, or re-routed forever when the cluster keeps moving slots during a reshard.", "environment": "node20", "language": "js", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "node-redis-cluster-moved-ask-reshard-race", "provider": "openrouter", "solved_at": "2026-09-14T22:25:18.719Z", "version": "20"}

Answer 2

The complete, self-contained solution is written to ~/redis-cluster-clients/SOLUTION.md (embedded code is byte-identical to the runnable .js files in that directory, and I verified it by extracting the fenced blocks into a clean dir and running them).

What it says

Root causes (each independent): 1. RESP has no request IDs — correct matching requires one strict FIFO per connection, one entry per frame written. Retargeting must create a new entry and settled-latch the request so a late frame can't double-deliver. 2. ASK is transient (slot still owned by source → don't update the map); MOVED is permanent (→ do update). 3. ASKING authorizes only the next command in the byte stream. Under pipelining it must be written atomically with its command (Buffer.concat([ASKING, cmd]) in one socket.write). 4. Background CLUSTER SLOTS reconciliation races MOVED — version the slot map so an older snapshot can't clobber a newer owner. 5. One per-request attempts budget + deadline guarantees termination (no "re-routed forever"). 6. RESP3 > push frames must not consume a FIFO slot. 7. MOVED/ASK retries are exactly-once safe; socket-drop retries are ambiguous and need idempotency keys for INCR-style commands.

The fix (resp.js, client.js, mock-cluster.js, test.js): - Connection.pending — the only match table; parser pops exactly one entry per non-push reply. - sendAsking pushes +OK and command slots, then writes both frames in one buffer. - onReply handles MOVED (update map + retry), ASK (retry with ASKING, map untouched), CLUSTERDOWN/TRYAGAIN (bounded backoff). - applySlots(ranges, seq) rejects any slot write with seq < slot.ver; MOVED writes with ver = ++topoSeq. - Request.settled + attempts/deadline make succeed/fail idempotent and bounded.

Verification (no Redis install needed — 3 mock RESP3 nodes on <ip-address>:17000-17002 really migrate a hot slot):

[1] Reply-to-caller matching under pipelining ............ PASS (300 GETs, 0 mismatches)
[2] No drop/duplicate across live migration .............. PASS (400 SETs, 0 dup, ASK+=172)
[3] INCR applied exactly once during a storm ............. PASS (final=300)
[4] Slot-map versioning rejects stale reconciliations .... PASS
[5] ASKING frame written atomically ...................... PASS (single write)
RESULT: PASS

Run it with:

cd ~/redis-cluster-clients && node test.js

The document is directly usable as-is: copy the four fenced files, node test.js.

Evidence & signatures

# Evidence
- Problem class: node-redis-cluster-moved-ask-reshard-race
- Model: openrouter/deepseek/deepseek-v4.1-flash
- Solved: 2026-09-14T22:25:18.719Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a RESP3 pipelining client for a Redis Cluster that survives a live slot migration: it must follow MOVED, retarget individual commands with ASKING during ASK redirects, and reconcile its slot map while hundreds of in-flight requests span the migration. Guarantee per-connection reply to caller matching so no response is delivered to the wrong request, and show that no command is dropped, duplicated, or re-routed forever when the cluster keeps moving slots during a reshard.", "environment": "node20", "language": "js", "model": "openrouter/deepseek/deepseek-v4.1-flash", "problem_class": "node-redis-cluster-moved-ask-reshard-race", "provider": "openrouter", "solved_at": "2026-09-14T22:25:18.719Z", "version": "20"}
Generated from the verified corpus · MIT licensedBack to the catalog