python-epoch-shuffle-seed-train-val-leak
Both the reproduction and regression test verify cleanly.
Delivered two files in ~:
SOLUTION.md — the self-contained markdown solution with title, root-cause analysis, exact fixed code, and verification.
leak_fix.py — runnable reproduction + fixed pipeline + regression test.
The pipeline reseeded one RNG from the global step and derived both training and validation from the same whole-set permutation each epoch. Two compounding bugs: (1) the seed advances every step, re-bucketing rows continuously, and (2) train/val come from the same correlated shuffle. Result: over 100 steps on 1000 rows, 1000/1000 rows get used as both training data and validation labels — eval metrics measure memorization.
Commit a fixed, disjoint train_pool/val_pool partition once, then use two independent, separately-seeded shuffle streams per epoch:
* train → shuffle only train_pool, seeded by (train_seed, epoch), fanned into per-worker shards;
* validation → shuffle only val_pool, seeded by (val_seed, epoch).
Because each stream never touches the other pool, train and validation are disjoint by construction for any epoch count and world size — no statistical overlap possible.
BUGGY: training rows seen 1000 | validation rows seen 1000 | OVERLAP 1000
FIXED: 50 epochs x 4 workers => train rows 1600, val rows 400, intersection 0 -> PASSED
The regression test (test_zero_intersection) asserts zero set intersection across 50 epochs × 4 workers, plus index/universe sanity checks. Reproduced on Python 3.14; compatible with the required 3.11.
# Evidence - Problem class: python-epoch-shuffle-seed-train-val-leak - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-08-24T22:17:38.471Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A distributed training pipeline reshuffles the whole dataset each epoch with an RNG seeded from the global training step and carves the validation split out of that same shuffle, so validation rows are exactly the rows just trained on and eval metrics are inflated by memorization instead of measuring generalization. Reproduce the overlap in pure Python with seeded shuffles, then fix the pipeline so train and validation draws come from independent, separately-seeded shuffle streams that stay disjoint across any epoch count and world size. Add a regression test asserting zero set intersection between train and validation rows across 50 epochs and 4 simulated workers.", "environment": "python3", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "python-epoch-shuffle-seed-train-val-leak", "provider": "openrouter", "solved_at": "2026-08-24T22:17:38.471Z", "version": "3.11"}Both the reproduction and regression test verify cleanly.
Delivered two files in ~:
SOLUTION.md — the self-contained markdown solution with title, root-cause analysis, exact fixed code, and verification.
leak_fix.py — runnable reproduction + fixed pipeline + regression test.
The pipeline reseeded one RNG from the global step and derived both training and validation from the same whole-set permutation each epoch. Two compounding bugs: (1) the seed advances every step, re-bucketing rows continuously, and (2) train/val come from the same correlated shuffle. Result: over 100 steps on 1000 rows, 1000/1000 rows get used as both training data and validation labels — eval metrics measure memorization.
Commit a fixed, disjoint train_pool/val_pool partition once, then use two independent, separately-seeded shuffle streams per epoch:
* train → shuffle only train_pool, seeded by (train_seed, epoch), fanned into per-worker shards;
* validation → shuffle only val_pool, seeded by (val_seed, epoch).
Because each stream never touches the other pool, train and validation are disjoint by construction for any epoch count and world size — no statistical overlap possible.
BUGGY: training rows seen 1000 | validation rows seen 1000 | OVERLAP 1000
FIXED: 50 epochs x 4 workers => train rows 1600, val rows 400, intersection 0 -> PASSED
The regression test (test_zero_intersection) asserts zero set intersection across 50 epochs × 4 workers, plus index/universe sanity checks. Reproduced on Python 3.14; compatible with the required 3.11.
# Evidence - Problem class: python-epoch-shuffle-seed-train-val-leak - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-08-24T22:17:38.471Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "A distributed training pipeline reshuffles the whole dataset each epoch with an RNG seeded from the global training step and carves the validation split out of that same shuffle, so validation rows are exactly the rows just trained on and eval metrics are inflated by memorization instead of measuring generalization. Reproduce the overlap in pure Python with seeded shuffles, then fix the pipeline so train and validation draws come from independent, separately-seeded shuffle streams that stay disjoint across any epoch count and world size. Add a regression test asserting zero set intersection between train and validation rows across 50 epochs and 4 simulated workers.", "environment": "python3", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "python-epoch-shuffle-seed-train-val-leak", "provider": "openrouter", "solved_at": "2026-08-24T22:17:38.471Z", "version": "3.11"}