python-transformer-manual-backprop-gradcheck
Everything is verified. Summary of the delivered solution:
~/)| File | Purpose |
|---|---|
transformer.py |
Pure-NumPy decoder-only transformer (d_model=32, 4 heads, 8-dim heads, GELU MLP 32→128→32, LayerNorm ×3, learned token + positional embeddings) with hand-written backward passes for every op and a bias-corrected Adam |
verify.py |
Per-element central-difference gradient check on all 18 parameter tensors + 200-step copy-task training check |
verify_extra.py |
Proves masked attention weights/gradients are exactly zero above the diagonal; re-runs everything on seeds 0, 7, 12345 |
solution.md |
The self-contained markdown deliverable: title, root-cause analysis, full code, verification + actual output |
solution.md)max|analytical − numerical| = 3.35e-08 (tolerance 1e-6, 30× margin) across ~15,000 parameter elements including:dlogits = p⊙(dp − Σ dp·p), masked entries hard-zeroed, exact zero gradient above the diagonal)-inf before exp)np.add.at, GELU, residuals, output projectiontarget[t] = input[t-1]) with Adam, loss 3.831 → 0.0069 = 99.8% reduction in 200 steps (bar: 60%)pos_idx was (T,) while the gradient is (B,T,D); fixed by tiling indices to (B,T) so np.add.at broadcasts correctly.# Evidence - Problem class: python-transformer-manual-backprop-gradcheck - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-09-04T22:35:36.698Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a single-block decoder-only transformer (causal self-attention, LayerNorm, two-layer MLP, residual stream, learned token embeddings) in pure NumPy with hand-written backward passes for every operation \u2014 no autograd library is permitted. Every parameter tensor's gradient must match central-difference numerical gradients within 1e-6 on a seeded d_model=32, 4-head configuration, including the correct gradient through the causal mask and numerically stable softmax. The implementation must then train on a character-level copy task with Adam and reduce loss by at least 60% within 200 steps, proving the backward pass is not only correct but usable.", "environment": "python3", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "python-transformer-manual-backprop-gradcheck", "provider": "openrouter", "solved_at": "2026-09-04T22:35:36.698Z", "version": "3.11"}Everything is verified. Summary of the delivered solution:
~/)| File | Purpose |
|---|---|
transformer.py |
Pure-NumPy decoder-only transformer (d_model=32, 4 heads, 8-dim heads, GELU MLP 32→128→32, LayerNorm ×3, learned token + positional embeddings) with hand-written backward passes for every op and a bias-corrected Adam |
verify.py |
Per-element central-difference gradient check on all 18 parameter tensors + 200-step copy-task training check |
verify_extra.py |
Proves masked attention weights/gradients are exactly zero above the diagonal; re-runs everything on seeds 0, 7, 12345 |
solution.md |
The self-contained markdown deliverable: title, root-cause analysis, full code, verification + actual output |
solution.md)max|analytical − numerical| = 3.35e-08 (tolerance 1e-6, 30× margin) across ~15,000 parameter elements including:dlogits = p⊙(dp − Σ dp·p), masked entries hard-zeroed, exact zero gradient above the diagonal)-inf before exp)np.add.at, GELU, residuals, output projectiontarget[t] = input[t-1]) with Adam, loss 3.831 → 0.0069 = 99.8% reduction in 200 steps (bar: 60%)pos_idx was (T,) while the gradient is (B,T,D); fixed by tiling indices to (B,T) so np.add.at broadcasts correctly.# Evidence - Problem class: python-transformer-manual-backprop-gradcheck - Model: openrouter/deepseek/deepseek-v4-flash-0731 - Solved: 2026-09-04T22:35:36.698Z - Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a single-block decoder-only transformer (causal self-attention, LayerNorm, two-layer MLP, residual stream, learned token embeddings) in pure NumPy with hand-written backward passes for every operation \u2014 no autograd library is permitted. Every parameter tensor's gradient must match central-difference numerical gradients within 1e-6 on a seeded d_model=32, 4-head configuration, including the correct gradient through the causal mask and numerically stable softmax. The implementation must then train on a character-level copy task with Adam and reduce loss by at least 60% within 200 steps, proving the backward pass is not only correct but usable.", "environment": "python3", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "python-transformer-manual-backprop-gradcheck", "provider": "openrouter", "solved_at": "2026-09-04T22:35:36.698Z", "version": "3.11"}