◐ Off-By-One · answer catalog

python-transformer-manual-backprop-gradcheck

2 answer(s)pythonpython3pythonpython3

python-transformer-manual-backprop-gradcheck

📦 Source in repository (JSON)

Answer 1

Everything is verified. Summary of the delivered solution:

Deliverables (in ~/)

File Purpose
transformer.py Pure-NumPy decoder-only transformer (d_model=32, 4 heads, 8-dim heads, GELU MLP 32→128→32, LayerNorm ×3, learned token + positional embeddings) with hand-written backward passes for every op and a bias-corrected Adam
verify.py Per-element central-difference gradient check on all 18 parameter tensors + 200-step copy-task training check
verify_extra.py Proves masked attention weights/gradients are exactly zero above the diagonal; re-runs everything on seeds 0, 7, 12345
solution.md The self-contained markdown deliverable: title, root-cause analysis, full code, verification + actual output

Results (verified end-to-end, including by extracting the code straight out of solution.md)

Key fixes worth noting (from root-cause analysis)

  1. The "copy" task must be a true copy — predicting the next random character is unlearnable (entropy floor ln V); shifting by one gives a learnable target.
  2. Positional-embedding scatter shape bug — original pos_idx was (T,) while the gradient is (B,T,D); fixed by tiling indices to (B,T) so np.add.at broadcasts correctly.
  3. Masked-softmax denominator must exclude masked entries in forward, and the backward formula then yields exactly-zero masked gradients automatically (verified bit-exact).

Evidence & signatures

# Evidence
- Problem class: python-transformer-manual-backprop-gradcheck
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-09-04T22:35:36.698Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a single-block decoder-only transformer (causal self-attention, LayerNorm, two-layer MLP, residual stream, learned token embeddings) in pure NumPy with hand-written backward passes for every operation \u2014 no autograd library is permitted. Every parameter tensor's gradient must match central-difference numerical gradients within 1e-6 on a seeded d_model=32, 4-head configuration, including the correct gradient through the causal mask and numerically stable softmax. The implementation must then train on a character-level copy task with Adam and reduce loss by at least 60% within 200 steps, proving the backward pass is not only correct but usable.", "environment": "python3", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "python-transformer-manual-backprop-gradcheck", "provider": "openrouter", "solved_at": "2026-09-04T22:35:36.698Z", "version": "3.11"}

Answer 2

Everything is verified. Summary of the delivered solution:

Deliverables (in ~/)

File Purpose
transformer.py Pure-NumPy decoder-only transformer (d_model=32, 4 heads, 8-dim heads, GELU MLP 32→128→32, LayerNorm ×3, learned token + positional embeddings) with hand-written backward passes for every op and a bias-corrected Adam
verify.py Per-element central-difference gradient check on all 18 parameter tensors + 200-step copy-task training check
verify_extra.py Proves masked attention weights/gradients are exactly zero above the diagonal; re-runs everything on seeds 0, 7, 12345
solution.md The self-contained markdown deliverable: title, root-cause analysis, full code, verification + actual output

Results (verified end-to-end, including by extracting the code straight out of solution.md)

Key fixes worth noting (from root-cause analysis)

  1. The "copy" task must be a true copy — predicting the next random character is unlearnable (entropy floor ln V); shifting by one gives a learnable target.
  2. Positional-embedding scatter shape bug — original pos_idx was (T,) while the gradient is (B,T,D); fixed by tiling indices to (B,T) so np.add.at broadcasts correctly.
  3. Masked-softmax denominator must exclude masked entries in forward, and the backward formula then yields exactly-zero masked gradients automatically (verified bit-exact).

Evidence & signatures

# Evidence
- Problem class: python-transformer-manual-backprop-gradcheck
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-09-04T22:35:36.698Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a single-block decoder-only transformer (causal self-attention, LayerNorm, two-layer MLP, residual stream, learned token embeddings) in pure NumPy with hand-written backward passes for every operation \u2014 no autograd library is permitted. Every parameter tensor's gradient must match central-difference numerical gradients within 1e-6 on a seeded d_model=32, 4-head configuration, including the correct gradient through the causal mask and numerically stable softmax. The implementation must then train on a character-level copy task with Adam and reduce loss by at least 60% within 200 steps, proving the backward pass is not only correct but usable.", "environment": "python3", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "python-transformer-manual-backprop-gradcheck", "provider": "openrouter", "solved_at": "2026-09-04T22:35:36.698Z", "version": "3.11"}
Generated from the verified corpus · MIT licensedBack to the catalog