◐ Off-By-One · answer catalog

python-numpy-transformer-gradcheck-train

2 answer(s)pythonpython3pythonpython3

python-numpy-transformer-gradcheck-train

📦 Source in repository (JSON)

Answer 1

Done. The solution is verified and delivered as ~/SOLUTION.md (with the exact code embedded, byte-identical to the runnable ~/solution.py).

Summary

Root causes found (4):

  1. Primary bug — dropped residual identity-path gradient. For h2 = h1 + z2, the backward pass kept only the ln_back term and omitted the dz2 = dh2 identity term, so dL/dh1 was under-computed. Signature: all MLP/head params passed gradcheck while every attention/embedding param failed with FD magnitudes ~2.3× larger and correlation ≈ 0.72.
  2. Transposed dWo einsum ("btd,bto->do" stored the gradient of Wo[o,d] at index [d,o]).
  3. Causal-mask softmax Jacobian — correct form dz = p⊙(dp − Σ p·dp) (masked columns automatically zero via p=0), 1/√d_head scaling applied exactly once, and forward/backward head-packing transposes kept inverse-consistent for dK/dV/dQ.
  4. Checker gotcha — the k-bias gradient is mathematically exactly zero (every softmax row sums Σ_s dscores[t,s] = 0), which breaks a pure relative-error metric (0 − 1e-19 → rel = 1.0). Fixed with the standard atol + rtol·max(|num|,|ana|) criterion.

Verification (actual runs):

Run with python3 solution.py — both assert gates enforce the requirements, and the backward pass includes the explicit causal-mask softmax Jacobian with correct per-head dK/dV/dQ bookkeeping.

Evidence & signatures

# Evidence
- Problem class: python-numpy-transformer-gradcheck-train
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-09-06T10:46:36.626Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a minimal decoder-only transformer (token+position embeddings, one causal multi-head self-attention block, pre-norm layernorm, MLP with GELU) in pure NumPy with hand-written backpropagation and no autograd library. Verify every parameter gradient against central finite differences (relative error < 1e-5), then train on a small char-level corpus for a fixed number of steps and show cross-entropy loss decreases past a stated floor. The backward pass must include the causal-mask softmax Jacobian and correct dK/dV/dQ bookkeeping across heads.", "environment": "python3", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "python-numpy-transformer-gradcheck-train", "provider": "openrouter", "solved_at": "2026-09-06T10:46:36.627Z", "version": "3.11"}

Answer 2

Done. The solution is verified and delivered as ~/SOLUTION.md (with the exact code embedded, byte-identical to the runnable ~/solution.py).

Summary

Root causes found (4):

  1. Primary bug — dropped residual identity-path gradient. For h2 = h1 + z2, the backward pass kept only the ln_back term and omitted the dz2 = dh2 identity term, so dL/dh1 was under-computed. Signature: all MLP/head params passed gradcheck while every attention/embedding param failed with FD magnitudes ~2.3× larger and correlation ≈ 0.72.
  2. Transposed dWo einsum ("btd,bto->do" stored the gradient of Wo[o,d] at index [d,o]).
  3. Causal-mask softmax Jacobian — correct form dz = p⊙(dp − Σ p·dp) (masked columns automatically zero via p=0), 1/√d_head scaling applied exactly once, and forward/backward head-packing transposes kept inverse-consistent for dK/dV/dQ.
  4. Checker gotcha — the k-bias gradient is mathematically exactly zero (every softmax row sums Σ_s dscores[t,s] = 0), which breaks a pure relative-error metric (0 − 1e-19 → rel = 1.0). Fixed with the standard atol + rtol·max(|num|,|ana|) criterion.

Verification (actual runs):

Run with python3 solution.py — both assert gates enforce the requirements, and the backward pass includes the explicit causal-mask softmax Jacobian with correct per-head dK/dV/dQ bookkeeping.

Evidence & signatures

# Evidence
- Problem class: python-numpy-transformer-gradcheck-train
- Model: openrouter/deepseek/deepseek-v4-flash-0731
- Solved: 2026-09-06T10:46:36.626Z
- Verification: solution produced by pi in sandbox; see signatures.json
{"description": "Implement a minimal decoder-only transformer (token+position embeddings, one causal multi-head self-attention block, pre-norm layernorm, MLP with GELU) in pure NumPy with hand-written backpropagation and no autograd library. Verify every parameter gradient against central finite differences (relative error < 1e-5), then train on a small char-level corpus for a fixed number of steps and show cross-entropy loss decreases past a stated floor. The backward pass must include the causal-mask softmax Jacobian and correct dK/dV/dQ bookkeeping across heads.", "environment": "python3", "language": "python", "model": "openrouter/deepseek/deepseek-v4-flash-0731", "problem_class": "python-numpy-transformer-gradcheck-train", "provider": "openrouter", "solved_at": "2026-09-06T10:46:36.627Z", "version": "3.11"}
Generated from the verified corpus · MIT licensedBack to the catalog