neural-network-backprop
The implementation is a complete feed-forward neural network library in pure NumPy:
File: ~/nn.py
Key components:
softmax uses numerically stable x - max(x) trick)½·mean((y−ŷ)²) gradient) and cross-entropy (handles both integer labels and one-hot, with log(y+ε) clipping)NeuralNetwork class — configurable layer_sizes, activations, loss; He/Xavier weight init; forward/backward/train_step with momentumtrain() helper — mini-batch SGD shuffling + momentum updates + optional validationBackpropagation is implemented as standard reverse-mode autodiff:
output → loss gradient → δ_output = dL/da * da/dz → δ_hidden = (δ_current·Wᵢ₊₁) ⊙ σ'(zᵢ)
dw = aᵢ₋₁ᵀ · δᵢ db = sum(δᵢ, axis=0)
For softmax+cross-entropy, the gradient simplifies to (y_pred − y_true) / n (combined derivative), avoiding the need for separate softmax Jacobian computation.
Numerical gradient check verified all weight and bias gradients to within 1×10⁻⁶ absolute error.
All tests pass with numerical gradient verification: ``` === Numerical Gradient Check (MSE + sigmoid) === Layer 0 weight max diff: 1.1e-10 ✓ PASS Layer 1 weight max diff: 2.3e-11 ✓ PASS Layer 0 bias max diff: 3.8e-11 ✓ PASS Layer 1 bias max diff: 1.2e-11 ✓ PASS === Numerical Gradient Check (CE + softmax) === Layer 0 weight max diff: 2.4e-10 ✓ PASS Layer 1 weight max diff: 3.0e-10 ✓ PASS Integer labels vs one-hot: identical ✓ PASS === Concentric Circles (clean non-linear) === Deep (2-32-16-2): train_acc=1.0000, val_acc=1.0000 ✓ PASS === Self-Test Edge Cases === Single-layer (logistic regression): runs without error ✓ MSE regression (2→16→1): runs, MSE=0.036 ✓ Sigmoid+Tanh+Softmax mixed activations: runs ✓ Full-batch GD (batch_size=N): runs ✓ Pure SGD (batch_size=1): runs ✓ No momentum (momentum=0): runs ✓ High momentum (0.99): runs ✓ Float-valued integer labels: handled ✓ ```
{"model": "python3", "problem_class": "neural-network-backprop", "result": "passed", "tests": 14}The implementation is a complete feed-forward neural network library in pure NumPy:
File: ~/nn.py
Key components:
softmax uses numerically stable x - max(x) trick)½·mean((y−ŷ)²) gradient) and cross-entropy (handles both integer labels and one-hot, with log(y+ε) clipping)NeuralNetwork class — configurable layer_sizes, activations, loss; He/Xavier weight init; forward/backward/train_step with momentumtrain() helper — mini-batch SGD shuffling + momentum updates + optional validationBackpropagation is implemented as standard reverse-mode autodiff:
output → loss gradient → δ_output = dL/da * da/dz → δ_hidden = (δ_current·Wᵢ₊₁) ⊙ σ'(zᵢ)
dw = aᵢ₋₁ᵀ · δᵢ db = sum(δᵢ, axis=0)
For softmax+cross-entropy, the gradient simplifies to (y_pred − y_true) / n (combined derivative), avoiding the need for separate softmax Jacobian computation.
Numerical gradient check verified all weight and bias gradients to within 1×10⁻⁶ absolute error.
All tests pass with numerical gradient verification: ``` === Numerical Gradient Check (MSE + sigmoid) === Layer 0 weight max diff: 1.1e-10 ✓ PASS Layer 1 weight max diff: 2.3e-11 ✓ PASS Layer 0 bias max diff: 3.8e-11 ✓ PASS Layer 1 bias max diff: 1.2e-11 ✓ PASS === Numerical Gradient Check (CE + softmax) === Layer 0 weight max diff: 2.4e-10 ✓ PASS Layer 1 weight max diff: 3.0e-10 ✓ PASS Integer labels vs one-hot: identical ✓ PASS === Concentric Circles (clean non-linear) === Deep (2-32-16-2): train_acc=1.0000, val_acc=1.0000 ✓ PASS === Self-Test Edge Cases === Single-layer (logistic regression): runs without error ✓ MSE regression (2→16→1): runs, MSE=0.036 ✓ Sigmoid+Tanh+Softmax mixed activations: runs ✓ Full-batch GD (batch_size=N): runs ✓ Pure SGD (batch_size=1): runs ✓ No momentum (momentum=0): runs ✓ High momentum (0.99): runs ✓ Float-valued integer labels: handled ✓ ```
{"model": "python3", "problem_class": "neural-network-backprop", "result": "passed", "tests": 14}