A tiny neural network
Wire up neurons and train them. Discover why one neuron can never learn XOR.
Wire up neurons and train them. Discover why one neuron can never learn XOR.
p = σ(0.75×0.50 + 0.87×0.50 − 0.31) = σ(0.50) = 0.623, fed by x₁ and x₂
Hover or tap the left plot to feed any point through the network. Hover or focus a neuron to see its sum.
| x₁ | x₂ | target | output p | Correct? |
|---|---|---|---|---|
| 0 | 0 | 0.424 | ||
| 1 | 0 | 0.637 | ||
| 0 | 0 | 0.608 | ||
| 1 | 1 | 0.788 |
Forward pass. A neuron multiplies each input by a weight, adds them up with a bias, and passes the total through an activation function. Here it is the sigmoid σ, which squashes any number into 0 to 1. In a network, the hidden neurons’ outputs become the inputs of the next neuron.
Loss. For each of the four corners the network outputs p, its belief that the answer is 1. Log loss charges −ln(p) when the answer is 1 and −ln(1 − p) when it is 0, so confident mistakes cost the most. We average over the four corners.
Training. Training is just gradient descent on the weights: for every weight, work out whether nudging it up or down lowers the loss, then nudge it a little that way (η sets how far). Backpropagation is the bookkeeping trick that works out all those nudges at once by passing the blame backwards from the output, layer by layer. One epoch = one nudge using all four corners.
Forward pass. A neuron multiplies each input by a weight, adds them up with a bias, and passes the total through an activation function. Here it is the sigmoid σ, which squashes any number into 0 to 1. In a network, the hidden neurons’ outputs become the inputs of the next neuron.
Loss. For each of the four corners the network outputs p, its belief that the answer is 1. Log loss charges −ln(p) when the answer is 1 and −ln(1 − p) when it is 0, so confident mistakes cost the most. We average over the four corners.
Training. Training is just gradient descent on the weights: for every weight, work out whether nudging it up or down lowers the loss, then nudge it a little that way (η sets how far). Backpropagation is the bookkeeping trick that works out all those nudges at once by passing the blame backwards from the output, layer by layer. One epoch = one nudge using all four corners.
Things to try