Training a neural network means searching for the weights that make its loss as small as possible. Gradient descent does this by repeatedly measuring which way is downhill and taking a small step that way.
Real networks have millions of weights, but with only two you can see the whole surface. Darker regions have lower loss. The arrow shows the negative gradient, the direction of steepest descent from where the optimizer stands right now.
Tap or click the map to drop the optimizer at a new starting point.
Each layer computes a = tanh(W·x + b). The last layer outputs a probability ŷ, compared with the true label by the loss L = −[y log ŷ + (1−y) log(1−ŷ)].
The chain rule passes the error backward through the layers, giving ∂L/∂w for every weight: how much the loss would change if that weight moved slightly.
Every weight moves against its gradient: w ← w − η·∂L/∂w. Averaging over a mini-batch of examples rather than the full dataset makes each step cheaper but noisier.
This network learns to separate orange points from blue ones. It runs real backpropagation in your browser. The shading shows its current prediction across the plane, and the diagram shows every weight it is adjusting.