Gradient Descent Lab (Copy)
Interactive · Optimization

Gradient Descent Lab

Training a neural network means searching for the weights that make its loss as small as possible. Gradient descent does this by repeatedly measuring which way is downhill and taking a small step that way.

θt+1 = θt − η · ∇L(θt)
Part 1 · Two weights

Walking down a loss landscape

Real networks have millions of weights, but with only two you can see the whole surface. Darker regions have lower loss. The arrow shows the negative gradient, the direction of steepest descent from where the optimizer stands right now.

θ₁ → · θ₂ ↑

Tap or click the map to drop the optimizer at a new starting point.

loss per step
step
0
θ
L(θ)
∇L
Δθ
–
  • Elongated bowl: push η above about 0.25 and watch plain descent bounce across the valley, then fly off.
  • Saddle point: plain descent crawls where the slope is nearly flat. Momentum and Adam roll off much sooner.
  • Several valleys: different starting points settle into different minima. Descent only finds a local low point.
  • Noisy gradients mimic mini-batches: each step sees a slightly wrong slope, so the path jitters.
The training loop

One step of training, in three moves

1 · Forward pass

Make predictions

Each layer computes a = tanh(W·x + b). The last layer outputs a probability ŷ, compared with the true label by the loss L = −[y log ŷ + (1−y) log(1−ŷ)].

2 · Backpropagation

Measure the slope

The chain rule passes the error backward through the layers, giving ∂L/∂w for every weight: how much the loss would change if that weight moved slightly.

3 · Update

Step downhill

Every weight moves against its gradient: w ← w − η·∂L/∂w. Averaging over a mini-batch of examples rather than the full dataset makes each step cheaper but noisier.

Part 2 · A real network

Training a small neural network

This network learns to separate orange points from blue ones. It runs real backpropagation in your browser. The shading shows its current prediction across the plane, and the diagram shows every weight it is adjusting.

epoch 0
class 0class 1shade strength = confidence
training loss per epoch
epoch
0
loss
accuracy
weights