From One Neuron to Many

From one neuron to many

A single artificial neuron computes a weighted sum, w·x + b, which is a straight line: enough to pass through any two points and no more. Pass many of those lines through a nonlinear activation and add them up, and the network can bend into almost any shape the data asks for.

Act 1 · one neuron

Two points, one line

Drag the points. The neuron's weight and bias solve exactly for the line through them. Add a third point and watch a single line run out of freedom.

y = w·x + b
Free parameters
2
Error (mean squared)
0.000

Two parameters, two points: the line fits them exactly.

Act 2 · a network

Many neurons, one curve

Each hidden neuron draws its own line, bends it with an activation, and scales it. The output adds every bent line together. Tap the plot to add points, pick a shape, and add neurons until the curve can follow it.

Network outputEach neuron's contributionData
6
Error (mean squared)
–
Parameters / points
–

Tap a hidden neuron in the diagram to highlight its piece of the curve.

Why one neuron is a lineIts pre-activation w·x + b has two numbers to set, so it can satisfy two constraints. A third point not on that line leaves it with irreducible error.
Why the activation mattersSums of straight lines are still a straight line. The activation bends each one, a hinge for ReLU, a jump for the step, an S for sigmoid, and sums of bends can trace curves.
A note on McCulloch–PittsThe 1943 neuron was a threshold unit: it fired 1 if its weighted sum crossed a threshold, else 0. That is the step option. It is smoothed slightly here so gradient descent can train it.
Toward deep networksMore neurons add more bends. Stacking layers lets bends be built from other bends, which is how networks fit far richer distributions than one wide layer can do efficiently.