In part one we said a neural network learns its own rules from examples. This post shows how. We will train a network so small you can follow every number with a calculator, and every idea in it, weights, loss, gradient, learning rate, is exactly the idea used to train a model with a trillion parameters. Only the size changes.
A network with one weight
Suppose we want a model that converts hours worked into pay, and we have three examples from a payslip:
| Hours (input) | Pay (correct output) |
|---|---|
| 1 | 20 |
| 2 | 40 |
| 3 | 60 |
You can see the answer is pay = 20 times hours. The network cannot see that. All it has is a single number called a weight, which we will call w, and a rule: output = w times input. Its job is to find a value of w that makes the outputs match the examples. That is the whole game, for this network and for every network: adjust the weights until the outputs match the data.
Start with a random guess. Say w = 5.
Step 1: measure how wrong we are
With w = 5, the network predicts 5, 10 and 15 for our three inputs. The correct answers were 20, 40 and 60. We need a single number that says how bad this is. That number is the loss, and the function that computes it is the loss function.
The most common loss for numbers is mean squared error: take each error, square it so negatives do not cancel positives, and average.
errors: 20-5=15, 40-10=30, 60-15=45
squared: 225, 900, 2025
mean: (225 + 900 + 2025) / 3 = 1050
Loss is 1050. A perfect model would have loss 0. The whole of training is: change w to make this number smaller.
Step 2: work out which direction to move
We could try w = 4 and w = 6, compute the loss for each, and keep whichever is lower. With one weight that works. With a billion weights it would take forever. Instead we use calculus to ask: if I nudge w up by a tiny amount, does the loss go up or down, and how steeply? That steepness is the gradient.
You do not need to do the calculus to follow along. For our loss the gradient works out as a simple formula: minus two times the average of (input times error).
input × error: 1×15=15, 2×30=60, 3×45=135
average: (15 + 60 + 135) / 3 = 70
gradient: -2 × 70 = -140
The gradient is negative. That means increasing w decreases the loss. So we should move w up. The size, 140, says the slope is steep, so we are far from the answer.
Step 3: take a step
We move w in the opposite direction to the gradient, because we want the loss to go down. How far? We multiply the gradient by a small number called the learning rate. Pick 0.01.
new w = old w - learning rate × gradient
= 5 - 0.01 × (-140)
= 5 + 1.4
= 6.4
That is one step of gradient descent. Now repeat from step 1 with w = 6.4.
Watching it converge
Here is what happens if you keep going. Each row is one pass over the three examples, which is called an epoch.
| Epoch | w | Loss |
|---|---|---|
| 0 | 5.00 | 1050 |
| 1 | 6.40 | 863 |
| 2 | 7.67 | 709 |
| 5 | 10.83 | 392 |
| 10 | 14.62 | 135 |
| 20 | 18.53 | 10 |
| 40 | 19.92 | 0.03 |
| 60 | 19.996 | 0.0001 |
The network never gets told "the answer is 20". It just keeps stepping downhill on the loss, and the steps get smaller as the slope flattens near the bottom. After 60 epochs it has found 19.996, which for practical purposes is the rule pay = 20 times hours. It learned the rule from examples. That is machine learning.
Why the learning rate matters
Try the same thing with learning rate 0.1. The first step takes w from 5 to 19, which looks great. The second step overshoots to 20.9, the third to 19.2, and it settles down after a wobble. Push it to 0.25 and each step overshoots by more than the last. The loss grows instead of shrinking and w flies off to infinity. Too small a learning rate and training takes forever. Too big and it explodes. Picking it, and adjusting it as training goes on, is one of the main knobs engineers turn.
Scaling up: more weights, more layers
Our network had one weight and no layers to speak of. A real network differs in three ways, none of which change the procedure:
- Many inputs and many weights. A house price model might take size, age, bedrooms and postcode. Each input gets its own weight, plus a bias, which is a weight for a constant input of 1 so the line does not have to pass through zero. The output is the weighted sum. This one unit is called a neuron.
- Non-linearity. If you only ever multiply and add, stacking layers gains nothing, since a sum of sums is still a sum. So after each neuron the result goes through a simple bend. The most common one, ReLU, is just "if negative, make it zero". That tiny kink is what lets a stack of layers represent curves, boundaries and eventually cat versus dog.
- Layers. The outputs of one layer of neurons are the inputs of the next. A model with hundreds of layers and billions of weights is still doing weighted sums with bends in between.
Training is identical. Compute the loss. Compute the gradient of the loss with respect to every one of the billions of weights. Nudge every weight a little against its gradient. Repeat.
Backpropagation: computing a billion gradients at once
The one piece of machinery we skipped is how you get the gradient for a weight buried ninety layers deep. The answer is the chain rule from calculus, applied backwards from the loss through every layer. That algorithm is called backpropagation. It works out how much the loss would change if each weight moved, by passing blame backwards: the output was too high, so the last layer's weights get blame in proportion to their inputs, then the layer before gets blame in proportion to how much it fed the last layer, and so on down to the first. Modern frameworks like PyTorch do this automatically. You write the forward computation and the framework works out every gradient.
Batches and stochastic gradient descent
We used all three examples every step. With ten million examples that is too slow, so real training grabs a random batch of, say, 256 examples, computes the loss and gradient on just those, and steps. The gradient is noisier but you take thousands of steps in the time one full pass would take. This is stochastic gradient descent, and nearly every model you have heard of was trained with a variant of it called Adam, which adapts the learning rate per weight.
Overfitting: learning the examples instead of the rule
Our tiny model could only learn straight lines through zero, so it had no choice but to find the real rule. A model with billions of weights has another option: memorise the training examples exactly, including their noise, and be useless on anything new. That is overfitting. The fix is to hold back some data the model never trains on, the validation set, and watch its loss. When training loss keeps falling but validation loss starts rising, the model has stopped learning the rule and started memorising. You stop there.
What you now know
A model is a pile of weights. Training means: make a prediction, measure the loss, compute the gradient of the loss for each weight, step each weight a little downhill, repeat over batches for many epochs, and stop when a held out set says you are overfitting. Every headline model, from the one behind face unlock to the one behind ChatGPT, was made this way.
What differs between models is the shape of the network and what the input and output are. For a language model the input is text, and text is not numbers. Part three covers how text is turned into numbers a network can multiply: tokens and embeddings.