Backpropagation
In simple words
The chain rule applied at scale. How a network learns which weights were at fault.
The fuller explanation
Gradient descent needs a gradient: for every weight in the network, the answer to how much the loss would change if that one weight moved slightly. A small network has thousands of weights and a large one has billions, so computing each answer separately is out of the question. Backpropagation is the trick that gets all of them in a single pass.
It works because a network is a chain of functions, and the chain rule says the derivative of a chain is the product of the derivatives of its links. So the loss is computed at the output, and then blame flows backwards: each layer receives the gradient of the loss with respect to its own output, hands back the gradient with respect to its input, and along the way reads off the gradients for its own weights. The cost of the whole thing is roughly the cost of the forward pass, which is why deep networks are trainable at all.
Two consequences matter in practice. A backward pass needs the activations the forward pass produced, which is why memory, not arithmetic, usually decides how large a batch you can fit. And because gradients are multiplied layer by layer, long chains can shrink them towards nothing or blow them up. Much of modern architecture design exists to keep that product well behaved.
Learn these first
Real prerequisites, taken from the map rather than guessed.
Sources
Where this came from, so you can go past us.
Where does this sit on your route?
The free assessment places you on the same map and names which terms stand between you and the role you want.
Take the free assessmentSee it in context
The Atlas shows this term with everything that leads into it and everything that follows, as one picture.
Open the map