Machine Learning
Gradient Descent
The optimisation procedure underneath essentially every model in this Atlas.
Grasp
Training a model means searching for parameters that make a loss small, and gradient descent is the search. Compute the gradient of the loss with respect to every parameter, take a step in the opposite direction, repeat. The gradient says which way is uphill, so the negative gradient is the direction that reduces the error fastest from where you currently stand.
Two practical details do most of the work. The first is that computing the gradient over the entire dataset for every step is wasteful, so in practice it is estimated from a small batch at a time. The estimate is noisy, and the noise turns out to help, because it keeps the search from settling into the first flat spot it finds. The second is the step size, the learning rate. Too small and training crawls; too large and the parameters overshoot and the loss climbs or oscillates instead of settling. Nearly every optimiser you will meet is a scheme for choosing that step size more cleverly.
The thing to hold on to is what this procedure does not know. It minimises the loss you wrote down, faithfully, whatever that loss actually rewards. If a model learns something you did not intend, the optimiser is rarely at fault. The objective is.