Gradient Descent

Machine Learningtechnique

First met at Foundation6h

In simple words

The optimisation procedure underneath essentially every model in this Atlas.

The fuller explanation

Training a model means searching for parameters that make a loss small, and gradient descent is the search. Compute the gradient of the loss with respect to every parameter, take a step in the opposite direction, repeat. The gradient says which way is uphill, so the negative gradient is the direction that reduces the error fastest from where you currently stand.

Two practical details do most of the work. The first is that computing the gradient over the entire dataset for every step is wasteful, so in practice it is estimated from a small batch at a time. The estimate is noisy, and the noise turns out to help, because it keeps the search from settling into the first flat spot it finds. The second is the step size, the learning rate. Too small and training crawls; too large and the parameters overshoot and the loss climbs or oscillates instead of settling. Nearly every optimiser you will meet is a scheme for choosing that step size more cleverly.

The thing to hold on to is what this procedure does not know. It minimises the loss you wrote down, faithfully, whatever that loss actually rewards. If a model learns something you did not intend, the optimiser is rarely at fault. The objective is.

Where does this sit on your route?

The free assessment places you on the same map and names which terms stand between you and the role you want.

Take the free assessment

See it in context

The Atlas shows this term with everything that leads into it and everything that follows, as one picture.

Open the map