Deep Learning
Neural Networks
Layers, weights and activations. The unit of construction for everything downstream.
Grasp
A neural network is a stack of layers. Each layer multiplies its input by a matrix of learned weights, adds a bias, and passes the result through a non-linear function. That is genuinely all a layer is.
The power comes from composition. One layer can only draw a linear boundary. Two layers with a non-linearity between them can approximate essentially any continuous function, and depth beyond that buys efficiency: deep networks represent complex functions with far fewer parameters than shallow ones would need.
Three things are worth internalising early. The non-linearity is not optional - without it, stacked layers collapse algebraically into a single linear layer. The weights start random and are entirely learned; nothing about the structure encodes knowledge of your problem. And the width and depth you choose are hyperparameters, meaning you will find them by measurement rather than derivation.