← all topics

Deep Learning

Neural Networks

Layers, weights and activations. The unit of construction for everything downstream.

Levels foundation / engineer / advanced
Depth 4
Time 8h
Kind concept
On AI Engineer, RAG Specialist, Agent Builder, AI Architect

Grasp

A neural network is a stack of layers. Each layer multiplies its input by a matrix of learned weights, adds a bias, and passes the result through a non-linear function. That is genuinely all a layer is.

The power comes from composition. One layer can only draw a linear boundary. Two layers with a non-linearity between them can approximate essentially any continuous function, and depth beyond that buys efficiency: deep networks represent complex functions with far fewer parameters than shallow ones would need.

Three things are worth internalising early. The non-linearity is not optional - without it, stacked layers collapse algebraically into a single linear layer. The weights start random and are entirely learned; nothing about the structure encodes knowledge of your problem. And the width and depth you choose are hyperparameters, meaning you will find them by measurement rather than derivation.