← all topics

Generative AI

Transformers

Attention, feed-forward layers, residuals and normalisation. The block that ate the field.

Levels foundation / engineer / advanced
Depth 6
Time 10h
Kind system
On AI Engineer, RAG Specialist, Agent Builder, AI Architect

Grasp

The transformer is not one idea, it is a stack of a repeated block: multi-head attention, a position-wise feed-forward network, residual connections around both, and layer normalisation. Positional information is injected explicitly, because attention itself is order-blind.

Its real advantage over what came before is not accuracy, it is parallelism. Recurrent models must process token t before token t+1, which pins training to sequential time. A transformer sees the whole sequence at once, which means training scales with the hardware you can buy. That property, more than any modelling insight, is why the architecture won.

Understand this node properly and large language models stop being magic. They are this block, repeated dozens to hundreds of times, trained on a very large corpus.