Generative AI
Transformers
Attention, feed-forward layers, residuals and normalisation. The block that ate the field.
Grasp
The transformer is not one idea, it is a stack of a repeated block: multi-head attention, a position-wise feed-forward network, residual connections around both, and layer normalisation. Positional information is injected explicitly, because attention itself is order-blind.
Its real advantage over what came before is not accuracy, it is parallelism. Recurrent models must process token t before token t+1, which pins training to sequential time. A transformer sees the whole sequence at once, which means training scales with the hardware you can buy. That property, more than any modelling insight, is why the architecture won.
Understand this node properly and large language models stop being magic. They are this block, repeated dozens to hundreds of times, trained on a very large corpus.