Transformers

Generative AIsystem

First met at Foundation10h

In simple words

Attention, feed-forward layers, residuals and normalisation. The block that ate the field.

The fuller explanation

The transformer is not one idea, it is a stack of a repeated block: multi-head attention, a position-wise feed-forward network, residual connections around both, and layer normalisation. Positional information is injected explicitly, because attention itself is order-blind.

Its real advantage over what came before is not accuracy, it is parallelism. Recurrent models must process token t before token t+1, which pins training to sequential time. A transformer sees the whole sequence at once, which means training scales with the hardware you can buy. That property, more than any modelling insight, is why the architecture won.

Understand this node properly and large language models stop being magic. They are this block, repeated dozens to hundreds of times, trained on a very large corpus.

Learn these first

Real prerequisites, taken from the map rather than guessed.

Sources

Where this came from, so you can go past us.

Where does this sit on your route?

The free assessment places you on the same map and names which terms stand between you and the role you want.

Take the free assessment

See it in context

The Atlas shows this term with everything that leads into it and everything that follows, as one picture.

Open the map