Attention

Generative AIconcept

First met at Foundation5h

In simple words

Letting every position look at every other position. The hinge the modern field turns on.

The fuller explanation

Recurrent networks process a sequence one step at a time, carrying a fixed-size memory forward. That memory is a bottleneck, and it degrades over distance: by token 500, the influence of token 1 has largely washed out.

Attention removes the bottleneck. Every position computes a weighted sum over every other position, where the weights are learned relevance scores. Nothing has to be squeezed through a running state, and any token can directly influence any other regardless of distance.

The cost is quadratic: attending over n positions requires n squared comparisons. Almost every architectural refinement of the last several years - sparse attention, sliding windows, linear attention, state-space models - is an attempt to keep the benefit while escaping that cost.

Learn these first

Real prerequisites, taken from the map rather than guessed.

Sources

Where this came from, so you can go past us.

Where does this sit on your route?

The free assessment places you on the same map and names which terms stand between you and the role you want.

Take the free assessment

See it in context

The Atlas shows this term with everything that leads into it and everything that follows, as one picture.

Open the map