← all topics

Generative AI

Attention

Letting every position look at every other position. The hinge the modern field turns on.

Levels foundation / engineer / advanced
Depth 5
Time 5h
Kind concept
On AI Engineer, RAG Specialist, Agent Builder, AI Architect

Grasp

Recurrent networks process a sequence one step at a time, carrying a fixed-size memory forward. That memory is a bottleneck, and it degrades over distance: by token 500, the influence of token 1 has largely washed out.

Attention removes the bottleneck. Every position computes a weighted sum over every other position, where the weights are learned relevance scores. Nothing has to be squeezed through a running state, and any token can directly influence any other regardless of distance.

The cost is quadratic: attending over n positions requires n squared comparisons. Almost every architectural refinement of the last several years - sparse attention, sliding windows, linear attention, state-space models - is an attempt to keep the benefit while escaping that cost.