← all topics

Large Language Models

Large Language Models

Transformers trained on enough text to become general-purpose reasoning surfaces.

Levels foundation / engineer / advanced / architect
Depth 7
Time 4h
Kind system
On AI Engineer, RAG Specialist, Agent Builder, AI Architect

Grasp

A large language model is a transformer trained to predict the next token over an enormous corpus. That objective sounds trivially narrow. It is not, because predicting the next token well enough, at sufficient scale, requires the model to internalise syntax, facts, style, reasoning patterns and a rough model of the world.

Three properties matter for anyone building on top of one. It is stateless: every call sees only what you put in the context window, and nothing persists unless you persist it. It is probabilistic: the same prompt can produce different outputs, and this is a feature of the sampling process, not a bug to be fixed. And it is confidently wrong: fluency is uncorrelated with accuracy, which is why retrieval, tool use and evaluation exist as entire disciplines downstream of this node.

Almost everything in the rest of this region of the Atlas is a technique for working around one of those three properties.