Large Language Models
Large Language Models
Transformers trained on enough text to become general-purpose reasoning surfaces.
Grasp
A large language model is a transformer trained to predict the next token over an enormous corpus. That objective sounds trivially narrow. It is not, because predicting the next token well enough, at sufficient scale, requires the model to internalise syntax, facts, style, reasoning patterns and a rough model of the world.
Three properties matter for anyone building on top of one. It is stateless: every call sees only what you put in the context window, and nothing persists unless you persist it. It is probabilistic: the same prompt can produce different outputs, and this is a feature of the sampling process, not a bug to be fixed. And it is confidently wrong: fluency is uncorrelated with accuracy, which is why retrieval, tool use and evaluation exist as entire disciplines downstream of this node.
Almost everything in the rest of this region of the Atlas is a technique for working around one of those three properties.