LLM Evaluation
In simple words
Measuring quality when there is no single correct output. The hardest unsolved problem in shipping.
The fuller explanation
Classical evaluation compares a prediction to a label. Generative evaluation usually cannot, because there are many acceptable outputs and no list of them exists.
Four approaches, in rough order of how much you should trust them. Deterministic checks - does it parse, does it match the schema, does the code run - are cheap, reliable and criminally underused. Reference-based metrics compare to a gold answer and are useful only where one genuinely exists. Model-graded evaluation uses a strong model as a judge against an explicit rubric; it scales well and carries real biases, notably toward length and toward its own style. Human review remains the ground truth and the thing everything else is calibrated against.
The practical failure is not choosing wrong among these. It is having no evaluation set at all, shipping on vibes, and discovering a regression from a user. Build a fixed set of a few hundred real examples before you build anything else.
Learn these first
Real prerequisites, taken from the map rather than guessed.
Sources
Where this came from, so you can go past us.
Where does this sit on your route?
The free assessment places you on the same map and names which terms stand between you and the role you want.
Take the free assessmentSee it in context
The Atlas shows this term with everything that leads into it and everything that follows, as one picture.
Open the map