Large Language Models
LLM Evaluation
Measuring quality when there is no single correct output. The hardest unsolved problem in shipping.
Grasp
Classical evaluation compares a prediction to a label. Generative evaluation usually cannot, because there are many acceptable outputs and no list of them exists.
Four approaches, in rough order of how much you should trust them. Deterministic checks - does it parse, does it match the schema, does the code run - are cheap, reliable and criminally underused. Reference-based metrics compare to a gold answer and are useful only where one genuinely exists. Model-graded evaluation uses a strong model as a judge against an explicit rubric; it scales well and carries real biases, notably toward length and toward its own style. Human review remains the ground truth and the thing everything else is calibrated against.
The practical failure is not choosing wrong among these. It is having no evaluation set at all, shipping on vibes, and discovering a regression from a user. Build a fixed set of a few hundred real examples before you build anything else.