LLM Evaluation

Large Language Modelspractice

First met at Engineer8h

In simple words

Measuring quality when there is no single correct output. The hardest unsolved problem in shipping.

The fuller explanation

Classical evaluation compares a prediction to a label. Generative evaluation usually cannot, because there are many acceptable outputs and no list of them exists.

Four approaches, in rough order of how much you should trust them. Deterministic checks - does it parse, does it match the schema, does the code run - are cheap, reliable and criminally underused. Reference-based metrics compare to a gold answer and are useful only where one genuinely exists. Model-graded evaluation uses a strong model as a judge against an explicit rubric; it scales well and carries real biases, notably toward length and toward its own style. Human review remains the ground truth and the thing everything else is calibrated against.

The practical failure is not choosing wrong among these. It is having no evaluation set at all, shipping on vibes, and discovering a regression from a user. Build a fixed set of a few hundred real examples before you build anything else.

Where does this sit on your route?

The free assessment places you on the same map and names which terms stand between you and the role you want.

Take the free assessment

See it in context

The Atlas shows this term with everything that leads into it and everything that follows, as one picture.

Open the map