← all topics

Machine Learning

Model Evaluation

The skill that separates people who ship models from people who publish notebooks.

Levels foundation / engineer / advanced
Depth 3
Time 8h
Kind practice
On RAG Specialist, Agent Builder, AI Architect

Grasp

Accuracy is almost always the wrong number. On a dataset where 99% of examples are negative, a model that predicts "negative" every time scores 99% and is worthless.

What you need instead is a metric that reflects the cost structure of your actual problem. Precision asks: of the things I flagged, how many were right. Recall asks: of the things I should have flagged, how many did I catch. These trade off against each other, and choosing where to sit on that curve is a product decision, not a modelling one.

The deeper skill is not metric selection, it is split hygiene. Any leakage between your training and evaluation data - a duplicated row, a feature computed over the whole dataset, a time-ordered problem split randomly - produces a number that is confidently, silently wrong. Assume leakage until you have proven otherwise.