Machine Learning
Model Evaluation
The skill that separates people who ship models from people who publish notebooks.
Grasp
Accuracy is almost always the wrong number. On a dataset where 99% of examples are negative, a model that predicts "negative" every time scores 99% and is worthless.
What you need instead is a metric that reflects the cost structure of your actual problem. Precision asks: of the things I flagged, how many were right. Recall asks: of the things I should have flagged, how many did I catch. These trade off against each other, and choosing where to sit on that curve is a product decision, not a modelling one.
The deeper skill is not metric selection, it is split hygiene. Any leakage between your training and evaluation data - a duplicated row, a feature computed over the whole dataset, a time-ordered problem split randomly - produces a number that is confidently, silently wrong. Assume leakage until you have proven otherwise.