← all topics

Computer Vision

Computer Vision

Getting structured meaning out of pixels: classification, detection, segmentation.

Levels engineer / advanced / architect
Depth 6
Time 4h
Kind concept

Grasp

An image arrives as a grid of numbers, and vision is the work of turning that grid into something a program can act on. The field is mostly organised by how specific the answer has to be. Classification returns one label for the whole image. Detection returns boxes, so it can say there are three of them and where. Segmentation returns a label per pixel, which is what you need when the shape itself matters, as it does for a tumour boundary or a drivable lane. Each step up is a harder labelling problem before it is a harder modelling problem.

That labelling cost shapes the whole discipline. Drawing a box takes seconds; outlining every pixel of every object takes minutes, and someone has to agree on where the edges are. So in practice you rarely train from scratch. A model pretrained on a large general corpus already has useful early features, and you fine-tune it on the few thousand examples you could actually afford to annotate.

What surprises people moving from benchmarks to production is that the model is seldom the weak link. Cameras change, lenses smear, lighting shifts, someone remounts a unit two degrees lower, and accuracy falls for reasons no retraining fixes. Vision systems are mostly judged on whether they survive the conditions their data collection never saw.