Computer Vision

Computer Visionconceptcv

First met at Engineer4h

In simple words

Getting structured meaning out of pixels: classification, detection, segmentation.

The fuller explanation

An image arrives as a grid of numbers, and vision is the work of turning that grid into something a program can act on. The field is mostly organised by how specific the answer has to be. Classification returns one label for the whole image. Detection returns boxes, so it can say there are three of them and where. Segmentation returns a label per pixel, which is what you need when the shape itself matters, as it does for a tumour boundary or a drivable lane. Each step up is a harder labelling problem before it is a harder modelling problem.

That labelling cost shapes the whole discipline. Drawing a box takes seconds; outlining every pixel of every object takes minutes, and someone has to agree on where the edges are. So in practice you rarely train from scratch. A model pretrained on a large general corpus already has useful early features, and you fine-tune it on the few thousand examples you could actually afford to annotate.

What surprises people moving from benchmarks to production is that the model is seldom the weak link. Cameras change, lenses smear, lighting shifts, someone remounts a unit two degrees lower, and accuracy falls for reasons no retraining fixes. Vision systems are mostly judged on whether they survive the conditions their data collection never saw.

Where does this sit on your route?

The free assessment places you on the same map and names which terms stand between you and the role you want.

Take the free assessment

See it in context

The Atlas shows this term with everything that leads into it and everything that follows, as one picture.

Open the map