Foundations

Computer Vision

The field of AI that enables computers to extract information from images and video: what is shown, where it is, and what is happening.

IMAGEvehicle · 98%pedestrian · 91%Classificationwhat is in this image?Object detectionwhere? draw a box around itSegmentationwhich pixels belong to the object?OCRread the text in the imageToday's multimodal models can do most of these with one model, asked in plain language.

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/computer-vision

In plain terms

To a computer a photo is a grid of numbers. Computer vision turns that grid into something useful: this is a cat; there is a pedestrian at this position; this weld is cracked; this is the total on the receipt. It is what unlocks a phone with a face, reads number plates at a car park and keeps a car in its lane.

Why it matters

It is among the most mature and widely deployed kinds of AI, with clear returns in quality inspection, medical imaging, retail, logistics, agriculture and security. It is also among the most sensitive: recognising faces and tracking people raise questions of privacy and discrimination, and the EU AI Act restricts or prohibits several such uses. Results depend heavily on conditions; a model trained in good light fails in bad.

Example

A bottling plant installs cameras above its line. A vision model checks every bottle for fill level, cap position and label alignment, at 600 bottles a minute, and ejects the faulty ones. Before that, a person sampled one bottle in two hundred. In the first quarter, returns for mislabelled bottles fall by 80%.

Most often confused with

Computer Vision vs. Multimodal AI

Computer VisionSpecialised models for specific visual tasks
Multimodal AIGeneral models that reason over images and text

A classic vision model does one thing, such as finding defects on one product, and does it very fast and very consistently. A multimodal model can be asked anything about any image in plain language, with no training, but it is slower, costs more and is less precise for measurement. For a fixed, high-speed task use the first; for varied, open questions the second.

Under the hood

Core tasks: image classification, object detection (bounding boxes), segmentation (pixel-level masks), keypoint and pose estimation, optical character recognition, tracking across video frames, and depth and 3D reconstruction. Models: convolutional neural networks have dominated since 2012; vision transformers and foundation models for vision, such as Segment Anything, have joined them. Well-known tools include the YOLO family for real-time detection. In practice, performance depends on labelled data that matches real conditions (lighting, angle, camera), so collecting and annotating data is most of the work; models often run on edge devices close to the camera for speed. Measures: accuracy, precision and recall, mean average precision for detection, and intersection over union for segmentation. Risks: uneven accuracy across demographic groups in face analysis, adversarial examples, and privacy law.

Written by Mehmet Erkek · Last updated: