
Computer Vision
Image classification, object detection, GANs, vision-language models. From CNNs to state-of-the-art architectures.

After this course, you'll be able to:
Who this is for
- →Engineers who need to ship a vision feature — a defect check, a document reader, a camera analytic — and want to understand the choices rather than copy a notebook.
- →Machine learning practitioners from other domains who need the vision-specific vocabulary: anchors, IoU, mAP, receptive fields, feature pyramids.
- →Robotics and embedded developers who will have to run a model on a device with a fixed power budget and no room for a full-precision network.
- →Researchers and analysts in imaging-heavy fields such as medical, agricultural or industrial inspection who want to evaluate vendor claims competently.
- →Not for you if you want to generate images. That is a different discipline with different tooling; this course covers generative vision as one topic, not as its subject.
- →Not for you yet if you have never trained any model. The material assumes you know what a loss curve is and why validation data exists.
What you need first
- ·Python at a comfortable working level, including reading a training script written by somebody else and modifying it without breaking it.
- ·Some prior exposure to neural network training: forward pass, loss, gradient step, epochs, overfitting. You do not need to have implemented backpropagation yourself.
- ·Basic linear algebra intuition. Convolution, pooling and attention are all easier to reason about if matrix shapes do not intimidate you.
- ·Access to a GPU for the training portions, whether local or rented. Inference and most of the analysis work run fine on CPU; training a detector on CPU is impractical.
- ·A tolerance for image data logistics: formats, color spaces, annotation files, and datasets measured in gigabytes rather than rows.
The task you choose determines everything that follows
The most expensive mistake in a vision project is made in the first week, before any model is trained, when somebody decides what the model is supposed to output.
Image classification assigns one label to a whole image. Annotation is fast, models are small, and the metric is straightforward. Object detection returns boxes with classes and confidences, which is a harder learning problem and a considerably more expensive labeling job. Semantic segmentation assigns a class to every pixel but does not separate touching instances. Instance segmentation separates them. Panoptic segmentation does both stuff and things at once. Keypoint estimation returns landmark coordinates. Depth estimation returns a value per pixel with no class at all. Tracking adds identity across frames, which introduces an entirely separate class of errors.
These are not interchangeable, and the annotation cost between them is not remotely comparable. Drawing a box around an object is quick and mostly mechanical. Producing an accurate pixel mask of something with an irregular or soft boundary is slow, tiring, and full of genuine judgment calls about where the object stops. Teams routinely commit to segmentation because it looks more capable, discover the labeling budget halfway through, and end up with a small, inconsistently annotated dataset that would have been better spent on a large, clean detection set.
The right question is what decision the output feeds. If the downstream system only needs to know whether a defect is present, classification is enough and everything else is waste. If it needs to count objects, detection is the minimum. If it needs an area measurement or a precise boundary for a robot to grasp, segmentation is genuinely required. If it needs to know that the object in frame two hundred is the same object as in frame one, you are building a tracker, and that is a systems problem as much as a modeling one.
Downgrading the task where you can is one of the few reliable ways to make a vision project cheaper and more accurate simultaneously. A classifier trained on a large, well-labeled set generally beats a segmentation model trained on the small, rushed set that the same budget bought. This course works through the taxonomy explicitly, because choosing correctly here saves more than any architectural improvement will.
Convolutions, transformers, and why the argument is not settled
Convolutional networks encode a set of assumptions about images directly into the architecture: that nearby pixels are related, that a useful feature is useful anywhere in the frame, and that meaning is built hierarchically from edges to textures to parts to objects. Those assumptions are constraints, and constraints are what let a model learn from a limited number of examples. Residual connections, introduced by Microsoft Research with ResNet, made very deep convolutional stacks trainable and remain a default backbone for a great deal of production work.
Vision transformers, from Google Research, discard most of that structure. An image is cut into patches, each patch becomes a token, and self-attention lets every patch attend to every other from the first layer. This buys global context immediately, where a convolutional network has to build up receptive field depth-wise. The cost is data hunger: with weaker built-in assumptions, the model has to learn spatial structure from examples, which is why transformer backbones typically depend on large-scale pretraining to be competitive. Attention over patches also scales quadratically with token count, which matters as soon as you want high-resolution input.
In practice the field converged on pragmatism rather than a winner. Hybrid designs use convolutional stems for efficient early feature extraction and attention in later stages. Hierarchical transformer variants reintroduce locality and multi-scale structure explicitly. Modern convolutional architectures have absorbed training recipes and design choices originally developed for transformers, and remain very strong.
For a practitioner, the decision usually reduces to three inputs. How much labeled data do you have, and are you fine-tuning from a strong pretrained checkpoint or training closer to scratch? What is your inference hardware, since convolutional operations map cleanly onto essentially every accelerator and mobile runtime while attention support is more uneven? And what input resolution does your problem actually need, because small-object detection often demands resolution that makes attention costs uncomfortable.
The habit worth building is to treat backbone choice as a constrained optimization rather than a fashion decision. Fix your latency budget and hardware first, then choose the largest model that fits it, then spend the remaining effort on data. The gap between backbone families on a given task is usually smaller than the gap between a mediocre dataset and a good one.
Detection and segmentation: the trade-offs that actually bite
Detection architectures split along a few axes worth understanding, because each choice moves a different cost around.
Two-stage designs propose candidate regions and then classify and refine them. They tend to be accurate and comparatively slow. Single-stage designs predict boxes and classes directly across a dense grid, trading some accuracy for a large speed gain, which is why single-stage families dominate real-time applications. Anchor-based methods predict offsets from a preset grid of box shapes, which means anchor sizes and aspect ratios become hyperparameters that must match your objects; a detector configured for pedestrians will do poorly on long thin scratches until someone notices. Anchor-free methods predict box geometry directly from points or centers and remove that tuning burden. Set-prediction approaches such as DETR, from Facebook AI Research, now Meta AI, reframe detection as directly predicting a set of objects with bipartite matching, which eliminates the non-maximum suppression post-processing step entirely, at the cost of slower convergence during training.
Non-maximum suppression deserves specific attention because it causes a category of bug that is easy to misattribute. It removes overlapping duplicate boxes based on an overlap threshold. Set it too aggressively and genuinely adjacent objects get merged, which looks like the model failing to detect crowded instances when in fact the model found them and post-processing deleted them. Anyone debugging a detector should visualize raw predictions before suppression at least once.
Evaluation is the other place people go wrong. Mean average precision aggregates precision and recall across confidence thresholds and object classes, and it is the standard reported metric, popularized by the COCO benchmark from Microsoft. It is also a summary that conceals a great deal. It averages over object sizes, so a detector that is excellent on large objects and poor on small ones can post a respectable score while failing at your actual use case. It averages over classes, so a rare but critical class can be invisible in the number. And it says nothing about the specific operating point you will deploy at. Always break the metric down by class and by object size, and always choose and report the confidence threshold you will actually run.
Segmentation adds its own considerations. Boundary quality is where models differ most and where the standard region-overlap metrics are least sensitive, because the boundary is a small fraction of the pixels. If a precise edge matters for your application, measure it directly rather than trusting an aggregate overlap score.
Data problems that no architecture will fix
Vision datasets fail in characteristic ways, and recognizing them early is worth more than any architectural upgrade.
Leakage through near-duplicates. Video is the worst offender. Sampling frames and splitting them randomly puts nearly identical images in both train and test, and the resulting evaluation number is fiction. Hold out whole recording sessions, whole cameras, whole production batches, whole patients or whole sites, whichever grouping the deployed system will genuinely be encountering for the first time. The same applies to any dataset built by photographing one object from many angles.
Class imbalance and the long tail. Real inspection and safety applications are mostly negatives. A model that predicts the majority class will look accurate and be useless. Handle this deliberately with sampling strategy, loss weighting, or by restructuring the problem, and evaluate with metrics that do not reward the trivial solution.
Label noise and boundary ambiguity. Two annotators will draw different boxes around the same partially occluded object and different masks around anything with a soft edge. Without a written annotation guideline and a measured level of agreement, you cannot distinguish a model that is wrong from labels that disagree with each other. This matters most in exactly the domains where accuracy matters most.
Domain shift. A model trained on images from one camera, one lens, one lighting setup and one facility will degrade on another, sometimes severely, and often without any obvious visual difference to a human. Sensor noise characteristics, white balance, JPEG compression level, focal length and mounting angle are all part of the training distribution whether you intended them to be. If deployment spans multiple sites, the validation set must span them too.
Augmentation as a blunt instrument. Flips, crops, color jitter and geometric transforms genuinely help, but each encodes an assumption. Horizontal flipping is fine for most natural scenes and wrong for anything where handedness or text orientation carries meaning. Aggressive color augmentation is counterproductive when color is the signal, as in many medical and agricultural tasks. Augmentation should be chosen from knowledge of the domain, not copied from a default configuration.
Synthetic and generated data. Rendering or generating training images is increasingly practical and closes some gaps well, particularly for rare classes. It also introduces a domain gap of its own between synthetic and real appearance. It is a supplement, and the validation set must remain real.
From a trained model to something that runs on the device
A checkpoint that scores well in a notebook is the smaller half of the work. The rest is making it run where the camera is, at the rate the camera produces frames, without a person watching it.
Quantization. Converting weights and activations to lower precision, commonly eight-bit integers, reduces memory and often dramatically increases throughput on hardware with integer acceleration. Post-training quantization needs a representative calibration set; calibrate on data that does not match deployment conditions and accuracy drops in ways that are hard to trace back to the cause. Quantization-aware training recovers more accuracy at the cost of a training cycle. Either way, re-evaluate the quantized model on your real validation set. The number you report must come from the artifact you actually ship.
Batch size one. Throughput benchmarks are usually measured with large batches. Real cameras deliver frames one at a time, and single-sample latency is a different measurement dominated by different bottlenecks: memory transfers, preprocessing, and framework overhead rather than raw compute. Preprocessing in particular is often the surprise, since resizing and color conversion on a weak CPU can cost more than the network inference.
Temporal instability. Running a per-frame model on video produces flicker: boxes that appear and vanish between frames, labels that oscillate. To a user this reads as the system being broken even when per-frame accuracy is good. Temporal smoothing, tracking, or hysteresis on the decision threshold are the standard remedies, and they belong in the design rather than bolted on after a complaint.
Monitoring without labels. In production nobody tells you the correct answer. You can still monitor the distribution of confidences, the rate of detections per frame, the fraction of inputs that look unlike training data, and image-level statistics such as brightness and sharpness that catch a dirty lens or a camera someone bumped. Most real production failures in vision are input failures rather than model failures, and they are detectable if you look.
The regulated cases. Medical imaging and autonomous driving change the standard entirely. Sensitivity and specificity at a chosen operating point matter more than aggregate accuracy, because the two error types have different consequences. Validation must be per-site and per-device, since scanner and sensor differences are substantial. Face recognition carries legal restrictions that vary considerably by jurisdiction and has well-documented differential error rates across demographic groups, which the United States National Institute of Standards and Technology has evaluated in its published face recognition vendor testing program. Building in these areas is a compliance and evaluation exercise before it is a modeling one.
Are you ready, and what comes after
You are ready for this material if you can train a simple neural network end to end, read a training script and change the optimizer or the augmentation without breaking it, and interpret a validation curve that stops improving. You do not need previous vision experience; convolution is introduced from the ground up.
You are probably not ready if you have never trained any model at all. Start with a general machine learning foundation first, then come back. The vision-specific content assumes the general vocabulary and moves quickly past it.
A useful sign that the timing is right: you have images and a question about them. Vision is a domain where the material only becomes durable when applied to pictures you understand well enough to know when the model is wrong.
Three directions follow naturally. Vision-language models connect an image encoder to a language model so a system can be asked open-ended questions about a picture instead of returning a fixed label set, which is the right next step if your problem is document understanding or anything requiring reasoning about scene content. Generative vision, meaning diffusion models and image synthesis, shares the encoder and training vocabulary but has its own tooling and workflow. And the deployment path leads into serving and hardware: quantization in depth, accelerator selection, and running inference on machines you own.
Whichever you pick, the discipline established here transfers. Choose the narrowest task that answers the question being asked, keep whole cameras and whole sessions out of training rather than shuffling frames, report the metric at the confidence threshold you will actually run in production, and plan for the lens, the lighting and the scene to change under you without warning.
Common questions
Do I need a GPU to learn computer vision?
For training, effectively yes. Training a detector or segmentation model on CPU is slow enough to be impractical for iteration. You do not have to buy one: a rented cloud GPU for a handful of hours per project is usually far cheaper than hardware, and free hosted notebook tiers are adequate for smaller experiments. Inference, evaluation, data inspection and annotation work all run fine on an ordinary machine, and that is a substantial share of real vision work.
Are CNNs obsolete now that vision transformers exist?
No. Convolutional networks remain competitive on many tasks, train from less data because their built-in assumptions about locality do real work, and map cleanly onto mobile and embedded runtimes where transformer support is patchier. Vision transformers are strong when there is large-scale pretraining behind them and global context matters. Most current systems are somewhere between the two, using convolutional stems with attention layers, and the practical decision is driven by your hardware and dataset size rather than by which family is newer.
How many labeled images do I need to train a detector?
There is no single figure, and anyone quoting one is guessing about your problem. The real drivers are how many classes you have, how visually distinct they are, how much variation exists in lighting and viewpoint, and whether you are fine-tuning from a pretrained backbone or starting cold. Fine-tuning from a good checkpoint on a narrow, visually consistent task needs far less than training broadly. The practical approach is to label a first batch, train, look at which cases fail, and label more of those rather than more of everything.
What is the difference between object detection and segmentation?
Detection returns a rectangle around each object with a class and a confidence. Segmentation assigns a class to individual pixels, so it captures the actual shape. Semantic segmentation does not distinguish two touching objects of the same class; instance segmentation does. Choose based on what the downstream system needs. Counting and locating objects only needs detection. Measuring an area, following a precise boundary, or masking a region for editing needs segmentation, and you should expect the labeling effort to be substantially higher.
Why does my model perform worse after deployment than in testing?
The most common causes, in rough order: the test split shared near-duplicate images with training so the reported score was never real; the deployment cameras, lighting or lens differ from the training conditions; preprocessing in the production pipeline does not match what was used during training, particularly resizing and color channel order; or the model was quantized for the target device and never re-evaluated afterward. Check preprocessing parity and the split methodology before assuming the model itself is at fault.
Can computer vision models run offline on local hardware?
Yes, and it is the normal arrangement for camera systems, industrial inspection and robotics, where sending video to a remote service is impractical on bandwidth, latency or confidentiality grounds. Detection and classification models are small compared with large language models, and quantized versions run on modest accelerators and even some microcontrollers. The engineering work moves toward the runtime: choosing an inference framework the device supports, quantizing correctly, and keeping preprocessing identical to training.
Related reading
Best local vision models
Which multimodal and vision models can be run on your own machine, and what each is suited to.
Local AI vision tasks
Concrete image understanding jobs — captioning, extraction, comparison — run without a cloud API.
DeepSeek OCR setup guide
Document image processing, the most common real-world entry point into vision work.
Video analysis with local AI
Frame sampling and temporal handling, the part that per-frame models get wrong.
Frigate and camera analytics
A working example of continuous detection on live camera streams at the edge.
Used GPU buying guide
Hardware sizing and second-hand pricing if you intend to train vision models rather than only run them.
Full syllabus
Advanced Architectures
Object Detection
Segmentation
Medical Imaging
Autonomous Vehicles
Face Recognition
Video Understanding
3D Vision
Generative Vision
Vision-Language Models
Edge Deployment
Production Computer Vision
Unlock all 13 chapters
Plus 24 other courses — 548 more chapters included.
Every course, every future course, the Python Lab and eight downloadable kits, nothing to renew. Or subscribe: Pro $8.99/month