←Home KnowML
Neural Nets & VisionChapter 05

CNNs & Vision Foundations

A convolution is a small, learned pattern-detector slid across every position of an image, reusing the same weights everywhere — that one design choice is most of what took computer vision from hand-engineered features to human-level accuracy inside a single decade.

26 min read Assumes: backprop & residual connections (04), dot products (01)
Start reading
TL;DR

A convolution is a small filter, a handful of learned weights, slid across every position of an image and computing a dot product at each spot. Reusing the same weights everywhere (weight sharing) buys two things a fully-connected layer cannot get: far fewer parameters, and translation invariance. A filter that learns to detect an edge in one corner recognizes that same edge anywhere else, for free.

Stack conv+ReLU layers, shrink space while growing channels, and you get a feature hierarchy: edges, then textures, then parts, then objects. ResNet's residual connections (04) fixed a genuine optimization failure in very deep plain CNNs and unlocked networks over 100 layers deep. After that, the field split into task-specific heads bolted onto the same backbone: detection (two-stage vs one-stage) and segmentation (semantic vs instance vs panoptic).

One sentence for an interview: a convolutional layer trades the full flexibility of a fully-connected layer for two strong, mostly-correct assumptions about images, locality and translation invariance, and that trade is what makes learning from images tractable at all.

01Intuition

Think of a convolution as a small stencil that gets stamped down on every patch of the image, one patch at a time.

A 3×3 edge-detector filter is nothing more than 9 numbers. Slide it over the top-left corner of an image and take the dot product between those 9 numbers and the 9 pixels underneath. You get one output value: high if that patch looks like an edge, low if it doesn't.

Now slide the exact same 9 numbers one pixel to the right, and repeat, over every position in the image. That is a convolution: one small, learned pattern-detector, reused at every location.

Reusing the filter at every location matters for two reasons:

  • Parameter count. That filter has 9 weights plus a bias, full stop. It does not matter whether the image is 32×32 or 4096×4096. The same 9 numbers are used everywhere, instead of every pixel position getting its own private set of weights.
  • Translation invariance. Because the same filter runs everywhere, an edge detector that works in the top-left corner works in the bottom-right corner too, automatically. The network never has to learn "edge at (5,5)" and "edge at (200,200)" as two unrelated facts.

A fully-connected layer on raw pixels has neither property. Every output unit has its own independent weight for every input pixel. So it has to see an object in every possible position during training to have any hope of recognizing it there.

A convolutional layer has a few more knobs. Worth having sharp intuitions for before the formula:

  • Stride is how far the stencil moves between stamps. Stride 1 checks every position. Stride 2 skips every other one and halves the output resolution.
  • Padding is a border of zeros added around the image before sliding, most often just to keep the output the same size as the input.
  • Dilation spreads the filter's 9 sample points further apart, skipping pixels between them. One filter can then "see" a wider area without adding a single extra weight. A cheap way to grow receptive field fast.
  • Channels are the depth dimension. A real conv layer doesn't run one filter. It runs dozens or hundreds of independent filters in parallel over the same input, each producing its own output map: edges, corners, color blobs, whatever gradient descent decides is useful. Stacked together, those are the "channels" the next layer sees.
Why 9 numbers are enough for an image of any size

The stencil analogy carries the whole section. A stencil has a fixed shape. You do not cut a new one for every place you want to stamp it.

So the filter's size follows the pattern it looks for, independent of how big the picture is. Nine numbers describe "a vertical edge", and that description does not get longer when the image does.

Everything else on this page follows from that one property. Parameter counts stay flat as resolution grows, a detector learned in one corner transfers to every other corner, and depth becomes affordable because each extra layer costs filters and not pixels.

A convolutional layer slides a small filter over an image and records how strongly each patch matches it. Here the filter looks for a jump from dark to bright, and the image has exactly one such edge.

Try it Slide an edge-detecting filter over a tiny image
import numpy as np

img = np.array([[0, 0, 9, 9, 9]] * 5, dtype=float)   # dark | bright
k = np.array([[-1, 0, 1]] * 3, dtype=float)          # dark-to-bright

out = np.zeros((3, 3))
for i in range(3):
    for j in range(3):
        out[i, j] = (img[i:i+3, j:j+3] * k).sum()

print(out.astype(int))
The filter answers 27 where the edge is and 0 where the image is flat. Every row is the same because every row of the image is the same. The first two windows contain the jump from 0 to 9, and the last window sees only bright pixels. The same nine filter weights are reused at every position, which is what keeps a convolution cheap.

02Timeline

Before

Vision pipelines fed hand-engineered features such as SIFT and HOG into a classical classifier like an SVM, which plateaued below what tasks needed and broke under changes in lighting, rotation or camera angle.

→
Innovation

AlexNet (Krizhevsky, Sutskever, Hinton, 2012) trained a deep convolutional network end to end on GPUs with ReLU and ImageNet and beat the hand-engineered pipelines by a large margin, because the features were learned from pixels.

→
After

VGG, Inception, ResNet and EfficientNet refined the design, task-specific heads added detection, segmentation and pose estimation, and CNNs stayed the default until Vision Transformers (06) began taking over at the largest scales after 2020.

03Architecture: how the data flows

A classification CNN is a straightforward pipeline once you've internalized the convolution itself: alternate conv+ReLU layers with downsampling, shrinking the spatial size while growing the number of channels, until the spatial map is small enough to flatten and hand to a plain classifier.

A simple CNN classification stack — spatial size shrinks, channel depth grows
Input image H × W × 3 Conv + ReLU 112×112×64 ↓ pool /2 Conv + ReLU 56×56×128 ↓ pool /2 Conv + ReLU 28×28×256 global pool Flatten 256-d FC layer → softmax Classifier head class scores
The dashed amber squares over the input mark how much of the original image a single output unit can "see" at increasing depth — its receptive field — even though every individual filter only ever looks at a small local window at a time. Spatial dimensions shown (112×112, 56×56, 28×28) are illustrative, roughly how a ResNet-style backbone shrinks space while channel depth grows.

Pooling vs. strided convolution — two ways to downsample

Both operations do the same job, shrinking the spatial size of the feature map, and they get there differently.

  • Pooling has zero learned parameters. Max pooling takes the largest value in each small window, average pooling takes the mean. It is a fixed, hard-coded rule, which makes it cheap. It also buys a small amount of translation invariance almost for free, since a small shift in the input often doesn't change which value was the max.
  • Strided convolution downsamples by moving its learned filter more than one pixel at a time (stride 2 instead of stride 1). The network gets to learn how to summarize each region instead of being handed a fixed rule.

Most modern backbones lean toward strided convolutions, or a mix, for exactly that reason. Letting gradient descent decide how to downsample beat the fixed rule often enough that plain max-pooling-everywhere fell out of fashion in the highest-performing designs. It is still cheap, simple, and common.

The historical arc: LeNet → AlexNet → VGG → Inception → ResNet

Each step in this lineage answered one specific question the previous architecture left open, and interviewers tend to ask "why did X matter" more often than "when was X," so the reasons are the part worth having cold.

Five architectures, each answering a different question:

  • LeNet (late 1980s/1990s, LeCun) established the template on small images like handwritten digits: conv layers, pooling, a small classifier head. The compute and data of the era could not scale it further.
  • AlexNet (2012) proved the template scales. Enough labeled data (ImageNet), enough compute (GPUs), and ReLU instead of saturating activations were enough to blow past hand-engineered features. ReLU mattered because it avoided the vanishing-gradient problem sigmoid/tanh networks suffered at depth.
  • VGG (2014) made a simplicity argument. Stop hand-tuning filter sizes per layer; stack small 3×3 convolutions repeatedly and go deeper. Three stacked 3×3 convolutions see as much of the input as one 7×7, with more non-linearity and roughly 45% fewer parameters: 27·C² against 49·C² at equal channel count C.
  • Inception/GoogLeNet (2014) argued the opposite axis mattered too. Instead of picking one filter size per layer, run several (1×1, 3×3, 5×5) in parallel within the same layer. Let the network combine multi-scale information itself rather than betting on one receptive field size being right everywhere.
  • ResNet (He et al., 2015) fixed something structural. Plain networks stacked past roughly 20 layers got harder to train, which is a different problem from overfitting. The exact failure and fix are covered in §05; it is the same residual-connection idea introduced in 04.

Beyond ResNet: the efficiency-focused lineage

Once depth stopped being the bottleneck, the next generation optimized for parameter and compute efficiency rather than raw accuracy alone.

  • DenseNet connects each layer to every previous layer's output, not just the one before it. That encourages feature reuse and eases gradient flow by a different mechanism than ResNet's: concatenation instead of summation.
  • EfficientNet scaled depth, width, and input resolution together in a fixed ratio (compound scaling) rather than growing one dimension. That gives a better accuracy-per-FLOP curve than scaling any single dimension alone.
  • MobileNet targets on-device inference, built from depthwise separable convolutions. These split a standard convolution into a per-channel spatial filter followed by a 1×1 filter that mixes channels, doing almost the same job for a fraction of the parameters and compute. The cost ratio against a standard k×k convolution with N output channels is 1/N + 1/k². For 3×3 kernels at typical channel counts, that is roughly 8–9× cheaper.
  • ConvNeXt (2022) is a more recent data point in the CNN vs ViT story. It took a standard ResNet, modernized the training recipe, and borrowed a handful of transformer design choices. A "pure" CNN can still match ViT-era accuracy when trained with modern tricks, which suggests some of the ViT-era gains came from training recipes rather than attention alone.

Feature pyramids solve a different problem. Early layers have fine spatial detail; late layers have strong semantic content but coarse resolution. Combining feature maps from multiple depths of the same backbone is the standard way detection and segmentation heads get both properties at once, instead of picking one depth and losing the other.

Detection and segmentation: two families of task heads

Both are built on the same convolutional backbone described above. They differ in what question they answer.

Object detection asks "what objects are here, and where": a class label plus a bounding box per object. Two families answer it differently.

  • Two-stage detectors (R-CNN → Fast R-CNN → Faster R-CNN) first propose candidate regions that might contain an object, then run a second network to classify and refine each region individually, reasoning region by region.
  • One-stage detectors (YOLO, SSD) skip the proposal stage entirely. They predict every box and class directly, everywhere on the image, in a single forward pass.

Most detectors in both families historically relied on anchor boxes: a fixed set of predefined box shapes tiled across the image, with the network predicting offsets from the nearest anchor rather than raw coordinates.

Anchor-free detectors predict box boundaries straight from each spatial location instead. That trades away a category of shape hyperparameters, at the cost of a different and often simpler matching problem during training. The speed/accuracy tradeoff this split creates is worked out in §06.

Segmentation asks a finer-grained question than detection: a label for every pixel where detection gives a box. It comes in three variants, and the distinction matters:

  • Semantic segmentation answers "what class is this pixel". Every pixel gets one of a fixed set of labels. Two adjacent cats are just "cat, cat", indistinguishable as separate objects.
  • Instance segmentation answers "which individual object does this pixel belong to", separating those two cats into two distinct instances.
  • Panoptic segmentation combines both. Every pixel gets a class label, and every pixel belonging to a countable "thing" (an object, and not a background region like sky or road) also gets an instance id, so one unified output replaces two separate tasks.

Two architectures cover most of this. Mask R-CNN extends a two-stage detector into instance segmentation: everything Faster R-CNN does, plus a small branch predicting a pixel mask per detected region.

U-Net (Ronneberger et al., 2015) is the classic semantic segmentation architecture. An encoder that downsamples like a normal CNN, a decoder that upsamples back to original resolution, and skip connections carrying fine spatial detail from encoder to decoder at matching resolutions. Why those skips are required is worked out in §06.

Named applications worth recognizing

Several other vision tasks reuse the same convolutional-backbone-plus-task-head pattern. Worth being able to name correctly, even without a deep dive:

  • Pose estimation. Predicts pixel locations of a fixed set of keypoints (joints, for a human body) rather than a box or a mask.
  • OCR and document understanding. Combines detection (where is the text) with sequence recognition (what does it say). Structured documents add layout analysis: what role does this block play, a title, a table cell, a caption?
  • Optical flow. Estimates for every pixel how it moved between two consecutive video frames. The output is a dense motion field, which is more than a single global camera motion.
  • Depth estimation. Predicts per-pixel distance from the camera. From a single image it is monocular, inherently ambiguous, and relies on learned priors. From a stereo pair you can triangulate directly.
  • Visual tracking. Follows a specific object's identity across video frames. Builds on detection, then adds the constraint that the same physical object should keep the same identity over time.
  • Face recognition and metric learning. Trains a network to produce embeddings where one identity's images land close together and different identities land far apart, in place of a fixed closed-set classifier. The same idea reappears constantly in retrieval and recommendation systems (16).

04The equations

$$W_{out} = \left\lfloor \frac{W - K + 2P}{S} \right\rfloor + 1$$
  • W the input's spatial width (or height, computed identically and usually equal for square feature maps) — e.g. 224 for a 224×224 input image, or whatever spatial size the previous layer output.
  • K kernel size — the width/height of the learned filter, most commonly 3, sometimes 5 or 7 for an early "stem" layer.
  • P padding — rows/columns of zeros added to each side of the input before sliding the kernel, most often chosen so the output stays the same spatial size as the input ("same" padding).
  • S stride — how many pixels the kernel moves between applications. S=1 checks every position; S=2 skips every other position and roughly halves the output size — this is exactly what a strided convolution uses to downsample instead of pooling.
  • ⌊⌋ + 1 the floor accounts for the kernel needing to fit fully inside the padded input (a fractional position isn't a valid convolution step); the +1 counts the starting position itself.

Concrete check: ResNet's first layer takes a 224×224 input through a 7×7 kernel, stride 2, padding 3. (224 − 7 + 2·3)/2 + 1 = (223)/2 + 1 = 111 + 1 = 112 — a 112×112 output, exactly matching the published architecture.

Parameter count

A 3×3 convolution with C_in input channels and C_out output channels has exactly $9 \cdot C_{in} \cdot C_{out} + C_{out}$ parameters. Each (input channel, output channel) pair contributes 9 weights, plus one bias per output channel, and the count has nothing to do with the image's width or height.

Run the same filter on a 32×32 image or a 4096×4096 image and the parameter count doesn't move. Only the compute does, because the same small filter gets evaluated at more positions.

Compare that to a fully-connected layer mapping a 224×224×3 image (150,528 input values) to a modest 1,000-unit hidden layer: that's over 150 million weights in a single layer.

That gap is, more than anything else, the entire reason convolutions work at all for images — there aren't enough labeled images in existence to learn 150 million independent weights per layer without catastrophic overfitting, but $9 \cdot C_{in} \cdot C_{out}$ weights, reused at every spatial position, is a completely different and tractable learning problem.

05Why this architecture

The obvious alternative is a plain fully-connected (MLP) network applied to raw pixels. It loses on the two properties that matter most for images: parameter efficiency and translation invariance.

A convolution bakes in two assumptions. Nearby pixels are related (locality), and a useful pattern stays useful wherever it appears (translation invariance). Both hold for almost every natural image. So a CNN never spends parameter budget or training data rediscovering them.

An MLP has to see a cat in every possible pixel position during training to have a real chance of recognizing a cat anywhere, because every output unit's weights are independent of every other position's. A CNN's shared filter recognizes it in one position and, for free, recognizes it everywhere.

Worth stating precisely, since it is a common interview probe and the mechanism is the same one covered in 04.

He et al. (2015) observed something that looks backwards: a plain, non-residual 56-layer CNN had higher training error than an 18-layer one. That is worse error on the data it was trained on, so overfitting cannot explain it.

A deeper plain network can always represent the shallower one exactly, by setting the extra layers to compute the identity and do nothing. So the cause had to be an optimization failure, since capacity was never the issue: plain SGD, propagated through dozens of stacked non-linear layers, could not reliably find that identity mapping.

A residual block computes $f(x) + x$ instead of $f(x)$. Identity becomes the default whenever the learned branch $f$ contributes little. The network starts at "pass the input straight through" and only learns a useful deviation, instead of learning identity the hard way through stacked layers.

That single change took ImageNet-winning networks from roughly 19 layers (VGG) to 152 in the same year, still improving as depth increased.

ResNet-18 full architecture diagram, layer by layer
Reference figureResNet-18's full layer-by-layer architecture — conv stem, four stages of residual blocks with increasing channels and decreasing spatial size, global pool, then a classifier head. Source: Zhang, Lipton, Li & Smola, Dive into Deep Learning, via Wikimedia Commons (CC BY-SA 4.0).

06Tradeoffs and what breaks

Two-stage (R-CNN family)One-stage (YOLO / SSD)
Core ideaPropose candidate regions first, then classify and refine each one individuallyPredict every box and class directly, everywhere, in one forward pass
SpeedSlower — two networks run in sequenceFaster — a single pass, real-time capable
Accuracy (historically)Higher, especially on small or overlapping objects, from the dedicated per-region refinement stepHistorically lower, though better losses (e.g. focal loss, addressing the huge foreground/background imbalance one-stage detectors face) and feature pyramids have closed most of that gap
Typical use caseOffline or high-accuracy pipelines where latency matters less — e.g. medical image reviewReal-time systems — video, robotics, mobile/edge deployment

Failure modes

Pooling (or a strided convolution) makes a network cheaper and more translation-invariant. It destroys exact spatial detail on the way. After five rounds of halving, a 224×224 input is down to 7×7, and the network no longer knows which original pixel any given feature came from.

This suits classification, since "is there a cat in this image" does not need pixel-exact detail. Segmentation does need it: a label for every original pixel, and that detail is gone from the bottleneck.

U-Net's skip connections are the fix: at each resolution level, the encoder's feature map from before it was downsampled is copied straight across to the matching level in the decoder and concatenated with the upsampled features.

Without the skip, the decoder has to reconstruct a precise object boundary from a heavily compressed, low-resolution bottleneck. That information was thrown away several pooling steps ago and is not recoverable. The skip connection hands the decoder back exactly the spatial detail the encoder saw before pooling, at exactly the resolution needed to draw a sharp edge.

CNNs' locality-and-translation-invariance assumption is a genuine strength in low-data regimes. The network does not have to spend training examples learning that images have local structure, because that is wired into the architecture before training starts. So a CNN reaches good accuracy from comparatively little labeled data.

But a baked-in assumption is also a ceiling. At very large scale, an architecture with fewer built-in assumptions and more freedom to learn spatial structure directly from data has empirically scaled further. That architecture is the Vision Transformer (08 covers the attention mechanism, 06 covers ViT specifically). The freedom becomes an advantage once dataset size and compute are large enough.

Common misconception

"Two-stage detectors are always meaningfully more accurate, one-stage is just the fast/cheap option." That gap was real and large around 2015–2017 (Faster R-CNN clearly beat early YOLO/SSD on accuracy). Better losses and multi-scale features closed most of it over the following years, and for most production systems today a well-tuned one-stage detector is the default and is no longer a compromise made only when speed is forced on you.

2026 status

CNN backbones remain the default wherever inference budget or labeled data is limited: mobile and edge deployment, and transfer-learning pipelines fine-tuned on a modest labeled dataset. The reasons are the low-data inductive bias described above and a decade of mature tooling built around it.

At the frontier end, ViT-based and hybrid conv-transformer backbones (06) are now the default for large-scale pretrained vision foundation models. "CNNs are obsolete" is an overclaim. It is a scale- and data-dependent tradeoff and not a strict replacement, and exactly where the line falls is what 06 covers next.

07Build this

Nobody tells a convolution to look for edges. It finds them anyway, and the fastest way to believe that is to plot the filters yourself.

Project Watch a conv layer grow its own edge detectors ~3 hours · PyTorch

Train a small CNN on CIFAR-10, then render its first convolutional layer's filters as images. The classification task is ordinary on purpose. The reveal is the filter grid. It starts as random noise and ends up looking like a page from a signal-processing textbook nobody handed the network.

  1. Write the convolution yourself first. Loop over output positions, dot the kernel with the patch underneath, and check the answer matches nn.Conv2d on a random input. Use the library version afterwards so training finishes.
  2. Build a small CNN: a 7×7 first conv with 32 filters, then three 3×3 blocks with stride-2 downsampling, global pool, one linear head. The first kernel is large on purpose, so the filters are legible when plotted. Size each layer with the formula from §04.
  3. Before any training, render those 32 filters as tiny RGB images on a grid and save that picture.
  4. Train for a few epochs. Re-render the same grid after every epoch. Then push one test image through and plot the 32 activation maps the first layer produces.
  5. Now break it: replace the whole conv stack with an MLP of matching parameter count and train it identically.
  6. Shift every test image two pixels to the right. Re-evaluate both models on the shifted set.
You'll know it worked when the filter grid stops looking like static. Oriented edges, opposing-colour blobs and a few centre-surround spots appear, and each one lights up a different region of the activation maps. Gradient descent chose those. You did not.
What the breakage teaches. The parameter-matched MLP scores worse, and the two-pixel shift costs it far more than it costs the CNN. That second gap is weight sharing made visible. The CNN learned one edge detector and got every position for free. The MLP had to learn each position separately, from whichever training examples happened to land there. §05 argues this in prose; here it is a number you measured.

Where this runs in production

Say the constraint is a phone-class NPU running object detection at 30 fps — roughly a 33ms budget per frame. Every choice covered above becomes a concrete lever and stops being abstract:

  • Backbone: a MobileNet-style stack built from depthwise separable convolutions rather than standard convolutions. Illustrative at C_in = C_out = 256: a standard 3×3 convolution costs 9·256·256 ≈ 590,000 parameters (and proportional FLOPs); splitting it into a depthwise 3×3 (one filter per channel, 9·256 ≈ 2,300 parameters) followed by a pointwise 1×1 that mixes channels (256·256 ≈ 65,500 parameters) does nearly the same job for roughly 8–9× fewer parameters and FLOPs — the difference between fitting the latency budget and not.
  • Detector head: one-stage (YOLO/SSD-style) rather than two-stage — a single forward pass fits inside the budget; a two-stage detector's second region-by-region pass does not.
  • Downsampling: strided convolutions folded into the backbone rather than separate pooling layers, so downsampling doesn't cost an extra pass over the data.
  • Anchors: a small anchor set tuned to the object scales that occur in the target use case, rather than the broad, expensive anchor sets tuned for general-purpose benchmarks — fewer anchors means fewer candidate boxes to score every frame.

None of this is an exotic trick list — every item is a direct, mechanical consequence of an idea already covered above: fewer parameters per layer, one forward pass instead of two, and letting the network fold downsampling into a convolution it's already computing rather than paying for a separate operation.

08Vision is not one task

"Computer vision" names a dozen different problems that happen to share an input type. The backbone is often identical. What changes is the head, the shape of the label, and what counts as a failure.

Section §03 introduced detection and segmentation as families of task head, and this is the practical version. For each problem: what goes in, what comes out, what makes it hard, and what you would reach for now.

  • Image classification. In: one image. Out: one label. It is the easiest member of the family, and the one where architecture choice matters least. The decision that matters is which pretrained weights you start from.
  • Object detection. Out: a variable-length list of boxes with labels. The variable length is the hard part: a network with a fixed-size output has to be coerced into emitting a set, and that coercion is the entire design problem.
  • Semantic segmentation. Out: a class label for every pixel. Hard because you need pixel-exact boundaries from a backbone that threw spatial detail away while downsampling. Skip connections (§06) exist for exactly this.
  • Instance segmentation. Out: one mask per object, plus a class. Two adjacent cats become two masks. Masks are allowed to overlap, and nothing requires you to label the background at all.
  • Panoptic segmentation. Out: exactly one class and one instance id for every pixel, background included. It is not a union of the two above, for reasons that deserve a subsection of their own.
  • Keypoint and pose estimation. Out: pixel coordinates for a fixed skeleton of joints. Joints get occluded constantly, and you still have to work out which person a detected elbow belongs to.
  • Depth estimation. Out: distance from the camera, per pixel. A stereo pair lets you triangulate geometrically. A single image does not: absolute scale is unrecoverable, so a monocular model leans on learned priors about how big things usually are.
  • Tracking. Out: detections plus a consistent identity across frames. Detection quality is rarely the bottleneck here; identity is, and the signature failure is the ID switch, where two objects cross paths and swap labels.

Why detection needed anchors, and then stopped

A convolutional head emits a fixed grid of predictions. Objects arrive as a variable-length set. Something has to decide which output slot is responsible for which real object. That assignment problem is the thing every generation of detector has been attacking.

Anchors were the first answer: tile a set of predefined box shapes across every location, then assign each ground-truth object to whichever anchor overlaps it most. The network predicts an offset from that anchor rather than raw coordinates, and regressing a small correction is far easier than regressing an absolute box from nothing.

Anchors bought tractability and charged for it:

  • Hyperparameters. How many shapes, which aspect ratios, which scales. All tuned per dataset, none of them learned.
  • Class imbalance. Most anchors match nothing, so easy background examples swamp the gradient. Focal loss was designed specifically to stop that drowning effect.
  • Duplicate predictions. Several anchors fire on the same object, so you need non-maximum suppression afterwards to clean up.

Anchor-free detectors predict box edges straight from each spatial location, deleting the shape hyperparameters outright. DETR went further still, reframing detection as direct set prediction: it matches predictions to ground truth with the Hungarian algorithm during training, and so removes NMS by construction rather than by post-processing.

The thread running through all of it

Anchors, anchor-free heads and set prediction look like three unrelated architectures. They are three answers to one question: which prediction should be held responsible for which object?

Anchors answer it with a hand-designed geometric prior. Anchor-free heads answer it with spatial position. DETR answers it by solving an explicit matching problem during training.

If an interviewer asks why detection is harder than classification, this is the answer. Classification has one output and one label. Detection has to invent a correspondence.

Why panoptic is not semantic plus instance

Kirillov et al. (2019) defined panoptic segmentation with a hard constraint: every pixel gets exactly one class label and one instance id, with no overlaps and no gaps.

You cannot reach that by running two models and merging their output:

  • Instance masks may overlap. Two predicted masks claiming the same pixel is legal in instance segmentation and illegal in panoptic, so something has to arbitrate.
  • Instance segmentation ignores the background. Sky, road and grass have no countable instances, so an instance model has nothing to say about them.
  • Semantic segmentation cannot split objects. It covers every pixel, but two touching cars come back as one undifferentiated "car" region.

The vocabulary is the part worth keeping. Things are countable: cars, people, bottles. Stuff is not: road, sky, vegetation. Panoptic gives things an instance id and leaves stuff without one.

Mask2Former resolved the split architecturally. Predict a set of binary masks, each carrying one class label. Semantic, instance and panoptic then become the same model output, read three different ways.

09Where you meet this in the wild

Almost none of these deployments look like ImageNet. The label format is usually the thing that changes, and the interesting engineering happens around the model.

Medical imaging
Segmenting an organ or a tumour in a CT volume
U-Net's home turf, and still competitive. nnU-Net surpassed most existing approaches, including highly specialised ones, across 23 public datasets from international segmentation competitions — without a novel architecture. Its contribution was automatic configuration: preprocessing, patch size, and training schedule derived from the dataset itself. The lesson generalises well beyond medicine.
Manufacturing
Catching a defect you have never seen before
The obvious framing fails here: you cannot train a defect classifier when defects are rare and the next one is a category nobody has photographed yet. So the task is usually inverted: train only on normal parts and flag deviation, as in the MVTec AD benchmark. Anomaly detection, not classification.
Satellite and aerial
Finding small objects in enormous images
A single scene can run to tens of thousands of pixels a side, so images get tiled and stitched, and objects of interest are often a handful of pixels across. Sensors also give you more than three channels. A convolution does not care: change the input channel count and the rest of the network is untouched.
Documents
Turning a scanned invoice into structured fields
Three problems wearing one coat: detection finds where the text is, sequence recognition reads it, and layout analysis works out what each block means, whether that is a total, a line item or a header. The vision half is usually the easy half.
Sports and retail
Following the same person across a whole video
Detection per frame is close to solved, but identity across frames is not. Players occlude each other, shoppers leave and re-enter the frame, and a single ID switch corrupts every downstream statistic. Most of the engineering budget goes to the association step, not the detector.
Reach for a CNN when
  • Local texture or shape decides the answer. Defects, tissue boundaries, edges. Locality is a correct assumption here, and you get it for free.
  • You have modest labelled data and a pretrained backbone. The inductive bias substitutes for examples you do not have.
  • Compute or latency is tight. Edge devices, real-time video, embedded hardware. The tooling is mature and the models are small.
  • Your input is not three-channel RGB: multispectral, depth or stacked medical slices all drop straight in.
Look elsewhere when
  • The answer depends on relating distant parts of the image. Attention models that assumption instead of fighting it. See 06.
  • You have large-scale pretraining available: past a certain data budget, fewer built-in assumptions scale further, which is the tradeoff §06 lays out.
  • The problem is geometric with known mathematics. Calibration, homographies, and stereo rectification are solved analytically, and learning them gives a worse answer.
  • Labels are the bottleneck. Promptable models like SAM can cut annotation cost far more than a better backbone would.

10Interview questions

BeginnerWhat does a convolution actually compute, and why does weight sharing matter?

A convolution slides a small learned filter over every position of the input, computing a dot product between the filter and the patch underneath at each position. Reusing the same filter (weight sharing) at every position gives two things: far fewer parameters than a fully-connected layer, since the filter's size doesn't depend on the image's size, and translation invariance, since a pattern learned in one position is automatically recognized in every other position.

BeginnerPooling vs. a strided convolution — what's the actual difference?

Both downsample the spatial size of a feature map. Pooling (max or average) applies a fixed, hard-coded rule with zero learned parameters — cheap, and gives a bit of translation invariance for free. A strided convolution downsamples using its learned weights, moving the filter more than one pixel per step, which lets the network learn how to summarize each region instead of following a fixed rule. Most high-performing modern backbones lean toward strided convolutions or a mix, rather than pure pooling everywhere.

IntermediateDerive the output spatial size of a convolutional layer.

W_out = ⌊(W − K + 2P) / S⌋ + 1, where W is input width, K is kernel size, P is padding on each side, and S is stride. The kernel needs K input positions to produce one output; padding effectively extends the usable width to W + 2P; dividing by stride and flooring counts how many valid starting positions fit; the +1 counts the first position. Concrete check: ResNet's stem uses a 224 input, K=7, P=3, S=2: (224 − 7 + 6)/2 + 1 = 111 + 1 = 112.

IntermediateWhy did ResNet's residual connections specifically fix the degradation problem — why wasn't it just an overfitting issue?

He et al. showed that a plain (non-residual) deeper network could have higher training error than a shallower one — not a generalization gap, worse error on the training set itself. Since a deeper plain network can always represent the shallower function exactly (set extra layers to identity), this had to be an optimization failure: plain SGD struggled to find the identity mapping through many stacked non-linear layers. A residual block computes $f(x)+x$ instead of $f(x)$, making identity the default whenever the learned branch contributes little, so the network only has to learn a useful deviation from "do nothing" rather than learning identity the hard way.

IntermediateWhen would you actually choose a two-stage detector over a one-stage one today?

When the latency budget genuinely allows for it and the task rewards the extra accuracy a dedicated per-region refinement step gives — small or heavily overlapping objects, or offline/high-stakes review pipelines like medical imaging — rather than real-time constraints. For most real-time or edge deployments, a well-tuned one-stage detector has closed most of the historical accuracy gap and is the default choice, not a fallback.

DeepWhy do skip connections matter specifically for segmentation?

An encoder built from repeated pooling/stride-2 convolutions destroys exact spatial detail as it downsamples — after several halvings, the network no longer knows precisely which original pixel a feature came from. That's harmless for classification but fatal for segmentation, which needs a label for every original pixel. U-Net's skip connections copy the encoder's feature map at each resolution directly across to the matching resolution in the decoder, concatenating it with the upsampled features. Without that connection, the decoder would have to reconstruct sharp object boundaries purely from a heavily compressed low-resolution bottleneck — detail that was already thrown away and isn't recoverable from the bottleneck alone.

DeepYou're choosing a backbone — CNN or ViT — for a new vision product with a modest labeled dataset, tens of thousands of images. Which do you pick, and why?

The deciding factor isn't "CNN vs. ViT" in the abstract, it's how much data and pretraining you actually have access to. A CNN's built-in locality and translation-invariance assumptions give it a strong head start on modest data, so a CNN — or better, a CNN pretrained on a large dataset and fine-tuned on yours (transfer learning) — is the safer default when data is limited. A ViT can still win, but only by leaning on large-scale pretraining (06) to learn spatial structure it doesn't get for free architecturally; training a ViT from scratch on tens of thousands of images, with no pretraining, typically underperforms a comparable CNN.

11Go deeper

●Now write it yourself

Reading the derivation and being able to produce it are different skills. These are Deep-ML problems that exercise what this page covers — each one is checked against real test cases, not multiple choice.

Matched to this page from Deep-ML's catalogue of 1,380 problems. More at deep-ml.com, and Where to practise covers the other platforms and what each one trains.

My Notes — 05 CNNs & Vision Foundations

Free notes

Highlights on this page