A convolutional neural network (CNN) is a neural network built for data laid out on a grid, above all images. Instead of connecting every pixel to every neuron, it slides small filters, often 3×3 pixels, across the image. This sliding operation is a convolution. Each filter learns to respond to one pattern, such as a vertical edge or a patch of fur, and produces a feature map showing where that pattern appears.
Why convolution suits images
Two design choices do most of the work.
- Local connections: each output looks at a small neighbourhood of the input. Nearby pixels are related; pixels in opposite corners mostly aren’t.
- Shared weights: the same filter is applied at every position, so the network learns one edge detector rather than a separate one for each location. A pattern learned in one corner is recognized anywhere.
The savings are large. A single fully connected neuron looking at a 224×224 colour image needs 150,528 weights. A 3×3 filter over the same three colour channels needs 27.
How depth builds meaning
Stacking layers widens how much of the image each unit can see, its receptive field. Early layers respond to edges and colours, middle layers combine them into textures and parts, and late layers respond to whole objects. Pooling layers shrink the feature maps along the way, which makes the network tolerant of small shifts and cheaper to run.
Yann LeCun and colleagues developed the approach for reading handwriting, summarized in their 1998 paper on document recognition. In 2012, AlexNet scaled it up on GPUs and beat the previous best ImageNet results by a wide margin, which turned computer vision toward deep learning. ResNet later added skip connections that made networks with over a hundred layers trainable.
Limits
A CNN’s built-in assumptions, that nearby pixels matter most and that a pattern means the same thing anywhere, help most when data is limited. They also constrain it: relating distant parts of an image takes many stacked layers. Vision transformers drop those assumptions and, given enough training data, match or beat CNNs. CNNs remain common where compute is tight, such as on phones and cameras.