Landmark 3 stops to get here · leads to 13

Convolutional Neural Network

A neural network that scans images with small learned filters, reusing the same weights at every position to build up from edges to whole objects.

Your route here

3 stops · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Convolution ✓ understood

    A mathematical operation that applies filters/kernels to input data to extract features like edges, textures, and patterns.

  4. Convolutional Neural Network · you are here ✓ understood

Picture it

  1. Class scores e.g. softmax over labels
  2. Fully connected Combines features
  3. Flatten Feature maps to a vector
  4. Deeper conv + pooling Detect parts, then objects
  5. Pooling Downsamples feature maps
  6. Convolution + ReLU Filters detect edges and textures
  7. Input image Grid of pixels
Reading upward, notice how each stage trades spatial detail for more abstract features until only class scores remain.

A convolutional neural network (CNN) is a neural network built for data laid out on a grid, above all images. Instead of connecting every pixel to every neuron, it slides small filters, often 3×3 pixels, across the image. This sliding operation is a convolution. Each filter learns to respond to one pattern, such as a vertical edge or a patch of fur, and produces a feature map showing where that pattern appears.

Why convolution suits images

Two design choices do most of the work.

  • Local connections: each output looks at a small neighbourhood of the input. Nearby pixels are related; pixels in opposite corners mostly aren’t.
  • Shared weights: the same filter is applied at every position, so the network learns one edge detector rather than a separate one for each location. A pattern learned in one corner is recognized anywhere.

The savings are large. A single fully connected neuron looking at a 224×224 colour image needs 150,528 weights. A 3×3 filter over the same three colour channels needs 27.

How depth builds meaning

Stacking layers widens how much of the image each unit can see, its receptive field. Early layers respond to edges and colours, middle layers combine them into textures and parts, and late layers respond to whole objects. Pooling layers shrink the feature maps along the way, which makes the network tolerant of small shifts and cheaper to run.

Yann LeCun and colleagues developed the approach for reading handwriting, summarized in their 1998 paper on document recognition. In 2012, AlexNet scaled it up on GPUs and beat the previous best ImageNet results by a wide margin, which turned computer vision toward deep learning. ResNet later added skip connections that made networks with over a hundred layers trainable.

Limits

A CNN’s built-in assumptions, that nearby pixels matter most and that a pattern means the same thing anywhere, help most when data is limited. They also constrain it: relating distant parts of an image takes many stacked layers. Vision transformers drop those assumptions and, given enough training data, match or beat CNNs. CNNs remain common where compute is tight, such as on phones and cameras.

Where it sits

Explore nearby

In the research

All papers →

4 papers that build on Convolutional Neural Network .

Sources

  1. LeCun, Bottou, Bengio and Haffner, "Gradient-Based Learning Applied to Document Recognition" . Proceedings of the IEEE, 1998
  2. Krizhevsky, Sutskever and Hinton, "ImageNet Classification with Deep Convolutional Neural Networks" . NeurIPS 2012 (AlexNet)
  3. Ian Goodfellow, Yoshua Bengio and Aaron Courville, Deep Learning . MIT Press, 2016; chapter 9, Convolutional Networks