9 LeNet / CNN
Context
A fully connected network on an image is a disaster: millions of weights and total disregard for structure (neighbouring pixels are related, an object can move). Yann LeCun (1989→1998) builds a network with that structure wired into the architecture.
The idea and the mechanism
Three principles. Local receptive fields: a neuron looks at a small patch. Weight sharing: one set of weights (a filter) slides across the whole image — this is the convolution operation:
One filter = a feature detector (an edge, a corner) that does not care where the feature sits. Pooling aggregates neighbouring responses → robustness to shifts and a drop in dimensionality. A stack of conv → pool → conv → pool → fc, trained end-to-end by backprop; early layers catch edges, later ones parts of objects.
linear algebra Weight sharing: how many parameters convolution saves
Take a 32×32 input and a hidden layer of 32×32 neurons. The fully connected version: each of the 1024 outputs is connected to each of the 1024 inputs →
The convolutional version with a 5×5 filter: the same filter is applied at every position →
A saving of tens of thousands of times, and the same thing gives you translational equivariance: move the object → the response moves the same way, because the filter is identical everywhere. Convolution encodes the prior knowledge "locality + invariance to position" directly into the structure — the strongest inductive bias there is for images.
NumPy Convolving one filter over an image
import numpy as np
def conv2d(img, K): # img: (H,W), filter K: (k,k)
k = K.shape[0]
H, W = img.shape
out = np.zeros((H - k + 1, W - k + 1))
for i in range(out.shape[0]):
for j in range(out.shape[1]):
patch = img[i:i+k, j:j+k]
out[i, j] = (patch * K).sum() # one filter at every position
return out # feature map
Why it matters
LeNet-5 read handwritten digits in production (cheques, postcodes) and became the template for every CNN. The first convincing demonstration that a learned hierarchy of features from pixels beats hand-crafted ones. The idea will reach full force in 2012 (AlexNet, #16), once GPUs and ImageNet arrive.
Connections
A CNN is the same network, trained by backprop, but with constraints wired in (locality, sharing). A convolution is simply a layer with shared weights; the gradient flows through it by the same chain rule.
AlexNet is LeNet grown up on GPUs and ImageNet: the same convolutions + pooling, but deeper, with ReLU and dropout. A 1998 idea waited for the data and the hardware — and then took off.
ViT drops the hard-wired inductive bias of convolutions and replaces it with data scale + self-attention. A CNN "knows" about locality from birth; a ViT "learns" it, but only on large data. The argument "built-in biases vs learn it from data" runs right through ML.
Questions worth asking
Convolution gives translational equivariance, not invariance — what is the difference and does it matter?
Equivariance: move the input → the output moves the same way (a property of convolution itself). Invariance: move the input → the output does not change (what you need for classification, "it's a cat wherever it is"). Invariance is picked up by pooling and by global aggregation at the end. Confusing the two is a common mistake; a CNN is equivariant by construction and invariant only after pooling.
If convolution is so good, why wasn't it invented and adopted widely before 2012?
It was there (LeNet was reading cheques in the 1990s), but it ran into data and compute: small datasets + weak CPUs gave deep CNNs no room, while SVMs were competitive on those tasks and simpler. ImageNet (#15) and GPUs removed both limits — and CNNs immediately pulled ahead.
Pooling throws away information about exact position — isn't that harmful?
It is a deliberate trade-off: we sacrifice precise localization for robustness to shifts and for compression. For classifying "what is in the picture" that is useful. But for tasks where position matters (segmentation, detection) aggressive pooling gets in the way — there people use strided convolutions, dilated convolutions, or drop pooling altogether to keep spatial precision.
What to read in the original
The 1998 paper is long (46 pp.) — read selectively, the sections on convolution, pooling and weight sharing; the graph-transformer part at the end can be skipped.